🤖 StepFun releases StepAudio 3 audio family

The flagship StepAudio 3 Realtime is a full-duplex model with Think-While-Speaking mode: reasoning runs in parallel with speech, so it asks clarifying questions and interrupts without pauses. Languages — Chinese and English.

🌍 StepFun covers the audio pipeline with a single lineup: ASR Max ($0.24 per hour), TTS ($0.36 per 10,000 characters), Gen with voice, effects, and music on a single timeline, and Realtime API (WebSocket, VAD, ASR, web_search). On the Full-Duplex Bench, Realtime scores 98.9 compared to 95.3 for GPT-realtime-2.

👤 Conversational demos are already working in Voice Studio (audio.stepfun.ai), and the stepaudio-3-gen-preview plan is free. For voice bots, the Realtime API saves code on VAD and ASR, while Music generates a 48 kHz stereo track up to 5:30 long from a description.

Source 1: https://stepaudiollm.github.io/step-audio-3-realtime/

Source 2: https://stepaudiollm.github.io/step-audio-3-music/