StepAudio 3 Realtime: think-while-speaking audio-language model with asynchronous tool calls

StepFun

Research official 2 src. ~1 min

StepAudio 3 Realtime is an audio-language foundation model built around a continuous listen-converse-think-act loop: deep acoustic perception, seamless duplex turn-taking with natural interruptions, 'think-while-speaking' parallel reasoning, and asynchronous tool calls inside voice conversation. It reports dialogue and reasoning quality comparable to dedicated reasoning models while speaking in real time.

Why it matters

90 upvotes on HuggingFace Daily Papers (2026-09-16); state of the art in closing the gap between full reasoning depth and real-time voice latency, the key constraint for voice agents.

Importance: 2/5

Notable technical report from StepFun

Sources