#speech
- Thinking Machines Lab Unveils TML-Interaction-Small: 276B MoE Real-Time Multimodal Model Thinking Machines Lab models-llm
- EVA-Bench: End-to-End Framework for Evaluating Voice Agents ServiceNow AI research
- Gemini 3.5 Live Translate: Real-Time Speech-to-Speech in 70+ Languages Google DeepMind audio
- Google launches Gemini 3.8 Live and 3.8 Live Extended Thinking speech models Google DeepMind models-llm
- MiniCPM-o 4.5: Real-Time Full-Duplex Omni-Modal AI on Edge Devices OpenBMB / Tsinghua University research
- NVIDIA Releases Nemotron-Labs-Audex-30B-A3B: Unified Audio-Text MoE Model NVIDIA audio
- OpenAI brings ChatGPT Voice to the desktop app with Codex control OpenAI tools
- Cartesia launches Sonic-3.6 TTS, claims top Elo over Eleven v3 Cartesia audio
- SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks ByteDance research
- Tencent open-sources AuK, a 1.5B speech foundation model for generation and editing Tencent Hunyuan audio
- Audio Interaction Model: Unified Streaming Framework Combining Offline and Real-Time Audio Instruction Following research
- xAI Launches Grok Voice Agent Builder in Beta xAI tools
- Gander: end-to-end omni interaction agent with Cerebellum-Brain architecture Tencent Hunyuan research
- Yandex opens medical AI assistant to all Russian doctors Yandex industry
- Meta Superintelligence Labs releases Muse Voice Transcribe, a real-time speech model Meta audio
- StepAudio 3 Realtime: think-while-speaking audio-language model with asynchronous tool calls StepFun research
- xAI Grok Voice Mode Coming to Apple CarPlay, App Build Reveals xAI tools