📰 Key Takeaways

Gemini just launched two new real-time voice models today, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, both built around near-instant reasoning to make voice assistants feel more like natural conversation. 3.8 Live is aimed at scale and cost efficiency, combining conversational intelligence, smooth responsiveness, and visual understanding, while Extended Thinking is built for high-complexity tasks with stronger multi-step reasoning.

On performance, 3.8 Live Extended Thinking took the #1 overall spot on Artificial Analysis’s Speech to Speech Quality Index (82.6 points), leading in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s banking scenario, τ-Voice-banking. On reasoning, it scored a striking 97.7% on Big Bench Audio, all while staying competitively priced against other frontier models. 3.8 Live ranked #2 in user preference on Speech Agent Arena, also holding onto strong cost efficiency, making it a solid fit for developers and enterprises scaling deployment. Both models pushed the Pareto frontier forward on ServiceNow’s EVA-Bench voice assistant benchmark, balancing accuracy and conversational quality in complex workflow scenarios.

On the feature side, 3.8 Live can process visual input near-instantly to enrich conversational context, automatically detect and switch between 97 supported languages mid-conversation, and run tool calls and API requests in the background — so the model can respond to a user right away while finishing the task behind the scenes. The Extended Thinking version can reason and speak at the same time, naturally picking up a request with verbal cues like “let me check on that,” and narrating progress in real time so users stay in the loop on multi-step background tasks. Both models are now available through the Gemini Live API, ready for developer platforms like Agora, LangChain, and LiveKit.


💬 JudyAI Lab’s Take

Gemini splitting “real-time voice” and “deep reasoning” into two separate product lines with 3.8 Live and 3.8 Live Extended Thinking — that division alone is worth paying attention to.

The trend we’re seeing is that voice assistant competition isn’t just about who answers faster anymore, it’s about who can think while talking. Extended Thinking uses conversational cues like “let me check on that” to pick up a request while running tool calls in the background, turning what used to be dead-air waiting into an experience where the user feels kept in the loop. That “make the latency visible” design thinking is a good reminder for any team building agent products: beyond the raw performance numbers, how you communicate during the wait matters just as much for the overall experience.

If your product also has long-task waiting scenarios, it’s worth considering whether you can add a progress cue instead of leaving users staring at a blank screen.


📅 Source Info


🔗 Further Reading