Gemini 3.8 Live and 3.8 Live Extended Thinking
Google announced two new live dialogue models, Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as its most advanced yet, with major upgrades in intelligence and parallel reasoning that make voice collaboration and complex task execution more intuitive. The announcement comes from Tom Ouyang (Principal Engineer) and Malini Jaganathan (Member of Technical Staff) on behalf of the Gemini Audio Team.
The two models are positioned for different workloads. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, offering increased intelligence and multi-step reasoning. For developers and enterprises, Google frames these as building blocks for reliable, production-ready voice agents, and they also improve speaking with Gemini across the Gemini app, Google Workspace, and Search.
Benchmark results favor the Extended Thinking model: it captured the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with 82.6, leads agentic task completion at 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking, and scored 97.7% on Big Bench Audio while maintaining a highly competitive price point versus other frontier models. Gemini 3.8 Live ranked second in the Speech Agent Arena and remains highly cost-effective. On ServiceNow's EVA-Bench, the models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality (run on the Live API on Gemini Enterprise Agent Platform).
Capability details: Gemini 3.8 Live processes visual inputs in near real-time to add conversational context, automatically detects and transitions between 97 supported languages mid-conversation, and executes tools and API calls in the background while the conversation continues — so it can acknowledge a request and keep chatting while tasks finish. For deeper reasoning, Extended Thinking reasons and speaks simultaneously, using early verbal cues such as "Let me check that…" to acknowledge prompts without interrupting conversational flow.