Back to News
ProductAI Understanding briefing

Google introduces Gemini 3.8 Live and Extended Thinking models

Google DeepMind has launched Gemini 3.8 Live and 3.8 Live Extended Thinking, new live dialogue models designed for near real-time reasoning, voice agents, and complex task execution across Google Workspace and Search.

5 min readRead the primary source
Source-provided image accompanying Google introduces Gemini 3.8 Live and Extended Thinking models
Primary-source documentSource recorded
Publisher
deepmind.google
Source link
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/
Source type
Primary document — an official announcement, paper, filing, or first-party page we read directly.
ContextUnderstand this in 60 seconds

Start here

Key terms

API (Application Programming Interface)
A structured way for one software system to send requests to and receive responses from another system.
Watermarking
Embedding a detectable signal in AI-generated text or media so it can later be identified as machine-produced.
Benchmark
A standardized test or dataset used to measure and compare model performance.
Test yourselfAI Agents Quiz

What happened

Google DeepMind announced the release of two new AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live dialogue, featuring near real-time reasoning capabilities that allow for simultaneous speaking and reasoning. The announcement highlights improvements in intelligence, parallel reasoning, and the ability to execute tools and API calls in the background while maintaining conversational flow. The models are available via the Gemini Live API for developers and are integrated into the Gemini app, Google Workspace, and Search.

Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as the company's most advanced live dialogue models. The primary innovation is the ability to perform near real-time reasoning, allowing the models to process complex tasks while maintaining an uninterrupted conversational flow. This is achieved through parallel reasoning, where the model can acknowledge requests with verbal cues like 'Let me check that…' and provide live progress narration for multi-step background tasks.

The models are designed to execute tools and API calls in the background while continuing the conversation. This capability is intended to make voice agents more reliable and production-ready for developers and enterprises. Gemini 3.8 Live processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. The Extended Thinking variant is positioned for enterprise-grade task completion, offering deeper reasoning capabilities for complex workflows.

According to the source, Gemini 3.8 Live Extended Thinking achieved the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark, and scored 97.7% on Big Bench Audio. Gemini 3.8 Live secured second place in the Speech Agent Arena. The source notes that these models push the Pareto Frontier for complex workflows on ServiceNow’s EVA-Bench by balancing accuracy with conversational quality.

The models are accessible via the Gemini Live API, with support from developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms manage real-time media streaming infrastructure, allowing developers to focus on user experience. Google also announced partnerships with companies like Salesforce, Genspark, and Lumeris, who are utilizing the models for their latency, fluidity, and tool-calling capabilities. All audio generated by these models is watermarked with SynthID to ensure detectability and prevent misinformation.

Source details: deepmind.google

Why it matters

This release represents a significant advancement in voice-based AI interaction, moving beyond simple chat to complex, agentic task completion. By enabling models to reason and speak simultaneously, Google addresses a key limitation in current voice interfaces, making AI more intuitive for handling multi-step workflows. The integration with enterprise platforms like Salesforce and developer tools like LiveKit and Vercel suggests a push toward production-ready voice agents. This development is critical for the evolution of AI from passive assistants to active collaborators in professional and personal contexts, potentially reshaping how users interact with complex software systems through natural language.

The introduction of simultaneous reasoning and speaking addresses a fundamental bottleneck in voice AI, where models typically pause to think before responding. By maintaining conversational flow during complex task execution, these models make AI interactions feel more natural and less robotic. This is particularly important for enterprise applications where users may need to manage multiple tasks or workflows through voice commands without losing context or momentum.

The emphasis on 'production-ready' voice agents signals a shift from experimental AI demos to practical, scalable solutions. The integration with established developer platforms and enterprise partners like Salesforce suggests that Google is positioning these models as core infrastructure for the next generation of voice-driven software. This could accelerate the adoption of voice interfaces in industries where hands-free operation is critical, such as logistics, healthcare, and field services.

The high benchmark scores, particularly in agentic task completion, indicate that these models are not just better at chatting but are significantly more capable at executing complex, multi-step tasks. This is a key differentiator in the AI market, where the ability to reliably complete tasks is often more valuable than raw conversational fluency. The cost-effectiveness mentioned in the source also suggests that Google is aiming to make these advanced capabilities accessible to a broader range of developers and enterprises, not just large tech companies.

The inclusion of SynthID watermarking in all generated audio is a notable safety measure. As voice AI becomes more prevalent, the risk of deepfakes and misinformation increases. By embedding an imperceptible watermark, Google is attempting to create a technical standard for verifying AI-generated audio, which could have broader implications for trust and security in digital communications.

What to watch next

Monitor the actual latency and reliability of the 'Extended Thinking' feature in real-world scenarios, as simultaneous reasoning and speaking is a technical challenge. Watch for third-party evaluations of the claimed benchmark scores, particularly the #1 spot on the Speech to Speech Quality Index. Observe how these models perform in enterprise deployments with partners like Salesforce, and whether the SynthID watermarking effectively prevents misuse of the generated audio. Additionally, track the adoption rate among developers using the Gemini Live API and the specific use cases that emerge from the new tool-calling capabilities.

The real-world performance of the 'Extended Thinking' feature will be crucial. While the source claims it reasons and speaks simultaneously, the actual latency, accuracy, and user experience in complex scenarios may vary. Independent testing and user feedback will be important to validate these claims.

The adoption of the Gemini Live API by developers and the specific use cases that emerge will provide insight into the practical value of these models. If developers are able to build robust, reliable voice agents that can handle complex workflows, it will demonstrate the true potential of the technology.

The impact of the SynthID watermarking on the audio quality and user perception will be worth monitoring. If the watermarking is too intrusive or if it can be easily removed, it may undermine its effectiveness as a safety measure. Additionally, the response of other AI companies to this watermarking standard will be interesting to observe.

The competitive landscape in voice AI is likely to intensify as other companies respond to Google's release. Watch for announcements from OpenAI, Anthropic, and other major players regarding their own voice model capabilities and benchmarks.

Related guides & quizzes

AI AgentsAI Models ExplainedWhat is AI?Test what you know — try a free AI quizLook up an AI term in our glossary
Found this useful?