Retour aux Actualités
ProduitBriefing AI Understanding

Google présente les modèles Gemini 3.8 Live et Extended Thinking

Google DeepMind a lancé Gemini 3.8 Live et 3.8 Live Extended Thinking, de nouveaux modèles de dialogue en direct conçus pour le raisonnement en temps quasi réel, les agents vocaux et l'exécution de tâches complexes dans l'espace de travail et la recherche Google.

5 min readRead the primary source
Source-provided image accompanying Google introduces Gemini 3.8 Live and Extended Thinking models
Document de source principaleSource enregistrée
Éditeur
deepmind.google
Lien source
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/
Type de source
Document principal : une annonce officielle, un document, un dépôt ou une page de première partie que nous lisons directement.
ContexteComprenez cela en 60 secondes

Commencez ici

Termes clés

API (interface de programmation d'applications)
Une manière structurée permettant à un système logiciel d'envoyer des requêtes et de recevoir des réponses d'un autre système.
Filigrane
Intégrer un signal détectable dans un texte ou un média généré par l'IA afin qu'il puisse ensuite être identifié comme étant produit par une machine.
Référence
Un test ou un ensemble de données standardisé utilisé pour mesurer et comparer les performances du modèle.
Testez-vousQuiz sur les agents IA

Que s'est-il passé

Google DeepMind announced the release of two new AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live dialogue, featuring near real-time reasoning capabilities that allow for simultaneous speaking and reasoning. The announcement highlights improvements in intelligence, parallel reasoning, and the ability to execute tools and API calls in the background while maintaining conversational flow. The models are available via the Gemini Live API for developers and are integrated into the Gemini app, Google Workspace, and Search.

Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as the company's most advanced live dialogue models. The primary innovation is the ability to perform near real-time reasoning, allowing the models to process complex tasks while maintaining an uninterrupted conversational flow. This is achieved through parallel reasoning, where the model can acknowledge requests with verbal cues like 'Let me check that…' and provide live progress narration for multi-step background tasks.

The models are designed to execute tools and API calls in the background while continuing the conversation. This capability is intended to make voice agents more reliable and production-ready for developers and enterprises. Gemini 3.8 Live processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. The Extended Thinking variant is positioned for enterprise-grade task completion, offering deeper reasoning capabilities for complex workflows.

According to the source, Gemini 3.8 Live Extended Thinking achieved the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking , and scored 97.7% on Big Bench Audio. Gemini 3.8 Live secured second place in the Speech Agent Arena. The source notes that these models push the Pareto Frontier for complex workflows on ServiceNow’s EVA-Bench by balancing accuracy with conversational quality.

The models are accessible via the Gemini Live API, with support from developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms manage real-time media streaming infrastructure, allowing developers to focus on user experience. Google also announced partnerships with companies like Salesforce, Genspark, and Lumeris, who are utilizing the models for their latency, fluidity, and tool-calling capabilities. All audio generated by these models is watermarked with SynthID to ensure detectability and prevent misinformation.

Détails de la source: deepmind.google ↗

Pourquoi c'est important

This release represents a significant advancement in voice-based AI interaction, moving beyond simple chat to complex, agentic task completion. By enabling models to reason and speak simultaneously, Google addresses a key limitation in current voice interfaces, making AI more intuitive for handling multi-step workflows. The integration with enterprise platforms like Salesforce and developer tools like LiveKit and Vercel suggests a push toward production-ready voice agents. This development is critical for the evolution of AI from passive assistants to active collaborators in professional and personal contexts, potentially reshaping how users interact with complex software systems through natural language.

The introduction of simultaneous reasoning and speaking addresses a fundamental bottleneck in voice AI, where models typically pause to think before responding. By maintaining conversational flow during complex task execution, these models make AI interactions feel more natural and less robotic. This is particularly important for enterprise applications where users may need to manage multiple tasks or workflows through voice commands without losing context or momentum.

The emphasis on 'production-ready' voice agents signals a shift from experimental AI demos to practical, scalable solutions. The integration with established developer platforms and enterprise partners like Salesforce suggests that Google is positioning these models as core infrastructure for the next generation of voice-driven software. This could accelerate the adoption of voice interfaces in industries where hands-free operation is critical, such as logistics, healthcare, and field services.

The high scores, particularly in agentic task completion, indicate that these models are not just better at chatting but are significantly more capable at executing complex, multi-step tasks. This is a key differentiator in the AI market, where the ability to reliably complete tasks is often more valuable than raw conversational fluency. The cost-effectiveness mentioned in the source also suggests that Google is aiming to make these advanced capabilities accessible to a broader range of developers and enterprises, not just large tech companies.

The inclusion of SynthID in all generated audio is a notable safety measure. As voice AI becomes more prevalent, the risk of deepfakes and misinformation increases. By embedding an imperceptible watermark, Google is attempting to create a technical standard for verifying AI-generated audio, which could have broader implications for trust and security in digital communications.

Interactive Mechanism

Mécanisme interactif : comment cela fonctionne réellement

Explorez de manière interactive la technologie sous-jacente à ce développement.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Vérification de concept interactive+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Que regarder ensuite

Monitor the actual latency and reliability of the 'Extended Thinking' feature in real-world scenarios, as simultaneous reasoning and speaking is a technical challenge. Watch for third-party evaluations of the claimed scores, particularly the #1 spot on the Speech to Speech Quality Index. Observe how these models perform in enterprise deployments with partners like Salesforce, and whether the SynthID effectively prevents misuse of the generated audio. Additionally, track the adoption rate among developers using the Gemini Live API and the specific use cases that emerge from the new tool-calling capabilities.

The real-world performance of the 'Extended Thinking' feature will be crucial. While the source claims it reasons and speaks simultaneously, the actual latency, accuracy, and user experience in complex scenarios may vary. Independent testing and user feedback will be important to validate these claims.

The adoption of the Gemini Live API by developers and the specific use cases that emerge will provide insight into the practical value of these models. If developers are able to build robust, reliable voice agents that can handle complex workflows, it will demonstrate the true potential of the technology.

The impact of the SynthID on the audio quality and user perception will be worth monitoring. If the watermarking is too intrusive or if it can be easily removed, it may undermine its effectiveness as a safety measure. Additionally, the response of other AI companies to this watermarking standard will be interesting to observe.

The competitive landscape in voice AI is likely to intensify as other companies respond to Google's release. Watch for announcements from OpenAI, Anthropic, and other major players regarding their own voice model capabilities and benchmarks.

Guides et quiz associés

Agents IAModèles d'IA expliquésQu’est-ce que l’IA ?Testez ce que vous savez : essayez un quiz gratuit sur l'IARecherchez un terme d'IA dans notre glossaireSuivez le suivi des versions du modèle AI
Vous avez trouvé cela utile ?