Вернуться к новостям
ПродуктAI Understanding брифинг

Google представляет модели Gemini 3.8 Live и Extended Thinking

Google DeepMind выпустил Gemini 3.8 Live и 3.8 Live Extended Thinking, новые модели живого диалога, предназначенные для рассуждений практически в реальном времени, голосовых агентов и выполнения сложных задач в Google Workspace и Search.

5 min readRead the primary source
Source-provided image accompanying Google introduces Gemini 3.8 Live and Extended Thinking models
ПервоисточникИсточник записан
Издатель
deepmind.google
Ссылка на источник
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/
Тип источника
Первичный документ — официальное объявление, документ, файл или собственная страница, которую мы читаем напрямую.
КонтекстПоймите это за 60 секунд

Начните здесь

Ключевые термины

API (интерфейс прикладного программирования)
Структурированный способ отправки одной программной системой запросов и получения ответов от другой системы.
Водяные знаки
Встраивание обнаруживаемого сигнала в текст или медиафайлы, сгенерированные ИИ, чтобы впоследствии его можно было идентифицировать как созданный машиной.
Контрольный показатель
Стандартизированный тест или набор данных, используемый для измерения и сравнения производительности модели.
Проверьте себяВикторина «Агенты ИИ»

Что случилось

Google DeepMind announced the release of two new AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live dialogue, featuring near real-time reasoning capabilities that allow for simultaneous speaking and reasoning. The announcement highlights improvements in intelligence, parallel reasoning, and the ability to execute tools and API calls in the background while maintaining conversational flow. The models are available via the Gemini Live API for developers and are integrated into the Gemini app, Google Workspace, and Search.

Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as the company's most advanced live dialogue models. The primary innovation is the ability to perform near real-time reasoning, allowing the models to process complex tasks while maintaining an uninterrupted conversational flow. This is achieved through parallel reasoning, where the model can acknowledge requests with verbal cues like 'Let me check that…' and provide live progress narration for multi-step background tasks.

The models are designed to execute tools and API calls in the background while continuing the conversation. This capability is intended to make voice agents more reliable and production-ready for developers and enterprises. Gemini 3.8 Live processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. The Extended Thinking variant is positioned for enterprise-grade task completion, offering deeper reasoning capabilities for complex workflows.

According to the source, Gemini 3.8 Live Extended Thinking achieved the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking , and scored 97.7% on Big Bench Audio. Gemini 3.8 Live secured second place in the Speech Agent Arena. The source notes that these models push the Pareto Frontier for complex workflows on ServiceNow’s EVA-Bench by balancing accuracy with conversational quality.

The models are accessible via the Gemini Live API, with support from developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms manage real-time media streaming infrastructure, allowing developers to focus on user experience. Google also announced partnerships with companies like Salesforce, Genspark, and Lumeris, who are utilizing the models for their latency, fluidity, and tool-calling capabilities. All audio generated by these models is watermarked with SynthID to ensure detectability and prevent misinformation.

Подробности об источнике: deepmind.google ↗

Почему это важно

This release represents a significant advancement in voice-based AI interaction, moving beyond simple chat to complex, agentic task completion. By enabling models to reason and speak simultaneously, Google addresses a key limitation in current voice interfaces, making AI more intuitive for handling multi-step workflows. The integration with enterprise platforms like Salesforce and developer tools like LiveKit and Vercel suggests a push toward production-ready voice agents. This development is critical for the evolution of AI from passive assistants to active collaborators in professional and personal contexts, potentially reshaping how users interact with complex software systems through natural language.

The introduction of simultaneous reasoning and speaking addresses a fundamental bottleneck in voice AI, where models typically pause to think before responding. By maintaining conversational flow during complex task execution, these models make AI interactions feel more natural and less robotic. This is particularly important for enterprise applications where users may need to manage multiple tasks or workflows through voice commands without losing context or momentum.

The emphasis on 'production-ready' voice agents signals a shift from experimental AI demos to practical, scalable solutions. The integration with established developer platforms and enterprise partners like Salesforce suggests that Google is positioning these models as core infrastructure for the next generation of voice-driven software. This could accelerate the adoption of voice interfaces in industries where hands-free operation is critical, such as logistics, healthcare, and field services.

The high scores, particularly in agentic task completion, indicate that these models are not just better at chatting but are significantly more capable at executing complex, multi-step tasks. This is a key differentiator in the AI market, where the ability to reliably complete tasks is often more valuable than raw conversational fluency. The cost-effectiveness mentioned in the source also suggests that Google is aiming to make these advanced capabilities accessible to a broader range of developers and enterprises, not just large tech companies.

The inclusion of SynthID in all generated audio is a notable safety measure. As voice AI becomes more prevalent, the risk of deepfakes and misinformation increases. By embedding an imperceptible watermark, Google is attempting to create a technical standard for verifying AI-generated audio, which could have broader implications for trust and security in digital communications.

Interactive Mechanism

Интерактивный механизм: как он на самом деле работает

Изучите технологию, лежащую в основе этой разработки, в интерактивном режиме.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Интерактивная проверка концепции+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

Что посмотреть дальше

Monitor the actual latency and reliability of the 'Extended Thinking' feature in real-world scenarios, as simultaneous reasoning and speaking is a technical challenge. Watch for third-party evaluations of the claimed scores, particularly the #1 spot on the Speech to Speech Quality Index. Observe how these models perform in enterprise deployments with partners like Salesforce, and whether the SynthID effectively prevents misuse of the generated audio. Additionally, track the adoption rate among developers using the Gemini Live API and the specific use cases that emerge from the new tool-calling capabilities.

The real-world performance of the 'Extended Thinking' feature will be crucial. While the source claims it reasons and speaks simultaneously, the actual latency, accuracy, and user experience in complex scenarios may vary. Independent testing and user feedback will be important to validate these claims.

The adoption of the Gemini Live API by developers and the specific use cases that emerge will provide insight into the practical value of these models. If developers are able to build robust, reliable voice agents that can handle complex workflows, it will demonstrate the true potential of the technology.

The impact of the SynthID on the audio quality and user perception will be worth monitoring. If the watermarking is too intrusive or if it can be easily removed, it may undermine its effectiveness as a safety measure. Additionally, the response of other AI companies to this watermarking standard will be interesting to observe.

The competitive landscape in voice AI is likely to intensify as other companies respond to Google's release. Watch for announcements from OpenAI, Anthropic, and other major players regarding their own voice model capabilities and benchmarks.

Сопутствующие руководства и викторины

ИИ-агентыОбъяснение моделей искусственного интеллектаЧто такое ИИ?Проверьте свои знания — пройдите бесплатную викторину по искусственному интеллектуНайдите термин ИИ в нашем глоссарии.Следите за трекером выпуска моделей AI
Нашли это полезным?