返回新闻
产品展示AI Understanding 简报

Google推出Gemini 3.8实时和扩展思维模型

Google DeepMind 推出了 Gemini 3.8 Live 和 3.8 Live Extended Thinking,这是专为近实时推理、语音代理以及 Google 工作区和搜索中的复杂任务执行而设计的新实时对话模型。

5 min readRead the primary source
Source-provided image accompanying Google introduces Gemini 3.8 Live and Extended Thinking models
主要来源文件来源记录
出版商
deepmind.google
来源链接
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

API(应用程序编程接口)
一种软件系统向另一个系统发送请求并接收响应的结构化方式。
水印
在人工智能生成的文本或媒体中嵌入可检测信号,以便稍后将其识别为机器生成的。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
测试一下自己AI 代理测验

发生了什么

Google DeepMind announced the release of two new AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live dialogue, featuring near real-time reasoning capabilities that allow for simultaneous speaking and reasoning. The announcement highlights improvements in intelligence, parallel reasoning, and the ability to execute tools and API calls in the background while maintaining conversational flow. The models are available via the Gemini Live API for developers and are integrated into the Gemini app, Google Workspace, and Search.

Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as the company's most advanced live dialogue models. The primary innovation is the ability to perform near real-time reasoning, allowing the models to process complex tasks while maintaining an uninterrupted conversational flow. This is achieved through parallel reasoning, where the model can acknowledge requests with verbal cues like 'Let me check that…' and provide live progress narration for multi-step background tasks.

The models are designed to execute tools and API calls in the background while continuing the conversation. This capability is intended to make voice agents more reliable and production-ready for developers and enterprises. Gemini 3.8 Live processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. The Extended Thinking variant is positioned for enterprise-grade task completion, offering deeper reasoning capabilities for complex workflows.

According to the source, Gemini 3.8 Live Extended Thinking achieved the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking , and scored 97.7% on Big Bench Audio. Gemini 3.8 Live secured second place in the Speech Agent Arena. The source notes that these models push the Pareto Frontier for complex workflows on ServiceNow’s EVA-Bench by balancing accuracy with conversational quality.

The models are accessible via the Gemini Live API, with support from developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms manage real-time media streaming infrastructure, allowing developers to focus on user experience. Google also announced partnerships with companies like Salesforce, Genspark, and Lumeris, who are utilizing the models for their latency, fluidity, and tool-calling capabilities. All audio generated by these models is watermarked with SynthID to ensure detectability and prevent misinformation.

来源详情: deepmind.google ↗

为什么这很重要

This release represents a significant advancement in voice-based AI interaction, moving beyond simple chat to complex, agentic task completion. By enabling models to reason and speak simultaneously, Google addresses a key limitation in current voice interfaces, making AI more intuitive for handling multi-step workflows. The integration with enterprise platforms like Salesforce and developer tools like LiveKit and Vercel suggests a push toward production-ready voice agents. This development is critical for the evolution of AI from passive assistants to active collaborators in professional and personal contexts, potentially reshaping how users interact with complex software systems through natural language.

The introduction of simultaneous reasoning and speaking addresses a fundamental bottleneck in voice AI, where models typically pause to think before responding. By maintaining conversational flow during complex task execution, these models make AI interactions feel more natural and less robotic. This is particularly important for enterprise applications where users may need to manage multiple tasks or workflows through voice commands without losing context or momentum.

The emphasis on 'production-ready' voice agents signals a shift from experimental AI demos to practical, scalable solutions. The integration with established developer platforms and enterprise partners like Salesforce suggests that Google is positioning these models as core infrastructure for the next generation of voice-driven software. This could accelerate the adoption of voice interfaces in industries where hands-free operation is critical, such as logistics, healthcare, and field services.

The high scores, particularly in agentic task completion, indicate that these models are not just better at chatting but are significantly more capable at executing complex, multi-step tasks. This is a key differentiator in the AI market, where the ability to reliably complete tasks is often more valuable than raw conversational fluency. The cost-effectiveness mentioned in the source also suggests that Google is aiming to make these advanced capabilities accessible to a broader range of developers and enterprises, not just large tech companies.

The inclusion of SynthID in all generated audio is a notable safety measure. As voice AI becomes more prevalent, the risk of deepfakes and misinformation increases. By embedding an imperceptible watermark, Google is attempting to create a technical standard for verifying AI-generated audio, which could have broader implications for trust and security in digital communications.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
交互式概念检查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下来看什么

Monitor the actual latency and reliability of the 'Extended Thinking' feature in real-world scenarios, as simultaneous reasoning and speaking is a technical challenge. Watch for third-party evaluations of the claimed scores, particularly the #1 spot on the Speech to Speech Quality Index. Observe how these models perform in enterprise deployments with partners like Salesforce, and whether the SynthID effectively prevents misuse of the generated audio. Additionally, track the adoption rate among developers using the Gemini Live API and the specific use cases that emerge from the new tool-calling capabilities.

The real-world performance of the 'Extended Thinking' feature will be crucial. While the source claims it reasons and speaks simultaneously, the actual latency, accuracy, and user experience in complex scenarios may vary. Independent testing and user feedback will be important to validate these claims.

The adoption of the Gemini Live API by developers and the specific use cases that emerge will provide insight into the practical value of these models. If developers are able to build robust, reliable voice agents that can handle complex workflows, it will demonstrate the true potential of the technology.

The impact of the SynthID on the audio quality and user perception will be worth monitoring. If the watermarking is too intrusive or if it can be easily removed, it may undermine its effectiveness as a safety measure. Additionally, the response of other AI companies to this watermarking standard will be interesting to observe.

The competitive landscape in voice AI is likely to intensify as other companies respond to Google's release. Watch for announcements from OpenAI, Anthropic, and other major players regarding their own voice model capabilities and benchmarks.

相关指南和测验

人工智能代理人工智能模型解释什么是人工智能?测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语关注 AI 模型发布跟踪器
觉得这有用吗?