返回新聞
產品展示AI Understanding 簡報

Google推出Gemini 3.8即時與擴展思維模型

Google DeepMind 推出了 Gemini 3.8 Live 和 3.8 Live Extended Thinking,這是專為近實時推理、語音代理以及 Google 工作區和搜尋中的複雜任務執行而設計的新即時對話模型。

5 min readRead the primary source
Source-provided image accompanying Google introduces Gemini 3.8 Live and Extended Thinking models
主要來源文件來源記錄
出版商
deepmind.google
來源連結
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-3-8-live-and-3-8-live-extended-thinking/
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
水印
在人工智慧生成的文字或媒體中嵌入可偵測訊號,以便稍後將其識別為機器生成的。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
測試一下自己AI 代理測驗

發生了什麼事

Google DeepMind announced the release of two new AI models: Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live dialogue, featuring near real-time reasoning capabilities that allow for simultaneous speaking and reasoning. The announcement highlights improvements in intelligence, parallel reasoning, and the ability to execute tools and API calls in the background while maintaining conversational flow. The models are available via the Gemini Live API for developers and are integrated into the Gemini app, Google Workspace, and Search.

Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, describing them as the company's most advanced live dialogue models. The primary innovation is the ability to perform near real-time reasoning, allowing the models to process complex tasks while maintaining an uninterrupted conversational flow. This is achieved through parallel reasoning, where the model can acknowledge requests with verbal cues like 'Let me check that…' and provide live progress narration for multi-step background tasks.

The models are designed to execute tools and API calls in the background while continuing the conversation. This capability is intended to make voice agents more reliable and production-ready for developers and enterprises. Gemini 3.8 Live processes visual inputs in near real-time and automatically detects and transitions between 97 supported languages mid-conversation. The Extended Thinking variant is positioned for enterprise-grade task completion, offering deeper reasoning capabilities for complex workflows.

According to the source, Gemini 3.8 Live Extended Thinking achieved the #1 overall spot on Artificial Analysis' Speech to Speech Quality Index with a score of 82.6. It also leads in agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking , and scored 97.7% on Big Bench Audio. Gemini 3.8 Live secured second place in the Speech Agent Arena. The source notes that these models push the Pareto Frontier for complex workflows on ServiceNow’s EVA-Bench by balancing accuracy with conversational quality.

The models are accessible via the Gemini Live API, with support from developer platforms such as Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents. These platforms manage real-time media streaming infrastructure, allowing developers to focus on user experience. Google also announced partnerships with companies like Salesforce, Genspark, and Lumeris, who are utilizing the models for their latency, fluidity, and tool-calling capabilities. All audio generated by these models is watermarked with SynthID to ensure detectability and prevent misinformation.

來源詳情: deepmind.google ↗

為什麼這很重要

This release represents a significant advancement in voice-based AI interaction, moving beyond simple chat to complex, agentic task completion. By enabling models to reason and speak simultaneously, Google addresses a key limitation in current voice interfaces, making AI more intuitive for handling multi-step workflows. The integration with enterprise platforms like Salesforce and developer tools like LiveKit and Vercel suggests a push toward production-ready voice agents. This development is critical for the evolution of AI from passive assistants to active collaborators in professional and personal contexts, potentially reshaping how users interact with complex software systems through natural language.

The introduction of simultaneous reasoning and speaking addresses a fundamental bottleneck in voice AI, where models typically pause to think before responding. By maintaining conversational flow during complex task execution, these models make AI interactions feel more natural and less robotic. This is particularly important for enterprise applications where users may need to manage multiple tasks or workflows through voice commands without losing context or momentum.

The emphasis on 'production-ready' voice agents signals a shift from experimental AI demos to practical, scalable solutions. The integration with established developer platforms and enterprise partners like Salesforce suggests that Google is positioning these models as core infrastructure for the next generation of voice-driven software. This could accelerate the adoption of voice interfaces in industries where hands-free operation is critical, such as logistics, healthcare, and field services.

The high scores, particularly in agentic task completion, indicate that these models are not just better at chatting but are significantly more capable at executing complex, multi-step tasks. This is a key differentiator in the AI market, where the ability to reliably complete tasks is often more valuable than raw conversational fluency. The cost-effectiveness mentioned in the source also suggests that Google is aiming to make these advanced capabilities accessible to a broader range of developers and enterprises, not just large tech companies.

The inclusion of SynthID in all generated audio is a notable safety measure. As voice AI becomes more prevalent, the risk of deepfakes and misinformation increases. By embedding an imperceptible watermark, Google is attempting to create a technical standard for verifying AI-generated audio, which could have broader implications for trust and security in digital communications.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

An agent must create a draft calendar event for Tuesday at 2 p.m. Which evidence would establish the requested result?

接下來看什麼

Monitor the actual latency and reliability of the 'Extended Thinking' feature in real-world scenarios, as simultaneous reasoning and speaking is a technical challenge. Watch for third-party evaluations of the claimed scores, particularly the #1 spot on the Speech to Speech Quality Index. Observe how these models perform in enterprise deployments with partners like Salesforce, and whether the SynthID effectively prevents misuse of the generated audio. Additionally, track the adoption rate among developers using the Gemini Live API and the specific use cases that emerge from the new tool-calling capabilities.

The real-world performance of the 'Extended Thinking' feature will be crucial. While the source claims it reasons and speaks simultaneously, the actual latency, accuracy, and user experience in complex scenarios may vary. Independent testing and user feedback will be important to validate these claims.

The adoption of the Gemini Live API by developers and the specific use cases that emerge will provide insight into the practical value of these models. If developers are able to build robust, reliable voice agents that can handle complex workflows, it will demonstrate the true potential of the technology.

The impact of the SynthID on the audio quality and user perception will be worth monitoring. If the watermarking is too intrusive or if it can be easily removed, it may undermine its effectiveness as a safety measure. Additionally, the response of other AI companies to this watermarking standard will be interesting to observe.

The competitive landscape in voice AI is likely to intensify as other companies respond to Google's release. Watch for announcements from OpenAI, Anthropic, and other major players regarding their own voice model capabilities and benchmarks.

相關指引和測驗

人工智慧代理人工智慧模型解釋什麼是人工智慧?測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器
覺得有用嗎?