返回新聞
產品展示AI Understanding 簡報

Google launches Gemini 3.8 Live with Live Avatar for enterprise

Google has introduced Gemini 3.8 Live with Live Avatar, a feature that adds real-time visual presence and lip-synced video to its conversational AI, available now in Gemini Enterprise.

5 min readRead the primary source
Source-provided image accompanying Google launches Gemini 3.8 Live with Live Avatar for enterprise
主要來源文件來源記錄
出版商
deepmind.google
來源連結
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-38-live-with-live-avatar/
來源類型
主要文件-我們直接閱讀的官方公告、文件、文件或第一方頁面。
背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
生成式 AI
產生文字、圖像、音訊、視訊或程式碼等新內容的人工智慧系統。
水印
在人工智慧生成的文字或媒體中嵌入可偵測訊號,以便稍後將其識別為機器生成的。
測試一下自己AI 代理測驗

發生了什麼事

Google DeepMind announced the release of Gemini 3.8 Live with Live Avatar, a new capability that integrates low-latency streaming video generation with native live dialogue. The feature enables AI agents to listen, see, and speak with a dynamic visual persona, supporting precise lip-syncing and natural expressions. It is available immediately in Gemini Enterprise, allowing organizations to deploy interactive virtual agents for customer service and walkthroughs. The system supports asynchronous tool calling, enabling background data fetching without interrupting the conversation, and features native multilingual speech-to-speech synchronization across 97 languages. Custom avatar creation is available via enterprise allowlisting, and all outputs are watermarked with SynthID.

Google DeepMind has released Gemini 3.8 Live with Live Avatar, a feature that couples live dialogue capabilities with low-latency streaming video. This update builds on the recent launch of Gemini 3.8 Live, adding a visual layer to the conversational experience. The system processes visual and audio inputs simultaneously, generating a dynamic visual persona that listens, sees, and speaks with precise lip-syncing and natural expressions.

The feature is designed for enterprise use cases, such as customer service and interactive walkthroughs. It supports asynchronous tool calling, which allows the AI to trigger background tasks and fetch data without pausing the active dialogue. This capability ensures that complex tasks, like checking in a hotel guest, can be handled while maintaining an uninterrupted conversational flow.

Live Avatar includes native multilingual speech-to-speech synchronization, allowing it to transition seamlessly across 97 languages without degrading video fidelity. Organizations can choose from a library of preset avatars or create custom ones using a high-quality reference image. Custom avatar creation is currently restricted to enterprise allowlisting, ensuring that brands can maintain specific visual identities and character likenesses.

Safety and transparency are central to the release. All AI-generated audio and video outputs are watermarked with SynthID, an imperceptible watermark designed to help detect AI-generated content and minimize misinformation. Google states that the feature is built with strict safeguards to respect identity and keep AI-generated content transparent, with further details available in the model card.

來源詳情: deepmind.google ↗

為什麼這很重要

This launch marks a significant shift in enterprise AI from text and audio-only interactions to fully multimodal, visually present agents. By combining real-time video generation with conversational reasoning, Google is addressing a key barrier to human-like digital interaction: the lack of visual feedback and non-verbal cues. The ability to maintain uninterrupted dialogue while executing complex background tasks (asynchronous tool calling) makes these agents practical for high-stakes workflows like hotel check-ins or complex customer support. The inclusion of SynthID and strict identity safeguards addresses growing concerns about AI-generated misinformation and deepfakes in professional settings. This move positions Gemini as a comprehensive platform for immersive enterprise communication, potentially reducing the need for human agents in routine, high-volume interactions while maintaining a sense of personal connection through visual presence.

The introduction of Live Avatar represents a maturation of enterprise AI from functional text/audio bots to immersive, multimodal agents. Visual presence is a critical component of human communication, and its integration into AI agents could significantly improve user trust and engagement in digital interactions. This is particularly relevant for customer-facing roles where non-verbal cues and visual confirmation are important.

The technical implementation of asynchronous tool calling within a multimodal framework is a notable engineering achievement. It allows AI agents to perform complex, multi-step tasks without the latency penalties that typically accompany real-time video generation. This makes the technology viable for practical business applications where efficiency and responsiveness are paramount.

The focus on multilingual support and custom branding addresses the global and enterprise-specific needs of large organizations. By allowing seamless language switching and brand-specific avatar creation, Google is positioning Gemini as a flexible tool for international businesses that need to maintain consistent brand identity across diverse markets.

The inclusion of SynthID and strict identity safeguards is a proactive response to the ethical and security challenges posed by . As AI-generated video becomes more realistic, the risk of deepfakes and misinformation increases. By embedding detectable watermarks and limiting custom avatar creation to allowlisted enterprises, Google aims to mitigate these risks while still enabling innovative use cases.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
互動式概念檢查+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

接下來看什麼

Monitor the rollout of custom avatar creation beyond the initial enterprise allowlist, as this could open the door for widespread brand-specific AI personas. Watch for third-party evaluations of the lip-syncing quality and latency in real-world, high-bandwidth scenarios, as the source claims 'near real-time' performance but does not provide specific technical benchmarks. Observe how competitors respond to the integration of visual presence with agentic tool calling, which may trigger a new wave of multimodal agent development. Finally, track any regulatory or public reaction to the use of AI-generated visual personas in customer-facing roles, particularly regarding transparency and user consent.

The availability of custom avatar creation is currently limited to enterprise allowlisting. Watch for any expansion of this feature to broader API access or self-service options, which could accelerate the adoption of branded AI personas across various industries.

While Google claims 'near real-time' performance and 'fluid turn-taking,' independent testing will be necessary to verify latency and quality under varying network conditions. Look for third-party benchmarks or user reports that assess the practical usability of the feature in real-world enterprise environments.

The integration of visual presence with agentic capabilities (tool calling) sets a new standard for AI agents. Monitor how other AI providers respond to this development, as it may lead to a competitive race to incorporate multimodal, visually present agents into their enterprise offerings.

Regulatory and public scrutiny of AI-generated visual personas may increase as they become more common in customer-facing roles. Watch for any new guidelines or regulations regarding the disclosure of AI identity in visual interactions, and how Google's SynthID is received by regulators and the public.

相關指引和測驗

人工智慧代理人工智慧模型解釋AI 倫理測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語
覺得有用嗎?