Dzokera kuNhau
ProductAI Understanding muchidimbu

Google launches Gemini 3.8 Live with Live Avatar for enterprise

Google has introduced Gemini 3.8 Live with Live Avatar, a feature that adds real-time visual presence and lip-synced video to its conversational AI, available now in Gemini Enterprise.

5 min readRead the primary source
Source-provided image accompanying Google launches Gemini 3.8 Live with Live Avatar for enterprise
Primary-source documentKwakanyorwa
Muparidzi
deepmind.google
Source link
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-38-live-with-live-avatar/
Source type
Gwaro rekutanga - chiziviso chepamutemo, bepa, faira, kana peji rebato rekutanga ratinoverenga zvakananga.
ContextNzwisisa izvi mumasekonzi makumi matanhatu

Tanga pano

Matemu akakosha

API (Application Programming Interface)
Nzira yakarongeka yeimwe software system yekutumira zvikumbiro uye kugamuchira mhinduro kubva kune imwe system.
AI inogadzira
AI masisitimu anoburitsa zvinyorwa zvitsva senge zvinyorwa, mifananidzo, odhiyo, vhidhiyo, kana kodhi.
Watermarking
Kupinza chiratidzo chinooneka muAI-yakagadzirwa zvinyorwa kana midhiya kuti igozoonekwa seyakagadzirwa muchina.
Zviedze iwe pachakoAI Agents Quiz

Chii chaitika

Google DeepMind announced the release of Gemini 3.8 Live with Live Avatar, a new capability that integrates low-latency streaming video generation with native live dialogue. The feature enables AI agents to listen, see, and speak with a dynamic visual persona, supporting precise lip-syncing and natural expressions. It is available immediately in Gemini Enterprise, allowing organizations to deploy interactive virtual agents for customer service and walkthroughs. The system supports asynchronous tool calling, enabling background data fetching without interrupting the conversation, and features native multilingual speech-to-speech synchronization across 97 languages. Custom avatar creation is available via enterprise allowlisting, and all outputs are watermarked with SynthID.

Google DeepMind has released Gemini 3.8 Live with Live Avatar, a feature that couples live dialogue capabilities with low-latency streaming video. This update builds on the recent launch of Gemini 3.8 Live, adding a visual layer to the conversational experience. The system processes visual and audio inputs simultaneously, generating a dynamic visual persona that listens, sees, and speaks with precise lip-syncing and natural expressions.

The feature is designed for enterprise use cases, such as customer service and interactive walkthroughs. It supports asynchronous tool calling, which allows the AI to trigger background tasks and fetch data without pausing the active dialogue. This capability ensures that complex tasks, like checking in a hotel guest, can be handled while maintaining an uninterrupted conversational flow.

Live Avatar includes native multilingual speech-to-speech synchronization, allowing it to transition seamlessly across 97 languages without degrading video fidelity. Organizations can choose from a library of preset avatars or create custom ones using a high-quality reference image. Custom avatar creation is currently restricted to enterprise allowlisting, ensuring that brands can maintain specific visual identities and character likenesses.

Safety and transparency are central to the release. All AI-generated audio and video outputs are watermarked with SynthID, an imperceptible watermark designed to help detect AI-generated content and minimize misinformation. Google states that the feature is built with strict safeguards to respect identity and keep AI-generated content transparent, with further details available in the model card.

Kwakabva mashoko: deepmind.google โ†—

Nei zvichikosha

This launch marks a significant shift in enterprise AI from text and audio-only interactions to fully multimodal, visually present agents. By combining real-time video generation with conversational reasoning, Google is addressing a key barrier to human-like digital interaction: the lack of visual feedback and non-verbal cues. The ability to maintain uninterrupted dialogue while executing complex background tasks (asynchronous tool calling) makes these agents practical for high-stakes workflows like hotel check-ins or complex customer support. The inclusion of SynthID and strict identity safeguards addresses growing concerns about AI-generated misinformation and deepfakes in professional settings. This move positions Gemini as a comprehensive platform for immersive enterprise communication, potentially reducing the need for human agents in routine, high-volume interactions while maintaining a sense of personal connection through visual presence.

The introduction of Live Avatar represents a maturation of enterprise AI from functional text/audio bots to immersive, multimodal agents. Visual presence is a critical component of human communication, and its integration into AI agents could significantly improve user trust and engagement in digital interactions. This is particularly relevant for customer-facing roles where non-verbal cues and visual confirmation are important.

The technical implementation of asynchronous tool calling within a multimodal framework is a notable engineering achievement. It allows AI agents to perform complex, multi-step tasks without the latency penalties that typically accompany real-time video generation. This makes the technology viable for practical business applications where efficiency and responsiveness are paramount.

The focus on multilingual support and custom branding addresses the global and enterprise-specific needs of large organizations. By allowing seamless language switching and brand-specific avatar creation, Google is positioning Gemini as a flexible tool for international businesses that need to maintain consistent brand identity across diverse markets.

The inclusion of SynthID and strict identity safeguards is a proactive response to the ethical and security challenges posed by . As AI-generated video becomes more realistic, the risk of deepfakes and misinformation increases. By embedding detectable watermarks and limiting custom avatar creation to allowlisted enterprises, Google aims to mitigate these risks while still enabling innovative use cases.

Interactive Mechanism

Interactive Mechanism: Iyo Inonyatsoshanda

Ongorora ari pasi tekinoroji kuseri kwekusimudzira uku uchipindirana.

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
Interactive Concept Check+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

Zvekutarisa zvinotevera

Monitor the rollout of custom avatar creation beyond the initial enterprise allowlist, as this could open the door for widespread brand-specific AI personas. Watch for third-party evaluations of the lip-syncing quality and latency in real-world, high-bandwidth scenarios, as the source claims 'near real-time' performance but does not provide specific technical benchmarks. Observe how competitors respond to the integration of visual presence with agentic tool calling, which may trigger a new wave of multimodal agent development. Finally, track any regulatory or public reaction to the use of AI-generated visual personas in customer-facing roles, particularly regarding transparency and user consent.

The availability of custom avatar creation is currently limited to enterprise allowlisting. Watch for any expansion of this feature to broader API access or self-service options, which could accelerate the adoption of branded AI personas across various industries.

While Google claims 'near real-time' performance and 'fluid turn-taking,' independent testing will be necessary to verify latency and quality under varying network conditions. Look for third-party benchmarks or user reports that assess the practical usability of the feature in real-world enterprise environments.

The integration of visual presence with agentic capabilities (tool calling) sets a new standard for AI agents. Monitor how other AI providers respond to this development, as it may lead to a competitive race to incorporate multimodal, visually present agents into their enterprise offerings.

Regulatory and public scrutiny of AI-generated visual personas may increase as they become more common in customer-facing roles. Watch for any new guidelines or regulations regarding the disclosure of AI identity in visual interactions, and how Google's SynthID is received by regulators and the public.

Related guides & Quizzes

AI AgentsAI Models InotsanangurwaTsika dzeAIEdza zvaunoziva - edza yemahara AI quizTarisa kumusoro izwi reAI mune yedu glossary
Wakawana izvi zvinobatsira?