ወደ ዜና ተመለስ
ምርትAI Understanding አጭር መግለጫ

Google Gemini 3.8 Live with Live Avatar for Enterprise አስጀመረ።

Google Gemini 3.8 Live with Live Avatar አስተዋውቋል፣ ይህ ባህሪ በአሁን ጊዜ በGemini ኢንተርፕራይዝ ውስጥ በእውነተኛ ጊዜ ምስላዊ መገኘት እና በከንፈር የተመሳሰለ ቪዲዮን ወደ የውይይት AI የሚጨምር ነው።

5 min readRead the primary source
Source-provided image accompanying Google launches Gemini 3.8 Live with Live Avatar for enterprise
ዋና-ምንጭ ሰነድምንጭ ተመዝግቧል
አታሚ
deepmind.google
ምንጭ አገናኝ
deepmind.googlehttps://deepmind.google/blog/introducing-gemini-38-live-with-live-avatar/
የምንጭ ዓይነት
ዋና ሰነድ - ኦፊሴላዊ ማስታወቂያ ፣ ወረቀት ፣ ፋይል ወይም የመጀመሪያ ወገን ገጽ በቀጥታ እናነባለን።
አውድይህንን በ60 ሰከንድ ውስጥ ይረዱት።

እዚ ጀምር

ቁልፍ ቃላት

ኤፒአይ (የመተግበሪያ ፕሮግራሚንግ በይነገጽ)
አንድ የሶፍትዌር ስርዓት ከሌላ ስርዓት ጥያቄዎችን ለመላክ እና ምላሽ የሚቀበልበት የተቀናጀ መንገድ።
አመንጪ AI
እንደ ጽሑፍ፣ ምስሎች፣ ኦዲዮ፣ ቪዲዮ ወይም ኮድ ያሉ አዲስ ይዘቶችን የሚያመርቱ AI ስርዓቶች።
የውሃ ምልክት ማድረግ
ሊታወቅ የሚችል ምልክት በ AI የመነጨ ጽሑፍ ወይም ሚዲያ ውስጥ በመክተት በኋላ በማሽን እንደተመረተ ሊታወቅ ይችላል።
እራስህን ፈትን።AI ወኪሎች ጥያቄዎች

ምን ተፈጠረ

Google DeepMind announced the release of Gemini 3.8 Live with Live Avatar, a new capability that integrates low-latency streaming video generation with native live dialogue. The feature enables AI agents to listen, see, and speak with a dynamic visual persona, supporting precise lip-syncing and natural expressions. It is available immediately in Gemini Enterprise, allowing organizations to deploy interactive virtual agents for customer service and walkthroughs. The system supports asynchronous tool calling, enabling background data fetching without interrupting the conversation, and features native multilingual speech-to-speech synchronization across 97 languages. Custom avatar creation is available via enterprise allowlisting, and all outputs are watermarked with SynthID.

Google DeepMind has released Gemini 3.8 Live with Live Avatar, a feature that couples live dialogue capabilities with low-latency streaming video. This update builds on the recent launch of Gemini 3.8 Live, adding a visual layer to the conversational experience. The system processes visual and audio inputs simultaneously, generating a dynamic visual persona that listens, sees, and speaks with precise lip-syncing and natural expressions.

The feature is designed for enterprise use cases, such as customer service and interactive walkthroughs. It supports asynchronous tool calling, which allows the AI to trigger background tasks and fetch data without pausing the active dialogue. This capability ensures that complex tasks, like checking in a hotel guest, can be handled while maintaining an uninterrupted conversational flow.

Live Avatar includes native multilingual speech-to-speech synchronization, allowing it to transition seamlessly across 97 languages without degrading video fidelity. Organizations can choose from a library of preset avatars or create custom ones using a high-quality reference image. Custom avatar creation is currently restricted to enterprise allowlisting, ensuring that brands can maintain specific visual identities and character likenesses.

Safety and transparency are central to the release. All AI-generated audio and video outputs are watermarked with SynthID, an imperceptible watermark designed to help detect AI-generated content and minimize misinformation. Google states that the feature is built with strict safeguards to respect identity and keep AI-generated content transparent, with further details available in the model card.

የምንጭ ዝርዝሮች: deepmind.google ↗

ለምን አስፈላጊ ነው።

This launch marks a significant shift in enterprise AI from text and audio-only interactions to fully multimodal, visually present agents. By combining real-time video generation with conversational reasoning, Google is addressing a key barrier to human-like digital interaction: the lack of visual feedback and non-verbal cues. The ability to maintain uninterrupted dialogue while executing complex background tasks (asynchronous tool calling) makes these agents practical for high-stakes workflows like hotel check-ins or complex customer support. The inclusion of SynthID and strict identity safeguards addresses growing concerns about AI-generated misinformation and deepfakes in professional settings. This move positions Gemini as a comprehensive platform for immersive enterprise communication, potentially reducing the need for human agents in routine, high-volume interactions while maintaining a sense of personal connection through visual presence.

The introduction of Live Avatar represents a maturation of enterprise AI from functional text/audio bots to immersive, multimodal agents. Visual presence is a critical component of human communication, and its integration into AI agents could significantly improve user trust and engagement in digital interactions. This is particularly relevant for customer-facing roles where non-verbal cues and visual confirmation are important.

The technical implementation of asynchronous tool calling within a multimodal framework is a notable engineering achievement. It allows AI agents to perform complex, multi-step tasks without the latency penalties that typically accompany real-time video generation. This makes the technology viable for practical business applications where efficiency and responsiveness are paramount.

The focus on multilingual support and custom branding addresses the global and enterprise-specific needs of large organizations. By allowing seamless language switching and brand-specific avatar creation, Google is positioning Gemini as a flexible tool for international businesses that need to maintain consistent brand identity across diverse markets.

The inclusion of SynthID and strict identity safeguards is a proactive response to the ethical and security challenges posed by . As AI-generated video becomes more realistic, the risk of deepfakes and misinformation increases. By embedding detectable watermarks and limiting custom avatar creation to allowlisted enterprises, Google aims to mitigate these risks while still enabling innovative use cases.

Interactive Mechanism

በይነተገናኝ ሜካኒዝም፡ በትክክል እንዴት እንደሚሰራ

ከዚህ ልማት በስተጀርባ ያለውን ቴክኖሎጂ በይነተገናኝ ያስሱ።

Thinking Budget (Test-Time Tokens):1,024 tokens
Complex Accuracy79%Math & Code Logic
Latency3.2sTime to first full output
Inference Cost$0.0092Per query estimated
Reasoning StyleStep VerificationInternal chain depth
Active Thinking Trace:
1Deconstruct user problem into formal constraints
2Propose candidate hypotheses & step-by-step calculation
3Self-correction: Backtrack and refute subtle edge cases
4Exhaustive consistency check & final output synthesis
Core takeaway: Test-time compute fundamentally changes AI economics. Instead of only scaling during pre-training, giving reasoning models more tokens at inference time allows them to systematically solve PhD-level STEM problems.
በይነተገናኝ ጽንሰ-ሐሳብ ቼክ+10 Points
AI Agents Quiz

What most distinguishes an AI agent from a basic chatbot?

ቀጥሎ ምን እንደሚታይ

Monitor the rollout of custom avatar creation beyond the initial enterprise allowlist, as this could open the door for widespread brand-specific AI personas. Watch for third-party evaluations of the lip-syncing quality and latency in real-world, high-bandwidth scenarios, as the source claims 'near real-time' performance but does not provide specific technical benchmarks. Observe how competitors respond to the integration of visual presence with agentic tool calling, which may trigger a new wave of multimodal agent development. Finally, track any regulatory or public reaction to the use of AI-generated visual personas in customer-facing roles, particularly regarding transparency and user consent.

The availability of custom avatar creation is currently limited to enterprise allowlisting. Watch for any expansion of this feature to broader API access or self-service options, which could accelerate the adoption of branded AI personas across various industries.

While Google claims 'near real-time' performance and 'fluid turn-taking,' independent testing will be necessary to verify latency and quality under varying network conditions. Look for third-party benchmarks or user reports that assess the practical usability of the feature in real-world enterprise environments.

The integration of visual presence with agentic capabilities (tool calling) sets a new standard for AI agents. Monitor how other AI providers respond to this development, as it may lead to a competitive race to incorporate multimodal, visually present agents into their enterprise offerings.

Regulatory and public scrutiny of AI-generated visual personas may increase as they become more common in customer-facing roles. Watch for any new guidelines or regulations regarding the disclosure of AI identity in visual interactions, and how Google's SynthID is received by regulators and the public.

ተዛማጅ መመሪያዎች እና ጥያቄዎች

AI ወኪሎችAI ሞዴሎች ተብራርተዋልየAI ሥነ ምግባርየሚያውቁትን ይሞክሩ - ነፃ የ AI ጥያቄዎችን ይሞክሩበእኛ የቃላት መፍቻ ውስጥ የ AI ቃልን ይፈልጉ
ይህ ጠቃሚ ሆኖ ተገኝቷል?