กลับไปที่ข่าว
สินค้าAI Understanding บรรยายสรุป

Google Gemini 3.8 Live เพิ่มตัวแทนวิดีโออวาตาร์สดสำหรับองค์กร

Google ประกาศว่า Gemini 3.8 Live with Live Avatar พร้อมใช้งานโดยทั่วไปแล้วผ่าน Gemini Enterprise ช่วยให้ธุรกิจปรับใช้เอเจนต์วิดีโอการสนทนาที่ลิปซิงค์กับเสียงพูดผ่านแอปบนเว็บ มือถือ และคีออสก์

4 min readRead the linked source
Source-provided image accompanying Google Gemini 3.8 Live adds live avatar video agents for enterprise
แหล่งอ้างอิงแหล่งที่มาบันทึกไว้
สำนักพิมพ์
qazinform.com
ลิงค์แหล่งที่มา
qazinform.comhttps://qazinform.com/news/google-gemini-introduces-live-avatar-for-conversational-ai-a3747f
ประเภทแหล่งที่มา
แหล่งที่มาที่เชื่อมโยง — ยังไม่ได้สร้างสถานะแหล่งที่มาหลัก
บริบทเข้าใจสิ่งนี้ใน 60 วินาที

เริ่มที่นี่

เงื่อนไขสำคัญ

API (อินเทอร์เฟซการเขียนโปรแกรมแอปพลิเคชัน)
วิธีการที่มีโครงสร้างสำหรับระบบซอฟต์แวร์หนึ่งในการส่งคำขอและรับการตอบกลับจากอีกระบบหนึ่ง
ลายน้ำ
การฝังสัญญาณที่ตรวจจับได้ในข้อความหรือสื่อที่สร้างโดย AI เพื่อให้สามารถระบุได้ว่าเป็นสัญญาณที่ผลิตด้วยเครื่องจักรในภายหลัง
คุณสมบัติ
ตัวแปรอินพุตที่แบบจำลองใช้เพื่อคาดการณ์
ทดสอบตัวเองแบบทดสอบอธิบายโมเดล AI

เกิดอะไรขึ้น

Google has launched Gemini 3.8 Live with Live Avatar as a generally available of Gemini Enterprise. The capability lets developers embed conversational video agents that generate synchronized lip‑movement avatars in real‑time. The feature, previewed at Google Cloud Next 2026, supports speech‑to‑speech interaction in 97 languages, can run background tool or API calls while maintaining dialogue, and can process live camera feeds and screen‑sharing streams. Users may select from a curated library of pre‑built avatars; custom avatars require enterprise allow‑listing and verification. All generated audio and video include an imperceptible SynthID watermark for provenance. The service is offered via US and EU endpoints with enterprise‑grade throughput, compliance and data‑governance controls.

Google’s Qazinform report states that Gemini 3.8 Live with Live Avatar is now generally available through Gemini Enterprise after a preview at Google Cloud Next 2026. The enables conversational video agents that generate avatars with synchronized lip movements across web, mobile, and interactive kiosk applications.

Gemini 3.8 Live adds native speech‑to‑speech capabilities, supporting fluid dialogue in 97 languages and automatic language detection. The model can continue interacting while executing background tool or API calls, preserving conversation context and ongoing backend transactions.

Live visual understanding lets the system process live camera feeds and screen sharing alongside audio. For misuse prevention, users must choose from a library of curated avatars; custom avatars require enterprise allow‑listing and verification. All generated streams carry an imperceptible SynthID watermark for transparency.

The service is currently offered through US and EU endpoints with provisioned throughput, enterprise compliance, and data‑governance features. No pricing or quota information is provided in the source.

รายละเอียดที่มา: qazinform.com ↗

ทำไมมันถึงสำคัญ

The launch marks the first time a major cloud AI provider has made real‑time, lip‑synced video avatars broadly available to enterprise customers. By combining speech‑to‑speech, live visual understanding, and background tool execution, Gemini 3.8 Live could reshape customer‑service, virtual tutoring, and remote assistance workflows that previously relied on static chat or text‑only bots. The inclusion of SynthID watermarks addresses growing regulatory pressure for AI‑generated content transparency. However, pricing, quota limits and the process for custom‑avatar approval remain undisclosed, leaving enterprises to negotiate terms directly with Google. The ’s availability only in the US and EU also limits immediate global rollout.

Real‑time video avatars open new interaction modalities for enterprises, potentially reducing friction in support, training, and sales scenarios where visual presence adds trust and engagement.

The ability to run background tasks while maintaining conversation flow differentiates Gemini 3.8 Live from earlier text‑only or static‑video bots, enabling more complex workflows without user interruption.

SynthID watermarks address transparency concerns, aligning with emerging regulations that require clear labeling of AI‑generated media, which could become a competitive advantage if adopted widely.

Limited regional availability and undisclosed pricing create uncertainty for global firms and may affect the speed of adoption, especially for organizations operating outside the US/EU.

Interactive Mechanism

กลไกเชิงโต้ตอบ: มันทำงานอย่างไร

สำรวจเทคโนโลยีเบื้องหลังการพัฒนานี้แบบโต้ตอบ

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
การตรวจสอบแนวคิดแบบโต้ตอบ+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

จะดูอะไรต่อไป.

Watch for pricing and quota details released by Google Cloud, as well as early adopter case studies that reveal performance and integration challenges. Regulators may scrutinize the approach, so any policy shifts around AI‑generated media could affect adoption. Competitors’ responses—especially from Microsoft and Anthropic—could accelerate similar avatar offerings. Finally, watch for extensions of the service to additional regions or to the consumer‑facing Gemini app.

Google’s forthcoming pricing tiers, usage quotas, and SLA details for Gemini Enterprise will clarify cost‑benefit calculations for potential customers.

Early enterprise deployments and case studies will reveal real‑world performance, integration effort, and user acceptance of live avatars.

Regulatory responses to SynthID could influence industry standards for AI‑generated content labeling.

Competitive moves from other AI vendors—such as Microsoft’s Copilot or Anthropic’s agents—may accelerate similar avatar capabilities, shaping the market landscape.

คำแนะนำและแบบทดสอบที่เกี่ยวข้อง

อธิบายโมเดล AIตัวแทนเอไอจริยธรรม AIPrompt Engineeringทดสอบสิ่งที่คุณรู้ — ลองแบบทดสอบ AI ฟรีค้นหาคำศัพท์ AI ในอภิธานศัพท์ของเราติดตามตัวติดตามการเปิดตัวโมเดล AI
พบว่าสิ่งนี้มีประโยชน์หรือไม่?