返回新闻
产品展示AI Understanding 简报

Google launches Gemini 3.8 Flash TTS models

Google DeepMind has released Gemini 3.8 Flash TTS and Flash-Lite TTS, new text-to-speech models available in Google AI Studio and via the Gemini API, featuring custom voice creation and SynthID watermarking.

4 min readRead the primary source
Source-provided image accompanying Google launches Gemini 3.8 Flash TTS models
主要来源文件来源记录
出版商
deepmind.google
来源链接
deepmind.googlehttps://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/
来源类型
主要文件——我们直接阅读的官方公告、文件、文件或第一方页面。
背景60 秒内了解这一点

从这里开始

关键术语

API(应用程序编程接口)
一种软件系统向另一个系统发送请求并接收响应的结构化方式。
水印
在人工智能生成的文本或媒体中嵌入可检测信号,以便稍后将其识别为机器生成的。
基准测试
用于测量和比较模型性能的标准化测试或数据集。
测试一下自己AI 模型解释测验

发生了什么

Google DeepMind announced the release of two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models are available immediately in Google AI Studio and through the Gemini API, with integrations for platforms like Agora, LiveKit, and Vercel. The release expands the Gemini Audio family, adding capabilities for generating custom character voices and directing scene dialogue with precise control over delivery.

Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as the most expressive audio generation models in the Gemini family. The announcement states that these models transform voice generation from static presets into a dynamic creative studio, allowing for the creation of custom character voices and the direction of scene dialogue.

The models are available starting today in Google AI Studio, where they function as a voice design workspace. Users can prompt new vocal identities from scratch or replicate their own voices, then use a dual-speaker screenplay editor to direct line-by-line delivery. Access is also provided via the Gemini API, enabling developer platforms such as Agora, LiveKit, Pipecat, and Vercel to build and deploy speech generation experiences.

Google claims that Gemini 3.8 Flash TTS secures the #1 overall spot on Hume AI’s Voice Design with a score of 71.4 and leads in accent modeling with a score of 60.8. Both models reportedly secure the #1 and #2 spots on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, the models are said to hold top positions in key global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi, with support for over 100 languages.

The release includes specific safety mechanisms for voice replication, requiring users to provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Additionally, every audio clip generated by these models is watermarked with SynthID, an imperceptible watermark designed to ensure AI-generated speech remains detectable to help prevent misinformation.

来源详情: deepmind.google

为什么这很重要

This launch marks a significant shift in AI voice generation from static presets to dynamic, customizable audio creation. By enabling developers to create entirely new vocal identities or replicate specific voices with consent verification, the models open new possibilities for media localization, conversational agents, and creative content production. The inclusion of SynthID addresses growing concerns about AI-generated audio misinformation, providing a technical safeguard for content transparency.

The introduction of these models provides developers and enterprises with tools to create richer, more expressive audio experiences without relying on fixed voice presets. This capability is particularly relevant for industries requiring nuanced regional accents for media localization or consistent brand voices for conversational agents.

The partnership with companies such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang indicates immediate practical application in accelerating global dubbing and powering conversational voice agents at scale. This suggests a move toward integrating advanced TTS capabilities into mainstream creative and enterprise workflows.

The emphasis on consent verification for voice replication and the use of SynthID addresses critical ethical and security concerns in AI audio generation. These safeguards aim to protect voice talent identity and ensure content transparency, which is increasingly important as AI-generated media becomes more prevalent.

Interactive Mechanism

互动机制:它实际上是如何运作的

以交互方式探索这一发展背后的基础技术。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
交互式概念检查+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

接下来看什么

Monitor the adoption of these models by partner companies like Figma and HeyGen for global dubbing and media localization. Watch for independent evaluations of the voice replication safeguards and the effectiveness of SynthID in detecting AI-generated speech. Additionally, observe how the dual-speaker screenplay editor in Google AI Studio is utilized by developers for complex audio narratives.

Observe how partner companies integrate these models into their products, particularly in the areas of global dubbing and media localization, to assess real-world performance and user reception.

Monitor independent third-party evaluations of the voice replication consent mechanisms and the detectability of SynthID watermarks, as these are critical for ensuring the safety and integrity of the technology.

Track the development of the dual-speaker screenplay editor in Google AI Studio, as this feature represents a new interface for directing AI-generated dialogue that could influence how creators approach audio storytelling.

相关指南和测验

人工智能模型解释AI 伦理人工智能代理测试你所知道的——尝试免费的人工智能测验在我们的词汇表中查找人工智能术语
觉得这有用吗?