Chuyện gì đã xảy ra
Google DeepMind announced the release of two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models are available immediately in Google AI Studio and through the Gemini API, with integrations for platforms like Agora, LiveKit, and Vercel. The release expands the Gemini Audio family, adding capabilities for generating custom character voices and directing scene dialogue with precise control over delivery.
Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as the most expressive audio generation models in the Gemini family. The announcement states that these models transform voice generation from static presets into a dynamic creative studio, allowing for the creation of custom character voices and the direction of scene dialogue.
The models are available starting today in Google AI Studio, where they function as a voice design workspace. Users can prompt new vocal identities from scratch or replicate their own voices, then use a dual-speaker screenplay editor to direct line-by-line delivery. Access is also provided via the Gemini API, enabling developer platforms such as Agora, LiveKit, Pipecat, and Vercel to build and deploy speech generation experiences.
Google claims that Gemini 3.8 Flash TTS secures the #1 overall spot on Hume AI’s Voice Design with a score of 71.4 and leads in accent modeling with a score of 60.8. Both models reportedly secure the #1 and #2 spots on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, the models are said to hold top positions in key global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi, with support for over 100 languages.
The release includes specific safety mechanisms for voice replication, requiring users to provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Additionally, every audio clip generated by these models is watermarked with SynthID, an imperceptible watermark designed to ensure AI-generated speech remains detectable to help prevent misinformation.
Chi tiết nguồn: deepmind.google ↗
Tại sao nó quan trọng
This launch marks a significant shift in AI voice generation from static presets to dynamic, customizable audio creation. By enabling developers to create entirely new vocal identities or replicate specific voices with consent verification, the models open new possibilities for media localization, conversational agents, and creative content production. The inclusion of SynthID addresses growing concerns about AI-generated audio misinformation, providing a technical safeguard for content transparency.
The introduction of these models provides developers and enterprises with tools to create richer, more expressive audio experiences without relying on fixed voice presets. This capability is particularly relevant for industries requiring nuanced regional accents for media localization or consistent brand voices for conversational agents.
The partnership with companies such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang indicates immediate practical application in accelerating global dubbing and powering conversational voice agents at scale. This suggests a move toward integrating advanced TTS capabilities into mainstream creative and enterprise workflows.
The emphasis on consent verification for voice replication and the use of SynthID addresses critical ethical and security concerns in AI audio generation. These safeguards aim to protect voice talent identity and ensure content transparency, which is increasingly important as AI-generated media becomes more prevalent.
Cơ chế tương tác: Nó thực sự hoạt động như thế nào
Khám phá công nghệ cơ bản đằng sau sự phát triển này một cách tương tác.
crm_get_transaction(id='4092').In AI, what are a model's "parameters"?
Xem gì tiếp theo
Monitor the adoption of these models by partner companies like Figma and HeyGen for global dubbing and media localization. Watch for independent evaluations of the voice replication safeguards and the effectiveness of SynthID in detecting AI-generated speech. Additionally, observe how the dual-speaker screenplay editor in Google AI Studio is utilized by developers for complex audio narratives.
Observe how partner companies integrate these models into their products, particularly in the areas of global dubbing and media localization, to assess real-world performance and user reception.
Monitor independent third-party evaluations of the voice replication consent mechanisms and the detectability of SynthID watermarks, as these are critical for ensuring the safety and integrity of the technology.
Track the development of the dual-speaker screenplay editor in Google AI Studio, as this feature represents a new interface for directing AI-generated dialogue that could influence how creators approach audio storytelling.