뉴스로 돌아가기
제품AI Understanding 브리핑

Google launches Gemini 3.8 Flash TTS models

Google DeepMind has released Gemini 3.8 Flash TTS and Flash-Lite TTS, new text-to-speech models available in Google AI Studio and via the Gemini API, featuring custom voice creation and SynthID watermarking.

4 min readRead the primary source
Source-provided image accompanying Google launches Gemini 3.8 Flash TTS models
기본 소스 문서녹음된 소스
출판사
deepmind.google
소스 링크
deepmind.googlehttps://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/
소스 유형
기본 문서 — 우리가 직접 읽는 공식 발표, 논문, 서류 또는 자사 페이지입니다.
맥락60초 안에 이해하세요

여기서 시작하세요

주요 용어

API(애플리케이션 프로그래밍 인터페이스)
한 소프트웨어 시스템이 다른 시스템에 요청을 보내고 응답을 받는 구조화된 방식입니다.
워터마킹
AI가 생성한 텍스트나 미디어에 감지 가능한 신호를 삽입하여 나중에 기계가 생성한 것으로 식별할 수 있습니다.
벤치마크
모델 성능을 측정하고 비교하는 데 사용되는 표준화된 테스트 또는 데이터 세트입니다.
자신을 테스트해 보세요AI 모델 설명 퀴즈

무슨 일이 일어났나요?

Google DeepMind announced the release of two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. These models are available immediately in Google AI Studio and through the Gemini API, with integrations for platforms like Agora, LiveKit, and Vercel. The release expands the Gemini Audio family, adding capabilities for generating custom character voices and directing scene dialogue with precise control over delivery.

Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, describing them as the most expressive audio generation models in the Gemini family. The announcement states that these models transform voice generation from static presets into a dynamic creative studio, allowing for the creation of custom character voices and the direction of scene dialogue.

The models are available starting today in Google AI Studio, where they function as a voice design workspace. Users can prompt new vocal identities from scratch or replicate their own voices, then use a dual-speaker screenplay editor to direct line-by-line delivery. Access is also provided via the Gemini API, enabling developer platforms such as Agora, LiveKit, Pipecat, and Vercel to build and deploy speech generation experiences.

Google claims that Gemini 3.8 Flash TTS secures the #1 overall spot on Hume AI’s Voice Design with a score of 71.4 and leads in accent modeling with a score of 60.8. Both models reportedly secure the #1 and #2 spots on Hume AI’s Overall Quality Index. In blind human preference evaluations on Voice Arena, the models are said to hold top positions in key global languages including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish, and Hindi, with support for over 100 languages.

The release includes specific safety mechanisms for voice replication, requiring users to provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Additionally, every audio clip generated by these models is watermarked with SynthID, an imperceptible watermark designed to ensure AI-generated speech remains detectable to help prevent misinformation.

소스 세부정보: deepmind.google

왜 중요한가요?

This launch marks a significant shift in AI voice generation from static presets to dynamic, customizable audio creation. By enabling developers to create entirely new vocal identities or replicate specific voices with consent verification, the models open new possibilities for media localization, conversational agents, and creative content production. The inclusion of SynthID addresses growing concerns about AI-generated audio misinformation, providing a technical safeguard for content transparency.

The introduction of these models provides developers and enterprises with tools to create richer, more expressive audio experiences without relying on fixed voice presets. This capability is particularly relevant for industries requiring nuanced regional accents for media localization or consistent brand voices for conversational agents.

The partnership with companies such as Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang indicates immediate practical application in accelerating global dubbing and powering conversational voice agents at scale. This suggests a move toward integrating advanced TTS capabilities into mainstream creative and enterprise workflows.

The emphasis on consent verification for voice replication and the use of SynthID addresses critical ethical and security concerns in AI audio generation. These safeguards aim to protect voice talent identity and ensure content transparency, which is increasingly important as AI-generated media becomes more prevalent.

Interactive Mechanism

대화형 메커니즘: 실제로 작동하는 방식

이 개발의 이면에 있는 기본 기술을 대화식으로 살펴보세요.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
대화형 개념 확인+10 Points
AI Models Explained Quiz

In AI, what are a model's "parameters"?

다음에 무엇을 볼 것인가

Monitor the adoption of these models by partner companies like Figma and HeyGen for global dubbing and media localization. Watch for independent evaluations of the voice replication safeguards and the effectiveness of SynthID in detecting AI-generated speech. Additionally, observe how the dual-speaker screenplay editor in Google AI Studio is utilized by developers for complex audio narratives.

Observe how partner companies integrate these models into their products, particularly in the areas of global dubbing and media localization, to assess real-world performance and user reception.

Monitor independent third-party evaluations of the voice replication consent mechanisms and the detectability of SynthID watermarks, as these are critical for ensuring the safety and integrity of the technology.

Track the development of the dual-speaker screenplay editor in Google AI Studio, as this feature represents a new interface for directing AI-generated dialogue that could influence how creators approach audio storytelling.

관련 가이드 및 퀴즈

AI 모델 설명AI 윤리AI 에이전트알고 있는 내용을 테스트해 보세요. 무료 AI 퀴즈를 시도해 보세요.용어집에서 AI 용어를 찾아보세요.
이것이 유용하다고 생각하시나요?