返回新聞
產品展示AI Understanding 簡報

Google 的 Gemini 3.5 轉錄文件闡明了語音轉文字模式和限制

Google 的开发者文档详细介绍了 Gemini 3.5 Transcribe 如何处理音频文件,包括自动语言检测、说话人标签、单词时间戳和清理的“智能”转录。該頁面還記錄了實際限制,包括當分類或時間戳......時較短的最大音訊持續時間

6 min readRead the linked source
Source-provided image accompanying Google’s Gemini 3.5 Transcribe docs spell out speech-to-text modes and limits
來源參考來源記錄
出版商
ai.google.dev
來源連結
ai.google.devhttps://ai.google.dev/gemini-api/docs/transcribe
來源類型
連結來源-主要來源狀態尚未確定。
還引用了

故事最後修訂

背景60 秒內了解這一點

從這裡開始

關鍵術語

API(應用程式介面)
一種軟體系統向另一個系統發送請求並接收回應的結構化方式。
基準測試
用於測量和比較模型性能的標準化測試或資料集。
延遲
發送請求和接收模型輸出之間的時間。
測試一下自己AI 模型解釋測驗

自發布以來發生了什麼變化

  1. 首次發表
  2. The Verge materially advances the continuing Gemini 3.5 Transcribe release by reporting specific capabilities and rollout details: automatic filler-word removal, customized vocabulary, attribution for up to three speakers, word-level timestamps, related Gemini 3.5 Live updates, macOS and selected Android availability, and public developer preview access. These details come from The Verge’s report and Google’s claims as quoted there; they are not independently confirmed in the supplied material.
  3. This Ars Technica report materially advances the existing Gemini 3.5 Transcribe update by adding Google’s reported performance figures, support for 85 languages and up to three speakers in pre-recorded audio, custom vocabulary support, and a broader rollout across macOS, Antigravity, AI Studio, the Gemini API, and planned Chrome integration.
  4. Google’s updated primary documentation materially expands the earlier Gemini 3.5 Transcribe announcement with API examples, two transcription modes, support for more than 85 locales, up to 1,000 custom vocabulary phrases, up to eight speaker labels, word-level timestamps and explicit duration and compatibility limits.

發生了什麼事

Google’s Gemini API documentation describes Gemini 3.5 Transcribe, a speech-to-text model for uploaded audio files. The page provides implementation details for verbatim and smart transcription, automatic language detection across more than 85 locales, custom vocabulary, speaker diarization and word-level timestamps.

Google’s page says the Gemini API converts uploaded audio files into text with the Gemini 3.5 Transcribe model, identified as gemini-3.5-transcribe. The documented workflow is to upload an audio file through the Files API, pass its URI to the Interactions API and read the returned transcript from interaction.output_text. The page includes Python, JavaScript and REST examples. For real-time, low- recognition from a microphone or live audio stream, Google directs developers to a separate Live API model called gemini-3.5-transcribe-live.

The model is documented as supporting automatic language recognition across more than 85 locales, including multiple varieties of English, Spanish, Portuguese, Bengali, Punjabi and Chinese. Google says the model can switch languages dynamically when speakers code-switch within or between sentences. Developers can instead supply BCP-47 language codes when the language is known. The page also describes custom vocabulary: up to 1,000 phrases can be supplied to bias recognition toward specialized terms, acronyms, brand names and proper nouns, although Google says the best results typically come from using up to 100 terms.

The page describes two transcription modes. Verbatim mode is the default and is intended to preserve what was said, including filler words, repetitions, pauses and false starts. Smart mode applies post-processing that removes disfluencies, resolves spoken self-corrections, adds punctuation and sentence casing, and formats spoken material into paragraphs, lists, dates, currencies and numbers. Google’s example changes a spoken correction from “Tuesday, actually no, Wednesday” into a cleaned-up Wednesday appointment. Smart mode cannot be combined with speaker diarization or word-level timestamps.

Google also documents speaker diarization and word-level timing as options within verbatim mode. Diarization labels distinct voices with identifiers such as spk_1 and spk_2, supporting up to eight speakers; attribution for three or more speakers is described as experimental. Word-level timestamps return start and end offsets for recognized words, but Google warns that enabling them may reduce overall transcription accuracy. Standard unary requests support audio files up to one hour, while audio processing is limited to 30 minutes when diarization or word-level timestamps are enabled. The page was last updated August 27, 2026, and refers developers elsewhere for pricing, token limits and file-management details.

來源詳情: ai.google.dev ↗

為什麼這很重要

The documentation turns the model’s previously announced capabilities into a more actionable specification for developers building transcription workflows. It also makes clear that accuracy and functionality involve tradeoffs: smart transcription cannot provide speaker labels or timestamps, while timestamps may reduce overall transcription accuracy.

The practical significance is that Google is presenting transcription as a configurable model workflow rather than a single undifferentiated speech-recognition endpoint. A developer can choose literal output for records or analysis, or cleaned-up output for reading. Language hints and custom vocabulary offer additional controls for specialized recordings. Those controls could reduce manual cleanup and correction in some workflows, but the source does not provide comparative accuracy measurements or evidence from production deployments.

Speaker labeling and word-level timing are particularly relevant to applications that need to navigate a recording rather than simply produce a block of text. Labels can separate turns in a multi-speaker recording, and offsets can support search or synchronization with audio. However, the labels identify distinct voices rather than named people, and Google explicitly marks attribution for three or more speakers as experimental. A transcript with timestamps may also be less accurate overall according to the documentation’s warning, creating a tradeoff between navigability and fidelity.

The language support described on the page could matter for multilingual meetings, interviews, customer-service recordings and other conversations in which speakers change languages. Automatic detection reduces the need to configure every recording in advance, while explicit language codes may improve results when the language is known. Still, the source establishes Google’s stated support and behavior, not uniform performance across all listed locales, accents, noise conditions or code-switching patterns. Its reference to diverse accents and background noise is a capability claim, not a supplied .

The documented limits also affect system design. Long recordings may require file-upload handling, and adding timestamps or diarization cuts the stated processing ceiling from one hour to 30 minutes. Smart transcription may be easier for readers but is unsuitable when exact wording, speaker attribution or timing is required. The page also separates transcription from broader audio question-answering and from text-to-speech synthesis, so developers cannot assume that this model is a general audio-analysis or voice-generation system. Pricing, token limits, privacy terms and retention practices remain outside the information provided here.

Interactive Mechanism

互動機制:它實際上是如何運作的

以互動方式探索這項發展背後的基礎技術。

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
互動式概念檢查+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

接下來看什麼

The important next questions are performance in real-world recordings, pricing, token limits, data handling and availability. Google does not provide results, retention terms, service-level commitments or detailed pricing on this page, so the documentation establishes capability and constraints rather than independently verified quality.

The first verification priority is real-world transcription quality. Independent testing would need to examine accents, overlapping speech, background noise, low-quality recordings, multilingual code-switching and specialized terminology across the listed locales. It should compare verbatim and smart outputs, especially where smart mode removes filler words or resolves corrections that may be important to the original meaning. The documentation gives no word-error rates, language-by-language results or examples from uncontrolled recordings.

Speaker attribution deserves separate scrutiny. Google supports up to eight speakers but calls attribution for three or more experimental. Tests should measure how often the system splits one speaker into multiple labels, merges different speakers or assigns words to the wrong person. Similar testing is needed for word timestamps, because the page warns that timestamps may degrade overall transcription accuracy. These tradeoffs are consequential for interview transcripts, accessibility tools, searchable archives and any workflow that treats the output as a record.

Developers and organizations should also watch the operational terms that this page does not specify. Google points readers to a pricing page for model pricing and token limits, but those figures are not included in the source. The page likewise does not state regional availability, service-level commitments, supported audio formats in a comprehensive list, retention periods, deletion controls or whether uploaded recordings are used for other purposes. Those unknowns could materially affect adoption, especially for sensitive recordings or high-volume services.

Finally, the relationship between this model and Google’s other audio tools will matter. The documentation directs real-time users to a separate Live transcription path, while broader audio analysis belongs to Audio understanding and voice generation belongs to Text-to-speech. That division may give developers clearer interfaces, but it can also require separate integrations for live capture, transcription, analysis and synthesis. Further release notes, pricing details, independent evaluations and evidence of use in consequential settings would show whether Gemini 3.5 Transcribe is more than a well-specified API capability.

相關指引和測驗

人工智慧模型解釋AI 倫理人工智慧培訓測試你所知道的—嘗試免費的人工智慧測驗在我們的詞彙表中尋找人工智慧術語關注 AI 模型發布追蹤器

更新和更正

當正在發生的事件發生重大變化時,這個典型的故事就會被更新。它的 URL 和原始發布日期永遠不會改變。

  • Google’s updated primary documentation materially expands the earlier Gemini 3.5 Transcribe announcement with API examples, two transcription modes, support for more than 85 locales, up to 1,000 custom vocabulary phrases, up to eight speaker labels, word-level timestamps and explicit duration and compatibility limits.
  • This Ars Technica report materially advances the existing Gemini 3.5 Transcribe update by adding Google’s reported performance figures, support for 85 languages and up to three speakers in pre-recorded audio, custom vocabulary support, and a broader rollout across macOS, Antigravity, AI Studio, the Gemini API, and planned Chrome integration.
  • The Verge materially advances the continuing Gemini 3.5 Transcribe release by reporting specific capabilities and rollout details: automatic filler-word removal, customized vocabulary, attribution for up to three speakers, word-level timestamps, related Gemini 3.5 Live updates, macOS and selected Android availability, and public developer preview access. These details come from The Verge’s report and Google’s claims as quoted there; they are not independently confirmed in the supplied material.
查看公開更正日誌
覺得有用嗎?