Subira ku makuru
IbicuruzwaAI Understanding ibisobanuro

Google ya Gemini 3.5 Kwandukura inyandiko zerekana imvugo-ku-buryo hamwe n'imbibi

Iterambere rya Google ibisobanuro birambuye byerekana uburyo Gemini 3.5 Transcript ikora dosiye zamajwi, zirimo gutahura ururimi rwikora, kuranga imvugo, igihe cyerekana ijambo hamwe no guhanagura "ubwenge". Urupapuro rwerekana imipaka ifatika, harimo igihe gito cyamajwi igihe cyo gutandukana cyangwa igihe cyerekana…

6 min readRead the linked source
Source-provided image accompanying Google’s Gemini 3.5 Transcribe docs spell out speech-to-text modes and limits
InkomokoInkomoko yanditse
Umwanditsi
ai.google.dev
Ihuza ry'inkomoko
ai.google.devhttps://ai.google.dev/gemini-api/docs/transcribe
Ubwoko bw'inkomoko
Inkomoko ihujwe - ibanze-isoko yimiterere ntabwo yashizweho.
Byatanzwe kandi

Inkuru iheruka gusubirwamo

ImirongoSobanukirwa ibi mumasegonda 60

Tangira hano

Amagambo y'ingenzi

API (Imigaragarire ya Porogaramu)
Inzira yuburyo bwa sisitemu imwe yohereza ibyifuzo no kwakira ibisubizo bivuye murindi sisitemu.
Ibipimo
Ikizamini gisanzwe cyangwa dataset ikoreshwa mugupima no kugereranya imikorere yicyitegererezo.
Ubukererwe
Igihe kiri hagati yo kohereza icyifuzo no kwakira icyitegererezo cyibisubizo.
IsuzumeModeri ya AI Yasobanuwe Ikibazo

Niki cyahindutse kuva cyatangazwa

  1. Byatangajwe bwa mbere
  2. The Verge materially advances the continuing Gemini 3.5 Transcribe release by reporting specific capabilities and rollout details: automatic filler-word removal, customized vocabulary, attribution for up to three speakers, word-level timestamps, related Gemini 3.5 Live updates, macOS and selected Android availability, and public developer preview access. These details come from The Verge’s report and Google’s claims as quoted there; they are not independently confirmed in the supplied material.
  3. This Ars Technica report materially advances the existing Gemini 3.5 Transcribe update by adding Google’s reported performance figures, support for 85 languages and up to three speakers in pre-recorded audio, custom vocabulary support, and a broader rollout across macOS, Antigravity, AI Studio, the Gemini API, and planned Chrome integration.
  4. Google’s updated primary documentation materially expands the earlier Gemini 3.5 Transcribe announcement with API examples, two transcription modes, support for more than 85 locales, up to 1,000 custom vocabulary phrases, up to eight speaker labels, word-level timestamps and explicit duration and compatibility limits.

Byagenze bite

Google’s Gemini API documentation describes Gemini 3.5 Transcribe, a speech-to-text model for uploaded audio files. The page provides implementation details for verbatim and smart transcription, automatic language detection across more than 85 locales, custom vocabulary, speaker diarization and word-level timestamps.

Google’s page says the Gemini API converts uploaded audio files into text with the Gemini 3.5 Transcribe model, identified as gemini-3.5-transcribe. The documented workflow is to upload an audio file through the Files API, pass its URI to the Interactions API and read the returned transcript from interaction.output_text. The page includes Python, JavaScript and REST examples. For real-time, low- recognition from a microphone or live audio stream, Google directs developers to a separate Live API model called gemini-3.5-transcribe-live.

The model is documented as supporting automatic language recognition across more than 85 locales, including multiple varieties of English, Spanish, Portuguese, Bengali, Punjabi and Chinese. Google says the model can switch languages dynamically when speakers code-switch within or between sentences. Developers can instead supply BCP-47 language codes when the language is known. The page also describes custom vocabulary: up to 1,000 phrases can be supplied to bias recognition toward specialized terms, acronyms, brand names and proper nouns, although Google says the best results typically come from using up to 100 terms.

The page describes two transcription modes. Verbatim mode is the default and is intended to preserve what was said, including filler words, repetitions, pauses and false starts. Smart mode applies post-processing that removes disfluencies, resolves spoken self-corrections, adds punctuation and sentence casing, and formats spoken material into paragraphs, lists, dates, currencies and numbers. Google’s example changes a spoken correction from “Tuesday, actually no, Wednesday” into a cleaned-up Wednesday appointment. Smart mode cannot be combined with speaker diarization or word-level timestamps.

Google also documents speaker diarization and word-level timing as options within verbatim mode. Diarization labels distinct voices with identifiers such as spk_1 and spk_2, supporting up to eight speakers; attribution for three or more speakers is described as experimental. Word-level timestamps return start and end offsets for recognized words, but Google warns that enabling them may reduce overall transcription accuracy. Standard unary requests support audio files up to one hour, while audio processing is limited to 30 minutes when diarization or word-level timestamps are enabled. The page was last updated August 27, 2026, and refers developers elsewhere for pricing, token limits and file-management details.

Ibisobanuro birambuye: ai.google.dev ↗

Impamvu ari ngombwa

The documentation turns the model’s previously announced capabilities into a more actionable specification for developers building transcription workflows. It also makes clear that accuracy and functionality involve tradeoffs: smart transcription cannot provide speaker labels or timestamps, while timestamps may reduce overall transcription accuracy.

The practical significance is that Google is presenting transcription as a configurable model workflow rather than a single undifferentiated speech-recognition endpoint. A developer can choose literal output for records or analysis, or cleaned-up output for reading. Language hints and custom vocabulary offer additional controls for specialized recordings. Those controls could reduce manual cleanup and correction in some workflows, but the source does not provide comparative accuracy measurements or evidence from production deployments.

Speaker labeling and word-level timing are particularly relevant to applications that need to navigate a recording rather than simply produce a block of text. Labels can separate turns in a multi-speaker recording, and offsets can support search or synchronization with audio. However, the labels identify distinct voices rather than named people, and Google explicitly marks attribution for three or more speakers as experimental. A transcript with timestamps may also be less accurate overall according to the documentation’s warning, creating a tradeoff between navigability and fidelity.

The language support described on the page could matter for multilingual meetings, interviews, customer-service recordings and other conversations in which speakers change languages. Automatic detection reduces the need to configure every recording in advance, while explicit language codes may improve results when the language is known. Still, the source establishes Google’s stated support and behavior, not uniform performance across all listed locales, accents, noise conditions or code-switching patterns. Its reference to diverse accents and background noise is a capability claim, not a supplied .

The documented limits also affect system design. Long recordings may require file-upload handling, and adding timestamps or diarization cuts the stated processing ceiling from one hour to 30 minutes. Smart transcription may be easier for readers but is unsuitable when exact wording, speaker attribution or timing is required. The page also separates transcription from broader audio question-answering and from text-to-speech synthesis, so developers cannot assume that this model is a general audio-analysis or voice-generation system. Pricing, token limits, privacy terms and retention practices remain outside the information provided here.

Interactive Mechanism

Uburyo bukoreshwa: Uburyo bukora

Shakisha ikoranabuhanga ryihishe inyuma yiri terambere.

Agent Lifecycle Stage:
1
User Intent & Planning: "Audit customer refund request #4092 and settle payment."
2
Tool Calling: Emits structured JSON call crm_get_transaction(id='4092').
3
Guardrail & Verification:🛡️ Paused: High-value action requires human operator sign-off.
4
Final Settlement: Refund recorded, email receipt dispatched, and audit log stored.
Core takeaway: An AI agent is not just a language model—it is a closed loop of planning, tool invocation, and environment feedback. Production systems require self-healing retries and strict human approval guardrails.
Kugenzura Ibitekerezo Byagenzuwe+10 Points
AI Models Explained Quiz

Which component of an AI application is the machine-learning model itself?

Ibyo kureba

The important next questions are performance in real-world recordings, pricing, token limits, data handling and availability. Google does not provide results, retention terms, service-level commitments or detailed pricing on this page, so the documentation establishes capability and constraints rather than independently verified quality.

The first verification priority is real-world transcription quality. Independent testing would need to examine accents, overlapping speech, background noise, low-quality recordings, multilingual code-switching and specialized terminology across the listed locales. It should compare verbatim and smart outputs, especially where smart mode removes filler words or resolves corrections that may be important to the original meaning. The documentation gives no word-error rates, language-by-language results or examples from uncontrolled recordings.

Speaker attribution deserves separate scrutiny. Google supports up to eight speakers but calls attribution for three or more experimental. Tests should measure how often the system splits one speaker into multiple labels, merges different speakers or assigns words to the wrong person. Similar testing is needed for word timestamps, because the page warns that timestamps may degrade overall transcription accuracy. These tradeoffs are consequential for interview transcripts, accessibility tools, searchable archives and any workflow that treats the output as a record.

Developers and organizations should also watch the operational terms that this page does not specify. Google points readers to a pricing page for model pricing and token limits, but those figures are not included in the source. The page likewise does not state regional availability, service-level commitments, supported audio formats in a comprehensive list, retention periods, deletion controls or whether uploaded recordings are used for other purposes. Those unknowns could materially affect adoption, especially for sensitive recordings or high-volume services.

Finally, the relationship between this model and Google’s other audio tools will matter. The documentation directs real-time users to a separate Live transcription path, while broader audio analysis belongs to Audio understanding and voice generation belongs to Text-to-speech. That division may give developers clearer interfaces, but it can also require separate integrations for live capture, transcription, analysis and synthesis. Further release notes, pricing details, independent evaluations and evidence of use in consequential settings would show whether Gemini 3.5 Transcribe is more than a well-specified API capability.

Ibijyanye nuyobora & ibibazo

Moderi ya AI YasobanuweImyitwarire ya AIAmahugurwa ya AIGerageza ibyo uzi - gerageza ikibazo cya AI kubuntuReba ijambo AI mumagambo yacuKurikiza icyerekezo cya AI cyo kurekura

Kuvugurura no gukosora

Iyi nkuru yemewe ivugururwa mugihe iyo iterambere ryiterambere rihindutse mubintu. URL yayo nitariki yo gusohora itariki ntizigera ihinduka.

  • Google’s updated primary documentation materially expands the earlier Gemini 3.5 Transcribe announcement with API examples, two transcription modes, support for more than 85 locales, up to 1,000 custom vocabulary phrases, up to eight speaker labels, word-level timestamps and explicit duration and compatibility limits.
  • This Ars Technica report materially advances the existing Gemini 3.5 Transcribe update by adding Google’s reported performance figures, support for 85 languages and up to three speakers in pre-recorded audio, custom vocabulary support, and a broader rollout across macOS, Antigravity, AI Studio, the Gemini API, and planned Chrome integration.
  • The Verge materially advances the continuing Gemini 3.5 Transcribe release by reporting specific capabilities and rollout details: automatic filler-word removal, customized vocabulary, attribution for up to three speakers, word-level timestamps, related Gemini 3.5 Live updates, macOS and selected Android availability, and public developer preview access. These details come from The Verge’s report and Google’s claims as quoted there; they are not independently confirmed in the supplied material.
Reba igitabo gikosora rusange
Basanze ari ingirakamaro?