Voltar às notícias
ProdutoInstruções AI Understanding

Google’s Gemini 3.5 Transcribe expands cleaned-up voice input beyond Gboard

Ars Technica reports that Google is extending Gemini 3.5 Transcribe, a speech-to-text model that removes verbal filler and handles corrections, from Pixel 11’s Gboard feature to selected Google products and developer tools. Google claims faster transcription and a lower live-speech error rate, but those figures…

Por 6 min read
The supplied source does not describe the accompanying Google-provided image.
A versão curta

Ars Technica reports that Google is extending Gemini 3.5 Transcribe, a speech-to-text model that removes verbal filler and handles corrections, from Pixel 11’s Gboard feature to selected Google products and developer tools. Google claims faster transcription and a lower live-speech error rate, but those figures…

Official primary-source video from arstechnica.com · shown with attribution.

O que aconteceu

Ars Technica reports that Google announced Gemini 3.5 Transcribe, an AI speech-to-text model designed to turn spoken input into polished text by removing filler words and handling self-corrections. The model already powers Gboard’s Rambler feature on Pixel 11 phones and is expanding to the Gemini app on macOS, Antigravity, AI Studio’s build model, and the Gemini API. Google also plans to bring it to Chrome soon.

Ars Technica reports that Google announced Gemini 3.5 Transcribe as a model for voice input and speech-to-text. Its main reported distinction from a conventional transcription system is that it edits the spoken stream: it can remove “ums” and “uhs,” account for corrections made while speaking, and produce text that is intended to reflect what the speaker meant. The report says the model already powers Google’s Gboard “Rambler” feature on Pixel 11 phones. Google’s announcement is the basis for the product and capability claims described by Ars Technica; the supplied report does not include an independent technical evaluation of the model.

Ars Technica says Google claims Gemini 3.5 Transcribe is about 70 percent faster than the company’s previous voice-to-text engine, Chirp 3, when measured from speech to final transcribed text. The report also gives Google’s stated live-speech error rates: 5.5 percent for Gemini 3.5 Transcribe compared with 7.32 percent for Chirp 3. Those figures are presented as Google’s measurements, not as results independently reproduced by Ars Technica or another evaluator. The report does not specify the test corpus, language distribution, hardware, definition of “final” text, or the conditions used to calculate the rates.

The model’s reported language and speaker support is broader than a single-device voice-input feature. According to Ars Technica, Gemini 3.5 Transcribe works in 85 languages and supports up to three speakers in pre-recorded audio. The report says users can provide custom vocabulary so the system can better handle specialized jargon. It does not provide language-by-language accuracy results, explain how speaker labels are generated, or establish whether all listed languages receive equivalent support.

The rollout described by Ars Technica is staged. Gemini 3.5 Transcribe is already used by Rambler in Gboard, although that feature remains limited to Pixel 11 phones. Google says Rambler will expand to more “Gemini Intelligence devices” later in the year. The Gemini app on macOS is reported to receive the voice-input benefit immediately, while Antigravity and AI Studio’s build model also gain access. Developers can use the model through the Gemini API. Google plans to bring it to Chrome soon, where it would support text entry in web fields, but the report gives no specific Chrome release date.

Leia a fonte primária: arstechnica.com

Por que isso importa

The reported change could make voice input more useful for writing, coding, and interacting with web applications, because the system is designed to clean up speech rather than reproduce every hesitation verbatim. It also illustrates a shift from basic transcription toward interpretation and editing. The tradeoff is that altering wording can introduce errors or change meaning, especially where an exact record of speech is important.

The practical significance is that voice input could become an editing layer rather than a literal transcript. A person who pauses, repeats a phrase, or corrects themselves would not necessarily need to revise every hesitation afterward. Ars Technica reports that its own testing with Rambler found the model effective at cleaning up inconsistencies and verbal stumbles in short blocks of text. That observation is limited to the reporter’s testing and does not establish performance across longer dictation, accents, noisy environments, technical vocabulary, or all 85 languages.

The product could matter to users who write or interact with software through speech. The reported expansion reaches a general-purpose macOS app, developer tools, an API, and eventually the Chrome browser. If the planned browser integration works as described, users could dictate into ordinary web fields, including email and AI-chat interfaces, without relying on a separate transcription step. For developers, API access may make the model available as a component in applications that accept spoken input. The source does not state pricing, rate limits, service-level commitments, regional availability, or whether API access is generally available.

The model also highlights an important distinction between transcription and interpretation. Removing filler words and correcting spoken text can make output easier to read, but it means the result is not necessarily a verbatim record. Ars Technica explicitly notes that the model technically changes the wording of what a person said and that this may be inappropriate in some situations. That limitation matters for interviews, legal or medical records, customer disputes, accessibility documentation, and any workflow where the original phrasing carries evidence or meaning.

The broader product implication is not that speech recognition has become error-free, but that Google is applying a generative model to a familiar input task. The reported capabilities combine recognition, contextual editing, custom vocabulary, and multi-speaker handling. Those functions may reduce friction for everyday dictation, yet they also shift responsibility to users to review the output. The supplied report offers no independent comparison with competing transcription systems and no evidence about how the model handles ambiguous wording, names, code, sensitive content, or dialectal variation.

O que assistir a seguir

The key questions are how broadly Google makes Gemini 3.5 Transcribe available, whether its reported accuracy and speed gains hold across languages and speakers, and how often cleanup changes users’ intended meaning. Availability remains limited to selected products and devices, while Chrome access is only described as coming soon. The report does not independently verify Google’s performance figures.

First, availability will determine whether this is a meaningful ecosystem change or a narrowly limited feature. Ars Technica reports current access through Pixel 11’s Rambler feature, the Gemini app on macOS, Antigravity, AI Studio’s build model, and the Gemini API. Google says support will expand to more Gemini Intelligence devices and Chrome, but the source provides no firm dates for those expansions. It also does not say whether access will vary by country, account type, subscription, or developer tier.

Second, independent testing should examine the performance claims in context. Google’s reported 70 percent speed improvement and the reduction from a 7.32 percent to a 5.5 percent live-speech error rate could be useful, but the source does not identify the test design or explain how errors were weighted. Future evaluations should compare literal transcription and cleaned-up output separately, because a system can appear more readable while still changing meaning. Tests should also cover long recordings, overlapping speakers, background noise, accents, specialized terminology, and less commonly used languages.

Third, users and organizations will need clearer controls over when editing is allowed. The report says the system can remove disfluencies and revise corrections on the fly, but it does not describe a setting that preserves an original transcript alongside the polished version. That distinction could be important wherever records must be auditable. Google’s custom-vocabulary support may help with jargon, but the source does not explain how vocabulary is stored, how personal or proprietary terms are handled, or whether audio and transcripts are retained.

Finally, the rollout should be watched for evidence about reliability and user review. Ars Technica says the model was useful in the reporter’s short-form testing, while also warning that its changes may not suit every situation. The report does not independently confirm Google’s claims, does not provide systematic test results, and does not describe user reactions beyond the article’s testing. Until those unknowns are addressed, Gemini 3.5 Transcribe is best understood as a potentially useful voice-input product change with meaningful limits, not as a replacement for careful review of important speech records.

Guias e questionários relacionados

ChatGPT e LLMModelos de IA explicadosPrompt EngineeringTeste o que você sabe – experimente um teste gratuito de IAProcure um termo de IA em nosso glossário
Achou isso útil?