Google's Gemini 3.5 Transcribe doesn't just transcribe you, it edits you
Google DeepMind's new speech-to-text model ships today with filler removal, self-correction resolution and 3-speaker attribution built in, a cleaner transcript that is an edited rendering, not a verbatim one. What that means for developers wiring two new APIs into production, and for anyone whose spoken words become an official record.

Google's Gemini 3.5 Transcribe doesn't just transcribe you. It edits you.
Google DeepMind's new speech-to-text model, Gemini 3.5 Transcribe, ships today with editing built in: it removes filler words, handles speakers' mid-sentence self-corrections, and attributes up to three speakers, at an average word error rate of 4.0% for streaming audio and 2.6% for recorded audio, as measured by Artificial Analysis 1. It is rolling out starting today in English for all macOS Gemini app users and in the Rambler dictation feature on Android in select countries and languages, with a developer public preview in the Gemini API via AI Studio and Antigravity
2.
The polish is the point, and it is also the catch. A transcript from Gemini 3.5 Transcribe is a rendering in which the model has already decided what "um" becomes (nothing), what a false start becomes (gone), and what a retracted date becomes (the corrected one). The finished text carries no sign of any of it; removal is the value proposition.
What Gemini 3.5 Transcribe actually does to your words
DeepMind's launch feature list is unusually concrete about the editing layer 1:
- Filler removal: the model "removes filler words", ums and ahs among them, and auto-formats the result.
- Self-correction handling: DeepMind's own worked example is a speaker who proposes meeting Tuesday and then corrects to Wednesday. The model writes Wednesday. The correction itself leaves no visible trace in the text.
- Attribution: recorded audio gets word-level timestamps and speaker labels for up to three speakers; support for more than three is marked experimental.
- Custom vocabulary: supplied jargon and unique spellings are folded into transcripts automatically.
- Function calling: in the Gemini macOS app, the transcription model can hand tasks such as image generation to other Gemini models mid-dictation.
Against Google's previous transcription model, Chirp 3, DeepMind claims new capabilities, improved word error rates, and a 70% improvement in time to final transcription as measured by Artificial Analysis, plus FLEURS results of 5.50% streaming and 5.04% recorded across a set of top languages 1. One number the announcement states but never works out: recorded audio's 2.6% error rate is 35% lower than streaming's 4.0%. For anything record-critical, that gap argues for the recorded path, not the live one.
When a cleaner transcript is a worse record
Three contexts turn the invisible edit from a consumer delight into a liability:
- Journalists quoting sources. A quote assembled from an edited rendering attributes a fluency the speaker never performed, and a source who corrects a shipped date mid-sentence appears in the record as having only ever said the final version. The hesitation was the human signal; deleting it is the model's job description.
- Depositions and courtrooms. Sworn testimony turns on what a witness retracted and when, sometimes more than on the final answer. A model that "seamlessly handles self-corrections"
1 produces precisely the record in which retraction is invisible.
- Compliance and medical notes. DeepMind's stated target workloads include "recorded audio, meetings, call logs" and post-call analytics pipelines
1. An edit applied before storage is an edit applied to the record itself.
What the announced feature list does not settle is whether the polished rendering is the only rendering. That list contains removal, resolution, attribution, and formatting; it contains no edit trail and no verbatim or original-audio mode. Whether the polish can be switched off is the first question a records-dependent buyer should put to Google before this model touches an official document.
The periphery ships while the flagship stalls
Gemini 3.5 Transcribe lands in the middle of a strange release cadence for the 3.5 model family, assembled here from both sources:
- June 2026: Google promised to roll out Gemini 3.5 Pro, a launch The Verge now calls overdue. It has not shipped
2.
- June 2026: Gemini 3.5 Live Translate, another audio model, did ship
3.
- August 2026: Gemini 3.7 Flash, a later-numbered model, gets its own announcement on DeepMind's news page while 3.5 Pro stays absent
3.
- August 26: Gemini 3.5 Transcribe launches. The same day, The Verge updated its story to report that Google had walked back plans, disclosed before publication, for same-day Gemini 3.5 Live and 3.5 Live Experimental updates: they are not launching yet, and no new date was given
2.
Audio and mid-tier models keep shipping on schedule while the 3.5 family's missing flagship drifts without a date. Periphery velocity is real progress; it is just not the model anyone was promised in June.
If your pipeline assumed verbatim, you now ship an editorial product
Developers wiring gemini-3.5-transcribe-live into the Live API or gemini-3.5-transcribe into the Interactions API are not buying faithful capture. They are buying a transcript with a built-in editor, delivered with sub-second latency for streaming use 1. That moves a decision onto the builder: labeling outputs as recorded audio versus rendered text. A live captioning tool can absorb the edit. A call-log archive that feeds audits, disputes, or care records cannot, and Chrome support, which The Verge reports is coming soon, will spread the same question to the browser
2.
The cleanest transcript is not the truest one. When polish is the product, the record needs a label.
References
Cite this story
ProvenBrief (2026). "Google's Gemini 3.5 Transcribe doesn't just transcribe you, it edits you." ProvenBrief. https://provenbrief.com/story/google-s-gemini-3-5-transcribe-doesn-t-just-transcribe-you-it-edits-you
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
WordsProduced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.