Google's Gemini 3.5 Transcribe turns speech into clean text
Your audio just got a translator and a typist. Google's Gemini 3.5 Transcribe handles live speech and recordings in 85+ languages.
Voice notes are getting smarter. Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to turn spoken audio into accurate, polished and well-formatted text.
Announced on 26 August 2026, Google says it is its most precise transcription system yet, built for both live conversations and recorded audio such as meetings and calls.
It does more than just hear words
Speech-to-text tools can struggle with background noise, technical terms, pauses and filler words such as “um” and “ah”. Gemini 3.5 Transcribe is designed to handle these situations more naturally.
Rather than simply converting every sound into text, it can clean up the transcript and make it easier to read.
For instance, if someone changes their mind halfway through a sentence, the model can tidy up the correction instead of leaving a messy transcript behind.
It can also recognise custom vocabulary and handle details such as order IDs, postal codes and file names.
One model for live audio, another for recordings
Google is offering two versions for developers. The gemini-3.5-transcribe-live model supports continuous streaming with sub-second latency, making it suitable for voice agents, real-time captions and interactive voice applications.
Gemini-3.5-transcribe AI is designed for recorded audio. It can process meetings, calls and other recordings while adding speaker attribution and word-level timestamps. Google says it currently supports up to three speakers, with support for more still experimental.
The accuracy numbers look promising
Google says Gemini 3.5 Transcribe has an average word error rate of 4.0% for streaming audio and 2.6% for non-streaming audio, based on measurements cited from Artificial Analysis. Word error rate measures how often a transcription gets words wrong, so a lower number is better.
Google also says the model improves time to final transcription by 70% compared with Chirp 3, its previous transcription model. The model can automatically detect and transcribe more than 85 languages, including regional accents and dialects.
Voice typing is coming to more Google products
Gemini 3.5 Transcribe is currently available in public preview for developers through the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
It is also appearing in products including the Gemini app on macOS in English, Rambler on Android in selected countries and languages, and Google Antigravity.
Google says the technology is also coming to Chrome, where users will be able to use their voice to type into web fields. If it delivers on Google's accuracy, voice input could start feeling less like dictation and more like having someone clean up your words as you speak.


