Google announced Gemini 3.5 Transcribe on August 26 -- a speech-to-text model it says converts raw audio into "accurate, polished, formatted text" across more than 85 languages, automatically stripping filler words and fixing a speaker's own self-corrections as it goes. It ships as two endpoints: `gemini-3.5-transcribe` for recorded audio (a 2.6% average word error rate, per Google), and a `gemini-3.5-transcribe-live` twin for real-time captioning, which runs at 4.0%. It's already running quietly inside products people use daily -- Gboard's dictation on Android, a "Speak to Window" tool in the Gemini app for Mac -- with Chrome support still to come.
Gemini 3.5 Transcribe, measured
- Announced
- August 26, 2026
- Accuracy, non-streaming
- 2.6% word error rate
- Accuracy, streaming
- 4.0% word error rate
- Independent leaderboard rank
- 5th of the models tracked
- API price
- $5.00 per 1,000 minutes
The headline number is credited to Artificial Analysis, an independent benchmark firm -- not a figure Google measured itself and simply published. That distinction matters more than it usually does here: Artificial Analysis runs a public, continuously updated leaderboard for exactly this category, and Gemini 3.5 Transcribe's score sits on it in the open, not just quoted in Google's own press materials.
On that leaderboard, Gemini 3.5 Transcribe ranks fifth, not first. ElevenLabs' Scribe v2 (2.2% WER), Microsoft's MAI-Transcribe-1.5 (2.4%), and Smallest AI's Pulse Pro (2.4%) all currently post lower error rates, and a fifth model, Fun-Realtime-ASR-preview, tops the board at 1.7%. Google's launch post doesn't mention where the model lands relative to competitors -- unsurprising, since no vendor's announcement leads with its own rank -- but the independent board is exactly where a reader can check the claim against the four models actually beating it.
The API prices at $5.00 per 1,000 minutes of audio on Artificial Analysis' listing, though Google hasn't published a consumer rate card of its own for the feature. For most people the model will show up invisibly either way -- as the engine quietly running behind a keyboard's dictation button or a video call's live captions, not as a product anyone chooses by name.
- Google's Gemini 3.5 Transcribe launched August 26 with a 2.6% word-error rate on recorded audio.
- It already powers Gboard's Android dictation and a Mac Gemini app tool; Chrome support is next.
- The accuracy score comes from Artificial Analysis, an independent benchmark firm with a public leaderboard.
- On that same leaderboard, four other models -- from ElevenLabs, Microsoft, and Smallest AI -- score better.
- Caveat: Google's launch post doesn't mention the 5th-place rank; the number checks out, the framing doesn't.