6 days ago
Google Launches Gemini 3.5 Transcribe for Faster, Smarter Speech
Google has made a new computer model called Gemini 3.5 Transcribe.
It listens to speech and turns it into written words.
Google says it can handle more than 85 languages, accents, dialects and some background challenges.
It can also notice when different people are speaking in recorded audio.
For now, it can identify up to three speakers reliably.
Developers can use it to build voice assistants and other audio tools.
Google is adding related features to products such as Gboard, Chrome and its Gemini macOS app.
The model is currently available to some users as a public preview.
Gemini 3.5 Transcribe is designed to improve speech-to-text accuracy for voice agents and audio applications.
The model supports more than 85 languages, including various accents and dialects.
Recorded audio can include up to three identified speakers and word-level timestamps; support for more speakers is experimental.
Google reports 4% average streaming Word Error Rate and 2.6% for non-streaming use, citing Artificial Analysis measurements.
The model is available in public preview through Google AI Studio and Google Antigravity, with integrations planned across Google products.
- Who
- Google launched Gemini 3.5 Transcribe for developers and users of selected Google products.
- What
- A speech-to-text model with multilingual support, speaker attribution, timestamps and APIs for streaming and recorded audio.
- Where
- It is available through Google AI Studio and Google Antigravity, with access also offered through Google's enterprise platform and selected products.
- When
- Why
- To make transcriptions more accurate and useful for voice agents, audio applications and spoken workflows.
Key facts
- Language support
- More than 85 languages, including accents and dialects.
- Speaker identification
- Up to three speakers in recorded audio; support for more than three is experimental.
- Streaming Word Error Rate
- Google cites a 4% average measurement from Artificial Analysis.
- Non-streaming Word Error Rate
- Google cites a 2.6% average measurement from Artificial Analysis.
- FLEURS benchmark
- The model recorded 5.50% WER in streaming and 5.04% in non-streaming tests.
- Speed
- Google says time to final transcription is 70% faster than Chirp 3.
- Developer access
- The Live API supports continuous bidirectional streaming, while the Interactions API supports recorded audio, speaker attribution and timestamps.






