Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to deliver more accurate, context-aware and polished transcriptions from natural speech. The company describes it as its most precise transcription model yet, with capabilities aimed at voice agents, real-time captioning, meeting transcription and other voice-driven applications.
Unlike conventional speech-recognition systems that can struggle with background noise, specialized terminology and speech disfluencies, Gemini 3.5 Transcribe is designed to process raw audio into clean, formatted text. It can handle self-corrections, such as changing a spoken date or word mid-sentence, while automatically removing filler words and improving formatting.
Google is making the model available to developers through two APIs. The Live API, using gemini-3.5-transcribe-live, provides continuous bidirectional streaming with sub-second latency for interactive voice applications. The Interactions API, using gemini-3.5-transcribe, supports pre-recorded audio such as meetings and call recordings, including speaker attribution and word-level timestamps.
According to measurements cited by Google from Artificial Analysis, Gemini 3.5 Transcribe achieves an average Word Error Rate (WER) of 4.0% for streaming transcription and 2.6% for non-streaming use cases. The model can also recognize customized vocabulary, helping applications handle specialized terminology, unique spellings, order numbers and other alphanumeric information.
The system supports automatic language detection and transcription across more than 85 languages, including regional accents and different dialects. Pre-recorded audio can identify up to three speakers, while support for more than three speakers remains experimental.
Google says the model also significantly improves on its previous Chirp 3 transcription technology. On selected benchmarks, the company reports improved accuracy and a 70% reduction in time to final transcription.
Beyond developer APIs, Gemini 3.5 Transcribe is being integrated into several Google experiences. On Android, the new Rambler feature on Gboard can turn spoken thoughts into formatted text while removing filler words and allowing voice-based corrections. The technology is also being used in Google Antigravity, Google AI Studio and the Gemini app on macOS.
Google said Gemini 3.5 Transcribe will eventually arrive in Chrome, enabling users to dictate text directly into web fields. Developers can currently access the model through Google AI Studio and Google Antigravity, while enterprise availability is offered through the Gemini Enterprise Agent Platform.

