Intelligent transcription with Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is Google DeepMind's latest speech-to-text model, designed for intelligent voice interactions. Unlike conventional models that struggle with background noise, jargon, and disfluency cleanup, it converts raw audio directly into accurate, polished, formatted text. The model is already powering consumer voice features like Rambler on Android and in the Gemini app on macOS, and is now available to developers via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Developers can integrate it into voice agents, real-time captioning, or post-call analytics. It is offered through two APIs: the Live API (using gemini-3.5-transcribe-live) for real-time, bidirectional streaming with sub-second latency, and the Interactions API (using gemini-3.5-transcribe) for pre-recorded audio with speaker attribution and word-level timestamps.
Key features include smart transcription that handles self-corrections (e.g., 'let’s meet Tuesday—no, Wednesday'), removes filler words, and auto-formats text; function calling to delegate tasks like image generation and file analysis to other Gemini models; and custom vocabulary for specialized jargon. Performance benchmarks from Artificial Analysis show an average Word Error Rate of 4.0% for streaming and 2.6% for non-streaming, with strong results in noisy environments and accurate capture of alphanumeric entities like postal codes.
The model automatically detects and transcribes over 85 languages across regional accents and dialects, and supports multi-speaker identification for up to three speakers in pre-recorded audio, with timestamps. This release expands Gemini's voice capabilities for both consumers and developers, enabling more precise and context-aware transcription in real-time applications.