Model Releases Hacker News (Gemini)

Gemini-3.5-Transcribe

Gemini 3.5 Transcribespeech-to-textGoogle AIvoice AI

Google has announced Gemini 3.5 Transcribe, its most precise speech-to-text model yet, designed for intelligent voice interactions. It converts raw audio into accurate, polished, formatted text while handling background noise, complex jargon, and disfluency cleanup. The model already powers voice capabilities like Rambler on Android and in the Gemini app on macOS, and is now available to developers via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.

Developers can use two separate APIs: the Live API (gemini-3.5-transcribe-live) for real-time bidirectional streaming with sub-second latency, and the Interactions API (gemini-3.5-transcribe) for pre-recorded audio with speaker attribution and word-level timestamps. Key features include smart transcription that handles self-corrections (e.g., "Tuesday—no, Wednesday") and removes filler words, function calling to delegate tasks like image generation and file analysis to other Gemini models, custom vocabulary support for specialized jargon, automatic detection of over 85 languages with regional accents and dialects, and multi-speaker identification for up to three speakers in pre-recorded audio.

As measured by Artificial Analysis, the model achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use cases, with strong performance in noisy, real-world environments and accurate capture of alphanumeric entities like postal codes and order IDs. This positions Gemini 3.5 Transcribe as a competitive option for building voice agents, real-time captioning tools, and post-call analytics pipelines.

Read original →

← Back to home