Gemini 3.8 text-to-speech
On September 23, 2026, Google's Gemini Audio team introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, described as their most expressive audio generation models yet, transforming voice generation from static presets into a dynamic creative studio. The models are rolling out across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids, targeting creators, developers, and enterprises. The announcement comes from Leland Rechis (Group Product Manager) and Alan Cowen (Director, Research Science) on behalf of the Gemini Audio Team.
Gemini 3.8 Flash TTS is built for deep creative direction and character design, allowing users to create entirely new voices from scratch using natural language prompts for gaming, immersive audiobooks, podcasts, and interactive media, with granular line-by-line control over acting cues, pacing, dialect shifts, and backchanneling. Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale, optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance. The two models complement the fast-growing Gemini Audio family, which includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Key capabilities include scaling from 30 original voices to an infinite library. Generative voice design lets users create bespoke voices from scratch by customizing role, accent, and voice characteristics across more than 100 languages and dialects via natural language prompting—whether a dramatic, fire-breathing dragon or a charismatic narrator with a distinct regional cadence. An expansive voice library offers 2,000+ production-ready voices with broad language coverage, including regional varieties like Mexican Spanish, Quebec French, and Scots English. Voice replication can recreate consistent vocal profiles from just a 30-second audio sample of your own voice or one you have rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect developers and vocal talent. Users can save and manage custom voices to ensure consistent performance and minimal drift across projects, and voice remixing is coming soon, allowing users to pick a voice from the library and fine-tune timbre, pitch, pace, and accent using prompts.