Gemini 3.8 text-to-speech says hello
Google DeepMind announced two new text-to-speech models in the Gemini family: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, described as the company's most expressive audio generation models yet. The models are meant to shift voice generation from static presets to a dynamic creative studio, letting creators, developers, and enterprises produce richer, more expressive audio and improve user experiences in products like Gemini Notebook and Google Vids. They join a fast-growing Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking, and are available across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
The two models target different use cases. Gemini 3.8 Flash TTS is built for deep creative direction and character design, allowing users to create entirely new voices from scratch with natural-language prompts for gaming, immersive audiobooks, podcasts, and interactive media, with line-by-line control over acting cues, pacing, dialect shifts, and backchanneling. Gemini 3.8 Flash-Lite TTS is built for high-volume, cost-efficient scale, optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance.
On voice creation, the models scale from 30 original voices to what Google describes as an infinite library. Generative voice design lets users create bespoke voices from scratch by customizing role, accent, and voice characteristics across more than 100 languages and dialects via natural-language prompting — for example, a dramatic fire-breathing dragon or a charismatic narrator with a distinct regional cadence. The voice library offers 2,000+ production-ready voices with broad language coverage, including regional varieties such as Mexican Spanish, Quebec French, and Scots English. Voice replication can recreate consistent vocal profiles from just a 30-second audio sample of your own voice or one you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent. Users can also save and manage custom voices for consistent performance and minimal drift across projects, and a voice remixing feature coming soon will let users pick a voice from the library and fine-tune timbre, pitch, pace, and accent via prompts.