AI for everyone in every language
Google says its technologies and products now power everyday interactions in more than 300 languages spoken by over 7 billion people, or 86% of the global population. The company frames that milestone as progress against a long-standing gap: for decades technology has served a handful of dominant languages, leaving thousands of living languages and dialects poorly represented or absent from the digital world. Google Translate launched in 2006 with the goal of breaking down language barriers, and AI advances have expanded it from a handful of languages to more than 250 today — but Google argues translating text alone isn't enough, and that systems must honor cultural nuance and how people actually communicate.
The company says it has moved beyond text transcripts to "native audio intelligence," training models such as Gemini to process audio directly rather than following the old rigid pipeline of transcribing audio to text, processing it, and synthesizing it back to speech — a pipeline that strips away tone, pacing, emotion, and context. The motivation is that people don't speak in neat grammatical sentences: they laugh, overlap, hesitate, and weave languages together mid-sentence, as in Spanglish or Hinglish.
On the product side, Google says Gemini 3.5 Live Translate powers real-time spoken translation across 70 languages and 2,000+ language pairs, naturally capturing code-switching and emotional cues. Gemini 3.5 Transcribe is described as its most precise speech-to-text model yet, turning raw audio into polished, formatted text even in noisy environments or with complex jargon; it also powers features like Rambler on Android Gboard, which removes filler words, fixes grammar and punctuation, allows editing or rewriting via voice commands, and supports seamless switching between languages.
Finally, Google outlines its 1,000 Languages Initiative: the goal is to support the world's 1,000 most-spoken languages, going beyond the languages where AI already performs best. To help achieve this, its Universal Speech Model — trained on 12 million hours of audio — used cross-lingual transfer learning, which lets models transfer what they learn from data-rich languages to improve speech understanding in languages with far less training data.