Model Releases Google DeepMind Blog

Introducing Gemini 3.8 Live with Live Avatar

GeminiLive AvatarGoogle DeepMindmultimodalenterprise AI

Building on last week's Gemini 3.8 Live launch, Google DeepMind introduced Gemini 3.8 Live with Live Avatar, which brings near real-time visual presence to the company's native live dialogue models. The feature natively couples live dialogue capabilities with low-latency streaming video, pairing near real-time video generation with speech so the system listens, sees, and speaks through a dynamic visual persona.

The experience relies on precise lip-syncing, natural expressions, and fluid turn-taking, and supports a wide range of characters, each with a distinct look, voice, and expressive presence. Because it processes visual and audio inputs simultaneously, it can respond with expressive audio and video in near real time for more natural, multimodal conversations. Enterprises can use it to expand virtual offerings more interactively, such as engaging customer service or interactive walkthroughs.

Beyond visual presence, the feature is backed by Gemini's advanced reasoning through asynchronous tool calling: Live Avatar can trigger tool calls and fetch data in the background while continuing active dialogue, keeping conversational flow uninterrupted. Google illustrates this with a complex task like checking a guest in at a hotel, where tools are called in the background as the conversation proceeds.

For global scale, Live Avatar features native multilingual speech-to-speech synchronization. It dynamically adapts its lip-sync and expressions and can transition across 97 languages without degrading video fidelity or introducing visual drift, including switching languages mid-conversation. Google also notes that organizations needing distinct visual identities can draw on a library of diverse avatars to fit their brand needs.

Gemini 3.8 Live with Live Avatar is available starting today in Gemini Enterprise.

Read original →

← Back to home