What Google Announced
Building on the Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models that launched September 15, Google says Live Avatar brings "real-time visual presence" to its conversational AI. The pitch, laid out in a joint post from research scientist Shuo-yiin Chang and software engineer CJ Zheng of the Gemini Audio Team, is an agent that listens, sees and speaks through a dynamic visual persona, aimed at use cases like customer service and interactive walkthroughs across web, mobile and kiosks.
Google Cloud's own post, from group product manager Fabien Blanc-paques, frames this as a shift in priority: after competing mainly on speed and cost, the company says its focus has moved to interaction quality. Live Avatar was first previewed at Google Cloud Next 2026 and is now positioned as production-ready.
How It Works
The core capability pairs low-latency streaming video with Gemini's existing native speech-to-speech foundation. Google highlights precise lip-syncing, natural facial expressions and fluid turn-taking as what separates this from a static avatar bolted onto a voice model.
A more technical capability sits underneath the visual layer: asynchronous tool calling. Live Avatar can trigger tool calls and retrieve data in the background while the conversation continues uninterrupted, which Google demonstrates with an example of an agent checking in a hotel guest while pulling reservation data mid-conversation.
On language, Live Avatar features native multilingual speech-to-speech synchronization across 97 languages, the same language count Google claimed for the underlying live-dialogue models at their September 15 launch. What's new here is that the lip-sync and expressions are said to adapt dynamically as the avatar switches languages, without degrading video fidelity.
Custom Avatars and Safeguards
Enterprises can choose from a library of pre-built avatars or create a custom persona from a single reference photo and an audio sample. Custom avatar creation is gated behind a verification and allowlisting process, which Google positions as a guardrail against misuse. Every output, audio and video alike, carries an imperceptible SynthID watermark, intended to preserve transparency about AI-generated content.
Availability
Live Avatar is generally available now in Gemini Enterprise, with endpoints in the United States and European Union, and includes provisioned throughput, enterprise compliance features and data governance controls. It's worth noting precisely what shipped: Live Avatar is not a separate model ID on the general-purpose Gemini API. The API's live-audio model list still shows only gemini-3.8-live and gemini-3.8-live-extended-thinking, with no avatar entry or rate-card line item. In practice, that means Live Avatar's audience right now is enterprise buyers and the developers integrating on their behalf, rather than individual developers calling the Live API directly. Gemini 3.8 Live Extended Thinking, meanwhile, remains in private preview.
Why It Matters
The launch lands at a competitive moment for Google. The company's shares reportedly fell roughly 20% in recent months amid delays tied to Gemini 3.5 Pro, even as OpenAI and Anthropic pushed out newer models. Live Avatar is also arriving into a market where the boundaries between conversational AI, video generation and enterprise software are increasingly blurred, with Microsoft deepening Copilot's enterprise integration, Meta pushing open-source Llama toward business use, and dedicated avatar startups like Synthesia and HeyGen already established in the space. Google's strategy appears to be integration: tying a real-time video face directly to its existing live-dialogue and enterprise infrastructure, rather than competing as a standalone avatar product.