Google made Gemini 3.8 Live with Live Avatar available to enterprise customers on September 24, 2026. The feature generates a speaking video avatar during a live conversation, with lip movements and expressions synchronized to the agent’s speech.
The release goes beyond visual presence. Google says the agent can interpret a user’s camera feed or screen share while listening, then call tools in the background without ending the conversation. Those capabilities could make a customer-service or guided-support agent more useful than a talking face alone. The production announcement does not, however, provide independent proof of latency, visual quality, or reliability.
The Avatar Is New; Gemini 3.8 Live Came First
Google announced the underlying Gemini 3.8 Live models the previous week. The September 24 development is the general availability of Live Avatar in Gemini Enterprise. The base voice model had already debuted, and Google says Live Avatar had previously been previewed at Google Cloud Next 2026.
The distinction affects where the feature is available. Google’s Gemini 3.8 Audio model card lists several distribution channels for the broader Live model family, including the Gemini app. The avatar launch announcement names Gemini Enterprise. It does not announce Live Avatar as a new feature in the consumer Gemini app.
General availability also has limits. Google Cloud says Gemini 3.8 Live Extended Thinking remains in private preview. An enterprise can build with Live Avatar without assuming it has access to the Extended Thinking variant.
The Agent Can Watch and Work While Its Avatar Speaks
Live Avatar pairs Gemini’s native speech-to-speech conversation with generated video. Google describes synchronized lip movements, expressions, and turn-taking, including when a conversation changes language. The company positions it for agents on websites, mobile devices, and interactive kiosks.
Users can also show the agent what they see. Google says Gemini 3.8 Live can process camera or screen-share input alongside audio. A support agent could discuss an item visible on camera or guide someone through an on-screen process while maintaining a spoken exchange. The generated avatar is the agent’s output; the camera feed or screen share is information the agent receives. They are separate capabilities working in the same conversation.
Asynchronous tool calling, Google says, lets the agent request data or trigger an API operation while dialogue continues. Its launch materials demonstrate a hotel check-in interaction built around that behavior. In a separate insurance claims-intake demo, a user describes damage on camera while an agent workflow checks policy information and prepares an adjuster packet.
These demos illustrate intended workflows. They do not show that a deployment will handle claims correctly or complete transactions reliably. An organization would still need to test its own tools, permissions, failure handling, and handoff to a person.
What the Live API Exposes to Developers
The Gemini Live API documentation lists gemini-3.8-live as generally available and includes Live Avatar among its features. It does not present the avatar as a separate model ID that developers must select instead of Gemini 3.8 Live.
The API uses a stateful WebSocket connection for bidirectional interaction. Its specifications list audio, text, and image or video input; audio and text output; and MP4 video output for live avatars. The documented visual input format is JPEG at one frame per second. Developers should check that limit before designing an application that depends on noticing brief events in a camera feed. “Live visual understanding” does not promise that the model examines every frame of full-motion video.
Google’s documentation offers SDK, WebSocket, and Agent Development Kit integration paths. Connecting a stream is only part of the technical work: the agent must keep its conversation, visual context, tool results, and displayed avatar in sync when a user interrupts or a backend request takes longer than expected. Google says its native speech-to-speech system supports interruption recovery, but the launch materials do not provide an independent end-to-end latency test for an avatar-based deployment.
General Availability Still Has Access Conditions
Google Cloud says Live Avatar is available through US and EU endpoints, with provisioned throughput and enterprise governance options. Teams planning a deployment should confirm the endpoint and capacity arrangement for their use case. “Generally available” is not a promise of unrestricted scale.
Google offers a library of preset characters that enterprise customers can use. Creating a custom avatar from a reference likeness, however, requires enterprise allowlisting and verification. Google’s custom-avatar demo shows a workflow involving system instructions, a reference photo, and an audio sample; it should not be mistaken for a self-service feature available to every developer.
For branded agents, that restriction means a company can assess the conversation and integration using the available feature but cannot assume it can immediately deploy a particular employee’s or performer’s likeness. Google’s API documentation describes authorized employee likenesses and paid talent with releases as possible corporate use cases. Permissions around a person’s image and voice remain a deployment responsibility, regardless of how easy avatar generation becomes.
Google directs customers to its pricing information and sales team for commercial planning. The cited launch materials do not establish a single verified cost for a complete avatar interaction, so it would be misleading to infer one from the base model’s availability.
The 97-Language Claim Needs a Documentation Check
Google says Gemini 3.8 Live can understand and speak 97 languages, detect language automatically, and switch languages during a conversation. Its avatar announcement further claims lip-sync and expressions adapt across those switches without visual drift. That is a substantial claim for organizations serving multilingual customers, though the demonstrations are Google’s own.
The documentation also contains a discrepancy worth resolving before procurement. The general Live API overview describes conversation in 24 supported languages, while the Gemini 3.8 Live launch materials say 97. The API page covers multiple models, so its broader feature summary may not describe every model-specific capability. The supplied documentation does not reconcile the counts. Developers should confirm support for their required languages and test switching in the actual avatar configuration. Neither number guarantees equal quality across languages.
Language recognition, spoken response quality, and believable lip-sync call for different tests. A system might switch languages successfully yet struggle with a particular accent, noisy input, or visual synchronization. Google has not provided an independent avatar-specific evaluation in the cited launch materials that resolves those questions.
Frequently Asked Questions
4 questions
1Is Gemini 3.8 Live Avatar available in the Gemini app?
Google has not announced Live Avatar as a consumer Gemini app feature. The September 24, 2026 launch makes it generally available in Gemini Enterprise. Google’s model card lists the broader Gemini 3.8 Live family among products distributed through the Gemini app, but that listing does not confirm the newly launched avatar feature is available there.
2
Sources
- Gemini 3.8 Live with Live Avatarblog.google
- general availability of Live Avatar in Gemini Enterprisecloud.google.com
- Gemini 3.8 Audio model carddeepmind.google
- Gemini Live API documentationdocs.cloud.google.com


