📰 Key Summary
This Gemini 3.8 Live Avatar news is general AI industry news with no direct connection to the current working directory’s project context, so here’s a straight translated summary:
Google DeepMind today launched Gemini 3.8 Live with Live Avatar, building on the momentum from last week’s Gemini 3.8 Live release by adding a near-real-time visual identity to the native live conversation model. By pairing near-instant video generation with speech, Live Avatar creates characters with a dynamic visual personality, all while being able to hear, see, and speak. The feature delivers precise lip sync, natural expressions, and smooth conversational turn-taking, letting businesses extend virtual services in a more interactive way — whether that’s customer service interactions or interactive tours — turning them into richer, more intuitive digital communication experiences. Starting today, Gemini 3.8 Live with Live Avatar is available on Gemini Enterprise. On the technical side, the system processes visual and audio input simultaneously to generate richer conversational content, and it supports asynchronous tool calling, letting Live Avatar trigger tool calls and pull data in the background without interrupting the conversation — handling complex tasks like the demoed hotel guest check-in scenario. On the multilingual front, the feature has native multilingual voice-to-voice sync, dynamically adjusting lip movement and expressions to switch seamlessly across 97 languages without degrading video quality or causing visual drift. For brand customization, businesses can draw on a diverse library of preset avatars, or generate a fully animated, responsive custom avatar from a high-quality reference image while preserving the reference image’s likeness, brand style, or character identity — though custom avatars are currently limited to enterprise allowlist applicants.
💬 JudyAI Lab Take
The launch of Gemini 3.8 Live with Live Avatar pushes real-time conversational AI from “understands you, answers correctly” to “looks like it’s actually responding” — a milestone worth noting in this round of the multimodal race.
What matters here isn’t how lifelike the avatar looks — it’s that enterprise-grade AI interaction is shifting from pure text/voice interfaces toward composite systems that process visual and audio input simultaneously while supporting background asynchronous tool calls. Scenarios like hotel check-in need the AI to quietly pull data and run processes without breaking the conversation — that “natural conversation up front, quiet work in the background” design thinking is what actually determines whether a product can ship, not surface features like lip sync or switching across 97 languages. For AI builders, this is a reminder that when evaluating a new release, you should look past the demo polish and check whether its tool integration and multitasking can hold up under real business workflows.
Next time you evaluate a similar product, ask first: how complex a task can its background tool-calling mechanism actually handle?
📅 Original Source Info
- Published: 2026-09-24T16:20
- Source: https://deepmind.google/blog/introducing-gemini-38-live-with-live-avatar/