首页
创新与人工智能
模型与研究
Gemini 模型
推出带有实时虚拟形象的 Gemini 3.8 Live
2026年9月24日
3分钟阅读
带有实时虚拟形象的 Gemini 3.8 Live 将实时视觉存在感带入 Gemini 的对话式人工智能。通过将我们的实时对话能力与低延迟流媒体视频原生耦合,Live Avatar(实时虚拟形象)为企业及其用户带来了更加自然直观的对话体验。
张硕茵
研究员
郑志远
软件工程师,代表 Gemini 音频团队
分享
继上周 Gemini 3.8 Live 发布的热潮之后,今天我们很高兴地推出带有实时虚拟形象的 Gemini 3.8 Live——为我们的原生实时对话模型带来近实时的视觉存在感。通过将近实时视频生成与语音相结合,Live Avatar 功能创建了一种能够倾听、观看并以动态视觉人格说话的体验。
凭借精确的口型同步、自然的表情和流畅的轮流发言,Live Avatar 使企业能够以更具互动性的方式扩展其虚拟服务。无论是提供引人入胜的客户支持还是交付交互式导览,它都将数字交流转化为更丰富、更易访问的体验。
从今天起,带有实时虚拟形象的 Gemini 3.8 Live 已在 Gemini Enterprise 中提供。
查看带有实时虚拟形象的 Gemini 3.8 Live 如何支持各种角色,每个角色都有独特的外观、声音和表现力。
更自然的多模态对话
对话本质上是多模态的:我们倾听、观看、说话并使用面部表情进行交流。Live Avatar 将这些能力带入了企业代理。通过同时处理视觉和音频输入,它生成丰富的对话,带来更全面的体验。
观看带有实时虚拟形象的 Gemini 3.8 Live 如何近乎实时地接收其看到和听到的内容,并通过富有表现力的音频和视频进行回应,从而实现更自然的对话。
具有持续存在感的异步工具执行
除了视觉存在感之外,该功能还由 Gemini 的高级推理能力提供支持。通过异步工具调用,Live Avatar 可以在继续活跃对话的同时在后台触发工具调用并获取数据,在处理复杂任务的同时确保不间断的对话流程。
查看带有实时虚拟形象的 Gemini 3.8 Live 如何处理诸如酒店客人入住等复杂任务。在对话持续进行的同时在后台调用工具。
面向全球规模的对话体验
对话存在感应感觉自然,不应受语言限制。Live Avatar 具备原生的多语言语音到语音同步功能。该功能动态调整其口型同步和表情,并能在 97 种语言之间无缝切换,而不会降低视频保真度或引入视觉漂移。
观看带有实时虚拟形象的 Gemini 3.8 Live 如何在对话中途切换语言,口型同步和表情在 97 种语言之间无缝适应。
符合您品牌需求的实时虚拟形象
组织通常需要独特的视觉身份以契合其品牌形象。除了多样化的预设虚拟形象库外,组织还可以定制其实时虚拟形象。通过高质量参考图像,开发人员可以生成完全动画化、响应式的虚拟形象,同时保留参考 likeness(相似性)、品牌风格或角色身份。自定义虚拟形象的创建目前仅对企业白名单用户开放。
以信任和透明为核心
我们构建了 Live Avatar,并配备了严格的安全措施,旨在尊重身份并保持人工智能生成内容的透明度。我们所有 AI 产品生成的输出均带有 SynthID 水印。这种不可见的水印直接编织在音频和视频输出中,有助于确保可检测 AI 生成的内容,从而最大限度地减少错误信息和误 attribution(归因)。要探索我们在安全和负责任部署方面的全面方法,请阅读我们的模型卡片。
开始使用带有实时头像的 Gemini 3.8 Live
带有实时头像的 Gemini 3.8 Live 现已在 Gemini Enterprise 中提供。请查阅 API 文档以开始使用。
发布于:
Gemini 模型
Home
Innovation & AI
Models & research
Gemini Models
Introducing Gemini 3.8 Live with Live Avatar
Sep 24, 2026
3 min read
Gemini 3.8 Live with Live Avatar brings real-time visual presence to Gemini’s conversational AI. By natively coupling our live dialogue capabilities with low-latency streaming video, Live Avatar enables a more natural and intuitive conversational experience for enterprises and their users.
Shuo-yiin Chang
Research Scientist
CJ Zheng
Software Engineer, on behalf of the Gemini Audio Team
Share
Building on the momentum of last week's Gemini 3.8 Live launch, today we are excited to introduce Gemini 3.8 Live with Live Avatar — bringing near real-time visual presence to our native live dialogue models. By pairing near real-time video generation with speech, the Live Avatar feature creates an experience that listens, sees, and speaks with a dynamic visual persona.
With precise lip-syncing, natural expressions, and fluid turn-taking, Live Avatar enables enterprises to expand their virtual offerings more interactively. Whether providing engaging customer service or delivering interactive walkthroughs, it transforms digital exchanges into richer, more accessible experiences.
Starting today, Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise.
See how Gemini 3.8 Live with Live Avatar supports a wide range of characters, each with a distinct look, voice, and expressive presence.
More natural and multimodal conversations
Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate. Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience.
Watch how Gemini 3.8 Live with Live Avatar takes in what it sees and hears in near real time, responding with expressive audio and video for a more natural conversation.
Asynchronous tool execution with continuous presence
Beyond visual presence, the feature is backed by Gemini’s advanced reasoning. With asynchronous tool calling, Live Avatar can trigger tool calls and fetch data in the background while continuing active dialogue, handling complex tasks while ensuring an uninterrupted conversational flow.
See how Gemini 3.8 Live with Live Avatar handles complex tasks like checking in a guest at a hotel. Calling tools in the background while the dialogue continues uninterrupted.
Conversational experiences built for global scale
Conversational presence should feel natural and not be limited by languages. Live Avatar features native multilingual speech-to-speech synchronization. The feature dynamically adapts its lip-sync and expressions and can seamlessly transition across 97 languages without degrading video fidelity or introducing visual drift.
Watch how Gemini 3.8 Live Avatar switches between languages mid-conversation, with lip-sync and expressions adapting seamlessly across 97 languages.
A Live Avatar to fit your brand needs
Organizations often need distinct visual identities to fit their brand. In addition to a library of diverse, preset avatars, organizations can customize their Live Avatars. From a high-quality reference image, developers can generate a fully animated, responsive avatar while preserving reference likeness, brand styling, or character identity. Custom avatar creation is currently available only through enterprise allowlisting.
Trust and transparency at its core
We built Live Avatar with strict safeguards designed to respect identity, and keep AI-generated content transparent. All output generated by our AI products is watermarked with SynthID. This imperceptible watermark is woven directly into the audio and video output, helping to ensure AI-generated content remains detectable to help minimise misinformation and misattribution. To explore our comprehensive approach to safety and responsible deployment, read our model card.
Get started with Gemini 3.8 Live with Live Avatar
Gemini 3.8 Live with Live Avatar is available in Gemini Enterprise. Explore the API documentation to get started.
Posted in:
Gemini models
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-04 | 9.4 | 17 | 入选 |
| 2026-09-25 | 11.66 | 7 | 入选 |