首页
创新与人工智能
模型与研究
Gemini 模型
Gemini 3.8 文本转语音向您问好
2026年9月23日
12分钟阅读
Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 是我们迄今为止最具表现力的音频生成模型。通过 Google AI Studio、Gemini API、Gemini Enterprise、Gemini Notebook 和 Google Vids,您可以生成自定义角色声音并直接进行场景对话。
Leland Rechis
集团产品经理
Alan Cowen
代表 Gemini 音频团队的研究科学总监
分享
阅读 AI 生成的摘要
今天,我们向 Gemini 家族推出了两款新的文本转语音模型,将语音生成从静态预设转变为动态创意工作室。这些模型使创作者、开发者和企业能够创造更丰富、更具表现力的音频体验,同时提升 Gemini Notebook 和 Google Vids 等产品中的用户体验。
Gemini 3.8 Flash TTS:专为深度创意指导和角色设计而打造。使用自然语言提示从零开始创建全新的声音,让角色在游戏、沉浸式有声书、播客和互动媒体中栩栩如生。通过细微控制表演提示、节奏、方言转换和背景回应,逐行指导每一句台词的表演。
Gemini 3.8 Flash-Lite TTS:专为高容量、成本高效的规模化应用而设计。针对大规模配音、音频内容创作以及具有精细音调、节奏和情感细微差别控制的表达性语音代理进行了优化。
这些模型补充了我们快速增长的 Gemini Audio 家族,此前我们已推出了 3.5 Live Translate、3.5 Transcribe、3.8 Live 和 3.8 Live Extended Thinking。
创建并自定义您自己的声音
00:00
从 30 个原始声音扩展到无限的声音库。无论您需要一个完全原创的角色声音,还是一个一致的品牌代言人,我们的 3.8 Flash TTS 模型都能提供完整的声乐工作室功能。这使您能够为每个时刻创建并使用富有表现力、听起来自然的声音,同时赋能开发者和企业轻松构建自定义音频体验。
生成式语音设计:借助 Gemini 3.8 Flash TTS,通过自然语言提示在 100 多种语言和方言中定制角色、口音和声音特征,从零开始创建专属声音——无论是让一条喷火的巨龙栩栩如生,还是塑造一位具有独特地域韵律的魅力旁白者。
听听 Gemini 3.8 Flash TTS 如何生成来自墨尔本的高能量 DJ 声音。
听听 Gemini 3.8 Flash TTS 如何生成极其尖细、单调的机器人声音。
听听 Gemini 3.8 Flash TTS 如何让一条日本龙栩栩如生。
广阔的声音库:访问 2,000 多个生产就绪的声音,涵盖广泛的语言——包括墨西哥西班牙语、魁北克法语和苏格兰英语等地域变体。
声音复制:仅凭您自己或您有权使用的声音的 30 秒音频样本,即可重建一致的声音档案,并内置同意验证、SynthID 水印和 C2PA 凭证,以保护开发者和他们的声音人才。
保存与扩展:保存和管理您设计的自定义声音,确保在持续的项目中保持一致的性能并将漂移降至最低。
声音混音:即将推出,从我们的声音库中选择一种声音,并微调音色、音高、节奏和口音。使用提示来调整特征(例如,“添加微妙的美国南部口音”或“柔化表达方式”)。
逐行指导表演
选定声音后,这两款 TTS 模型都让您能够精确控制每句台词的呈现方式。
逐行指导表演:编写自己的舞台指示,或让 Gemini 通过自然脚本提示来引导交付——从冷静的客户服务代表到低声耳语的悬疑场景。
听听 Gemini 3.8 Flash TTS 如何为交互式语音代理实现自然、高度富有表现力的对话。
观看并聆听 Gemini 3.8 Flash TTS 如何利用细粒度的脚本控制,构建深度引人入胜、沉浸式的音频体验。
长篇幅生成:在数小时的连续音频中保持高音质、自然的语速和角色音色,同时最小化说话人漂移——非常适合播客和有声书。
原生双说话人场景编排:从单一脚本无缝引导多轮对话——无论是播客还是戏剧性叙事——同时通过自然的对话轮流机制,使两种声音保持清晰分离。
脚本化的语音爆发与背景回应:使用非语言线索(如
观看 Gemini 3.8 Flash TTS 如何将自然语言提示从零开始转化为定制化的声音人格。
观看 Gemini 3.8 Flash TTS 如何使创作者能够设计自定义场景,让动画对话栩栩如生。
观看 Gemini 3.8 Flash TTS 如何将脚本转化为完整的表演对话场景,让创作者指导语音表达和自然的轮流机制。
获取专为全球规模打造的富有表现力的高质量语音生成
Gemini 3.8 Flash TTS 提供领先的语音定制能力,在 Hume AI 的 Voice Design Benchmark(71.4)中稳居综合排名第一,并在口音建模方面以 60.8 分领先。
Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 在确保可靠性的同时实现真正富有表现力的表演,分别在 Hume AI 的综合质量指数(Overall Quality Index)中占据第一和第二的位置。与 Gemini 3.1 Flash TTS 相比,该模型在长篇幅内容和双说话人剧本控制等广泛用例方面显示出重大改进。
在 Voice Arena 的盲测人类偏好评估中,Gemini 3.8 Flash 和 Flash-Lite TTS 在包括日语、巴西葡萄牙语、越南语、现代标准阿拉伯语(MSA)、墨西哥西班牙语和印地语在内的关键全球语言中,在竞争对手中占据领先地位。凭借对 100 多种语言的支持,这些模型赋能创作者、开发者和企业,在全球范围内构建高质量的_multilingual_语音体验。
以信任、同意和透明为基础进行构建
我们构建了严格的保障措施来保护声音人才、尊重身份并确保内容透明,从而打造我们的声音创建和复制能力。对于声音复制,我们的系统利用同意验证:用户必须提供来自声音所有者的口头同意录音,且该录音需与参考说话人匹配,才能创建声音。
更广泛地说,由我们的 Gemini Audio 模型生成的每个音频片段都带有 SynthID 水印。这种不可见的水印直接编织在音频输出中,确保 AI 生成的语音可被检测,从而帮助防止虚假信息。有关我们安全性和责任方法的更多详细信息,请查阅模型卡片(model card)。
试用我们的全新 Google AI Studio 音频游乐场
从今天起,开发者可以在 Google AI Studio 中体验这些新的语音生成能力。该工具构建为声音设计工作区,您可以从零开始提示完全新的声音身份,或复制您自己的声音
1
,然后将其直接带入双说话人剧本编辑器,逐行指导表达。
在 Google AI Studio 中试用声音复制功能。
轻松部署高性能语音界面
通过使用 Gemini API,Agora、LiveKit、Pipecat、Vercel 等开发者平台使开发者能够轻松构建和部署高性能的语音生成体验。
我们与 Figma、HeyGen、Linguana、Wondercraft、99.co 和 Ollang 等公司合作,这些公司正在集成我们最新的 TTS 模型,以帮助加速全球配音、通过细微的区域口音本地化媒体,并大规模支持对话式语音代理。
开始使用我们最新的 Gemini Audio 模型:
Gemini 3.8 Flash TTS 从今天开始推出:
面向开发者:在 Gemini API 和 Google AI Studio 中
面向企业:即将通过 Gemini Enterprise 的 API 提供
面向所有人:在 Gemini Notebook 中
Gemini 3.8 Flash-Lite TTS 今日开始推出:
面向开发者:在 Gemini API 和 Google AI Studio 中
面向企业:即将通过 Gemini Enterprise 的 API 提供
面向所有人:在 Google Vids 中
发布于:
Gemini 模型
Home
Innovation & AI
Models & research
Gemini Models
Gemini 3.8 text-to-speech says hello
Sep 23, 2026
12 min read
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS are our most expressive audio generation models yet. Generate custom character voices and direct scene dialogue across Google AI Studio, Gemini API, Gemini Enterprise, Gemini Notebook, and Google Vids.
Leland Rechis
Group Product Manager
Alan Cowen
Director, Research Science, on Behalf of the Gemini Audio Team
Share
Read AI-generated summary
Today, we’re introducing two new text-to-speech models to the Gemini family, transforming voice generation from static presets into a dynamic creative studio. These models enable creators, developers, and enterprises to create richer, more expressive audio experiences, while enabling improved user experiences in products like Gemini Notebook and Google Vids.
Gemini 3.8 Flash TTS: Built for deep creative direction and character design. Create entirely new voices from scratch using natural language prompts to bring characters to life across gaming, immersive audiobooks, podcasts, and interactive media. Direct every performance line by line with granular control over acting cues, pacing, dialect shifts, and backchanneling.
Gemini 3.8 Flash-Lite TTS: Built for high-volume, cost-efficient scale. Optimized for high-volume dubbing, audio content creation, and expressive voice agents with fine-grained control over tone, pacing, and expressive nuance.
These models complement our fast-growing Gemini Audio family, following 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking.
Create and customize your own voices
00:00
Scale up from 30 original voices to an infinite library. Whether you need an entirely original character voice or a consistent brand ambassador, our 3.8 Flash TTS model powers a full vocal studio. This enables you to create and use expressive, natural-sounding voices for every moment, while empowering developers and enterprises to easily build custom audio experiences.
Generative voice design: With Gemini 3.8 Flash TTS, create bespoke voices from scratch by customizing role, accent and voice characteristics across more than 100 languages and dialects using natural language prompting — whether you're bringing a dramatic, fire-breathing dragon to life or crafting a charismatic narrator with a distinct regional cadence.
Hear how Gemini 3.8 Flash TTS generates a high-energy DJ voice from Melbourne.
Hear how Gemini 3.8 Flash TTS generates a super-tinny, monotone robot voice.
Hear how Gemini 3.8 Flash TTS brings a Japanese dragon to life.
Expansive voice library: Access 2,000+ production-ready voices with broad language coverage — including regional varieties like Mexican Spanish, Quebec French, and Scots English.
Voice replication: Recreate consistent vocal profiles from just a 30-second audio sample of your voice or a voice you have the rights to use, backed by built-in consent verification, SynthID watermarking, and C2PA credentials to protect both developers and their vocal talent.
Save and scale: Save and manage the custom voices you designed to ensure consistent performance and minimal drift across ongoing projects.
Voice remixing: Coming soon, pick a voice from our voice library and fine-tune timbre, pitch, pace, and accent. Use prompts to dial in characteristics (e.g. “add subtle Southern US accent” or “soften the delivery”).
Direct the performance, line by line
Once you've selected your voices, both TTS models give you precise control over how each line is delivered.
Direct performance line by line: Write your own stage directions or let Gemini steer delivery with natural script cues — from a calm customer service agent to a whispered suspense scene.
Hear how Gemini 3.8 Flash TTS enables natural, highly expressive conversations for interactive voice agents.
Watch and hear how Gemini 3.8 Flash TTS uses granular script control to build a deeply engaging, immersive audio experience.
Long-form generation: Maintain high voice quality, natural pacing, and character timbre across hours of continuous audio with minimal speaker drift — ideal for podcasts and audiobooks.
Native two-speaker scene staging: Direct multi-turn conversations seamlessly from a single script —whether for a podcast or dramatic storytelling—while keeping both voices distinctly separated with natural conversational turn-taking.
Scripted vocal bursts & backchanneling: Add realistic conversational texture using non verbal cues (like
See how Gemini 3.8 Flash TTS turns natural language prompts into bespoke vocal personas from scratch.
Watch how Gemini 3.8 Flash TTS enables creators to design custom scenes to bring animated dialogue to life.
See how Gemini 3.8 Flash TTS turns scripts into fully performed dialogue scenes, letting creators direct vocal delivery, and natural turn-taking.
Get expressive high-quality speech generation built for global scale
Gemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and also leading in accent modeling (60.8).
Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.
In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash and Flash-Lite TTS secure top positions amongst competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.
Build with trust, consent, and transparency
We built our voice creation and replication capabilities with strict safeguards to help protect voice talent, respect identity, and ensure content transparency. For voice replication our system leverages consent verification: users must provide a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created.
More broadly, every audio clip generated by our Gemini Audio models is watermarked with SynthID. This imperceptible watermark is woven directly into the audio output, ensuring AI-generated speech remains detectable to help prevent misinformation. For more details on our approach to safety and responsibility, review the model card.
Try our new Google AI Studio audio playground
Starting today, developers can experience these new speech generation capabilities in Google AI Studio. Built like a voice design workspace, you can prompt entirely new vocal identities from scratch or replicate your own voice
1
, then bring them directly into a dual-speaker screenplay editor to direct line-by-line delivery.
Try voice replication in Google AI Studio.
Deploy high-performance voice interfaces with ease
By using the Gemini API, developer platforms such as Agora, LiveKit, Pipecat, Vercel enable developers to build and deploy high-performance speech generation experiences with ease.
We’re partnering with companies like Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang, who are integrating our latest TTS models to help accelerate global dubbing, localize media with nuanced regional accents, and power conversational voice agents at scale.
Start using our latest Gemini Audio models:
Gemini 3.8 Flash TTS is rolling out starting today:
For developers: In the Gemini API and Google AI Studio
For enterprises: Coming soon via API in Gemini Enterprise
For everyone: In Gemini Notebook.
Gemini 3.8 Flash-Lite TTS is rolling out starting today:
For developers: In the Gemini API and Google AI Studio
For enterprises: Coming soon via API in Gemini Enterprise
For everyone: In Google Vids
Posted in:
Gemini models
| 刊期 | 得分 | 排名 | 结果 |
|---|---|---|---|
| 2026-10-05 | 9.4 | 14 | 入选 |
| 2026-10-04 | 9.4 | 19 | 未入选 |
| 2026-10-02 | 9.4 | 33 | 未入选 |
| 2026-09-24 | 11.62 | 15 | 入选 |