谷歌正在推出 Gemini 3.8 Flash TTS 和 Flash-Lite TTS,这是两款用于语音生成的新模型。Flash TTS 能够根据文本描述创建全新的声音,且这两款模型均支持超过 100 种语言,并允许用户为对话的每一行添加舞台指示(stage directions)。
据谷歌介绍,Gemini 3.8 Flash TTS 专为游戏角色、有声读物和播客等创意项目设计,而 Gemini 3.8 Flash-Lite TTS 则侧重于以低成本大规模生成语音,适用于配音、音频内容和语音代理。
Flash TTS 允许用户根据文本描述创建声音
借助 Gemini 3.8 Flash TTS,用户可以从零开始设计声音。据谷歌称,文本提示可以定义声音的角色、口音以及在广泛的语言和方言范围内的发声特征。对于不想从零开始的用户,谷歌提供了一个包含超过 2,000 个预设声音的库,其中包括墨西哥西班牙语、魁北克法语和苏格兰英语等地区变体。
一项语音克隆功能可以通过 30 秒的音频样本构建声音档案。要使用此功能,被克隆声音的人必须录制一段口头同意的声明,且录音中的声音必须与样本匹配。据谷歌称,Gemini 音频模型生成的每一段片段都带有不可听的 SynthID 水印,以帮助检测 AI 生成的语音。
谷歌还宣布了“语音混音”(Voice Remixing)功能,该功能将允许用户调整库中声音的音色、音调、节奏和口音,但该功能目前尚不可用。
脚本指示赋予用户对对话和演绎的控制权
两款模型都允许用户为每一行编写指示,或者让模型自行解读脚本提示。谷歌表示,这些模型可以生成数小时的音频,且“说话者漂移”(speaker drift)极小,这意味着声音随时间的变化微乎其微。
据该公司称,双声模式可以从单一脚本生成对话,同时保持声音的区分度。用户还可以编写笑声、叹息以及“嗯哼”等声音,以便将反应和停顿精确放置在所需位置。
在我自己的两项测试中,我使用了一个预设声音,并附带风格提示,要求它模仿一个说着带有浓重德国口音英语的恼怒柏林人。风格控制产生了令人信服口音和语调,但在某些时刻,两项测试的背景中都出现了高频嗡嗡声。在其中一项测试中,声音在片段结束时也发生了变化。
风格提示:一位来自柏林的德语母语者以外语身份说英语,带着浓厚且 unmistakable 的德国口音。他显然不是英语母语者:他以德语方式发音英语单词,将德语的节奏和语调应用于英语句子。“Th”变成“z”或“d”(如“ze”、“sink”、“dat”),“w”变成“v”(如“vat”、“vell”),词尾辅音变得更硬(如用“goot”、“bat”代替“bad”)。“r”是喉音且低沉,绝非英语中的“r”。元音扁平且短促,缺乏美式或英式英语中的柔和感。
声音带有鼻音且略显紧张,处于中音区,带有沙哑的颤动和类似 Kermit 的晃动,但更粗糙、更少木偶感。听起来像是凌晨 2 点时疲惫的柏林人。
演绎:快速、简短、断断续续。他在寻找英语单词时会短暂犹豫,有时甚至会生硬地插入德语单词而不进行翻译。偶尔会出现因沮丧而音调上扬的情况。干燥、讽刺、无动于衷,并略带恼火,因为他不得不解释任何事情——而且更恼火的是他必须用英语来做这件事。
谷歌表示,其新模型在 Hume AI 文本转语音基准测试的大多数类别中领先。| 图片来源:谷歌
谷歌开始推出这两款模型,企业 API 访问权限将随后开放
谷歌正通过 Gemini API 和 Google AI Studio 推出这两款模型。Flash TTS 也可在 Gemini Notebook 中使用,而 Flash-Lite TTS 可在 Google Vids 中使用。谷歌表示,两款模型的 Gemini Enterprise API 访问权限也将很快推出。
包括Agora、LiveKit、Pipecat和Vercel在内的开发者平台已经通过Gemini API支持集成。尽管早期的文本转语音(TTS)模型提供了欧盟数据处理服务,但谷歌尚未列出新模型的区域端点。
谷歌表示,免费层级的数据用于改进其产品,而付费层级的数据则不会。付费价格以每百万个token的美元金额列出,其中文本token按输入计费,音频token按输出计费。
计费标准
Flash TTS(至2026年底)
Flash TTS(2027年起)
Flash-Lite TTS(至2026年底)
Flash-Lite TTS(2027年起)
文本输入
$0.50
$1.00
$0.50
$1.00
音频输出
$9.00
$18.00
$6.00
$12.00
谷歌表示,一秒钟生成的音频等于25个音频token,这意味着一小时相当于90,000个token。这意味着在2026年底前,一小时的音频输出成本为Flash TTS的$0.81和Flash-Lite TTS的$0.54。从2027年1月1日起,这些成本将分别上升至$1.62和$1.08,文本输入将单独计费。
无炒作的人工智能新闻 – 由人工策划
订阅THE DECODER,享受无广告阅读、每周人工智能通讯、每年六次的独家“AI雷达”前沿报告、完整档案访问权限以及评论区的访问权。
立即订阅
Google is introducing Gemini 3.8 Flash TTS and Flash-Lite TTS, two new models for speech generation. Flash TTS can create new voices from text descriptions, and both models support more than 100 languages and let users add stage directions to individual lines of dialogue.
Gemini 3.8 Flash TTS is designed for creative projects such as game characters, audiobooks, and podcasts, while Gemini 3.8 Flash-Lite TTS focuses on low-cost speech generation at scale for dubbing, audio content, and voice agents, according to Google.
Flash TTS lets users create voices from text descriptions
With Gemini 3.8 Flash TTS, users can design voices from scratch. According to Google, a text prompt can define a voice's role, accent, and vocal traits across a wide range of languages and dialects. For users who don't want to start from zero, Google offers a library of more than 2,000 preset voices, including regional variants such as Mexican Spanish, Quebec French, and Scottish English.
A voice cloning feature can build a voice profile from a 30-second audio sample. To use it, the person whose voice is being cloned has to record a spoken statement of consent, and the voice in that recording must match the sample. Every clip the Gemini audio models generate carries an inaudible SynthID watermark to help detect AI-generated speech, according to Google .
Google has also announced "Voice Remixing," a feature that will let users adjust the timbre, pitch, tempo, and accent of library voices, but it isn't available yet.
Script directions give users control over dialogue and delivery
Both models let users write directions for each line or have the model interpret script cues on its own. Google says the models can generate hours of audio with minimal "speaker drift," meaning the voice barely changes over time.
A two-voice mode generates dialogue from a single script while keeping the voices distinct, according to the company. Users can also script laughter, sighs, and sounds like "mhm" to place reactions and pauses exactly where they want them.
In two of my own tests, I used a preset voice with a style prompt asking it to imitate an annoyed Berliner speaking English with a thick German accent. The style controls produced a convincing accent and intonation, but both tests had a high-pitched whine in the background at some points. In one of them, the voice also changed at the end of the clip.
Style prompt: A native German man from Berlin speaking English as a foreign language, with a thick, unmistakable German accent. He is clearly not a native English speaker: he pronounces English words the German way, applying German rhythm and intonation to English sentences. "Th" becomes "z" or "d" ("ze," "sink," "dat"), "w" becomes "v" ("vat," "vell"), and final consonants become harder ("goot," "bat" for "bad"). The "r" is guttural and throaty, never the English "r." Vowels are flat and short, lacking the softness found in American or British English.
The voice is nasal and slightly strained, in the mid-range, with a raspy quiver and a Kermit-like wobble, but grittier and less puppet-like. It sounds like a tired Berliner at 2 a.m.
Delivery: quick, clipped, choppy. He hesitates briefly when searching for an English word and sometimes inserts the German word flatly without translating it. Occasional exasperated upward pitch breaks. Dry, sardonic, unimpressed, and slightly annoyed that he has to explain anything at all—and even more annoyed that he has to do it in English.
Google says its new models lead in most categories of Hume AI's text-to-speech benchmark. | Image: Google
Google starts rolling out both models, with enterprise API access to follow
Google is rolling out both models through the Gemini API and Google AI Studio . Flash TTS is also available in Gemini Notebook , while Flash-Lite TTS is available in Google Vids . Google says access through the Gemini Enterprise API will follow soon for both models.
Developer platforms including Agora , LiveKit , Pipecat , and Vercel already support integration through the Gemini API. Google hasn't listed regional endpoints for the new models yet, though earlier TTS models offered EU data processing .
According to Google, data from the free tier is used to improve its products, while data from the paid tier isn't. Paid pricing is listed in US dollars per million tokens, with text tokens billed for input and audio tokens for output.
Billing
Flash TTS (through the end of 2026)
Flash TTS (starting in 2027)
Flash-Lite TTS (through the end of 2026)
Flash-Lite TTS (starting in 2027)
Text Input
$0.50
$1.00
$0.50
$1.00
Audio Output
$9.00
$18.00
$6.00
$12.00
Google says one second of generated audio equals 25 audio tokens, which puts an hour at 90,000 tokens. That means an hour of audio output costs $0.81 with Flash TTS and $0.54 with Flash-Lite TTS through the end of 2026. On January 1, 2027, those costs rise to $1.62 and $1.08, with text input billed separately.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-09-24 · 10.73 分