微软人工智能部门发布了MAI-Transcribe-2-Streaming,这是一款用于实时转录的新模型。微软表示,该模型在Artificial Analysis的准确性排名中位列第一。该模型支持60种语言的转录,并能在一百多毫秒内输出首个部分结果。微软称,这使得语音代理能够在用户尚未说完一句话时就做出响应。到今年年底,按入门价格计算,一小时音频的费用为0.54美元。
微软还发布了两个新的文本转语音模型。MAI-Voice-2.1旨在以同一种声音讲述23种语言,并在每种语言中保持地道的口音。据微软介绍,MAI-Voice-2.1-Flash变体的延迟低至150毫秒,每百万字符的费用为15美元,而非之前的22美元。
这两款语音模型仅需几秒钟的参考音频即可克隆声音,内置的安全措施旨在防止滥用。这些模型可通过Microsoft Foundry和MAI Playground等平台获取,两款语音模型也在OpenRouter上提供。在一次测试中,约4,000名参与者中的一半认为这些声音来自真人。
去伪存真的人工智能新闻——由人类精心策划
订阅THE DECODER,享受无广告阅读、每周人工智能通讯、每年六次的独家“AI雷达”前沿报告、完整档案访问权限以及评论区的参与资格。
立即订阅
Microsoft AI has released MAI-Transcribe-2-Streaming, a new model for real-time transcription. Microsoft says it ranks first for accuracy on Artificial Analysis. The model transcribes 60 languages and delivers its first partial results in just over 100 milliseconds. Microsoft says this lets voice agents respond while someone is still mid-sentence. Through the end of the year, an hour of audio costs $0.54 at the introductory price.
Microsoft also released two new text-to-speech models. MAI-Voice-2.1 is designed to speak 23 languages in the same voice, with a native accent in each one. According to Microsoft, the MAI-Voice-2.1-Flash variant hits a latency of 150 milliseconds and costs $15 per million characters instead of $22.
Both voice models can clone a voice from just a few seconds of reference audio, and built-in safeguards are meant to prevent misuse. The models are available through Microsoft Foundry and the MAI Playground , among other platforms, and the two voice models are also on OpenRouter . In one test, about half of the 4,000 participants thought the voices belonged to a real person.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe now
首次收录 · 2026-10-03 · 10.33 分