Gemini 3.8 text-to-speech

原始链接: https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/

Gemini 3.8 Flash TTS 具备领先的语音定制能力,在 Hume AI 的语音设计基准测试中以 71.4 分位列总榜第一,并在口音建模(60.8 分)方面同样处于领先地位。 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 能够在不牺牲可靠性的前提下实现极具表现力的语音演绎,并在 Hume AI 的综合质量指数中分别占据第一和第二名。与 Gemini 3.1 Flash TTS 相比,该模型在长文本内容和双人剧本控制等多种使用场景中均有显著提升。 在 Voice Arena 的盲测评估中,Gemini 3.8 Flash 和 Flash-Lite TTS 在日语、巴西葡萄牙语、越南语、现代标准阿拉伯语 (MSA)、墨西哥西班牙语和印地语等全球主要语言中,均在竞争对手中名列前茅。这些模型支持超过 100 种语言,助力创作者、开发者和企业在全球范围内构建高质量的多语言语音体验。

这篇 Hacker News 帖子讨论了谷歌发布的 Gemini 3.8 文本转语音功能,该功能允许使用 30 秒的音频样本进行声音克隆。谷歌通过同意验证、SynthID 水印和 C2PA 凭证来强调安全性。 评论者反应不一。一些用户,特别是业余作家,对使用该工具为个人项目制作有声读物式的旁白感到兴奋。然而,另一些人对隐私、数据训练的不透明度以及冒充攻击的滥用潜力表示担忧。 讨论还强调了人们对本地开源替代方案(如“KeenLore”)日益增长的兴趣,这些方案无需云端依赖或基于代币的成本,即可提供隐私保护和创作控制。总体而言,该帖子反映了富有表现力的高质量合成语音的快速发展与人们对人工智能生成内容在日常生活中常态化的正当焦虑之间的紧张关系。
相关文章

原文

Gemini 3.8 Flash TTS delivers leading voice customization capabilities, securing the #1 overall spot on Hume AI’s Voice Design Benchmark (71.4) and also leading in accent modeling (60.8).

Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS enable truly expressive performances without sacrificing reliability, also securing the #1 and #2 spots respectively on Hume AI’s Overall Quality Index. The model shows major improvements on a wide range of use cases such as long-form content and dual-speaker screenplay control compared to Gemini 3.1 Flash TTS.

In blind human preference evaluations on Voice Arena, Gemini 3.8 Flash and Flash-Lite TTS secure top positions amongst competitors in key global languages, including Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic (MSA), Mexican Spanish and Hindi. With support for over 100 languages, these models empower creators, developers, and enterprises to build high-quality, multilingual voice experiences worldwide.

联系我们 contact @ memedata.com