Google宣布推出Gemini 3.5 Transcribe,用于AI驱动的语音转文本

Ars Technica 模型

摘要

Google已宣布Gemini 3.5 Transcribe,这是一款用于语音转文本的AI模型,通过移除如“嗯”这样的填充词并支持85种语言来提高速度和准确性,正在其生态系统中推出。

<p>在我们等待(<a href="https://arstechnica.com/ai/2026/08/google-says-gemini-has-reached-1b-users-faster-than-any-other-google-product/">可能徒劳</a>)Gemini 3.5 Pro推出的同时,Google正在3.5系列中发布另一个模型。该公司宣布了Gemini 3.5 Transcribe,这是一款旨在通过编辑掉“嗯”和修正来简化语音输入、输出精炼AI文本的AI模型。该模型已经驱动了Pixel 11上的Gboard“Rambler”功能,但即将出现在整个Google生态系统中。</p> <p>根据Google的说法,Gemini 3.5 Transcribe比其之前的语音转文本引擎Chirp 3更快、更准确。这款新的AI模型从语音到最终转录文本的速度应该快约70%,实时语音错误率已降至5.5%。这只比Chirp 3稍好,Google测量为7.32%。尽管如此,使用语音输入时修复打字错误很麻烦,所以这里的任何改进都是有益的。</p> <p><a href="https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp.png"><img width="1000" height="562" src="https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp.png" class="fullwidth full" alt="" decoding="async" loading="lazy" srcset="https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp.png 1000w, https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp-640x360.png 640w, https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp-768x432.png 768w, https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp-384x216.png 384w, https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp-980x551.png 980w" sizes="auto, (max-width: 1000px) 100vw, 1000px"> 图片来源: Google </a></p><p><a href="https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/">阅读全文</a></p> <p><a href="https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/#comments">评论</a></p>
查看原文
查看缓存全文

缓存时间: 2026/08/26 21:18

# 谷歌发布Gemini 3.5 Transcribe:AI驱动的语音转文本模型 来源:https://arstechnica.com/ai/2026/08/google-announces-gemini-3-5-transcribe-for-ai-powered-speech-to-text/ 在我们等待(可能徒劳无功(https://arstechnica.com/ai/2026/08/google-says-gemini-has-reached-1b-users-faster-than-any-other-google-product/))Gemini 3.5 Pro发布的同时,谷歌在3.5系列中推出了另一款新模型。该公司宣布推出Gemini 3.5 Transcribe,这是一款旨在通过剪除“嗯”“啊”等语气词和修正语,输出精炼AI文本来优化语音输入的AI模型。这款模型已在Pixel 11 (https://arstechnica.com/gadgets/2026/08/google-pixel-11-series-review-is-the-magic-fading/)的Gboard“随言”功能中投入使用,且即将在谷歌全生态系统内推广。 据谷歌介绍,Gemini 3.5 Transcribe在速度和准确性上均显著超越其前代语音转文字引擎——Chirp 3。这款新AI模型从语音到最终转录文本的速度提升约70%,实时语音错误率已降至5.5%。虽然这一数字仅略优于谷歌测得的Chirp 3的7.32%,但考虑到使用语音输入时修正拼写错误的繁琐,任何改进都具有实用价值。 https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp.png[](https://cdn.arstechnica.net/wp-content/uploads/2026/08/gemini-3.5-audio-transcribe_fleu.width-1000.format-webp.png) *图片来源:谷歌* 新模型不仅在识别语音方面表现更优,还能更精准地把握用户真实意图。在语音输入过程中,Gemini 3.5 Transcribe能智能剔除意识流中杂乱的语气词,并支持实时文本编辑(当用户需要自我修正时)。它还能调用用户自定义词汇库处理“专业术语”,支持85种语言,最多可识别预录音频中的三位说话者。 当然,其局限性在于需要依赖AI准确理解语音内涵。从我的“随言”功能测试来看,该模型在处理短文本时能有效清理语言不连贯和口误,但AI本质上改变了用户原话的措辞——这在某些正式场合可能并不适用。

相似文章

Gemini 3.5 Transcribe 智能转录

Google DeepMind Blog

Google 推出 Gemini 3.5 Transcribe,这是一款新的 AI 模型,提供精确且智能的实时语音转文本转录,开发者可通过 API 获取。

借助 Gemini 3.5 Live Translate 实现流畅自然的语音翻译

Google DeepMind Blog

Google 发布了 Gemini 3.5 Live Translate,这是一款音频模型,支持超过 70 种语言的近乎实时的语音到语音翻译,并保留说话者的语调和节奏。该功能正在 Google 产品中逐步推出,包括 Gemini Live API、Google Meet 和 Google Translate。