今天,我们推出了 Gemini 3.5 Transcribe,这是我们迄今为止最精确的语音转文字模型,专为智能语音互动而设计。与在背景噪音、复杂术语和语音不流畅清理方面表现不佳的传统语音识别模型不同,Gemini 3.5 Transcribe 可以将原始音频直接转换为准确、精炼且格式化的文本。
在我们的产品中,例如 Gemini 应用程序和 Android 平台,我们已经看到消费者从这个转录模型中受益,体验到新的语音功能,例如 Android 上的 Rambler:https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/ 以及 macOS 上的 Gemini 应用程序。现在,开发者可以在 Google AI Studio 的 Gemini API(https://aistudio.google.com/live?model=gemini-3.5-transcribe-live)和 Gemini 企业代理平台(https://console.cloud.google.com/agent-platform/studio/multimodal-live?model=gemini-3.5-transcribe-live-preview)中使用 Gemini 3.5 Transcribe 构建类似功能。
我们构建 3.5 Transcribe 以无缝融入开发者工作流程,无论您是在构建语音代理、实时字幕工具还是通话后分析管道。该模型可通过两个独立的 API 使用:
Gemini 3.5 Transcribe 旨在捕捉您的自然说话风格,更好地理解您的意图并识别自定义词汇,从而使您能够通过语音执行任务。
Gemini 3.5 Transcribe 支持实时语言切换和无缝流式转录。
观看 Gemini 3.5 Transcribe 利用智能转录功能清理语音不流畅的表现。
3.5 Transcribe 提供多说话人标注和逐词时间戳的转录。
Gemini 3.5 Transcribe 的性能相比我们之前的转录模型 Chirp 3 有重大进步,提供了新的功能、改进的词错误率和显著更低的延迟。例如,根据 Artificial Analysis 的测量,最终转录时间提高了 70%。在 FLEURS 基准测试中,涵盖多种主要语言和地区,该模型提供了精确的多语言性能,相比 Chirp 3 有所提高,在流式模式下实现 5.50% 的词错误率(WER),非流式使用场景下实现 5.04% 的词错误率(WER)。
除了Google AI Studio和Gemini企业代理平台中的Gemini API外,3.5 转录比标准的语音转文字功能更进一步,让在Google生态中工作感觉更加自然和直观。通过将上下文感知理解直接引入Gboard、Antigravity、Gemini应用和Chrome等日常界面,它可以轻松捕捉细微差别、意图和内联编辑。
Gemini 3.5 转录让您只用语音就可以在macOS上的Gemini应用中分析文件、生成图像和进行搜索。
看看Gemini 3.5 转录如何在Android上的Rambler中自动删除填充词并清理语音。
Gemini 3.5 转录利用Google Antigravity上的屏幕上下文来确保转录的准确性。
通过利用 Gemini Live API,开发者平台如 Agora:https://docs.agora.io/en/ai/models/asr/gemini、Fishjam:https://docs.fishjam.io/tutorials/gemini-live-integration、LangChain:https://docs.langchain.com/langsmith/trace-gemini-live、LiveKit:https://docs.livekit.io/agents/models/stt/gemini/、Pipecat:https://docs.pipecat.ai/api-reference/server/services/stt/google、Vercel:https://vercel.com/docs/ai-gateway/modalities/speech-to-text 以及 Vision Agents:https://visionagents.ai/integrations/stt/gemini,使开发者能够轻松构建和部署高性能的语音驱动界面。这些平台在后台管理复杂的实时媒体流基础设施,使开发者可以完全专注于打造用户体验。
Vivo、Intellitek Health和Lingopal等公司也分享了对3.5 转录的积极反馈,强调其令人印象深刻的延迟、准确性及广泛的语言支持。
检查您的收件箱以确认订阅。
Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.
Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android:https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/ and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio:https://aistudio.google.com/live?model=gemini-3.5-transcribe-live and Gemini Enterprise Agent Platform:https://console.cloud.google.com/agent-platform/studio/multimodal-live?model=gemini-3.5-transcribe-live-preview.
We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:
Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.
Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription
Watch Gemini 3.5 Transcribe clean up speech disfluencies with smart transcription capabilities.
3.5 Transcribe delivers transcription with multi-speaker attribution and word-level timestamps.
Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.
In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.
Gemini 3.5 Transcribe lets you analyze files, generate images, and search in the Gemini app on macOS using just your voice.
See how Gemini 3.5 Transcribe uses Rambler on Android to automatically remove filler words and clean up speech.
Gemini 3.5 Transcribe leverages screen context on Google Antigravity to ensure accurate transcription accuracy.
By leveraging the Gemini Live API, developer platforms such as Agora:https://docs.agora.io/en/ai/models/asr/gemini, Fishjam:https://docs.fishjam.io/tutorials/gemini-live-integration, LangChain:https://docs.langchain.com/langsmith/trace-gemini-live, LiveKit:https://docs.livekit.io/agents/models/stt/gemini/, Pipecat:https://docs.pipecat.ai/api-reference/server/services/stt/google, Vercel:https://vercel.com/docs/ai-gateway/modalities/speech-to-text, and Vision Agents:https://visionagents.ai/integrations/stt/gemini enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.
Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.
Check your inbox to confirm your subscription.