Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android:https://blog.google/products-and-platforms/platforms/android/gemini-intelligence/ and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio:https://aistudio.google.com/live?model=gemini-3.5-transcribe-live and Gemini Enterprise Agent Platform:https://console.cloud.google.com/agent-platform/studio/multimodal-live?model=gemini-3.5-transcribe-live-preview.

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.

Gemini 3.5 Transcribe handles live language switches and seamless streaming transcription

Watch Gemini 3.5 Transcribe clean up speech disfluencies with smart transcription capabilities.

3.5 Transcribe delivers transcription with multi-speaker attribution and word-level timestamps.

Gemini 3.5 Transcribe’s performance represents a major advancement from our previous transcription model, Chirp 3, offering new capabilities, improved word error rates, and significantly better latency. As measured by Artificial Analysis, time to final transcription, for example, improves by 70%. On the FLEURS benchmark across a set of top languages and locales, the model delivers precise multilingual performance, improving over Chirp 3, and achieving a 5.50% WER in streaming mode and 5.04% WER in non-streaming use-cases.

In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.

Gemini 3.5 Transcribe lets you analyze files, generate images, and search in the Gemini app on macOS using just your voice.

See how Gemini 3.5 Transcribe uses Rambler on Android to automatically remove filler words and clean up speech.

Gemini 3.5 Transcribe leverages screen context on Google Antigravity to ensure accurate transcription accuracy.

By leveraging the Gemini Live API, developer platforms such as Agora:https://docs.agora.io/en/ai/models/asr/gemini, Fishjam:https://docs.fishjam.io/tutorials/gemini-live-integration, LangChain:https://docs.langchain.com/langsmith/trace-gemini-live, LiveKit:https://docs.livekit.io/agents/models/stt/gemini/, Pipecat:https://docs.pipecat.ai/api-reference/server/services/stt/google, Vercel:https://vercel.com/docs/ai-gateway/modalities/speech-to-text, and Vision Agents:https://visionagents.ai/integrations/stt/gemini enable developers to build and deploy high-performance voice-driven interfaces with ease. These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus entirely on crafting the user experience.

Companies like Vivo, Intellitek Health, and Lingopal have also shared positive feedback on 3.5 Transcribe, highlighting its impressive latency, accuracy, and expansive language support.

Check your inbox to confirm your subscription.