Gemini 3.5 Transcribe
Gemini 3.5 Transcribe is a speech-to-text model designed for precise, real-time transcription that handles complex audio environments, filler words, and multi-speaker attribution.
Gemini 3.5 Transcribe is a speech-to-text model designed for precise, real-time transcription that handles complex audio environments, filler words, and multi-speaker attribution.
What the product does and how it is positioned
Gemini 3.5 Transcribe is an intelligent speech-to-text model engineered to convert raw audio into polished, formatted text. It is designed to manage background noise, specialized jargon, and natural speech patterns effectively.
The model supports developers through real-time streaming and pre-recorded audio processing APIs. It integrates into various Google surfaces to provide context-aware transcription and voice-command capabilities.
Source-supported ways to use the product
Developers use the Live API to build high-performance voice agents and real-time captioning tools.
Organizations process recorded meetings and call logs to generate transcripts with speaker attribution and timestamps.
Gemini 3.5 Transcribe offers significant improvements in latency and accuracy compared to previous models. It is optimized to understand intent and recognize custom vocabulary, supporting high performance in diverse environments.
Checks to run with your own material and workflow
What was checked and when
Answers based on the source-checked product record
The model supports both real-time continuous streaming for interactive voice applications and the processing of pre-recorded audio files such as meetings and call logs.
Gemini 3.5 Transcribe features smart transcription capabilities that automatically identify and remove filler words like 'ums' and 'ahs' while cleaning up self-corrections.
Yes, the model provides multi-speaker identification for up to three speakers in pre-recorded audio, including word-level timestamps for each attribution.
The model features global language support, capable of automatically detecting and transcribing over 85 languages, including various regional accents and dialects.
Developers can access the model in public preview via the Gemini API in Google AI Studio, the Gemini Enterprise Agent Platform, and Google Antigravity.