Intelligent transcription with Gemini 3.5 Transcribe

Gemini
Google introduced Gemini 3.5 Transcribe, a precise speech-to-text model featuring real-time streaming, smart formatting, custom vocabulary, and multi-speaker identification.

Summary

Google has announced Gemini 3.5 Transcribe, its most advanced speech-to-text model designed for precise and intelligent voice interactions. The model converts raw audio directly into polished, formatted text while handling background noise, self-corrections, and filler words. Available to developers through Google AI Studio and the Gemini Enterprise Agent Platform, it supports real-time streaming and pre-recorded audio processing with speaker attribution. Key features include low word error rates, custom vocabulary support, function calling, and automatic detection of over 85 languages. Integrations span across Android, macOS, Gboard, and upcoming Chrome features, empowering both developers and consumers with natural voice capabilities.

(Source:Gemini)