Google introduced Gemini 3.5 Transcribe, a real-time speech-to-text model that automatically recognizes more than 85 languages, cleans up filler words, corrects minor speech errors, and formats transcripts for both live and recorded audio. The system achieves a reported 4.0 percent word error rate for streaming and 2.6 percent for recorded audio, while reducing latency by 70 percent compared with the earlier Chirp 3 model, and supporting function calling to other Gemini models. Gemini 3.5 Transcribe is available through Live and Interactions APIs, integrated into Google AI Studio, Gemini Enterprise Agent Platform, Android’s Gboard, and the Gemini macOS app, with browser-based Chrome support expected to follow soon.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.