Meta Superintelligence Labs has released Muse Voice Transcribe, a real-time audio perception model that unifies streaming ASR, speaker diarization for over 20 speakers, and endpointing in a single autoregressive system. Instead of three stitched-together models, Muse Voice Transcribe processes 80ms audio chunks as soft tokens, uses reinforcement learning to balance transcription accuracy against latency, and inserts special tokens to handle speaker turns and speech boundaries. Benchmarks show leading word error and diarization rates at lower prices than several commercial rivals. The model supports 70+ languages with native code-switching and is currently available only as a hosted Meta Model API service.
This update represents a notable development in the Ai sector. Organizations and founders tracking this space should evaluate potential strategic and technical implications on their operations.