Meta has unveiled Muse Voice Transcribe, its first real-time speech recognition model developed by Meta Superintelligence Labs. Designed to transcribe speech as it happens, the model combines speed, accuracy, and multilingual capabilities in a single system, opening up new possibilities for digital assistants, transcription services and enterprise applications.
Unlike conventional speech-to-text systems that often struggle with multiple speakers or mixed-language conversations, Muse Voice Transcribe is built to handle real-world communication. It supports more than 70 languages, with 25 validated at launch, and can seamlessly process code-switching, a common practice where speakers alternate between languages within the same conversation. This feature is especially valuable in multilingual regions such as South Asia, where conversations frequently blend English with local languages.
Another standout capability is speaker “diarization”, allowing the model to distinguish between more than 20 speakers in lengthy conversations while accurately identifying who said what. Meta has also incorporated real-time end-pointing, enabling the system to determine when a speaker has finished talking, which helps AI assistants respond more naturally and with minimal delay.
Rather than relying on fixed listening windows, the model dynamically adjusts how long it processes each word, balancing responsiveness with transcription accuracy.
According to Meta, Muse Voice Transcribe currently ranks at the top of the Artificial Analysis streaming speech-to-text benchmark. The company has already integrated the technology into Meta AI for Mac and its Muse Code platform, while developers can access it through the Meta Model API.
The launch reflects the growing importance of voice as the next frontier for artificial intelligence. As AI assistants become more conversational, reliable speech recognition will be essential for everything from customer support and meeting transcription to accessibility tools and multilingual communication. By combining live transcription, multilingual understanding and advanced speaker recognition in one model, Meta is positioning Muse Voice Transcribe as a foundation for the next generation of voice-enabled AI experiences.


