Meta introduces Muse Voice Transcribe, a real-time AI transcription model
Meta introduced Muse Voice Transcribe on September 1, 2026, and the model is already available in the Meta AI Mac app. It is Meta's first real-time audio perception model.
Meta says the model combines streaming speech-to-text, speaker diarization, and endpointing in a single system. It can transcribe more than 20 speakers at once, handle multiple languages simultaneously, and even follow code-switching, including sentences that mix words from different languages.
According to the research blog, Muse Voice Transcribe was trained on 70+ languages, with 25 validated at launch. Meta also says it can handle hour-long sessions with 20+ speakers, making it suitable for meetings, interviews, podcasts, and other live conversations.
The model is available now in the Meta AI Mac app, where it can also power dictation features in other apps. Developers can access it through Muse Code and Meta's Model API.
Meta priced Muse Voice Transcribe at $3 for 1,000 audio minutes. Mark Zuckerberg summed up the model's approach this way:
The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy.
Do real-time transcription tools become much more useful once they can follow multiple speakers, languages, and code-switching at the same time?