Radio
Now Playing
Quickyla Radio โ€” Click to play
Open โ†’
3 min left

Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time

It's the latest release from Meta Superintelligence Lab (MSI). Meta has introduced its first real-time audio model, Muse Voice Transcribe. The model can handle dictation and transcription for more tโ€ฆ

Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time
Engadget โ€” 1 September 2026
Text:
5 0 0

It's the latest release from Meta Superintelligence Lab (MSI).

Meta has introduced its first real-time audio model, Muse Voice Transcribe. The model can handle dictation and transcription for more than 20 speakers and can seamlessly handle multiple languages at once, Meta says.

Meta CEO Mark Zuckerberg, who recently returned to X after three years of not posting on the platform, shared an example of the model's ability to handle multiple speakers and languages at once. In the video, the transcription is able to automatically distinguish between multiple speakers and switch between languages. It's even able to pick up on "code-switching" and transcribe sentences that use words from multiple languages.

Muse Voice Transcribe is MSL's first real-time audio perception model โ€” rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model. pic.twitter.com/LViMDSkbim

"The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy," he explained. "It holds up on messy, real audio too โ€” trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers.

Meta's release comes less than a week after Google Gemini 3.5 Transcribe, its own audio model that boasts similar capabilities . But while Google is baking its audio model into Android and (eventually) Chrome, it's not clear if Meta has plans to integrate Muse Voice Transcribe into its flagship services.

For now though, people can experience the new mode's capabilities in Meta's recently released Meta AI Mac app. Because the Mac app is able to power voice-enabled features on other apps, Muse Voice Transcribe will now power dictation features on other services. The model is also available to developers within Muse Code and Meta's Model API. It's priced at $3 for 1,000 audio minutes. Additionally, theres a demo version of the Muse Transcribe on Meta's research blog .

3/ already powering dictation in the meta desktop app and muse code. live now via meta model api https://t.co/MFosERCV0E pic.twitter.com/UESTQhKJrH

Read Full Story at Engadget โ†’
Advertisement
"The model decides when to listen. It waits a little longer on hard words and commits faster on easy ones, using adaptive delay to predict each token and increase accuracy,"
โ€” Engadget
React:
Sources
Sponsored

More to Read

Flock Safety develops AI tool for police, raising privacy cโ€ฆ
๐Ÿ’ป Technology
Flock Safety develops AI tool for police, raising privacy concerns
Wired ยท 14 days ago
New Fire TV devices with Android 16 could be coming very soโ€ฆ
๐Ÿ’ป Technology
New Fire TV devices with Android 16 could be coming very soon
Android Authority ยท 14 days ago
What's the difference between Android's Qi2 And Qi2.2 wirelโ€ฆ
๐Ÿ’ป Technology
What's the difference between Android's Qi2 And Qi2.2 wireless charging?
Engadget ยท 14 days ago
Whistleblower Arturo Bรฉjar leads testimony in landmark triaโ€ฆ
๐ŸŒ World News
Whistleblower Arturo Bรฉjar leads testimony in landmark trial against Meta
NPR News ยท 14 days ago
USS Washington arrives in Arabian Sea after aircraft carrieโ€ฆ
๐ŸŒ World News
USS Washington arrives in Arabian Sea after aircraft carrier outcry
Al Jazeera ยท 13 days ago
Nigeria's jet fuel conundrum: Scarcity at home, abundance aโ€ฆ
๐ŸŒ World News
Nigeria's jet fuel conundrum: Scarcity at home, abundance abroad
DW World ยท 8 days ago
Full view