The Falcon-ASR speech recognition model processes Arabic audio with a word error rate of 20.92% and contains 1.6 billion parameters. Developed at the Technology Innovation Institute (TII) in Abu Dhabi, the system focuses on the Emirati dialect but also handles English, French, Spanish and Portuguese.
In this article
Performance against existing systems
Testing against six Arabic datasets shows an average word error rate of 20.92%. This figure beats the best published result of 23.17% available in the leaderboard snapshot used for comparison. An internal evaluation of Emirati speech recorded the lowest word and character error rates among the systems compared.
The model outputs word-level timestamps. Each transcribed word links to its specific position in the audio file.
Handling spoken Arabic
Arabic speech changes by region, speaker and recording environment. A model trained on formal news broadcasts may still struggle with conversations in Emirati or recordings captured over a phone line. Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), which makes training and evaluation harder.
Training data included Emirati, MSA, other Gulf and Arabic dialects, and English. The goal is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.
Arabic benchmark results
The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average word error rate across six test sets. It also reports character error rate (CER). Lower values are better for both metrics. This evaluation followed that protocol.
WER = Word Error Rate; CER = Character Error Rate. A lower value indicates better performance.
The team evaluated Falcon-ASR on the same six benchmarks using the leaderboard’s pinned manifests. Competitor figures are the published leaderboard averages checked on 30 September 2026. Falcon-ASR’s average WER is 2.25 percentage points better than the best published result in that snapshot.
Evaluating Emirati speech
Public evaluation data already includes Emirati: Casablanca has a UAE subset. The team complemented that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess transcription accuracy beyond the public UAE subset.
In the internal Emirati evaluation, Falcon-ASR achieved a 22.73% WER and 10.19% CER.
Falcon-ASR has the lowest WER and CER among the systems compared here. Its WER is 4.07 percentage points below Qwen3-Omni, the next best result. The results show improved transcription accuracy at both the word and character level on this evaluation.
Training for different recording conditions
Training data included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch. The same treatment applied to Emirati recordings exposed the model to a range of conditions it may encounter in meetings, calls and other everyday recordings.
English and other languages
Falcon-ASR also transcribes English with the same model weights. In the evaluation on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.
The model also supports French, Spanish and Portuguese. All five languages use the same weights, without requiring a language flag. The output is a transcript in the language spoken.
Model foundation
Falcon-ASR builds on Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.
Try Falcon ASR
The Hugging Face Demo Space lets users try Falcon-ASR and explore its transcription capabilities. API access and native applications are planned. The team invites users to try the Demo with their own recordings.
Acknowledgments
The team thanks the group behind the Falcon-Emirati model for their support with Arabic foundation models. Readers can find their latest work in the Falcon-Emirati blog post.
The team also thanks Mikhail Lubinets for continued support with the compute infrastructure.
What it means
Users working with Arabic audio now have a tool that handles dialectal variation better than previous public options. The system processes five languages with a single set of weights, removing the need to switch models or flags. Word-level timestamps allow for precise alignment of text with audio tracks, which aids in editing and analysis.




