TII introduces Falcon-ASR for Arabic and Emirati speech
Technology Innovation Institute (TII) has introduced Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic, including the Emirati dialect. TII says it supports English, French, Spanish and Portuguese, offers word-level timestamps, and achieved a 20.92% average WER across six Arabic test sets in its evaluation.
Why it matters: It adds a multilingual speech model with a specific focus on dialectal Arabic and Emirati transcription.
Introducing Falcon ASR
Published October 7, 2026

Arabic WER: 20.92% · Parameters: 1.6B · Emirati WER (TII evaluation): 22.73%
We’re introducing Falcon-ASR, our 1.6 billion parameter speech recognition model for Arabic, with a particular focus on the Emirati dialect. Developed at the Technology Innovation Institute (TII) in Abu Dhabi, it also supports English, French, Spanish and Portuguese.
In our evaluation, Falcon-ASR achieved an average word error rate of 20.92% across six Arabic test sets, compared with the best published result of 23.17% in the leaderboard snapshot we used. On our internal Emirati evaluation, it recorded the lowest word and character error rates among the systems we compared.
We also support word-level timestamps for transcriptions, linking each transcribed word to its position in the audio.
You can try Falcon-ASR in our Hugging Face Demo.
Recognising spoken Arabic
Arabic speech varies by region, speaker and setting. A model that handles a formal news broadcast may still struggle with a conversation in Emirati or with speech recorded over a phone line. Dialectal Arabic also has fewer transcribed resources than Modern Standard Arabic (MSA), which makes training and evaluation harder.
We trained Falcon-ASR on Emirati, MSA, other Gulf and Arabic dialects, and English. Our aim is to transcribe the words people use in everyday speech, including dialectal forms and changes between languages.
Arabic benchmark results
The Open Universal Arabic ASR Leaderboard, maintained by the ELM Research Center, ranks systems by the equal-weight average WER across six test sets. It also reports character error rate (CER). Lower values are better for both metrics. Our Falcon-ASR evaluation follows this protocol.
| Model | Parameters | Avg WER (%) | Avg CER (%) |
|---|---|---|---|
| Falcon-ASR | 1.6B | 20.92 | 8.79 |
| Audar-ASR-V1-Turbo | 2.35B | 23.17 | 9.23 |
| Cohere Transcribe Arabic (07-2026) | 2.0B | 25.87 | 11.80 |
| omniASR LLM 7B | 7.0B | 28.32 | 12.52 |
WER = Word Error Rate; CER = Character Error Rate. A lower value indicates better performance.
We evaluated Falcon-ASR on the same six benchmarks using the leaderboard’s pinned manifests. Competitor figures are the published leaderboard averages checked on 30 September 2026. Falcon-ASR’s average WER is 2.25 percentage points better than the best published result in that snapshot.
Evaluating Emirati speech
Public evaluation data already includes Emirati: Casablanca has a UAE subset. We complement that coverage with an internal evaluation of additional Emirati and Gulf speech, using held-out recordings and human-validated transcripts to assess transcription accuracy beyond the public UAE subset.
In our internal Emirati evaluation, Falcon-ASR achieved 22.73% WER and 10.19% CER:
| Model | Parameters | WER (%) | CER (%) |
|---|---|---|---|
| Falcon-ASR | 1.6B | 22.73 | 10.19 |
| Qwen3-Omni-30B-A3B-Instruct | 30.0B (3.0B active) | 26.80 | 12.72 |
| Audar-ASR-V1-Turbo | 2.35B | 27.89 | 13.75 |
| Cohere Transcribe Arabic (07-2026) | 2.0B | 31.05 | 18.07 |
| Qwen3-ASR-1.7B-hf | 2.0B | 31.52 | 13.35 |
| Audar-ASR-V1-Flash | 0.78B | 32.87 | 15.36 |
Falcon-ASR has the lowest WER and CER among the systems compared here. Its WER is 4.07 percentage points below Qwen3-Omni, the next best result. The results show improved transcription accuracy at both the word and character level on this evaluation.
Training for different recording conditions
We included background noise, overlapping speech, music, room reverberation and telephony effects, as well as variations in speed and pitch. We applied the same treatment to Emirati recordings, exposing the model to a range of conditions it may encounter in meetings, calls and other everyday recordings.
English and other languages
Falcon-ASR also transcribes English with the same model weights. In our evaluation on the seven public English test sets used by the Hugging Face Open ASR Leaderboard, it achieved a mean WER of 5.74%.
| Test set | WER (%) |
|---|---|
| LibriSpeech clean | 1.75 |
| LibriSpeech other | 4.21 |
| SPGISpeech | 2.02 |
| VoxPopuli | 3.87 |
| GigaSpeech | 8.15 |
| AMI | 8.33 |
| Earnings-22 | 11.86 |
The model also supports French, Spanish and Portuguese. All five languages use the same weights, without requiring a language flag. The output is a transcript in the language spoken.
Model foundation
Falcon-ASR builds on our Falcon3-Audio work. The architecture and training approach for Falcon3-Audio are described in Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data.
Try Falcon ASR
Our Hugging Face Demo Space lets you try Falcon-ASR and explore its transcription capabilities. API access and native applications are planned. We invite you to try the Demo with your own recordings.
نقدّم Falcon ASR
نقدّم Falcon-ASR، نموذجنا للتعرف على الكلام العربي بحجم 1.6 مليار معلمة، مع اهتمام خاص باللهجة الإماراتية. طوّرنا النموذج في معهد الابتكار التكنولوجي (TII) في أبوظبي، وهو يدعم أيضًا اللغات الإنجليزية والفرنسية والإسبانية والبرتغالية.
في تقييمنا، حقق Falcon-ASR متوسط معدل خطأ في الكلمات بلغ 20.92% عبر ست مجموعات اختبار باللغة العربية، مقارنةً بأفضل نتيجة منشورة بلغت 23.17% في نسخة لوحة المتصدرين التي استخدمناها. وفي تقييمنا الداخلي للهجة الإماراتية، سجّل النموذج أدنى معدلات خطأ في الكلمات والأحرف بين الأنظمة التي قارناها.
ندعم أيضاً الطوابع الزمنية على مستوى الكلمات في عمليات التفريغ النصي، بحيث ترتبط كل كلمة بموضعها في التسجيل الصوتي.
يمكنكم تجربة Falcon-ASR عبر عرضنا التجريبي على Hugging Face.
التعرف على العربية المنطوقة
يختلف الكلام العربي باختلاف المنطقة والمتحدث وظروف التسجيل. فقد ينجح نموذج في تفريغ نشرة إخبارية رسمية، ثم يجد صعوبة في تفريغ محادثة باللهجة الإماراتية أو تسجيل عبر الهاتف. كما أن الموارد الصوتية المفرّغة نصيًا للهجات العربية أقل من تلك المتاحة للعربية الفصحى، مما يزيد صعوبة التدريب والتقييم.
درّبنا Falcon-ASR على اللهجة الإماراتية والعربية الفصحى ولهجات خليجية وعربية أخرى، إلى جانب الإنجليزية. وهدفنا هو تفريغ الكلمات التي يستخدمها الناس في حديثهم اليومي، بما في ذلك الصيغ اللهجية والانتقال بين اللغات.
نتائج الاختبارات العربية
ترتّب لوحة المتصدرين المفتوحة الشاملة للتعرف على الكلام العربي، التي يديرها مركز ELM للأبحاث، الأنظمة وفق متوسط معدل خطأ الكلمات عبر ست مجموعات اختبار، بوزن متساوٍ لكل مجموعة. وتعرض أيضًا معدل خطأ الأحرف (CER). وكلما انخفضت قيمة أي من المقياسين، كان الأداء أفضل. ويتبع تقييمنا للنموذج هذا البروتوكول.
| Model | Parameters | Avg WER (%) | Avg CER (%) |
|---|---|---|---|
| Falcon-ASR | 1.6B | 20.92 | 8.79 |
| Audar-ASR-V1-Turbo | 2.35B | 23.17 | 9.23 |
| Cohere Transcribe Arabic (07-2026) | 2.0B | 25.87 | 11.80 |
| omniASR LLM 7B | 7.0B | 28.32 | 12.52 |
WER هو معدل خطأ الكلمات؛ وCER هو معدل خطأ الأحرف. تشير القيمة الأقل إلى أداء أفضل.
قيّمنا Falcon-ASR على مجموعات الاختبار الست نفسها، باستخدام قوائم العينات المثبّتة في لوحة المتصدرين. وأرقام النماذج المنافسة هي المتوسطات المنشورة في اللوحة، والتي جرى التحقق منها في 30 سبتمبر 2026. وكان متوسط خطأ الكلمات للنموذج أفضل بمقدار 2.25 نقطة مئوية من أفضل نتيجة منشورة في تلك النسخة.
تقييم اللهجة الإماراتية
تتوافر بالفعل بيانات عامة لتقييم اللهجة الإماراتية، إذ تضم مجموعة Casablanca قسمًا خاصًا بالإمارات. ونكمّل هذه التغطية بتقييم داخلي لتسجيلات إضافية من الكلام الإماراتي والخليجي، باستخدام تسجيلات مخصّصة للاختبار ونصوص مرجعية خضعت لمراجعة بشرية، لتقييم دقة التفريغ على مواد إضافية إلى جانب البيانات الإماراتية العامة.
حقق Falcon-ASR في تقييمنا الإماراتي الداخلي معدل خطأ كلمات قدره 22.73٪ ومعدل خطأ أحرف قدره 10.19٪:
| Model | Parameters | WER (%) | CER (%) |
|---|---|---|---|
| Falcon-ASR | 1.6B | 22.73 | 10.19 |
| Qwen3-Omni-30B-A3B-Instruct | 30.0B (3.0B active) | 26.80 | 12.72 |
| Audar-ASR-V1-Turbo | 2.35B | 27.89 | 13.75 |
| Cohere Transcribe Arabic (07-2026) | 2.0B | 31.05 | 18.07 |
| Qwen3-ASR-1.7B-hf | 2.0B | 31.52 | 13.35 |
| Audar-ASR-V1-Flash | 0.78B | 32.87 | 15.36 |
سجّل Falcon-ASR أقل معدل لخطأ الكلمات والأحرف بين الأنظمة المقارَنة هنا. وكان معدل خطأ الكلمات أقل بمقدار 4.07 نقطة مئوية من Qwen3-Omni، صاحب النتيجة التالية. وتُظهر النتائج تحسنًا في دقة التفريغ على مستوى الكلمات والأحرف في هذا التقييم.
التدريب على ظروف تسجيل مختلفة
ضمّنا بيانات التدريب ضوضاء خلفية وكلامًا متداخلًا وموسيقى وصدى الصوت وتأثيرات الاتصالات الهاتفية، إلى جانب تغيّرات في سرعة الكلام وحدّة الصوت. وطبّقنا المعالجة نفسها على التسجيلات الإماراتية، لتهيئة النموذج للتعامل مع ظروف مختلفة قد يواجهها في الاجتماعات والمكالمات والتسجيلات اليومية الأخرى.
الإنجليزية واللغات الأخرى
يفرّغ Falcon-ASR الكلام الإنجليزي باستخدام أوزان النموذج نفسها. وفي تقييمنا على مجموعات الاختبار الإنجليزية العامة السبع المستخدمة في لوحة Hugging Face المفتوحة للتعرف على الكلام، بلغ متوسط معدل خطأ الكلمات 5.74٪.
| Test set | WER (%) |
|---|---|
| LibriSpeech clean | 1.75 |
| LibriSpeech other | 4.21 |
| SPGISpeech | 2.02 |
| VoxPopuli | 3.87 |
| GigaSpeech | 8.15 |
| AMI | 8.33 |
| Earnings-22 | 11.86 |
يدعم النموذج أيضًا الفرنسية والإسبانية والبرتغالية. وتستخدم اللغات الخمس أوزانًا واحدة، دون الحاجة إلى تحديد اللغة مسبقًا. ويكون الناتج تفريغًا نصيًا باللغة المنطوقة.
أساس النموذج
يستند Falcon-ASR إلى أعمالنا في Falcon3-Audio. وتعرض ورقة Competitive Audio-Language Models with Data-Efficient Single-Stage Training on Public Data بنية Falcon3-Audio ونهج تدريبه.
تجربة Falcon ASR
يتيح عرضنا التجريبي على Hugging Face تجربة Falcon-ASR واستكشاف قدراته في تفريغ الكلام. أما الوصول عبر واجهة API والتطبيقات الأصلية فهو مخطّط له. ندعوكم إلى تجربة النموذج باستخدام تسجيلاتكم.
More from this author
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance
20
October 6, 2026Falcon OCR Arabic: 270M Parameters State-of-the-Art Arabic OCR
13
October 6, 2026Community
Upload images, audio, and videos by dragging in the text input, pasting, or clicking here.
Tap or paste here to upload images
How we got here
- A taxonomy aims to make AI agent sandbox security measurableTechRadar · Hugging Face
- Transformers v5.19.0 adds EmbeddingGemma 2 supportTransformers Releases · Hugging Face
- Insurers Prepare for AI Agent Liability ClaimsThe Decoder · Hugging Face
- Interconnects: open-weight cyber-risk discourse is brokenInterconnects · Hugging Face
- Why AI sandbox escapes are inevitable: four incidents, a defense checklist钛媒体 · Hugging Face
- OpenAI Brakes on Agent Safety as Meta Keeps Pushing虎嗅 AI · Hugging Face