Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Falcon-Emirati-7B scores 84.83% on Alyah, beating larger multilingual models on Emirati dialect benchmarks Arabic is a family of languages sharing one name,…

By Vane October 6, 2026 8 min read
Falcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Falcon-Emirati-7B scores 84.83% on Alyah, beating larger multilingual models on Emirati dialect benchmarks

Arabic is a family of languages sharing one name, yet Modern Standard Arabic is rarely how people speak. In the UAE, daily conversation relies on Emirati Arabic, a Gulf dialect with its own vocabulary and rhythm. Emirati poetry and proverbs carry meanings that vanish in a literal translation. A model trained only on Modern Standard Arabic can translate every word of an Emirati sentence and still miss the point.

Falcon-Emirati-7B is built to close that gap. It is a dialect-specialized model based on Falcon-H1-Arabic, designed to understand and generate Emirati Arabic as a native speaker would: the vocabulary, the tone, and the cultural context behind it.

Built on Falcon-H1-Arabic

The team did not start from scratch. Falcon-Emirati-7B sits on top of Falcon-H1-Arabic, a model family that set new benchmarks for the language earlier this year. Falcon-H1-Arabic uses a hybrid architecture combining State Space Models and Transformer attention running in parallel. Their outputs fuse before each block’s projection. This setup gives linear-time efficiency on long sequences while keeping the precision of attention for long-range dependencies, which matters for a morphologically rich language like Arabic.

The family spans three scales with context windows up to 128K and 256K tokens. It was already trained on a broad mix of Modern Standard Arabic and dialectal Arabic alongside English and multilingual data.

That provided a strong starting point: a model that understood Arabic broadly, handled long context well, and had some dialectal exposure baked in. Falcon-Emirati-7B takes that foundation and pushes it specifically toward the Emirati dialect, the vocabulary, the grammar, and the cultural knowledge that a general Arabic model does not pick up on its own.

The team built Falcon-Emirati-7B on the 7B variant. It is the sweet spot in the family: large enough to hold onto the nuance that dialect adaptation needs, but small enough that both training and inference stay practical. The 34B model would likely push quality a bit further, but at a training and serving cost that does not make sense for a dialect-specialized chat model. The 3B model does not leave enough headroom for the depth of cultural and linguistic understanding the team was after. 7B gave the best balance of quality against training and inference cost.

Why Dialect Adaptation Is Hard

Turning a general Arabic model into an Emirati-dialect specialist sounds like a smaller job than building the base model. It is not. A few things make it genuinely difficult.

  • Emirati is mostly a spoken dialect. It shows up far less in writing online than Modern Standard Arabic, or even other Gulf and Levantine dialects, so there is not as much raw text to learn from.
  • Meaning is often non-literal. Idioms, proverbs, and poetic references lean on shared cultural context, not surface vocabulary.
  • There is no established playbook. There is not a well-documented recipe for how much dialectal data is enough, how to mix it with Modern Standard Arabic and general Arabic, or which training stage matters most for picking up a dialect.

That last point shaped how the team worked. A lot of building Falcon-Emirati-7B came down to trial and error: testing different data mixes, training stages, and supervision strategies, and using both human judgment and benchmark scores to figure out what actually moved the needle.

Our Approach to Data

The team built a dedicated Emirati data pipeline on top of Falcon-H1-Arabic’s pretraining, drawing on three complementary sources.

1. Authentic Emirati-Dialect Web Data

The team crawled and curated content from Emirati websites and forums written natively in the dialect, not translated or transliterated from Modern Standard Arabic. This is where they got their ground truth: how Emiratis actually write and speak online, the everyday phrasing, the colloquial expressions, and the natural back-and-forth between Emirati and Modern Standard Arabic that shows up in real usage.

2. MSA Data About Emirati Culture and Identity

Alongside the dialectal text, the team pulled in Modern Standard Arabic-language material specifically about Emirati culture, heritage, and language: articles and references on local customs, values, history, and social norms, including how Emiratis are perceived and stereotyped. This does not teach the model to write in dialect, but it teaches the model what it is talking about when Emirati topics come up, things like heritage, etiquette, and the context a native speaker just knows.

3. Synthetic Data, Guided by Glossaries and Style Rules

Authentic dialectal text alone was not enough to cover the range of topics a chat model actually needs to handle day to day. So the team generated a large amount of synthetic Emirati-dialect data to fill the gaps. They did not just let a generator model improvise in Gulf-ish Arabic. They constrained it with strict rules and glossaries and dictionaries built specifically for Emirati vocabulary and grammar. Those guardrails made the difference between synthetic output that reads as authentically Emirati and output that is grammatically fine but sounds off to anyone who actually speaks the dialect.

Finding the Right Adaptation Recipe

Since there is no standard recipe for Modern Standard Arabic-to-dialect adaptation, the team treated the training strategy itself as something to figure out experimentally. They ran ablations on how much dialectal data to inject and at which stage of training, how to balance authentic crawled data against synthetic data without the model overfitting to synthetic patterns, and how much Modern Standard Arabic cultural context was actually needed to keep it culturally grounded rather than just fluent on the surface. At each step they leaned on a mix of automatic scoring and native-speaker review, since automatic metrics alone do not capture naturalness, tone, or cultural fit well enough to trust on their own.

Evaluation Methodology

The team tracked progress throughout training with two complementary approaches:

Manual Evaluation by Native Speakers

Emirati native speakers reviewed model outputs directly, judging not just whether an answer was correct but whether it sounded right: naturalness, tone, cultural appropriateness. These are the things a benchmark score will not tell you but a native ear catches immediately.

Automatic Evaluation on Alyah

For quantitative tracking, the team used Alyah, a benchmark they and the community released specifically to evaluate Emirati-dialect capability in Arabic LLMs. Alyah is a fully native multiple-choice benchmark of 1,173 samples, collected manually from native Emirati speakers and spanning categories from everyday greetings and etiquette to figurative language, heritage knowledge, and Emirati poetry: the categories where dialect and culture matter most and where generic Arabic models tend to struggle. Full details on Alyah’s construction and category breakdown are available in their benchmark blog post, and background on the base model family is available in the Falcon-H1-Arabic announcement.

Results

Falcon-Emirati-7B scores 84.83% on Alyah, ahead of every other Arabic and multilingual model they compared it against, including several models many times its size. The chart below shows where it lands next to a representative set of leading instruction-tuned models on the Alyah leaderboard.

Alyah accuracy, instruction-tuned models. Falcon-H1-Arabic family models are excluded from this comparison since Falcon-Emirati-7B is built on top of them.

What the Results Tell Us

A couple of things jump out from this comparison. Size alone does not buy you dialect competence. Some of the largest multilingual models score well below smaller, more dialect-aware ones, which tells you Emirati proficiency has to be trained for on purpose, not picked up as a side effect of scale. The models that do best also tend to be Arabic-native or Arabic-focused to begin with, which lines up with what the team saw during their own ablations: general Arabic and dialect coverage is a necessary starting point, but it still takes targeted, dialect-specific work to close the rest of the gap, particularly on the hardest parts of Alyah, like poetry, heritage knowledge, and the language-and-dialect category itself.

This also matches what came out of the Alyah benchmark release more broadly: even strong models show real degradation once you move into genuinely dialectal, culturally embedded content. That gap does not close on its own with bigger models. It takes data and evaluation built specifically for the dialect.

Beyond Multiple Choice: LLM-as-Judge Evaluation

Multiple-choice accuracy tells you whether a model can recognize the right answer among four options. It does not tell you whether the model will actually produce Emirati Arabic on its own when someone just talks to it. So alongside Alyah, the team ran a second evaluation: open-ended generation on the same 1,173 Alyah questions, scored by an LLM judge (Gemini 3.7 Flash) against five models, Falcon-Emirati-7B, ALLaM-7B-Instruct-preview, gemma-3-27b-it, Jais-2-8B-Chat, and Fanar-2-27B-Instruct, chosen as the strongest competing models from the Alyah leaderboard.

The judge scored each answer on two separate dimensions: whether the content was correct, and, independently, whether the answer actually came back in Emirati dialect rather than Modern Standard Arabic. The team reports both a partial-credit score and a stricter pass/fail version, plus how often each model abstained instead of answering.

LLM-judged correctness on the 1,173 Alyah questions, open-ended generation, Gemini 3.7 as judge.

LLM-judged dialect fidelity on the same questions: does the answer actually come back in Emirati, or does the model default to Modern Standard Arabic?

Falcon-Emirati-7B leads on correctness, but the real gap is in the second chart. On dialect fidelity, Falcon-Emirati-7B scores 0.52 against 0.05 for ALLaM, 0.03 for gemma-3-27b-it, 0.02 for Jais-2-8B-Chat, and effectively 0.00 for Fanar-2-27B-Instruct. That is not a small edge, it is close to two orders of magnitude at the low end. In practice, this means the other models often know the right answer but say it in Modern Standard Arabic by default, even when asked directly in Emirati. Falcon-Emirati-7B is the only one of the five that reliably answers back in the dialect it was asked in.

Fanar-2-27B-Instruct stands out for a second reason too: it abstains far more than any other model, declining to answer 26.2% of the time, versus under 5% for every

Scroll to Top