MikhbarMIKHBAR
Artificial Intelligence

Falcon-Emirati-7B Launched for Emirati Dialect and Culture

A new language model designed to understand and generate Emirati Arabic with native-like nuance has been detailed in the Hugging Face Blog.

Falcon-Emirati-7B Launched for Emirati Dialect and Culture

Bridging the Gap in Arabic Dialects

Arabic encompasses a wide family of distinct languages and spoken variations under a single name. While Modern Standard Arabic appears in formal texts and news broadcasts, day-to-day conversation, humor, and storytelling rely heavily on regional expressions. In the UAE, daily communication occurs through Emirati Arabic, a Gulf dialect featuring its own vocabulary, distinct rhythm, and deep cultural ties. Traditional elements like nabati poetry, proverbs, and local anecdotes carry meanings that are frequently missed by literal, word-for-word translations from standard Arabic. Further details are available from Hugging Face Blog in the original source material.

To close this linguistic gap, developers introduced Falcon-Emirati-7B. Detailed extensively via the Hugging Face Blog, this dialect-specialized model operates on top of existing Arabic frameworks to understand and generate Emirati Arabic with the precision of a native speaker, capturing proper tone, vocabulary, and cultural context.

Architecture and Foundation

Falcon-Emirati-7B is not built from scratch; instead, it leverages the Falcon-H1-Arabic model family, which previously established new benchmarks for the language. The underlying Falcon-H1 hybrid architecture runs State Space Models, specifically Mamba, alongside Transformer attention in parallel within every block. The outputs are then fused before each block's projection.

This hybrid design delivers the linear-time efficiency of Mamba for handling long sequences while maintaining the analytical precision of attention mechanisms for long-range dependencies. Such capabilities are crucial when processing a morphologically rich language like Arabic. The model family spans three different scales—3B, 7B, and 34B parameters—supporting context windows up to 128K and 256K tokens, and was initially trained on a mixture of Modern Standard Arabic, various regional dialects including Gulf, Levantine, Egyptian, and Maghrebi, alongside English and multilingual data.

Selecting the Optimal Model Scale

Among the available scales within the model family, the developers selected the 7B variant as the ideal foundation for Falcon-Emirati-7B. This specific size represents a practical sweet spot for development and deployment.

The 7B parameter count provides sufficient capacity to retain the intricate nuances required for dialect adaptation, while remaining compact enough to ensure that both training and inference stay practical. While a larger 34B model might yield marginal quality improvements, its significantly higher training and serving costs are impractical for a dedicated dialect-focused chat model. Conversely, the smaller 3B variant lacks the necessary headroom to capture deep cultural and linguistic understanding.

Overcoming Dialect Data Challenges

Adapting a general Arabic model into an Emirati-dialect specialist presents unique hurdles. Because Emirati is predominantly a spoken dialect, it appears far less frequently in online written formats compared to Modern Standard Arabic or other regional Gulf and Levantine variants, resulting in a scarcity of raw training text.

Furthermore, dialect meaning often relies on non-literal interpretations. Idioms, proverbs, and poetic references lean heavily on shared cultural understanding rather than surface-level vocabulary. Because no established playbook exists for MSA-to-dialect adaptation—such as determining optimal data mix ratios or training stages—the development process required extensive trial and error with data mixes and supervision strategies.

Building a Dedicated Emirati Data Pipeline

To address data scarcity, creators constructed a dedicated Emirati data pipeline operating on top of the pretraining data. This pipeline draws from three complementary sources to ensure comprehensive coverage of the dialect and its cultural backdrop.

The first source involves crawled and curated content sourced natively from Emirati websites and forums written directly in the dialect, avoiding MSA translations. This provides ground truth regarding everyday phrasing and colloquial expressions. The second source incorporates Modern Standard Arabic material focused specifically on Emirati heritage, customs, values, history, and social etiquette, supplying the contextual knowledge a native speaker naturally possesses.

Finally, the team utilized synthetic data generated under strict constraints. Using specialized dictionaries, glossaries, and style rules, the generation process was kept within rigid guardrails to prevent the output from sounding artificial or grammatically awkward to native speakers.

Evaluation and Native Review

Tracking progress throughout the training process required a dual approach combining quantitative benchmarks with qualitative human evaluation. Automated metrics alone cannot adequately capture conversational naturalness, tone, or cultural fit.

To ensure high quality, Emirati native speakers directly reviewed model outputs to judge naturalness and cultural appropriateness. For quantitative tracking, developers utilized Alyah, a specialized native multiple-choice benchmark featuring 1,173 samples collected from native speakers. Covering categories from everyday greetings to figurative language, heritage knowledge, and poetry, Alyah evaluates the specific areas where generic Arabic models traditionally struggle.

Sources

  • Hugging Face BlogFalcon-Emirati: When an LLM Learns the Dialect, the Culture, and the Nuance

Continue chronologically