A voice application that sounds natural in Cairo does not automatically sound natural in Beirut. And both of those will sound formal and stilted to everyday speakers if you synthesize them in Modern Standard Arabic. Yet most TTS providers for Arabic offer exactly one variant: MSA. This post explains why that is a problem, how we approached covering three variants, and what the training data challenges look like in practice.
The Arabic diglossia problem for TTS
Arabic is a diglossic language. Modern Standard Arabic (MSA, also called Fusha) is the formal register used in news broadcasts, official documents, religious text, and formal education. Virtually no one speaks MSA as their first or home language. The spoken varieties that people actually use daily are the regional dialects: Egyptian Arabic (Masri), Levantine Arabic (spoken across Syria, Lebanon, Jordan, and Palestine with sub-variants), Gulf Arabic, Moroccan Darija, and others.
The gap between MSA and spoken dialects is not accent-level; it is closer to the gap between Latin and modern Italian. Phonology, morphology, vocabulary, and sentence structure all differ substantially. A TTS system trained on MSA data will produce speech that a native Egyptian Arabic speaker recognizes as correct but immediately identifies as formal, written-language speech. For a banking IVR or a consumer application, that register mismatch creates friction. Users perceive it as cold or bureaucratic.
Why we chose three variants specifically
We cover MSA, Egyptian Arabic, and Levantine Arabic. This is a deliberate scope decision, not a limitation we are working to eliminate. Here is the reasoning.
MSA is non-negotiable for news, government, and any application targeting pan-Arab audiences. Egyptian Arabic is the most widely understood spoken dialect across the Arab world, partly due to the reach of Egyptian cinema and television. If you can only ship one spoken dialect, Egyptian has the broadest comprehensibility. Levantine Arabic represents the second major cluster and is specifically important for applications targeting the Mashreq region (Lebanon, Syria, Jordan). Gulf Arabic (Khaleeji) is our next planned addition; we have training data collection underway but have not reached the quality bar we require for public release.
We are not claiming that covering three variants solves the dialectal TTS problem. It is a starting point. Moroccan Darija, for instance, is substantially different from eastern dialects and is not covered by our Egyptian model. Users in Morocco synthesizing with our Egyptian variant will notice the mismatch.
Training data challenges
Annotated speech data for spoken Arabic dialects is scarce compared to MSA and is substantially scarcer than equivalent data for European languages. The challenges break down into several categories.
Orthographic inconsistency is the largest problem. MSA has a standardized orthography; Egyptian and Levantine Arabic do not. Written dialect text on the internet uses idiosyncratic spelling conventions that vary by writer. The word "what" in Egyptian Arabic might be spelled four different ways in the same corpus. This creates severe challenges for G2P training because the mapping from text to phoneme is inconsistent.
We addressed this by building normalization layers specific to each dialect that map common spelling variants to a standardized phonemic representation. This required manual review of roughly 15,000 high-frequency word forms for Egyptian Arabic and a similar number for Levantine. The normalization layer is one of our key engineering investments and something we have not seen adequately addressed in public Arabic TTS systems.
Code-switching is the second challenge. Real Egyptian Arabic speech frequently embeds English words, especially in technical and commercial contexts. A synthesized sentence about a mobile app in Egyptian Arabic might include "download," "update," or "settings" in English pronunciation. Our model needs to handle these correctly without switching to a full English voice register for the embedded word. We train on code-switched examples and have a lightweight language identification pass that routes phoneme prediction appropriately for embedded foreign words.
Phonological differences that matter for TTS quality
A few specific phonological differences between variants that TTS models must handle correctly:
The phoneme represented by the letter qaf (ق) in MSA is a uvular stop. In Egyptian Arabic, this sound is dropped or replaced by a glottal stop in most words. In Levantine Arabic it becomes a glottal stop in urban speech but remains uvular in some rural dialects. Getting this right is the single most important marker of dialectal authenticity for Arabic TTS.
The phoneme represented by the letter jim (ج) is a voiced palato-alveolar affricate in MSA and Levantine. In Egyptian Arabic, it is pronounced as a voiced velar stop (similar to the g in "get"). This single change affects hundreds of common words and is immediately noticeable when wrong.
Vowel length contrast is phonemic in MSA (long vs. short vowels change meaning) and is maintained in formal Egyptian Arabic but is reduced in casual Levantine speech. Our models are trained with explicit vowel length annotations in the phoneme layer rather than relying on the acoustic model to learn length implicitly.
How the three-variant architecture works at the API level
The language parameter in our API accepts ar-MSA, ar-EG, and ar-LEV. Each maps to a different acoustic model checkpoint and a different G2P pipeline. The vocoder is shared across variants, which keeps our inference infrastructure simpler and reduces latency from switching.
Input text normalization runs dialect-specific rules before G2P. If you pass Egyptian dialect text to the MSA endpoint, the output will sound formally correct but will have MSA pronunciation of dialect-specific words. If you pass MSA text to the Egyptian endpoint, the model will apply the qaf and jim substitutions that characterize Egyptian speech, which may sound inconsistent. We recommend matching input text variety to the endpoint variety for best results.
What good dialectal coverage actually sounds like to users
The best feedback we have received from early-access developers is when they report that their users stopped commenting on the TTS voice at all. A customer service IVR in Egyptian Arabic that sounds like MSA reads back generated text attracts comments. One that sounds like a natural Egyptian speaker becomes invisible in the interaction. Getting TTS to the point where it stops being noticed is the target, and dialect accuracy is the biggest lever for Arabic applications.