Back to The Maya Lab
Research Ananya Krishnamurthy

Hindi and Urdu Script Divergence: How One Spoken Language Becomes Two TTS Voices

Spoken Hindi and Urdu are nearly identical, but their scripts are completely different. We explain how we handle this divergence in our TTS models and the phoneme-level decisions involved.

Hindi and Urdu share a spoken root. Linguists call it Hindustani: a spoken language mutually intelligible across communities that refer to it by different names, with some vocabulary differences but largely the same phonology, morphology, and syntactic patterns. A Hindi speaker and an Urdu speaker conducting a conversation have no difficulty understanding each other.

The scripts are completely different. Hindi uses Devanagari; Urdu uses Nastaliq, an Arabic-script calligraphic style written right-to-left. And the divergence is not just visual. The two scripts have different orthographic conventions that encode the same underlying spoken language differently, and those differences create specific challenges for TTS pipelines.

Where the orthographic conventions diverge

Vowel representation is the most significant difference. Devanagari is an abugida: consonants carry an inherent vowel sound (a short /a/) that is modified by explicit diacritic marks. The inherent vowel can also be deleted in specific phonological contexts (schwa deletion), and this deletion must be inferred from context because it is not written. Short vowels in Devanagari are typically written with diacritics (matras), though in common Hindi text, short interior vowels are often omitted and must be inferred.

Urdu in Nastaliq script uses the Arabic-script vowel system: consonants are written without inherent vowels, and short vowels (harakat: zabar, zer, pesh) are diacritics that are optional in most written text. In everyday Urdu writing, these short vowel marks are almost never written. Readers and TTS systems must infer vowel values from context, morphological knowledge, and the lexicon.

This means both scripts require vowel inference, but the inference rules operate differently and the ambiguity patterns are different. A Devanagari G2P system trained on Hindi cannot be straightforwardly applied to Nastaliq Urdu text, even though the underlying spoken output would be nearly identical.

Vocabulary divergence and its TTS implications

The formal vocabularies of Hindi and Urdu diverge significantly, reflecting their different literary traditions. Formal Hindi draws heavily on Sanskrit-origin (tatsama) vocabulary. Formal Urdu draws heavily on Persian and Arabic vocabulary. In practice, contemporary conversational Hindustani mixes both freely, but formal texts and domain-specific content lean strongly toward one tradition or the other.

For TTS, this matters because a Hindi G2P model trained on Sanskrit-origin vocabulary may handle Persian and Arabic loanwords poorly. The phonological rules differ: Arabic phonemes like the voiced pharyngeal fricative (ayin) and the emphatic consonants do not exist in the Sanskrit phonological tradition. When Urdu text includes words of Arabic origin with these sounds, a Hindi G2P system may produce incorrect phoneme sequences.

Our approach separates the G2P pipelines for Hindi and Urdu at the language level, even though the acoustic model shares substantial weights. The Urdu G2P includes Persian and Arabic phoneme handling; the Hindi G2P is optimized for the Sanskrit-dominant phonological inventory. Both systems share a common phoneme representation that can be fed to the same acoustic model, allowing weight sharing where the underlying speech patterns are similar.

Nastaliq rendering for developer input text

A practical challenge for developers building Urdu TTS applications is text input handling. Nastaliq Arabic script requires right-to-left text direction and specific Unicode character handling. Arabic-script languages use Unicode's Bidirectional Algorithm, and incorrect handling of text direction can produce reversed character order or incorrect ligature formation before the text even reaches the G2P layer.

When integrating Urdu synthesis, ensure your text pipeline handles the following correctly:

  • Right-to-left Unicode text direction (use Unicode bidirectional control characters correctly, or ensure your text is stored in logical order rather than visual order)
  • Normalization of Arabic character variants: certain Arabic letters have multiple Unicode representations (with and without specific diacritics, with and without hamza). Normalize to canonical forms before sending to the API.
  • Handling of mixed-direction content: Urdu text with embedded Roman numerals or Latin-script brand names requires careful bidirectional handling. Test these cases explicitly.

Our API accepts Nastaliq Urdu text directly and handles normalization internally for common cases, but text that has been incorrectly handled upstream (reversed byte order, character form confusion) will produce incorrect synthesis. Garbage in, garbage out applies to script encoding as much as to content.

One model or two?

A common question from developers building bilingual Hindi/Urdu applications is whether they need separate model configurations or can use a single endpoint. Our answer is: separate G2P pipelines are required; separate acoustic model checkpoints are optional depending on quality requirements.

The G2P pipeline must be separate because the script conventions and phoneme inference rules are different enough that a single pipeline cannot reliably handle both. Using Hindi G2P on Urdu text produces errors on Persian-origin vocabulary; using Urdu G2P on Hindi text may handle sandhi junctions and schwa deletion less reliably.

The acoustic model can be shared if you are primarily targeting the Hindustani spoken intersection (conversational language that is natural in both traditions). If you are targeting formal Urdu with significant Persian/Arabic vocabulary, a separate acoustic model fine-tuned on Urdu speech data produces noticeably better results on that vocabulary. We provide both configurations: a shared acoustic model that works acceptably for mixed Hindustani content, and Urdu-specific acoustic weights for applications where Urdu formal register quality is a priority.

Script as identity, not just encoding

It is worth noting that for many speakers, the choice of script is not a technical decision but a cultural and sometimes political one. Requesting Hindi synthesis using Nastaliq script or Urdu synthesis using Devanagari would feel wrong to many users, not because the output would necessarily be unintelligible, but because script identity is part of how communities relate to their language. Build your application accordingly: if you are targeting Hindi-speaking users, use Devanagari input; if you are targeting Urdu-speaking users, use Nastaliq input. Do not conflate the scripts just because the spoken output is similar.

Try the Maya Research API
Stream multilingual TTS in 5 minutes. Free tier, no credit card required.
Get API Key

More from the lab