Back to The Maya Lab
Research Dheemanth Reddy

Measuring Phoneme Accuracy in Devanagari TTS: Our Internal Benchmark Setup

Phoneme error rate tells you more about TTS quality than MOS scores for morphologically-rich languages. Here is how we measure it for Hindi and Marathi.

When we started building TTS for Hindi and Marathi, we evaluated model quality the same way most teams do: Mean Opinion Score (MOS) from listening panels. We still run MOS evaluations, but we have come to rely more heavily on phoneme error rate as our primary diagnostic metric. This post explains why, and walks through the benchmark setup we use internally.

Why MOS is insufficient for Indic languages

MOS is a perceptual metric. Listeners score synthesized speech on a 1-5 scale for naturalness, and you average the scores. It works well for English because English phoneme coverage is relatively forgiving: mispronouncing one phoneme in a common English word usually still sounds acceptable to a naive listener.

Devanagari script encodes a 46-phoneme consonant inventory plus 16 vowels in standard Hindi. Critically, several phoneme distinctions that are contrastive (meaning they change word meaning) are also perceptually subtle to listeners unfamiliar with the language. An evaluator from outside the target language community might rate a sentence as 4.2 MOS even when the model has systematically confused aspirated and unaspirated stops. A Hindi native speaker would notice immediately.

More practically: we found that our early model versions scored 3.9-4.1 MOS on Hindi but had measurable confusion rates on the retroflex consonant series (ta ta tha tha da da dha dha na) because our training data did not have sufficient coverage of retroflex-heavy words in natural sentence contexts. MOS did not surface this; phoneme error rate did.

Our benchmark corpus structure

We maintain a test corpus of 1,200 sentences for Hindi and 800 for Marathi. The sentences are constructed to specifically probe challenging phonological contexts rather than being sampled randomly from text data. Coverage dimensions include:

  • Schwa deletion contexts: Hindi orthography writes inherent vowels that are silent in specific phonological environments. Correct deletion of the inherent schwa in word-final position and before consonant clusters is required for intelligibility.
  • Aspirated vs. unaspirated stops: ka/kha, ga/gha, ca/cha, ja/jha, ta/tha, da/dha, pa/pha, ba/bha. Each contrastive pair is tested in word-initial, word-medial, and word-final positions.
  • Retroflex vs. dental consonants: This is where non-specialist TTS systems tend to fail most visibly on Hindi. We have 140 sentence pairs that differ only in retroflex vs. dental articulation.
  • Anusvara and visarga handling: The nasalization dot (anusvara) is assimilated differently depending on the following consonant's place of articulation. The aspiration marker (visarga) at word ends varies by dialect.
  • Conjunct consonant clusters: Standard Hindi has 35+ commonly used conjuncts. We test each in at least 4 sentence contexts.
  • Numerals and mixed-script fragments: A significant portion of real-world Hindi TTS input includes embedded English words, numerals in Devanagari or Arabic script, and punctuation patterns that affect prosody.

The evaluation pipeline

We synthesize all test sentences with the model under evaluation. We then run the audio through an ASR system trained specifically on Hindi and Marathi (we use a Wav2Vec2-based model fine-tuned on a clean newsreading corpus) to transcribe the synthesized audio back to text. The ASR transcription is then aligned with the original text at the phoneme level using forced alignment, and we compute phoneme error rate as the number of substitutions, deletions, and insertions divided by the total number of reference phonemes.

This TTS-ASR loop approach is not a perfect measure of perceptual quality. ASR systems have their own error patterns, and a phoneme that the ASR gets wrong does not necessarily sound wrong to a human listener. We treat the metric as a diagnostic tool, not a ground truth. The value is in tracking direction over model versions and spotting systematic confusion patterns rather than producing a single quality number.

Current numbers and what they reveal

On our current production Hindi model (v2.3 in our internal versioning), phoneme error rate on the benchmark corpus is 4.1%. For context, a human speaker reading the same sentences through the same ASR pipeline scores around 2.8%, so we are targeting a ceiling around that value. The gap between our model and human performance concentrates heavily in three categories: retroflex consonants (accounting for 38% of our errors), anusvara assimilation (21%), and schwa deletion at morpheme boundaries (18%).

Marathi is harder. Our Marathi model (v1.8) has a 7.3% phoneme error rate on the equivalent corpus. Marathi has additional phonological features including a retroflex lateral consonant and distinctive vowel length contrasts that Hindi does not have in the same form. Our training data for Marathi is roughly 40% the size of our Hindi corpus, and that data imbalance is visible in the numbers. We are actively collecting more Marathi data from native speaker recordings, targeting 200 additional hours of read speech.

Boundary conditions on this approach

We want to be precise about what this methodology measures and what it does not. Phoneme error rate in the TTS-ASR loop measures articulatory accuracy, not prosody. A model can produce every phoneme correctly but place incorrect stress, flatten intonation contours, or produce unnatural rhythm at clause boundaries. Prosody quality requires separate evaluation: we use a combination of pitch contour comparison (against reference speaker recordings) and human listener studies focused specifically on naturalness and sentence-level rhythm.

We are also not claiming that phoneme error rate is the right primary metric for all TTS use cases. For very short, formulaic output (OTPs, order totals, navigation prompts), users tolerate a wider range of phoneme accuracy because context is strong. For long-form narration, educational content, or news reading, phoneme accuracy matters significantly more. The benchmark setup described here is optimized for the conversational and narration use cases that most of our early-access partners are building.

Making the benchmark available

We are planning to release the Hindi phoneme benchmark corpus as an open dataset. The sentences were written by our team with native speaker review; they are not scraped from web text and are designed to be clean enough for TTS evaluation without copyright issues. If you are building TTS for Indic languages and want to use a consistent evaluation harness, we expect to have a public release ready by the end of this quarter. The Marathi corpus will follow once we are satisfied with the coverage of the additional phonological features unique to that language.

Try the Maya Research API
Stream multilingual TTS in 5 minutes. Free tier, no credit card required.
Get API Key

More from the lab