We ran a blind listening study comparing TTS output quality across nine Indic languages. The results were useful internally and surprising in places. This post covers the methodology, what we found, and what the variance across languages tells us about where the hard problems actually are.
Study setup
We recruited 120 listeners, stratified by native language. For each language in the study, we recruited at least 12 listeners with that language as their primary or dominant language. Listeners evaluated synthesized speech without knowing which system produced it, on a 5-point naturalness scale. Each listener evaluated only the languages they are native or near-native in; we did not ask Hindi speakers to evaluate Tamil synthesis.
We used 50 test sentences per language, drawn from our internal evaluation corpus. Sentences were balanced across three registers: conversational, informational (news-style), and instructional (step-by-step directions). We included both short sentences (under 10 words) and long sentences (15-25 words) to surface prosody differences at clause boundary handling.
We evaluated our current production models across nine languages: Hindi, Tamil, Telugu, Kannada, Malayalam, Bengali, Marathi, Gujarati, and Punjabi. We are not presenting this as a comparison against other systems in this study; that comparison is in a separate benchmark post. This is about understanding the variance within our own system.
Results summary
The spread was larger than we expected. Hindi scored highest at 4.1 MOS. Tamil was second at 3.9. Telugu, Kannada, and Malayalam clustered between 3.6 and 3.8. Bengali scored 3.5. Marathi, Gujarati, and Punjabi were at 3.2, 3.1, and 2.9 respectively.
The pattern roughly follows training data availability, but not exactly. Hindi and Tamil have the largest high-quality training corpora we have assembled. Malayalam scores higher than its training corpus size would predict, and Bengali scores lower despite having more transcribed training data than Malayalam. This is the kind of discrepancy that motivates deeper analysis.
Why Malayalam outperforms its training data size
Malayalam has an unusually regular relationship between its script and phonological surface forms. While it is morphologically complex (agglutinative with long compound words), the Brahmi-derived Malayalam script has a very consistent letter-to-phoneme correspondence. Schwa insertion and deletion rules are more predictable than in Hindi, and there are fewer consonant cluster ambiguities that require context-dependent G2P resolution.
For TTS, this means the G2P layer introduces fewer errors, which means the acoustic model works with cleaner phoneme sequences, which means the prosody model has a more consistent input. The compounding of G2P accuracy into acoustic model quality is often underestimated. Malayalam's script regularity partially compensates for its smaller training corpus.
Why Bengali underperforms its training data
Bengali phonology includes a set of vowel distinctions that are rapidly neutralizing in modern spoken Bengali: historical distinctions between inherent vowel sounds that are maintained in the orthography but largely merged in contemporary speech. This creates a challenge for TTS: training data sourced from news reading and formal speech will preserve the historical distinctions, but a contemporary speaker will hear the synthesis as overly formal or slightly archaic.
Additionally, the Bengali retroflex series is realized differently from Hindi retroflexes in a way that our shared Indic phoneme representation does not fully capture. The phonetic implementation borrowed from our Hindi model was insufficiently adapted for Bengali-specific articulation. This is a known limitation we are actively working to address in the next model version.
The performance gap at sentence boundaries
Across all nine languages, listeners gave lower scores on long sentences with multiple embedded clauses compared to short sentences of equivalent word type. This is a prosody problem. TTS systems trained primarily on short utterances tend to produce flat intonation contours on long sentences; they do not model the pitch reset and phrase-level prosodic structure that native speakers use to organize complex sentences.
The gap between short-sentence and long-sentence MOS scores was largest for Tamil (0.6 MOS points), smallest for Hindi (0.3 points). Tamil has a highly regular sentence-final prosody that is relatively easy to model on short sentences, but the intonation structure of embedded relative clauses in Tamil is complex and our current model does not handle it as well. This is a specific area of active work for us.
What these numbers mean for language selection decisions
Developers often ask which languages are "production ready" in our system. We do not love that framing, but the MOS data gives you a way to calibrate. Hindi, Tamil, and Telugu-Kannada-Malayalam are reliably above 3.5 MOS on informational content, which is the threshold at which most listeners in our study described synthesis as "acceptable for a product." Bengali and Marathi are close to that threshold. Gujarati and Punjabi are below it for long-form content.
That said, MOS on our evaluation corpus is not the same as quality on your specific content. Formulaic content (OTPs, order confirmations, navigation prompts) scores higher than our corpus average across all languages because the sentences are short and predictable. Long-form narration and conversational synthesis are where the gap between strong and weaker models is most visible. Match your evaluation to your actual content type before making a production decision.
What comes next
The clearest action items from this study are: Bengali prosody adaptation (specifically the vowel contrast issue), Punjabi corpus expansion (we need at least 50% more clean training audio), and Tamil clause-boundary prosody. These three will absorb most of our model improvement effort over the next six months. We expect to rerun this study with updated models and share the delta.