Back to The Maya Lab
Research Ananya Krishnamurthy

Building TTS for Low-Resource Indic Languages: Data, Transfer, and Tradeoffs

Languages like Punjabi and Gujarati have far less training data than Hindi or Tamil. We describe how we use cross-lingual transfer learning to bootstrap reasonable quality without massive corpora.

Punjabi, Gujarati, and Marathi have substantial speaker populations but relatively limited transcribed speech data compared to Hindi. Building usable TTS for these languages requires a different approach than scaling up what works for Hindi. This post describes our strategy and the tradeoffs we made.

Defining "low resource" in the Indic TTS context

Low resource is relative. In global TTS terms, Hindi has substantial data. In Indic TTS terms, Punjabi has dramatically less than Hindi. Our Hindi training corpus is approximately 800 hours of clean, transcribed speech. Punjabi is around 120 hours, Gujarati around 90 hours, and early Marathi data was around 180 hours (we have added more since). These numbers are not a complaint; they reflect the reality of what has been recorded and published for research use.

The quality ceiling for TTS is roughly correlated with training data size, but not linearly. The first 50 hours of clean data produce the largest quality gain per hour. Hours 50-200 produce diminishing but still meaningful improvement. Beyond 300 hours, gains are smaller and the model starts to be limited by other factors: phonological coverage of rare contexts, prosody model training data, vocoder quality.

For Punjabi at 120 hours, we are not in catastrophically low-resource territory. We can train a model that works well on common vocabulary and natural sentence structures. The challenge is coverage of uncommon phoneme contexts and naturalness on content far from the training distribution.

Transfer learning from Hindi as a starting point

Hindi and Punjabi are closely related Indo-Aryan languages with significant phonological overlap. Punjabi has a tonal system that Hindi lacks (three phonemic tones: high, low, mid), but the consonant inventory and vowel system are otherwise similar.

We initialize Punjabi acoustic model training from Hindi model weights rather than from scratch. The lower encoder layers that have learned general Indic phoneme representations transfer well; the variance adaptor (which predicts duration, pitch, and energy) needs more significant adaptation because of the tonal system. We freeze the lower encoder layers early in training and allow the upper layers and variance adaptor to adapt to Punjabi-specific patterns.

The empirical result: initializing from Hindi weights reduces required training data by approximately 40-50% for equivalent phoneme error rate compared to training Punjabi from scratch. We reach comparable quality to scratch-trained models at around 70-80 hours of Punjabi data versus the 120-150 hours scratch training would require.

For Gujarati, which is also closely related to Hindi but with distinct phonological features (retroflexes realized differently, vowel length distinctions), transfer from Hindi is similarly effective. For Marathi, which is a Deccan-branch Indo-Aryan language with distinct features, transfer from Hindi is helpful but less dramatic; Marathi has phonological patterns that Hindi does not represent well.

Data collection strategy

When we need more data for a low-resource language, we cannot wait for academic dataset releases. We collect it ourselves using a combination of approaches.

Sentence selection for recording: we use coverage-maximizing sentence selection algorithms that choose sentences to maximize the phoneme and phoneme-pair coverage of recording sessions. A speaker recording 10 hours of speech selected for phoneme coverage produces a more useful dataset than 10 hours of random sentences from a news corpus. This is particularly important for rare phoneme pairs: in Punjabi, the tonal contrasts need to appear in enough contexts that the model can learn the relationship between tone markings and acoustic realization.

Native speaker recording quality matters enormously. We have worked with speakers in noisy environments, speakers with regional accents far from the standard, and speakers who change delivery style mid-session. Any of these problems in even a small fraction of training files degrades model quality disproportionately. We run automated quality screening (SNR estimation, speaker consistency checks, transcript alignment verification) before any recordings enter training. Rejection rates are typically 8-15% of raw recordings.

We also use data augmentation to extend effective dataset size: speed perturbation (slightly varying playback rate of existing recordings to create variants), pitch perturbation (within a narrow range), and noise augmentation with room impulse responses from clean rooms. These augmentations help robustness but are not a substitute for phoneme coverage; they expand the acoustic diversity without adding new phoneme contexts.

Handling tones in Punjabi

Punjabi's three-tone system is the most distinctive challenge for a speaker of Hindi or other non-tonal Indic languages building Punjabi TTS. The tones (uddatt, anadatta, svarit in linguistic terminology, though different romanization systems use different names) are phonemic: they change word meaning. "Kora" meaning "whip" and "kora" meaning "horse" are distinguished by tone in spoken Punjabi.

Gurmukhi script (used for Punjabi) encodes tonal information through a system of tone letters and diacritic marks. The letter "ha" and the diacritics "bindi" and "tippi" interact with surrounding consonants to indicate tone. Our G2P pipeline for Punjabi includes explicit tone parsing that extracts the tonal category for each syllable before passing to the acoustic model. The acoustic model then predicts pitch contours conditioned on both the phoneme sequence and the tonal annotation.

This is an area where our phoneme error rate metric is particularly useful: tone errors are audible to native Punjabi speakers but largely invisible to non-Punjabi listeners rating naturalness. Without explicit phoneme-level evaluation by native speakers, it is easy to ship a Punjabi TTS system that scores acceptably on generic MOS but systematically gets tone wrong.

Honest assessment of current quality

Punjabi at 2.9 MOS in our internal comparison study is below our quality threshold for unreserved production recommendation. Gujarati at 3.1 is borderline. Marathi at 3.2 is improving faster since we added more training data.

We ship these languages because developers building Punjabi, Gujarati, and Marathi applications need something to work with, and our current models are better than no TTS or than using English-trained models adapted for Indic text. But we are direct about the quality gap. If you are building a use case where synthesis quality is critical (educational narration, high-stakes customer communication), evaluate the output on your specific content before committing. If you are building a prototype or a use case where intelligibility is the primary requirement and naturalness is secondary, our current Punjabi and Gujarati models are usable.

Try the Maya Research API
Stream multilingual TTS in 5 minutes. Free tier, no credit card required.
Get API Key

More from the lab