We ship open weights alongside our API. This post is a practical guide to fine-tuning the Hindi model on domain-specific speech, using financial call-center speech as the example. The same workflow applies to other Indic languages and domains, with adjustments for the language-specific G2P pipeline.
Before getting into the steps, a quick framing. Fine-tuning TTS on a small domain corpus is not the same as fine-tuning a language model. You are not trying to teach the model new linguistic patterns; you are adapting the acoustic characteristics and prosodic style of the output to match a target domain. The base model already knows Hindi phonology. You are teaching it a speaking style: the measured pace of a financial narrator, the specific register of a call-center agent, or the animated delivery of an educational explainer.
What the open weights package includes
When you download the open weights for the Hindi model, you get three components: the acoustic model checkpoint (FastSpeech2-style architecture, ~90M parameters), the vocoder checkpoint (our HiFi-GAN variant), and the language-specific G2P model. Fine-tuning typically targets only the acoustic model; the vocoder and G2P are usually left unchanged unless you have a specific reason to modify them.
The acoustic model checkpoint is saved in PyTorch format. We provide a training configuration file that specifies layer freezing, learning rate schedules, and batch sizes tested to work with small datasets. We also provide a validation script that computes phoneme error rate on a held-out set during training, which is your primary signal for whether fine-tuning is converging correctly.
Data requirements for the financial call-center example
For this example, the goal was to adapt the Hindi model to the speaking style used by financial service agents: measured pace, emphasis on numbers and financial terms, and a professional but approachable tone. We worked with a dataset of 4 hours of recorded call-center speech from a Bengaluru-based financial services firm, transcribed in-house by native Hindi speakers.
4 hours is on the low end of what fine-tuning reliably works with. Our general guidance is 3-8 hours for style adaptation. Below 3 hours, you risk the model memorizing the speaker-specific characteristics of your recordings rather than generalizing the style. Above 8 hours, you get diminishing returns from additional data unless you are doing full retraining rather than fine-tuning.
Audio quality requirements: 16kHz or higher sample rate, minimal background noise, single speaker per file. Call recordings often have compression artifacts from telephony codecs. We provide a preprocessing script that applies spectral denoising and resamples to 22050Hz (the model's native rate). Files with significant noise, heavy codec artifacts, or overlapping speech should be excluded; even a small proportion of low-quality files can degrade fine-tuning results disproportionately.
The fine-tuning configuration
We freeze the lower encoder layers of the acoustic model (layers 1-4 in the transformer stack) and fine-tune only the upper layers and the variance adaptor. This preserves the base phonological knowledge while allowing the model to adapt duration, pitch, and energy prediction to the target style.
Key training parameters for the financial domain example:
- Learning rate: 1e-4 with a linear warmup over 1000 steps, then cosine decay. Higher rates cause instability on small datasets; lower rates make convergence too slow to be practical.
- Batch size: 16 audio segments. We segment audio files into 3-10 second clips during preprocessing.
- Fine-tuning duration: typically 15,000-25,000 steps on a 4-hour dataset. Validation phoneme error rate plateaus around step 18,000 in our financial example.
- Layer freezing: freeze layers 1-4 (0-indexed) of the main encoder. This is documented in the training config as
freeze_encoder_layers: [0, 1, 2, 3].
Training on a single A100 takes about 3-4 hours for 20,000 steps on a 4-hour dataset. On a T4 (a more commonly available cloud option), expect 8-10 hours. We have not tested on consumer GPU hardware below T4 class.
Evaluating the fine-tuned model
After fine-tuning, run the validation script against a held-out set of test sentences from the target domain. For the financial example, we used 150 sentences specifically covering financial vocabulary, number sequences, and the formal-but-accessible register typical of Indian financial advisory speech.
Two things to check: phoneme error rate (should be at or below baseline), and a listening test specifically for prosodic style. The phoneme accuracy check catches regressions from fine-tuning (if you have inadvertently damaged the model's phonological accuracy by training with low-quality data or incorrect hyperparameters). The prosody listening test is where you evaluate whether the adaptation actually worked.
In our financial example, phoneme error rate stayed at 4.2% (essentially unchanged from baseline) and the domain-specific listening test rated the fine-tuned model 0.4 MOS points higher on the style match dimension compared to the base model. The most noticeable improvement was in number sequences: financial contexts involve frequent reading of multi-digit numbers (account numbers, amounts, percentage rates), and the base model's pacing on these sequences was less convincing than after fine-tuning on domain data that included similar sequences.
Deployment after fine-tuning
The fine-tuned acoustic model checkpoint can be used with the Maya SDK or our containerized inference server in the same way as the base checkpoint. You specify the checkpoint path in the model configuration. The G2P pipeline and vocoder are unchanged, so inference latency is similar to the base model.
One practical consideration: the fine-tuned model is optimized for the domain distribution it was trained on. If you route general-purpose Hindi synthesis through it, the output may sound slightly less natural on text that is far from the financial register. We recommend maintaining the base model for general synthesis and routing only domain-specific text through the fine-tuned variant. The model parameter in the API lets you specify which checkpoint to use per request, making this routing straightforward.
What fine-tuning does not fix
Fine-tuning on small domain data cannot fix limitations in the base model's phonological coverage. If the base Hindi model has systematic errors on a specific phoneme, those errors will likely persist through fine-tuning because the fine-tuning dataset is too small to provide enough signal to correct them. Fix phonological issues at the base model level; use fine-tuning only for style and domain adaptation. This is the most important boundary to understand before starting a fine-tuning project.