SSML (Speech Synthesis Markup Language) is the standard way to override TTS prosody programmatically. The <prosody> tag in SSML can modify rate, pitch, and volume. If you have used SSML with English TTS systems and are now synthesizing Indic languages, the behavior is different in ways that matter. This post is a practical guide to what actually works and what does not on our models.
How SSML prosody interacts with Indic acoustic models
English TTS systems trained heavily on expressive, varied speech corpora tend to have a wide dynamic range in the underlying model. Pitch can be raised 20% and the output sounds natural because the model has seen many examples of high-pitched speech. Rate can be slowed by 30% and the output maintains naturalness because the model has internalized how duration expansion affects phoneme blending.
Indic TTS models, including ours, are trained on corpora that are generally less varied in prosodic range. News reading and formal narration dominate most available transcribed speech data for Hindi, Tamil, and other Indic languages. The practical result: the acoustic model's "operating range" for prosody is narrower. SSML tags that request values outside that range produce output that sounds increasingly unnatural, not just louder or faster.
The safe ranges we have established through testing are: rate from 0.75 to 1.3 (outside this, phoneme blending degrades), pitch from -15% to +15% (outside this, voice quality artifacts appear), volume modifications work reliably but are mostly handled by the audio pipeline rather than the model, so they are less interesting for quality tuning.
Rate control: what changes and what does not
Setting rate="slow" or rate="0.85" on Hindi synthesis works well for most content. Duration is stretched relatively uniformly, and the output sounds like a slightly deliberate reading pace, which is often appropriate for IVR prompts, instructional content, and accessibility-focused applications.
Rate increases are more problematic. At rate="1.25", Hindi output sounds acceptably faster but conjunct consonant clusters start to compress in ways that reduce intelligibility. Retroflexes, which require slightly longer voice onset time for naturalness, are the first phonemes to degrade at high rates. Tamil is more sensitive to rate increases than Hindi; we do not recommend going above 1.15 for Tamil without listening carefully to your specific content.
<speak>
<prosody rate="0.9">
<!-- Slightly slower for an IVR confirmation prompt -->
आपका ऑर्डर बुक हो गया है। ऑर्डर नंबर है
</prosody>
<say-as interpret-as="digits">7492</say-as>
</speak>
Pitch control: practical limits
Pitch modification in SSML (pitch="+10%" or pitch="-5st") applies a frequency shift to the synthesized output. For small adjustments (up to ±10%), this produces natural-sounding results and can help differentiate voices or match the intended register of content.
For larger adjustments, pitch modification starts to interfere with Indic language tonal and intonational patterns. Hindi is not a tonal language, but it has systematic sentence-level intonation contours (rising on questions, falling at declarative sentence ends) that are baked into the model's prosody. Large pitch shifts can flatten or exaggerate these contours in unnatural ways.
One useful application of small pitch adjustment: if your application uses TTS for both male and female voice options and you want to approximate a second voice register without separate model checkpoints, a pitch adjustment of -3 to -5 semitones applied to a female-voiced model can give a passable lower register for low-stakes applications. This is not a substitute for separate voice models, but it works for prototyping.
Emphasis tags: mixed results
The <emphasis> tag is supposed to modify prosody to signal emphasis on a word. In English TTS, this typically increases pitch, duration, and energy on the emphasized word to mimic spoken stress. In our Indic models, the behavior is less reliable.
For Hindi, <emphasis level="strong"> consistently increases duration and energy on the emphasized word, which is the correct physical effect. Pitch modification is inconsistent; sometimes it increases, sometimes not, depending on where in the sentence the emphasis falls. <emphasis level="reduced"> works as expected.
For Tamil and Malayalam, emphasis tags have limited effect. The prosodic system in Dravidian languages handles sentence-level focus differently from Indo-Aryan languages, and our current models have not been trained with explicit emphasis annotation in Tamil/Malayalam corpora. We treat this as a known limitation.
Break tags for natural pacing
The <break time="Xms"/> tag is the most reliably effective SSML control for Indic TTS. Inserting explicit pauses at major clause boundaries, after listed items, or between a header and its content consistently improves perceived naturalness. This is partly compensating for prosodic deficiencies in the model's handling of long sentences.
For IVR content, we recommend inserting a 400-600ms break before reading an account number, OTP, or important figure. This matches natural speech pacing for important information and gives the listener a moment to prepare to remember the value.
<speak>
आपका वन-टाइम पासवर्ड है
<break time="500ms"/>
<say-as interpret-as="digits">384920</say-as>
<break time="300ms"/>
यह पासवर्ड दस मिनट में एक्सपायर हो जाएगा।
</speak>
What SSML cannot fix
SSML prosody control operates on top of what the acoustic model produces. It cannot fix systematic errors in the model's phonological output or add naturalness to a voice that the model does not have. If a specific phrase sounds flat or unnatural in your content, check first whether the G2P is producing the correct phoneme sequence before reaching for SSML adjustments. A prosody tag applied on top of an incorrect phoneme sequence will produce a modified but still-wrong output. Debug phoneme accuracy first, then tune prosody.