This post walks through a complete integration of Maya TTS into a Twilio-based IVR for Hindi voice. We cover the API integration, audio format handling, streaming buffer management, and the specific network jitter challenges you will encounter on Indian mobile networks. This is a practical guide, not a product tour.
The basic architecture
A Twilio IVR (Interactive Voice Response) built with TwiML uses the <Play> verb to play audio and the <Gather> verb to collect DTMF or voice input. The simplest approach is to pre-generate audio files for each IVR prompt and host them on a CDN, then reference them in your TwiML. This works and is reliable, but it does not scale to dynamic content: account balances, appointment times, personalized greetings.
For dynamic Hindi TTS in an IVR, you have two options: generate audio on-demand during the call and serve it from a URL, or stream audio directly using TwiML's streaming features. We will cover both approaches, but the on-demand URL approach is more reliable for most IVR use cases on Indian mobile networks.
Approach 1: on-demand audio URL generation
When a call reaches a point requiring dynamic synthesis, your webhook server makes a synchronous call to the Maya API, saves the returned audio to a temporary location (S3, GCS, or a local path with an accessible URL), and returns TwiML referencing that URL.
import maya_tts
import boto3
import uuid
def synthesize_and_upload(text: str) -> str:
client = maya_tts.Client(api_key=MAYA_API_KEY)
audio = client.synthesize(
text=text,
language="hi-IN",
voice="meera",
format="mp3",
sample_rate=8000 # Twilio's preferred rate for telephony
)
key = f"ivr/audio/{uuid.uuid4()}.mp3"
s3 = boto3.client("s3")
s3.put_object(Bucket=IVR_BUCKET, Key=key, Body=audio, ContentType="audio/mpeg")
url = f"https://{IVR_BUCKET}.s3.amazonaws.com/{key}"
return url
The TwiML response then uses that URL in a <Play> verb. The caller hears a brief pause (typically 800-1200ms on a good connection) while synthesis and upload complete. For account balance or OTP prompts where a 1-second pause is acceptable, this approach is reliable and simple to debug.
One important detail: Twilio's telephony network expects 8kHz mono audio for lowest-latency playback. Our API accepts a sample_rate parameter; set it to 8000 for IVR use. The higher sample rates sound better but add unnecessary file size and playback latency on Twilio's end.
Approach 2: streaming with TwiML streams
Twilio's <Stream> verb supports bidirectional audio streaming via WebSocket. Your server can push synthesized audio bytes to Twilio in real time, and Twilio plays them as they arrive. This enables lower perceived latency for long dynamic prompts, since Twilio starts playing audio before the full synthesis is complete.
async def stream_synthesis_to_twilio(ws, text: str):
client = maya_tts.AsyncClient(api_key=MAYA_API_KEY)
async for chunk in client.synthesize_stream(
text=text,
language="hi-IN",
voice="meera",
format="mulaw", # Twilio streams use mulaw encoding
sample_rate=8000
):
# Encode chunk as Twilio media message
payload = base64.b64encode(chunk).decode("utf-8")
await ws.send_json({
"event": "media",
"streamSid": stream_sid,
"media": {"payload": payload}
})
Twilio's streaming interface requires mu-law encoded audio at 8kHz. The Maya API supports format="mulaw" directly; you do not need to transcode on your server. Transcoding in Python adds latency and is a common mistake in streaming IVR integrations.
Network jitter on Indian mobile networks
This is where most IVR integrations hit unexpected problems. Indian mobile network latency is highly variable, especially on 4G/LTE connections in tier-2 and tier-3 cities. RTT between a caller's handset and your server can vary from 40ms to 300ms within a single call, with occasional spikes over 500ms.
For the on-demand URL approach, network variability mostly affects how long the caller waits before hearing audio. Build in a minimum pause before the dynamic prompt to mask synthesis and upload time: a TwiML <Pause> of 500ms before a <Play> gives synthesis time to complete and avoids the perception of silence.
For streaming synthesis, network jitter is more serious. Twilio's streaming infrastructure has a jitter buffer, but it is sized for phone-to-phone audio, not server-to-server media streams. We have seen streaming playback become choppy or lose synchronization when the server pushes audio chunks faster than Twilio can buffer them. The mitigation is to push chunks at a controlled pace relative to audio duration: for 8kHz mulaw, each 20ms chunk is 160 bytes. Send chunks with a minimum 15ms interval between sends, not as fast as they arrive from the Maya streaming API.
Handling Hindi text preparation for IVR
IVR prompts often contain mixed content: a standard Hindi sentence with an embedded account number, date, or currency amount. You need to normalize these before passing to the TTS API.
Numbers in Hindi read differently from English. "502" in an account context reads "panch sau do" (five hundred two) in standard Hindi. Currency amounts have their own conventions: "1,450 rupees" reads "ek hazaar chaar sau pachaas rupaye." Our text normalizer handles most of these cases, but if you are constructing dynamic strings programmatically, test each number format explicitly. Edge cases include numbers with trailing zeros (1,500 vs. 1,200), large amounts in lakhs and crores, and dates in DD/MM/YYYY format.
The Maya API also accepts inline SSML for cases where you need explicit control over how an embedded number or abbreviation is pronounced. For an OTP prompt, wrapping the digits in <say-as interpret-as="digits"> forces digit-by-digit pronunciation rather than reading the OTP as a number.
Testing before production
Test your IVR synthesis on an actual phone call, not just through your browser or a laptop speaker. Telephony audio at 8kHz sounds noticeably different from synthesized audio at 22kHz. The compressed bandwidth removes high-frequency content that carries naturalness, and a voice that sounds excellent at full quality may sound flat or muffled on a phone call. Adjust your voice selection and optionally your SSML prosody settings based on listening through an actual handset.
We also recommend testing on a 4G connection from a non-metro location if your IVR is targeting broad Indian geography. The latency and packet loss characteristics are meaningfully different from a fiber connection in a major city, and they will expose streaming buffer issues that are invisible on a low-latency connection.