Stream natural-sounding Hindi and Arabic speech from one API call
Maya Research ships a streaming TTS API plus open weights for 12 Indian and Arabic languages, with sub-200ms first-byte latency, running on your infra or ours.
import mayaresearch
client = mayaresearch.Client(
api_key="myr_live_..."
)
stream = client.tts.stream(
text="नमस्ते, मैं आपकी कैसे मदद कर सकता हूँ?",
language="hi-IN",
voice="ananya",
format="pcm_16khz"
)
for chunk in stream:
audio_player.write(chunk)
# First byte in ~170ms
# 24 voices across 12 languages
Three steps to streaming speech
Authenticate, call the stream endpoint, play audio. No preprocessing, no phoneme dictionaries, no pronunciation files.
$ pip install mayaresearch
import mayaresearch
client = mayaresearch.Client(
api_key="myr_live_..."
)
One pip install gets you the Python client. Node.js, Go, and REST examples in the docs.
stream = client.tts.stream(
text="مرحباً بك في عالم الصوت",
language="ar-MSA",
voice="khalid"
)
Pass text, language tag, and voice ID. The response streams PCM or MP3 chunks immediately.
for chunk in stream:
speaker.play(chunk)
# Or buffer to file:
with open("output.mp3", "wb") as f:
f.write(stream.bytes())
Iterate chunks for real-time playback or collect to a buffer for file export. Both patterns are under 10 lines.
12 languages, one endpoint
Every language shares the same streaming API signature. Switch from Hindi to Tamil by changing one parameter.
Scripts from Devanagari to Nastaliq, all on one platform.
Our open weights are trained on speaker-verified recordings at 24kHz, with separate phoneme models per script family. Devanagari, Perso-Arabic, and Brahmic scripts each follow their own normalization pipeline before reaching the acoustic model.
Built for South Asian and MENA developers
Generic TTS systems treat Indic languages as edge cases. We built the platform the other way around.
Streaming-first architecture
Our inference pipeline starts emitting audio before the full sentence is synthesized. You hear the first chunk in under 200ms on a typical Bengaluru or Dubai data centre connection.
Open weights, your infra
Download the acoustic models and run them on-premises. All major Indian cloud providers supported. No vendor lock-in, no per-character overage if you self-host.
Dialect-aware models
Arabic includes MSA, Egyptian, and Levantine variants. Hindi covers standard, Bhojpuri-inflected, and formal registers. Each variant has dedicated voice personas, not a single model stretched thin.
SSML prosody control
Rate, pitch, and emphasis tags that actually work on Indic corpora. We benchmark prosody changes against human raters for every new model release so you can trust the output.
REST and WebSocket APIs
Standard HTTP for batch synthesis, WebSocket for real-time interactive use. SDKs for Python and Node.js ship on the same day as API changes.
DPDP-aligned data practices
We process text on servers in India by default. Enterprise plans include contractual data residency, DPA, and deletion guarantees under India's DPDP Act and equivalent frameworks.
What developers are building
We switched from a general-purpose TTS to Maya Research for our Hindi podcast app. The prosody on long sentences is significantly more natural. Latency dropped our buffering complaints by roughly half.
Our Arabic IVR previously used a European TTS vendor. The MSA voice was technically correct but clearly foreign. Maya's Egyptian Arabic voice landed completely differently with our users in Cairo.
The open weights let us fine-tune on our call-centre recordings in two days. No data left our servers. The resulting voice sounds like it was trained on our own agents, which is exactly what we needed.
Start streaming speech today
Free tier, no credit card. First byte in under 200ms. 12 languages from one endpoint.