Developer API

Stream natural-sounding Hindi and Arabic speech from one API call

Maya Research ships a streaming TTS API plus open weights for 12 Indian and Arabic languages, with sub-200ms first-byte latency, running on your infra or ours.

Sub-200ms TTFB 12 Languages Open Weights
stream_speech.py
import mayaresearch

client = mayaresearch.Client(
    api_key="myr_live_..."
)

stream = client.tts.stream(
    text="नमस्ते, मैं आपकी कैसे मदद कर सकता हूँ?",
    language="hi-IN",
    voice="ananya",
    format="pcm_16khz"
)

for chunk in stream:
    audio_player.write(chunk)

# First byte in ~170ms
# 24 voices across 12 languages
TTFB P50: 172ms
<200ms
First-byte latency (P50)
12
Languages supported
24
Distinct voice personas
Open
Weights for self-hosting
How it works

Three steps to streaming speech

Authenticate, call the stream endpoint, play audio. No preprocessing, no phoneme dictionaries, no pronunciation files.

01
Install and authenticate
$ pip install mayaresearch

import mayaresearch
client = mayaresearch.Client(
    api_key="myr_live_..."
)

One pip install gets you the Python client. Node.js, Go, and REST examples in the docs.

02
Call the stream endpoint
stream = client.tts.stream(
    text="مرحباً بك في عالم الصوت",
    language="ar-MSA",
    voice="khalid"
)

Pass text, language tag, and voice ID. The response streams PCM or MP3 chunks immediately.

03
Consume audio chunks
for chunk in stream:
    speaker.play(chunk)

# Or buffer to file:
with open("output.mp3", "wb") as f:
    f.write(stream.bytes())

Iterate chunks for real-time playback or collect to a buffer for file export. Both patterns are under 10 lines.

Languages

12 languages, one endpoint

Every language shares the same streaming API signature. Switch from Hindi to Tamil by changing one parameter.

Hindi हिन्दी GA
Arabic عربي GA
Tamil தமிழ் GA
Telugu తెలుగు GA
Bengali বাংলা GA
Kannada ಕನ್ನಡ GA
Malayalam മലയാളം GA
Punjabi ਪੰਜਾਬੀ GA
Gujarati ગુજરાતી Beta
Marathi मराठी Beta
Urdu اردو Beta
Odia ଓଡ଼ିଆ Beta

Scripts from Devanagari to Nastaliq, all on one platform.

Our open weights are trained on speaker-verified recordings at 24kHz, with separate phoneme models per script family. Devanagari, Perso-Arabic, and Brahmic scripts each follow their own normalization pipeline before reaching the acoustic model.

Why Maya Research

Built for South Asian and MENA developers

Generic TTS systems treat Indic languages as edge cases. We built the platform the other way around.

Streaming-first architecture

Our inference pipeline starts emitting audio before the full sentence is synthesized. You hear the first chunk in under 200ms on a typical Bengaluru or Dubai data centre connection.

Open weights, your infra

Download the acoustic models and run them on-premises. All major Indian cloud providers supported. No vendor lock-in, no per-character overage if you self-host.

Dialect-aware models

Arabic includes MSA, Egyptian, and Levantine variants. Hindi covers standard, Bhojpuri-inflected, and formal registers. Each variant has dedicated voice personas, not a single model stretched thin.

SSML prosody control

Rate, pitch, and emphasis tags that actually work on Indic corpora. We benchmark prosody changes against human raters for every new model release so you can trust the output.

REST and WebSocket APIs

Standard HTTP for batch synthesis, WebSocket for real-time interactive use. SDKs for Python and Node.js ship on the same day as API changes.

DPDP-aligned data practices

We process text on servers in India by default. Enterprise plans include contractual data residency, DPA, and deletion guarantees under India's DPDP Act and equivalent frameworks.

From the community

What developers are building

We switched from a general-purpose TTS to Maya Research for our Hindi podcast app. The prosody on long sentences is significantly more natural. Latency dropped our buffering complaints by roughly half.

Rohan K.
Engineering lead, audio content platform

Our Arabic IVR previously used a European TTS vendor. The MSA voice was technically correct but clearly foreign. Maya's Egyptian Arabic voice landed completely differently with our users in Cairo.

Sara A.
Product manager, fintech CX team

The open weights let us fine-tune on our call-centre recordings in two days. No data left our servers. The resulting voice sounds like it was trained on our own agents, which is exactly what we needed.

Preethi M.
ML engineer, enterprise SaaS
Pricing

Start free, scale as you grow

Free tier includes 1,000 characters per day. No credit card required.

Free
$0 /month

For side projects and evaluation. No credit card required.


  • 1,000 characters/day
  • 2 languages
  • REST API access
  • Community support
Most Popular
Growth
$49 /month

For production apps with steady usage. All 12 languages included.


  • 2 million characters/month
  • All 12 languages
  • WebSocket streaming
  • Email support, 24h SLA
Pro
$149 /month

For high-volume apps and teams who need priority support and custom voices.


  • 10 million characters/month
  • Custom voice fine-tuning
  • Dedicated model endpoint
  • Priority support, 4h SLA

Start streaming speech today

Free tier, no credit card. First byte in under 200ms. 12 languages from one endpoint.