The Maya Lab
Technical notes on multilingual TTS, acoustic modeling, and language engineering from the Maya Research team.
Lab articles
How We Get First-Byte Audio Under 180ms in Our Streaming TTS API
A look at the inference pipeline changes that brought our median TTFB from 380ms down to under 180ms, and what we had to give up to get there.
All articles
Measuring Phoneme Accuracy in Devanagari TTS: Our Internal Benchmark Setup
Phoneme error rate tells you more about TTS quality than MOS scores for morphologically-rich languages. Here is how we measure it for Hindi and Marathi.
MSA, Egyptian, Levantine: Why Dialect Coverage Changes Everything for Arabic TTS
Standard Arabic and spoken Egyptian Arabic are not interchangeable for a voice application. We explain our three-variant approach and the training data challenges.
Fine-Tuning Maya TTS on Domain-Specific Speech: A Practical Guide
Our open weights are designed to be fine-tunable on small datasets. This guide walks through a real example: adapting the Hindi model to financial call-center speech.
Building a Hindi Voice IVR With Maya TTS: From API Call to Phone Audio
A step-by-step walkthrough of integrating Maya TTS into a Twilio-based IVR flow, handling streaming audio, and dealing with network jitter on Indian mobile networks.
Comparing TTS Output Quality Across Nine Indic Languages: What We Found
Not all Indic languages have the same model maturity. We ran a blind listening study across Hindi, Tamil, Telugu, Kannada, Malayalam, and four others. Results inside.
SSML Prosody Tags for Indic TTS: What Works, What Does Not
SSML rate, pitch, and emphasis tags behave differently on models trained primarily on Indic corpora. Here is a practical guide to what actually changes the output.
Why EdTech Platforms in India Need Regional-Language TTS (And What to Look For)
Students comprehend regional-language explanations 40% faster than English-translated content in our partner trials. Here is what that means for TTS selection in edtech.
Streaming TTS Latency Benchmarks: Maya vs Three Open-Source Alternatives
We ran head-to-head latency tests against three open-source TTS systems on equivalent hardware. The results were not what we expected on shorter inputs.
Voice Cloning Ethics: Our Principles for Open Weights and User-Generated Voices
Open weights create real misuse risks. We explain the consent framework we require for voice cloning features and why we think platform-level controls matter more than model-level restrictions.
When to Use MSA vs Dialectal Arabic in Your Voice Application
Formal Arabic and conversational Egyptian Arabic sound completely different to native speakers. This guide explains the right choice for news, customer service, and consumer apps.
Hindi and Urdu Share a Spoken Form but Not a Script: What That Means for TTS
Hindustani spoken language is largely mutual, but Devanagari and Nastaliq scripts encode different spelling conventions. Here is how we handle the divergence in the TTS pipeline.
Building TTS for Low-Resource Indic Languages: Data, Transfer, and Tradeoffs
Punjabi, Gujarati, and Marathi have far less transcribed speech data than Hindi. This is our strategy for getting usable TTS quality with limited corpora.
Neural vs Statistical TTS for Indian Languages: A 2025 Perspective
Statistical parametric TTS was the standard for Indic languages until recently. This article examines where neural models have genuinely surpassed it and where the gap is still surprising.