Back to The Maya Lab
Benchmarks Ananya Krishnamurthy

Streaming TTS Latency Benchmarks: Maya vs Three Open-Source Alternatives

We measured time-to-first-byte, P50, P95, and P99 latency across Maya and three open-source streaming TTS systems under equivalent conditions. Here are the numbers.

We ran head-to-head latency tests comparing Maya's streaming TTS API against three open-source TTS systems on equivalent hardware. The results were not uniformly in our favor, and a few of them surprised us. This post covers the methodology, results, and what we think explains the less expected findings.

Why we ran this benchmark

Latency claims in TTS marketing materials are almost always measured on the vendor's own optimized infrastructure, with pre-warmed models and favorable input sizes. We wanted to know how Maya performs relative to open-source alternatives on a level playing field: same hardware, same input text, same measurement methodology.

We are not claiming this is a definitive comparison. Open-source systems can be optimized in ways we did not apply. Our Maya API has infrastructure that open-source self-hosting cannot easily replicate. This benchmark is useful for developers deciding whether to use a hosted API versus self-hosted open-source, and for understanding where the tradeoffs land on short versus long inputs.

Systems compared

We compared four systems: Maya streaming API (served from our standard production endpoint), VITS (a neural TTS architecture commonly used as a baseline), Coqui TTS with the YourTTS backbone, and a FastSpeech2 + HiFi-GAN combination that represents a standard open-source production setup. All three open-source systems were run on an A100 GPU instance (same hardware class as our inference servers) with TorchScript models and CUDA 12.1.

We measured time-to-first-audio-byte (TTFB): elapsed from API request sent to first audio byte received at the client. For open-source systems, we implemented equivalent streaming interfaces and measured the same way. Network RTT between the benchmark client and our API endpoint was ~8ms (same AWS region, same datacenter). We ran 500 trials per condition and report P50 and P95.

Results: standard inputs (200 characters, Hindi)

SystemP50 TTFBP95 TTFB
Maya API162ms240ms
FastSpeech2 + HiFi-GAN198ms290ms
VITS244ms380ms
Coqui TTS (YourTTS)310ms520ms

On standard 200-character Hindi inputs, Maya's streaming pipeline is fastest at P50 and P95. This is expected given our pipeline optimizations (described in a separate post). The FastSpeech2 + HiFi-GAN setup is competitive at P50; the gap at P95 is larger, likely because our pipeline has more consistent tail latency due to queue management on our inference servers.

Results: short inputs (under 80 characters)

This is where the results surprised us.

SystemP50 TTFBP95 TTFB
Maya API205ms310ms
FastSpeech2 + HiFi-GAN175ms260ms
VITS190ms295ms
Coqui TTS (YourTTS)240ms410ms

On short inputs, self-hosted FastSpeech2 + HiFi-GAN is actually faster than our API. We think this is explained by our pipeline chunking overhead. The streaming vocoder architecture we built to handle medium and long inputs efficiently has a fixed setup cost (buffers, CUDA context, pipeline initialization) that dominates on very short inputs where there is little audio to generate. A simpler sequential inference path on self-hosted hardware does not pay this setup cost.

We are working on input-length routing: short inputs below a threshold will take a different inference path that avoids the streaming pipeline overhead. We do not have a target date for this, but the benchmark data made it a clear priority.

Results: long inputs (500+ characters)

SystemP50 TTFBP95 TTFB
Maya API168ms248ms
FastSpeech2 + HiFi-GAN350ms490ms
VITS480ms680ms
Coqui TTS (YourTTS)580ms890ms

On long inputs, the streaming pipeline advantage compounds. VITS and Coqui TTS process the full input before returning any audio, so TTFB scales linearly with input length. Our streaming vocoder returns the first audio chunk while continuing to process the rest of the input, so TTFB on 500-character inputs is essentially the same as on 200-character inputs.

What self-hosting is actually competitive for

If your application primarily synthesizes short prompts (OTPs, status notifications, short confirmations) and you have the engineering capacity to maintain a self-hosted TTS stack, open-source FastSpeech2 + HiFi-GAN on GPU hardware is a legitimate option with competitive latency. You pay the maintenance cost in exchange for the infrastructure cost savings.

For medium to long inputs, for applications that need multi-language support without maintaining separate model checkpoints per language, or for applications where model quality (especially on morphologically complex Indic content) matters more than latency, the hosted API approach becomes more attractive. The benchmark numbers are one input to that decision; the total cost of model maintenance, G2P pipeline maintenance, and infrastructure is another.

Try the Maya Research API
Stream multilingual TTS in 5 minutes. Free tier, no credit card required.
Get API Key

More from the lab