Back to The Maya Lab
Engineering Ananya Krishnamurthy

How We Get First-Byte Audio Under 180ms in Our Streaming TTS API

A look at the inference pipeline changes that brought our median TTFB from 380ms down to under 180ms, and what we had to give up to get there.

Last quarter we shipped a change to our streaming TTS pipeline that cut median first-byte audio latency from 380ms to 162ms on standard 200-character inputs. This post explains exactly what we changed, the tradeoffs we made, and one optimization path we tried and abandoned.

A quick framing note before diving in: TTFB for audio is different from TTFB for HTTP responses. We define it as the elapsed time from when the API receives the synthesis request to when the first decodable audio chunk arrives at the client. That includes G2P, acoustic model inference, vocoder, and the transport stack. Reducing just one of those does not move the number enough to matter.

What the pipeline looked like before

The original streaming pipeline had a sequential waterfall. Text came in, we ran a text normalizer (handling numerals, abbreviations, mixed-script inputs), a grapheme-to-phoneme converter for the target language, then batched the full phoneme sequence into our acoustic model (FastSpeech2-style architecture), got the full mel-spectrogram back, then ran the neural vocoder in a single forward pass to produce a PCM audio buffer, and finally streamed that buffer in 20ms chunks.

The problem is obvious when you lay it out: the client does not receive a single byte of audio until the vocoder has produced the entire utterance. For a 200-character Hindi sentence (roughly 2-3 seconds of synthesized speech), that meant waiting for the full forward pass before streaming began. Our P50 TTFB was 380ms and P95 was around 520ms on our primary inference nodes.

The change that mattered most: streaming vocoder chunks

Neural vocoders like HiFi-GAN and our internal variant operate on mel-spectrogram frames, typically 256-sample hop sizes at 22050Hz. The acoustic model produces a full spectrogram in a single inference call, but the vocoder does not have to consume it all at once.

We restructured the vocoder to process incoming spectrogram frames in windows of 40 frames (roughly 465ms of audio at 22050Hz with 256-sample hops). As soon as the acoustic model finishes producing the first 40 frames of spectrogram, the vocoder starts generating audio. While the vocoder is processing frames 1-40, the acoustic model continues producing frames 41-80, and so on. The acoustic model and vocoder now run as a pipeline, not sequentially.

This alone dropped our P50 TTFB to 198ms. The first audio bytes now arrive after only the first spectrogram chunk is processed, not after the entire utterance.

G2P pre-processing: shifting work earlier

For Devanagari text, our G2P pipeline handles schwa deletion, conjunct consonant expansion, and nukta-modified consonant resolution. This is not a trivial pass; it can take 15-40ms depending on sentence complexity and the density of sandhi (phoneme junction) rules.

We moved G2P to happen in parallel with connection setup and request parsing. When a client opens a WebSocket connection and sends the synthesis request, we begin G2P processing immediately on a separate thread rather than waiting for the request to be fully parsed and validated. In practice this means G2P output is ready roughly 12-18ms before the acoustic model is waiting for it. On a 200-character input that savings is meaningful at the tail.

We also added a G2P result cache keyed on normalized text. Repeated synthesized strings (OTPs, standard IVR prompts, confirmation messages) skip G2P entirely. Cache hit rate in our early-access developer traffic is around 34%, which is higher than we expected.

What we tried and discarded: smaller acoustic model

One obvious path to lower TTFB is a smaller acoustic model. We trained a distilled variant with 40% fewer parameters that ran roughly 80ms faster per inference. The TTFB improvement was real, but MOS scores on morphologically complex sentences dropped from 4.0 to 3.6 on our internal Hindi test set. The distilled model struggled with longer conjunct consonant clusters and produced flat prosody on embedded clauses.

We are not saying distillation is a bad idea for all use cases. For short, formulaic inputs (OTPs, status messages, order confirmations) the distilled model sounds fine and the latency win is significant. We are saying that for a general-purpose API targeting both conversational and long-form synthesis, the quality tradeoff was not acceptable. We are keeping the distilled variant as an opt-in parameter for latency-critical applications.

Current benchmarks and honest caveats

On 200-character Hindi inputs, P50 TTFB is now 162ms and P95 is 240ms measured from our Bengaluru inference cluster. For Tamil and Marathi inputs of the same length, numbers are slightly higher (P50 ~175ms) because those G2P pipelines are more complex.

Two honest caveats. First, these numbers are from our own infrastructure. Client-side TTFB adds network RTT, which on Indian mobile networks can be 40-120ms. We measure our own latency contribution, not total perceived latency. Second, very short inputs (under 80 characters) are actually slower in our current architecture because the pipeline-setup overhead dominates. For short inputs, the original sequential approach with the distilled model is faster. We are working on an input-length routing heuristic that selects the pipeline automatically.

What we are working on next

The next target is speculative decoding on the acoustic model side. We have early experimental results showing that a small "draft" acoustic model can predict spectrogram frames that the full model then verifies, similar to how speculative decoding works in LLM inference. In experiments on 15% of our test set, this improves throughput on longer inputs without increasing TTFB. It does not help TTFB on short inputs, which is still the harder problem for us.

Streaming TTS latency is a systems problem as much as a modeling problem. The vocoder chunking change required rethinking our memory management (we had to handle partial spectrogram buffers across CUDA kernel calls), and the G2P threading required careful handling of Devanagari Unicode normalization across thread boundaries. If you are building something latency-sensitive on top of our API and want to discuss the streaming configuration options, the stream_chunk_frames parameter in the API exposes some of this control directly.

Try the Maya Research API
Stream multilingual TTS in 5 minutes. Free tier, no credit card required.
Get API Key

More from the lab