Two years ago, our text-to-speech API worked the way you'd expect: send text, generate the full clip, return it. Fine for an audiobook chapter or a batch of ad voiceovers. Not fine for a voice agent on a live call, where a 900ms pause before the first word feels broken, not polite.

This post covers how we rebuilt our inference stack to consistently hit sub-200ms time-to-first-audio-byte: streaming inference, kernel-level serving work, and a routing layer that treats network distance as a latency budget line item.

The problem with batch inference

The naive pipeline is: receive text, run the full model to produce a complete waveform, encode it, send the response. Each stage waits for the last to finish. A 10-second utterance might take 600–900ms of generation alone, before encoding, before a single byte leaves our servers.

That gets worse as utterances get longer — exactly backwards from what real-time systems need. A voice agent doesn't need the last second of audio quickly; it needs the first 200ms quickly, so playback can start while the rest generates. Batch inference optimizes for total generation time. Real-time products need to optimize for time-to-first-byte, and those are different problems.

We also learned that a faster model alone doesn't solve this. Even a model generating audio 2x faster in aggregate still forces the client to wait for the whole clip if responses return atomically. The bottleneck wasn't purely compute — it was the shape of the response.

The shift to streaming inference

The fix is conceptually similar to how large language models stream tokens instead of a full response: we stream audio frames as they're produced. As soon as our model generates and vocodes roughly 40–80ms of audio, that chunk flushes to the client over a persistent connection, while generation of the next chunk continues in parallel.

This meant rethinking two things. First, the model needed to emit audio incrementally — we restructured our decoder to output fixed-size frame windows as soon as they're stable, instead of requiring full-utterance context before emitting anything. Second, our API layer needed a transport that supports partial delivery: chunked HTTP over a keep-alive connection, with a WebSocket option for clients that want bidirectional control like barge-in or early termination.

Here's a simplified client consuming a streamed response from our API, measuring time-to-first-byte along the way:

const requestStartedAt = performance.now();

const response = await fetch('https://api.play.ht/v1/tts/stream', {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${process.env.PLAYHT_API_KEY}`,
    'Content-Type': 'application/json',
  },
  body: JSON.stringify({
    text: 'Streaming audio as it is generated, not after.',
    voice: 'en-US-Standard-A',
    output_format: 'mp3',
  }),
});

const reader = response.body.getReader();
let firstByteAt = null;

while (true) {
  const { done, value } = await reader.read();
  if (done) break;

  if (!firstByteAt) {
    firstByteAt = performance.now();
    console.log(`time-to-first-byte: ${Math.round(firstByteAt - requestStartedAt)}ms`);
  }

  // Append each chunk to a playback buffer so audio can start
  // before the stream finishes.
  sourceBuffer.appendBuffer(value);
}

The client never waits for a complete file. It starts playing audio the moment the first chunk arrives, and the rest fills in behind it.

Model-serving optimizations

Streaming buys you the ability to return audio early, but it doesn't automatically make that first chunk fast. Two areas of low-level work mattered most.

Attention kernel efficiency

Our voice models lean heavily on attention layers, and generic implementations spend a surprising amount of time on memory movement rather than math. We built fused, memory-aware attention kernels that keep intermediate activations resident in fast GPU memory instead of round-tripping through slower global memory between steps, and we maintain a persistent key-value cache per active stream so each new frame only attends against cached context instead of recomputing it. Together, these cut per-frame latency substantially and, just as important for streaming, made it far more consistent — a single slow frame stalls playback even if the average is fine.

GPU batching strategies

Serving many simultaneous streams efficiently means batching them on one GPU without making any individual stream wait for a batch to fill. We use continuous batching: incoming requests join an in-flight batch at frame boundaries rather than waiting for a fixed window to close, and a completed stream's slot is freed immediately for the next request. We pair this with padding-aware grouping — batching requests of similar length and cache size together — and run quantized inference (int8/fp8 where model quality allows) for the bulk of production traffic, which increases concurrent streams per GPU without a perceptible quality hit.

Edge and regional routing

None of the above matters if a request from Singapore round-trips to a single data center in Virginia before generation even starts. Network latency is often the largest and most invisible component of time-to-first-byte, and it's pure overhead.

We run inference pools in multiple regions and route each request to the nearest healthy pool using latency-based routing rather than static geo-DNS rules. Connections stay warm: we maintain pre-established, TLS-resumed connection pools between our edge layer and inference nodes, so a request doesn't pay a full handshake on top of a cold start. For customers on our streaming TTS endpoints, regional routing alone removed more round-trip time than any single model optimization — latency engineering is as much about where compute happens as how fast it happens.

Results

We track time-to-first-audio-byte (TTFB-audio) as our primary latency metric, since it's what actually determines whether a conversation feels responsive:

  • Before (batch inference, single region): p50 780ms, p95 1,850ms
  • After (streaming + kernel work, nearest region): p50 140ms, p95 195ms
  • After, cross-region worst case: p50 210ms, p95 340ms

The p95 number matters as much as the median. A system fast on average but occasionally spiking to two seconds still breaks the feel of a live conversation. Closing that tail required the same continuous-batching work described above — a request queued behind a slow neighbor is the most common cause of p95 outliers, and eliminating queuing stalls closed most of that gap.

The goal was never "fast on average." It was "never noticeably slow," which is a much harder bar to clear.

What's next

We're working on predictive pre-warming — spinning up regional capacity ahead of anticipated demand so cold starts disappear during traffic spikes — and extending frame-level streaming to voice cloning and dubbing workloads that have historically run in batch mode. If you're building a voice agent or anything where a caller is waiting on the other end, the streaming API is documented in our developer docs, with current plan limits on our pricing page. Latency work is never really finished — it's a budget you keep spending down, one bottleneck at a time.