Most teams' first encounter with text-to-speech is a simple request/response call: send text, wait, get back an audio file. That works fine for narrating a podcast or generating a voicemail greeting, but it falls apart the moment you're building something conversational — a voice agent, a live translation tool, a real-time assistant that needs to sound like it's actually listening. In those contexts, the thing users notice first isn't voice quality. It's latency, and specifically the gap between "the system has something to say" and "the user hears the first sound."

We built PlayHT's streaming API to close that gap, and along the way learned a lot about what makes real-time audio pipelines architecturally different from batch generation. This post is a practical walkthrough of that architecture, for engineers building voice agents, live dubbing tools, or conversational products on our text-to-speech API.

Why batch TTS assumptions break in real time

A batch integration looks like this: call the endpoint, wait, get back a complete file, play it. The entire clip is generated before the client sees a single byte — fine for pre-recorded content, since nobody is waiting live on the other end.

Conversational applications invert nearly every one of those assumptions:

  • First-byte latency matters more than total latency. Users don't perceive "how long until the whole sentence is ready" — they perceive "how long until I hear anything at all." A pipeline that takes 300ms to deliver the first frame feels instant, even if the full response takes two seconds to finish.
  • Audio has to arrive in chunks, not as one file. The client needs to play frame one while frame forty is still being synthesized upstream.
  • The conversation can be interrupted at any moment. A real person doesn't wait politely for the assistant to finish a sentence — they talk over it. Your pipeline needs a way to cut generation off mid-stream without leaving dangling connections or wasted compute behind.

Designing for this means moving from request/response to streaming end to end: text streamed in as it becomes available, synthesis streamed out, and a client built to play audio incrementally rather than all at once.

A reference architecture

Here's the shape of the pipeline we see teams converge on when building real-time voice agents on PlayHT:

  1. The client establishes a streaming connection. Instead of a one-off HTTP request per utterance, the client opens a persistent WebSocket (or HTTP/2 streaming) connection once, at the start of a session, and keeps it open for the life of the conversation. A fresh connection per turn adds a TLS handshake and setup cost you don't want to pay when chasing milliseconds.
  2. Text arrives incrementally, not as one blob. If text is coming from a language model or translation step, stream it to the TTS endpoint as it arrives, ideally chunked at clause or sentence boundaries, so synthesis can begin on the first clause while the rest is still being produced upstream.
  3. PlayHT synthesizes and returns chunked audio frames. Rather than generating the whole utterance and returning one finished file, the endpoint streams back small frames as they become available, often before the sentence has fully finished synthesizing. Each frame is independently playable and encoded (commonly raw PCM, or a compressed format when bandwidth is a concern) so the client never waits for a container header at the end of the stream.
  4. The client buffers frames into a short playback queue. Frames arrive at irregular intervals — network jitter, chunk size, and synthesis timing all vary — so the client needs a small buffer between "received" and "played." This is the single most important tuning knob in the whole pipeline, covered below.
  5. An audio player drains the queue in real time. Whether that's the Web Audio API, a native mobile audio engine, or a raw PCM sink on embedded hardware, the player pulls frames off the queue at the device's native rate, and queue depth becomes your control surface for latency versus smoothness.

The key shift is that generation and playback become two independently paced processes connected by a queue, rather than one synchronous call — a decoupling that also makes graceful interruption possible.

Handling interruptions gracefully (barge-in)

In any spoken conversation, people interrupt. They talk over the assistant, change their mind mid-sentence, say "wait, stop" before a thought finishes. If your pipeline can't react within a turn or two, the interaction feels like a phone tree, not a conversation.

Handling this well means treating cancellation as a first-class operation, not an afterthought:

  • Detect the interruption early, usually via a voice-activity-detection signal on the client or in your speech-to-text layer, firing the moment the user starts talking — not after they finish.
  • Cancel generation immediately, not just playback. Stopping the audio player isn't enough — if you don't also tell PlayHT to stop generating, you're paying for audio nobody will hear, and those frames can pile up and get played anyway if the "stop" signal races against still-arriving network frames.
  • Flush every buffer in the chain. The synthesis stream, the playback queue, and any jitter buffer in between all need to clear together. A half-flushed pipeline is exactly how you get the assistant seeming to keep talking half a second after the user already started.
  • Reset cleanly for the next turn without tearing down and reopening the connection, which would reintroduce the setup latency you worked to avoid in the first place.

The pattern we recommend: an explicit cancel message sent over the same connection used for synthesis, acknowledged by the server before you stream new text. Racing a "stop" against still-arriving old frames is the single most common source of "the assistant talked over itself" bugs we see in the wild.

Buffering strategy: smooth, without feeling slow

Buffering is the central tension in every real-time audio pipeline. Buffer too little and playback stutters whenever a frame arrives a few milliseconds late. Buffer too much and you've quietly added hundreds of milliseconds of latency — smooth, but slow, which for a conversational product is often worse than a little roughness delivered fast.

A few principles that hold up in production:

  • Buffer in time, not frame count. Frame count represents a variable amount of audio if sizes vary. Target a small, fixed duration — tens of milliseconds, not seconds — and let frame count float.
  • Make the buffer adaptive. Start small. Grow it slightly on observed underruns; shrink it back down when generation and network consistently run ahead of playback. A static buffer sized for worst-case conditions punishes every user with a good connection.
  • Prioritize first-frame delivery over steady-state depth. It's fine, even good, to start on a thinner buffer than steady state, then let it fill as the stream continues. Users forgive a thin first second far more readily than a slow start.

A simplified illustration of that on the client side, combining adaptive buffering with barge-in handling:

// Pseudo-code: streaming playback with adaptive buffering and barge-in

const stream = playht.connectStream({ voice: "olivia", format: "pcm_16000" });
const playback = new AudioQueue({ targetBufferMs: 80 });

stream.onAudioFrame((frame) => {
  playback.push(frame);
});

stream.onError((err) => {
  log.warn("stream error, reconnecting", err);
  reconnectWithBackoff(stream);
});

function speak(textChunk) {
  stream.sendText(textChunk);
}

function onUserStartedSpeaking() {
  // 1. Stop generation on the server, not just local playback
  stream.cancel();

  // 2. Drop everything already queued or in flight
  playback.flush();

  // 3. Only accept new text once cancellation is acknowledged
  stream.onCancelAck(() => {
    readyForNextTurn = true;
  });
}

// Adaptive buffer: grow on underrun, shrink on sustained surplus
playback.onUnderrun(() => { playback.targetBufferMs += 20; });
playback.onSustainedSurplus(() => { playback.targetBufferMs -= 10; });

Common pitfalls

A handful of mistakes come up often enough in real-time integrations that they're worth calling out:

  • Over-buffering "just to be safe." Every millisecond of buffer is a millisecond of latency the user feels before the assistant responds. Start with the smallest buffer you can get away with, and grow it only in response to observed underruns — not in anticipation of them.
  • Ignoring backpressure. If your client-side consumer is slower than the incoming frame rate — a busy render thread, a blocked main loop, a saturated downlink — frames pile up in memory. Without a backpressure signal to the producer, you end up exhausting memory or playing increasingly stale audio. Make sure your client honors flow-control signals and can pause consumption without dropping the connection.
  • Not handling stream errors and reconnects. Long-lived connections will occasionally drop — a Wi-Fi handoff, a load balancer recycling a socket, a device switching towers. Treat this as normal, not exceptional: implement reconnect-with-backoff, and decide up front whether a mid-utterance reconnect resumes synthesis or restarts the turn cleanly. Silently swallowing stream errors is how "the assistant just went silent" tickets happen.
  • Treating cancellation as client-only. Stopping local playback without canceling server-side generation wastes compute and risks stale frames racing back into the queue.
  • Testing only on fast, stable networks. A pipeline that looks flawless on a wired office connection is exactly where over-tight buffers and missing backpressure handling hide. Test on throttled, lossy connections before you ship, not after.

Closing thoughts

Real-time voice pipelines reward a different kind of engineering discipline than batch generation: everything is measured in tens of milliseconds, every buffer is a latency decision, and interruption has to be a designed-for state rather than an edge case handled later. If you're building a voice agent, a live dubbing tool, or anything conversational, design generation and playback as two independently paced processes from day one — retrofitting streaming onto a request/response architecture later is a much harder problem than building for it up front.

You can find the full details of our streaming endpoints, connection setup, and cancellation semantics in the developer docs, and current usage-based pricing for streaming synthesis on our pricing page.