Kafka Labs

Research note—October 2026—11 min read

A latency budget for live speech

Where the time goes between a caller's mouth and the agent's voice, which costs are physics and which are design, and why a live speech system has to stream at every stage.

Kafka Labs · San Francisco


Two hundred milliseconds is not a lot of time. It is less than the blink of an eye takes to complete, and it is roughly the gap a listener expects between the end of your sentence and the beginning of theirs. An agent that answers in that window feels present. One that answers in a second feels like it is somewhere else, consulting something, and the caller starts to speak slower and louder to compensate, the way people do on a bad satellite link.

This note treats responsiveness as a budget to be spent deliberately rather than a benchmark to be won. It walks through where the time goes between a caller's mouth and the agent's voice, which of those costs are physics and which are design choices, why a live speech system has to stream at every stage, what the telephone channel does to all of this, and how we intend to measure what callers actually experience. The numbers in the budget below are targets and typical values from public standards, not measurements of our system. We will report our own when we have them.

Why two hundred milliseconds

Two independent lines of evidence converge on roughly the same figure.

The first is human conversation. The modal gap between turns across languages is around two hundred milliseconds (Stivers et al., 2009). Responses that arrive much later than that are not just slow; they are heard as meaningful. A long pause before "yes" sounds like "no". A long pause before an answer sounds like the answer is uncertain or being looked up. An agent cannot escape this reading of its delays, so it has to either answer within the human window or account for its delay in a way the caller understands.

The second is telephony engineering. ITU-T Recommendation G.114 sets a planning target of 150 milliseconds of one-way mouth-to-ear delay for networks carrying conversation, with quality degrading noticeably beyond about 400 milliseconds. That figure was arrived at empirically by measuring how delay affects the ease of conversation on telephone calls, and it describes the transport alone. An agent lives inside that constraint: its own processing is added to whatever the network already costs.

Put together, the design target is clear. The whole loop, from the caller finishing to the caller hearing the first sound of a reply, should be in the low hundreds of milliseconds, and every stage in between is spending from the same account.

Where the time goes

Here is a rough accounting of a telephone call into a live speech agent. Each line is a cost that has to be paid once per response, and some are paid continuously.

StageTypical costWhat it is
Carrier and transport20 to 80 ms each wayThe phone network and the path to our gateway, varying with distance and carrier.
Packetization20 msTelephony audio arrives in fixed packets, usually twenty milliseconds each.
Jitter buffer40 to 100 msPackets do not arrive evenly; a buffer smooths them at the cost of delay.
Audio framing for the model40 to 80 msNeural codecs operate on frames; the model cannot see a frame until it is complete.
Model: perceiving the end of the turnvariableIn a cascade this is the silence threshold, 500 ms or more. In a full-duplex model it is a prediction and can be near zero or even negative.
Model: first output frametens of ms to a few hundredTime from deciding to speak to producing the first frame of audio.
Output framing and playout40 to 100 msThe first frame has to be encoded, packetized, and buffered on the far side before it sounds.

Add the columns and the lesson is uncomfortable. Before the model has done anything at all, transport, packetization, buffering, and framing have spent somewhere between one hundred and two hundred milliseconds of the budget. The model's own work has to fit in what remains, and the single largest item in a conventional design, waiting for silence to confirm the turn is over, is bigger than everything else combined.

This is why the architecture matters more than the hardware. A faster accelerator shrinks one line. Removing the endpointing wait removes the largest line, and only a model that predicts the end of the turn rather than detecting it can do that.

Geography is part of the budget

One line in the table above is set by physics rather than engineering, and it deserves a closer look. Signals travel through fiber at roughly two thirds of the speed of light, which puts a round trip across a continent at several tens of milliseconds and a round trip across an ocean well above a hundred. A caller in one region talking to a model in another pays that toll twice on every exchange, once for their voice to arrive and once for the reply to return, before any processing has happened.

Two consequences follow for anyone building live speech. The model has to run close to the callers, which means more than one place as soon as the callers are in more than one place, and the path from the carrier's edge to the model has to be short and predictable rather than routed through whatever happens to be convenient. Neither is exotic; both are easy to postpone, and postponing them quietly spends a third of the budget before a single frame reaches the model. We treat placement as a first-order design decision, not an operations detail, and we measure the transport leg separately from the model leg so that we always know which one is costing us.

Streaming all the way down

The second principle follows from the first: no stage may wait for a later boundary than it needs.

In a conventional pipeline, boundaries accumulate. The recognizer waits for the end of the utterance. The language model waits for the recognizer's final transcript. The synthesizer waits for the language model to finish a sentence, because it needs the sentence to plan prosody. Playback waits for enough synthesized audio to fill a buffer. Each wait is individually reasonable and collectively ruinous.

A live system has to be incremental at every seam:

  • Encoding should operate on short frames with bounded lookahead, so that audio becomes model input a few tens of milliseconds after it is spoken, not after the sentence ends.
  • The model should consume frames as they arrive and emit frames as it decides, so that its first output can begin before the caller's final word has fully decayed if the content warrants it.
  • Decoding to audio should produce playable sound from the first frame, not after a sentence of text exists. This is one of the structural advantages of generating audio tokens directly: the first frame of a reply is available the moment the model emits it.
  • Transport out should send that first frame immediately and keep the playout buffer as shallow as the network allows.

The practical test for whether a design is streaming is to ask, for every stage, "what is the earliest moment this stage could start, and does it?" In our experience the answer in most systems is "much later than it could", and the reasons are usually convenience rather than necessity: a library that only accepts complete utterances, a synthesizer that insists on sentence boundaries, a buffer sized for safety rather than measured need.

The telephone channel

Most of the public work on speech models uses studio or wideband audio. The telephone is a different instrument, and a model that has not learned it will be surprised by it.

A call on the public network is narrowband: 8 kHz sampling, which keeps only frequencies up to about 4 kHz. It is quantized with μ-law or A-law companding under ITU-T G.711, which was designed in the 1970s to fit a voice into 64 kilobits per second. Consonants that distinguish "f" from "s" live partly above the cutoff. Sibilance, breath, and much of what makes a voice recognizable are attenuated or gone. On top of the codec, real calls carry packet loss, jitter, echo from the far end, compression from a mobile handset, hold music, keypad tones, and the acoustics of wherever the caller happens to be standing.

Two design consequences follow.

First, the model has to be trained on this channel, not merely tested on it. Augmenting training audio to imitate the telephone path, with band limiting, companding, loss, and realistic noise, is cheap and matters a great deal. Real telephone recordings matter more.

Second, the agent's own voice has to be designed for the channel. Prosody that reads clearly in wideband audio can turn muddy at 8 kHz. Speaking slightly slower, with clearer consonant onsets, is a stylistic choice that pays for itself in intelligibility on a bad line, and it is the kind of thing a speech-native model can learn to do when it hears how its output is being received.

Honest filler

Sometimes the agent genuinely needs time. It has to look up a booking, call a tool, or wait for a system that is slower than a conversation. The question is what the caller hears while that happens.

The dishonest answers are silence and theater. Silence reads as a dropped call or a confused agent. Theater, fake typing sounds, "let me just pull that up" delivered identically every time, reads as exactly what it is once the caller has heard it twice.

The honest answer is what a person does: say that you are checking, in a way that reflects how long it will take, and come back when you have it. "One second" means one second. "Let me look, this might take a moment" means more. A speech-native model can produce these naturally, with the hesitations and fillers of real speech rather than a canned clip, and can time its return to the actual completion of the task rather than to a scripted pause. We treat this as part of the latency budget: delay that is acknowledged is spent well; delay that is hidden is spent badly.

Measuring what the caller experiences

The only measurement that counts is taken at the caller's ear, so our instrumentation is built around two-channel recordings of real calls with timestamps on both tracks, plus loopback calls placed over real carriers to our own number.

What we measure:

  • Response offset. Time from the end of the caller's turn to the first audible frame of the agent's reply, as a distribution with its tail, not as a mean. The tail is what callers remember.
  • Time to first audio after the model decides to speak, separated from the offset so we can tell model delay from transport delay.
  • Gap and overlap distribution over the whole call, compared against the human figures, so that we are measuring the shape of the conversation and not only its first exchange.
  • Yield latency when the caller interrupts, measured from the onset of the caller's speech to the agent's silence.
  • Playout stalls: moments where the agent's audio broke up because the pipeline fell behind, which callers hear as stuttering.

How we report: against the human and telephony targets above, in milliseconds, with distributions, on real calls, with the carrier and region stated. We are not going to publish a single latency number measured on a laptop with a microphone, because that number describes a situation no caller is ever in.

The realtime path at Kafka Labs

For completeness, the shape of what we are building, with the status we have stated elsewhere on this site.

Calls land on a gateway written in Rust that terminates the WebSocket or telephony stream, handles the codecs in play (μ-law and A-law for the phone network, PCM and Opus for everything else), resamples, runs a lightweight voice activity detector at the edge for bookkeeping, and applies backpressure so that a slow consumer cannot silently grow the buffers. From there, audio frames go over a persistent connection to the model runtime, which keeps state for each conversation and streams frames back the moment they exist. Everything in that path is designed around one rule: nothing waits for a boundary it does not need.

Both pieces are in development. Turtle, our phone agent, is in private beta on real numbers, and it is where these measurements are taken. We will add NVIDIA's streaming speech and model optimization tooling to this path as planned; we describe that on the main page as a plan because that is what it is.

What we are not claiming

We have not published latency figures for our system, and this note does not contain any. The standards and research cited here set the target. The budget table is a reasoned decomposition of where time goes in a telephone speech agent, using typical values from public specifications, not a measurement of ours. When we have numbers we are prepared to defend, measured the way described above, they will appear in a note of their own.

References

  • ITU-T Recommendation G.114 (2003). One-way transmission time. International Telecommunication Union.
  • ITU-T Recommendation G.711 (1988). Pulse code modulation (PCM) of voice frequencies. International Telecommunication Union.
  • Valin, J.-M., Vos, K., and Terriberry, T. (2012). Definition of the Opus Audio Codec. IETF RFC 6716.
  • Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. (2022). High fidelity neural audio compression. arXiv:2210.13438.
  • Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., and Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26).
  • Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M. (2021). SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing.