Kafka Labs

Research note—October 2026—11 min read

Turn-taking is a model problem, not a timer problem

What people actually do when they hand the floor back and forth, why silence thresholds cannot reproduce it, and what a model that handles turn-taking natively has to learn.

Kafka Labs · San Francisco


Ask anyone who has built a voice product what the hardest tuning knob is, and they will tell you about the silence threshold. Set it too short, and the agent barges in while the caller is still thinking. Set it too long, and every exchange has a dead beat in it. Engineers move it in fifty-millisecond steps, add heuristics for trailing "um"s, and never find a setting that works for everyone, because no such setting exists. The knob is a symptom. The disease is that the system decides whose turn it is with a timer, and conversation does not run on timers.

This note is about turn-taking as a modeling problem. It covers what people actually do when they hand the floor back and forth, why endpointing by silence cannot reproduce it, what a model that handles turn-taking natively has to represent, where the training signal for that can come from, and how we intend to measure whether it works. As with our other notes, it describes a direction and the reasoning behind it rather than results.

The timer

The standard design is simple. A voice activity detector classifies each short frame of audio as speech or not. When the detector has seen a run of non-speech frames longer than a threshold, the system declares that the user's turn has ended, and the rest of the pipeline begins. While the agent is speaking, the detector is typically either switched off, so the agent cannot be interrupted at all, or used as a barge-in trigger that stops playback the moment any sound is heard.

Each half of this design fails in a characteristic way. On the listening side, silence is a poor proxy for completion. People pause inside turns all the time, to think, to find a word, to look something up, and those pauses are often longer than the pauses between turns. A threshold long enough to ride through a mid-turn pause is long enough to feel sluggish at a real transition. Conversely, people frequently finish a turn without any silence at all; the next speaker simply begins. A silence-based endpointer cannot even represent that case.

On the speaking side, barge-in by sound level treats every noise as an interruption. A caller who says "mm-hm" to signal they are following, a cough, a television in the background, or the caller's own echo returning down a bad line will all stop the agent mid-sentence. So products add a louder threshold, or require the sound to last a while, and now the agent talks over genuine interruptions for half a second before it notices. There is no setting that distinguishes "go on" from "stop" by amplitude, because the difference is not acoustic. It is pragmatic.

What people actually do

Conversation analysis has studied the mechanics of turn-taking for fifty years, and the findings are consistent enough to serve as design requirements.

The foundational account (Sacks, Schegloff, and Jefferson, 1974) describes turns as built from units whose possible completion points listeners can anticipate. At each such point, a transition may occur: the current speaker may select the next, someone may self-select, or the current speaker may continue. Transitions are negotiated at these points, not at arbitrary moments, and listeners are very good at predicting where they fall.

The timing data bear this out. Gaps between turns cluster around two hundred milliseconds across a wide range of languages (Stivers et al., 2009), far too short for the next speaker to have started planning after hearing the end of the previous turn. Levinson and Torreira (2015) lay out the implication: listeners must be predicting the content and the end of the current turn while it is still in progress, and launching their own production ahead of time. Overlap is not an error in this picture. Heldner and Edlund (2010) measured the distribution of gaps and overlaps in natural conversation and found that a substantial share of transitions involve some overlap, most of it brief.

Three further observations matter for an agent.

First, the cues that let listeners project a turn's end are multimodal and mostly linguistic. Syntax tells you whether the sentence is complete. Intonation tells you whether the speaker is winding down or setting up a continuation. Pragmatics tells you whether the question has been answered. Silence is the weakest of these signals, and the one humans rely on least.

Second, a great deal of what listeners say is not a turn at all. Backchannels, the small vocalizations and words produced while the other person holds the floor, signal attention and agreement without claiming the floor. A system that treats them as interruptions will stop talking every few seconds. A system that ignores them loses a channel of feedback that humans use to adjust what they are saying.

Third, when overlap does occur, people resolve it fast. A speaker who is interrupted competitively usually drops out within a few hundred milliseconds, often mid-word, and then does something sensible afterward: resumes, restarts, or abandons the point. Yielding late is heard as rudeness. Yielding and then losing the thread is heard as incompetence.

Skantze (2021) reviews the attempts to build these behaviors into dialogue systems and robots. The consistent lesson is that treating turn-taking as a prediction problem over rich input, rather than a detection problem over silence, is what moves the needle.

A taxonomy of events

If the agent is going to behave well, it needs to tell these situations apart, in real time, from the audio alone:

  • Completion. The caller has finished a turn and expects a reply. Respond promptly, ideally within the human range.
  • Pause within a turn. The caller has stopped but is not done: a thinking pause, a word search, a trailing "and, um". Wait, and signal attention if the pause grows long.
  • Backchannel. The caller says "right" or "okay" while the agent is talking. Continue, perhaps with a slight acknowledgment in prosody.
  • Cooperative overlap. The caller finishes the agent's sentence or answers before the question is fully asked. Yield gracefully and move on; do not repeat the completed part.
  • Competitive interruption. The caller wants the floor now. Stop within a few hundred milliseconds, even mid-word, and listen.
  • Barge-in on a list. The agent is reading options and the caller picks one. Stop and act on it; do not finish the list.
  • Hold. The caller says "hang on" and goes away. Wait, tolerate long silence, and re-engage gently when sound returns.
  • Noise. A door, a cough, a passing truck. Ignore it.

Every one of these turns on what is being said and how, not on how loud it is or how long the silence lasts. That is the sense in which turn-taking is a model problem.

Full-duplex modeling

The architectural answer is a model that listens and speaks at the same time, in the way the data says people do. Concretely, that means a model that consumes the caller's audio stream and produces its own audio stream frame by frame, with each output frame conditioned on everything heard up to that instant, including what the caller said while the agent was mid-sentence. Moshi (Défossez et al., 2024) demonstrated this multi-stream design publicly, modeling the user's and the system's audio as parallel token streams with a shared backbone.

Within that design, turn-taking stops being a separate module and becomes a set of things the model predicts:

  • Whether to produce silence or speech on the next output frame. This is the moment-to-moment floor decision, and it replaces the endpointer.
  • Whether the current input is a turn, a backchannel, or noise, as a latent the model infers rather than a label it is given.
  • When the caller's turn will end, so the model can prepare content before the end arrives and respond at human speed rather than after a wait.
  • When to stop. If the caller starts a competitive interruption, the model's own next frames should become silence quickly, which requires that the model's generation be interruptible at the frame level, not the sentence level.

The last point has a systems consequence that is easy to miss. A design that generates a sentence of text and then synthesizes it cannot stop mid-word in any meaningful sense; it can only cut playback. A model that generates audio frame by frame can stop generating, and the words it did not say were never committed. That difference is what allows the agent to resume coherently afterward: it knows exactly what the caller heard.

Where the training signal comes from

Behavior like this is learned, so the question is where the examples come from.

The best source is recorded two-channel conversation, where each participant's audio is on its own track. From the timing alone, without any manual annotation, you can derive who was speaking when, where the gaps and overlaps fell, which overlaps were brief and which were takeovers, and how quickly the yielding party stopped. Telephone conversations are especially valuable because they match the channel our agent lives on: narrowband audio, no visual cues, and all the real-world mess of lines and environments.

Synthetic dialog, generated from text and voiced by a synthesizer, is useful for content and coverage but nearly useless for timing. Whatever timing you script is the timing the model will learn, and scripted timing is not human timing. We use it where it helps and keep it out of the parts of training that are about when to talk.

The deployed agent is the third source, and over time probably the most important one. Every real call is a two-channel recording with natural timing, and every moment where the agent interrupted, hesitated too long, or talked through a backchannel is a labeled failure once you measure the caller's reaction. With consent and careful handling, those recordings close the loop between what the model does and what it should have done.

Auxiliary objectives help the model learn the right structure. Predicting, from the input stream, whether the caller's turn will end within the next few hundred milliseconds gives the model an explicit handle on projection. Predicting whether a given input segment is a backchannel gives it a handle on floor state. These are cheap to derive from timing and make the learned behavior easier to inspect.

Evaluating turn-taking

Because there is no single number, we are building a small battery of measurements, all of which require two-channel recordings with reliable timestamps.

  • Response offset. The time from the end of the caller's turn to the start of the agent's. The target is a distribution, not a mean: humans center around two hundred milliseconds with a long tail, and an agent that is always exactly on time will feel mechanical in its own way.
  • False cut-in rate. How often the agent begins speaking during a within-turn pause. This is the cost of being fast.
  • Yield latency. When the caller interrupts competitively, how long until the agent is silent. Hundreds of milliseconds is the human range; a full sentence is a failure.
  • Backchannel tolerance. How often the agent stops for an acknowledgment that was not a bid for the floor.
  • Recovery quality. After an interruption, does the agent resume sensibly, repeat the right amount, and avoid assuming the caller heard the cut-off part?

Alongside the measurements, a set of adversarial scripts: the caller who trails off, the caller who thinks aloud for three seconds, the caller who answers with "yeah, no, I mean", the caller with a television on, the caller who says "okay" after every clause. These do not produce scores so much as failure modes, and failure modes are what we need in order to decide what to train on next.

Recovery

One detail deserves its own section because it is where most systems that get interrupted fall apart. After the agent yields, what should it do?

The agent has two histories that have now diverged: what it intended to say and what the caller actually heard. If it was cut off in the middle of a sentence, the caller heard a fragment. If the caller's interruption was a question, the agent should answer it and then decide whether the abandoned sentence still matters. If the interruption was a correction, the abandoned sentence was probably wrong and should be dropped. If it was cooperative, the sentence was already understood and should not be repeated.

A model that generates audio frame by frame knows the boundary precisely, because it stopped at a frame. A model that generates text and then speech does not; it has to guess how much of its sentence was played before playback was cut. We think precise knowledge of what was heard is a prerequisite for good recovery, and it is one of the quieter arguments for the speech-native design.

Where we are

We are building toward a full-duplex model trained on two-channel telephone conversation, with explicit timing objectives and the evaluation battery above. The turn-taking behavior of our deployed agent today is the thing we are least satisfied with and most focused on, and it is the main reason the agent runs on real calls in a private beta rather than on a demo page. We will publish measurements against the human figures cited here when we can stand behind them.

References

  • Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N. (2024). Moshi: A speech-text foundation model for real-time dialogue. Kyutai technical report.
  • Heldner, M., and Edlund, J. (2010). Pauses, gaps and overlaps in conversations. Journal of Phonetics, 38(4).
  • Levinson, S. C., and Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6:731.
  • Sacks, H., Schegloff, E. A., and Jefferson, G. (1974). A simplest systematics for the organization of turn-taking for conversation. Language, 50(4).
  • Skantze, G. (2021). Turn-taking in conversational systems and human-robot interaction: A review. Computer Speech & Language, 67.
  • Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., and Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26).