Kafka Labs

Research note—October 2026—11 min read

The case for speech-to-speech

Why the transcribe, think, synthesize pipeline cannot be fixed with faster parts, what a speech-native model has to learn instead, and where the hard problems are.

Kafka Labs · San Francisco


Every voice assistant you have ever talked to has a tell. You finish a sentence, and there is a pause. Not a thoughtful pause, not the breath a person takes before answering a hard question, but a dead, flat silence of half a second or more, followed by a reply read in a voice that has clearly never heard what you just said. The reply may be correct. It may even be helpful. It still does not feel like a conversation, and the reason is architectural, not cosmetic.

This note lays out the case for building speech-to-speech models: systems that take audio in and put audio out, with no transcript in between. It explains why the dominant pipeline cannot be fixed by making its parts faster, what a speech-native model has to learn instead, and where the genuinely hard problems are. It is a statement of a research bet, not a report of results. We will publish results when we have them.

The pause that gives it away

Human conversation runs on a remarkably tight clock. Across languages and cultures, the typical gap between one speaker finishing and the next beginning is on the order of two hundred milliseconds, and a large fraction of transitions have no gap at all or a slight overlap (Stivers et al., 2009). That number is striking because producing even a short utterance takes far longer than that. Planning a response, retrieving words, and preparing articulation takes somewhere around six hundred milliseconds and often more (Levinson and Torreira, 2015). The arithmetic only works one way: people begin planning their reply while the other person is still talking. Listening and speaking are not phases that alternate. They run concurrently, and the listener is continuously predicting when the current turn will end.

Now consider the pipeline that almost every voice product uses today. Audio is captured and run through a voice activity detector. When the detector sees enough silence, typically five hundred to eight hundred milliseconds of it, it declares the turn over. The accumulated audio is sent to a speech recognizer, which produces a final transcript. The transcript goes to a language model, which begins generating text. The first sentence of that text goes to a speech synthesizer, which produces audio, which is buffered and played back.

Each stage can be made fast, and vendors compete hard on exactly that. But notice what the structure forces. Nothing downstream can begin until the endpointer has decided the person is done, and the endpointer cannot decide that until a silence has elapsed. That silence is not a processing delay you can engineer away. It is the system waiting to find out whether you have stopped. A human listener does not wait for the silence; they project the end of your turn from your syntax, your intonation, and what you are plausibly trying to say. A cascade has no mechanism to do that, because the component that understands language only sees text, and only after the fact.

So the floor on a cascade's response time is roughly the endpointing wait plus the recognizer's finalization plus the language model's time to first token plus the synthesizer's time to first audio plus buffering. Make every stage instantaneous and you are still left with the endpointing wait, which is already two to four times the human gap. This is why a faster cascade still feels like a cascade.

What is lost in the transcript

Latency is the obvious cost. The subtler cost is what the transcript throws away.

A transcript is a projection of speech onto text, and the projection is lossy in exactly the dimensions that make conversation feel human. Prosody carries emphasis, contrast, and attitude: "I said the Tuesday appointment" and "I said the Tuesday appointment" are the same string. Rate and hesitation carry uncertainty. Laughter, sighs, and in-breaths carry state. The small vocal particles that listeners produce while someone else is talking, the "mm-hm" and "right" and "okay" that conversation analysts call backchannels, are not requests for the floor and should not be treated as turns, but a transcript flattens them into words. Overlap, which is common and usually cooperative in human talk, becomes a transcription error to be cleaned up.

The loss runs in the other direction too. A text-to-speech system receives a string and must invent prosody for it. It cannot know which word the language model meant to stress, because the language model never decided; it produced tokens, not an intonation contour. The result is the flat read that everyone recognizes: fluent, clear, and emotionally disconnected from what was just said. The pipeline has two models that each understand half of the problem and a text bottleneck between them that guarantees neither can use what the other knows.

None of this is an argument that text is useless. Text is an extraordinary compression of meaning, and the best reasoning we know how to build runs on it. The argument is that text should be a representation inside the model, not a wall between two models.

What a speech-native model is

A speech-to-speech model is a sequence model whose inputs and outputs include audio directly. The practical route to that, established over the last few years, has three parts.

The first part is a discrete representation of audio. Neural audio codecs such as SoundStream (Zeghidour et al., 2021) and EnCodec (Défossez et al., 2022) learn to encode a waveform into a small number of integer tokens per frame, using residual vector quantization so that a handful of codebooks captures first the coarse structure of the sound and then progressively finer detail. Audio language models such as AudioLM (Borsos et al., 2023) showed that a transformer trained on these tokens can continue speech coherently, and introduced the useful distinction between semantic tokens, which capture what is being said, and acoustic tokens, which capture how it sounds.

The second part is a model that treats these audio tokens as a language. Once speech is a token stream, the same architectures that model text can model it, and the same scaling behavior appears to apply. The model can be trained to predict the next frame of audio given everything that has been heard and said so far.

The third part, and the one that matters most for a conversational agent, is keeping text in the loop without putting it in the way. A model that only ever sees audio tokens has a hard time inheriting the knowledge and reasoning of a text-trained language model. The approach that has worked in practice is to have the model predict a text stream alongside its audio stream, so that the words it is about to say are represented explicitly as a kind of inner monologue, and the audio is generated conditioned on them. Moshi (Défossez et al., 2024) is the clearest public demonstration of this design: a full-duplex model that generates its own speech while continuously listening to the user's, with an aligned text stream that anchors the content of what it says.

The important property of this design is that there is no endpointer and no handoff. The model is always listening and always deciding, frame by frame, whether to stay silent, produce a backchannel, or begin a turn. Timing becomes something the model learns from data rather than something a threshold imposes.

The hard parts

Describing the architecture is the easy part. Here is where we expect to spend our time.

The sample rate of thought. Audio tokens arrive at tens of frames per second, far more than the handful of text tokens per second a person speaks. A conversation that fits comfortably in a few thousand text tokens becomes tens of thousands of audio tokens. Context length, memory, and cost all scale with that. Codec design, token rate, and how much of the history to keep in audio versus text are open engineering questions with real trade-offs between fidelity and tractability.

Keeping the mind while changing the mouth. Training a text-capable model on speech can erode what it knew. The model has to learn a new modality without forgetting how to follow instructions, use tools, and reason about the caller's problem. How much text to interleave, which parameters to adapt, and how to schedule the mixture of text and speech during training are decisions that determine whether you end up with a model that talks well or one that talks well and is also useful.

Data. There is a great deal of recorded speech and a great deal of text, but comparatively little recorded, two-channel, naturally timed conversation of the kind a phone agent has to handle, with the interruptions and overlaps and hold music left in. Synthesizing dialog from text with voice generation helps with content but teaches nothing about timing, because the timing in synthetic dialog is whatever you scripted. We think the honest path is a mixture: synthetic data for breadth, real conversational recordings for timing, and the model's own deployed conversations, with consent, for the long tail of what actually happens on a call.

Evaluation. There is no single number. Word error rate measures the wrong thing. Listening tests measure naturalness but not behavior. What we care about is behavioral: how quickly the model responds relative to the end of the user's turn, how often it cuts people off, whether it yields when interrupted, whether it continues through a backchannel, whether its prosody matches its content. Each of these needs its own measurement, and most of them need two-channel recordings with timestamps to measure at all. We have more to say about this in the companion notes on turn-taking and latency.

Safety and control. A model that generates speech directly can also generate speech that sounds like a particular person. Voice identity has to be a controlled input, not an emergent property, and the model's willingness to say things has to be governed with the same care as a text model's, with the added wrinkle that speech is harder to filter after the fact. We treat this as a first-class design constraint rather than a feature to add later.

What speech-native does not mean

It is worth being precise about the claim, because "end-to-end" has become a slogan that covers several different designs.

It does not mean the model has no notion of words. The inner-monologue design keeps an explicit text stream, and we think that is correct: words are how the model reasons, plans tool calls, and stays grounded in what it knows. What changes is that the text is produced in lockstep with the audio rather than ahead of it, so the model can decide to stress a word, slow down, or stop mid-sentence because the caller has started talking, and the words it has already committed to are the words it actually said.

It does not mean one monolithic network does everything. There may be a codec, a backbone, and a lightweight audio decoder, and in a telephone deployment there will certainly be a gateway handling the network, the codecs, and the call itself. What matters is that no stage in the path waits for a later stage to finish before the earlier one can start, and that no stage discards information the next one needs.

It does not mean abandoning recognition and synthesis as components. Streaming recognizers and synthesizers are mature, and there are places in a product where a transcript is exactly what you want: a summary of the call, a search over past conversations, an accessibility feature. We plan to use them for those jobs. The claim is narrower: the live loop between a caller's mouth and the agent's voice should not pass through text as a bottleneck.

And it does not mean the problem is solved by architecture alone. A full-duplex model with bad timing is still rude; one with good timing and shallow content is still useless. The architecture removes a ceiling. Clearing the space under it is the research.

Why ship early

A research lab could reasonably spend years on the problems above before letting a model near a telephone. We have chosen not to, for a reason that is more empirical than commercial. The distribution of what happens on real calls is not something you can imagine your way into. People trail off. They answer a different question than the one asked. They put the phone down to find a document and come back mid-sentence. The hold music has a voice in it. A model trained only on clean, scripted dialog will be confidently wrong about all of this, and the only way to find out what the long tail contains is to be in it.

So our products are not a separate activity from the research; they are the instrument. Turtle, our voice agent for live phone calls, is where our models meet real callers, and the behaviors we measure there feed directly back into what we train next.

What we are not claiming

We do not have a model that matches human timing. We have not measured our systems against the figures cited above in a way we would be comfortable publishing. The public work we cite belongs to other groups, and we are building on it, not improving on it, until we can show otherwise. This note describes a direction and the reasoning behind it. The results, when they exist, will get their own notes.

References

  • Borsos, Z., Marinier, R., Vincent, D., Kharitonov, E., Pietquin, O., Sharifi, M., Roblek, D., Teboul, O., Grangier, D., Tagliasacchi, M., and Zeghidour, N. (2023). AudioLM: A language modeling approach to audio generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing.
  • Défossez, A., Copet, J., Synnaeve, G., and Adi, Y. (2022). High fidelity neural audio compression. arXiv:2210.13438.
  • Défossez, A., Mazaré, L., Orsini, M., Royer, A., Pérez, P., Jégou, H., Grave, E., and Zeghidour, N. (2024). Moshi: A speech-text foundation model for real-time dialogue. Kyutai technical report.
  • Levinson, S. C., and Torreira, F. (2015). Timing in turn-taking and its implications for processing models of language. Frontiers in Psychology, 6:731.
  • Stivers, T., Enfield, N. J., Brown, P., Englert, C., Hayashi, M., Heinemann, T., Hoymann, G., Rossano, F., de Ruiter, J. P., Yoon, K.-E., and Levinson, S. C. (2009). Universals and cultural variation in turn-taking in conversation. Proceedings of the National Academy of Sciences, 106(26).
  • Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., and Tagliasacchi, M. (2021). SoundStream: An end-to-end neural audio codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing.