Kafka Labs · San Francisco
Talking to a machine should feel like talking to a person.
Kafka Labs is an AI research lab and software company building speech-to-speech models and real-time voice agents.
Built by grads from
01—Why
Voice AI still sounds like a machine.
Most voice AI hears you, writes it down, thinks in text, then reads a reply aloud. Every hop adds a pause and strips the tone of voice.
Our models work in speech directly, so a conversation keeps its timing and its tone, and holds up on a real phone call.
02—Research
Field notes from the work.
Five directions, one question: what does it take for a model to hold a natural conversation in real time?
- 01
Speech-to-speech modeling
One pass, no text in between.
- 02
Real-time streaming
Answer within a human pause.
- 03
Turn-taking
Know when to speak and when to yield.
- 04
Expressive voice
Pacing and emphasis that fit the moment.
- 05
Tool use
Act without breaking the rhythm.
Research directions, not shipped product claims.
03—Technology
How our speech models run in real time.
What runs today, and what we plan to build with. None of this is a benchmark claim.
Today
Realtime gateway
In developmentA Rust WebSocket gateway for two-way audio.
Model runtime
In developmentStreams live audio through our own speech-to-speech models, keeping state for each conversation.
Telephony
Private betaPhone-network audio for Turtle, on real numbers.
NVIDIA technologies we plan to use
Streaming ASR and TTS
PlannedNVIDIA Riva and NVIDIA NeMo for streaming recognition and synthesis alongside our models.
Optimizing our models on NVIDIA GPUs
PlannedNVIDIA NIM, TensorRT-LLM, and Triton to optimize and run our speech-to-speech models on NVIDIA GPUs through cloud providers.

