Kafka Labs

Kafka Labs  ·  San Francisco

Talking to a machine should feel like talking to a person.

Kafka Labs is an AI research lab and software company building speech-to-speech models and real-time voice agents.

Built by grads from

01—Why

Voice AI still sounds like a machine.

Most voice AI hears you, writes it down, thinks in text, then reads a reply aloud. Every hop adds a pause and strips the tone of voice.


Our models work in speech directly, so a conversation keeps its timing and its tone, and holds up on a real phone call.

02—Research

Field notes from the work.

Five directions, one question: what does it take for a model to hold a natural conversation in real time?

  1. 01

    Speech-to-speech modeling

    One pass, no text in between.

  2. 02

    Real-time streaming

    Answer within a human pause.

  3. 03

    Turn-taking

    Know when to speak and when to yield.

  4. 04

    Expressive voice

    Pacing and emphasis that fit the moment.

  5. 05

    Tool use

    Act without breaking the rhythm.

Research directions, not shipped product claims.

03—Technology

How our speech models run in real time.

What runs today, and what we plan to build with. None of this is a benchmark claim.

Today

  • Realtime gateway

    In development

    A Rust WebSocket gateway for two-way audio.

  • Model runtime

    In development

    Streams live audio through our own speech-to-speech models, keeping state for each conversation.

  • Telephony

    Private beta

    Phone-network audio for Turtle, on real numbers.

NVIDIA technologies we plan to use

  • Streaming ASR and TTS

    Planned

    NVIDIA Riva and NVIDIA NeMo for streaming recognition and synthesis alongside our models.

  • Optimizing our models on NVIDIA GPUs

    Planned

    NVIDIA NIM, TensorRT-LLM, and Triton to optimize and run our speech-to-speech models on NVIDIA GPUs through cloud providers.