End-to-End Speech-to-Speech (S2S) Models vs. Cascaded STT-LLM-TTS: Latency Budgets & Prosody in Voice AI

Cascaded voice pipelines suffer from cumulative serialization latency and strip away emotional inflection. Unpack the architecture of native Speech-to-Speech models delivering sub-300ms turnarounds.

The Fundamental Limit of Cascaded Voice Architectures

Over the past four years, conversational voice AI has predominantly relied on a cascaded tripartite architecture. Incoming user speech is processed in three distinct sequential steps:

User Audio ---> [ 1. ASR / Whisper ] ---> Text Tokens ---> [ 2. LLM / GPT-4 ] ---> Text Stream ---> [ 3. TTS / ElevenLabs ] ---> Output Audio

While engineering teams have heavily optimized individual components—using streaming ASR lookahead chunking, speculative decoding, and streaming WebSocket audio—the cascaded paradigm faces an insurmountable barrier: compounding serialization latency.

Pipeline Stage Minimum Realistic Production Latency Average Real-World Telephony Latency
Audio Capture & Acoustic VAD Frame 60 ms 120 ms
ASR Word Finalization 120 ms 180 ms
LLM Time to First Token (TTFT) 140 ms 250 ms
TTS Sentence Boundary Chunking 180 ms 300 ms
Network Jitter Buffer & WebRTC Playback 40 ms 80 ms
Total End-to-End Turnaround 540 ms 930 ms (Noticeable Hesitation)

1. The Information Loss Problem: Text as an Inadequate Bottleneck

Beyond latency, serializing speech into plain text discards critical human communicative information. Written text strips away pitch, cadence, hesitation, sarcastic inflection, whisper tones, and background emotional distress. If a caller sighs heavily or speaks with urgency, an ASR model produces the same flat string: "I need help right now."

Similarly, text-to-speech synthesis must guess the appropriate prosody from bare punctuation marks, frequently producing robotic or emotionally mismatched intonations.

2. The Speech-to-Speech (S2S) Architecture

Next-generation systems—such as Kyutai's Moshi, Mini-Omni, and OpenAI's GPT-4o Voice—eliminate text as the intermediate transport medium. Instead, these models operate as Audio-In / Audio-Out native multimodal transformers.

User Audio Stream ---> [ Neural Audio Codec (SNAC / EnCodec) ] ---> Acoustic Tokens ---> [ Dual-Stream S2S Transformer ] ---> Output Acoustic Tokens ---> [ Neural Vocoder ] ---> Audio Stream

3. Discrete Neural Audio Tokenization

Native S2S models convert continuous 24kHz audio waveforms into discrete tokens using Residual Vector Quantization (RVQ) codecs such as EnCodec, Mimi, or SoundStream. A single 20ms audio frame is compressed into 8 hierarchical codebook levels, representing both phonetic semantic content (lower levels) and fine acoustic prosody/speaker identity (higher levels).

# audio/s2s_gateway.py
import asyncio
import numpy as np

class SpeechToSpeechStreamer:
    """Simultaneous full-duplex audio token streaming gateway."""
    def __init__(self, s2s_client, codec_decoder):
        self.client = s2s_client
        self.decoder = codec_decoder
        self.audio_out_queue = asyncio.Queue()

    async def ingest_audio_chunks(self, pcm_stream):
        """Streams 20ms audio frames directly into the neural audio tokenizer."""
        async for frame in pcm_stream:
            # Codec encodes 20ms of audio (320 samples at 16kHz) into discrete tokens
            acoustic_tokens = self.decoder.encode_frame(frame)
            # S2S model processes audio tokens and yields response tokens simultaneously
            async for out_token in self.client.generate_stream(acoustic_tokens):
                synthesized_pcm = self.decoder.decode_token(out_token)
                await self.audio_out_queue.put(synthesized_pcm)

4. Architectural Comparison: Cascaded vs. Native S2S

Engineering Dimension Cascaded (STT $ o$ LLM $ o$ TTS) Native Speech-to-Speech (S2S)
Conversational Turnaround 650ms to 1,200ms 180ms to 280ms (Human Parity)
Prosody & Emotion Preservation Zero (Stripped by text normalization) Native (Acoustic token cross-attention)
Full-Duplex Interruption (Barge-In) Requires external VAD halting loops Native (Model listens while speaking)
Compute & Tool Calling Complexity Easy (Standard JSON structured outputs) Challenging (Requires parallel text heads)

While native S2S models present new challenges for deterministic tool execution and safety guardrails, they fundamentally redefine conversational latency, delivering sub-300ms turnaround with human-level expressive prosody. Learn more about telephony pipelines in our guide to WebRTC TWCC Congestion Control.

Real-Time Voice Agent Latency Budget Estimator

// Full-Duplex WebRTC Pipeline Waterfall
WebRTC / SFU Telemetry

Calculate end-to-end voice turnaround time across each stage of a bidirectional voice agent pipeline (User stops speaking → First synthetic audio byte received).

160 ms
Turn detection window
110 ms
Deepgram / Whisper chunking
210 ms
vLLM / Groq / OpenAI stream
130 ms
Cartesia / ElevenLabs / Melo
40 ms
Edge SFU Gateway
Estimated Total Turnaround Latency
650 ms
✓ Natural Conversational Flow (<700ms)
VAD STT LLM TTFT TTS 1st Packet WebRTC RTT
Engineering a low-latency voice pipeline? We build full-duplex WebRTC SFU systems with sub-700ms round trips.
// Real-Time Audio Telemetry • Voice AI Architecture Review

Engineering Real-Time Voice Agents or Low-Latency LLM Serving?

Achieving sub-700ms full-duplex conversational latency requires careful orchestration between WebRTC media gateways, continuous batching (vLLM), and neural TTS streaming. Let's inspect your pipeline waterfall together.

All Insights
Chat on WhatsApp