The Fundamental Limit of Cascaded Voice Architectures
Over the past four years, conversational voice AI has predominantly relied on a cascaded tripartite architecture. Incoming user speech is processed in three distinct sequential steps:
User Audio ---> [ 1. ASR / Whisper ] ---> Text Tokens ---> [ 2. LLM / GPT-4 ] ---> Text Stream ---> [ 3. TTS / ElevenLabs ] ---> Output Audio
While engineering teams have heavily optimized individual components—using streaming ASR lookahead chunking, speculative decoding, and streaming WebSocket audio—the cascaded paradigm faces an insurmountable barrier: compounding serialization latency.
| Pipeline Stage | Minimum Realistic Production Latency | Average Real-World Telephony Latency |
|---|---|---|
| Audio Capture & Acoustic VAD Frame | 60 ms | 120 ms |
| ASR Word Finalization | 120 ms | 180 ms |
| LLM Time to First Token (TTFT) | 140 ms | 250 ms |
| TTS Sentence Boundary Chunking | 180 ms | 300 ms |
| Network Jitter Buffer & WebRTC Playback | 40 ms | 80 ms |
| Total End-to-End Turnaround | 540 ms | 930 ms (Noticeable Hesitation) |
1. The Information Loss Problem: Text as an Inadequate Bottleneck
Beyond latency, serializing speech into plain text discards critical human communicative information. Written text strips away pitch, cadence, hesitation, sarcastic inflection, whisper tones, and background emotional distress. If a caller sighs heavily or speaks with urgency, an ASR model produces the same flat string: "I need help right now."
Similarly, text-to-speech synthesis must guess the appropriate prosody from bare punctuation marks, frequently producing robotic or emotionally mismatched intonations.
2. The Speech-to-Speech (S2S) Architecture
Next-generation systems—such as Kyutai's Moshi, Mini-Omni, and OpenAI's GPT-4o Voice—eliminate text as the intermediate transport medium. Instead, these models operate as Audio-In / Audio-Out native multimodal transformers.
User Audio Stream ---> [ Neural Audio Codec (SNAC / EnCodec) ] ---> Acoustic Tokens ---> [ Dual-Stream S2S Transformer ] ---> Output Acoustic Tokens ---> [ Neural Vocoder ] ---> Audio Stream
3. Discrete Neural Audio Tokenization
Native S2S models convert continuous 24kHz audio waveforms into discrete tokens using Residual Vector Quantization (RVQ) codecs such as EnCodec, Mimi, or SoundStream. A single 20ms audio frame is compressed into 8 hierarchical codebook levels, representing both phonetic semantic content (lower levels) and fine acoustic prosody/speaker identity (higher levels).
# audio/s2s_gateway.py
import asyncio
import numpy as np
class SpeechToSpeechStreamer:
"""Simultaneous full-duplex audio token streaming gateway."""
def __init__(self, s2s_client, codec_decoder):
self.client = s2s_client
self.decoder = codec_decoder
self.audio_out_queue = asyncio.Queue()
async def ingest_audio_chunks(self, pcm_stream):
"""Streams 20ms audio frames directly into the neural audio tokenizer."""
async for frame in pcm_stream:
# Codec encodes 20ms of audio (320 samples at 16kHz) into discrete tokens
acoustic_tokens = self.decoder.encode_frame(frame)
# S2S model processes audio tokens and yields response tokens simultaneously
async for out_token in self.client.generate_stream(acoustic_tokens):
synthesized_pcm = self.decoder.decode_token(out_token)
await self.audio_out_queue.put(synthesized_pcm)
4. Architectural Comparison: Cascaded vs. Native S2S
| Engineering Dimension | Cascaded (STT $ o$ LLM $ o$ TTS) | Native Speech-to-Speech (S2S) |
|---|---|---|
| Conversational Turnaround | 650ms to 1,200ms | 180ms to 280ms (Human Parity) |
| Prosody & Emotion Preservation | Zero (Stripped by text normalization) | Native (Acoustic token cross-attention) |
| Full-Duplex Interruption (Barge-In) | Requires external VAD halting loops | Native (Model listens while speaking) |
| Compute & Tool Calling Complexity | Easy (Standard JSON structured outputs) | Challenging (Requires parallel text heads) |
While native S2S models present new challenges for deterministic tool execution and safety guardrails, they fundamentally redefine conversational latency, delivering sub-300ms turnaround with human-level expressive prosody. Learn more about telephony pipelines in our guide to WebRTC TWCC Congestion Control.