The Latency Divide: Telephony PSTN vs. Modern WebRTC
Deploying conversational AI voice agents on the web is relatively straightforward thanks to native browser WebRTC APIs and WebSockets. However, enterprise adoption demands phone connectivity: allowing customers to dial a regular phone number and converse with an autonomous AI voice agent in real time. Bridging the Public Switched Telephone Network (PSTN) into modern AI pipelines presents a unique architectural challenge.
Traditional PSTN telephony relies on Session Initiation Protocol (SIP) negotiating uncompressed 8kHz narrowband audio (G.711 μ-law or A-law) over Real-time Transport Protocol (RTP). In contrast, modern WebRTC engines (like LiveKit, Janus, or Mediasoup) expect 48kHz wideband Opus audio over encrypted DTLS-SRTP. Without an optimized media bridge, transcoding between these protocols introduces 200ms to 400ms of latency—destroying conversational fluidness before your LLM even receives audio.
1. Architecture of the SIP-to-WebRTC Ingestion Pipeline
An enterprise-grade telephony bridge bypasses heavy intermediary audio servers by routing inbound calls through Twilio Elastic SIP Trunking directly into a dedicated LiveKit SIP Ingress service:
- Inbound PSTN Termination: A user dials your business phone number. Twilio terminates the PSTN call and converts it into a SIP INVITE sent over secure TLS to your LiveKit SIP Gateway.
- Room Dispatch & Participant Joining: The LiveKit SIP server extracts the caller's phone number from the SIP headers, dynamically creates a virtual WebRTC room, and admits the caller as a simulated WebRTC participant.
- Audio Transcoding Pipeline: Hardware-accelerated G.711 to Opus transcoding converts the 8kHz incoming audio stream into 48kHz stereo frames with sub-10ms packetization latency.
- AI Worker Connection: An asynchronous Python agent worker (built with FastAPI and LiveKit Agents SDK) subscribes to the room's incoming audio track and pumps raw PCM buffers into real-time speech models.
2. Orchestrating the Inbound Call Handler in Python
Using the LiveKit Agents framework, your Python worker automatically joins the room upon caller ingress, spins up speech-to-text, queries your LLM, and begins real-time speech synthesis:
import asyncio
import logging
from livekit import agents, rtc
from livekit.agents import JobContext, WorkerOptions, cli
from livekit.plugins import deepgram, openai, silero
logger = logging.getLogger("telephony-bridge")
async def entrypoint(ctx: JobContext):
# Connect to the dynamically provisioned telephony room
await ctx.connect(auto_subscribe=agents.AutoSubscribe.AUDIO_ONLY)
# Identify caller telephone number from participant attributes
caller = await ctx.wait_for_participant()
caller_phone = caller.attributes.get("sip.phoneNumber", "Unknown Caller")
logger.info(f"Incoming call connected from: {caller_phone}")
# Initialize low-latency voice pipeline components
vad = silero.VAD.load()
stt = deepgram.STT(model="nova-2", sample_rate=16000)
llm = openai.LLM(model="gpt-4o-mini", temperature=0.3)
tts = openai.TTS(voice="alloy")
# Construct the bidirectional voice assistant
assistant = agents.VoiceAssistant(
vad=vad,
stt=stt,
llm=llm,
tts=tts,
fnc_ctx=agents.FunctionContext(),
)
# Attach assistant to the room audio track
assistant.start(ctx.room, caller)
await assistant.say("Thank you for calling devManue Engineering. How can I direct your project inquiry today?")
if __name__ == "__main__":
cli.run_app(WorkerOptions(entrypoint_fnc=entrypoint))
3. Handling DTMF Signals & Telephone Line Disconnections
Unlike browser WebRTC connections where users simply close a tab, telephone users press keypad numbers (DTMF) to navigate menus or abruptly hang up. Your bridge must capture SIP INFO packets or RFC 2833 out-of-band DTMF digits:
- Keypad Tone Capture: Listen for
dtmf_receivedevents on the SIP participant to instantly route callers or capture sensitive numerical data (such as account numbers) without exposing them to speech recognition errors. - Clean Call Termination: When the caller hangs up, Twilio transmits a SIP
BYEmessage. The LiveKit bridge must immediately tear down the virtual room, cancel in-flight LLM token generation, and flush transcription logs to PostgreSQL to avoid lingering cloud compute costs.
"Telephony voice agents live and die by packet turnaround. Squeezing your audio ingestion and transcoding pipeline down to under 30 milliseconds is what enables a sub-800ms natural conversational rhythm on regular cell phones."
For related production architectures and system implementations, explore these companion guides:
- WebRTC SFU Architecture: Voice AI & Jitter Buffers — Manage packet loss and jitter compensation when bridging SIP calls to WebRTC audio tracks.
- Deploying Real-Time WebSockets & Voice AI — Handle full-duplex audio stream connections between telephony providers and media servers.
- Sub-Second Audio Chunking & Turn Detection for Real-Time LLM Voice — Stream low-latency audio chunks into real-time speech models for instant response.
Key Architectural Takeaways
Building scalable telephone voice AI agents requires eliminating traditional Asterisk or FreeSWITCH complexity in favor of modern cloud-native WebRTC SFUs. By terminating SIP trunks directly into LiveKit, transcoding audio at the network boundary, and managing caller state with asynchronous Python workers, you deliver high-fidelity, human-grade conversational telephony to standard smartphones globally.