The Self-Interruption Problem in Full-Duplex Voice AI
The hallmark of human conversation is full-duplex interactivity: both parties can speak simultaneously, interject, or interrupt seamlessly. When implementing voice agents in browsers or smart speakers, enabling "barge-in" is essential. The moment a user speaks while the AI is talking, the agent must instantly halt synthesis, clear its audio buffer, and listen.
However, without dedicated headphones, a critical physical flaw arises: Acoustic Speaker-to-Microphone Bleed. The audio emitted by the laptop or smartphone speakers travels through the physical air and strikes the built-in microphone. To a naive Voice Activity Detection (VAD) model running in the browser or on the edge server, this audio signal is indistinguishable from human speech. The agent detects "speech", assumes the user is interrupting, cuts itself off mid-sentence, and feeds its own synthesized voice into the speech-to-text pipeline as a hallucinated user query.
Understanding Acoustic Echo Cancellation (AEC3)
Eliminating this acoustic loop requires an adaptive Acoustic Echo Canceller (AEC). In modern WebRTC implementations, Google's AEC3 algorithm models the physical acoustic path between the speaker output and the microphone input as an adaptive Finite Impulse Response (FIR) filter.
The mechanism relies on two concurrent audio streams:
- The Near-End Signal (Microphone Input): Contains user speech, ambient room noise, AND the acoustic echo of the agent's speaker output.
- The Far-End Reference Signal (Speaker Output): The exact digital audio buffer that the application sent to the sound card for playback.
By delaying the reference signal to match the physical latency of the sound hardware and room acoustics, the AEC algorithm continuously calculates filter coefficients and subtracts the reference audio from the microphone stream in real time, outputting a clean audio stream containing only the user's authentic voice.
Browser Implementation via AudioWorklet and MediaStream Constraints
Standard HTML5 navigator.mediaDevices.getUserMedia enables browser-level hardware echo cancellation, but developers frequently break it by routing audio incorrectly. Below is the proper architecture using an AudioWorklet to maintain continuous, phase-aligned reference channels:
// 1. Request strict system AEC and disable AGC auto-gain fluctuations
const constraints = {
audio: {
echoCancellation: { ideal: true },
noiseSuppression: { ideal: true },
autoGainControl: { ideal: false }, // Prevent gain surging during AI speech
channelCount: 1,
sampleRate: 48000
}
};
const micStream = await navigator.mediaDevices.getUserMedia(constraints);
const audioCtx = new AudioContext({ sampleRate: 48000 });
const micSource = audioCtx.createMediaStreamSource(micStream);
// 2. Load dedicated VAD & Cross-Correlation Worklet
await audioCtx.audioWorklet.addModule('/static/js/audio/barge-in-processor.js');
const bargeInNode = new AudioWorkletNode(audioCtx, 'barge-in-processor');
// 3. Connect microphone to processor
micSource.connect(bargeInNode);
// Handle clean voice detection events from Worklet
bargeInNode.port.onmessage = (event) => {
if (event.data.type === 'AUTHENTIC_USER_BARGE_IN') {
console.log('Real user interruption confirmed! Signal server to abort TTS.');
websocket.send(JSON.stringify({ event: 'INTERRUPT_TRIGGERED' }));
}
};
Server-Side Secondary Residual Cancellation
Even with browser AEC enabled, non-linear speaker distortion (especially on low-cost mobile hardware with overdriven micro-speakers) leaves harmonic residual echo. On the Python backend, we implement a secondary spectral subtraction filter before passing audio to Silero VAD:
import numpy as np
def compute_normalized_cross_correlation(mic_chunk: np.ndarray, ref_chunk: np.ndarray) -> float:
# Computes Pearson correlation coefficient between microphone frame and reference playback.
# A correlation > 0.65 indicates residual speaker bleed, suppressing false barge-in triggers.
if np.std(mic_chunk) == 0 or np.std(ref_chunk) == 0:
return 0.0
corr = np.corrcoef(mic_chunk, ref_chunk)[0, 1]
return float(corr)
By pairing client-side WebRTC AEC3 with server-side cross-correlation validation, false barge-in triggers drop from over 35% of turns to under 0.2%. For enterprise telephony teams designing resilient audio pipelines, review our specialized Voice AI & Real-Time Telephony Services.