The machinery
Acoustic echo cancellation, explained simply
Somewhere inside every calling app you use is a small, relentless piece of software whose only job is to stop the other person's voice from making a round trip through your room. It's called an acoustic echo canceller — AEC — and it succeeds so often that most people don't know it exists. You only meet it on the day it fails. This page explains what it does, how, and why it sometimes loses — in plain language, no equations.
The problem AEC solves
When you're on speakers, your device does two things at once: it plays the far side's voice into your room, and it listens to your room through your mic. Inevitably, the mic hears the speakers. Without intervention, everything the far side says would be captured and sent straight back to them — they'd hear their own voice returning, delayed by the network. That's call echo, and the loop happens on your device even though they're the ones who suffer it. (This asymmetry — the sufferer is not the cause — is the single most useful diagnostic fact about echo; the start-here guide is built on it.)
The trick: subtract what you know you played
Here's the insight that makes AEC possible: your device has perfect knowledge of what it just sent to the speakers. So in principle, echo removal is subtraction — take the mic signal, subtract the speaker signal, and what remains is only what's genuinely new in the room: your voice.
In practice, you can't subtract the speaker signal directly, because that's not what the mic hears. The mic hears the speaker signal after your room has had its way with it — delayed by the trip across your desk, colored by your speakers' character, smeared by reflections off the wall, the window, the mug. The canceller must first answer: "when I play X, what does my mic hear?" That answer is called the echo path, and AEC's core is an adaptive filter that learns it continuously — playing audio, watching what comes back, updating its model hundreds of times a second. Once the model is good, the canceller predicts the echo before it's captured and subtracts the prediction. What survives subtraction gets a final polish from a suppressor that mops up the residue the model missed.
All of this happens in a few milliseconds per frame of audio, forever, invisibly, on every call you take.
Why such a clever system still fails
- The model can't keep up. The echo path isn't static — you lean back, move the laptop, someone opens a door, and the room's response changes. The filter re-learns fast, but a burst of echo can leak through right after any change. Echo that flares when you shift around is this.
- Loudness breaks the math. The adaptive filter assumes the echo is a linear transformation of what was played. Push small laptop speakers hard and they distort — the mic now hears sounds that were never in the reference signal, and subtraction can't remove what was never modeled. This is why "turn your speakers down" is real engineering advice, not a brush-off.
- The reference signal goes missing. Route your audio through virtual devices, mixers, or loopback tools, and the app's canceller may no longer know what's actually being played — no reference, no subtraction. A surprising amount of suddenly-appearing echo is exactly this.
- Double-talk. When both sides speak at once, the mic carries echo and real speech, and the canceller must subtract one without gouging the other. Cheap implementations freeze or clip; that's why some calls chop your first word off whenever you interrupt. We wrote it up separately: the double-talk problem.
- Two clocks, one signal. Playback and capture sometimes run on subtly different hardware clocks (common with mismatched devices). The reference drifts against reality and the model chases it forever. This one produces stubborn, low-level echo that no setting seems to fix — and is a classic case where a different canceller in the chain behaves differently.
Acoustic echo vs line echo
Older telephony had a different echo: line echo, born not from speakers and mics but from electrical reflections in the phone network itself — signals bouncing at the junction where two-wire home circuits met four-wire trunk lines. Carriers deployed line echo cancellers inside the network, and the discipline of echo cancellation grew up there decades before video calls existed. Modern app-to-app calls don't traverse those junctions, so the echo you fight today is almost always acoustic — made of air, in somebody's room. The name "acoustic echo cancellation" exists to mark that distinction.
AEC is not noise suppression (and not de-reverb)
Three technologies ride together in modern voice tools and get blurred into one word, but they solve different problems:
| Layer | Removes | How it knows what to remove |
|---|---|---|
| Acoustic echo cancellation | Copies of the far side's voice re-entering your mic | Has the reference signal — it knows exactly what was played |
| Noise suppression | Keyboards, fans, traffic, chatter — sound that isn't speech | Learned/statistical models of "voice vs everything else" — no reference exists |
| De-reverberation | Your own room's smear on your own voice (the hollow sound) | Models of how rooms smear speech — again, no clean reference |
The practical consequence: a tool can be superb at one and mediocre at another, and "the call sounds bad" needs splitting into which problem before choosing a fix. Echo that people hear as their own voice needs AEC. A hollow, bathroom-like ring on your voice is reverb — related but different, and we untangle the two in echo vs reverb.
Where AEC lives in your stack
On a typical setup there are up to four places echo cancellation can run, and knowing the order helps you debug:
- The calling app. Every major meeting platform ships AEC, on by default. Tuned for common setups; the layer that fails when your setup is uncommon.
- The browser. Web calls use WebRTC, whose echo cancellation is likewise on by default — one reason a browser join sometimes behaves differently from the same platform's desktop app.
- The OS and drivers. Varies wildly. macOS ships a Voice Isolation mic mode; Windows machines have driver-level "enhancements" of very mixed quality. We deliberately make no promises for this layer.
- A dedicated virtual-mic layer. Tools like Krisp insert themselves before any app, cancel echo once, system-wide, and hand every app a clean signal. When the built-in layers lose, this is the reinforcement — compared honestly here.
Layers generally coexist politely — a canceller that receives echo-free audio simply has nothing to do. The known bad interaction is doubled noise processing thinning a voice out, which is why the usual advice is: pick one layer to lead.
Quick answers
Is AEC why my music sounds terrible when I share it on a call?
Partly, yes — alongside noise suppression. The whole pipeline is built to preserve one nearby human voice and treat everything else as removable. Music is "everything else." Apps with a music or original-sound mode work by relaxing these layers.
Why does echo vanish when I put on headphones?
Because the loop needs a speaker a mic can hear. Headphones deliver the far side straight to your ears; your mic never hears them; there is nothing to cancel. It's the one fix that works by removing the problem instead of solving it.
Can AEC remove echo the OTHER side is causing?
No — cancellation needs the reference signal, and only the looping device has it. Your tools can clean your side; the other side's loop needs fixing there. Find the device, then fix that device.
Does "echo cancellation" in tool marketing always mean this?
Mostly, but read closely: some tools bundle de-reverberation under the same label (removing room ring from your outgoing voice), which is a different job than stopping speaker-to-mic re-transmission. Good vendors name them separately.