Off-Hook · entry 001
Two noise filters, one call: when suppression stacks fight
By the LineSide desk ·
There's a failure mode that produces more "VoIP audio is just bad" resignation than any network problem: a voice that turns watery, clips the ends off words, and pumps between sentences. It isn't the codec and it isn't your connection. It's two noise filters fighting over the same audio — and the strongest evidence that stacking is a real fault, not an audiophile superstition, comes from the vendors themselves.
The vendors already told you
Collect the warnings scattered across softphone documentation and a pattern appears (all read 2026-09-20):
- Aircall says it in writing. Its help center notes that many headsets and devices apply their own noise processing, which "could interfere" with the app's noise cancellation — and recommends not enabling the additional layer in such cases. That's a vendor telling you its own feature performs worse behind another filter. (Our Aircall guide has the full fine print.)
- Dialpad engineered around the adjacent problem. AI Voice Isolation automatically switches itself off under low system performance — an explicit acknowledgment that real-time filtering competes for resources and that a degraded filter is worse than none. Two filters double that bill. (Details here.)
- Every virtual-device setup guide ends the same way. The standard instruction when putting a device-level filter like Krisp under a softphone is to disable the app's own suppression — it's the closing step in our setup order for a reason.
Why stacked suppression sounds worse than either filter alone
Plain-language version, no signal-processing degree required. A noise filter works by continuously estimating which parts of the incoming audio are "voice" and which are "noise," then subtracting its noise estimate. The estimate is never perfect — every filter leaves faint artifacts: slightly hollowed consonants, a whisper of residue in pauses, tiny spectral holes where noise used to be.
Feed that output into a second filter and you've handed it a signal unlike anything its models were trained on. Natural rooms don't contain spectral holes and synthetic residue. The second filter misclassifies: it treats the first filter's artifacts as noise and carves deeper, treats damaged consonants as not-voice and deletes them entirely. The audible results are exactly the complaints in the first paragraph:
- Word endings vanish — final consonants were already the weakest part of the signal after pass one; pass two finishes them off.
- The "underwater" tail — both filters' residue-carving stacks into an unnatural, warbling emptiness behind the voice.
- Pumping — two gain-and-gate decisions oscillating against each other between speech and silence.
And that's before the resource cost: two real-time inference processes on one laptop, which is precisely the condition that makes adaptive filters (and softphones generally) degrade.
Count your layers — there are more than you think
The insidious part is that nobody decides to run two filters. They accumulate. A quick audit of places suppression can be active on one ordinary call:
- The OS — platform "enhancement" processing on the mic.
- The browser — WebRTC apps like the 3CX Web Client get default echo cancellation and mild suppression from the browser itself.
- The softphone — Dialpad's default-on isolation, RingCentral's toggle, Aircall's desktop setting, Zoiper's default-on classic DSP.
- A virtual-device filter — Krisp or similar, if installed.
- The headset's own firmware — many comms headsets process the mic before the OS ever sees it (which is why Aircall's warning names headsets specifically).
The fault: two layers on
- Headset firmware
- OS enhancement
- Browser DSP
- Softphone suppression
- Virtual-device filter
The fix: one owner
- Headset firmware
- OS enhancement
- Browser DSP
- Softphone suppression
- Virtual-device filter
Five candidate layers. It only takes two being active to produce the fault — and a default-on app filter plus a processing headset gets there with zero user action.
The one-owner rule, and how to pick the owner
The fix costs nothing: exactly one layer owns noise removal per direction (mic-out and speaker-in are separate jobs and can have different owners). Choosing the owner:
- Quiet room, steady noise: let the app's built-in own it if yours has one — free, zero setup. Two of our six tracked apps ship nothing, so check the matrix.
- Moving noise — voices, keyboards, dogs: AI-class filtering owns it; classic DSP layers (Zoiper, browser defaults) go off where you can switch them.
- Multiple apps all day: a device-level owner beats per-app toggles — one filter, one configuration, every softphone downstream. Switch each app's own suppression off as you adopt it.
- Processing headset: pick it or the software layer, not both — and judge by a recorded test in your actual room, not by spec sheets. Ours is the rare corner of audio where "turn the feature off" is often the upgrade.
The five-minute audit, concretely
How to actually run the count on your own machine, once: open your headset vendor's companion software (if one is installed) and note whether any noise or voice processing is enabled. Check the OS input settings for enhancement processing. Open your softphone's audio or hardware settings and note the state of its suppression toggle — the per-app locations are in the matrix. If a virtual-device filter is installed, note which directions it's filtering. Now count the active layers per direction: more than one, switch the extras off, starting with the weakest (classic DSP first, per the ownership rules above). Finish with a thirty-second recording while typing — the before/after is usually audible enough to end the argument with whoever configured the stack in the first place.
This entry is the long-form version of a warning that appears on nearly every page of this site. If you only remember one sentence from LineSide, make it this one: one filter per direction, chosen on purpose.