Architecture, not tutorials
The home VO software chain: measure, treat via settings, clean via software
A home voice-over chain that works is three moves in a fixed order: measure the raw floor, remove what settings and scheduling can remove, then let software subtract only the steady remainder — a high-pass filter first, gentle broadband reduction second, and nothing gated to silence, ever. This page is the architecture: what goes where, how much is safe, and which tool class belongs on which job.
Why architecture beats presets
Every noisy recording is a different mixture, so no downloaded preset can know yours. But the structure of a good chain is nearly universal, because it follows one physical rule: remove the easiest, most separable noise first, so each later stage works on cleaner material. Rumble is trivially separable from voice (it lives below where voice lives), so filtering it comes first. Steady hiss is statistically separable, so broadband reduction comes second. What's left — noise woven through the voice itself — is barely separable at all, which is why the honest last stage of every chain is “stop, you've done enough.”
Before any of it: measure. A chain built against a measured floor is engineering; a chain built by ear at midnight is mood. You need the raw number to know how much work software must do — and 6 dB of work sounds clean from almost any tool, while 25 dB sounds robotic from every tool.
- Stage one: high-passAround 70–80 Hz — rumble your voice doesn’t use
- Stage two: gentle reductionBroadband, 6–12 dB — the steady remainder only
- No stage three: the gateNothing gated to silence, ever
Stage one: the high-pass filter
Set a high-pass (low-cut) filter around 70–80 Hz for most voices — a touch higher for higher voices, a touch lower for genuinely deep ones. Everything below that line is traffic, HVAC thrum, footfalls, desk bumps traveling up the stand — energy your listener never needed. It's the highest-value, lowest-risk processing in all of VO: the voice loses nothing, the floor drops immediately, and every meter downstream (including the RMS math in ACX's spec) stops counting energy that was never performance. If your interface or DAW offers a slope choice, a moderate 12–18 dB/octave slope is plenty.
Stage two: gentle broadband reduction
Broadband noise reduction learns the steady texture of your floor and subtracts it everywhere. Two rules keep it invisible. First, work in single-digit-to-low-teens dB: 6–12 dB of reduction is the polite range where voices stay human; every tool's slider goes far beyond it, and far beyond is where the underwater artifacts live. Second, feed it honest material: reduction tools work from your noise's fingerprint, so the same discipline that makes measurement valid — real gain, real room state — makes reduction clean. If 12 dB of reduction still doesn't reach spec, the answer is upstream (find the source), not a bigger number on the slider.
We deliberately stop short of per-editor walkthroughs — the click-this-menu genre is well served elsewhere, and the settings above translate to any tool you own. Our lane is making sure the number you're chasing and the order you chase it in are right.
The stage that isn't there: the gate
Noise gates mute everything below a threshold — and on narration masters they're a trap. A gate doesn't remove noise; it removes noise only when you aren't talking, which means your floor audibly breathes: silence between phrases, hiss under them. Worse, a gate slammed down produces stretches of pure digital silence, and audiobook QA — human and automated — reads dead silence as an edit artifact. ACX's own requirement of 1–5 seconds of room tone (not silence) at file boundaries tells you what the format expects: a continuous, natural, quiet floor. If your untreated pauses are loud enough to need a gate, that loudness is a diagnosis, and the diagnosis page is where to take it.
The live/offline split
The most consequential chain decision isn't a setting — it's routing the right tool class to the right half of the job.
| Tool class | How it works | Belongs on | Keep away from |
|---|---|---|---|
| High-pass filter | Removes energy below a set frequency | Everything — masters, auditions, calls | Nothing; it's the free lunch |
| Offline broadband reduction | Learns steady noise, subtracts it file-wide, fully undoable | Recorded masters, auditions, deliverables | Live calls (it isn't real-time) |
| Real-time AI suppression (Krisp-class, RTX-class) | Isolates voice from everything else, milliseconds at a time | Directed sessions, agent calls, coaching — the live half | Masters: optimized for call intelligibility, not natural room tone |
| Noise gate | Mutes below a threshold | Almost nothing in narration | Anything ACX-bound |
The split exists because the two halves optimize for different listeners. A call needs maximum intelligibility right now, and aggressive suppression is the correct trade. A master needs to survive ten relaxed hours in a listener's ears, where naturalness beats aggression every time. Run both — suppression guarding your sessions, a gentle offline chain polishing your files — and never let them swap jobs.
Two chains to steal
The audiobook master chain: high-pass at 70–80 Hz → broadband reduction, 6–12 dB → RMS normalization into ACX's −23 to −18 dB window → limiter at −3.2 dB → re-measure the floor on the finished file. Five stages, every one boring, which is the compliment a mastering chain should aspire to.
The one-hour audition chain: high-pass → reduction only if your measured floor demands it → level to sit comfortably, peaks around −3 to −6 dB → export. Speed matters on auditions and over-processing costs bookings, so this chain is deliberately shorter — the reasoning is in the audition audio guide.
The audiobook master chain — five stages
- High-pass, 70–80 Hz
- Reduction, 6–12 dB
- RMS into −23 to −18 dB
- Limiter at −3.2 dB
- Re-measure the floor
The one-hour audition chain — deliberately shorter
- High-pass
- Reduction, only if the floor demands it
- Peaks around −3 to −6 dB
- Export
Both assume the upstream work is done: sources silenced, gain staged, technique tight. Software is the last 10 dB, never the first 30 — that proportion is the whole of the fix order, and chains that respect it stay invisible.
Chain questions
- Where does EQ beyond the high-pass belong?
- After noise work, and sparingly. Tonal EQ is voicing, not noise control — mixing the two makes both harder to judge. Get the floor right first; then decide whether the voice needs shaping at all.
- Real-time suppression would simplify everything. Why not record through it?
- Because you can't un-process a take. Record raw and you own every option forever; record through a suppressor and you own its artifacts forever. The one exception is live-to-client delivery where there is no “later” — which is precisely the call scenario, not the recording one.
- How do I know if my chain is over-processing?
- Listen to your pauses at high volume: if the floor pumps, flutters, or drops to nothing between phrases, you've crossed the line. And check the reverse tell — a measured floor near digital silence on a home recording usually means the chain erased the room instead of quieting it. Natural-but-quiet wins; QA ears agree.