How to get clean, professional voice audio for video
Viewers forgive a slightly soft image or an average color grade, but they click away from bad audio within seconds. The good news is that clear voice audio is mostly about recording well and applying a handful of standard processing steps in the right order. You don't need to be a sound engineer.
Start at the source
No plugin can fully rescue a bad recording, so a few minutes of preparation saves hours later.
- Get the microphone close. Distance is the biggest factor in audio quality. A cheap lavalier clipped to a shirt usually sounds better than an expensive microphone two meters away. For a desk microphone, aim for a hand's width from the mouth.
- Choose a soft room. Hard, empty rooms create echo. Rooms with curtains, carpets, sofas and bookshelves sound much better. Recording in a closet full of clothes is an old voice-over trick that works.
- Kill background noise. Turn off fans, air conditioners and fridges for the recording if you can. Constant hum is removable later; it's still better not to have it.
- Set levels with headroom. Aim for speech peaking around −12 dB on the recorder. Clipping (hitting 0 dB) is permanent distortion that cannot be repaired.
- Record 10 seconds of silence in the room before speaking. That "room tone" is useful for noise reduction and for filling gaps in the edit.
The processing chain
Apply these effects to the voice track in this order. Every major editor has built-in versions of each.
1. Noise reduction
Remove constant background noise like hum, hiss or air conditioning. Use a light setting: aggressive noise reduction makes voices sound robotic and underwater. AI-based voice isolation tools (in DaVinci Resolve, Premiere, CapCut, and standalone tools) have become very good at separating speech from noise, but still check the result for artifacts.
2. EQ (equalization)
- High-pass filter around 80–100 Hz: removes low rumble from traffic, handling noise and desk bumps. It's almost always the first EQ move.
- Cut a little around 200–400 Hz if the voice sounds muddy or boxy.
- A small boost around 3–5 kHz adds clarity and presence.
- Reduce harsh "s" sounds (sibilance, around 5–8 kHz) with a de-esser rather than EQ.
3. Compression
A compressor evens out the difference between loud and quiet words so the voice stays understandable. A good starting point for speech: ratio 3:1, a threshold where the gain reduction is around 3–6 dB on louder words, medium attack and release. Over-compression sounds flat and makes breaths too loud.
4. Loudness normalization
Finally, bring the whole mix to a standard loudness. Streaming platforms normalize audio, so mixes much louder than their target just get turned down. A common target for YouTube and social media is around −14 LUFS integrated, with true peaks below −1 dBTP. Most editors have a loudness meter and a normalize-to-LUFS option on export.
Music and sound effects
Music should support the voice, never compete with it. Keep background music roughly 15–25 dB below the voice while someone is speaking, and let it rise in pauses. Most editors can do this automatically with "ducking". If the music has vocals, keep it even lower, because two voices at once are tiring to listen to.
Editing tricks that improve audio
- Add very short fades (a few frames) at every audio cut to prevent clicks.
- Use your recorded room tone to fill gaps instead of leaving dead digital silence, which sounds unnatural.
- Reduce or remove loud breaths, but don't delete every breath — people sound unnatural without them.
Always check on different speakers
Listen to the final export on headphones, laptop speakers and a phone speaker. Phone speakers have almost no bass, so if the voice disappears there, it needs more presence in the 2–5 kHz range. Your audience mostly listens on phones and earbuds, so that's the test that matters.
Summary
Record close, in a soft and quiet room, with healthy levels. Then apply noise reduction, high-pass and corrective EQ, gentle compression, and loudness normalization to about −14 LUFS. Duck the music under speech. That simple chain gets most voice recordings to a clean, professional sound.