How to Voice Over Videos: Script, Record, Edit, Mix
How I record voice overs for video: writing for the ear, treating the room, setting clean levels, and mixing narration under the picture so it stays clear.
To voice over a video well, I write the script for the ear first, record clean audio in a soft, quiet room with a decent microphone and a pop filter, then edit out the noise, level the read with a compressor and a little EQ, and mix it under the picture so the words stay clear. That sequence is the whole job. Rush the script and the read fights you the entire session; ignore the room and no plugin later will save the take.
I have recorded narration for product demos, brand films, and long training modules, and the same four or five habits carry every time. Here is how I work through them.
Why the voice carries more than people expect
A viewer trusts your narrator before they trust anything on screen. A steady, warm read in a software demo makes the product feel finished. A rushed, muffled read makes even a beautiful edit feel amateur, and people click away in the first few seconds.
There is a practical reason to sweat the audio, and it runs the other way too. Most social video now plays on mute, and HubSpot’s 2025 data found roughly 85 percent of younger viewers watch on silent by default. So I treat the voice over and the captions as one deliverable: the read has to work for anyone who turns the sound on, and the captions have to carry the meaning for everyone who does not. If you are weighing whether the extra polish pays off, our note on video marketing ROI has the arguments I use with clients.
Write for the ear, not the page
A sentence that reads fine on paper often trips the tongue out loud. My first rule of scripting is to read every line aloud as I write it. If I stumble, the narrator will too, so I cut the clause or split the sentence right then.
I also bake performance notes straight into the script. A [pause] before a key point, a [slower] on a number people need to catch, a tone cue like [warm] at the top of a section. These take seconds to add and save real time in the booth. Our full walkthrough on how to write a video script goes deeper on structure, but the ear-first habit is the part that changes your recording day.
The scripting mistake I see most often is cramming. People speak at about 150 words a minute in a normal register, a rate the National Center for Voice and Speech has cited for years, so a 60-second read is around 150 words, not 220. Slower, detail-heavy narration runs lower than that. When I write to length, I count words before I ever hit record.
Based on an average conversational speaking rate of about 150 words per minute; slower, detail-heavy narration runs lower.
Build a room that does not fight the mic
Your biggest enemy is reverb, the hollow echo you get when your voice bounces off hard walls and windows. A $500 microphone in a bare room loses to a $60 one in a soft one, so I fix the space before I think about gear.
Small rooms beat large ones because there is less air for sound to ricochet through. A closet full of clothes is the classic recording booth for a reason. If I cannot get a closet, I drape thick blankets on the wall I face, close the curtains, and stand near soft furniture rather than in the middle of an empty room.
For gear, a USB condenser microphone covers most of what a marketing team needs, and it plugs straight in. I keep my mouth 6 to 12 inches back and slightly off-axis so my breath does not slam the diaphragm, and I always use a pop filter to soften the hard “p” and “b” sounds. Closed-back headphones let me hear plosives and background hum as they happen instead of discovering them in the edit. A lot of the time you are recording narration over a screen capture, so it is worth pairing this setup with one of the free screen recording tools that handle audio cleanly.
Get the take right so the edit stays short
Before I record a real word, I run a test read of my loudest lines and watch the meter. I aim for peaks between -12 dB and -6 dB. If the meter hits 0 dB the audio clips, and that crackly distortion cannot be fixed afterward.
Consistency does the rest. I keep the same distance and roughly the same energy across a session, because a read that wanders in level is a nightmare to smooth later. When I flub a line, I do not stop the recording; I pause, clap once to leave a sharp spike in the waveform as a visual marker, and repeat the line. Those claps point me straight to the fixes.
I save every take as an uncompressed WAV at 48kHz, 24-bit. Recording straight to MP3 throws away audio data before you have even started editing, and you cannot add that quality back.
Edit like an engineer, with a light hand
Editing is where a clean take becomes a finished one, and you can do all of it in a free tool like Audacity. I work in three moves: clean, level, shape.
First I clean. I snip the flubs, remove mouth clicks, and run noise reduction by sampling a few seconds of silent room tone so the software knows what hum to strip. A light touch matters here, because heavy noise reduction makes a voice sound thin and watery. Then I level with a compressor, which nudges the quiet parts up and the loud parts down so the listener is not reaching for the volume. Last I shape with EQ, a small lift in the low mids for warmth and a gentle high-end boost for clarity.
Adobe’s own video team recorded a walkthrough of tracking and cleaning up a voice over inside their editor, and it is worth watching if you want to see these moves happen in real time rather than read them.
If you would rather see the whole post-production picture, our guide to video editing covers where the audio pass fits alongside the picture cut.
Mix the voice under the picture
Now the voice meets everything else. I drop the edited WAV onto its own audio track, separate from music and effects, so I control each element on its own.
Then I sync. If the script says “with one click,” I nudge the clip until “click” lands on the cursor click on screen. Once the timing is right, I set the balance with one priority in mind: the narration reads clearly at all times. Music sits underneath, usually around -18 dB to -24 dB below the voice, quiet enough to add feeling without covering a word. Rather than ride that fader by hand, I use audio ducking, which automatically dips the music whenever the voice is talking and lifts it in the gaps.
The last step before export is loudness. Platforms normalize volume to a target measured in LUFS, and mastering louder than the target buys you nothing because the platform just turns it back down. YouTube normalizes to roughly -14 LUFS integrated with a true peak near -1 dBTP, a reference documented in the standard loudness normalization guidance, while feed-first apps like Instagram tend to sit a little hotter. I check my mix against a loudness meter and hit the target for wherever the video will live. If you are handing this stage to someone else, our piece on how to hire a video editor covers what a good audio-literate editor should already know.
When a synthetic voice is fine, and when it is not
Synthetic voices have improved, and I do reach for them. For scratch reads while I lock timing, or for a quick internal draft, a generated voice saves a booth session. For a customer-facing hero video, I still book a real read, because the small human choices, a breath before a hard word or a lift on the payoff, are the part a listener feels without knowing why. If you want to compare what is available, our rundown of voice generators ranked walks through the current tools. On the work my team ships at Moonb, the audio pass gets its own review before anything leaves the door, and that habit is what keeps a whole series sounding like one voice.
None of this needs a treated studio or a rack of hardware. A soft room, a modest microphone, a script written for the ear, and a patient edit will get you a voice over that sounds like you meant it.
Frequently asked questions
Around 150 words. Most people read narration at roughly 150 words a minute in a normal register, so a full minute lands near 150 words, not 200-plus. If the script is dense with numbers or technical terms, aim lower, closer to 120 to 130 words, so the read has room to breathe and the viewer can follow.
A USB condenser microphone with a pop filter, used in a soft room, will beat a phone every time, and it stays inexpensive to set up. That said, the room matters more than the mic. A modest microphone in a closet full of clothes sounds better than a high-end one in a bare, echoey office, so treat the space first.
For YouTube, aim for about -14 LUFS integrated with a true peak near -1 dBTP, since the platform normalizes anything louder back down anyway. Feed-first apps like Instagram and TikTok tend to prefer slightly hotter mixes. Use a loudness meter in your editor and match the target for wherever the video will actually play.