Good audio mixing is less about technical perfection and more about making every word clear, every sound purposeful, and the overall loudness consistent from beginning to end.
Most viewers will forgive shaky footage before they will forgive bad audio. A mix that is too loud, too quiet, muddy, or unbalanced pulls people out of your video the moment it starts. Audio mixing is the process of adjusting every sound in your edit — dialogue, music, and effects — so they work together cleanly instead of fighting each other. You do not need a professional studio to get a solid mix. You need to understand a few core concepts and apply them methodically every time you finish an edit.
Every audio track in your editing software has a level — a measurement of how loud it is at any given moment. Levels are shown in decibels, abbreviated dB. The key reference point to understand is 0 dBFS, which stands for zero decibels full scale. This is not silence; it is the absolute ceiling. Any audio that pushes past 0 dBFS clips, which means it distorts. Clipping sounds like a harsh crackle or crunch, and it is nearly impossible to fix after the fact.
Loudness is a related but different idea. Where level describes a single moment in time, loudness describes how loud something feels across a longer stretch of audio — how your ears perceive the average intensity of a piece. A track can have brief loud peaks but still feel relatively quiet overall, or it can sit at a consistent moderate level and feel quite present. Good mixing manages both: keeping peaks safely below the ceiling while maintaining a consistent, comfortable loudness throughout.
Think of it like driving. Your speedometer shows your speed at this exact second (that is level). Your average speed for the whole trip (that is loudness). Both matter for a safe journey.
Dialogue is almost always the most important element in a video. If the viewer cannot hear and understand what is being said, nothing else in the mix matters. Always set your dialogue level before touching anything else.
A practical starting point: bring your dialogue up until it sits comfortably in the mid-range of your meters — roughly in the area of -12 dB to -6 dB on your level display during speech. This leaves space above it so that peaks do not clip, and it gives you room below for music and effects to sit under the voice without drowning it out. This space above your working level is called headroom, and protecting it is one of the most important habits you can build.
Listen to your dialogue at a normal, comfortable volume. If you have to strain to hear the words, it is too low. If it feels aggressive or tiring, it is too hot. Your ear is a reliable guide once you are not judging in a silent room at maximum volume — mix at the volume your audience will actually use.
Once your dialogue is sitting where you want it, bring in your music and sound effects underneath it. The goal is a hierarchy: dialogue on top, effects filling in the middle layer, music supporting from below. Every element should be audible, but none should compete for the same space as the spoken word.
Music under dialogue typically needs to sit noticeably lower than you might expect. In isolation, the music track may sound fine at a moderate level. The moment you layer it under speech, it can suddenly feel intrusive. Bring it down until it feels like texture beneath the voice rather than a separate competing track. If you find yourself having to choose between hearing the music or hearing the words, the music is too loud.
Sound effects occupy the middle ground. Ambient sounds and room tone should sit low and blend naturally into the background. Specific effects that are meant to be noticed — a door slamming, a notification sound, a key hit — can sit closer to dialogue level for the brief moment they occur, then recede. Let the purpose of each sound guide how present it needs to be.
EQ, short for equalization, is a tool for adjusting the tone of an audio track by boosting or cutting specific frequency ranges. Think of it like a multi-band tone control. The low end covers bass and rumble. The midrange is where most of a voice lives. The high end contains clarity, air, and sibilance — the sharpness in S sounds.
On dialogue, EQ is most commonly used to remove problems rather than add qualities. A low-cut filter — sometimes called a high-pass filter — rolls off the very bottom of the frequency range and eliminates hum, air conditioning rumble, and handling noise that muddy up a voice. Most dialogue benefits from a gentle low-cut. Beyond that, if a voice sounds boxy or nasal, a small reduction in the midrange often opens it up. If it sounds harsh, a reduction in the upper midrange or high end softens it.
Compression reduces the dynamic range of a track — the gap between the loudest and quietest moments. When a speaker's voice goes from a soft sentence to a sudden loud word, compression automatically turns down the loud word, keeping the overall level more consistent. The result is dialogue that is easier to follow without constantly riding the volume fader. Most editors apply light compression to dialogue as a matter of course, not as a creative choice but as basic housekeeping. Keep the settings gentle: heavy compression makes voices sound pumped, unnatural, and fatiguing.
Ducking is the technique of automatically or manually lowering the music and effects whenever dialogue is present, then bringing them back up during pauses or music-only passages. It is one of the most practical mixing moves in video production and the reason why professionally mixed videos feel clear and professional even when music is playing throughout.
The simplest form of ducking is manual: you draw keyframes on your music track's volume envelope, reducing the level when a speaker starts talking and returning it to a higher level when they stop. Most editing software makes this straightforward with automation or keyframe tools. Set the music lower during speech, bring it up in gaps, and crossfade the transitions smoothly so the changes do not sound abrupt.
Some editors use a sidechain compressor or an auto-ducking plugin that detects the dialogue track and reduces the music automatically in response. This is faster on longer projects but requires some adjustment to get the timing and depth of the duck sounding natural rather than mechanical. Either approach works — the goal is simply that your music steps back when someone is talking.
Once your mix sounds balanced through headphones and speakers, the final step is making sure the overall loudness is consistent and appropriate for where the video will be watched. Streaming platforms, broadcast, and social media all have their own loudness expectations, and they will normalize your audio if it arrives too loud or too quiet — often in ways that shift your carefully balanced mix.
The most practical approach is to use a loudness meter in your editing or export process, if your software supports it, and aim for a level that feels broadcast-like rather than maximized. Louder is not better in a delivery context. A mix that has been artificially pushed to the loudest possible level will actually sound quieter than a well-balanced one on many platforms, because the platform's normalization will pull it down more aggressively.
Before exporting, listen to your final mix from beginning to end at normal volume. Check for moments where the music suddenly spikes, where dialogue gets buried, where a sound effect startles unexpectedly. These are the edits to make. A full-pass listen at the volume your audience will use catches problems that metering alone will not reveal.
A repeatable workflow removes guesswork and helps you catch problems before they reach a client or audience. The order matters: setting dialogue first gives you a reference point for everything else.
Start by cleaning your dialogue — apply a low-cut filter, remove obvious noise, and level out inconsistent performances with light compression. Then set the overall dialogue level for the sequence. Bring in sound effects next, placing them in the hierarchy below dialogue. Finally, add your music and duck it under speech. Once the balance feels right, run a loudness check and make any final adjustments before export.
Keep this workflow consistent and your mixes will improve quickly. Most audio problems in beginner edits come not from lack of equipment or skill but from skipping steps — adding music before dialogue is set, or never ducking at all. The process is straightforward; the discipline of following it every time is what makes the difference.
There is no single correct number, but a practical rule is to have your dialogue peaks sitting comfortably between roughly -12 dB and -6 dB on your level meter during normal speech. This keeps you well clear of clipping at the top and leaves room for music and effects to sit below the voice without being inaudible. Trust your ears at normal listening volume as much as your meters.
This is one of the most common mixing problems in video editing. Bring the music down further than feels comfortable in isolation — often significantly further. Music under dialogue typically needs to sit 15 to 20 dB below your dialogue level to stay out of the way. Use ducking to pull it down even more during speech and allow it to rise slightly in gaps. If the music still feels competitive, check whether it occupies a similar frequency range as the voice; if so, a gentle midrange cut on the music track can help them coexist.
Headroom is the gap between your average working level and the absolute ceiling of 0 dBFS. If your dialogue sits at around -10 dB on average, you have roughly 10 dB of headroom before a sudden loud moment would clip. That buffer is essential because audio has transient peaks — brief spikes that are much louder than the average — and without headroom those peaks distort. Protecting headroom keeps your audio clean through the entire production and delivery process.
Not necessarily. If your dialogue was recorded cleanly in a treated space, with good microphone placement and consistent performance, it may need very little processing. A light low-cut filter is almost always useful on spoken audio regardless of recording quality, but heavy EQ and compression are tools you reach for when something needs fixing or evening out. The goal is a natural-sounding result, not a processed one.
Mixing solely on headphones or solely on one set of speakers is a common trap. Headphones can exaggerate low end and stereo width in ways that do not translate to speakers, and a single pair of laptop speakers may emphasize midrange harshness. The solution is to check your mix on multiple playback systems before delivery: headphones, laptop speakers, a phone, and ideally a decent pair of monitor speakers if you have access. A mix that holds up across different listening environments is a solid mix.
Studio432 edits concert footage, music videos, gaming content, vlogs and short-form for creators worldwide. Send the footage, get a quote by email.
Start a project →