Intro to Vocal Cleanup
When cleaning up vocal audio, phrases like "de-noise" or "de-hum" can already spark the imagination enough to give you an idea of what's wrong, or what needs to be fixed, but there are plenty of words that could use some extra context. In this post, I thought I would go over some other terms which are important in the context of voice work cleanup, and have tried to bold them to help draw your attention.
The Human Element
For those of us that aren't robots, we generate sound and speech by using air from our lungs, blown through our larynx and over our vocal cords, and shaped by our mouth and nasal passages. We are all built a little dissimilar, and us-as-an-instrument all have innate differences.
Formants, for example, are the acoustic resonances of the human vocal tract, and are relatively unique depending on the person. Things like a longer/short vocal tract, the shape of your sinus cavities, and even bone/tissue density can all affect it. We shift the first and second formants to form words, the higher ones tend to rain more stable, all coming together to be a sort-of vocal fingerprint.
But as much as we are different, we are the same! We all use our unique voices to communicate with each other in a shared way. We produce vowel sounds by unrestricted airflow, and consonant sounds by restricting/stopping airflow with different parts of our mouth. Vowels are typically voiced, but consonants can be voiced or unvoiced.
If something is voiced, it means that the sound involves the usage of our vocal cords - something is resonating/vibrating. If something is un-voiced, it means the sound is made by air only. For example, do a quick octave of Do, Re, Mi, Fa, So, La, Ti, Do out loud. Those are all ending with voiced vowel sounds - now try to do the octave making only an hissing S consonant sound. Hard to do without those resonances, right? We can also intentionally de-voice something that's typically voiced. For example, try whispering Do, Re, Mi, Fa, So, La, Ti, Do.
Sibilance
A sibilant is a sound produced by pushing air through a narrow gap between the tongue and the roof of the mouth, creating a turbulent, high-frequency hiss. It can have a more meaning in other contexts, like literature, but in regards to audio, sibilance is likely to mean a harsh, high frequency sound relates to S and T sounds. Knowing this, it becomes less-surprising to hear that the audio tool to address these types of issues is called a de-esser. You can think of these tools like narrow band compressors, where they try and tame the excessive volume of these sounds where they typically occur in the 4 to 10 kHz frequency range. The goal is to lower their peaks more in-line with the rest of the vocals, but keeping it musical, and not completely drowning out the S sound. If you de-ess too aggressively, it can make the audio sound like the person has a lisp by making the S sound more like a Th.
Plosives
A plosive a sound produced by blocking the flow of air in the vocal tract and quickly releasing it, which makes a small burst of sound. These are sounds like b, d, g (voiced) and p, t, k (unvoiced). We're not going to get so far into the details, but these can be further categorized by where the blockage occurs in your mouth. For example, try enunciating the word bodega or potluck, and you'll hit all three.
These "bursts of sound" and air can overpower a microphone, and even cause clipping. Keeping some distance, and using a pop filter can help, but sometimes you don't have those handy, or they still make their way through. A filter or process labeled something like de-plosive is used to clean up these types of sounds.
Clicks and Breath
Mouth-clicks (non-linguistically) are just what the term sounds like - unwanted clicks from the mouth that make a distracting sound, like a lip smack or a wet tongue separating from the roof of the mouth. They can unexpectedly happen when your mouth is too dry or when saliva buildup causes tissues to stick together - both of which are bound to happen on long sessions of dialog, or many multiple takes of vocals for a song. A process or filter called de-click or mouth-de-click are geared to addressing these issues.
Loud inhales or exhales can sneak their way into records as well. Maybe a narrator didn't catch a breath before going into the next section, or a singer got a little too caught up in a take - it happens! Cleaning these types of audio artifacts are usually targeted in a process or filter called breath-control.
Both clicking and breathing can sometimes be less of an issue in a raw recording, but become more pronounce by compression and processing further down the pipeline. Sometimes it can pay off to clean up these types of noises a bit, even if they don't seem overpowering at first.
Wrapping Up
If a lot of these terms were new to you, hopefully this post helped inspire you to think about them more in your next song/voice-over/podcast, and keep in mind how you can up your game in your next creative piece of work.
Member discussion