Fix Terrible Audio in Videos: Professor Bear's Guide to Using FFmpeg and 11labs
A volunteer's unusable audio gets extracted with FFmpeg, cleaned in Audacity, then completely rebuilt with an 11Labs voice, turning an unwatchable recording into a usable video.
A lot of good videos never get published because of one fixable problem: the audio is unusable. Volunteers at Humanitarians AI are graduates with real engineering credentials, but most of them aren't professional YouTubers or voiceover talent, and most don't own professional microphones. Professor Bear's fix isn't to reshoot anything. It's to strip the bad audio out entirely and rebuild it with an AI voice, using two free tools and one paid one.
The core idea
Rather than trying to salvage a recording made on a laptop mic in a noisy room, the process removes the original audio from the video entirely and remaps it onto a cleaner voice using 11Labs. If the speaker has a good-quality voice clone of their own, that's the ideal choice. Without one, the fallback is picking a similar voice, matching language, accent, and gender as closely as possible, so the final result still sounds like a plausible match for the person on screen. In the example worked through here, the original speaker didn't have a voice clone available, so the fix was selecting another voice with a similar accent and gender rather than trying to force an exact match.
Step one: pull the audio out with FFmpeg
The first tool is FFmpeg, a free command-line utility. On a Mac, it installs through Homebrew; on Windows, the recommended path is simply asking ChatGPT how to install it, since the process differs by system. Once installed, a single command line pulls the audio track out of the video file and saves it as an intermediate MP3, a working file that isn't the final product, just the raw material for the next step.
Step two: clean it up in Audacity
That raw audio gets opened in Audacity, a free tool praised here as genuinely serious software, on par with paid options that cost far more. The waveform itself tells you a lot before you even listen: a badly recorded track shows a visibly poor, cramped vocal range. Zooming into the waveform can also reveal specific defects, like a click or blip from a mic bump, that are easy to isolate, select, and silence, leaving a natural pause in place of an audible pop rather than cutting the audio outright.
From there, the fix for a weak vocal range is Audacity's amplify effect, reached through Effect, Volume and Compression. The tool will show the maximum boost available, sometimes close to a sevenfold increase, but pushing all the way to that ceiling causes distortion, since it drives the loudest peaks into clipping. Backing off slightly from the maximum gives a genuinely improved range without introducing new artifacts. The cleaned track is then exported as a higher-quality WAV file, replacing the rough intermediate MP3.
Step three: split the audio if it's too long
11Labs' voice changer tool has a practical limit of roughly five minutes of spoken audio per upload. If the cleaned track is close to that ceiling, it needs to be split into two pieces at a natural point where nobody is speaking, cut cleanly so the two halves line up exactly with no gap or overlap once they're rejoined later.
Step four: pick a voice and regenerate
Inside 11Labs, voices can be filtered by language, accent, gender, and other traits, English, an Indian accent, and a female voice in this case, since matching those attributes for the original speaker mattered more than any other factor. With a close voice selected, each audio chunk gets uploaded to the voice changer and regenerated in the new voice. The difference is dramatic: a original line so garbled it was barely intelligible in spots becomes clean and fully understandable, and importantly, all the natural imperfections of real human speech remain, since a real person is still speaking, just without the background noise and weak mic that made the original nearly unusable.
Step five: rebuild the video
Once both regenerated chunks are downloaded, using clear names rather than 11Labs' default unreadable filenames, they get brought into a video editor like Premiere Pro or Adobe Rush. The original audio track is separated from the video and removed entirely, and the new AI-generated audio is placed in its exact position, since the timing is a one-to-one match with the original recording. From there, the video exports normally, ready to be titled, described, and published.
Key takeaways
- The fix isn't repairing bad audio in place, it's extracting it, cleaning it, and rebuilding it as a new track in a matched AI voice.
- FFmpeg extracts audio from video for free on both Mac and Windows, using a single terminal command.
- Audacity can silence clicks and blips and boost a weak vocal range, but pushing amplification to its maximum introduces distortion.
- 11Labs' voice changer has roughly a five-minute limit per upload, so longer recordings need to be split into matched chunks first.
- Choosing a voice with a similar language, accent, and gender to the original speaker keeps the final video feeling like a natural match, even without an exact voice clone.
Who this is for
This workflow is built specifically for Humanitarians AI fellows who have access to 11Labs but not to professional recording equipment, giving them a practical way to publish videos that would otherwise be unusable because of audio quality alone. It's a direct, repeatable process anyone in a similar position can follow end to end, from a rough recording to a finished, watchable video.
Full transcript(auto-generated, with timestamps)
[0:03]UI builder designed to help developers and designers create responsive and production ready UIs effortlessly. It integrates NLP. >> Okay, Professor Bearer here. Uh we have a lot of volunteers at humanitarians. They're not professional YouTube people or voiceover people, but they are graduates who know something about engineering. They got master's degree in engineering. when they make videos, it's fairly common for the audio to be horrible. Most of them do not have professional mics. Uh, you know, it's unfortunate because this is a nice little video which is unusual, unusable because of the audio. So, what I'm going to show you now is how to correct that. how to take absolutely terrible unusable
[0:55]Audio, remove the audio from the video, and then use a tool called 11 Labs to remap it to a better audio. This can be done with your own voice if you make a voice clone, assuming you have good quality recordings of your voice, or you can just pick a similar voice. In this case, this is an Indian woman who's making this. I'm just going to pick another Indian woman's voice. uh because we don't have this person, this volunteer, this fellow's particular voice as a voice clone. So what do we do? We we take the um the video and we use a free tool here. It's called afmpeg.
[1:43]And what we're doing here is we are basically just using command line. Use brew to install this on your Mac. Use chachiver to ask it how to install ffmpeg which is a free tool on a Windows machine. But this is the name of the video. And I'm going to output the the audio to uh just a file called audio mp3 because that right now is in an intermediate file. which we will not uh you know use directly. We're going to make a better version of that. So you know you just install it and hit return. So there we go. So I'm now going to open
[2:30]Up the audio in a free tool called Audacity which is an awesome tool by the way. I love it when the free tool is also the best tool. Uh, we'll make some more stuff on Audacity, but Audacity is great and it's free. It's a serious serious tool as any audio tool out there other than maybe ProTools, which cost a gazillion dollars and it's awesome. So, whoever made uh Audacity, thank you. Uh, because it's free and as good as it exists for for audio stuff. So, we look at the audio. We can see how poorly uh poorly this was done just by the vocal range of the
[3:12]Audio. We can also see there's likely a blip here. This is like a click or a blit. So, I'm going to zoom in there and listen to that. I'm likely to delete this because if I increase the range, this will limit my max range by the upper range of that. So, let's listen to that and see what is that because this is very unusual. So, I'm talking talking boom something. >> Instead of data, let's put software. >> So, that was a click. So, what I'm going to do is I'm going to highlight that. I'm going to silence it. And silence audio. There we go. Now, if we listen to it,
[3:55]>> exactly what all I need instead of >> have a little pause there rather than a click. If we wanted to be more sophisticated, we could cut it, but we're not going to do that. The next thing is we see the vocal range here is horrible. So, we want to increase it. And we can see that immediately from the just the wave file here. So, what we do is we go to effect, we go to volume compression, we go to amplify. This says we can bring it up nearly by a factor of seven. We don't want to do that. We don't want to bring it all the
[4:25]Way to the top. If you bring it all the way to the top, you'll see it it go all the way up and down, and that'll have the tendency of sort of vibrating your mic when it hits the max. So, this is 6.9. So, I'm going to take it to say 6.7. So, we can see we have a better range now. And I'm just going to save that. So, I'm going to export the audio. Now, I'm just going to overwrite the the original u the file is an MP3 file. So I'm creating a higher version wave of that file. Now we need to look at the size of this
[5:02]Thing. So 11 Labs, which we're about to use, can do about five minutes of video max. This is close. My guess I'm going to try this, but my guess this is too much audio for um just a little bit like by 30 seconds. too much audio for 11 Labs to handle. But what's let's just go ahead and try it. So, we are going to go to 11 Labs here, wherever my 11 Labs is. And we're going to do what's called voice changer. And we're going to upload. My bet here is it's going to say this is too much, but will you well see it's too large. Uh, and it's probably
[5:49]Too large just by a little bit. So, I'm just going to cut it in half. For this one, it doesn't really matter. I'll pick a region where there's not really any speaking. And I'll trim the audio. And I will create uh two versions of it. I'll create the A version. So, this is the first half. And because I want it exact, I'm just going to reverse this. And now I'm just going to cut this. The reason why I do that is I want to cut it exactly. I want to just piece one, piece two. So I'm just going to then do a cut. And now here's my second piece. I'm just
[6:38]Going with B. So now we have a couple chunks that 11 Labs should be able to handle. Again, my experience is around 5 minutes of audio, just sort of spoken audio. It may be different for music or something. Is this now we need to pick a voice. So let me go to my um rather than history, let's go to settings. So this is a So normally I would do this. I would because this is a female voice. I would go to voices. I would search. So I'm going to go to filter. I want this to be English. English is one of those hard words to
[7:27]Spell. Uh I want an accent. I'm going to pick a Indian accent. I'm going to pick a female because the female speaking and I'm not going to worry about the age. and I'll apply filters and I'll find a bunch of voices. Um, I just try to pick something that's sort of close to the original speaker because we don't really have the original speaker's voice clone for this one. For example, this is the one I'll likely go with because I know it and I like it. >> In the land of sprinklewood, every cupcake grew wings. >> I find he kind of funny. So, that's a kind of funny nice little voice. So,
[8:08]We're going to go to the voice changer again. So, this is the voice. We upload our audio and we generate the speech. This is about half of 2 minutes and 42 seconds. It will generate. It'll take a couple minutes to do this. I will show you how to make a portfolio website using VO the AI powered UI builder. So what is VI builder designed to help developers and designers create responsive and production ready UIs? >> Let's go back to the original. >> I will show you how to make a portfolio website using VO the AI powered. >> That's a big difference. And what's nice about sort of going with
[8:57]This method is because it's a human being speaking, all the imperfections of her speaking aren't kept in the audio. But you can hear the background noise here. You can hear the poor mic. You can just This is just terrible audio. Beware. >> Well done. So what is view? It is an AIdriven UI builder designed to help. The noise in her system is almost as strong as her voice. And that's a problem in a video because people want to hear what you got to say. So, let's listen to this again. >> I will show you how to make a portfolio website using VO, the AI powered.
[9:40]>> That's just so much clearer. So, I haven't download it with the original name. Otherwise, 11 Labs will give it a unreadable number for a name. And now we need to do the second half. And so I just named it A and B just to keep it super simple. And we did the same thing. >> You could see it has given us P.TX TSX where uh it has called all the pages. This is the about me section where uh we have given about me uh we need to change this according to us. Just put whatever details about you. This is a contact page. So just change
[10:30]This and give >> So it does this. It's actually human. See, right now what it's doing is it's allowing me to listen to it before it's actually finished rendering. Uh I'll I'll do this as quick as possible because I'm impatient. Uh but you might need to hit this a couple times because it actually won't it sort of pretends it's done when it's not really done. Once you can download it, it's actually done. And so then what we do is we bring it into any recording, Premiere Pro, Adobe Rush, whatever you want. Um, I'm This is from a another thing here. I'm just going to delete
[11:17]That. So, we just bring in our media. So, uh, I'm going to also just delete that from the timeline. So we bring in our video which is here. We put our video on the timeline and we separate the audio in the video. How you do this will depend on your tool but we need to remove the old v the old audio from it. Here again is a replaying of what it sounds like. I will show you how to beat a portfolio website using VO, the AI powered UI. Well done. >> So, we don't want that. It's just a, you know, a terrible quality recording. It's
[12:00]Not the fault of the volunteer. It's just they're not an expert in um, you know, recording. They're not professional recorders. Uh, and typically a recent graduate, somebody who's just graduated from college, does not have expensive recording equipment. So uh so for those reasons, you know, it's a typically a low quality recording. So we just because this is exact match, we're literally just going a one to one correspondence with the original um recording and then the 11 Labs enhanced recording. Uh we just lay it there exactly. and we play >> I will show you how to make a portfolio website using VO the AI powered UI
[12:52]Builder. So what is VO? It is an AIdriven UI builder designed to help developers and designers create response. >> So that's it. Now we just export our video. Give it some name. Everyone a portfolio. I'm just going to call P and I'll put it in uh you know create a little um YouTube description and put it on YouTube. So that's it for this video. It's basically this is for fellows who have sort of terrible recordings but have access to our 11 labs. This is the simplest and easiest and often best way of taking sort of terrible audio into decent audio. So that's it. Uh I guess I go through the
[13:48]Like, subscribe, whatever thing for the humanitarians.ai YouTube, but this is primarily for fellows making videos. How to improve the audio quality without expensive recording equipment. Take care.
More videos
2:08Bridging the Pixel Gap in Browser Automation.
2:23How One Narrow Safety Rule Can Make an AI Less Safe Everywhere Else.
2:04Why splitting a chunk from its document makes it retrieve for the wrong question
4:20Three You Can Take Back. One You Can't.
2:21Why a 50-turn agent pays for the same screenshot 35 times unless it caches the pixels
1:53