How to Remove Crowd Noise from a Video or Recording

Everything worth knowing before you clean up a noisy recording: how to get the audio out of a video, what the AI removes and what it protects, which situations it handles best, and why the noise-reduction slider in your editor cannot do this.

Removing crowd noise from a video, step by step

CrowdRemoval works on the audio track, so if your recording is a video you extract the sound first, clean it here, then put it back. It takes about two minutes and every tool you need is free.

  1. Extract the audio. In VLC on a computer: Media → Convert / Save, add your video, and convert it to Audio - MP3. On a phone, CapCut and iMovie both let you detach the audio from a clip. Any free online video-to-MP3 converter works too.
  2. Clean it here. Upload the extracted MP3 or WAV and let the AI strip out the applause and crowd noise. Your first file is free and needs no account.
  3. Put it back. In your video editor, mute or delete the original clip's audio, drop the cleaned track underneath, and line the start up with the video.

One tip that saves a lot of frustration: extract the audio before you trim the video, not after. If the cleaned track is the same length as the original, it drops straight back in and stays in sync with no nudging.

What it takes out — and what it leaves alone

The model is trained to recognise the sound of an audience specifically, rather than "noise" in general. That means it can pull out sounds that are loud, sudden and overlapping with the music without flattening the music itself.

It removes applause and clapping, cheering and shouting, whistling, screaming, laughter, and the general hum of a room full of people talking between and during pieces.

It keeps music and singing, spoken dialogue and announcements, instruments, and the natural sound of the room. If a soloist is speaking while the audience claps, the speech stays and the clapping goes.

What people use it for

Every one of these is a slightly different problem, and the model handles each differently.

Dance recitals and school shows

The classic case. A recital recording is usually a clean music bed played through a PA with sharp bursts of clapping on top, often right at the emotional peak of a piece. Because the clapping is short and loud while the music underneath is continuous, this is the scenario the tool handles best. If you recorded the whole show in one take, you can process it in one go and split the pieces afterwards.

Concerts, gigs and live music

A concert has two different noise problems: the applause between songs, and the continuous roar of a crowd during them. The between-song applause comes out almost completely. A crowd singing along with the band is harder, since their voices sit in the same range as the vocals you want to keep — expect it to be reduced rather than erased.

Theatre, school plays and musicals

Here the thing you need to protect is quiet dialogue, often recorded from the back of a hall, under audience laughter and applause. The model preserves speech, so lines delivered through a laugh usually survive. If the actors were not mic'd, the recording will still be quiet after cleaning — removing the audience makes the dialogue clearer, not louder.

Weddings, speeches and toasts

Speeches get buried under cheering, clinking glasses and a room full of side conversations. Cleaning the audio makes the speaker intelligible again, which matters most for the parts of the video people actually rewatch.

Sports matches and stadium recordings

A stadium is a continuous wall of noise rather than distinct bursts, which is the hardest case for any tool. If you are trying to recover commentary or a coach's voice, this will help. If you are trying to make a stadium sound empty, it will not get all the way there.

Church services, conferences and lectures

Congregational responses, applause after a talk, and audience shuffling all get removed while the speaker and any music stay intact. Useful when the recording is going to be published as a podcast or uploaded for people who missed it.

How to fix audio that came out too noisy

If a recording came back and you can barely hear the performance over the audience, the instinct is to reach for a noise reduction slider. That is usually the wrong tool, and it is why so many recordings end up sounding worse after an attempt to fix them.

CrowdRemoval uses AI source separation instead. Rather than measuring a noise floor and subtracting it, the model has learned what an audience sounds like and pulls that layer out of the mix as its own signal — much closer to un-mixing a recording than to filtering it. That is why it can remove a burst of clapping that lands directly on top of the music without leaving a hole where the music was.

Practically: upload the file, wait roughly half its length, and download the cleaned version. A ten-minute recording takes around five minutes. Nothing to install, and it runs in the browser on a phone as well as a computer.

The honest limit is that it can only recover what was captured in the first place. If the audience was so loud that the microphone clipped, or the music is barely present under the noise, cleaning will improve the balance but cannot restore detail that never made it onto the recording. Trying it costs nothing, so the fastest way to find out is to run the file through.

Why noise reduction in Audacity or iMovie doesn't fix this

People usually try a general-purpose editor first, and it usually disappoints. There is a specific reason.

Audacity's Noise Reduction is built around a noise profile: you highlight a passage containing only the unwanted sound, and it subtracts that fingerprint from the rest of the file. That works beautifully for steady, unchanging noise — tape hiss, an air conditioner, mains hum. Applause is the opposite of steady. It is broadband, impulsive, and its shape changes several times a second, so there is no single profile to capture. Push the sliders hard enough to dent it and the music turns watery and metallic, the artefact people describe as sounding underwater.

iMovie, CapCut and most phone editors give you a single "reduce background noise" slider aimed at exactly the same steady-hiss problem, with even less control. Premiere and DaVinci Resolve have stronger tools, but they are still working on the whole mix rather than separating the audience from it.

The difference is what the tool is looking for. A noise reducer looks for a constant background. CrowdRemoval looks for people.

What this tool doesn't do

It removes audience noise. It is worth being clear about what it is not, so you do not spend a credit finding out.

Wind noise

The low rumble from recording outdoors is not audience noise and will not be removed. A high-pass filter in any editor, rolling off below roughly 80 Hz, is the usual fix.

Hiss, hum and electrical buzz

Constant background noise from a cheap microphone or a mains loop is exactly what conventional noise reduction is good at. Audacity's Noise Reduction, using a profile taken from a silent passage, will do a better job here than this tool will.

Echo and reverb

A recording made in a large hall carries the reflections of the room. That is baked into the signal and needs a dedicated de-reverb tool; removing the audience will not make a boomy room sound dry.

Separating music from a voice

If you want to strip the backing track and keep only the vocal, or the reverse, that is stem separation rather than crowd removal. Tools built for that will serve you better.