Sound adds together
A microphone measures air pressure over time. When two people speak, their waves add into one signal; the microphone does not label who made which part.
Learn with PaperEdits · Lesson 1
Preparing the two voices…
A microphone measures air pressure over time. When two people speak, their waves add into one signal; the microphone does not label who made which part.
A short-time Fourier transform divides the recording by time and frequency. Each tile asks which voice has more energy at that moment and pitch.
Click bright yellow tiles where both speakers are strong. Listen to each contribution. Those collisions are where clean separation becomes hardest.
It is the challenge of focusing on one speaker when several voices reach the same ears or microphone. Humans use pitch, timing, location, and familiarity; software searches for related patterns in the signal.
Noise removal often assumes the unwanted sound is steady or unlike speech. A second speaker changes constantly and shares the same frequencies as the speaker you want, so it requires source separation instead.
No. Both voices can occupy the same tile. A binary mask makes a hard choice, which is simple to see and hear but can discard part of the quieter speaker. Modern systems can estimate softer masks instead.
sample voices: LibriSpeech (openslr.org/12), CC-BY 4.0 — read by Heather Barnett & Anders Lankford · STFT 1024 / Hann / 4× overlap / ideal binary mask · built while making paperedits.com