Finding a real recording of two people talking at once, where you also know exactly what each of them said, is nearly impossible. So we build them. Every file below is constructed: the overlap boundaries, the noise levels and the reference words are known to the decimal, because we placed them. It is how we test PaperEdits, and you are welcome to test anything with it.
Every file plays here, and every file opens straight in the editor, transcription and analysis run in your browser, nothing uploads.
Two different speakers, equal level, fully simultaneous for 10.5 seconds. The hardest everyday case, a plain transcriber returns fluent text that neither person said.
Test this in the editor →One person presenting; a second cuts in twice with short remarks. The common podcast/interview shape.
Test this in the editor →Clean turn-taking except two seconds where both talk across the handoff, the moment meeting recordings always garble.
Test this in the editor →Fifteen seconds of one speaker before a second joins. Built to catch tools that behave differently when overlap starts late.
Test this in the editor →Three voices piling in two seconds apart. Most separation models output exactly two voices, this file is the test of whether a tool admits that limit or hides it.
Test this in the editor →The second voice is pitch-shifted toward the first’s register. A crude similarity proxy, labelled as such, separating alike voices is much harder than separating different ones.
Test this in the editor →The lead-in case with a real HVAC recording underneath, mixed at a known level.
Test this in the editor →Background crowd noise, itself made of voices, under the two speakers. Noise that competes with speech in its own band.
Test this in the editor →Impulsive keystrokes over fully overlapped speech. Bursty noise is what fixed filters provably cannot follow.
Test this in the editor →The full-overlap pair through an echo chain. Measured honestly: our own overlap detector goes nearly blind here, published as a known limit, not hidden.
Test this in the editor →A single speaker with echo. Sounds like two people to detection models, the echo is a phantom second voice. Also a known limit, published.
Test this in the editor →The lead-in case after an AAC round trip, what audio actually looks like when it arrives inside a phone video. The encode alone costs about 3 dB in the voice band.
Test this in the editor →One clean voice over synthesized 50 Hz hum with two harmonics, the classic bad-cable sound. The case notch filters fully fix.
Test this in the editor →One voice over a synthetic chord progression. Suppression models treat music as noise and remove it, sometimes that is what you want, sometimes not.
Test this in the editor →A synthesized voice reading a script with counted fillers: um ×3, uh ×4, hmm ×1, like ×1. We measured transcribers silently rewriting these into real words, this file is how you check yours.
Test this in the editor →Voices come from Microsoft’s open MS-SNSD speech set, placed on exact timelines. Noise is mixed at stated signal-to-noise ratios with normalization disabled, an ffmpeg default that silently halves speech and once cost us a day. Reverb and codec cases are labelled with what they do. Two files document our own blind spots: the echo cases defeat our overlap detector, and this page says so, because a test set that hides its maker’s failures is marketing.
The generator is open, scripts/make-test-corpus.mjs in the repo rebuilds everything from source, with a manifest carrying every ground-truth number. Source recordings: MS-SNSD, MIT license.