5 min read
Converting audio to text: what to prepare before you hit transcribe
Transcription engines are only as good as the file. Ten minutes of file prep beats an hour of correcting a transcript.
Automatic transcription has become the backbone of documentary and interview editing. Accuracy varies enormously, and most of the variance comes from the file, not the engine.
One speaker per track beats one mixed track
If your interview was recorded with a lav on each person, transcribe each channel separately. Speaker separation is the hardest part of the job, and you already solved it on set. Split the poly WAV into named mono files and run them individually.
Mono is fine, stereo is often worse
Most engines downmix to mono anyway. If your stereo file has a boom on the left and a lav on the right, the sum can create phase cancellation on the voice. Choose the best channel instead of summing blindly.
Sample rate: do not resample twice
Engines typically work internally at 16 kHz. Feed them your 48 kHz original and let them downsample once, rather than exporting a low rate copy yourself and having it converted again.
Level and noise
- Normalise quiet interviews so speech sits well above the noise floor.
- Notch out constant hum, it confuses word boundaries.
- Do not heavily denoise before transcription, artefacts hurt more than hiss.
- Trim long silences to save processing time and cost.
Naming so the transcript is usable later
Name the exported files with speaker, scene and take, because most tools reuse the filename as the transcript title. A folder of INTERVIEW_MARIE_01.wav is searchable six months later. A folder of track_3.wav is not.
You can do all of this here: split the interview poly WAV, keep the recorder's track names, adjust levels per channel and export mono WAVs ready to transcribe, without uploading confidential rushes anywhere.