WAV is the lossless cousin of MP3: same audio, but stored uncompressed. It's what professional studios record into, what broadcast equipment outputs, what archival projects use to preserve recordings without generational loss. The downside is file size — a 10-minute stereo WAV is ~100 MB. The upside is that transcription accuracy on WAV is the ceiling for what's possible from any source format.
Why WAV is the gold standard for transcription
WAV files contain the original PCM audio waveform without any compression artefacts. For a transcription engine, that means: no compression noise that could be misheard as sibilants, no missing high frequencies that consonants depend on, no codec ringing around plosives. If you have the choice between a WAV and an MP3 of the same recording, WAV will produce a cleaner transcript every time. The difference is small for clean studio audio (already easy to transcribe) and larger for noisy field recordings (where every artefact-free decibel of signal-to-noise helps).
Studio and broadcast workflows
Radio stations, podcast studios, audiobook producers, and field journalists record to WAV by default. The workflow is: capture raw to WAV, edit the WAV in a DAW, and export a final MP3 only for distribution. MDisBetter accepts WAV directly so you can transcribe before the lossy export step happens — useful when you want a transcript of the unedited material (for journalist use: every word the source actually said) rather than the polished cut.
Structured Markdown output
The output structure is the same as for any other format: # H1 for the recording title, ## H2 at topic shifts, punctuated paragraphs, and [HH:MM:SS] timestamps at section boundaries. The transcription engine doesn't care that the input was WAV vs MP3; the structural pass that produces the Markdown runs the same way regardless of source format.
Before / After
Before (PDF):
[WAV audio file]
Uncompressed 16-bit 44.1 kHz stereo, 95 MB, 9 minutes
Metadata: BWF (Broadcast Wave Format) chunk with recording date, source, engineer
(no extractable text — raw PCM waveform)
After (Markdown):
# News Interview — Mayoral Candidate Q&A
## Opening
[00:00:02] I'm here at City Hall with mayoral candidate Maria Lopez. Maria, thanks for joining us. My pleasure, always happy to talk to constituents.
## On Housing Policy
[00:00:18] Housing has dominated this race. What's your specific proposal? Three things. First, we permit-by-right for any project under 50 units…
Frequently asked questions
Is WAV transcription more accurate than MP3 transcription?
Marginally yes, especially on noisy or low-volume source material. For clean studio audio the gap is negligible (modern engines are robust to typical MP3 compression). For field recordings, voicemails, or anything captured in a noisy environment, the lossless WAV gives the engine a few extra dB of clean signal that translates to fewer confused words per minute.
What's the maximum WAV file size I can upload?
WAV files are large — a 1-hour stereo recording is roughly 600 MB. Our primary transcription path handles large files via direct URL streaming (no reupload), so episode-length WAVs under that ceiling work fine. For multi-hour single files, splitting into 1-hour WAV chunks before upload keeps the workflow predictable.
Are BWF (Broadcast Wave Format) metadata chunks preserved?
BWF metadata (recording date, originator, engineer notes) is read for context but not preserved in the Markdown output by default. If your workflow depends on round-tripping that metadata, keep the original WAV alongside the Markdown — the transcript adds searchable text without replacing the source archive.
Does WAV transcription cost the same as MP3?
Yes — pricing is by audio duration, not file size or format. A 30-minute WAV and a 30-minute MP3 of the same recording cost identically in our credit system. The only difference is upload time (WAV is bigger), not the transcription cost itself.
Can I transcribe 24-bit or 96 kHz studio WAVs?
Yes — high-resolution WAVs (24-bit, 96 kHz, even 192 kHz) are accepted. Internally the audio is downsampled to the rate the transcription engine uses (16 kHz mono is sufficient for speech recognition), so the extra fidelity isn't leveraged for accuracy, but the upload still works without conversion on your side.