You have audio. You want the words from it as text. Drop an MP3, WAV, M4A or any common audio format into the converter and get clean plain-text transcription back. Useful when the downstream consumer is a script, a search index, or any tool that wants flat UTF-8 strings rather than structured Markdown. For AI input where structure helps, use the Markdown variant.
How "Audio to Text" actually works here
Every audio file goes through the same Whisper-class speech recognition pipeline as our Markdown converter, but the structural formatting (H2 sections, timestamps) is stripped at the end. Output is flat plain text — paragraphs separated by blank lines, no other structural markers. Copy-paste-ready, search-indexable, ready for any tool that wants UTF-8 string input.
Format support
MP3, WAV, M4A, FLAC, OGG, AAC, WebM, AMR — basically every common audio container. File size limit on the free tier is generous enough for typical podcast episodes and interviews; Pro tier handles multi-hour recordings without splits. Audio quality matters more than format: a 64kbps MP3 transcribes just as accurately as a 320kbps version for speech content.
Single file workflow
Paste the file in, click convert, get the text back. For batch transcription of large back-catalogues (hundreds of files), use openai-whisper or faster-whisper locally — same model class, MIT-licensed, runs on CPU or GPU, processes hundreds of hours overnight. mdisbetter's web tool is the right choice for one-at-a-time conversions where the per-file workflow is acceptable.
Frequently asked questions
What's the difference between Audio to Text and Audio to Markdown?
Audio to Text gives you flat UTF-8: paragraphs, no headings, no timestamps in the output. Audio to Markdown splits at topic shifts as ## H2 sections, keeps punctuation and paragraphing, and embeds timestamps inline. Neither identifies speakers. Use Text for plain-text scripts, search indexing, or NLP pipelines that don't need structure; use Markdown when structure matters (AI input, show notes, interview transcripts).
How accurate is the transcription?
92-97% word accuracy on clean recordings (single mic, native or fluent English speaker, quiet background). Phone-quality audio, accented speech, or noisy environments see 5-15% lower accuracy. Same Whisper-class model as the major commercial transcription services, including the OSS Whisper that everyone uses.
What audio formats work?
MP3, WAV, M4A, FLAC, OGG, AAC, WebM, AMR — every common container. The audio is auto-converted to the format the speech recogniser needs internally. Format choice doesn't affect accuracy; bitrate and recording quality do.
Is there a length limit?
Free tier handles typical podcast episodes and interviews (up to ~60 minutes per file). Pro handles multi-hour recordings without splits. For files much longer than 60 minutes on free tier, split before upload (Audacity, ffmpeg, any audio editor handles this in under a minute).
Free Audio to Text vs paid services — what's the catch?
No catch — our free tier covers most personal use. The catch with "free transcription" elsewhere is usually watermarks, mandatory account creation that hands your email to a marketing pipeline, output truncated to N minutes with "go pro for the rest" walls, or quietly uploading your audio to a training dataset. We don't do any of those. Audio is processed in memory and deleted immediately after conversion.