Convert Recording to Text — Transcribe Any Recording
You have a recording — interview, lecture, meeting, voice memo, podcast, call, conference talk. You want the text out of it. Upload the recording (audio or video), get the transcript back in minutes. Works on virtually any audio or video format. For longer recordings where structure matters, the Markdown variant adds H2 topic sections and inline timestamps.
Recording → text is one job, many sources
The umbrella case: you have audio (or video, where audio is auto-extracted) and you want the words written down. The source could be: an iPhone voice memo, a Zoom meeting recording, a downloaded podcast episode, a recorded interview from a research project, a lecture you recorded with permission, a recorded customer support call, a video file from your camera, an audiobook chapter, an old cassette tape you digitised, anything. mdisbetter handles all of these through the same upload-and-convert flow.
Audio and video both work
For pure audio files (MP3, WAV, M4A, FLAC, OGG, AAC, WebM, AMR), upload to /convert/audio-to-text or directly /convert/audio-to-markdown. For video files (MP4, MOV, MKV, WebM, AVI), audio is extracted automatically — you can upload the video directly, or use /convert/video-to-markdown for the video-specific page. For YouTube URLs, /convert/youtube-to-markdown handles the download + transcribe in one step.
For specific recording types, see the focused pages
For batch transcription of many recordings, use OSS
If you have 50+ recordings to transcribe in one go, mdisbetter's web tool is the wrong shape — it's one-upload-at-a-time. Run faster-whisper locally on a GPU for batch processing. MIT-licensed, free, processes hundreds of hours overnight. Use mdisbetter's web tool for one-off conversions where the manual workflow is acceptable.
Frequently asked questions
What kinds of recordings work?
Any audio or video recording with audible spoken word. Phone voice memos, meeting recordings (Zoom/Teams/Meet), podcast episodes, interviews, lectures, recorded calls, conference talks, voicemails, video files (audio auto-extracted), digitised cassette tapes, anything. The audio quality and speaker count matter more than the recording source — single-mic single-speaker = highest accuracy; phone-call multi-party = lower accuracy; conference-room with distant mic = lowest of the common cases.
Audio recording vs video recording — which page do I use?
Free tier handles up to ~60 minutes per file. Pro handles multi-hour recordings in a single pass. For longer recordings on free tier, split with any audio editor before upload. ffmpeg one-liner: ffmpeg -i input.mp3 -f segment -segment_time 1800 -c copy out%03d.mp3 splits into 30-minute chunks.
How accurate is the transcription?
92-97% on clean recordings (single mic, native or fluent speaker, quiet environment). Lower on phone audio, accented speech, or noisy backgrounds. Same Whisper-class model accuracy as major commercial transcription services. For challenging recordings, run audio cleanup first (Adobe Podcast Enhance, Krisp, Auphonic) — typical accuracy boost is 5-15%.
Where can I get more help on specific recording types?