Voice Memo to Markdown — Transcribe Phone Recordings
A voice memo is the fastest possible way to capture an idea, an interview, a meeting, or a thought you don't want to lose. The slow part is going back and listening to it later — most people record voice memos and never re-open them. MDisBetter turns voice memos into searchable, scannable Markdown so the recording you made on a walk in May is something you can actually use in June.
What "voice memo" means in practice
The term covers any short-to-medium-length audio recording made on a phone, tablet, or dedicated recorder, regardless of the underlying file format. iPhone Voice Memos output M4A. Android Recorder apps typically output M4A or MP3. Samsung Voice Recorder outputs M4A. Pixel Recorder outputs M4A. Standalone digital voice recorders (Olympus, Sony, Tascam consumer models) usually output WAV or MP3. MDisBetter accepts all of those directly — see also format-specific pages: M4A, MP3, WAV.
Common voice-memo use cases
Idea capture while walking: the rambling 5-minute voice memo of "I should write about X, then Y, then maybe Z" turns into a Markdown bullet outline you can paste into a draft. Field interviews: journalists, researchers, qualitative-research practitioners hit Record, conduct a 30-minute conversation, transcribe it for review and quote extraction. Meeting capture without a bot: place phone on conference table, hit Record, transcribe afterwards. Doctor/lawyer/exec dictation: see Dictation to Markdown for the structured-dictation use case.
Why structured Markdown beats plain transcription
A 20-minute voice memo as one wall of text is hard to use. Structured Markdown — section headings at topic shifts, real punctuation and paragraphs, timestamps at major junctures — lets you scan the recording in 30 seconds, jump back to the audio for a specific quote, and feed the transcript to an LLM for summarisation without it drowning in an undifferentiated wall of text. That's the difference between a voice memo as a graveyard of unprocessed audio and a voice memo as part of an actual workflow.
Before / After
Before (PDF):
[Voice memo audio file]
M4A from iPhone Voice Memos, 6.2 MB, 22 minutes mono
Metadata: recording date, location, device model
(no extractable text — phone-captured waveform)
After (Markdown):
# Voice Memo — Walk on May 9
## On the Q3 Plan
OK so thinking about Q3, the three big things are going to be: one, finishing the migration off legacy infra; two, hiring two more engineers on the platform team; three, shipping the analytics rebuild. Of those, the migration is the highest-risk because of the dependency on the vendor timeline, and the hiring is the slowest because Q2 was so quiet on the recruiting side.
## On the Migration
Biggest open question is whether we cut over in one big bang or do it in phases. Phases is safer but takes 2x as long…
Frequently asked questions
Will voice memos recorded while walking outdoors transcribe accurately?
Mostly yes, with caveats — modern phone microphones do well at capturing the speaker's voice and rejecting moderate background noise (wind, traffic, distant voices). Heavy wind directly on the mic, very loud street noise, or recording in a moving vehicle degrades accuracy. For best results outdoors, hold the phone close to your mouth and speak in normal volume; the transcript will recover from minor noise without much loss.
How long can a voice memo be?
Practically, no hard ceiling for typical phone-captured audio — phone Voice Memos and similar apps will record for hours into a single file. The MDisBetter pipeline handles long files via streaming. For very long single recordings (2+ hours), splitting into ~1-hour chunks before upload makes the workflow more predictable and produces more navigable per-chunk Markdown.
Can I transcribe a voice memo without uploading the original audio?
No — transcription requires the audio. The MDisBetter web tool uploads the file (or accepts a publicly-accessible URL), processes it on our servers, and returns the Markdown. For self-hosted privacy-sensitive workflows, the OSS path is faster-whisper running locally on your laptop with no network involved at all.
Does a two-person recorded interview come back with each voice labelled?
No, we do not identify speakers. A two-person interview transcribes as one continuous, properly punctuated document with ## H2 sections at topic shifts and inline timestamps. The words are accurate; the attribution is not there. If you need to know who said each line, run the same file through a dedicated diarisation tool (WhisperX, pyannote) and line the two outputs up by timestamp.
Can I include voice memos in my notes vault (Obsidian, Logseq, etc.)?
Yes — saving the converted .md alongside the original .m4a/.mp3 in your vault gives you both the searchable text (for grep, for Obsidian backlinks, for AI summarisation plugins) and the original audio (for tone, for verification, for re-listening). Many notes systems support audio file embedding so you can play the original directly from a Markdown link.