You have a recording of someone speaking. You want the words written down. Speech-to-text is the simplest possible framing of the transcription job: audio in, text out. Upload an audio file, click convert, get plain UTF-8 text back. For interviews and conversations where you want topic structure and timestamps, the Markdown variant is more useful, but for the basic "give me the words" job, plain text is the cleanest output.
What "Speech to Text" means here
Spoken words become written text. The pipeline: upload audio file → speech recognition (Whisper-class model, 50+ languages auto-detected) → punctuation and capitalisation restoration → paragraph break insertion → flat plain-text output. No structural markers, no headings, no timestamps in the output. Just the words, in paragraphs, ready to paste anywhere.
What it works on
Voice memos from your phone. Recorded interviews. Podcast episodes. Lecture recordings. Voicemails. Conference talks. Recorded video calls (audio extracted automatically). Single-speaker dictation for note-taking. Multi-speaker conversations (every voice is transcribed, but nothing marks who is talking, in either variant). Anything with audible spoken word.
What it doesn't do well
Music transcription (the model is trained on speech, not melody). Singing (close to speech but lyrics often get garbled). Heavy crosstalk where multiple people speak simultaneously (single speaker comes through, others get clipped). Extremely noisy environments where signal-to-noise is poor. For these cases, the right tool is a dedicated audio-cleanup pass first (Adobe Podcast, Krisp, Auphonic), then transcription.
Frequently asked questions
What's the difference between speech-to-text and audio transcription?
Same thing, different framing. "Speech to text" emphasises the model — turning spoken words into written text. "Audio transcription" emphasises the file workflow — processing an audio file end-to-end. The technology is identical: a speech recognition model converts the audio into text. We use both terms to match how different users search for the capability.
Can it handle multiple speakers?
It captures all speech but doesn't label who said what, output is flat text without speaker attribution. The same is true of the Markdown variant: mdisbetter does no speaker identification anywhere in the product. Audio to Markdown gives you punctuation, paragraphs, ## H2 sections at topic shifts and inline timestamps, but the text runs continuously with no per-speaker labels. If knowing who spoke each line genuinely matters, use a tool built for diarisation (AssemblyAI, Deepgram, WhisperX, pyannote-audio) and bring the labelled result back into your workflow.
What languages are supported?
Auto-detection across 50+ languages including English, Spanish, French, German, Portuguese, Italian, Mandarin, Japanese, Korean, Hindi, Arabic, Russian. Accuracy varies by language: top tier (English, Spanish, French) hits 92-97% on clean audio; tier two (most European, major Asian languages) hits 85-92%; tier three (low-resource languages) hits 70-85%. Mixed-language audio in a single file works but accuracy drops compared to single-language input.
How long can the audio be?
Free tier handles up to ~60 minutes per file. Pro handles multi-hour recordings in a single pass. For longer files on free tier, split with any audio editor (ffmpeg one-liner, Audacity, online splitters) before upload. Quality and accuracy don't change with length.
Is the output good enough for legal or medical use?
For routine use cases (interview reference, note-taking, content repurposing), yes. For high-stakes legal records (depositions, court proceedings) or clinical PHI, no — use certified court reporters for legal records and HIPAA-compliant medical dictation services (Suki, Nuance Dragon Medical) for clinical use. mdisbetter is appropriate for working transcripts and general transcription, not for the certified-record or PHI-handling use cases.