Subtitles are timestamped chunks of text aligned to audio segments. The standard format is SRT (SubRip) — used by YouTube, video editors, and most playback software. Generating subtitles from audio is a common need: captioning a video, adding closed captions for accessibility, creating a translated subtitle track. mdisbetter currently outputs structured Markdown with inline timestamps; for SRT specifically, see the workflow below for converting our Markdown output to SRT format.
What mdisbetter outputs today
Audio to Markdown produces structured transcripts with inline timestamps — typically [12:34] markers at each topic shift and paragraph boundary, plus the actual text content. This is dramatically more useful than flat text for most workflows (show notes, content repurposing, search) but it's not SRT format directly. For SRT specifically, you need a conversion step from our Markdown output.
From mdisbetter Markdown to SRT
The conversion is mechanical: each timestamp + following text chunk becomes one SRT cue. A simple Python script (10-15 lines using a Markdown parser to extract timestamp/text pairs and emitting SRT format) handles it cleanly. Or paste the Markdown into ChatGPT/Claude with "convert this timestamped Markdown transcript to SRT subtitle format with 3-second cues" — works for most files in one pass.
For direct SRT generation, use OSS Whisper
OpenAI's open-source whisper command-line tool generates SRT directly: whisper input.mp3 --output_format srt produces the SRT file as part of normal operation. Same Whisper-class model as the major commercial transcription services. MIT-licensed, runs locally. For one-off SRT generation jobs, this is the simplest path. For larger batch SRT work, faster-whisper is the speed-optimised variant.
SRT export coming to mdisbetter
Direct SRT output is on our roadmap — currently the Markdown-with-timestamps output handles the underlying need for most users (transcription with timing data), and the conversion to SRT is mechanical for those who need it. As demand justifies, we'll add a one-click SRT export option to the audio converter UI.
Frequently asked questions
Do you output SRT directly?
Not yet — currently we output structured Markdown with inline timestamps ([12:34] markers). For SRT specifically, either: (1) convert our Markdown output to SRT with a small script or one ChatGPT prompt ("convert this timestamped Markdown to SRT subtitle format"), or (2) use OpenAI's open-source whisper CLI tool which outputs SRT directly with whisper input.mp3 --output_format srt. Direct SRT export is on our roadmap.
How do I generate YouTube subtitles?
YouTube accepts SRT, VTT, and several other subtitle formats. Workflow with mdisbetter: get the audio transcript via Audio to Markdown, convert the timestamped Markdown to SRT with a script or AI prompt, upload the SRT to YouTube via the video's subtitle settings. Alternative: use OSS whisper directly for SRT output and skip the conversion step. YouTube also has built-in auto-captions for videos uploaded to its platform; for higher accuracy than YouTube's built-in captions, the Whisper-class transcription mdisbetter uses outperforms YouTube's default in most cases.
What's the difference between SRT, VTT, and other subtitle formats?
SRT (SubRip) is the most universal — supported by virtually every video player and platform. VTT (WebVTT) is the web-native format, with more styling options, used by HTML5 video. SSA/ASS are advanced subtitle formats with positioning and styling, used by anime fan-subs and high-end film captioning. For most use cases, SRT is the right choice. Convert between formats with free online converters if needed.
Can I generate subtitles for non-English audio?
Yes — speech recognition auto-detects 50+ languages and transcribes in the source language. For translated subtitles (transcribe in source language, output subtitles in a different language), the workflow is two-step: transcribe with mdisbetter to get the source-language Markdown, then translate with our Markdown translator or paste into ChatGPT/Claude with "translate to [target language]". Convert the translated Markdown to SRT for the final captions.
Are auto-generated subtitles accurate enough for accessibility?
Auto-generated subtitles hit 92-97% accuracy on clean recordings, which is close to but not at the WCAG-recommended threshold for accessibility (typically 99%+ for compliance). For ADA-compliant accessibility captioning, plan a human review pass over auto-generated subtitles to catch and fix the 3-8% error rate, especially around proper nouns, technical terminology, and homophones. Auto-generated subtitles are a strong starting point that dramatically reduces caption-creation time vs from-scratch typing; the human review pass is what ensures accessibility compliance.