What you get back
The output is a single .md file. The video title becomes the H1. Each chapter (if the creator added them) becomes an H2 with its starting timestamp. Speech is split into paragraphs by topic shift. We do not identify speakers, so a multi-person video comes back as one continuous transcript rather than labelled turns.
Why this beats YouTube's built-in transcript
The native transcript is auto-generated by YouTube's captioning system: no punctuation, no structure, 15-20% word-error rate on technical or accented content. Our pipeline re-transcribes the audio with a higher-accuracy model and emits structured Markdown — usable directly in Obsidian, Notion, or as input to ChatGPT/Claude.
What it doesn't do
It doesn't download the video file itself. It doesn't identify speakers (Whisper transcribes speech, it doesn't tell voices apart). It doesn't handle private or auth-required videos.