AAC (Advanced Audio Coding) is the format Apple, YouTube, and most streaming platforms use under the hood. iTunes purchases come as AAC. YouTube audio streams are AAC. Spoken-word podcasts distributed by Apple Podcasts are typically AAC. The bare <code>.aac</code> extension is less common than M4A (which is AAC inside an MP4 container) but the underlying audio is the same — and MDisBetter handles both.
AAC, M4A, and the container question
AAC refers strictly to the audio codec — the compression algorithm. It usually rides inside one of two containers: a bare .aac file (rare, no metadata wrapper) or an MP4 container with an .m4a extension (common, supports tags, chapter markers, and album art). The transcription pipeline handles both — drop a .aac or a .m4a, the engine decodes the audio and emits Markdown identically.
Where AAC files come from
iTunes Store purchases from Apple are AAC (the iTunes Plus tier at 256 kbps; older iTunes purchases were 128 kbps DRM-locked AAC, now mostly upgraded to DRM-free). YouTube audio streams are AAC at multiple quality tiers; tools that download YouTube audio (yt-dlp, etc.) typically output AAC or repackage to MP3. Apple Podcasts and Apple Music serve AAC. Most TVs, streaming devices, and broadcast workflows use AAC for audio because of its quality-per-bit advantage over MP3.
Practical use cases
The most common path: someone rips a YouTube interview, lecture, or podcast to AAC for offline reading and wants a transcript. The MDisBetter web tool takes the AAC, transcribes it into sectioned and timestamped Markdown, and gives you the file in a few minutes. For batch ripping (e.g., archiving a YouTube channel's spoken-word content), the OSS path is yt-dlp + faster-whisper running locally; the web tool is the no-setup path for one-off files.
# Audiobook Sample — Chapter 1
## Opening
[00:00:03] Chapter one. The morning of October the third arrived without any warning that the world was about to change.
[00:00:14] At least, that's how it would be remembered later. At the time, the morning seemed completely ordinary…
Frequently asked questions
Are AAC files from YouTube transcribed correctly?
Yes — YouTube audio rips (whether saved as .aac, .m4a, or repackaged to MP3) transcribe with the same accuracy as any other source. The output is continuous text with headings and timestamps rather than labelled speaker turns; for many-speaker panel discussions the text stays accurate but you will not see who said what. For YouTube videos directly without ripping first, see YouTube to Markdown.
Is bare .aac different from .m4a for the transcription pipeline?
For the transcription itself, no — the audio decodes identically. The difference is metadata: M4A files (AAC inside an MP4 container) carry tags, chapter markers, and album art; bare .aac files don't. We use any chapter markers in M4A as section break hints in the output Markdown.
Can I transcribe DRM-protected iTunes purchases?
No — DRM-locked AAC files (the older iTunes Music Store format pre-2009, and any current Apple Music subscription content) can't be decoded outside the Apple ecosystem. Most iTunes Store purchases sold today are DRM-free and transcribe normally; Apple Music subscription tracks are DRM-locked and can't be transcribed by any tool that doesn't bypass the DRM (which we don't).
Does AAC give better quality transcription than MP3?
Marginally yes at the same bitrate — AAC is more efficient than MP3, so a 128 kbps AAC sounds better than a 128 kbps MP3 and transcribes slightly more accurately. The gap is small for clean speech and larger for music-on-top-of-speech tracks. For most use cases the difference is negligible.
Can I transcribe an audiobook from an AAC file?
Yes — audiobooks distributed as AAC (iTunes audiobooks, Audible files exported via authorised tools, public-domain audiobooks from LibriVox in AAC) all transcribe normally. Long files (hours of audio in a single AAC) are processed by streaming, but for very long single files, splitting by chapter beforehand keeps the workflow predictable.