Image to Markdown — OCR Screenshots and Photos to MD
Drop an image and get the text inside it as clean Markdown. Works on screenshots, scans, photos of documents, whiteboards, and slides.
OCR built for AI workflows
Most OCR tools dump unstructured text. Ours uses layout-aware OCR plus an LLM post-processor: it reads the image, identifies headings vs body vs lists vs code vs tables, and outputs structured Markdown that’s ready for LLM ingestion.
Drop a screenshot of a webpage, a photo of a printed page, a scan of a hand-written note, a slide from a webinar, or a whiteboard photo from a meeting — you get back a Markdown document with semantic structure rather than a wall of broken text.
What it handles
Screenshots of articles, dashboards, and code editors
Photos of printed documents (books, receipts, contracts)
Scans, including multi-page PDFs of scanned pages
Hand-written notes (English and major Latin-script languages)
Whiteboards and meeting notes
Tables — converted to Markdown tables, not flat text
Code screenshots — wrapped in fenced code blocks with language detected
You can also caption images with AI: pass a photo and get a Markdown description plus alt text — perfect for accessible documentation and SEO.
How it works
Upload one image or a batch of images
Layout-aware OCR detects regions and reading order
An LLM post-processor structures the output as Markdown
Tables and code blocks are formatted natively
Copy the Markdown or download a .md file
Use cases
Turn screenshots of articles into clean Markdown for note-taking
OCR scanned contracts for AI-driven legal review
Extract code from screenshots into proper Markdown code blocks
Digitize whiteboard photos after meetings
Pull tables out of dashboard screenshots
Convert printed receipts and invoices into structured Markdown
Frequently asked questions
Which image formats work?
PNG, JPG, JPEG, WEBP, BMP, GIF, HEIC, HEIF, and TIFF. HEIC and TIFF are decoded server-side via libvips, so iPhone photos and legal/medical scans work without any client-side conversion.
Does it work on hand-writing?
For clean Latin-script handwriting, yes. Cursive and complex scripts produce more errors.
How accurate is it?
Print accuracy is typically >98%. Photo-of-screen accuracy depends on lighting and resolution.
Are tables detected?
Yes. Visible row and column boundaries are converted into Markdown tables when possible.
What about code, math, or charts?
Code screenshots become fenced code blocks with detected language. Equations become LaTeX ($inline$ / $$display$$). Charts become a data table plus a one-sentence summary.
Are images stored?
No. Files are processed in memory and deleted immediately after the request.
Can I OCR multiple images at once?
Yes. Drop several images and each one is processed sequentially with its own progress, then download every result as a ZIP. The ceiling is 50 files per batch, the same on every plan.
What happens if I upload the same image twice?
The second upload hits a content-addressed cache (SHA-256 keyed) — you get the result instantly with zero credits charged.