Markdown is for humans (and LLMs). JSON is for code. When your pipeline needs to programmatically navigate a converted document — pull every H2, list every table, count code blocks — JSON output is the right shape. We emit a structured tree where each PDF maps to a JSON object.
The output schema
{ metadata: { title, author, pages }, sections: [{ heading, level, content, subsections, tables, lists }, ...] }. Each section is a recursive node — sections nest inside sections by heading depth. Tables are emitted as 2D arrays of cell text. Lists are arrays of items. Code blocks carry their language tag.
Use cases
Building a pipeline that needs to extract specific sections programmatically (e.g. "every Methodology section across 500 papers"). Feeding structured data to a downstream model that expects JSON. Building a CMS importer that maps PDF structure to your own schema. Comparing document structures across versions.
Frequently asked questions
What does the JSON output schema look like?
Top-level: { metadata, sections }. Each section: { heading, level, content, subsections, tables, lists, code_blocks }. The full schema is documented in the API reference; the structure is recursive so deeply-nested PDF content maps cleanly.
Are tables represented as arrays in JSON?
Yes — tables come through as 2D arrays of cell text, with a separate headers field for the first row when detected. Cell formatting (bold, italic) is dropped in JSON output; if you need it, request Markdown instead.
How is metadata extracted?
Title, author, creation date, page count, and any embedded keywords are pulled from the PDF's metadata block when present. Falls back to inferring title from the first H1 in the document if metadata is empty.
Can I get JSON via the API?
Yes — pass ?format=json on the API endpoint, or set Accept: application/json with the appropriate parameter. Same auth flow as the Markdown endpoint.
When should I use JSON vs Markdown output?
JSON when your consumer is code that needs to navigate the structure programmatically. Markdown when your consumer is an LLM (cheaper tokens, better comprehension) or a human (readable). They're the same content, different shapes.