Document AI data
PDF Parsing Training Data: Page Images Paired with Structured Markdown or HTML
Quick answer
PDF to Markdown training data is a set of page images, or the PDFs that render them, each paired with a complete structured transcription: headings, lists, tables, formulas, captions and reading order expressed in Markdown, HTML or JSON. The strongest labels come from native source files (DOCX, PPTX, LaTeX, HTML) rendered to PDF, supplemented by human-verified transcriptions of scanned pages. Buyers should fix markup conventions, coverage by document type and evaluation metrics before sourcing anything.
By SourceX Editorial · Updated
Why page-to-markup pairs are a distinct dataset type
Page-to-markup pairs label the whole document structure, which neither OCR ground truth nor layout boxes do on their own. OCR ground truth tells a model which characters sit on a line; layout datasets such as DocLayNet tell it where a table or caption region is, with 80,863 pages labeled in 11 classes [3]. An end-to-end parser that emits markup from a page image needs a target that encodes hierarchy, table cells and math in one sequence. That target is the product you are buying.
Public supply is thin for this exact pairing. A December 2025 paper on end-to-end PDF-to-Markdown conversion states that no existing PDF-to-Markdown dataset had been released as PDFs, so the authors compiled arXiv LaTeX into 180,146 labeled pages [1]. They also excluded hard pages, including full-page images, long tables and bibliographies [1]. Those exclusions are exactly where business documents differ from papers, which is why academic-only corpora transfer poorly to invoices, board decks and loan files.
Most search results for this query are conversion tools rather than datasets. Course notes compare PyMuPDF4LLM, Markitdown, Unstructured, Docling and GROBID as converters [5], and Adobe's PDF Extract documentation lists training-data preparation, with review, cleanup and labeling of the output, as a PDF-to-Markdown use case [4]. Converter output is a starting point for labels, not ground truth.
Native-source pairs versus transcribed scans
Native-source pairs give exact structure cheaply, while transcribed scans give realism; a production parser needs both. When a company holds the original DOCX, PPTX, XLSX or HTML that produced a PDF, you can render the PDF and derive Markdown or HTML from the source's own object model: heading styles, list levels, table grids and merged cells come straight from the file rather than from inference. PubTabNet used the same idea at table level, aligning PubMed Central XML with the PDF to pair about 568k table images with HTML [2].
Native pairs have known fidelity gaps that you should test, not assume away. Common failure modes include:
- Rendering drift: the PDF printer substitutes fonts, reflows text boxes or drops hidden slides, so the label contains text the image lacks.
- Style misuse: authors fake headings with bold body text, or build lists with manual numbers, so the source model under-labels structure.
- Floating objects: text boxes, SmartArt and anchored images serialize in insertion order, not visual reading order.
- Headers, footers and tracked changes: these may appear in the source but not on the rendered page, or the reverse.
Scanned and photographed pages have no native source, so labels must be transcribed or corrected by people against the image. They teach robustness to skew, stamps, fax noise and handwriting margins that rendered PDFs never show; see degraded document capture data for the capture conditions to specify. A reasonable working hypothesis is to train mostly on native pairs for structure and reserve a human-verified scanned slice for both training and a held-out test.
Markup conventions to fix before labeling
A parsing dataset is only consistent if the label grammar is written down before the first page is labeled. Markdown cannot express merged cells, multi-row headers or nested tables, so most teams emit HTML tables inside Markdown, which is how PubTabNet represents structure [2]. Equations need one convention, usually LaTeX between delimiters, because a PDF stores positioned glyphs rather than formula structure, so math is where extracted text loses the most meaning.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Element | Recommended label convention | Decision you must record |
|---|---|---|
| Headings | # to ####, depth from source style or visual hierarchy | Maximum depth; whether numbered section labels stay in text |
| Lists | - and 1. with two-space nesting | How to label manually numbered paragraphs |
| Tables | HTML <table> with rowspan/colspan, <thead> for header rows | Whether empty cells, footnote markers and units stay inline |
| Formulas | $...$ inline, $$...$$ display, LaTeX | Treatment of equation numbers and chemical notation |
| Figures |  plus caption text | Whether chart data is transcribed (see chart data below) |
| Headers and footers | Excluded from body, kept in a page_furniture field | Page numbers, running titles, confidentiality legends |
| Footnotes | [^n] placed at reference point, text at page end | Cross-page footnote continuation |
| Reading order | Sequence order of blocks in the label | Rule for multi-column, sidebars and callouts |
| Redaction tokens | [NAME], [ACCOUNT] style placeholders | Must match tokens in the page image, if image is also masked |
Charts deserve their own decision: a caption-only label teaches nothing about the plotted values, while chart understanding data paired with source tables teaches extraction. Tables are similar, and if table accuracy is your bottleneck, a dedicated table structure recognition dataset may be cheaper than relabeling whole pages.
A record format that supports training and audit
Ship each page as one record that keeps image, label, provenance and label method together. JSON Lines works well for the manifest because each line is one UTF-8 JSON value with no byte order mark [6], and a Croissant JSON-LD descriptor can declare the file resources and record fields so loaders do not guess [7]. Keep page images as lossless PNG or the original PDF page, and record render DPI.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"page_id": "doc0412_p003", "doc_id": "doc0412", "page_no": 3, "page_count": 11,
"image": "images/doc0412_p003.png", "pdf": "pdf/doc0412.pdf", "dpi": 200,
"doc_type": "vendor_contract", "capture": "native_render",
"label_md": "labels/doc0412_p003.md", "label_html": "labels/doc0412_p003.html",
"label_method": "source_docx_conversion+human_review", "source_format": "docx",
"reviewer_pass": 2, "reading_order_verified": true,
"page_furniture": {"header": "Master Services Agreement", "footer": "Page 3 of 11"},
"deidentification": "names_accounts_replaced_tokens", "split": "train"}
The label_method and capture fields matter most. They let you report accuracy separately for native renders and scans, and they let you drop machine-only labels if a licensing or quality issue appears later. Keep doc_id so splits are grouped by document, not by page; a page-level split leaks near-identical templates between train and test.
Evaluating a parser and a candidate dataset
Evaluate with several metrics broken out by element and document type, because a single overall score hides where parsers fail. Report normalized edit distance for text and formulas, a table-structure score, and a reading-order score, each split by document type and by native render versus scan. Tables are typically scored with TEDS, the tree-edit-distance similarity introduced with PubTabNet [2].
Use the same metrics to judge a dataset before you buy it. Run your current parser on a sample, compute normalized edit distance and TEDS against the vendor's labels, then inspect the worst pages by hand: large errors are either your model's weakness, which is what you want to buy, or label errors, which you do not. Ask for double-labeled pages, as DocLayNet did to measure inter-annotator agreement [3], so you know the ceiling any model can reach. More on test design is in the training data quality assessment guide.
Pseudo-labels, provenance and rights
Labels produced by another provider's parser or model can carry that provider's output terms, so record which tool produced every draft label. If a vendor bootstrapped labels with a commercial extraction API or a hosted VLM, ask for the tool name and version, the terms in force at the time, and what share of pages a human corrected. Pages labeled purely by an open-source converter still need correction, since many converters target LLM-ready text rather than faithful structure [5].
The page images themselves raise the usual document-rights questions: who owns the documents, whether customer or employee personal data appears, and whether the documents came from public web crawls subject to text and data mining opt-outs. Business documents carry names, signatures and account numbers in headers and tables, so ask how masking was applied to both the image and the label, and check a sample yourself. The de-identified data guide covers methods and residual risk.
Sourcing checklist for page-to-markup data
Write the request around document types, label grammar and verification, not around volume alone.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Document mix: target types (contracts, statements, slide decks, forms, manuals), page-count distribution, share of multi-page and long documents.
- Capture mix: share of native renders versus scans, faxes and phone photos; render DPI and scan DPI.
- Source files: whether DOCX, PPTX, XLSX or HTML originals exist and are included; see accepted file formats.
- Label grammar: your convention table, attached, with three worked pages.
- Label method: native conversion, human transcription, or model draft plus human review, recorded per page.
- Quality evidence: double-labeled subset, TEDS and edit-distance scores against a second pass, list of excluded page types.
- Splits: document-level grouping, a held-out scanned slice, no template overlap.
- Rights and privacy: ownership, consents, masking method, upstream tool terms for any pseudo-labels.
- Delivery: JSONL manifest, images, labels and a Croissant descriptor.
Where business documents come from
The highest-value pages for business parsers include the ones academic corpora exclude, such as long tables and full-page images [1], plus scanned attachments, forms and slide decks. Those sit in company archives such as contract repositories, finance close binders, engineering document control systems and support knowledge bases, often alongside the native files that produced them. SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request; nothing is held in stock and a request does not guarantee a match. You can describe the documents and label grammar you need through the SourceX buyer request page, and the enterprise document datasets overview explains the archive types involved.
If your goal is retrieval quality rather than parser training, start with RAG evaluation datasets from real company documents. For the wider map of document tasks and public sets, see the Document AI datasets hub and the AI data hub.
Request PDF to Markdown training data from real business documents
SourceX looks for US businesses that hold the documents you describe, reviews ownership and consents, and delivers each approved dataset under a license defining records, uses, term and delivery. Personal details are removed or replaced before delivery and the method is recorded. Describe your document types, capture mix and markup conventions at sourcex.si/buyers.
Sources
- arXiv, "Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation" (2025). https://arxiv.org/pdf/2512.18122
- arXiv (IBM Research), "Image-based table recognition: data, model, and evaluation (PubTabNet)" (2019). https://arxiv.org/pdf/1911.10683v5
- arXiv (IBM Research; KDD 2022), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- Adobe, "PDF to Markdown API how-to (PDF Extract API)". https://developer.adobe.com/document-services/docs/overview/pdf-extract-api/howtos/pdf-to-markdown-api
- Tools in Data Science course notes (s-anand.net), "Converting PDFs to Markdown" (2025). https://tds.s-anand.net/2025-05/convert-pdfs-to-markdown/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- arXiv (MLCommons), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.