Skip to content

Document AI data

Licensed PDF Corpora for Multimodal Pre-Training: Scale, Rights and Composition

Quick answer

A licensed PDF corpus for VLM pre-training is a collection of non-public business documents, usually page images plus extracted text and layout, licensed under a written grant that names pre-training use. It adds what web-crawled PDFs underrepresent: scanned forms, internal reports, statements, contracts and engineering drawings. Judge any offer on four things: a composition report, the born-digital to scanned ratio, deduplication against corpora you already hold, and rights terms that survive model release.

By SourceX Editorial · Updated

Why web-crawled PDFs leave gaps in document pre-training

Web PDFs skew toward documents that someone chose to publish, so the hard cases a document model meets in production are thin. Open work shows how labs build PDF data today: the Idefics3 authors released Docmatix, a PDF-derived document understanding set about 240 times larger than earlier open equivalents, and used OCR-IDL PDFs as pre-training examples [1]. NVIDIA's NeMo Data Designer note describes filtering web-crawled PDFs down to roughly 8.2 million page images and generating about 11.4 million synthetic VQA pairs on top of them [2].

Both pipelines start from public documents. What they rarely contain is the operational long tail: fax-quality scans with stamps, multi-generation photocopies, filled and signed forms, handwritten margin notes, internal spreadsheets printed to PDF, and long packets that mix twenty document types in one file. Those are the pages where OCR-free models fail on reading order, table structure and key-value grounding. For task-specific labeled data on those failure modes, see OCR ground truth data and multi-page and long document data.

Licensed business PDFs are also where rights questions get sharper. Public scrapes carry inherited uncertainty; the Data Provenance Initiative found license omission above 70% and license error rates above 50% on popular dataset hosting sites [5]. A licensed corpus should replace that uncertainty with a named licensor, a written grant and per-document provenance.

What a composition report should contain

A composition report is the single most useful document a supplier can give you, because it lets you size mixture weights before you see a page. Ask for it at the collection level and as per-document metadata you can query. Without it, you cannot tell whether "2 million pages" is 2 million distinct invoices or 40,000 copies of one template.

Request at least these dimensions:

  • Document type taxonomy with counts per class (invoice, bank statement, contract, SOP, engineering change order, inspection report, slide deck export). The taxonomy should be the supplier's real one; mapping work is covered in document classification training data.
  • Production mode: born-digital, scanned, photographed, or hybrid (a scan with an OCR text layer added later).
  • Page statistics: pages per document distribution, page size, DPI for raster pages, color mode.
  • Language and script per page, not per document; bilingual forms are common.
  • Date range of creation, by year, so you can avoid a corpus dominated by one system migration.
  • Source system and industry: document management system, ERP export, shared drive, mailroom scanner.
  • Templates: number of distinct layouts and the share of pages from the top ten templates.
  • Redaction footprint: share of pages with redaction boxes or replaced values, since that changes what the model learns.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters for pre-training
doc_idd_00041932Stable key for dedup, takedown and audit
sha256_file9f2c...e1Exact-duplicate detection across deliveries
doc_typevendor_invoiceMixture weighting by class
production_modescanned_with_ocr_layerText layer may be wrong; decide whether to trust it
page_count7Long-context sampling
dpi / color200 / grayscaleResolution floor for render pipeline
languages["en", "es"]Per-page language routing
created_year2014Temporal balance
template_clustertpl_0187Caps on near-identical layouts
pii_treatmentnames, account numbers replaced; method v2What was altered, and how
license_refschedule_A_batch_3Ties each file to its grant

Machine-readable dataset documentation helps here. The MLCommons Croissant-RAI vocabulary extends Croissant with responsible-AI fields so provenance and usage metadata travel with the data rather than in a separate PDF [10].

Balancing born-digital and scanned PDFs

The right mix depends on what the model must read, but most labs want scanned pages overrepresented relative to the web, because the web already supplies born-digital text. Born-digital PDFs give you a reliable text layer, font information and vector graphics, which makes them cheap supervision for page-to-markdown and layout targets; see PDF parsing training data. Scanned pages teach robustness to skew, noise, bleed-through, stamps, signatures and handwriting.

Three failure modes show up when the mix is wrong. A corpus that is mostly born-digital produces models that read clean exports well and collapse on mailroom scans. A corpus that is mostly scans with vendor OCR layers can teach the model to reproduce OCR errors if you use the embedded text as a target. Hybrid files, where an OCR layer was added years later, are often mislabeled as born-digital, so ask how production mode was detected (producer string, presence of image-only pages, font embedding) rather than accepting a label.

Rendering choices matter as much as the mix. Fix the DPI and color profile you render at, keep the original PDF alongside the PNG or JPEG render, and record the renderer version. The NVIDIA pipeline stores per-page PNG renders referenced from parquet [2], a pattern that keeps the page image, extracted text and metadata joinable by page key.

Deduplicating against the corpora you already hold

Deduplication against your existing web PDF holdings is a condition of value, not a cleanup step, because a licensed corpus that overlaps public crawls adds cost without coverage. Lee et al. showed that common training sets contain many near-duplicate examples and long repeated substrings, and that removing them reduces memorized verbatim output [4]. Business PDFs add their own duplication patterns: the same contract template across thousands of counterparties, monthly statements that differ by a few numbers, and the same file stored in five shared-drive folders.

Run dedup at three levels:

  1. Exact file: SHA-256 of the raw PDF bytes, then of the normalized page render, so re-saved files with new metadata still match.
  2. Near-duplicate text: MinHash with locality-sensitive hashing over extracted text shingles, per document and per page.
  3. Near-duplicate layout: perceptual hashes of page renders, or embedding similarity from a document encoder, to catch template families whose text differs.

Decide in advance what you keep. For templates, a cap per template cluster usually beats removal, because the model still needs to see a form family more than once with different filled values. Ask the supplier to run exact-hash matching against a hash list you provide, so you can measure overlap with public sources such as OCR-IDL-derived data [1] before you sign, without either party sharing documents.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckMethodAccept if
Overlap with lab web PDF corpusBuyer-supplied SHA-256 list, supplier-side matchOverlap reported per doc_type, below agreed threshold
Within-corpus exact duplicatesFile and render hashesRemoved before delivery, count reported
Template concentrationLayout clusteringNo cluster above agreed share of pages
Text near-duplicatesMinHash LSH on page textPairs above similarity cut reported with doc_ids

Rights terms a pre-training license must cover

A pre-training license must name the use explicitly, because a generic "AI" or "analytics" grant can leave the status of trained weights unclear. The detailed grant structure is covered in licensing data for foundation-model pre-training; the PDF-specific points are below.

  • Who holds the rights to each document. Business PDFs mix the supplier's own material with third-party content: customer-submitted forms, vendor invoices, embedded copyrighted images, and documents the supplier received under NDA. Ask how third-party documents were identified and excluded or cleared.
  • Use scope. Pre-training, continued pre-training and fine-tuning are different uses; name each one you need, along with internal evaluation.
  • Retention after term. State whether trained weights, checkpoints and derived artifacts (tokenized shards, embeddings, synthetic QA generated from the pages) may be kept after the license ends.
  • Disclosure. EU AI Act Article 53 duties for general-purpose AI providers, including the copyright policy in 53(1)(c), apply [6]. The Commission's 24 July 2025 template for the public summary of training content requires describing data sources [7], and the Code of Practice copyright chapter expects a maintained copyright policy [8]. California AB 2013 requires developers to post training data documentation for generative AI made available to Californians, due by 1 January 2026 [9]. Confirm with the supplier what you may say publicly about the source.

Personal data is the other rights axis. Business PDFs carry names, signatures, account numbers and addresses in headers, footers and stamps, not only in body text. Screening at scale needs automated detection on both the text layer and the page image, because a scanned page has no reliable text to scan, followed by sampled manual review. Under GDPR Recital 26, only data that is anonymous given the means reasonably likely to be used falls outside data protection rules [11]; replaced names on a page whose letterhead still identifies a small business may not meet that bar.

Formats and delivery for page-image corpora

Deliver PDFs and renders together, in a sharded format your loaders already read. A common layout is the original PDF, one image per page at a fixed DPI, a page-level parquet or JSONL table with text, bounding boxes and metadata, and a manifest keyed by doc_id and page number. For interleaved image-text pre-training, keep reading order and page boundaries so you can build sequences that alternate page images and text without re-parsing. Streaming formats and transfer costs are covered in streaming training data from object storage and dataset delivery egress costs.

If you plan layout supervision, align class labels to an existing scheme. DocLayNet's 11 classes in COCO format across 80,863 manually annotated pages [3] are a practical reference; see also document layout analysis datasets.

How SourceX approaches business PDF sourcing

SourceX sources operational datasets from US companies, including documents and finance and legal workflows, on request rather than from stock; a request does not guarantee a match. Buyers describe the documents they need, SourceX looks for US businesses that hold them, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

SourceX does not source scraped web content, so it is not a route to more crawled PDFs. For the archive view, see enterprise document datasets, the short answer to do AI labs buy PDFs, training data for multimodal document models and proprietary data for AI labs. To start a request, describe your target corpus on the buyer page. For the wider cluster, see the document AI data hub, licensed text corpora for LLM pre-training and sourcing for foundation-model pre-training teams.

Source a licensed PDF corpus for multimodal pre-training

SourceX looks for US companies that hold the business documents you describe, runs assessment of the data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Nothing is contracted until a supplier agrees, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the PDF corpus your pre-training run needs.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Frequently asked questions

Should I use the PDF's embedded text layer as a training target?

Only for files you have verified as born-digital. Scanned files with a later OCR layer carry that engine's errors, so use them as image inputs and generate targets with your own pipeline or a verified transcription set.

How do I measure overlap before signing?

Send the supplier a list of SHA-256 hashes for your existing PDF holdings and ask for a match count by document type. This tests exact overlap without exchanging documents; near-duplicate overlap needs a sample under NDA.

Do redactions hurt pre-training value?

They change it. Black boxes and replaced values teach the model that those regions exist, so ask for the redaction method and the share of affected pages, and prefer consistent replacement over mixed styles.

Sources

  1. arXiv (Laurencon et al., Hugging Face), "Building and better understanding vision-language models: insights and future directions" (2024). https://arxiv.org/pdf/2408.12637
  2. NVIDIA NeMo Data Designer documentation, "Training a VLM to Understand Long Documents: An Iterative SDG Story". https://docs.nvidia.com/nemo/datadesigner/latest/dev-notes/vlm-long-document-understanding
  3. arXiv (Pfitzmann et al., IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  4. arXiv (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/pdf/2107.06499
  5. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  6. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  7. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  8. European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
  9. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  10. arXiv (Jain et al., MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  11. European Parliament and Council of the European Union (Official Journal of the EU, via EUR-Lex), "Regulation (EU) 2016/679 (General Data Protection Regulation)" (2016). https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data