Document AI data
Public Document AI Datasets: Which Ones Allow Commercial Use?
Quick answer
Few of the classic document AI datasets are clean for commercial training. As of October 2026, FUNSD is reported as non-commercial, research and educational use only; RVL-CDIP and DocVQA have no standalone license we could locate and rest on tobacco-litigation document archives; DocLayNet's card names the permissive CDLA-Permissive-1.0. Even a permissive annotation license does not clear the underlying page images, so check both layers on the official source before a commercial run.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Two layers of rights: annotations versus the document images
A commercial clearance needs two answers: who may license the annotations, and who holds rights in the pages those annotations sit on. Dataset authors usually own their bounding boxes, entity labels, question-answer pairs and class assignments, so the license on a card typically governs that layer. The page images are often third-party documents: tobacco-industry memos, scientific articles, annual reports, patents or government filings. A license from the annotators cannot grant more rights in those pages than the annotators held.
A license-compliance study of publicly available datasets makes this point directly: the license string on a dataset is the start of the analysis, and commercial-use rights depend on tracing every upstream source the data was built from [8]. For document AI, that tracing ends in an archive, a publisher or a crawl, not at the GitHub repository.
Per-dataset status table for common document AI benchmarks
The table below is a starting point for your own review, not a clearance: each row cites what a source in this session reported, and every row must be re-read on the official dataset page before you rely on it. Last checked: 9 October 2026.
| Dataset | Task | Source documents | Stated license (as reported) | Commercial-use reading | Provenance caveat |
|---|---|---|---|---|---|
| RVL-CDIP | Page classification, 16 classes, 400,000 grayscale images [1] | Subset of IIT-CDIP, from tobacco-litigation documents [1][2] | No standalone license located; hub cards defer to the UCSF Industry Documents copyright terms | Not established by the card; governed by archive terms you must read | Litigation-disclosed documents with many original rights holders |
| IIT-CDIP | OCR and pre-training source collection | Tobacco-litigation document scans [2] | No standalone dataset license found in this review | Treat as unresolved | Parent of RVL-CDIP, FUNSD and many pre-training mixes |
| FUNSD | Form entity labeling and linking [2] | Noisy scanned forms sampled from RVL-CDIP [2] | Reported project-page terms limit use to non-commercial, research and educational purposes [3] | Not for commercial training or commercial evaluation without separate permission | Inherits the tobacco-archive lineage underneath the annotation license |
| DocVQA | Visual question answering on document images | UCSF Industry Documents Library pages [7] | Confirm on the challenge site; no license text was located in this review | Unresolved until terms are read | Same archive family as RVL-CDIP |
| DocLayNet | Layout analysis, 11 classes, human-annotated [4] | Pages from varied categories such as financial reports, manuals, patents and laws [4] | Repository LICENSE file is CDLA-Permissive-1.0 [5] | Annotation layer permissive; page rights still need review | Underlying pages come from many publishers |
| PubLayNet | Layout analysis, auto-generated labels [6] | PubMed Central articles matched XML to PDF [6] | Confirm the repository license directly | Annotation layer depends on repository terms | Article-level licenses in PubMed Central vary, so image rights vary page by page |
Two patterns stand out. Three of the most-cited sets (RVL-CDIP, FUNSD, DocVQA) sit on the same tobacco-litigation archive, so one unresolved upstream question propagates across classification, extraction and question answering benchmarks. And the two "permissive" layout sets pair an open annotation license with page images drawn from many third parties.
Why the license string on a hub is not enough
Hosting-site license fields are frequently wrong or missing, so verify at the original source. The Data Provenance Initiative audit of more than 1,800 text datasets found license omission above 70% and license error rates above 50% on popular hosting sites [10]. DocLayNet illustrates the need to check the primary source directly: its GitHub repository LICENSE file states CDLA-Permissive-1.0 [5], a term not reliably surfaced on every dataset-hub mirror.
Document-specific surveys point the same way. The BigDocs authors report surveying 133 document datasets and finding roughly 80% with non-permissive or unclear licenses, leaving 16 they judged fully accessible and permissive [9]. If your team is building a commercial mix from public sets, expect most candidates to fall out at the license step.
Mirrors add a further failure mode. Re-uploads of RVL-CDIP or DocLayNet under personal accounts may carry a different license tag, a converted format (Parquet, COCO JSON) and no link back to the original terms. Pin the canonical source and record its URL and commit hash in your data card.
What "permissive" covers under CDLA-Permissive-1.0
CDLA-Permissive is a data license from the Linux Foundation's Community Data License Agreement family, written for datasets rather than code [11]. Its FAQ explains how the agreements treat results produced by computational use of data, which is the clause most relevant to model training [11]. Read the version that applies: DocLayNet's repository names version 1.0, while the Linux Foundation has since published 2.0.
What the license does not do is warrant third-party rights in the content. When annotators license their labels under CDLA, a downstream user still has to decide whether training on the underlying annual report or patent page is permissible in the target jurisdiction. That analysis sits outside the license text.
Commercial evaluation usually counts as commercial use
Using a non-commercial benchmark to score a product model is still a commercial use under most readings of "non-commercial, research and educational purposes" terms. FUNSD's reported terms [3] do not carve out evaluation. Teams often clear training data carefully and then run release-gating evals on restricted benchmarks without review.
The practical answer is a held-out set you are licensed to use for product decisions. The comparison on private evaluation sets versus public benchmarks covers the tradeoffs, and the general method for reading eval terms is on the public benchmark license page.
Benchmark fitness, not just rights
Rights are not the only reason to look past the classic sets. Much published document AI work reports results on a small set of public benchmarks such as DocLayNet and FUNSD, and RVL-CDIP's single-page, 16-class framing does not reflect production document streams. RVL-CDIP pages are grayscale scans from decades-old corporate archives, while production intake mixes born-digital PDFs, phone photos and multi-page packets.
If your target is invoices, bank statements, bills of lading or loss runs, the public sets mostly measure transfer from tobacco-era memos and scientific layouts. See document classification training data and layout analysis datasets beyond DocLayNet and PubLayNet for task-specific gaps.
Regulatory documentation: what a GPAI provider records
Providers of general-purpose AI models placed on the EU market owe Article 53 duties. The Code of Practice copyright chapter asks signatories to maintain a copyright policy covering those models [12], and the Commission's 24 July 2025 template sets the baseline for a public summary of training content [13]. A per-dataset clearance record like the one below feeds both.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_clearance_record:
dataset: DocLayNet
canonical_source: <official repository URL>
version_or_commit: <hash or release tag>
last_checked: 2026-10-09
annotation_license: CDLA-Permissive-1.0 # as named in the repository LICENSE file
metadata_license_tag: <confirm per source> # not all mirrors surface a license tag
underlying_content:
origin: mixed third-party pages (reports, manuals, patents, regulations)
rights_basis: <counsel conclusion per jurisdiction>
intended_uses: [pretraining, fine_tuning, internal_eval]
release_gating_eval_allowed: <yes/no, with reason>
attribution_required: <text and placement>
pii_review: <method and sample size>
reviewer: <name, role>
decision: approved | approved_with_conditions | rejected
Clearance checklist for a public document dataset
Run every candidate through the same questions before it enters a training or eval mix:
- Locate the canonical source (author repository or paper page), not a mirror, and record URL plus version.
- Read the annotation license text in full, not the hub tag.
- Identify the origin of the page images: archive, publisher, crawl or synthetic generator.
- Find the rights terms for that origin, such as an archive copyright page or article-level licenses.
- Check whether the dataset is a derivative of another set (FUNSD from RVL-CDIP, RVL-CDIP from IIT-CDIP) and clear each parent.
- Decide separately for pre-training, fine-tuning, internal evaluation and release-gating evaluation.
- Scan pages for personal data: names, signatures, addresses and account numbers appear on real forms.
- Record the decision and a re-check date, because terms and cards change.
Pair this with SourceX's AI training data due diligence checklist for the broader vendor and rights questions, and with licensed vs synthetic vs scraped data when deciding what replaces a set that fails.
When public sets fail: options for commercial document models
If the clearance fails, the realistic routes are synthetic documents, licensed PDF corpora, or licensed operational documents from businesses that hold them. Synthetic generation avoids third-party page rights but has known distribution gaps; see where generated documents break. For broad multimodal pre-training, see licensed PDF corpora.
For extraction targets tied to real workflows, licensed business documents carry a rights basis you can document. SourceX sources operational datasets from US companies on request, including documents and finance and legal workflows; each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Buyers can describe the documents they need; a request does not guarantee a match. The document AI data hub maps tasks to annotation layers.
Licensed business documents for commercial document AI
SourceX finds US businesses that hold the documents you describe, and every release is approved by the supplying company. The process runs Find, Assess, Agree, Transact, Manage, and nothing is contracted until a supplier agrees. Describe your document data need to SourceX.
Frequently asked questions
Can I train a commercial model on FUNSD?
Not on the reported terms. The FUNSD project page, as reported, limits use to non-commercial, research and educational purposes [3], and the forms themselves come from the RVL-CDIP tobacco archive [2]. Commercial use would need separate permission and a view on the underlying pages.
Is RVL-CDIP licensed for commercial use?
No standalone RVL-CDIP license was located in this review; the set is a subset of IIT-CDIP, built from tobacco-litigation documents [1][2], and hub cards point to the UCSF Industry Documents copyright terms. Read those terms and take counsel's view before any commercial training.
Does DocLayNet's CDLA-Permissive-1.0 license cover the page images?
It covers what the licensor could license, which is clearly the annotations [5]. The pages come from many publishers and document types [4], so rights in the images still need a separate assessment.
Sources
- IndiaAI AIKosh, "RVL-CDIP dataset listing". https://aikosh.indiaai.gov.in/home/datasets/details/rvl_cdip.html
- Jaume, Ekenel and Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- Papers with Code, "FUNSD dataset page". https://ml.paperswithcode.com/dataset/funsd
- Pfitzmann et al., IBM Research, "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/pdf/2206.01062
- DS4SD / Deep Search (IBM Research), GitHub, "DocLayNet LICENSE". https://github.com/DS4SD/DocLayNet/blob/main/LICENSE
- Zhong, Tang and Jimeno Yepes, "PubLayNet: largest dataset ever for document layout analysis" (2019). https://arxiv.org/pdf/1908.07836
- Mathew, Karatzas and Jawahar, "DocVQA: A Dataset for VQA on Document Images" (2020). https://arxiv.org/pdf/2007.00398
- arXiv:2111.02374, "Can I use this publicly available dataset to build commercial AI software? A case study on publicly available image datasets" (2021). https://arxiv.org/pdf/2111.02374v4
- OpenReview, "BigDocs: An Open Dataset for Training Multimodal Models on Document and Code Tasks". https://openreview.net/pdf/129a5f01a9d23187409b611b416fa4e2c40f0720.pdf
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- The Linux Foundation (cdla.dev), "Community Data License Agreement FAQ". https://cdla.dev/faq-resources/faq/
- European Commission (AI Office), "General-Purpose AI Code of Practice: Contents of the Code (Copyright chapter)" (2025). https://digital-strategy.ec.europa.eu/policies/contents-code-gpai
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.