Data for document AI, enterprise search and RAG evaluation
Document AI, enterprise search and RAG systems are trained and evaluated on real company corpora: messy file collections in native formats with folder structure, version history and access metadata, plus questions with known answers and known source passages. SourceX sources document archives, spreadsheets, decks, SOPs and contract histories from established businesses, licensed with agreed handling of personal, confidential and third-party content, so you can build parsing training data and private retrieval benchmarks.
Dataset types to start with
- Enterprise document archives
Whole drive and site exports keep what a production corpus looks like: duplicates, drafts, scanned PDFs, mixed languages, folder hierarchy and version history, which is what retrieval and permission-aware search have to cope with.
- Real-world spreadsheets and financial models
Workbooks with formulas, cross-sheet references, merged headers and pivot tables test table parsing and numeric question answering, where flattening a sheet to text loses the structure that makes the answer correct.
- Business presentation decks
Decks put meaning in layout, charts and speaker notes rather than prose. Native files give layout-aware parsers and multimodal retrieval real targets, and revision history shows how content was reworked between drafts.
- SOPs, playbooks and internal knowledge bases
Versioned wikis and procedures are the classic enterprise RAG corpus, and their edit history lets you test whether retrieval returns the current page rather than an older or near-duplicate one.
- Contract negotiation and redline histories
Long agreements with tracked changes and defined terms stress clause-level retrieval, cross-references and version reasoning, such as answering from the executed version rather than a draft.
Why this data is hard to get
Public documents are the clean end of the distribution
Public corpora lean on published papers, filings, manuals and web pages that were edited for outside readers. Internal files are drafts, notes, half-filled templates and near-duplicates, which is where retrieval and parsing break.
Benchmark questions echo their passages
Many public QA sets were written by annotators who had the passage in front of them, so questions reuse its wording. Employees ask without knowing which file holds the answer, in their own vocabulary, and some questions have no answer in the corpus at all.
Confidential content and third-party rights are everywhere
Company archives mix the company's own work with client deliverables, vendor contracts, licensed reports and embedded images, each with its own owner. That content has to be excluded or cleared file by file.
Personal data hides in places scanners miss
Names and contact details sit in tracked-change authors, comment threads, speaker notes, spreadsheet cells, document properties and scanned signatures, not only in body text.
Choosing the corpus before the questions
Retrieval quality depends on the corpus as much as on the questions. A RAG benchmark built on one tidy folder of final documents flatters most systems; the same questions over a whole department's drive, drafts and duplicates included, separate good retrieval from bad. Ask for complete slices of an archive, such as one team's drive across several years, rather than hand-picked files. Keep near-duplicates and superseded versions in the evaluation corpus, since choosing the current, authoritative version is part of what is being tested.
Document AI training has the opposite need. For parsing and extraction, diversity of templates, scan quality and layouts matters more than completeness, so it can draw on a wider spread of sources with fewer files from each.
Writing questions with known answers
Answer keys for enterprise documents need people who understand the domain. A finance analyst can say which tab of a model holds the assumption behind a forecast; a general annotator often cannot. Record the supporting evidence at the finest level that makes sense: passage, table cell, slide or clause, with the file version. Mix lookups, multi-hop questions across documents, numeric questions over spreadsheets, version questions ("what changed between the draft and the final?") and questions with no answer in the corpus.
Request the evaluation slice in its own delivery, restricted to the people who run evals, so it is not swept into a training mix later. See private evaluation sets for ways to stop a benchmark leaking into training over time.
Formats and preparation to specify
Name the formats you need, whether native files or extracted text with layout coordinates, and what has to survive: tracked changes, comments, speaker notes, formulas or embedded images. De-identification changes documents, so agree how replacements appear, consistent pseudonyms or typed placeholders, because that affects whether extraction labels still line up. Supply is not guaranteed, since it rests on businesses with suitable files agreeing to license them; each candidate SourceX finds comes with a manifest and sample files to review.
What good data looks like
- Files arrive in native formats (DOCX, XLSX, PPTX, PDF, scans) with folder paths, timestamps and version history intact.
- The manifest reports the file-type mix, language share, duplicate rate and the share of scanned versus born-digital files.
- Access groups are kept as pseudonymous labels, so permission-aware retrieval can be tested.
- Evaluation questions cite the passage, cell or slide that supports each answer, and some have no answer in the corpus.
- De-identification covers document properties, comments, tracked changes, speaker notes, cell contents and images, with a documented method.
- Third-party material, such as client deliverables, licensed reports and vendor documents, is excluded or cleared.
Questions buyers ask
How do I build a RAG evaluation set from licensed documents?
Choose a slice of the corpus to serve as the retrieval index and keep it, and everything from the same projects, out of all fine-tuning data. Then write or collect questions grounded in specific passages, cells or slides, recording the source file and version for each. Include multi-document questions and some with no answer, so abstention is tested. Domain experts should write or check the questions, and the license must permit that derived work and state who owns the eval set.
Can I get the questions people actually asked, not only documents?
Sometimes. Search logs, help desk tickets and chat threads that point to internal documents contain real questions. When the business agrees to license them, they help build evaluation sets with realistic phrasing. They are licensed and de-identified separately from the documents, because the questions themselves can contain personal or confidential details.
Can I fine-tune embedding or reranking models on licensed documents?
Yes, if the permitted use names it. Question-to-passage pairs make positives, and an archive's near-duplicates and superseded versions make the hard negatives that public pairs rarely supply. Train and evaluate on different departments, projects or periods, because documents from one project share wording and figures that let a retriever match on surface overlap.
Can document permissions be preserved for access-control testing?
Yes, in pseudonymous form. Sharing metadata can be delivered as group labels mapped to files, without real names or email addresses, so you can test whether retrieval respects who may see what. Some partners prefer to remove sharing metadata entirely, so if permission-aware search is part of your evaluation, say so in the request.
Can a corpus include scanned and handwritten documents?
It can. Older archives commonly include scanned contracts, signed forms, faxes and annotated printouts, but the share varies by partner and period, so list the formats and scanned share you need, then confirm both on the sample. Signatures, handwriting and stamps can identify people and are reviewed during de-identification.
Tell us what you are building
Describe the model or agent, the tasks it must handle, and the volume, format and permitted use you need. SourceX will match it to partner data.
Updated 3 October 2026.