Enterprise document archive datasets for AI training
An enterprise document dataset is an export of the files a company actually works from, such as reports, memos, proposals, policies, PDFs, spreadsheets and decks, kept with folder structure, version history, authorship roles and sharing metadata. SourceX sources these archives, which often hold 3–15+ years of files, from established businesses' Google Drive, SharePoint, OneDrive, Box or Dropbox. Third-party and confidential material is screened out, personal data is de-identified, and the license can be scoped by department, file type and date range.
Dataset manifest
Sourced to your spec- What it is
- A company's working files with folders, versions and sharing metadata
- Typical systems
- SharePoint, Google Drive, OneDrive, Box, Dropbox, network file shares
- Typical history
- Company archives often hold 3–15+ years of files; varies by partner
- Modality
- Office files and PDFs with extracted text, structure and file metadata
- Delivery formats
- Agreed per order; native files plus extracted text and metadata in JSONL
- Preparation
- Personal data de-identified; third-party and privileged material screened out
- Licensing
- Scoped by department, file type and date range; permitted use agreed in writing
- Availability
- Only where a partner holds a matching archive and agrees to license it
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| file_id | string | Pseudonymous file identifier, reused wherever another file links to, embeds or duplicates this one. |
| path | string | Folder path with client, person and project names replaced by placeholders; folder names work as human-assigned labels. |
| file_type | object | Native format such as docx, Google Doc, pdf, xlsx, pptx or saved email, and whether a PDF is born-digital or scanned. |
| language | string | Detected language of the body text; mixed-language files list each language with its share. |
| timestamps | object | Created and modified times in UTC, taken from the storage platform rather than editable document properties. |
| doc_class | enum | Document type such as policy, proposal, report, minutes, contract or form, from the partner's taxonomy or an agreed classifier. |
| sections | array | Extracted headings with levels, paragraphs, lists and page numbers, in reading order. |
| tables | array | Tables extracted as cell grids, with header rows and merged cells marked. |
| annotations | array | Comments, replies and tracked changes with author roles and resolution status. |
| versions | array | Saved revisions of the document, each marking which sections changed and the role of the editor. |
| sharing | object | Permission scope at export time, such as private, team, company-wide or external link, with groups pseudonymized. |
| links | array | Hyperlinks, linked objects and embedded files that point to other files in the delivery. |
| dedup | object | Content hash and near-duplicate cluster, so copies saved in several folders can be collapsed or kept. |
| ocr | object | For scanned pages, OCR text with confidence scores and a reference to the page image. |
Example record
{
"file_id": "f_9b27e4",
"path": "/Operations/Distribution network/2022/[REGION] DC site selection.docx",
"file_type": { "format": "docx", "origin": "born_digital" },
"language": "en",
"timestamps": { "created": "2022-05-03T14:20:11Z", "modified": "2022-06-17T09:05:48Z" },
"doc_class": "decision_memo",
"sharing": { "scope": "department", "groups": ["grp_ops_leads"], "external_link": false },
"sections": [
{ "level": 1, "heading": "Recommendation", "page": 1,
"text": "Lease the [SITE_B] facility from Q1 2023 and close [SITE_A] after peak season." },
{ "level": 1, "heading": "Options considered", "page": 2, "table_ref": "t1" },
{ "level": 1, "heading": "Risks and open questions", "page": 4,
"text": "Labor availability near [SITE_B] is unproven. [REDACTED] to confirm by July." }
],
"tables": [
{ "id": "t1", "header": ["option", "annual_cost", "lead_time_weeks", "avg_drive_hrs"],
"rows": [["[SITE_A] expand", 1840000, 10, 3.1], ["[SITE_B] lease", 2120000, 16, 2.4]] }
],
"annotations": [
{ "type": "comment", "anchor": "Options considered", "author_role": "finance_lead",
"text": "The [SITE_B] cost excludes fit-out. Add it before this goes to the exec team.",
"resolved": true }
],
"versions": [
{ "v": 1, "t": "2022-05-03T14:20:11Z", "author_role": "ops_analyst", "words": 1140 },
{ "v": 4, "t": "2022-06-01T16:42:09Z", "author_role": "ops_director", "words": 1720,
"changed": ["Recommendation", "Options considered"] },
{ "v": 6, "t": "2022-06-17T09:05:48Z", "author_role": "ops_director", "words": 1655,
"changed": ["Risks and open questions"] }
],
"links": [
{ "file_id": "f_3c81aa", "format": "xlsx", "relation": "linked_chart" },
{ "file_id": "f_70d5e2", "format": "pptx", "relation": "hyperlink" }
],
"dedup": { "sha256": "[HASH]", "near_dup_cluster": "nd_0412", "copies_elsewhere": 2 }
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Build realistic RAG and enterprise search benchmarks
Questions written against a real archive face what production retrieval faces — superseded drafts, near-duplicate copies, conflicting versions and files the asker should not see.
Train document understanding on messy layouts
Born-digital and scanned files with tables, footnotes, headers and forms cover layouts that curated public document sets tend to underrepresent.
Answer questions that span several files
Linked memos, workbooks and decks let a model work across files, for example tracing a figure in a memo back to the model that produced it.
Learn drafting and revision
Version chains pair early drafts with approved finals and show how documents are restructured, cut and corrected in review.
Classify and route documents
Folder paths and document classes act as human-assigned labels for classification, routing and records-management models.
Use-case guides: Document AI, enterprise search and RAG, Enterprise and computer-use agents, Legal AI, Sales agents
What makes this data valuable
Version depth
Several revisions per file show how drafts became the documents people approved and used.
Intact folders
Original paths keep the organization employees actually used to file and find their work.
Cross-file links
Linked charts, hyperlinks and embedded files tie memos to the workbooks and decks behind them.
Permission snapshot
Sharing scopes make it possible to test whether retrieval respects access boundaries.
Born-digital share
Native files keep structure that scanned pages lose; a known mix lets you plan OCR work.
Departmental breadth
Files from finance, operations, sales and other functions widen the range of document types and vocabulary.
Why the archive is the unit, not the file
An archive is worth more than its files taken one by one, because the relationships between files, such as drafts and finals, copies, links and permissions, are what enterprise search and document agents get wrong. Public document collections are mostly finished, published files: filings, papers, manuals, government forms. Each one stands alone, and each is the final version. A company drive is a different object. The same proposal exists as a draft, a revised draft and a signed-off final in three folders. A policy was rewritten twice and the old copies were never deleted. A board memo quotes a figure from a workbook two folders away, and many files are visible to only one team.
That mess is the point. An enterprise search system or document agent is deployed into one company's archive at a time, with all of its duplicates, stale versions, odd formats and access boundaries. A corpus assembled from clean files picked out of many sources removes exactly the conditions it will be judged on. Delivering the archive with its folder paths, version chains, cross-file links and sharing scopes keeps those conditions intact, so tests can be written against them with a known right answer.
Splitting an archive without leaking answers
Near-duplicates are the main hazard when one archive feeds both training and evaluation. The same memo saved in three folders, a proposal reused for the next client with the names swapped, a policy stored beside its four earlier versions: split by file and these land on both sides, so a held-out question can be answered from memory. Split by near-duplicate cluster and version chain instead, so every copy and revision of a document sits on one side, and keep linked files, such as a memo and the workbook it quotes, together.
Some files in a company archive were also published, such as annual reports, press releases, product manuals and public filings. Those may already be in a model's pretraining data, so flag them and keep them out of held-out sets. Templates and boilerplate cause a related problem in training: standard terms, cover pages, disclaimers and recurring report shells can dominate an archive while adding little, and clustering lets you downweight them on purpose rather than find them later in model outputs.
Both steps depend on metadata that should be fixed in scope before delivery: content hashes, near-duplicate clusters, version links and a flag for files that were published externally. Computing them after delivery is possible, but by then the partner's own knowledge of which documents went public is no longer at hand.
What to check before licensing
- Ask for the file-type and language mix measured on unique content, after duplicates, templates and auto-generated files are collapsed, not on raw file counts.
- Confirm that material the partner does not own was screened out — client-owned deliverables, documents received under NDA, purchased research, and manuals or standards saved to shared drives.
- Ask which areas were excluded, such as HR, legal, board and deal folders, and whether exclusion relied on folder paths, sensitivity labels or a content scan.
- Open sample files in their native format and check de-identification in headers and footers, comments, tracked changes, document properties, alt text and embedded objects.
- Confirm what the export method preserved. Version history, comments, sharing settings and links survive some export paths and are lost in others.
- Check the date distribution. Archives moved from an older system can be thin before the migration or carry migration dates in place of original timestamps.
- For scanned material, review OCR quality on the sample and confirm whether page images are delivered or only the extracted text.
- Agree whether question sets, labels and search indexes built from the archive count as derivatives, and which deletion obligations apply to them when the license ends.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
What file types does an enterprise document archive include?
Usually a mix of word-processing files, PDFs, spreadsheets, presentations, images and saved emails, with the balance set by the partner's industry and departments. A professional services firm's archive leans toward proposals and reports; a manufacturer's leans toward specifications, forms and supplier documents. Buyers typically assess file-type mix and language share when reviewing a candidate dataset, so give the mix you are after when you describe the dataset.
Do I get version history, or only final files?
Version history can be included when the source platform kept it and the export captures it. Google Docs, SharePoint and OneDrive keep revision histories, subject to the organization's versioning and retention settings, while downloaded or copied files usually carry only their latest state. Many archives also hold manual version chains, such as files saved as draft, v2 and final, which can be linked during preparation.
How is confidential and third-party content handled?
It is screened before delivery, and much of it is excluded. Documents a client owns, files received under NDA, purchased research and copyrighted manuals or standards are not the partner's to license, and folders such as HR, legal and board material are commonly left out. What remains is de-identified, and the exclusion rules are documented so you know what the corpus does not contain.
Can an archive be scoped to certain departments or document types?
Yes, and narrow scopes are common. An archive can be cut by department or folder, document class, file type, language and years covered. Scope also shapes how much review the partner must do before release, since a finance team's reports carry different risks from a sales team's proposals. A focused request, such as policies and reports in English and German, is easier to match than a request for everything.
Can I build a RAG or enterprise search benchmark from a licensed archive?
Yes, if the license permits evaluation use. A licensed archive supports test items that public corpora cannot, such as questions whose answer changed between two versions of a policy, questions that need two linked files, and permission tests where the right response is to withhold an answer the user may not see. Those items depend on version, link and sharing metadata, so keep all three in scope when you request the archive.
How many years of files do archives cover?
Company archives often hold 3–15+ years of files, and the range varies by partner. Older years tend to be thinner and less consistent, particularly before a move to cloud storage, when files were sometimes migrated without their version history. If your work depends on change over time, such as how a policy evolved, name the years that matter and check them against each candidate manifest.
Are scanned documents included?
They can be, usually as page images with OCR text alongside. Scanned material adds signed forms, older records and paper-first processes that born-digital files miss. Personal data is harder to remove from page images than from text, so some partners release scanned documents as text only, or exclude them. Say whether you need the images, the text or both.
Related datasets
- Real-world spreadsheets and financial models
Working Excel and Google Sheets files with formulas, links and version history
- Business presentation decks
Real slide decks with layouts, chart data, speaker notes and revisions
- SOPs, playbooks and internal knowledge bases
Written procedures with page history, ownership and links to execution records
- Approved workplace email and chat exports
Approved, de-identified exports of team email threads and chat channels
- Contract negotiation and redline histories
Contract version chains with tracked changes, comments, approvals and executed versions
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.