Skip to content

Enterprise document archive datasets for AI training

An enterprise document dataset is an export of the files a company actually works from, such as reports, memos, proposals, policies, PDFs, spreadsheets and decks, kept with folder structure, version history, authorship roles and sharing metadata. SourceX sources these archives, which often hold 3–15+ years of files, from established businesses' Google Drive, SharePoint, OneDrive, Box or Dropbox. Third-party and confidential material is screened out, personal data is de-identified, and the license can be scoped by department, file type and date range.

Dataset manifest

Sourced to your spec
What it is
A company's working files with folders, versions and sharing metadata
Typical systems
SharePoint, Google Drive, OneDrive, Box, Dropbox, network file shares
Typical history
Company archives often hold 3–15+ years of files; varies by partner
Modality
Office files and PDFs with extracted text, structure and file metadata
Delivery formats
Agreed per order; native files plus extracted text and metadata in JSONL
Preparation
Personal data de-identified; third-party and privileged material screened out
Licensing
Scoped by department, file type and date range; permitted use agreed in writing
Availability
Only where a partner holds a matching archive and agrees to license it

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
file_idstringPseudonymous file identifier, reused wherever another file links to, embeds or duplicates this one.
pathstringFolder path with client, person and project names replaced by placeholders; folder names work as human-assigned labels.
file_typeobjectNative format such as docx, Google Doc, pdf, xlsx, pptx or saved email, and whether a PDF is born-digital or scanned.
languagestringDetected language of the body text; mixed-language files list each language with its share.
timestampsobjectCreated and modified times in UTC, taken from the storage platform rather than editable document properties.
doc_classenumDocument type such as policy, proposal, report, minutes, contract or form, from the partner's taxonomy or an agreed classifier.
sectionsarrayExtracted headings with levels, paragraphs, lists and page numbers, in reading order.
tablesarrayTables extracted as cell grids, with header rows and merged cells marked.
annotationsarrayComments, replies and tracked changes with author roles and resolution status.
versionsarraySaved revisions of the document, each marking which sections changed and the role of the editor.
sharingobjectPermission scope at export time, such as private, team, company-wide or external link, with groups pseudonymized.
linksarrayHyperlinks, linked objects and embedded files that point to other files in the delivery.
dedupobjectContent hash and near-duplicate cluster, so copies saved in several folders can be collapsed or kept.
ocrobjectFor scanned pages, OCR text with confidence scores and a reference to the page image.

Example record

{
  "file_id": "f_9b27e4",
  "path": "/Operations/Distribution network/2022/[REGION] DC site selection.docx",
  "file_type": { "format": "docx", "origin": "born_digital" },
  "language": "en",
  "timestamps": { "created": "2022-05-03T14:20:11Z", "modified": "2022-06-17T09:05:48Z" },
  "doc_class": "decision_memo",
  "sharing": { "scope": "department", "groups": ["grp_ops_leads"], "external_link": false },
  "sections": [
    { "level": 1, "heading": "Recommendation", "page": 1,
      "text": "Lease the [SITE_B] facility from Q1 2023 and close [SITE_A] after peak season." },
    { "level": 1, "heading": "Options considered", "page": 2, "table_ref": "t1" },
    { "level": 1, "heading": "Risks and open questions", "page": 4,
      "text": "Labor availability near [SITE_B] is unproven. [REDACTED] to confirm by July." }
  ],
  "tables": [
    { "id": "t1", "header": ["option", "annual_cost", "lead_time_weeks", "avg_drive_hrs"],
      "rows": [["[SITE_A] expand", 1840000, 10, 3.1], ["[SITE_B] lease", 2120000, 16, 2.4]] }
  ],
  "annotations": [
    { "type": "comment", "anchor": "Options considered", "author_role": "finance_lead",
      "text": "The [SITE_B] cost excludes fit-out. Add it before this goes to the exec team.",
      "resolved": true }
  ],
  "versions": [
    { "v": 1, "t": "2022-05-03T14:20:11Z", "author_role": "ops_analyst", "words": 1140 },
    { "v": 4, "t": "2022-06-01T16:42:09Z", "author_role": "ops_director", "words": 1720,
      "changed": ["Recommendation", "Options considered"] },
    { "v": 6, "t": "2022-06-17T09:05:48Z", "author_role": "ops_director", "words": 1655,
      "changed": ["Risks and open questions"] }
  ],
  "links": [
    { "file_id": "f_3c81aa", "format": "xlsx", "relation": "linked_chart" },
    { "file_id": "f_70d5e2", "format": "pptx", "relation": "hyperlink" }
  ],
  "dedup": { "sha256": "[HASH]", "near_dup_cluster": "nd_0412", "copies_elsewhere": 2 }
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Build realistic RAG and enterprise search benchmarks

Questions written against a real archive face what production retrieval faces — superseded drafts, near-duplicate copies, conflicting versions and files the asker should not see.

Train document understanding on messy layouts

Born-digital and scanned files with tables, footnotes, headers and forms cover layouts that curated public document sets tend to underrepresent.

Answer questions that span several files

Linked memos, workbooks and decks let a model work across files, for example tracing a figure in a memo back to the model that produced it.

Learn drafting and revision

Version chains pair early drafts with approved finals and show how documents are restructured, cut and corrected in review.

Classify and route documents

Folder paths and document classes act as human-assigned labels for classification, routing and records-management models.

Use-case guides: Document AI, enterprise search and RAG, Enterprise and computer-use agents, Legal AI, Sales agents

What makes this data valuable

Version depth

Several revisions per file show how drafts became the documents people approved and used.

Intact folders

Original paths keep the organization employees actually used to file and find their work.

Cross-file links

Linked charts, hyperlinks and embedded files tie memos to the workbooks and decks behind them.

Permission snapshot

Sharing scopes make it possible to test whether retrieval respects access boundaries.

Born-digital share

Native files keep structure that scanned pages lose; a known mix lets you plan OCR work.

Departmental breadth

Files from finance, operations, sales and other functions widen the range of document types and vocabulary.

Why the archive is the unit, not the file

An archive is worth more than its files taken one by one, because the relationships between files, such as drafts and finals, copies, links and permissions, are what enterprise search and document agents get wrong. Public document collections are mostly finished, published files: filings, papers, manuals, government forms. Each one stands alone, and each is the final version. A company drive is a different object. The same proposal exists as a draft, a revised draft and a signed-off final in three folders. A policy was rewritten twice and the old copies were never deleted. A board memo quotes a figure from a workbook two folders away, and many files are visible to only one team.

That mess is the point. An enterprise search system or document agent is deployed into one company's archive at a time, with all of its duplicates, stale versions, odd formats and access boundaries. A corpus assembled from clean files picked out of many sources removes exactly the conditions it will be judged on. Delivering the archive with its folder paths, version chains, cross-file links and sharing scopes keeps those conditions intact, so tests can be written against them with a known right answer.

Splitting an archive without leaking answers

Near-duplicates are the main hazard when one archive feeds both training and evaluation. The same memo saved in three folders, a proposal reused for the next client with the names swapped, a policy stored beside its four earlier versions: split by file and these land on both sides, so a held-out question can be answered from memory. Split by near-duplicate cluster and version chain instead, so every copy and revision of a document sits on one side, and keep linked files, such as a memo and the workbook it quotes, together.

Some files in a company archive were also published, such as annual reports, press releases, product manuals and public filings. Those may already be in a model's pretraining data, so flag them and keep them out of held-out sets. Templates and boilerplate cause a related problem in training: standard terms, cover pages, disclaimers and recurring report shells can dominate an archive while adding little, and clustering lets you downweight them on purpose rather than find them later in model outputs.

Both steps depend on metadata that should be fixed in scope before delivery: content hashes, near-duplicate clusters, version links and a flag for files that were published externally. Computing them after delivery is possible, but by then the partner's own knowledge of which documents went public is no longer at hand.

What to check before licensing

  • Ask for the file-type and language mix measured on unique content, after duplicates, templates and auto-generated files are collapsed, not on raw file counts.
  • Confirm that material the partner does not own was screened out — client-owned deliverables, documents received under NDA, purchased research, and manuals or standards saved to shared drives.
  • Ask which areas were excluded, such as HR, legal, board and deal folders, and whether exclusion relied on folder paths, sensitivity labels or a content scan.
  • Open sample files in their native format and check de-identification in headers and footers, comments, tracked changes, document properties, alt text and embedded objects.
  • Confirm what the export method preserved. Version history, comments, sharing settings and links survive some export paths and are lost in others.
  • Check the date distribution. Archives moved from an older system can be thin before the migration or carry migration dates in place of original timestamps.
  • For scanned material, review OCR quality on the sample and confirm whether page images are delivered or only the extracted text.
  • Agree whether question sets, labels and search indexes built from the archive count as derivatives, and which deletion obligations apply to them when the license ends.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

What file types does an enterprise document archive include?

Usually a mix of word-processing files, PDFs, spreadsheets, presentations, images and saved emails, with the balance set by the partner's industry and departments. A professional services firm's archive leans toward proposals and reports; a manufacturer's leans toward specifications, forms and supplier documents. Buyers typically assess file-type mix and language share when reviewing a candidate dataset, so give the mix you are after when you describe the dataset.

Do I get version history, or only final files?

Version history can be included when the source platform kept it and the export captures it. Google Docs, SharePoint and OneDrive keep revision histories, subject to the organization's versioning and retention settings, while downloaded or copied files usually carry only their latest state. Many archives also hold manual version chains, such as files saved as draft, v2 and final, which can be linked during preparation.

How is confidential and third-party content handled?

It is screened before delivery, and much of it is excluded. Documents a client owns, files received under NDA, purchased research and copyrighted manuals or standards are not the partner's to license, and folders such as HR, legal and board material are commonly left out. What remains is de-identified, and the exclusion rules are documented so you know what the corpus does not contain.

Can an archive be scoped to certain departments or document types?

Yes, and narrow scopes are common. An archive can be cut by department or folder, document class, file type, language and years covered. Scope also shapes how much review the partner must do before release, since a finance team's reports carry different risks from a sales team's proposals. A focused request, such as policies and reports in English and German, is easier to match than a request for everything.

Can I build a RAG or enterprise search benchmark from a licensed archive?

Yes, if the license permits evaluation use. A licensed archive supports test items that public corpora cannot, such as questions whose answer changed between two versions of a policy, questions that need two linked files, and permission tests where the right response is to withhold an answer the user may not see. Those items depend on version, link and sharing metadata, so keep all three in scope when you request the archive.

How many years of files do archives cover?

Company archives often hold 3–15+ years of files, and the range varies by partner. Older years tend to be thinner and less consistent, particularly before a move to cloud storage, when files were sometimes migrated without their version history. If your work depends on change over time, such as how a policy evolved, name the years that matter and check them against each candidate manifest.

Are scanned documents included?

They can be, usually as page images with OCR text alongside. Scanned material adds signed forms, older records and paper-first processes that born-digital files miss. Personal data is harder to remove from page images than from text, so some partners release scanned documents as text only, or exclude them. Say whether you need the images, the text or both.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify