Retrieval, RAG and grounding data
Metadata a licensed RAG corpus should ship with
Quick answer
A licensed RAG corpus should ship with stable document and chunk identifiers, version and supersession links, created, modified and effective timestamps, the originating system and canonical locator, language and content type, a per-document license reference with permitted uses and expiry, and access-control entries for pseudonymized users and groups. Without these fields you cannot filter by date or source, cite precisely, honor license scope, enforce permissions or reproduce an evaluation run. Write them into the data spec before the supplier exports anything.
By SourceX Editorial · Updated
Why text alone is not enough for retrieval
Metadata is what lets a retriever answer questions the text cannot answer on its own. Enterprise benchmarks now include questions that depend on dates, authors and sources rather than passage content alone [1]. Consider a question such as "what did the latest revision of the returns policy say": a dense retriever cannot resolve "latest" unless every document carries a reliable modified or effective date and a link to the version it replaces.
The same logic applies to citation and audit. Research on enterprise RAG evaluation treats evidence traceability, the chain from answer to chunk to source document, as a core requirement [2], and test corpora that ship ground-truth citations only work because each document has a stable identifier [3]. If a supplier regenerates IDs on every export, every citation, qrel and cached embedding you built becomes orphaned.
For the cluster view of what to license and how, start at the retrieval, RAG and grounding data hub. This page covers only the metadata contract.
Document-level fields to require
Every document record needs identity, time, origin, rights and access fields, and each one maps to a concrete retrieval or compliance function. The table below is a spec you can paste into a supplier request and adapt.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Type | Purpose in your pipeline | Common failure if missing |
|---|---|---|---|
doc_id | string, immutable | Join key for chunks, embeddings, qrels, citations | IDs regenerate per export; caches and eval labels break |
version_id / version_seq | string / integer | Pin snapshots, dedupe revisions | Two revisions indexed as independent truths |
supersedes_doc_id | string or null | Rank current over stale | Retriever surfaces a withdrawn policy |
created_at, modified_at | ISO 8601 with offset | Recency filters and decay | Export date stamped as modified date |
effective_from, effective_to | ISO 8601 date | "As of" questions for policies, price lists, contracts | Valid-time confused with edit-time |
source_system | enum (e.g. confluence, sharepoint, zendesk, fileshare) | Source filters, per-system quality weighting | Mixed provenance cannot be separated |
source_locator | string (URL or path, pseudonymized if internal) | Human-checkable citation | Citation points to nothing reviewable |
title, section_path root | string | Display and heading-aware chunking | Chunker invents headings |
language | BCP 47 tag | Language routing and analyzer choice | Multilingual docs embedded with wrong model |
mime_type, original_format | string | Parser selection | DOCX tables flattened as prose |
content_sha256 | hex string | Exact-duplicate detection, tamper check | Silent re-delivery of changed content |
license_id | string | Ties document to the license schedule | Cannot prove which terms cover a passage |
permitted_uses | enum array (e.g. retrieval_internal, display_snippet, eval) | Enforce scope at query time | Content shown in a channel the license excludes |
license_expires_at | ISO 8601 date | Scheduled purge from index and caches | Expired content stays retrievable |
acl | array of entries (see below) | Permission-aware retrieval | All users see all documents |
pii_treatment | enum plus method note | Know which fields were masked or replaced | Re-identification risk unexamined |
Keep time semantics explicit. A modified timestamp tells you when a file changed; an effective window tells you when its content was true. Contracts, rate cards and HR policies need both, and suppliers often only have the first.
Rights fields belong on the document, not only in a contract PDF. When license_id, permitted_uses and license_expires_at travel with each record, your ingestion job can drop or quarantine records automatically. Clause-level negotiation is covered in RAG content license terms, and retention mechanics in caching and retention limits.
Chunk-level fields and citation anchors
Chunks need their own stable IDs and a pointer back to an exact location in a specific document version. If the supplier pre-chunks, require the fields below; if you chunk yourself, require the structural markers that let you generate them.
chunk_id: deterministic, for example a hash ofdoc_id+version_id+ start offset, so re-chunking the same version reproduces the same IDs.parent_doc_idandparent_version_id: never just the document ID, or a citation silently drifts to a newer revision.section_path: an ordered heading list such as["Warranty", "Exclusions", "Water damage"], used for heading-prefixed embeddings and breadcrumb citations.char_start,char_endagainst a declared canonical text rendition, pluspage_numberand bounding boxes for PDFs and scans.chunk_type:body,table,list,code,captionorform_field, so table chunks can go to a table-aware retriever.prev_chunk_id,next_chunk_id: allows context expansion around a hit without re-reading the whole document.
Offsets only mean something against a fixed text rendition. Ask the supplier to deliver that rendition (normalized UTF-8 text or structured JSON) alongside the original file, and to state the extraction tool and version. Format choices that survive chunking are covered in delivery formats for RAG.
Access-control metadata for permission-aware retrieval
ACL metadata should mirror how enterprise search connectors already model permissions: one entry per principal with an explicit allow or deny decision. A workable shape is an acl array on each document whose entries carry access (grant or deny), principal_type (user, group, everyone) and principal_id, with a stated rule that deny overrides grant. This is close to the item ACL shape used by Microsoft Graph connectors and maps readily to document-level security filters in search engines such as Elasticsearch, so the corpus stays loadable into common search stacks and testable for leakage.
Require four things. First, pseudonymized principal IDs that are consistent across documents, so one user maps to the same token everywhere. Second, a separate group membership table rather than groups expanded into every document ACL, which multiplies record updates whenever membership changes and hides the group structure you need for testing. Third, inheritance flags showing whether an ACL was inherited from a folder or space or set directly. Fourth, a sharing-change log with timestamps, so you can test what a user should have seen at a given date.
Designing the actual leakage tests (which user, which query, what counts as a leak) belongs to permission-aware RAG evaluation. Pseudonymization choices that keep retrieval intact are covered in de-identifying a RAG corpus.
Enrichment metadata: useful, but label it as derived
Enrichment fields such as entities, topics and event clusters improve filtering, but they must be marked as derived, with the method that produced them. News and grounding feeds already ship entities, topics, sentiment and event grouping as part of the record [4]. In operational corpora, the useful equivalents are product SKU, customer segment, ticket category, jurisdiction and document type.
The risk is treating a supplier's classifier output as ground truth. Require a derived_fields block listing each field, its producer (human, rule, or model:<name>@<version>) and a confidence where one exists. Then you can decide whether to filter on it or merely boost on it, and you can rerun your own enrichment without colliding with theirs.
Packaging, validation and manifests
The metadata contract should be machine-checkable before ingestion, not discovered during it. Specify the record schema as JSON Schema Draft 2020-12, the current release as of October 2026, which separates Core from Validation keywords [7], and reject deliveries that fail required, type or enum checks. Deliver records as JSON Lines or Parquet, one document record per line or row, with chunks in a sibling file keyed by parent_doc_id.
At the dataset level, a Croissant JSON-LD descriptor gives tools a standard way to read file resources and record structure [5], and the Croissant-RAI extension adds machine-readable provenance and responsible-AI fields [6]. Pair that with a datasheet-style narrative covering motivation, composition and collection process [8]. Neither replaces per-record fields: dataset cards describe the corpus, while the record metadata above drives retrieval. See Croissant metadata for licensed datasets, dataset cards for licensed enterprise data and the sample manifest.
A practical acceptance check before signing off a delivery:
- 100% of records validate against the agreed schema.
- Each
doc_id+version_idpair is unique, anddoc_idis unchanged against the previous delivery for every document that persists. - Every
supersedes_doc_idresolves to a record in the corpus or an explicit withdrawn list. modified_atvalues are not clustered at the export timestamp (a sign that original dates were lost).- Every principal in an
aclentry exists in the group or user table. - Every
license_idmatches a schedule in the executed license. - A random sample of chunk offsets reproduces the cited text exactly.
Metadata that supports reproducible evaluation
Evaluation needs the same metadata plus a pinned snapshot identifier. Qrels, nugget labels and citation ground truth reference doc_id and chunk_id, so a snapshot ID and a per-delivery changelog (added, modified, withdrawn) let you rerun last quarter's test against last quarter's corpus. Metadata-filtered evaluation, for example "only documents effective before 2025", also depends on the time fields above being trustworthy [1].
Snapshot practices are covered in pinned corpus snapshots for RAG evaluation, and stale-version and duplicate problems in knowledge-corpus quality for RAG.
What suppliers can and cannot usually provide
Most operational systems can export IDs, timestamps, authors, paths and permissions, but effective dates and clean version chains often need reconstruction. Ticketing tools such as Zendesk or Jira expose created, updated and status history; wikis such as Confluence keep page versions; file shares often keep only file-system timestamps, which copy and migration operations can reset. Ask the supplier which fields are native, which are reconstructed and how.
When you source through SourceX, you describe the data you need and SourceX looks for US businesses that hold it; datasets are sourced on request, not held in stock, so a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. You can put your metadata spec in a request through the buyer intake. For document categories, see enterprise document datasets and how licensed data is delivered.
Specify the metadata for your RAG corpus
SourceX sources operational datasets such as support histories, engineering records and documents from US companies, and every release is approved by the supplying company. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the corpus and the metadata fields you need at sourcex.si/buyers.
Sources
- Onyx, "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
- arXiv, "Enterprise RAG benchmark paper emphasizing evidence traceability (arXiv:2604.02640)" (2026). https://arxiv.org/pdf/2604.02640
- arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation" (2026). https://arxiv.org/pdf/2604.04936
- NewsAPI.ai, "Powering Chatbots with Real-Time News Data". https://newsapi.ai/blog/powering-chatbots-with-real-time-news-data/
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- arXiv (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- JSON Schema, "JSON Schema Specification". https://json-schema.org/specification
- arXiv (Gebru et al.), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.