Skip to content

Retrieval, RAG and grounding data

Metadata a licensed RAG corpus should ship with

Quick answer

A licensed RAG corpus should ship with stable document and chunk identifiers, version and supersession links, created, modified and effective timestamps, the originating system and canonical locator, language and content type, a per-document license reference with permitted uses and expiry, and access-control entries for pseudonymized users and groups. Without these fields you cannot filter by date or source, cite precisely, honor license scope, enforce permissions or reproduce an evaluation run. Write them into the data spec before the supplier exports anything.

By SourceX Editorial · Updated

Why text alone is not enough for retrieval

Metadata is what lets a retriever answer questions the text cannot answer on its own. Enterprise benchmarks now include questions that depend on dates, authors and sources rather than passage content alone [1]. Consider a question such as "what did the latest revision of the returns policy say": a dense retriever cannot resolve "latest" unless every document carries a reliable modified or effective date and a link to the version it replaces.

The same logic applies to citation and audit. Research on enterprise RAG evaluation treats evidence traceability, the chain from answer to chunk to source document, as a core requirement [2], and test corpora that ship ground-truth citations only work because each document has a stable identifier [3]. If a supplier regenerates IDs on every export, every citation, qrel and cached embedding you built becomes orphaned.

For the cluster view of what to license and how, start at the retrieval, RAG and grounding data hub. This page covers only the metadata contract.

Document-level fields to require

Every document record needs identity, time, origin, rights and access fields, and each one maps to a concrete retrieval or compliance function. The table below is a spec you can paste into a supplier request and adapt.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldTypePurpose in your pipelineCommon failure if missing
doc_idstring, immutableJoin key for chunks, embeddings, qrels, citationsIDs regenerate per export; caches and eval labels break
version_id / version_seqstring / integerPin snapshots, dedupe revisionsTwo revisions indexed as independent truths
supersedes_doc_idstring or nullRank current over staleRetriever surfaces a withdrawn policy
created_at, modified_atISO 8601 with offsetRecency filters and decayExport date stamped as modified date
effective_from, effective_toISO 8601 date"As of" questions for policies, price lists, contractsValid-time confused with edit-time
source_systemenum (e.g. confluence, sharepoint, zendesk, fileshare)Source filters, per-system quality weightingMixed provenance cannot be separated
source_locatorstring (URL or path, pseudonymized if internal)Human-checkable citationCitation points to nothing reviewable
title, section_path rootstringDisplay and heading-aware chunkingChunker invents headings
languageBCP 47 tagLanguage routing and analyzer choiceMultilingual docs embedded with wrong model
mime_type, original_formatstringParser selectionDOCX tables flattened as prose
content_sha256hex stringExact-duplicate detection, tamper checkSilent re-delivery of changed content
license_idstringTies document to the license scheduleCannot prove which terms cover a passage
permitted_usesenum array (e.g. retrieval_internal, display_snippet, eval)Enforce scope at query timeContent shown in a channel the license excludes
license_expires_atISO 8601 dateScheduled purge from index and cachesExpired content stays retrievable
aclarray of entries (see below)Permission-aware retrievalAll users see all documents
pii_treatmentenum plus method noteKnow which fields were masked or replacedRe-identification risk unexamined

Keep time semantics explicit. A modified timestamp tells you when a file changed; an effective window tells you when its content was true. Contracts, rate cards and HR policies need both, and suppliers often only have the first.

Rights fields belong on the document, not only in a contract PDF. When license_id, permitted_uses and license_expires_at travel with each record, your ingestion job can drop or quarantine records automatically. Clause-level negotiation is covered in RAG content license terms, and retention mechanics in caching and retention limits.

Chunk-level fields and citation anchors

Chunks need their own stable IDs and a pointer back to an exact location in a specific document version. If the supplier pre-chunks, require the fields below; if you chunk yourself, require the structural markers that let you generate them.

  • chunk_id: deterministic, for example a hash of doc_id + version_id + start offset, so re-chunking the same version reproduces the same IDs.
  • parent_doc_id and parent_version_id: never just the document ID, or a citation silently drifts to a newer revision.
  • section_path: an ordered heading list such as ["Warranty", "Exclusions", "Water damage"], used for heading-prefixed embeddings and breadcrumb citations.
  • char_start, char_end against a declared canonical text rendition, plus page_number and bounding boxes for PDFs and scans.
  • chunk_type: body, table, list, code, caption or form_field, so table chunks can go to a table-aware retriever.
  • prev_chunk_id, next_chunk_id: allows context expansion around a hit without re-reading the whole document.

Offsets only mean something against a fixed text rendition. Ask the supplier to deliver that rendition (normalized UTF-8 text or structured JSON) alongside the original file, and to state the extraction tool and version. Format choices that survive chunking are covered in delivery formats for RAG.

Access-control metadata for permission-aware retrieval

ACL metadata should mirror how enterprise search connectors already model permissions: one entry per principal with an explicit allow or deny decision. A workable shape is an acl array on each document whose entries carry access (grant or deny), principal_type (user, group, everyone) and principal_id, with a stated rule that deny overrides grant. This is close to the item ACL shape used by Microsoft Graph connectors and maps readily to document-level security filters in search engines such as Elasticsearch, so the corpus stays loadable into common search stacks and testable for leakage.

Require four things. First, pseudonymized principal IDs that are consistent across documents, so one user maps to the same token everywhere. Second, a separate group membership table rather than groups expanded into every document ACL, which multiplies record updates whenever membership changes and hides the group structure you need for testing. Third, inheritance flags showing whether an ACL was inherited from a folder or space or set directly. Fourth, a sharing-change log with timestamps, so you can test what a user should have seen at a given date.

Designing the actual leakage tests (which user, which query, what counts as a leak) belongs to permission-aware RAG evaluation. Pseudonymization choices that keep retrieval intact are covered in de-identifying a RAG corpus.

Enrichment metadata: useful, but label it as derived

Enrichment fields such as entities, topics and event clusters improve filtering, but they must be marked as derived, with the method that produced them. News and grounding feeds already ship entities, topics, sentiment and event grouping as part of the record [4]. In operational corpora, the useful equivalents are product SKU, customer segment, ticket category, jurisdiction and document type.

The risk is treating a supplier's classifier output as ground truth. Require a derived_fields block listing each field, its producer (human, rule, or model:<name>@<version>) and a confidence where one exists. Then you can decide whether to filter on it or merely boost on it, and you can rerun your own enrichment without colliding with theirs.

Packaging, validation and manifests

The metadata contract should be machine-checkable before ingestion, not discovered during it. Specify the record schema as JSON Schema Draft 2020-12, the current release as of October 2026, which separates Core from Validation keywords [7], and reject deliveries that fail required, type or enum checks. Deliver records as JSON Lines or Parquet, one document record per line or row, with chunks in a sibling file keyed by parent_doc_id.

At the dataset level, a Croissant JSON-LD descriptor gives tools a standard way to read file resources and record structure [5], and the Croissant-RAI extension adds machine-readable provenance and responsible-AI fields [6]. Pair that with a datasheet-style narrative covering motivation, composition and collection process [8]. Neither replaces per-record fields: dataset cards describe the corpus, while the record metadata above drives retrieval. See Croissant metadata for licensed datasets, dataset cards for licensed enterprise data and the sample manifest.

A practical acceptance check before signing off a delivery:

  1. 100% of records validate against the agreed schema.
  2. Each doc_id + version_id pair is unique, and doc_id is unchanged against the previous delivery for every document that persists.
  3. Every supersedes_doc_id resolves to a record in the corpus or an explicit withdrawn list.
  4. modified_at values are not clustered at the export timestamp (a sign that original dates were lost).
  5. Every principal in an acl entry exists in the group or user table.
  6. Every license_id matches a schedule in the executed license.
  7. A random sample of chunk offsets reproduces the cited text exactly.

Metadata that supports reproducible evaluation

Evaluation needs the same metadata plus a pinned snapshot identifier. Qrels, nugget labels and citation ground truth reference doc_id and chunk_id, so a snapshot ID and a per-delivery changelog (added, modified, withdrawn) let you rerun last quarter's test against last quarter's corpus. Metadata-filtered evaluation, for example "only documents effective before 2025", also depends on the time fields above being trustworthy [1].

Snapshot practices are covered in pinned corpus snapshots for RAG evaluation, and stale-version and duplicate problems in knowledge-corpus quality for RAG.

What suppliers can and cannot usually provide

Most operational systems can export IDs, timestamps, authors, paths and permissions, but effective dates and clean version chains often need reconstruction. Ticketing tools such as Zendesk or Jira expose created, updated and status history; wikis such as Confluence keep page versions; file shares often keep only file-system timestamps, which copy and migration operations can reset. Ask the supplier which fields are native, which are reconstructed and how.

When you source through SourceX, you describe the data you need and SourceX looks for US businesses that hold it; datasets are sourced on request, not held in stock, so a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. You can put your metadata spec in a request through the buyer intake. For document categories, see enterprise document datasets and how licensed data is delivered.

Specify the metadata for your RAG corpus

SourceX sources operational datasets such as support histories, engineering records and documents from US companies, and every release is approved by the supplying company. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the corpus and the metadata fields you need at sourcex.si/buyers.

Sources

  1. Onyx, "EnterpriseRAG-Bench". https://onyx.app/enterpriserag-bench
  2. arXiv, "Enterprise RAG benchmark paper emphasizing evidence traceability (arXiv:2604.02640)" (2026). https://arxiv.org/pdf/2604.02640
  3. arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation" (2026). https://arxiv.org/pdf/2604.04936
  4. NewsAPI.ai, "Powering Chatbots with Real-Time News Data". https://newsapi.ai/blog/powering-chatbots-with-real-time-news-data/
  5. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  6. arXiv (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  7. JSON Schema, "JSON Schema Specification". https://json-schema.org/specification
  8. arXiv (Gebru et al.), "Datasheets for Datasets" (2021). https://arxiv.org/pdf/1803.09010

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data