Skip to content

Provenance, rights and permitted use

Dataset Bills of Materials: SPDX and CycloneDX for Training Data

Quick answer

A dataset bill of materials is a machine-readable inventory of every dataset a model was trained, tuned or evaluated on, linked to the model component and to each dataset's supplier, version hash, license and handling constraints. As of October 2026, the two practical formats are SPDX 3.0, whose Dataset and AI profiles add a DatasetPackage class and trainedOn relationships, and CycloneDX, whose ML-BOM adds machine-learning-model and data components. Pick one per pipeline and populate it from your training data register.

By SourceX Editorial · Updated

What a dataset BOM must answer that a dataset card does not

A dataset BOM answers "exactly which bytes, under which license, went into this model build," which a prose dataset card cannot answer reliably. Dataset cards and Croissant files describe a dataset for people and loaders: purpose, biases, schema and record structure [5][6]. A BOM is a graph. It ties a pinned dataset version to a model artifact and makes the license and provenance fields queryable by tools such as policy engines and license scanners.

The questions a reviewer will put to the BOM are narrow. Which dataset versions (by hash) fed checkpoint X? What license governs each, and does any forbid commercial use or redistribution of derived weights? Which inputs carry personal data, and how were they de-identified? Which supplier can attest to each answer? If your BOM cannot answer those four questions without a human opening a PDF, it is an inventory, not a bill of materials.

Keep the two artifacts linked rather than merged. Our guide to dataset cards for licensed enterprise data covers the narrative layer; the BOM should reference the card by URL or hash.

How SPDX 3.0 represents a training dataset

SPDX 3.0 represents a dataset as a DatasetPackage element in the Dataset profile and links it to an AIPackage through typed relationships. SPDX 3.0 was released with the AI and Dataset profiles in April 2024, and the profiles were designed to stay compatible with the rest of the SPDX model and its tooling. Verify the current patch release (3.0.x) before you pin a schema.

Practitioner guides list datasetType, releaseTime, suppliedBy and the declared and concluded license relationships among the DatasetPackage's expected fields; confirm cardinality against the published model. Optional fields that matter most for licensed training data are:

  • dataCollectionProcess and dataPreprocessing: how records were gathered and transformed (deduplication, redaction, tokenization).
  • anonymizationMethodUsed: the de-identification method, recorded as text, for example "regex plus NER replacement of names, emails, phones."
  • confidentialityLevel: a traffic-light style marking for how the dataset itself may be shared.
  • intendedUse, knownBias, datasetSize, datasetUpdateMechanism and datasetAvailability.

On the model side, an AIPackage connects to datasets with relationships such as trainedOn and testedOn, so a single SPDX document can express "model M v2.1 trainedOn dataset D v2026-09-30". Use hasDeclaredLicense for what the supplier states and hasConcludedLicense for what your legal review concluded; keeping both exposes disagreement instead of hiding it.

How CycloneDX represents datasets in an ML-BOM

CycloneDX represents a dataset as a component of type "data" and a model as a component of type "machine-learning-model," with the model card's dataset references pointing at the data components. Machine-learning support arrived in CycloneDX 1.5 and is described in the OWASP CycloneDX ML-BOM guidance. That guidance covers model metadata, training datasets and lineage; check the current CycloneDX release and its Ecma standard edition before fixing a schema version.

Useful data-component fields include a data type (for example "dataset"), contents (a URL or attachment reference), classification, sensitiveData, and governance roles for owners, stewards and custodians. Licenses attach through the standard license block (an SPDX id, a name, or an expression), and CycloneDX also offers a licensing object for commercial terms such as licensor, licensee and expiration. Hashes on each component pin the exact version.

ML-BOM practice stresses model lineage (fine-tuning, quantization, adapters). For datasets, mirror that: record "derived from" when you filter or relabel a licensed corpus, so the derived component inherits the parent's license constraints.

SPDX or CycloneDX: a decision table

Choose the format your security tooling already consumes, then map the dataset fields you need; neither format is missing anything essential for a training-data BOM.

NeedSPDX 3.0 Dataset/AI profilesCycloneDX ML-BOM
Dataset elementDatasetPackageComponent type "data"
Model-to-data linktrainedOn, testedOn relationshipsModel card dataset references
De-identification recordanonymizationMethodUsedsensitiveData plus properties
Sharing restrictionconfidentialityLevelclassification
License, stated vs reviewedhasDeclaredLicense / hasConcludedLicenselicense block (no built-in split; use properties)
Commercial termsCustom license referencelicensing object
Best fitLicense-compliance teams, legal reviewAppSec teams with existing SBOM pipelines

Licenses: identifiers are not enough

License identifiers make automated checks possible, but a dataset BOM must also carry a pointer to the governing license text. One license-drift study released a rule engine covering almost 200 SPDX and model-specific clauses, showing that SPDX identifiers let tools check obligations across datasets and models [1]. A 2026 study on license integrity reports that required license text was missing for most datasets and models it examined [2].

For a proprietary training license there is no SPDX list identifier. Declare a custom reference (for example LicenseRef-acme-support-tickets-2026) and point it at the executed agreement in your contract system, not at a public URL. Record the allowed uses as structured fields too; our guide to encoding permitted uses per record shows a schema that a BOM entry can reference.

Treat confidentiality as a BOM design decision. A customer-facing BOM may need to show "proprietary licensed dataset, supplier withheld" while the internal BOM holds the supplier identity, agreement ID and hash. Generate both from one source of truth so they cannot drift.

Populate the BOM from the training data register

The BOM should be generated, not hand-written, from the training data register that tracks each dataset's ID, version, license and allowed uses. If you keep a training data use register, use its dataset ID as the BOM element identifier (spdxId or bom-ref) so the two never disagree. NIST SP 800-218A adds AI-specific secure development practices, including the integrity of training, testing, fine-tuning and alignment data [3], which a pinned, hashed BOM supports directly.

Regulatory pressure points the same way. As of October 2026, providers placing general-purpose AI models on the EU market must maintain a copyright compliance policy that identifies and honors text-and-data-mining reservations under Article 53(1)(c), which has applied since 2 August 2025 [4]. A BOM that records each dataset's source and license makes that policy auditable; see also DSM Article 4 opt-outs.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "type": "dataset_DatasetPackage",
  "spdxId": "urn:example:tdr:DS-0417:v2026-09-30",
  "name": "support-tickets-redacted",
  "suppliedBy": "urn:example:supplier:withheld",
  "releaseTime": "2026-09-30T00:00:00Z",
  "verifiedUsing": [{ "algorithm": "sha256", "hashValue": "9f2c...e1" }],
  "dataset_datasetType": ["text", "structured"],
  "dataset_dataCollectionProcess": "Exported from ticketing system, 2021-2025, English only",
  "dataset_dataPreprocessing": ["dedup by ticket_id", "PII replacement", "HTML stripped"],
  "dataset_anonymizationMethodUsed": ["NER + regex replacement; 500-record manual check"],
  "dataset_confidentialityLevel": "amber",
  "dataset_intendedUse": "SFT and evaluation for support-agent model",
  "comment": "License: LicenseRef-DS-0417 -> contract system ID AGR-2026-118"
}

Build checklist for a dataset BOM

A usable dataset BOM passes these checks before a model release.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Every training, tuning and eval dataset has one element with a content hash, not a "latest" tag.
  • Element IDs equal training data register IDs.
  • Declared and concluded licenses are both recorded; proprietary terms use LicenseRef with a pointer to the executed agreement.
  • Allowed uses (pre-training, SFT, eval, redistribution of weights) are structured, not only in comments.
  • De-identification method and validation sample are recorded for any personal data.
  • Derived datasets link to their parents.
  • The model element links to every dataset with trainedOn or testedOn (SPDX) or model card references (CycloneDX).
  • An external redacted BOM is generated from the same source as the internal one.
  • The BOM validates against the pinned schema version in CI.

Failure modes to test for: a filtered subset shipped without a parent link; a supplier license tag ("CC-BY-4.0") that contradicts the executed agreement; an eval set silently reused for training; and BOMs regenerated per build with new IDs, which breaks diffing between releases.

Asking suppliers for BOM-ready documentation

Ask suppliers for the facts a BOM needs at contract time, because reconstructing them later is slow and often impossible. Request the dataset version identifier and hash, collection process, preprocessing steps, de-identification method, record counts and the license text or agreement reference. The sample manifest shows an example of delivery-side manifest fields, and our guide to chain of title for AI training data covers the documents behind the license field. For the wider picture, start from the provenance hub or the AI data hub.

SourceX sources operational datasets from US companies on request and prepares diligence materials per dataset covering source, rights, preparation and allowed use. Every dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, which gives your BOM's license and provenance fields a documented basis. You can describe the data you need on the buyers page.

Sourcing licensed training data for your dataset BOM

SourceX finds US companies that hold the operational data you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Personal details are removed or replaced before delivery, the method is recorded, and delivery runs through private, access-controlled workflows. Start a data request at sourcex.si/buyers.

Sources

  1. arXiv, "From Hugging Face to GitHub: Tracing License Drift in the Open-Source AI Ecosystem" (2025). https://arxiv.org/abs/2509.09873v1
  2. arXiv (Jewitt et al.), "Permissive-Washing in the Open AI Supply Chain: A Large-Scale Audit of License Integrity" (2026). https://arxiv.org/abs/2602.08816
  3. National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
  4. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  5. arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  6. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data