Skip to content

Schemas, packaging and delivery

Croissant Metadata for Licensed Datasets: What to Ask Suppliers For

Quick answer

Croissant is a JSON-LD metadata format, published by MLCommons and built on the schema.org Dataset vocabulary, that tells ML tools what files a dataset contains, how records are laid out inside them, and what the data means [1][2]. For licensed data, ask suppliers for a Croissant file that declares every delivered file with a checksum, defines each RecordSet and Field with types and sources, carries the license reference and version, and fills the core Responsible AI (RAI) fields. Then validate it yourself, knowing validation proves structure, not rights or quality.

By SourceX Editorial · Updated

How Croissant is layered: dataset, resources, structure, semantics

Croissant separates what a dataset is from where its bytes live and how its records are shaped, which is why one file can drive loading and documentation at once [1][3]. The layers you will see in a supplier's JSON-LD are:

  • Dataset metadata. schema.org properties such as name, description, license, version, datePublished, creator and citeAs, plus a conformsTo value naming the Croissant version.
  • Resources. distribution lists FileObject entries (one file, with contentUrl, encodingFormat and a sha256 or md5 checksum) and FileSet entries (many files matched by an includes glob, usually containedIn an archive FileObject) [3].
  • Structure. recordSet entries hold field entries; each Field has a dataType (for example sc:Text, sc:Integer, sc:Date) and a source that points at a FileObject or FileSet and an extract rule such as a column name, a JSON path or file content [3].
  • Semantics. Fields can reference external vocabularies, declare keys, and link to other RecordSets, so a ticket_id in one table can join to a message table.

The practical consequence: a loader such as the mlcroissant library can iterate records from a Croissant file without bespoke parsing code, provided resources and fields resolve [2]. Exact property names and versions should be checked against the current specification in the MLCommons repository before you write them into a contract exhibit, because the format is still community-developed as of October 2026 [2].

Which Croissant fields to require from a supplier

Require the fields that make a delivery loadable and auditable; treat everything else as nice to have. The table below is a starting intake standard an ML platform team can adapt. It sits alongside, not instead of, a human-readable dataset card for licensed enterprise data.

LayerFieldRequire?Why it matters for licensed data
Datasetname, descriptionRequiredMatches the dataset named in the license schedule
DatasetversionRequiredTies every training run to one delivery; see versioning licensed datasets
DatasetlicenseRequiredPoints to the agreement or license document, not a guessed SPDX ID
DatasetdatePublished or dateModifiedRequiredSeparates snapshots in recurring deliveries
DatasetconformsToRequiredPins the Croissant version your validator expects
ResourcecontentUrl, encodingFormatRequiredLoader needs both; MIME types such as application/x-parquet or application/jsonl
Resourcesha256 per FileObjectRequiredReconciles with your manifest and checksum check
Resourceincludes on FileSetsRequired for sharded dataGlobs must match exactly the delivered shards
StructurerecordSet.field.dataTypeRequiredCatches string-typed timestamps and IDs early
Structurefield.source.extractRequiredMakes the column-to-field mapping explicit
Structurekey on RecordSetsRecommendedEnables joins and deduplication across tables
RAIcollection, limitations, biases, personal data, intended usesRequired for training setsFeeds governance records and model documentation

If the supplier delivers Parquet, the column schema already lives in each file footer, so the Croissant RecordSet should mirror it rather than restate it loosely [7]. Mismatches between the Parquet physical schema and Croissant dataType are a common failure worth testing first. For format choice itself, see Parquet vs JSONL for licensed training data.

The RAI extension: what responsible-AI fields add

The Croissant RAI vocabulary turns the prose sections of a Data Card or Datasheet into machine-readable properties attached to the same JSON-LD file [4][6]. Its design use cases include data life cycle documentation, data labeling and participatory data, and it can describe collection methods, annotation processes, known limitations, biases, personal or sensitive information and intended uses [4]. Property names in published examples carry an rai: prefix, such as rai:dataCollection and rai:dataLimitations; confirm the current list in the repository [2].

RAI fields are becoming a practical default rather than an academic option. For 2026, the NeurIPS Evaluations and Datasets Track asks submitters for a minimal set of RAI fields covering limitations, potential biases and intended use [5]. That is a conference policy, not a legal requirement, but it gives buyers a ready-made floor to ask for.

For enterprise operational data, the RAI fields that matter most are the ones a supplier is uniquely placed to answer: how records were generated (system of record, export method), which de-identification method was applied to personal details, what was excluded and why, and what the data is not representative of. If your team later prepares a public training-content summary under the EU AI Act template for general-purpose models, structured collection and source fields reduce the reconstruction work [8].

Carrying license, permitted uses and version for private files

Croissant was designed around openly hosted datasets, so licensed private deliveries need a few conventions agreed up front. The license property accepts a URL or a schema.org CreativeWork; for a commercial agreement, use a CreativeWork with the agreement name and an internal identifier rather than an SPDX URL that implies an open license. contentUrl values should be relative paths inside the delivery bundle or bucket prefix, never presigned URLs that expire or leak.

Permitted uses do not have a single core Croissant property. Record them in the RAI intended-use fields and in your own namespaced properties if needed, and treat the signed license as authoritative whenever the metadata and the contract disagree. The RAI paper discusses provenance and usage-condition vocabularies at dataset and record level; check the repository before relying on any one vocabulary [4][2].

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "@context": { "@vocab": "https://schema.org/", "cr": "http://mlcommons.org/croissant/", "rai": "http://mlcommons.org/croissant/RAI/", "sc": "https://schema.org/" },
  "@type": "sc:Dataset",
  "conformsTo": "http://mlcommons.org/croissant/1.1",
  "name": "support-tickets-2021-2025",
  "description": "Resolved B2B support tickets with agent replies, de-identified.",
  "version": "2.1.0",
  "datePublished": "2026-09-30",
  "license": { "@type": "CreativeWork", "name": "Data License Agreement, Schedule A", "identifier": "DLA-EXAMPLE-001" },
  "rai:dataCollection": "Exported from the supplier ticketing system via API; closed tickets only.",
  "rai:personalSensitiveInformation": "Names, emails, phone numbers and account numbers replaced with typed placeholders before delivery.",
  "rai:dataLimitations": "English only; enterprise customers; no chat or voice channels.",
  "distribution": [
    { "@type": "cr:FileSet", "@id": "tickets-parquet", "encodingFormat": "application/x-parquet", "includes": "tickets/part-*.parquet" }
  ],
  "recordSet": [
    { "@type": "cr:RecordSet", "@id": "tickets", "key": { "@id": "tickets/ticket_id" },
      "field": [
        { "@type": "cr:Field", "@id": "tickets/ticket_id", "dataType": "sc:Text", "source": { "fileSet": { "@id": "tickets-parquet" }, "extract": { "column": "ticket_id" } } },
        { "@type": "cr:Field", "@id": "tickets/opened_at", "dataType": "sc:Date", "source": { "fileSet": { "@id": "tickets-parquet" }, "extract": { "column": "opened_at" } } }
      ] }
  ]
}

Pair the file with a byte-level manifest, such as a sample manifest, because Croissant FileSets carry globs, not per-shard checksums.

Validating Croissant files, and what validation cannot prove

Run the reference validator on every delivery, then run your own checks, because a valid Croissant file only proves the metadata is well-formed and resolvable [2]. The mlcroissant library provides a validate command that takes the JSON-LD file and reports schema and reference errors; loading records through mlcroissant.Dataset exercises the sources and extract rules against the real files [2].

An intake gate worth automating:

  1. mlcroissant validate --jsonld metadata.json passes with no errors.
  2. Every FileSet glob matches the delivered object list exactly, with no extra or missing shards.
  3. Checksums in FileObjects and the manifest match recomputed sha256 values.
  4. Iterating the first N records of every RecordSet succeeds and types match the declared dataType.
  5. version and license.identifier match the delivery note and the license schedule.
  6. Required RAI fields are non-empty and specific, not boilerplate.

Validation cannot tell you whether the supplier owns the data, whether consents cover AI training, whether de-identification held up, or whether labels are accurate. Those questions belong to rights review and quality sampling; the data provenance standards explainer and questions to ask a training data vendor cover them.

Generating Croissant from Parquet and other formats

You can generate most of a Croissant file from the data itself and should only require the supplier to author the parts a machine cannot infer. For Parquet, read each footer schema with PyArrow or DuckDB, map physical and logical types to Croissant dataType values, and emit one RecordSet per table with extract.column sources [7]. Some dataset hubs generate Croissant automatically for hosted data, but private deliveries usually need a script in your intake pipeline.

What generation cannot fill: descriptions, keys and joins, RAI fields, license references and the meaning of coded values. Ask the supplier for those as a short questionnaire, then merge answers into the generated file. For JSONL, nested objects map to extract.jsonPath; for WebDataset tar shards, a FileSet over the shard glob plus file-content extraction is the usual pattern.

Where Croissant fits next to dataset cards and table formats

Croissant is the machine-readable layer; it complements rather than replaces a dataset card, a manifest, and the license. Teams delivering through lakehouse tables can still emit Croissant for snapshot exports, though table-native metadata covers much of the structure; see Iceberg and Delta Lake delivery. The delivery formats hub maps the remaining packaging choices, and the AI data hub covers the wider buying process.

When you request data through SourceX for AI data buyers, datasets are sourced on request from US companies, rights-reviewed for ownership and consents, and delivered under a license defining records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, which is the raw material for Croissant descriptions and RAI fields.

Request licensed datasets with Croissant metadata

SourceX sources operational datasets from US companies on request and manages licensing through Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Describe the data and the metadata your pipeline expects, and delivery runs through private, access-controlled workflows after an executed agreement. Start a buyer request.

Frequently asked questions

Is Croissant a replacement for a dataset card?

No. A Data Card is written for people making decisions about a dataset, while Croissant is written for tools that load and index it [6][1]. The RAI extension narrows the gap by making card content machine-readable, but most governance teams still want both [4].

Does a Croissant file need public URLs?

No. contentUrl can be a relative path resolved against the delivery location, which suits private buckets and access-controlled transfers. Avoid embedding signed URLs, which expire and can expose access.

Which Croissant version should we standardize on?

Pin whatever version your validator and loaders support and record it in conformsTo. As of October 2026 the format is still evolving in the MLCommons repository, so re-check before changing your intake standard [2].

Sources

  1. Akhtar et al., MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  2. MLCommons, via Algolia DocSearch, "mlcommons/croissant repository documentation (spec and mlcroissant library)". https://docsearch.algolia.com/mcp/docs/repo/mlcommons/croissant
  3. Dublin Core Metadata Initiative, "Croissant (metadata standards entry)". https://msi.dublincore.org/standards/croissant
  4. Jain et al., MLCommons Croissant RAI task force, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  5. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  6. Pushkarna, Zaldivar, Kjartansson, Google Research, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  7. The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data