Schemas, packaging and delivery
Technical Delivery Specification Template for Licensed Datasets
Quick answer
A data delivery specification is the technical exhibit attached to a data license that states exactly what arrives, in what shape, by which channel, how often, and how it leaves your systems at the end. At minimum it should fix ten things: file format and compression, schema and data dictionary, documentation, folder layout and manifest, transfer channel and credentials, encryption, cadence, versioning, change notice, and deletion or return. Remedies and service levels belong in the contract body, not here.
By SourceX Editorial · Updated
This page gives a fill-in template and the reasoning behind each field. It sits inside the dataset delivery formats and transfer hub and assumes you already know which records you want; if you are still scoping the data itself, start with how to write a data request for suppliers.
What a delivery exhibit covers and what it leaves to the contract
A delivery exhibit covers the mechanics of handover, while price, warranties, remedies and service credits stay in the license and any SLA schedule. Data marketplaces already treat delivery as its own structured object: Informatica's Data Marketplace defines a delivery template as a delivery format, a delivery method and a target location [1]. Older public-sector protocols go further, listing units and formats, collection procedures, responsible personnel, delivery schedule and mode, confidentiality and contacts [2].
For licensed AI training data you need more than either model, because the ML platform team will load the files straight into a pipeline. A vague "CSV via secure transfer" clause leaves room for delimiter disputes, silent schema drift and orphaned copies after termination. Keep the exhibit factual and testable, and point to acceptance criteria for licensed training data for what counts as a conforming delivery and to data supplier SLAs for timing and remedies.
The template: ten sections to fill in
The fastest way to draft the exhibit is to fill one row per section and attach the data dictionary and a sample manifest as annexes. Every value should be checkable by a script or a named person on delivery day.
Illustrative example: invented to show structure; it does not describe an available dataset.
| # | Section | Fields to specify | Example value | Common failure mode |
|---|---|---|---|---|
| 1 | Format and compression | Container, compression codec, encoding, max file size | Parquet, Snappy-compressed, UTF-8, files 256 MB to 1 GB | Supplier ships XLSX exports with merged header rows |
| 2 | Schema and data dictionary | Field names, types, nullability, enums, units, time zone | Annex A dictionary; all timestamps ISO 8601 UTC | Free-text "status" field with 40 undocumented values |
| 3 | Documentation | Dataset card, Croissant JSON-LD, preparation notes | README.md card plus croissant.json per release | Card describes a sample, not the delivered population |
| 4 | Layout and manifest | Folder convention, partitioning, manifest columns, checksum algorithm | /v{version}/{table}/dt=YYYY-MM-DD/; SHA-256 per file | Manifest generated before late files were added |
| 5 | Channel and credentials | Channel, who owns the destination, identity model, key rotation | Supplier writes to buyer-owned bucket via cross-account role | Shared long-lived access keys sent by email |
| 6 | Encryption | In transit, at rest, key ownership | TLS 1.2+ in transit; server-side encryption with buyer-managed key | Archive password sent in the same message as the link |
| 7 | Cadence and increments | Full vs incremental, schedule, late-arriving data, deletes | Monthly increment by updated_at, tombstone file for deletions | Increments that drop deleted records silently |
| 8 | Versioning | Version ID scheme, immutability, release notes | 2026.10.0; prior versions retained read-only | Files overwritten in place under the same path |
| 9 | Change notice | Schema change classes, notice period, sample in advance | Breaking changes announced one cycle ahead with a sample file | Column renamed mid-contract with no notice |
| 10 | Deletion and return | Trigger, scope (including derived copies), method, certificate | Deletion within an agreed window after term; written certificate citing method | "Deleted" means moved to a trash folder |
The example values are placeholders. Replace each with what your platform team can actually ingest and what the supplier can actually produce, then have both sides initial the annex.
Format, compression and schema: write them so a parser can enforce them
Format clauses should name a specification, not a file extension, because "CSV" and "JSON" each cover many incompatible dialects. RFC 4180 documents a common CSV format and the text/csv MIME type, but real exports vary, so state the delimiter, quote character, line ending, header row and encoding explicitly. Our CSV delivery specification guide lists the dialect fields worth pinning down.
For columnar data, Parquet is the usual default for tabular training corpora. The official documentation separates the format spec from the implementation status page, which matters if your reader cannot handle a newer encoding or logical type the supplier's writer emits [3]. Name the codec (Snappy, ZSTD or GZIP) and a target file size, since thousands of tiny files slow every downstream scan; the trade-offs against line-delimited JSON are covered in Parquet vs JSONL for licensed training data.
The schema itself belongs in an annex, not in prose. Use the data dictionary template for licensed deliveries and require, per field: name, physical type, nullability, allowed values, unit, source system, and whether it was transformed during de-identification.
Documentation: dataset card and Croissant metadata
Require human-readable and machine-readable documentation for every release, not just the first one. A dataset card is the human-readable layer; Hugging Face's guidance frames it as the place to promote responsible use and disclose potential biases [6]. Ask that the card describe the delivered population, collection window, preparation steps and known gaps.
The machine-readable layer is increasingly Croissant, a schema.org-based JSON-LD vocabulary that describes dataset metadata, file resources and record structure so ML tooling can load data without custom glue code [4]. The Croissant-RAI extension adds responsible-AI fields that build on Data Cards and Datasheets for Datasets, including life cycle and labeling documentation [5]. Specify which fields are mandatory; do not assume a supplier will populate optional provenance or labeling sections unless the exhibit says so.
Layout, manifest and verification
Each delivery should arrive with a manifest that lets you prove completeness before anyone trains on it. A practical manifest lists relative path, byte size, row count, checksum algorithm and value, schema version and delivery ID. SHA-256 is the common choice for checksums; MD5 values from storage ETags are not a reliable integrity check for multipart uploads.
Write the verification rule into the exhibit: the delivery is "received" only when every manifest entry matches, and the manifest itself is the last file written. The manifest and checksum verification guide has a step-by-step check. Fix the path convention too, for example version, table, then date partition, so increments never collide with prior releases.
Channel, credentials and who pays for transfer
The channel clause should name who owns the destination, who can write to it, and who pays for data movement. The cleanest pattern for many buyers is a supplier writing into a buyer-owned bucket through a scoped cross-account role, described in cross-account cloud bucket delivery. SFTP, warehouse sharing and table formats such as Iceberg or Delta Lake are alternatives with different operational costs.
Transfer costs need an explicit owner. On Amazon S3, a Requester Pays bucket shifts request and download charges to the requester while the bucket owner still pays storage, and anonymous access is not permitted [7]. Spell out who bears egress, storage during the retention window and any re-delivery costs; the egress cost guide works through common splits.
For very large datasets, check that any offline transfer option named in the exhibit is still orderable. As of October 2026, AWS documentation states that Snowball Edge has been available only to existing customers since November 7, 2025, and no Snow Family device can be ordered by new customers [8]. Never accept "we will email it" for anything containing records; credentials should be per-person or per-role, time-limited and revocable.
Encryption, cadence, versioning and change notice
Encryption, update rhythm and change control are where long-running data purchases most often break. State encryption in transit and at rest, and who holds the keys; buyer-managed keys let you revoke access to your own copy and support crypto-shredding at termination. If files are additionally PGP-encrypted, the exhibit should say how public keys are exchanged and rotated.
For cadence, choose between full refreshes and increments, and define how updates and deletions in the source are represented. A monthly increment keyed on updated_at with a separate tombstone file is easier to audit than a full replacement that silently drops records; see incremental vs full refresh deliveries. Every release needs an immutable version ID and release notes, as covered in versioning licensed datasets.
Change notice should classify changes. Additive columns may need only release notes, while renames, type changes, enum changes and granularity changes should require advance notice and a sample file before the changed release ships. The schema evolution guide offers a classification you can paste in.
Deletion and return at the end of the license
The deletion section should name the trigger, the scope, the method and the evidence. Triggers usually include expiry, termination and a supplier withdrawal of specific records; scope should say whether working copies, caches, feature stores and backups are included, and how derived artifacts are treated under the license.
For method, reference a recognized standard instead of the word "delete." NIST SP 800-88, Guidelines for Media Sanitization, distinguishes clear, purge and destroy methods [9]; name the method and the revision in force when you sign. In cloud environments, destroying the encryption key under a purge method is often the only practical way to reach every replica, which is another reason to fix key ownership up front. Close with the form of evidence: a written certificate naming the systems, method and date.
Worked example: a recurring support-ticket delivery
Here is how the exhibit reads for a recurring purchase of de-identified support conversations.
Illustrative example: invented to show structure; it does not describe an available dataset.
delivery_exhibit:
dataset: support_conversations_deidentified
format: { container: parquet, codec: zstd, encoding: utf-8, target_file_mb: 512 }
schema_annex: annex_a_data_dictionary_v3.xlsx
documentation: [README.md, croissant.json, preparation_notes.md]
layout: "/v{version}/{table}/dt={YYYY-MM-DD}/part-{n}.parquet"
manifest: { file: manifest.csv, checksum: sha256, written_last: true }
channel: { type: cross_account_write, destination_owner: buyer, credential_ttl_hours: 12 }
encryption: { in_transit: "TLS 1.2+", at_rest: "SSE with buyer-managed key" }
cadence: { mode: incremental, schedule: monthly, key: updated_at, deletes: tombstones.parquet }
versioning: { scheme: "YYYY.MM.patch", immutable: true, release_notes: required }
change_notice: { breaking: "one delivery cycle ahead plus sample file", additive: "release notes" }
transfer_costs: { egress: supplier, destination_storage: buyer }
deletion: { trigger: [expiry, termination, record_withdrawal], method: "NIST SP 800-88 Rev. 2 purge", evidence: written_certificate }
Notice what is absent: price, warranties and remedies. Those live in the license and SLA, which keeps the exhibit stable when commercial terms are renegotiated. For how licenses structure rights and allowed uses, see the AI training data licensing guide and the SourceX licensing terms page.
How SourceX handles delivery
SourceX sources operational datasets from US companies on request and manages the commercial process, including the licensing agreement and ongoing purchases. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval. The general flow is explained in how licensed data is delivered. If you already have an exhibit drafted, you can share your delivery requirements with SourceX alongside the description of the data you need.
Get a delivery specification matched to your data request
Describe the operational data you need and the format, channel and cadence your platform team expects; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and nothing is contracted until a supplier agrees. Start a buyer request at sourcex.si/buyers.
Sources
- Informatica, "Delivery templates (Data Marketplace bulk upload documentation)". https://docs.informatica.com/data-governance-and-quality-cloud/data-marketplace/current-version/set-up-data-marketplace/create-new-items/delivery-options/creating-delivery-templates.html
- UNFCCC Clean Development Mechanism, "QA/QC Process (includes a Data Delivery Protocol template)". https://cdm.unfccc.int/sunsetcms/storage/contents/stored-file-20150701152654040/QA%20QC%20Process.pdf
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
- MLCommons Croissant working group (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets". https://arxiv.org/pdf/2403.19546
- MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI". https://arxiv.org/pdf/2407.16883
- Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- Amazon Web Services, "AWS Snowball Edge availability change". https://docs.aws.amazon.com/snowball/latest/developer-guide/snowball-edge-availability-change.html
- NIST, "SP 800-88 Rev. 2, Guidelines for Media Sanitization". https://csrc.nist.gov/pubs/sp/800/88/r2/final
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.