Schemas, packaging and delivery
Dataset Manifests and Checksums: Verifying a Delivery Byte for Byte
Quick answer
To verify a dataset delivery, require a manifest that lists every file's relative path, byte size, SHA-256 digest and record count, plus the dataset version and a delivery ID. Recompute SHA-256 locally (or compare against cloud-native object checksums), reconcile file and row counts against the manifest, and verify a detached signature over the manifest itself. MD5 detects accidental corruption but is not tamper evidence. Quarantine the delivery until every check passes.
By SourceX Editorial · Updated
Integrity checking is the first gate in any pipeline that consumes licensed data, before profiling, PII audits or training runs. The basic principle is old archive practice: a checksum computed before and after a copy should match, and any single-byte change shows up as a mismatch [2]. Commercial data customers ask for it explicitly; buyers of one provider's downloadable datasets have requested published checksums so they can confirm files arrived intact [7]. This page covers the method. For a worked example of the document itself, see the sample delivery manifest, and for the wider set of delivery decisions, start at the delivery formats and transfer hub.
What a delivery manifest must list
A delivery manifest must let you prove two things independently: that the delivery is complete (every expected file is present and no unexpected file appears) and that it is valid (every file's content matches its recorded digest). That split comes straight from BagIt, the IETF-documented packaging format used by archives: a bag is complete when every file listed in its manifests is present, and valid when every checksum in every manifest verifies. Borrow the vocabulary in your delivery specification even if you never ship an actual bag.
At minimum, each file entry should carry the fields below. Dataset-level fields belong in a header block, or in a sidecar metadata file such as a Croissant JSON-LD description, which already models file resources and record structure [4].
| Field | Level | Why it matters |
|---|---|---|
delivery_id | Dataset | Ties the files to one transfer event and to the license schedule that defines the records |
dataset_version | Dataset | Distinguishes a re-delivery from a new release; see versioning licensed datasets |
schema_version | Dataset or file | Lets you reject files written against a schema you have not mapped |
hash_algorithm | Dataset | Removes ambiguity; never infer the algorithm from digest length |
path | File | Relative POSIX path from the delivery root, case-sensitive, no .. |
bytes | File | Cheap first check that catches truncation before hashing |
sha256 | File | Lowercase hex digest of the exact bytes delivered (after compression, before decryption is a common convention; state which) |
record_count | File | Basis for row reconciliation |
content_type / compression | File | e.g. application/vnd.apache.parquet, jsonl+zstd |
part_of | File | Shard group, partition or table name |
Two details cause most disputes. First, state whether digests cover the encrypted, compressed or plain file: if the supplier hashed plaintext and you receive OpenPGP-encrypted objects, nothing will match until after decryption (see encryption and key exchange for deliveries). Second, the manifest must list itself out, or be covered by a separate signature, because a manifest cannot contain its own digest.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"delivery_id": "dlv-2026-10-0042",
"dataset_version": "3.1.0",
"schema_version": "tickets-v4",
"hash_algorithm": "sha256",
"digest_scope": "file as delivered (zstd-compressed, not encrypted)",
"file_count": 3,
"total_records": 1842210,
"files": [
{"path": "tickets/part-00000.parquet", "bytes": 268431201, "sha256": "9f2c...e41a", "record_count": 614070},
{"path": "tickets/part-00001.parquet", "bytes": 268399874, "sha256": "41b7...0c3d", "record_count": 614070},
{"path": "tickets/part-00002.parquet", "bytes": 268102556, "sha256": "d08e...77f2", "record_count": 614070}
]
}
MD5 vs SHA-256 for file transfer
Use SHA-256 for delivery manifests; treat MD5 as a corruption check only, never as evidence that files were not altered. The IETF's updated security considerations for MD5 (RFC 6151) conclude that MD5 is no longer acceptable where collision resistance is required, such as digital signatures; practical collision attacks have been public since 2004 and run quickly on commodity hardware. A manifest is effectively a signature over your training data, so collision resistance is exactly the property you need.
In practice the risk is not random bit flips, which MD5 and even CRCs catch, but a party in the chain substituting a file that hashes identically. That matters more now that frameworks such as NIST SP 800-218A add practices addressing the integrity of training, testing, fine-tuning and alignment data for foundation models [3]. If a legacy supplier can only produce MD5, accept it as a transport check, recompute SHA-256 on receipt, and record your own digest as the reference value for lineage.
Fast non-cryptographic checksums (CRC32C, CRC64NVME, xxHash) still have a role: they are cheap enough to run on every hop. Keep them as transport checks alongside a cryptographic digest, not as its replacement.
Verifying S3 object checksums and multipart ETags
Cloud object checksums can replace a full download-and-hash pass, but only when you compare like with like. Amazon S3 supports several checksum algorithms for upload integrity, with CRC64NVME as the default, and offers a Batch Operations compute-checksum job that calculates checksums for objects already in a bucket and writes a completion report [1]. That lets you verify a multi-terabyte delivery in place, then compare the report against the manifest.
The trap is the ETag. For single-part uploads that are unencrypted or use SSE-S3 (not SSE-KMS or SSE-C), the ETag is typically the MD5 of the object, but for multipart uploads AWS documents a composite value derived from the parts, shown with a -N part-count suffix, not a hash of the whole object [1]. A supplier who pastes ETags into a manifest has given you numbers you cannot reproduce unless you also know the exact part size. Composite (checksum-of-checksums) values have the same property.
Rules that avoid false mismatches:
- Ask the supplier to request a full-object checksum at upload and to record the algorithm and checksum type in the manifest. As of October 2026, S3 offers the full-object type for multipart uploads only with CRC-based algorithms (CRC64NVME, CRC32, CRC32C); SHA-256 on a multipart upload is a composite value, so a whole-file SHA-256 from the manifest must be recomputed rather than read from S3.
- Never compare an ETag to a SHA-256. Compare SHA-256 to SHA-256, CRC64NVME to CRC64NVME.
- If a copy crosses accounts or regions, re-verify at the destination; a server-side copy can recompute multipart layout. See cross-account bucket delivery patterns.
- For deliveries over SFTP or network tools, hash after the transfer completes, not while temp files are still being renamed (network transfer tools and tuning).
Signing the manifest so the file list is tamper-evident
Sign the manifest with a detached signature so that an attacker cannot swap a file and update its digest at the same time. Checksums only prove that files match the manifest; if the manifest travels on the same channel as the data, anyone who can alter one can alter both. A detached signature (for example an OpenPGP .asc or .sig file produced with the supplier's published key, or a Sigstore or minisign signature) binds the file list to a key you obtained out of band.
Agree on three things in the technical delivery specification: which key or identity signs, how you received the public key (never in the same bucket as the data), and what happens on key rotation. Verify the signature first; if it fails, do not run any other check, because every subsequent comparison would be against untrusted values. BagIt achieves a similar effect with a tag manifest that hashes the payload manifest, though that still needs an external anchor.
File count and row count reconciliation
Reconciliation catches failures that checksums cannot, such as a supplier exporting the wrong date range into perfectly intact files. Run it after digests pass, at three levels.
File set. Compute the set difference between paths on disk and paths in the manifest. Missing files mean an incomplete delivery; extra files (stray _SUCCESS markers, .crc sidecars, .DS_Store, temp uploads) are either ignored by an agreed allowlist or treated as a defect. Unknown extra files in a licensed delivery also need a scoping question: are they inside the licensed records?
Format validity. A truncated Parquet file usually fails fast because readers look for the PAR1 magic number at both ends and the footer metadata written after the data [5]. JSONL is less forgiving: a file cut mid-line may still parse up to the break, so count lines and confirm each one is valid JSON; blank lines and a UTF-8 byte order mark are both invalid under the convention [6]. Choose formats with this in mind; see Parquet vs JSONL for training data.
Rows. Compare Parquet footer row counts or JSONL line counts with record_count per file and with total_records. Then reconcile against the license schedule: if the agreement defines records as tickets closed in a date window, check min and max timestamps and distinct primary keys, not just totals. Duplicates across shards often signal an overlapping incremental extract (see incremental vs full refresh deliveries).
Verification runbook and handling partial deliveries
A verification run should be scripted, deterministic and logged, so that its output can serve as evidence in a dispute and as the first event in your lineage record. The sequence below is ordered so that cheap and trust-establishing checks come first.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Step | Check | Tooling example | On failure |
|---|---|---|---|
| 1 | Land in a quarantine prefix with no training-job read access | Bucket policy or IAM deny | n/a |
| 2 | Verify detached signature on manifest | gpg --verify manifest.json.asc manifest.json | Stop; request re-issue over a separate channel |
| 3 | File set: on disk vs manifest | Script diffing path lists | Missing: request re-send; extra: allowlist or reject |
| 4 | Byte sizes | stat or object listing | Re-transfer the file |
| 5 | SHA-256 per file | sha256sum -c manifest.sha256 or S3 compute-checksum report [1] | Re-transfer; if repeat mismatch, escalate as content defect |
| 6 | Format validity | Parquet footer read; JSONL line parse | Request regenerated file |
| 7 | Row counts and key ranges | Footer num_rows, line counts, min/max timestamps | Reconcile against license schedule |
| 8 | Record results; promote to curated zone | Verification log keyed by delivery_id | n/a |
For partial deliveries, never promote a subset silently. Either hold the whole delivery until it is complete, or agree in advance that shards are independently usable and record which shards were promoted under which delivery_id. When a file is re-sent, the replacement must carry the same path and a digest matching the original manifest; a different digest means a new version, which needs a new manifest. Write the verification outcome into your lineage system so that later audits can show which bytes a model saw (see OpenLineage for training data).
Once integrity passes, content checks follow: residual PII sampling (auditing residual PII in a delivered dataset) and access restrictions after landing (access controls for licensed training data).
How SourceX handles delivery
SourceX sources operational datasets from US companies on request; datasets are not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, and diligence materials covering source, rights, preparation and allowed use are prepared per dataset. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval. The delivery overview explains the handover, and buyers can describe the data they need so the manifest and verification steps above can be written into the request.
Request a dataset with a verifiable delivery
If your team needs operational data such as support histories, engineering records or finance workflows, describe the data rather than the businesses that might hold it. SourceX looks for US companies that hold it, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is transacted. Start a buyer request at sourcex.si/buyers.
Sources
- Amazon Web Services (Amazon S3 User Guide), "Checking object integrity in Amazon S3". https://docs.aws.amazon.com/hi_in/AmazonS3/latest/userguide/checking-object-integrity.md
- UK Data Service, "Checksums for data integrity (data archive guidance)". https://ukdataservice.ac.uk/?p=3785
- National Institute of Standards and Technology, "Secure Software Development Practices for Generative AI and Dual-Use Foundation Models (NIST SP 800-218A)" (2024). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-218A.pdf
- Akhtar et al., MLCommons (arXiv:2403.19546), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- The Apache Software Foundation (Apache Parquet), "File Format". https://parquet.apache.org/docs/file-format/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- IPinfo Community, "Checksum endpoint for data downloads". https://community.ipinfo.io/t/checksum-endpoint-for-data-downloads/1667
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.