Schemas, packaging and delivery
Dataset Delivery Formats, Schemas and Transfer for Licensed AI Data
Quick answer
The right dataset delivery formats follow the data type and the loader that will read it: JSONL for conversations and nested records, Parquet for typed tables, original files with sidecar metadata for documents, audio and images, and tar shards for large multimodal training sets. Format is only one of nine delivery decisions. A licensed dataset also needs a written specification, a data dictionary and machine-readable metadata, a checksummed manifest, an agreed transfer channel, encryption and access rules, an update plan, version IDs and an end-of-term deletion record.
By SourceX Editorial · Updated
Nine delivery decisions, in the order to make them
Write every delivery decision into a technical exhibit attached to the license before any data moves, because each one is expensive to change after the first file lands. The table lists them in working order.
| Decision | What to fix in writing | Failure it prevents | Detailed guide |
|---|---|---|---|
| 1. Specification | Scope, schema, field types, null rules, time zones, ID scheme | Supplier exports whatever its system produces | Delivery specification template |
| 2. File format | Format per data type, compression, encoding, shard size | A conversion pass before every run | Parquet vs JSONL |
| 3. Documentation | Data dictionary, dataset card, machine-readable metadata | Provenance questions nobody can answer later | Data dictionary template |
| 4. Packaging | Manifest with byte counts, record counts and checksums | Truncated files or missing partitions | Manifests and checksums |
| 5. Channel | Push, pull, share, SFTP or shipment; who pays transfer | Stalled transfers, surprise egress bills | Cross-account bucket delivery |
| 6. Encryption and access | Keys, named readers, logs | Licensed data open to every engineer | Access controls after delivery |
| 7. Updates | Refresh or increments, deletions, schema change notice | Duplicates and deleted records that persist | Incremental vs full refresh |
| 8. Versioning | Immutable release IDs, release notes | No record of which release trained which model | Versioning licensed datasets |
| 9. End of term | Deletion scope, method, evidence | Copies outlive the license | Certificates of data destruction |
Three adjacent questions sit outside this section: measurable thresholds belong in acceptance criteria for licensed training data, failed deliveries in remedies for a defective data delivery, and label accuracy or contamination in the training data quality hub.
When SourceX manages a purchase, the license defines which records are included, what they can be used for, how long the license runs and how delivery happens, and nothing is delivered until the agreement is executed and the supplier approves the terms. Buyers who already know their format and channel requirements can include them when they submit a data request to SourceX.
Default delivery formats by data type
Most licensed datasets fit one of a handful of delivery patterns, set by the shape of the source records rather than by a supplier's default export. Override the defaults when your training stack reads something else.
| Data type | Default delivery | Must travel with it | Go deeper |
|---|---|---|---|
| Conversations, tickets, chat | JSONL, one thread per line, ordered turns | Speaker role, timestamp with UTC offset, outcome | Transcript schema |
| Tables and transactions | Parquet with explicit types; CSV only with written encoding and quoting rules | Dictionary, units, code lists | CSV specification |
| Workflow events | Event tables, or OCEL 2.0, whose exchange formats are SQLite, XML and JSON [1] | Case or object IDs, activities, event times | Linked records |
| Business documents | Originals plus extracted text and a sidecar JSON per file | Document type, link to the system-of-record entry | Sidecar metadata |
| MBOX or EML originals, or JSONL threads | Thread keys, attachment references | Email archive formats | |
| Speech and call audio | Native sample rate plus a segment manifest; diarization as RTTM, plain text with one speaker turn per line [2] | Sample rate, channels, speaker labels | Speech packaging, audio specs |
| Labeled images | Originals plus annotations, often as COCO JSON which requires info and licenses fields in its annotation file [3] | Per-image license, box coordinate convention | WebDataset shards |
| Video | Original containers plus a clip index; multi-sensor sets add aligned streams such as audio, eye gaze or stereo, as in Ego4D [4] | Frame rate, timecodes, stream alignment | Video packaging |
| Code | Git repositories with full history | Review and issue links, secrets-scan record | Repositories with history |
| Robotics and sensor logs | Native log formats, per-topic timestamps | Calibration, coordinate frames | Robotics formats |
SourceX's short answer on which file formats AI buyers accept states the rule behind the matrix: consistency and documentation matter more than the specific format.
JSONL, Parquet or tar shards: match the file to the loader
Pick the container by how the training job reads data: JSONL for nested records read line by line, Parquet for column scans over wide tables, and tar shards when one sample combines several files and must stream from object storage.
JSONL. JSON Lines requires UTF-8 with no byte order mark and one valid JSON value per line, with \n as the line terminator; a blank line is invalid [5]. The usual extension is .jsonl, and the format's site recommends a stream compressor such as gzip to save space [5]. It keeps threads intact, but every reader parses every field.
Parquet. Parquet writes its metadata in a footer that records where each column chunk starts, so readers fetch only the columns they need [6]. Implementations do not all support the same features of the format [7], so name the compression codec and logical types your readers handle instead of writing "Parquet" alone.
Tar shards. WebDataset stores samples in numbered tar shards, treats files sharing a basename as one sample, and reads shards sequentially from disk or any pipe, including cloud object stores [8]. According to Hugging Face, shards are often about 1 GB while a full dataset can reach several terabytes [9]. Specify shard size, naming and whether samples are shuffled across shards.
A workable default for mixed purchases: keep originals untouched, take records as JSONL or Parquet with a stable ID on every record, and build training shards yourself, where you control tokenization and shuffling. Related pages cover Apache Arrow formats, streaming from object storage and Iceberg and Delta Lake tables.
What must travel with the files
A delivery is incomplete without a data dictionary, a checksummed manifest, a dataset card or datasheet, and machine-readable metadata. Together they let an engineer verify the bytes and let counsel or a model team answer provenance questions months later.
- Dictionary. One row per field: type, unit, meaning of null, allowed values, source-system field and any transformation, such as replaced names.
- Datasheet or card. Datasheets document motivation, composition, collection process and recommended uses [10]. SourceX's guide to dataset cards for licensed enterprise data covers business records, and documentation frameworks compared weighs the alternatives.
- Machine-readable metadata. Croissant is a schema.org-based vocabulary in JSON-LD describing dataset metadata, file resources and record structure [11]; Croissant-RAI adds responsible-AI fields [12]. NeurIPS 2026 sets Croissant-RAI-based metadata requirements for its Evaluations and Datasets Track [13]. See Croissant metadata to ask suppliers for.
- License fields. An audit of more than 1,800 text datasets reported license omission above 70% and license errors above 50% on popular hosting sites [14]. Put the license reference and release ID in the manifest so terms survive format conversion.
Downstream disclosure duties make this metadata worth demanding. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of the datasets used, first due January 1, 2026 [15]. Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content on the AI Office template [16]. Collection dates, source categories and personal-information flags captured at delivery make both easier.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"release_id": "supplier-a-tickets-2026-10-r3",
"previous_release_id": "supplier-a-tickets-2026-07-r2",
"license_ref": "Agreement 0042, Schedule B",
"schema_version": "2.1.0",
"delivery_type": "incremental",
"generated_at": "2026-10-01T14:05:00Z",
"deidentification": {"method": "entity replacement", "sample_checked": true},
"files": [
{"path": "tickets/part-0000.jsonl.gz", "bytes": 734003200,
"records": 412337, "sha256": "<64-character hex digest>"},
{"path": "tickets/deletes-0000.jsonl.gz", "bytes": 18211,
"records": 152, "sha256": "<64-character hex digest>"}
],
"totals": {"files": 2, "records": 412489}
}
The deletes file lists record IDs to remove from earlier releases. SourceX's illustrative package summary and field guide shows the human-readable layer. Send the manifest by a different route than the data, or sign it, and recompute every checksum on arrival; the intake checklist for data engineers covers the next steps.
Transfer channels compared: push, pull, share or ship
Choose the channel by volume, by which side should control storage and cost, and by whether you need a physical copy at all. Shares avoid copying; shipment suits volumes the network cannot move in time.
| Channel | How it works | Watch for |
|---|---|---|
| Push to your bucket | Supplier writes to a prefix you own, with a credential you issue; you control retention and deletion | Limit write access to one prefix and revoke it after the final release |
| Pull from the supplier's bucket | The owner's bucket policy grants your account specific actions, your administrator delegates them, and an explicit deny overrides any allow [17] | Under S3 Requester Pays you pay requests and downloads; every request must be authenticated and carry x-amz-request-payer or RequestPayer [18] |
| Warehouse or table share | Snowflake Secure Data Sharing copies no data, and shared objects are read-only [19]. Delta Sharing is an open REST protocol over Delta Lake and Parquet with a provider-run sharing server [20] | Open-protocol recipients authenticate with a bearer token that may expire [21]; check whether the license lets you materialize a training copy |
| SFTP | Supplier uploads to a server you run, or you pull from theirs | Fine for gigabytes; slow and restart-prone for large media sets |
| Physical shipment | Encrypted drives or a cloud transfer service, with chain of custody for devices and keys | AWS Snowball Edge is no longer available to new customers; AWS points to DataSync, Data Transfer Terminal or partners [22] |
| API or streaming feed | Supplier publishes records continuously | Replay and backfill rules; see API feeds vs batch files |
For purchases SourceX manages, delivery runs through private, access-controlled workflows, never email attachments; see its answer on how licensed data is delivered. Detailed pages cover zero-copy warehouse sharing, SFTP limits, multi-terabyte network transfers, encrypted drives, offline appliances and egress costs.
Keeping licensed data controlled after it lands
Licensed data stays under license after it lands, so the receiving side needs keys handled apart from the data, access limited to named roles, a log of every read, and a record of which models used which release.
- Encryption. Encrypt in transit and at rest, agree a file-level method such as OpenPGP, age or a cloud KMS key shared across accounts, and send keys by a different channel than files. See encryption and key exchange and SourceX's answer on encryption in transit.
- Landing zone. Land each release in a quarantine prefix, verify the manifest, then promote it to a restricted location; audit logging covers what to record.
- Personal data. For data sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing. No method is perfect, so keep re-identification controls; the de-identified data hub covers residual risk.
- Lineage. Record the release ID in every training and evaluation run, or emit it as OpenLineage events, so dataset-to-model traceability holds up to a deletion request or license audit.
- End of term. Deletion must reach derived shards, embeddings, feature stores, caches and backups. Model retention after termination is a contract question for the licensing hub.
Recurring purchases: increments, corrections and schema changes
Ongoing purchases usually fail on identity and change control rather than on format: records need IDs that stay stable across releases, deletions need their own channel, and schema changes need notice and a version number.
Increments need a stable key, a change type on each record (insert, update or delete) and a high-water mark such as last-modified time. Pseudonymized IDs must stay consistent across releases or increments cannot join to earlier data; see stable record IDs and propagating deletions and corrections.
Treat schema changes like API changes: additive fields raise a minor version, while renamed, retyped or removed fields need notice and a major version. Data contracts turn those rules into checks, and schema change handling covers migrations. Delta Sharing supports time-travel queries on shared tables [21]; for copies you hold, see versioning tools such as DVC and lakeFS.
Delivery mistakes that surface after the data arrives
Most delivery problems surface in the first training or evaluation job, after the files were accepted, so name these failure modes in the specification.
- Timestamps without a UTC offset, mixing local times across sites; see timestamps and time zones in delivered datasets.
- CSV exports with locale-specific decimal separators or date formats.
- Attachments referenced by ID in the records but missing from the package.
- Image metadata stripped for privacy along with the orientation flag, leaving rotated images.
- Narrowband telephone audio upsampled to match wideband recordings, with no field saying so.
- Checksums computed before compression while the manifest lists compressed files.
- License fields dropped when files are repackaged or converted; audits of hosted datasets find licenses often missing [14].
- Training from a read-only warehouse share without a written right to copy.
SourceX's guide on how to procure enterprise training data places delivery in the buying sequence; the procurement hub and AI data overview cover the surrounding steps.
Planning how a licensed dataset reaches your team
If you need operational data from US companies, you can describe the records you need on SourceX's buyer page, along with your format, metadata and delivery requirements. SourceX looks for US businesses that hold the data, checks the data and the supplier's licensing permissions, and manages the license and the coordination of delivery and payment. Datasets are sourced on request, so a request does not guarantee a matching dataset. Describe your format and delivery needs.
Guides in this section
- Access Control for Licensed Training Data After DeliveryTurn license terms into enforceable access controls: per-license storage, license-ID tags, ABAC policies, write blocks, lineage checks and access reviews.
- Certificate of Data Destruction for Licensed AI DatasetsHow to scope, execute and certify deletion of a licensed dataset at license end: every copy, backup and cache, NIST SP 800-88 methods and a template.
- Cross-Account Bucket Delivery for Licensed DatasetsHow to set up cross-account S3, GCS and Azure Blob access for dataset delivery: push vs pull, bucket and KMS key policies, SAS expiry and access logging.
- Data Dictionary Template for Licensed Dataset DeliveriesA field-level data dictionary template to require from data suppliers: types, code lists, source mapping, transformations, redaction status and versioning.
- Dataset Delivery Specification Template for Data LicensesA technical delivery exhibit for licensed AI datasets: format, schema, documentation, manifest, channel, encryption, cadence, versioning and deletion.
- Dataset Manifests and Checksums: Verify Every Delivered FileHow to verify a licensed dataset delivery: manifest fields, SHA-256 vs MD5, S3 checksums, multipart ETags, signed manifests and row count reconciliation.
- Dataset Versioning for Licensed Data: IDs and Release NotesDataset versioning best practices for licensed AI data: immutable snapshots, version IDs, release notes, run pinning and license-to-version mapping.
- Dataset-to-Model Traceability for Licensed Training DataBuild lineage records that show which checkpoints, derived datasets and products used a licensed dataset version, from run configs to the model registry.
- Incremental vs Full Refresh Deliveries for Licensed DataChoose between full snapshots and incremental deltas for recurring dataset purchases, and specify how inserts, updates, deletes and watermarks are signaled
- Licensed Dataset Intake Checklist for Data EngineersA step-by-step intake checklist for licensed datasets: quarantine, malware scan, checksum and count reconciliation, schema checks and catalog registration.
- Packaging Linked Records from CRM, Ticketing and BillingHow to deliver relational data for machine learning: keyed tables, crosswalks, per-case bundles, integrity checks and pseudonymous keys across systems.
- Parquet vs JSONL for Training Data DeliveriesWhen to require Parquet or JSONL from a data supplier: schema enforcement, nested conversations, compression, validation and conversion pitfalls.
- PGP, age and KMS Encryption for Dataset DeliveriesHow to encrypt licensed dataset files with OpenPGP, age or cloud KMS envelope encryption, exchange and verify keys, and rotate keys on recurring feeds.
- Schema Changes in Recurring Data Deliveries: A PlaybookHow to classify, detect and absorb schema changes between recurring dataset deliveries: compatibility modes, renames, type changes and code-list drift.
- Shipping Datasets on Encrypted Drives: Keys and CustodyHow to receive a multi-terabyte dataset on an encrypted drive: FIPS 140-3 checks, out-of-band keys, custody logs, manifest checks and drive sanitization.
- Zero-Copy Data Sharing for Licensed AI Training DataHow Snowflake shares, Delta Sharing and BigQuery listings deliver licensed data without copying, and what training exports and revocation mean for buyers.
Sources
- OCEL standard authors (ocel-standard.org), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- MISP 2022 challenge organizers, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
- CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
- Grauman et al. (Ego4D consortium), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
- The Apache Software Foundation (Apache Parquet project), "Apache Parquet Documentation". https://parquet.apache.org/docs
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
- Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
- Akhtar et al. (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Jain et al. (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- Amazon Web Services, "Example 2: Bucket owner granting cross-account bucket permissions" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-walkthroughs-managing-access-example2.html
- Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
- Snowflake Inc., "About Secure Data Sharing" (Snowflake Documentation). https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021, vendor launch post). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- delta-io/delta-sharing documentation (rendered by Mintlify), "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
- Amazon Web Services, "AWS Snowball Edge availability change" (AWS Snowball Edge Developer Guide). https://docs.amazonaws.cn/en_us/snowball/latest/developer-guide/snowball-edge-availability-change.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.