Skip to content

Schemas, packaging and delivery

Dataset Delivery Formats, Schemas and Transfer for Licensed AI Data

Quick answer

The right dataset delivery formats follow the data type and the loader that will read it: JSONL for conversations and nested records, Parquet for typed tables, original files with sidecar metadata for documents, audio and images, and tar shards for large multimodal training sets. Format is only one of nine delivery decisions. A licensed dataset also needs a written specification, a data dictionary and machine-readable metadata, a checksummed manifest, an agreed transfer channel, encryption and access rules, an update plan, version IDs and an end-of-term deletion record.

By SourceX Editorial · Updated

Nine delivery decisions, in the order to make them

Write every delivery decision into a technical exhibit attached to the license before any data moves, because each one is expensive to change after the first file lands. The table lists them in working order.

DecisionWhat to fix in writingFailure it preventsDetailed guide
1. SpecificationScope, schema, field types, null rules, time zones, ID schemeSupplier exports whatever its system producesDelivery specification template
2. File formatFormat per data type, compression, encoding, shard sizeA conversion pass before every runParquet vs JSONL
3. DocumentationData dictionary, dataset card, machine-readable metadataProvenance questions nobody can answer laterData dictionary template
4. PackagingManifest with byte counts, record counts and checksumsTruncated files or missing partitionsManifests and checksums
5. ChannelPush, pull, share, SFTP or shipment; who pays transferStalled transfers, surprise egress billsCross-account bucket delivery
6. Encryption and accessKeys, named readers, logsLicensed data open to every engineerAccess controls after delivery
7. UpdatesRefresh or increments, deletions, schema change noticeDuplicates and deleted records that persistIncremental vs full refresh
8. VersioningImmutable release IDs, release notesNo record of which release trained which modelVersioning licensed datasets
9. End of termDeletion scope, method, evidenceCopies outlive the licenseCertificates of data destruction

Three adjacent questions sit outside this section: measurable thresholds belong in acceptance criteria for licensed training data, failed deliveries in remedies for a defective data delivery, and label accuracy or contamination in the training data quality hub.

When SourceX manages a purchase, the license defines which records are included, what they can be used for, how long the license runs and how delivery happens, and nothing is delivered until the agreement is executed and the supplier approves the terms. Buyers who already know their format and channel requirements can include them when they submit a data request to SourceX.

Default delivery formats by data type

Most licensed datasets fit one of a handful of delivery patterns, set by the shape of the source records rather than by a supplier's default export. Override the defaults when your training stack reads something else.

Data typeDefault deliveryMust travel with itGo deeper
Conversations, tickets, chatJSONL, one thread per line, ordered turnsSpeaker role, timestamp with UTC offset, outcomeTranscript schema
Tables and transactionsParquet with explicit types; CSV only with written encoding and quoting rulesDictionary, units, code listsCSV specification
Workflow eventsEvent tables, or OCEL 2.0, whose exchange formats are SQLite, XML and JSON [1]Case or object IDs, activities, event timesLinked records
Business documentsOriginals plus extracted text and a sidecar JSON per fileDocument type, link to the system-of-record entrySidecar metadata
EmailMBOX or EML originals, or JSONL threadsThread keys, attachment referencesEmail archive formats
Speech and call audioNative sample rate plus a segment manifest; diarization as RTTM, plain text with one speaker turn per line [2]Sample rate, channels, speaker labelsSpeech packaging, audio specs
Labeled imagesOriginals plus annotations, often as COCO JSON which requires info and licenses fields in its annotation file [3]Per-image license, box coordinate conventionWebDataset shards
VideoOriginal containers plus a clip index; multi-sensor sets add aligned streams such as audio, eye gaze or stereo, as in Ego4D [4]Frame rate, timecodes, stream alignmentVideo packaging
CodeGit repositories with full historyReview and issue links, secrets-scan recordRepositories with history
Robotics and sensor logsNative log formats, per-topic timestampsCalibration, coordinate framesRobotics formats

SourceX's short answer on which file formats AI buyers accept states the rule behind the matrix: consistency and documentation matter more than the specific format.

JSONL, Parquet or tar shards: match the file to the loader

Pick the container by how the training job reads data: JSONL for nested records read line by line, Parquet for column scans over wide tables, and tar shards when one sample combines several files and must stream from object storage.

JSONL. JSON Lines requires UTF-8 with no byte order mark and one valid JSON value per line, with \n as the line terminator; a blank line is invalid [5]. The usual extension is .jsonl, and the format's site recommends a stream compressor such as gzip to save space [5]. It keeps threads intact, but every reader parses every field.

Parquet. Parquet writes its metadata in a footer that records where each column chunk starts, so readers fetch only the columns they need [6]. Implementations do not all support the same features of the format [7], so name the compression codec and logical types your readers handle instead of writing "Parquet" alone.

Tar shards. WebDataset stores samples in numbered tar shards, treats files sharing a basename as one sample, and reads shards sequentially from disk or any pipe, including cloud object stores [8]. According to Hugging Face, shards are often about 1 GB while a full dataset can reach several terabytes [9]. Specify shard size, naming and whether samples are shuffled across shards.

A workable default for mixed purchases: keep originals untouched, take records as JSONL or Parquet with a stable ID on every record, and build training shards yourself, where you control tokenization and shuffling. Related pages cover Apache Arrow formats, streaming from object storage and Iceberg and Delta Lake tables.

What must travel with the files

A delivery is incomplete without a data dictionary, a checksummed manifest, a dataset card or datasheet, and machine-readable metadata. Together they let an engineer verify the bytes and let counsel or a model team answer provenance questions months later.

  • Dictionary. One row per field: type, unit, meaning of null, allowed values, source-system field and any transformation, such as replaced names.
  • Datasheet or card. Datasheets document motivation, composition, collection process and recommended uses [10]. SourceX's guide to dataset cards for licensed enterprise data covers business records, and documentation frameworks compared weighs the alternatives.
  • Machine-readable metadata. Croissant is a schema.org-based vocabulary in JSON-LD describing dataset metadata, file resources and record structure [11]; Croissant-RAI adds responsible-AI fields [12]. NeurIPS 2026 sets Croissant-RAI-based metadata requirements for its Evaluations and Datasets Track [13]. See Croissant metadata to ask suppliers for.
  • License fields. An audit of more than 1,800 text datasets reported license omission above 70% and license errors above 50% on popular hosting sites [14]. Put the license reference and release ID in the manifest so terms survive format conversion.

Downstream disclosure duties make this metadata worth demanding. As of October 2026, California AB 2013 requires developers of generative AI systems offered to Californians to post a high-level summary of the datasets used, first due January 1, 2026 [15]. Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content on the AI Office template [16]. Collection dates, source categories and personal-information flags captured at delivery make both easier.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "release_id": "supplier-a-tickets-2026-10-r3",
  "previous_release_id": "supplier-a-tickets-2026-07-r2",
  "license_ref": "Agreement 0042, Schedule B",
  "schema_version": "2.1.0",
  "delivery_type": "incremental",
  "generated_at": "2026-10-01T14:05:00Z",
  "deidentification": {"method": "entity replacement", "sample_checked": true},
  "files": [
    {"path": "tickets/part-0000.jsonl.gz", "bytes": 734003200,
     "records": 412337, "sha256": "<64-character hex digest>"},
    {"path": "tickets/deletes-0000.jsonl.gz", "bytes": 18211,
     "records": 152, "sha256": "<64-character hex digest>"}
  ],
  "totals": {"files": 2, "records": 412489}
}

The deletes file lists record IDs to remove from earlier releases. SourceX's illustrative package summary and field guide shows the human-readable layer. Send the manifest by a different route than the data, or sign it, and recompute every checksum on arrival; the intake checklist for data engineers covers the next steps.

Transfer channels compared: push, pull, share or ship

Choose the channel by volume, by which side should control storage and cost, and by whether you need a physical copy at all. Shares avoid copying; shipment suits volumes the network cannot move in time.

ChannelHow it worksWatch for
Push to your bucketSupplier writes to a prefix you own, with a credential you issue; you control retention and deletionLimit write access to one prefix and revoke it after the final release
Pull from the supplier's bucketThe owner's bucket policy grants your account specific actions, your administrator delegates them, and an explicit deny overrides any allow [17]Under S3 Requester Pays you pay requests and downloads; every request must be authenticated and carry x-amz-request-payer or RequestPayer [18]
Warehouse or table shareSnowflake Secure Data Sharing copies no data, and shared objects are read-only [19]. Delta Sharing is an open REST protocol over Delta Lake and Parquet with a provider-run sharing server [20]Open-protocol recipients authenticate with a bearer token that may expire [21]; check whether the license lets you materialize a training copy
SFTPSupplier uploads to a server you run, or you pull from theirsFine for gigabytes; slow and restart-prone for large media sets
Physical shipmentEncrypted drives or a cloud transfer service, with chain of custody for devices and keysAWS Snowball Edge is no longer available to new customers; AWS points to DataSync, Data Transfer Terminal or partners [22]
API or streaming feedSupplier publishes records continuouslyReplay and backfill rules; see API feeds vs batch files

For purchases SourceX manages, delivery runs through private, access-controlled workflows, never email attachments; see its answer on how licensed data is delivered. Detailed pages cover zero-copy warehouse sharing, SFTP limits, multi-terabyte network transfers, encrypted drives, offline appliances and egress costs.

Keeping licensed data controlled after it lands

Licensed data stays under license after it lands, so the receiving side needs keys handled apart from the data, access limited to named roles, a log of every read, and a record of which models used which release.

  • Encryption. Encrypt in transit and at rest, agree a file-level method such as OpenPGP, age or a cloud KMS key shared across accounts, and send keys by a different channel than files. See encryption and key exchange and SourceX's answer on encryption in transit.
  • Landing zone. Land each release in a quarantine prefix, verify the manifest, then promote it to a restricted location; audit logging covers what to record.
  • Personal data. For data sourced through SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded per dataset, and a sample is checked after processing. No method is perfect, so keep re-identification controls; the de-identified data hub covers residual risk.
  • Lineage. Record the release ID in every training and evaluation run, or emit it as OpenLineage events, so dataset-to-model traceability holds up to a deletion request or license audit.
  • End of term. Deletion must reach derived shards, embeddings, feature stores, caches and backups. Model retention after termination is a contract question for the licensing hub.

Recurring purchases: increments, corrections and schema changes

Ongoing purchases usually fail on identity and change control rather than on format: records need IDs that stay stable across releases, deletions need their own channel, and schema changes need notice and a version number.

Increments need a stable key, a change type on each record (insert, update or delete) and a high-water mark such as last-modified time. Pseudonymized IDs must stay consistent across releases or increments cannot join to earlier data; see stable record IDs and propagating deletions and corrections.

Treat schema changes like API changes: additive fields raise a minor version, while renamed, retyped or removed fields need notice and a major version. Data contracts turn those rules into checks, and schema change handling covers migrations. Delta Sharing supports time-travel queries on shared tables [21]; for copies you hold, see versioning tools such as DVC and lakeFS.

Delivery mistakes that surface after the data arrives

Most delivery problems surface in the first training or evaluation job, after the files were accepted, so name these failure modes in the specification.

  • Timestamps without a UTC offset, mixing local times across sites; see timestamps and time zones in delivered datasets.
  • CSV exports with locale-specific decimal separators or date formats.
  • Attachments referenced by ID in the records but missing from the package.
  • Image metadata stripped for privacy along with the orientation flag, leaving rotated images.
  • Narrowband telephone audio upsampled to match wideband recordings, with no field saying so.
  • Checksums computed before compression while the manifest lists compressed files.
  • License fields dropped when files are repackaged or converted; audits of hosted datasets find licenses often missing [14].
  • Training from a read-only warehouse share without a written right to copy.

SourceX's guide on how to procure enterprise training data places delivery in the buying sequence; the procurement hub and AI data overview cover the surrounding steps.

Planning how a licensed dataset reaches your team

If you need operational data from US companies, you can describe the records you need on SourceX's buyer page, along with your format, metadata and delivery requirements. SourceX looks for US businesses that hold the data, checks the data and the supplier's licensing permissions, and manages the license and the coordination of delivery and payment. Datasets are sourced on request, so a request does not guarantee a matching dataset. Describe your format and delivery needs.

Guides in this section

Sources

  1. OCEL standard authors (ocel-standard.org), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
  2. MISP 2022 challenge organizers, "The Multimodal Information based Speech Processing (MISP) 2022 Challenge: Audio-Visual Diarization and Recognition" (2023). https://arxiv.org/pdf/2303.06326
  3. CVAT.ai, "COCO (CVAT format documentation)". https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
  4. Grauman et al. (Ego4D consortium), "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1
  5. jsonlines.org, "JSON Lines". https://jsonlines.org/
  6. The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
  7. The Apache Software Foundation (Apache Parquet project), "Apache Parquet Documentation". https://parquet.apache.org/docs
  8. WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
  9. Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
  10. Gebru et al., "Datasheets for Datasets" (2018; Communications of the ACM 2021). https://arxiv.org/pdf/1803.09010
  11. Akhtar et al. (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  12. Jain et al. (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  13. NeurIPS blog, "Responsible AI metadata requirements for the Evaluations and Datasets Track NeurIPS 2026" (2026). https://blog.neurips.cc/?p=1527
  14. Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (arXiv 2023; journal version in Nature Machine Intelligence 6, 2024). https://arxiv.org/abs/2310.16787
  15. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  16. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  17. Amazon Web Services, "Example 2: Bucket owner granting cross-account bucket permissions" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-walkthroughs-managing-access-example2.html
  18. Amazon Web Services, "Using Requester Pays general purpose buckets for storage transfers and usage" (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
  19. Snowflake Inc., "About Secure Data Sharing" (Snowflake Documentation). https://docs.snowflake.com/en/user-guide/data-sharing-intro.html
  20. Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021, vendor launch post). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  21. delta-io/delta-sharing documentation (rendered by Mintlify), "Delta Sharing protocol: REST APIs". https://www.mintlify.com/delta-io/delta-sharing/protocol/rest-apis
  22. Amazon Web Services, "AWS Snowball Edge availability change" (AWS Snowball Edge Developer Guide). https://docs.amazonaws.cn/en_us/snowball/latest/developer-guide/snowball-edge-availability-change.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data