Schemas, packaging and delivery
Receiving a Licensed Dataset: Intake Checklist for Data Engineers
Quick answer
A licensed dataset intake checklist moves a new delivery from untrusted bytes to a governed, cataloged asset in a fixed order: land it in an isolated quarantine location, scan it for malware and unsafe archives, reconcile checksums and record counts against the supplier manifest, validate formats and schema against the data dictionary, spot-check de-identification, then register it in your catalog with the license ID, version and permitted uses before granting access. Record the handover time and pass it to quality acceptance.
By SourceX Editorial · Updated
This page covers technical intake only. Pre-purchase checks belong in the training data due diligence checklist, and statistical quality acceptance belongs to your quality process; intake is the gate between them. The broader context for formats and transfer patterns sits in the delivery hub for licensed AI data.
Why licensed data needs its own intake gate
Licensed data needs a dedicated intake gate because the license, not your general data policy, defines who may touch it, for what purpose and for how long. A file that lands in a shared lake before anyone records its license ID becomes impossible to scope later: it gets copied into feature tables, sampled into notebooks and pulled into a fine-tuning mix without a trail.
Research repositories already treat intake as a formal step, collecting data types, data-sharing documentation and use conditions before data is accepted [9]. Buyers should apply the same discipline in reverse. ISO/IEC 5259-3 frames data quality for ML as a managed process with defined requirements rather than an ad hoc check [6], and an intake gate is where that process starts for third-party data.
Stage 1: Land the delivery in quarantine
The first step is to land every delivery in an isolated quarantine location that no training job, notebook or BI tool can read. In practice that means a dedicated bucket or prefix (for example s3://intake-quarantine/<supplier>/<delivery-id>/) with write access for the transfer mechanism and read access only for the intake pipeline's service role.
For cross-account cloud handovers, the owning account grants specific operations through a bucket policy and the receiving account delegates them through its own IAM policy [1]. Scope both sides to the delivery prefix, and prefer pulling into your quarantine bucket over letting the supplier write directly into production storage. Patterns and trade-offs are covered in cross-account bucket delivery for licensed datasets.
Quarantine controls worth setting before the first byte arrives:
- Block public access, enable default encryption with a key your team controls, and turn on object-level access logging.
- Deny
s3:GetObject(or the equivalent in GCS or Azure Blob) to every principal except the intake role. - Set a lifecycle rule that deletes unpromoted quarantine objects after a fixed window, so rejected deliveries do not linger.
- If the delivery arrives encrypted with PGP or age, decrypt inside the quarantine boundary only; key handling is covered in encrypting dataset deliveries.
Stage 2: Scan for malware and unsafe archives
Malware scanning belongs in every intake because licensed operational data often includes attachments, exported documents and archives that can carry executable content. Support ticket exports, email archives and document repositories routinely contain .docm, .xlsm, PDFs with embedded JavaScript, and nested ZIPs.
Run a signature scanner such as ClamAV, or your cloud provider's object malware scanning, across every object, including members of archives. Add structural checks that signature engines miss:
- Archive safety: reject entries with absolute paths or
../segments (the "zip slip" pattern), cap decompression ratios to catch zip bombs, and cap nesting depth. - Type truth: compare file magic bytes against extensions; a
.csvthat starts withPKis a ZIP, and an.xlsxwith a VBA project part is macro-enabled. - Active content policy: decide in advance whether macro-enabled Office files are converted to text, stripped, or held for manual review. Never open them on an analyst workstation.
Log scanner name, signature database version and verdict per object. If anything is flagged, stop the pipeline, keep the delivery in quarantine and notify the supplier through the agreed channel rather than deleting evidence.
Stage 3: Reconcile the manifest, checksums and counts
Manifest reconciliation proves that you received exactly the files the supplier sent, byte for byte, and nothing extra. The supplier manifest should list every path, size in bytes, a SHA-256 digest and, for structured data, a record count; a sample manifest shows the shape to request.
Recompute digests independently rather than trusting transfer-layer metadata. Amazon S3 can compute and store additional checksums such as SHA-256 or CRC32C on upload [2], which is useful, but a multipart upload's ETag is not an MD5 of the whole file, so compare against the manifest digest you computed yourself. The full procedure, including how to handle split files, is in dataset manifests and checksum verification.
Reconcile four things and record each result: files in the manifest but missing on disk, files on disk but absent from the manifest, digest mismatches, and record-count deltas. Count records with the same parser you will use downstream, because a naive line count on JSONL with embedded newlines or CSV with quoted line breaks will disagree with the supplier's figure.
Stage 4: Validate formats and schema against the data dictionary
Format and schema validation confirms that each file parses as the declared format and matches the contracted data dictionary. Check physical validity first, then logical structure.
For Parquet, a valid file begins and ends with the 4-byte magic PAR1 and carries its metadata in the footer [3], so a truncated upload usually fails at footer read. For JSON Lines, every line must be a valid JSON value in UTF-8 [4]; reject files with a byte-order mark, mixed encodings or trailing partial lines. For CSV, pin delimiter, quote character, encoding and header row in the delivery specification.
Then compare the observed schema with the data dictionary for the delivery: column names, types, nullability, enumerations, key uniqueness and foreign-key coverage between linked tables. If the supplier provides Croissant metadata, a JSON-LD vocabulary built on schema.org that describes file resources and record structure [5], you can drive these checks from it directly; see Croissant metadata for licensed datasets. On recurring feeds, treat any added, dropped or retyped field as a schema-change event under your schema evolution process, not a silent pass.
Stage 5: Spot-check de-identification before anyone browses
A de-identification spot check confirms the supplier's stated method held on your copy before analysts or training jobs see the data. Run pattern detectors (emails, phone numbers, card and account numbers, government IDs) and a named-entity pass over free-text fields, attachment text and file metadata such as PDF author fields and DICOM headers.
No method catches everything. The DICOM standard itself notes that its confidentiality profiles do not guarantee removal of all identifying information and are only one part of a full process [7]. Log hit rates by field, route true positives back to the supplier, and do not attempt to reverse pseudonyms or tokens. Under California's definition, deidentified status depends partly on the holder committing not to reidentify and on contractual obligations passed to recipients [8], so re-identification attempts at intake can undermine the basis on which you received the data.
Stage 6: Register in the catalog with license metadata
Catalog registration turns the delivery into a governed asset by attaching license and lineage metadata that every downstream user and job inherits. Whether you use Unity Catalog, DataHub, AWS Glue Data Catalog or an internal registry, the entry should carry license fields as first-class attributes, not free-text descriptions.
Illustrative example: invented to show structure; it does not describe an available dataset.
dataset_id: lic-support-tickets-2026q3
delivery_id: dlv-0007
version: 1.2.0
license_id: LIC-2026-014
license_term_end: 2028-06-30
permitted_uses: [fine_tuning, evaluation]
prohibited_uses: [redistribution, re_identification]
source_supplier_ref: supplier-ref-A
received_at: 2026-10-02T14:07:31Z
manifest_sha256: 9f2c...e41a
record_count: 412530
scan: {engine: clamav, db_version: "27410", verdict: clean}
schema_check: pass
deid_spotcheck: {sample_size: 2000, residual_hits: 3, status: returned_to_supplier}
owner: data-platform@buyer.example
access_group: grp-lic-2026-014-readers
status: promoted_pending_quality
Bind access to the license, not the team: create a group per license ID and grant read only to named projects whose purpose matches permitted_uses. Implementation options are in access controls for licensed training data, and the dataset_id plus version pair is what lets you later answer which models trained on which licensed dataset.
Stage 7: Record the handover and pass to quality acceptance
The final intake step records an auditable handover and passes the promoted copy to quality acceptance. Capture the received timestamp, the person or service that promoted the data, every check result and any open exceptions, and store the record alongside the catalog entry.
Promotion should be a copy or atomic move from quarantine to a versioned, read-only location, never an in-place rename that leaves writers attached. Then hand off to statistical acceptance: completeness, duplicates, label agreement and distribution checks. If you are still evaluating a supplier, align these gates with the acceptance criteria in running a data pilot with a supplier.
The intake checklist at a glance
The checklist below condenses each stage into a gate with an evidence artifact you can store with the catalog entry.
Illustrative example: invented to show structure; it does not describe an available dataset.
| # | Gate | Pass condition | Evidence to keep |
|---|---|---|---|
| 1 | Quarantine landing | Objects only in isolated prefix; intake role is sole reader | Bucket policy version, access logs |
| 2 | Malware and archive scan | No detections; no path traversal; ratios under cap | Engine, signature version, per-object verdicts |
| 3 | Manifest reconciliation | No missing, extra or mismatched files; counts match | Recomputed SHA-256 list, diff report |
| 4 | Format and schema | All files parse; schema matches data dictionary | Validator output, schema diff |
| 5 | De-identification spot check | Residual hits within agreed tolerance or returned | Sample size, detector hits by field |
| 6 | Catalog registration | License ID, version, permitted uses, owner set | Catalog entry ID |
| 7 | Handover and promotion | Read-only versioned copy; quality ticket opened | Handover record, promotion timestamp |
How intake connects to sourcing
Intake is far simpler when the data arrives with documented rights, a manifest and a recorded de-identification method. SourceX sources operational datasets from US companies on request and manages the licensing process; every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval, and personal details such as names, emails, phones and account numbers are removed or replaced first, with the method recorded and a sample checked. You can describe the data your team needs, and more on documenting source and rights is in the AI data provenance guide.
Sourcing licensed datasets ready for intake
If your team needs licensed operational data such as support histories, engineering records or document workflows, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. A request does not guarantee a match, and nothing is contracted until a supplier agrees. Describe the licensed dataset you need.
Sources
- Amazon Web Services (Amazon S3 User Guide), "Example 2: Bucket owner granting cross-account bucket permissions". https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-walkthroughs-managing-access-example2.html
- Amazon Web Services (Amazon S3 User Guide), "Checking object integrity in Amazon S3". https://docs.aws.amazon.com/AmazonS3/latest/userguide/checking-object-integrity.html
- The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Akhtar et al. (MLCommons Croissant working group), arXiv:2403.19546, "Croissant: A Metadata Format for ML-Ready Datasets". https://arxiv.org/pdf/2403.19546
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 3: Data quality management requirements and guidelines". https://www.iso.org/standard/81092.html
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Annex E: Attribute Confidentiality Profiles". https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- AD Knowledge Portal, "Complete Data Intake". https://help.adknowledgeportal.org/apd/complete-data-intake
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.