Schemas, packaging and delivery
Using JSON Schema to Specify and Validate Dataset Deliveries
Quick answer
A JSON Schema for dataset validation is a versioned file, written against the 2020-12 draft, that states which fields every record must carry, their types, allowed values and patterns. You send it to the supplier before the first delivery, require each delivery to name the schema $id it conforms to, and run a streaming validator over every JSONL line on arrival. The output is a per-field error report that decides whether the batch is accepted, returned or quarantined.
By SourceX Editorial · Updated
- 01Your data spec
- 02Matched to businesses
- 03Rights and samples reviewed
- 04Licensed delivery
Why a schema file beats a field list in the license exhibit
A machine-readable schema removes the ambiguity that a prose field list leaves for the supplier's engineers to interpret. "Timestamp of ticket creation" can arrive as epoch milliseconds, a local time without an offset, or an ISO 8601 string with Z; a schema with "type": "string", "format": "date-time" and a pattern closes that gap before anyone writes an export job. The same file becomes the acceptance test, so the supplier can run it before shipping and you run it again on receipt.
Keep the schema separate from the data dictionary. The data dictionary template explains meaning, provenance and business rules for humans; the schema encodes only what a validator can check. Dataset-level metadata such as Croissant, a schema.org-based JSON-LD vocabulary for files and record structure, can point at both, but it is not a substitute for record-level validation [5].
The 2020-12 keywords that matter for training records
Draft 2020-12 is the current JSON Schema release, published in two parts: Core, which defines $id, $ref, $defs and the vocabulary system, and Validation, which defines the assertion keywords [1]. For dataset records, a small set of keywords does almost all the work.
type:string,integer,number,boolean,object,array,null. Note that the 2020-12 spec treats any number with a zero fractional part as aninteger, so1.0passes.- Nullability through type arrays:
"type": ["string", "null"]. JSON Schema has nonullablekeyword; that is an OpenAPI 3.0 extension, and 2020-12 validators ignore it as an unknown keyword, so a supplier who copies it from an API spec gets a schema that silently rejects the nulls it meant to allow. required: lists keys that must be present. Presence is not non-emptiness, so pair it withminLength: 1orminItems: 1.enumandconst: close the value set for fields likechannel,roleorlabel. Add new values by schema version bump, never by tolerance.pattern: ECMA-262 regular expressions, unanchored by default. Write^...$explicitly orTKT-123will matchxxTKT-123yy.format:date-time,email,uri,uuid. In 2020-12,formatis an annotation unless the validator is configured to assert it; the python-jsonschema library, for example, does no format checking unless you pass aFormatChecker, and some formats need optional dependencies [3].additionalProperties: falseorunevaluatedProperties: false: rejects undeclared keys. UseunevaluatedPropertieswhen the schema composes subschemas withallOforif/then, becauseadditionalPropertiescannot see properties declared in sibling subschemas.prefixItems(new in 2020-12, replacing array-formitems),minItems,maxItems,uniqueItems: shape arrays such as chatmessagesor label lists.if/then/elseanddependentRequired: conditional rules, such as "ifroleistool, thentool_call_idis required."
Declare the dialect with "$schema": "https://json-schema.org/draft/2020-12/schema" at the top. Validators pick their rule set from that line, and a missing or older $schema value is a common reason two parties get different results from the same file.
An illustrative record schema for a support-conversation delivery
The schema below is the contract artifact itself, written for a JSONL delivery of support conversations intended for SFT and evaluation. It shows versioning through $id, closed enums, pseudonymized identifiers with anchored patterns, and a conditional rule.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://schemas.example.com/support-conversation/2.1.0/record.schema.json",
"title": "support_conversation_record",
"type": "object",
"required": ["record_id", "schema_version", "created_at", "channel", "messages", "resolution"],
"unevaluatedProperties": false,
"properties": {
"record_id": { "type": "string", "pattern": "^conv_[0-9a-f]{32}$" },
"schema_version": { "const": "2.1.0" },
"created_at": { "type": "string", "format": "date-time", "pattern": "Z$" },
"channel": { "enum": ["email", "chat", "phone_transcript"] },
"product_area": { "type": ["string", "null"], "maxLength": 64 },
"messages": {
"type": "array", "minItems": 2, "maxItems": 400,
"items": { "$ref": "#/$defs/message" }
},
"resolution": { "enum": ["resolved", "escalated", "abandoned"] },
"redaction_method": { "enum": ["placeholder_tokens", "synthetic_replacement"] }
},
"$defs": {
"message": {
"type": "object",
"required": ["turn", "role", "text"],
"unevaluatedProperties": false,
"properties": {
"turn": { "type": "integer", "minimum": 0 },
"role": { "enum": ["customer", "agent", "system"] },
"speaker_id": { "type": ["string", "null"], "pattern": "^spk_[0-9a-f]{12}$" },
"text": { "type": "string", "minLength": 1, "maxLength": 20000 }
},
"if": { "properties": { "role": { "const": "agent" } } },
"then": { "required": ["speaker_id"] }
}
}
}
The schema_version field repeats the version inside every record, so a line copied out of its file still identifies its contract. If your records need a richer conversation layout with tool calls and loss masking, start from the conversation transcript delivery schema and the chat fine-tuning data format guide.
Versioning the schema with $id and shipping it beside the data
Every delivery should carry the exact schema it claims to meet, identified by an immutable $id that includes a semantic version. Put the schema file in the delivery root next to the manifest, list it with its own checksum, and reject a delivery whose manifest names a schema $id you have not approved. The manifest and checksum guide covers the file-level half of that check.
Use version numbers that signal compatibility. A patch bump fixes descriptions; a minor bump adds optional fields or enum values that existing consumers can ignore; a major bump removes or renames fields, tightens a type, or makes a field required. How to roll those changes across a recurring feed is a separate problem covered in handling schema changes across recurring deliveries; this page only requires that each batch names one version.
Validating large JSONL as a stream with per-field error counts
JSONL validation should read one line at a time, validate each object independently, and aggregate every error by field path rather than stopping at the first failure. JSON Lines requires UTF-8 with no byte order mark and a valid JSON value on every line, and a blank line is invalid, so decode and parse failures are their own error class before schema checks run [2]. In python-jsonschema, iter_errors lazily yields all errors for an instance, which is what makes counting by path possible [3].
Illustrative example: invented to show structure; it does not describe an available dataset.
import json, collections
from jsonschema import Draft202012Validator, FormatChecker
schema = json.load(open("record.schema.json"))
Draft202012Validator.check_schema(schema)
v = Draft202012Validator(schema, format_checker=FormatChecker())
counts, samples, bad_lines, total = collections.Counter(), {}, 0, 0
with open("part-0001.jsonl", "rb") as f:
for n, raw in enumerate(f, 1):
total += 1
try:
rec = json.loads(raw.decode("utf-8"))
except (UnicodeDecodeError, json.JSONDecodeError):
counts[("$line", "parse")] += 1; bad_lines += 1; continue
errs = list(v.iter_errors(rec))
bad_lines += bool(errs)
for e in errs:
key = ("/".join(map(str, e.absolute_path)) or "$root", e.validator)
counts[key] += 1
samples.setdefault(key, n)
Normalize array indexes out of paths (messages/17/role becomes messages/*/role) before reporting, or one systematic defect fans out into hundreds of rows. Run check_schema first so a malformed schema fails loudly instead of passing every record. For multi-gigabyte files, shard by file and run validators in parallel processes; the work is embarrassingly parallel because records are independent.
A useful report lists, per path and keyword, the error count, rate, and first offending line number:
Illustrative example: invented to show structure; it does not describe an available dataset.
| Path | Keyword | Errors | Rate | First line |
|---|---|---|---|---|
created_at | pattern | 3,412 | 2.8% | 118 |
messages/*/role | enum | 41 | 0.03% | 9,950 |
$root | unevaluatedProperties | 120,004 | 100% | 1 |
$line | parse | 2 | <0.01% | 77,301 |
The first and third rows show typical patterns: a 100% unevaluatedProperties rate usually means the supplier added an internal column such as _export_ts, not that every record is broken. A 2.8% pattern failure on timestamps usually means one source system emits offsets instead of Z.
Acceptance rules that turn validation results into a decision
Write the accept, return and quarantine thresholds into the delivery spec before the first file arrives, so the validator output maps to an action without negotiation. A reasonable starting rule set:
- Reject the batch on any
$line/parseerror above a stated tolerance, any schema$idmismatch, or anyrequiredfailure on identifier fields. - Quarantine records that fail
enum,patternorformatbelow an agreed rate, and load the rest. - Return to supplier with the report table and sample line numbers when any single path exceeds its threshold.
- Log, do not block on annotation-only keywords such as
descriptionorexamples.
For sampling plans that decide how many records a human must inspect after the schema passes, see acceptance sampling for dataset deliveries.
Mapping JSON Schema types to Parquet and Arrow on ingest
Teams that convert JSONL to Parquet on ingest should derive the Arrow schema from the JSON Schema rather than letting a reader infer it from the first rows. Inference picks int64 for a field that later holds a float, or null type for a column that happens to be empty in the first block. The Parquet documentation and its separate parquet-format specification define the physical and logical types you are mapping into [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
| JSON Schema | Arrow type | Parquet logical type | Watch for |
|---|---|---|---|
string | string / large_string | STRING | Long text over 2 GB per batch needs large_string |
string, format: date-time | timestamp[us, tz=UTC] | TIMESTAMP (UTC adjusted) | Offsets must be normalized first |
integer | int64 | INT(64) | 1.0 passes JSON Schema but fails a strict cast |
number | float64 | DOUBLE | Monetary values belong in decimal strings |
enum of strings | dictionary<int32, string> | STRING, dictionary encoded | New values need a version bump |
["T", "null"] | nullable T | OPTIONAL | Arrow nullability is per field, not per value |
array of objects | list<struct<...>> | LIST of group | unevaluatedProperties keeps structs stable |
Which format to ship in the first place is a separate decision, covered in Parquet vs JSONL for licensed training data and the broader question of what file formats AI buyers accept.
Where structural validation stops and quality checks begin
A passing schema proves the records are well-formed, not that they are useful, deduplicated or correctly labeled. Missingness rates on optional fields, range plausibility, referential integrity between files and label accuracy belong to the quality layer described in validation checks for structured dataset deliveries. ISO/IEC 5259-3 frames that broader work as a data quality management process for ML data rather than a fixed list of metrics, which is a useful reminder that a schema is one control within it [6].
Stable identifiers deserve both layers. The schema enforces the record_id pattern; the quality layer checks uniqueness within a batch and continuity across batches, as described in stable record IDs and join keys.
Agreeing a dataset schema with a supplier
Send the schema with your request, before pricing, because the export a supplier can produce determines what the license should describe. A practical sequence:
- Draft the schema from your training or eval loader, not from the supplier's source tables.
- Ask the supplier to run it against a small, approved sample and return the error report.
- Resolve each failing path by changing the export, relaxing the schema with a version bump, or dropping the field.
- Freeze the
$id, reference it in the delivery terms, and require every batch manifest to name it.
When you work with SourceX, you describe the data you need and SourceX looks for US businesses that hold it; every release is approved by the supplying company, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded, which is why a field like redaction_method belongs in your schema. You can start a buyer request with SourceX once your schema draft is ready. For the rest of the delivery stack, the dataset delivery hub and the AI data guide index link every related page.
Sourcing JSONL records that meet your schema
SourceX sources operational datasets, such as support and sales histories, engineering records, documents and finance and legal workflows, from US companies on request; categories are not inventory, and a request does not guarantee a match. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the records and schema you need.
Sources
- JSON Schema (json-schema.org), "JSON Schema Specification". https://json-schema.org/specification
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Tessl registry (python-jsonschema docs mirror), "jsonschema 4.25.0 validators documentation". https://tessl.io/registry/tessl/pypi-jsonschema/4.25.0/files/docs/validators.md
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
- Akhtar et al., MLCommons, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and ML, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.