Schemas, packaging and delivery
Parquet vs JSONL for Licensed Training Data Deliveries
Quick answer
Ask for Parquet when a licensed dataset is large, typed and recurring: tabular operational records, ticket histories with fixed fields, or pre-tokenized corpora you will filter by column. Ask for JSONL when records are irregular, need human review, or feed a fine-tuning API that expects one chat example per line. Many buyers want both: a small JSONL sample for inspection and Parquet for the full delivery, described by one data dictionary and validated against a pinned schema version.
By SourceX Editorial · Updated
This page covers the head-to-head trade-offs and the decision rules. For the short list of formats buyers commonly accept, see what file formats AI buyers accept. For the rest of the packaging stack, start at the dataset delivery hub.
How JSONL and Parquet differ on disk
JSONL is a row-oriented text format with no schema, while Parquet is a column-oriented binary format whose schema travels inside the file. A JSON Lines file is UTF-8 without a byte order mark, and each line holds exactly one valid JSON value; a blank line is not valid [1]. JSON Lines is a community convention rather than an IETF or ISO standard, so "JSONL" in a contract should point to the jsonlines.org rules explicitly [1].
A Parquet file starts and ends with the four-byte magic number PAR1, and its metadata sits in a footer written after the data so a writer can produce the file in a single pass [2]. That footer records where every column chunk starts, which is what lets readers such as PyArrow, DuckDB, Spark or Polars read only the columns and row groups a query needs [2]. Logical types, encodings and compression codecs are defined in the separate parquet-format specification, and support for newer features varies by library [3].
The practical consequences for a buyer:
- Inspection. You can
head -n 5a JSONL file or open it in any editor. Parquet needs a reader, thoughduckdb -c "DESCRIBE 'file.parquet'"orparquet-tools schemamakes that quick. - Types. In JSONL,
"amount": "12.50"in one row and"amount": 12.5in the next is legal. Parquet fixes the column asDECIMAL(12,2)orDOUBLEfor the whole file. - Appending. JSONL appends by concatenation. Parquet files are immutable once written; new data arrives as new files.
- Partial reads. Reading one field from JSONL means parsing every line in full. Parquet lets you read
ticket_idandresolution_codewithout decompressing the longbodycolumn.
When JSONL is the better delivery format
JSONL wins when the consumer is a fine-tuning endpoint, a human reviewer, or a pipeline that needs irregular nested records. Fine-tuning services commonly take chat and instruction data as JSONL with one example per line; Together AI, for example, documents JSONL for conversational and instruction data [4]. If the licensed data will go straight into such an endpoint, asking the supplier for Parquet only adds a conversion step back to JSONL.
JSONL also fits records whose structure varies legitimately from row to row, such as engineering incident timelines where some events carry attachments, some carry diffs and some carry neither [6]. Reviewers in counsel, privacy and data-quality roles can read a JSONL sample directly, which matters when you are checking that names, emails and account numbers were replaced before signing off on a full delivery.
The cost is drift. Without an enforced schema, a supplier's export job can silently change created_at from ISO 8601 strings to epoch milliseconds, or start emitting null where it used to omit the key. Each change is valid JSON and breaks your loader on the next delivery.
When Parquet is the better delivery format
Parquet wins for large, wide or repeatedly delivered datasets where types must hold and you will filter or join before training. Support histories, sales activity logs, finance and legal workflow records are typically dozens of fixed fields plus one or two long text fields; Parquet stores repetitive categorical columns like status, queue or product_line efficiently with dictionary encoding and compresses each column independently [3]. You can then prune to the columns a given experiment needs.
Parquet is also the native format of most lakehouse and sharing systems. Open sharing protocols such as Delta Sharing move large tables as Parquet on S3, ADLS or GCS [7], so if your team plans a zero-copy warehouse share or a cross-account cloud bucket delivery, Parquet avoids a translation layer.
For training pipelines, Parquet is the usual home for pre-tokenized data. Together AI documents Parquet for pre-tokenized inputs with custom loss masks, and notes that format choice affects how examples are packed [4]. Treat that as a reason to state the target format in the delivery spec rather than leave it to the supplier. If you are requesting licensed operational records, name the format and schema in the request so they can be written into the license's delivery terms. For why typed operational fields behave differently from free text, see structured vs unstructured training data.
Conversations and nested records: nest or flatten
Parquet can store a conversation as a LIST of STRUCT values, so one row per conversation with a nested messages column is usually the right default. Flatten to one row per message when you need to filter, join or sample at the message level, for example to drop all agent turns from one queue or to join messages to a separate CSAT table.
A nested schema for a support conversation looks like this in Arrow notation:
Illustrative example: invented to show structure; it does not describe an available dataset.
conversation_id: string (not null)
channel: dictionary<string>
opened_at: timestamp[us, tz=UTC]
messages: list<struct<
turn: int32,
role: string, -- "customer" | "agent" | "system"
sent_at: timestamp[us, tz=UTC],
text: string,
redactions: list<struct<start: int32, end: int32, label: string>>
>>
Three failure modes show up when suppliers export nested data:
- Stringified JSON. The exporter writes
messagesas onestringcolumn holding serialized JSON. It looks nested in a preview but loses all typing and column statistics. Reject it in the spec. - Inconsistent struct fields. One file has
messages.attachments, the next does not. Schema merging may pass on read but leaves silent nulls; handle it through your schema evolution policy. - Lost ordering. Flattened message tables without an explicit
turninteger rely on file order, which Spark and other parallel writers do not guarantee.
If you flatten, require a stable parent key on every row. The stable record IDs guide covers how to keep conversation_id consistent across refreshes.
Compression, sharding and file size
Compress JSONL with zstd or gzip and shard it; compress Parquet internally per column and size row groups for your readers. A single multi-gigabyte .jsonl.gz file is the most common delivery mistake: gzip streams are not splittable, so one worker must decompress the whole file sequentially and the rest of the cluster waits.
Practical defaults to write into a spec:
- JSONL:
part-00000.jsonl.zstshards of a size your workers can each handle, newline-terminated final line, no blank lines [1]. - Parquet: zstd or snappy codec, row groups sized so a reader can skip irrelevant ones, and an explicit statement of which codec and Parquet writer version produced the files [3].
- Both: a manifest with per-file checksums and row counts, as covered in dataset manifests and checksums.
Do not judge compression from one sample. Repetitive operational fields compress well in Parquet's columnar layout, while long free-text columns compress about as well in either format because the text dominates the bytes. Measure on a representative shard before choosing.
Validating each format against a pinned schema
Validate JSONL line by line with JSON Schema and validate Parquet by comparing its embedded schema to an expected schema, then pin both to the dataset version. JSON Schema draft 2020-12 is the current release, split into Core and Validation documents [5]; use type, required, enum, format and additionalProperties: false so unexpected keys fail loudly. The JSON Schema validation guide has worked examples.
For Parquet, the footer schema is authoritative [2]. Read it with pyarrow.parquet.read_schema() and diff it against a committed expected schema, checking logical types (timestamp unit and timezone, decimal precision and scale), nullability and field order. Then run value-level checks the schema cannot express: allowed role values, non-empty text, monotonic turn.
Pin the schema file and the data dictionary to the same version ID as the data. A data dictionary that describes opened_at as local time while the Parquet type says UTC is a defect, not a documentation nit. Machine-readable metadata such as MLCommons Croissant can describe the files and record structure of either format in one JSON-LD document [8].
Converting JSONL to Parquet without losing fidelity
Conversion is safe only when you declare the target schema up front instead of letting the library infer it from the first rows. Inferred types depend on what the library sees first (PyArrow's JSON reader, for example, infers from the first block it reads), so a field that is an integer for 10,000 rows and a string in row 10,001 either fails late or gets coerced.
Known traps when converting a supplier's JSONL:
- Large integers. Account-style IDs above 2^53 lose precision if any tool in the chain parses them as JSON numbers into doubles. Deliver IDs as strings.
- Null vs missing. JSONL distinguishes
"field": nullfrom an absent key; Parquet stores both as null. If the difference carries meaning, add an explicit flag column. - Timestamps. Mixed
Z, offset and naive strings must be normalized to one unit and timezone before writingtimestamp[us, tz=UTC]. - Duplicates. Conversion is a good moment to hash
textand count exact duplicates; training corpora routinely contain near-duplicate records that affect model behavior [9].
Run the reverse check too: convert back to JSONL and compare record counts and a sample of field values.
Decision table and delivery spec template
The choice follows from the consumer, the size and the stability of the schema. Use the table to decide, then put the result in the license's delivery terms so both sides test against the same definition.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Situation | Require | Why |
|---|---|---|
| Chat or instruction examples going to a hosted fine-tuning API | JSONL | Matches the endpoint's expected input [4] |
| Fixed-field operational records, recurring monthly | Parquet | Enforced types, column pruning, cheaper refreshes |
| Pre-tokenized corpus with loss masks | Parquet | Typed arrays, documented by fine-tuning services [4] |
| Rights and privacy review sample | JSONL | Readable without tooling |
| Irregular nested event logs | JSONL, or Parquet with explicit STRUCT schema | Avoids forcing sparse columns |
| Warehouse or lakehouse share | Parquet | Native to sharing protocols [7] |
Illustrative example: invented to show structure; it does not describe an available dataset.
Delivery format spec: support-conversations v2026.10
Full delivery: Parquet, zstd, one row per conversation, schema file conv_v3.arrow.json
Review sample: JSONL (jsonlines.org rules), 500 records, same fields, validated by conv_v3.schema.json (2020-12)
IDs: strings; conversation_id stable across refreshes
Timestamps: microseconds, UTC
Nested fields: messages as list<struct>; no stringified JSON
Sharding: files named part-NNNNN; manifest.json with sha256 and row counts
Reject if: schema diff, unknown enum value, empty text, broken turn order
Request licensed records in the format your pipeline needs
SourceX sources operational datasets such as support and sales histories, engineering records, documents and finance or legal workflows from US companies on request, and every dataset is delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows, never email attachments. Describe the data and the format your pipeline needs at SourceX for buyers.
Sources
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
- The Apache Software Foundation (Apache Parquet project), "Apache Parquet Documentation". https://parquet.apache.org/docs
- Together AI, "Fine-tuning data preparation". https://docs.together.ai/docs/fine-tuning-data-preparation.md
- JSON Schema, "JSON Schema Specification". https://json-schema.org/specification
- Hugging Face (community blog), "LLM Dataset Formats 101: A No-BS Guide for Hugging Face Devs". https://huggingface.co/blog/tegridydev/llm-dataset-formats-101-hugging-face
- Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
- Akhtar et al. (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.