Skip to content

Schemas, packaging and delivery

Parquet vs JSONL for Licensed Training Data Deliveries

Quick answer

Ask for Parquet when a licensed dataset is large, typed and recurring: tabular operational records, ticket histories with fixed fields, or pre-tokenized corpora you will filter by column. Ask for JSONL when records are irregular, need human review, or feed a fine-tuning API that expects one chat example per line. Many buyers want both: a small JSONL sample for inspection and Parquet for the full delivery, described by one data dictionary and validated against a pinned schema version.

By SourceX Editorial · Updated

This page covers the head-to-head trade-offs and the decision rules. For the short list of formats buyers commonly accept, see what file formats AI buyers accept. For the rest of the packaging stack, start at the dataset delivery hub.

How JSONL and Parquet differ on disk

JSONL is a row-oriented text format with no schema, while Parquet is a column-oriented binary format whose schema travels inside the file. A JSON Lines file is UTF-8 without a byte order mark, and each line holds exactly one valid JSON value; a blank line is not valid [1]. JSON Lines is a community convention rather than an IETF or ISO standard, so "JSONL" in a contract should point to the jsonlines.org rules explicitly [1].

A Parquet file starts and ends with the four-byte magic number PAR1, and its metadata sits in a footer written after the data so a writer can produce the file in a single pass [2]. That footer records where every column chunk starts, which is what lets readers such as PyArrow, DuckDB, Spark or Polars read only the columns and row groups a query needs [2]. Logical types, encodings and compression codecs are defined in the separate parquet-format specification, and support for newer features varies by library [3].

The practical consequences for a buyer:

  • Inspection. You can head -n 5 a JSONL file or open it in any editor. Parquet needs a reader, though duckdb -c "DESCRIBE 'file.parquet'" or parquet-tools schema makes that quick.
  • Types. In JSONL, "amount": "12.50" in one row and "amount": 12.5 in the next is legal. Parquet fixes the column as DECIMAL(12,2) or DOUBLE for the whole file.
  • Appending. JSONL appends by concatenation. Parquet files are immutable once written; new data arrives as new files.
  • Partial reads. Reading one field from JSONL means parsing every line in full. Parquet lets you read ticket_id and resolution_code without decompressing the long body column.

When JSONL is the better delivery format

JSONL wins when the consumer is a fine-tuning endpoint, a human reviewer, or a pipeline that needs irregular nested records. Fine-tuning services commonly take chat and instruction data as JSONL with one example per line; Together AI, for example, documents JSONL for conversational and instruction data [4]. If the licensed data will go straight into such an endpoint, asking the supplier for Parquet only adds a conversion step back to JSONL.

JSONL also fits records whose structure varies legitimately from row to row, such as engineering incident timelines where some events carry attachments, some carry diffs and some carry neither [6]. Reviewers in counsel, privacy and data-quality roles can read a JSONL sample directly, which matters when you are checking that names, emails and account numbers were replaced before signing off on a full delivery.

The cost is drift. Without an enforced schema, a supplier's export job can silently change created_at from ISO 8601 strings to epoch milliseconds, or start emitting null where it used to omit the key. Each change is valid JSON and breaks your loader on the next delivery.

When Parquet is the better delivery format

Parquet wins for large, wide or repeatedly delivered datasets where types must hold and you will filter or join before training. Support histories, sales activity logs, finance and legal workflow records are typically dozens of fixed fields plus one or two long text fields; Parquet stores repetitive categorical columns like status, queue or product_line efficiently with dictionary encoding and compresses each column independently [3]. You can then prune to the columns a given experiment needs.

Parquet is also the native format of most lakehouse and sharing systems. Open sharing protocols such as Delta Sharing move large tables as Parquet on S3, ADLS or GCS [7], so if your team plans a zero-copy warehouse share or a cross-account cloud bucket delivery, Parquet avoids a translation layer.

For training pipelines, Parquet is the usual home for pre-tokenized data. Together AI documents Parquet for pre-tokenized inputs with custom loss masks, and notes that format choice affects how examples are packed [4]. Treat that as a reason to state the target format in the delivery spec rather than leave it to the supplier. If you are requesting licensed operational records, name the format and schema in the request so they can be written into the license's delivery terms. For why typed operational fields behave differently from free text, see structured vs unstructured training data.

Conversations and nested records: nest or flatten

Parquet can store a conversation as a LIST of STRUCT values, so one row per conversation with a nested messages column is usually the right default. Flatten to one row per message when you need to filter, join or sample at the message level, for example to drop all agent turns from one queue or to join messages to a separate CSAT table.

A nested schema for a support conversation looks like this in Arrow notation:

Illustrative example: invented to show structure; it does not describe an available dataset.

conversation_id: string (not null)
channel: dictionary<string>
opened_at: timestamp[us, tz=UTC]
messages: list<struct<
    turn: int32,
    role: string,          -- "customer" | "agent" | "system"
    sent_at: timestamp[us, tz=UTC],
    text: string,
    redactions: list<struct<start: int32, end: int32, label: string>>
>>

Three failure modes show up when suppliers export nested data:

  1. Stringified JSON. The exporter writes messages as one string column holding serialized JSON. It looks nested in a preview but loses all typing and column statistics. Reject it in the spec.
  2. Inconsistent struct fields. One file has messages.attachments, the next does not. Schema merging may pass on read but leaves silent nulls; handle it through your schema evolution policy.
  3. Lost ordering. Flattened message tables without an explicit turn integer rely on file order, which Spark and other parallel writers do not guarantee.

If you flatten, require a stable parent key on every row. The stable record IDs guide covers how to keep conversation_id consistent across refreshes.

Compression, sharding and file size

Compress JSONL with zstd or gzip and shard it; compress Parquet internally per column and size row groups for your readers. A single multi-gigabyte .jsonl.gz file is the most common delivery mistake: gzip streams are not splittable, so one worker must decompress the whole file sequentially and the rest of the cluster waits.

Practical defaults to write into a spec:

  • JSONL: part-00000.jsonl.zst shards of a size your workers can each handle, newline-terminated final line, no blank lines [1].
  • Parquet: zstd or snappy codec, row groups sized so a reader can skip irrelevant ones, and an explicit statement of which codec and Parquet writer version produced the files [3].
  • Both: a manifest with per-file checksums and row counts, as covered in dataset manifests and checksums.

Do not judge compression from one sample. Repetitive operational fields compress well in Parquet's columnar layout, while long free-text columns compress about as well in either format because the text dominates the bytes. Measure on a representative shard before choosing.

Validating each format against a pinned schema

Validate JSONL line by line with JSON Schema and validate Parquet by comparing its embedded schema to an expected schema, then pin both to the dataset version. JSON Schema draft 2020-12 is the current release, split into Core and Validation documents [5]; use type, required, enum, format and additionalProperties: false so unexpected keys fail loudly. The JSON Schema validation guide has worked examples.

For Parquet, the footer schema is authoritative [2]. Read it with pyarrow.parquet.read_schema() and diff it against a committed expected schema, checking logical types (timestamp unit and timezone, decimal precision and scale), nullability and field order. Then run value-level checks the schema cannot express: allowed role values, non-empty text, monotonic turn.

Pin the schema file and the data dictionary to the same version ID as the data. A data dictionary that describes opened_at as local time while the Parquet type says UTC is a defect, not a documentation nit. Machine-readable metadata such as MLCommons Croissant can describe the files and record structure of either format in one JSON-LD document [8].

Converting JSONL to Parquet without losing fidelity

Conversion is safe only when you declare the target schema up front instead of letting the library infer it from the first rows. Inferred types depend on what the library sees first (PyArrow's JSON reader, for example, infers from the first block it reads), so a field that is an integer for 10,000 rows and a string in row 10,001 either fails late or gets coerced.

Known traps when converting a supplier's JSONL:

  • Large integers. Account-style IDs above 2^53 lose precision if any tool in the chain parses them as JSON numbers into doubles. Deliver IDs as strings.
  • Null vs missing. JSONL distinguishes "field": null from an absent key; Parquet stores both as null. If the difference carries meaning, add an explicit flag column.
  • Timestamps. Mixed Z, offset and naive strings must be normalized to one unit and timezone before writing timestamp[us, tz=UTC].
  • Duplicates. Conversion is a good moment to hash text and count exact duplicates; training corpora routinely contain near-duplicate records that affect model behavior [9].

Run the reverse check too: convert back to JSONL and compare record counts and a sample of field values.

Decision table and delivery spec template

The choice follows from the consumer, the size and the stability of the schema. Use the table to decide, then put the result in the license's delivery terms so both sides test against the same definition.

Illustrative example: invented to show structure; it does not describe an available dataset.

SituationRequireWhy
Chat or instruction examples going to a hosted fine-tuning APIJSONLMatches the endpoint's expected input [4]
Fixed-field operational records, recurring monthlyParquetEnforced types, column pruning, cheaper refreshes
Pre-tokenized corpus with loss masksParquetTyped arrays, documented by fine-tuning services [4]
Rights and privacy review sampleJSONLReadable without tooling
Irregular nested event logsJSONL, or Parquet with explicit STRUCT schemaAvoids forcing sparse columns
Warehouse or lakehouse shareParquetNative to sharing protocols [7]

Illustrative example: invented to show structure; it does not describe an available dataset.

Delivery format spec: support-conversations v2026.10
Full delivery:   Parquet, zstd, one row per conversation, schema file conv_v3.arrow.json
Review sample:   JSONL (jsonlines.org rules), 500 records, same fields, validated by conv_v3.schema.json (2020-12)
IDs:             strings; conversation_id stable across refreshes
Timestamps:      microseconds, UTC
Nested fields:   messages as list<struct>; no stringified JSON
Sharding:        files named part-NNNNN; manifest.json with sha256 and row counts
Reject if:       schema diff, unknown enum value, empty text, broken turn order

Request licensed records in the format your pipeline needs

SourceX sources operational datasets such as support and sales histories, engineering records, documents and finance or legal workflows from US companies on request, and every dataset is delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, and delivery runs through private, access-controlled workflows, never email attachments. Describe the data and the format your pipeline needs at SourceX for buyers.

Sources

  1. jsonlines.org, "JSON Lines". https://jsonlines.org/
  2. The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
  3. The Apache Software Foundation (Apache Parquet project), "Apache Parquet Documentation". https://parquet.apache.org/docs
  4. Together AI, "Fine-tuning data preparation". https://docs.together.ai/docs/fine-tuning-data-preparation.md
  5. JSON Schema, "JSON Schema Specification". https://json-schema.org/specification
  6. Hugging Face (community blog), "LLM Dataset Formats 101: A No-BS Guide for Hugging Face Devs". https://huggingface.co/blog/tegridydev/llm-dataset-formats-101-hugging-face
  7. Databricks, "Introducing Delta Sharing: An Open Protocol for Secure Data Sharing" (2021). https://www.databricks.com/blog/2021/05/26/introducing-delta-sharing-an-open-protocol-for-secure-data-sharing.html
  8. Akhtar et al. (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
  9. Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data