Skip to content

Schemas, packaging and delivery

Apache Arrow and Arrow-Based Dataset Formats for ML Teams

Quick answer

Apache Arrow is a columnar in-memory specification with two serialized forms: the IPC stream format and the IPC file format, which is what Feather v2 now is. Parquet is the compressed storage and exchange format. For licensed training data, the practical rule is to receive Parquet (or JSONL for nested text [3]) from suppliers and build Arrow inside your own pipeline as a memory-mapped cache for pre-training, SFT and eval loaders. Accept raw Arrow files from outside only with pinned versions and an intake check.

By Noah Loul, Founder & CEO, SourceX · Updated

What Apache Arrow actually specifies

Arrow defines how columns sit in memory, so every Arrow-aware process can read the same buffers without converting them. A table is a schema plus a sequence of record batches; each column is a set of contiguous buffers (validity bitmap, offsets for variable-length types, values). Inter-process communication (IPC) moves those record batches either synchronously as a stream or asynchronously by persisting them as a file.

The two IPC flavors matter for intake:

  • IPC stream format (conventional extension .arrows): batches written one after another with no footer, designed for pipes and sockets, read front to back.
  • IPC file format (conventional extension .arrow, also Feather v2): the same batches plus a footer that indexes them, which allows random access to any record batch [4].

Because the on-disk bytes match the in-memory layout, a reader can memory-map an IPC file and use it without a deserialization step or an extra copy. That property, not compression ratio, is the reason ML tooling uses Arrow.

Arrow IPC vs Parquet for training data

Parquet optimizes bytes on disk and over the network, while Arrow IPC optimizes the time from file to usable tensor. Parquet uses dictionary and run-length encodings plus page-level compression, writes metadata in a footer, and marks files with the PAR1 magic number at both ends [1]. An Arrow IPC file is closer to a raw memory dump: it can apply optional buffer compression, but it is usually larger than the equivalent Parquet file.

The table below summarizes the trade-off as most ML platform teams experience it. Use it to decide where each format sits in your ingest path, not as a benchmark.

ConcernParquetArrow IPC file (Feather v2)Arrow IPC stream
Primary roleStorage and exchange between organizationsLocal cache, fast reload, inter-process handoffPipes, RPC, Arrow Flight-style transport
Random accessRow groups and column chunks via footer [1]Record batches via footerNone; sequential only
Memory-mappingNot directly; decode requiredYes, zero-copy readsPossible but no index
Size on diskSmallest (encoding plus compression)Larger; optional LZ4 or ZSTD buffersSimilar to file format
Ecosystem for buyers' lakehousesNative to Iceberg, Delta Lake, Spark, DuckDB, BigQuery loadsRead by pyarrow, Polars, DuckDB, R arrowSame as file format, via stream readers
Long-term archiveReasonable, widely implemented [2]Workable, less commonPoor fit

Arrow vs Parquet for training is therefore rarely either/or. Teams typically store the licensed delivery as Parquet, convert once to Arrow, and let the dataloader memory-map that copy for every epoch.

Feather: v1, v2 and what the name means today

Feather v2 is exactly the Arrow IPC file format [4]; the name was kept for backward compatibility. Feather v1 was an earlier, simplified container for a subset of Arrow, and v1 files lack features such as support for all Arrow data types and compression [4]. If a supplier says they will send "Feather," ask which version, because old pandas and R scripts may still emit v1.

Writers expose a few parameters that change what you receive. The R write_feather function defaults to version 2, and its chunk size setting controls record batch size, where smaller chunks give faster random row access. In pyarrow, pyarrow.feather.write_feather exposes compression (lz4, zstd or uncompressed) and chunksize [6]; a compressed file can no longer be memory-mapped with zero copy, because buffers must be decompressed first.

How Hugging Face Datasets uses an Arrow cache

Hugging Face Datasets keeps every loaded or transformed dataset as Arrow files on local disk and memory-maps them, which is why a large corpus can load using only megabytes of RAM. Loading from CSV, JSON or Parquet triggers a conversion into Arrow cache files under the default cache directory, typically ~/.cache/huggingface/datasets [5]. Each transform in Dataset.map() or filter() writes a new cache file keyed by a fingerprint derived from the input Arrow data and the transform.

Those cache directories are working state, not a delivery format. They hold cache-<fingerprint>.arrow files and dataset_info.json (a save_to_disk directory adds state.json) that reflect one library version and one transform history, and cleanup functions such as cleanup_cache_files() delete them. The datasets library writes these files with the Arrow stream writer rather than the file writer, so a generic IPC file reader can fail on them; check with pyarrow.ipc.open_stream before concluding the file is corrupt.

For buyers this has a direct consequence: do not ask suppliers to ship a save_to_disk directory or a copy of their ~/.cache. Ask for Parquet shards plus a manifest, then let your own load_dataset("parquet", data_files=...) call build the cache in your environment.

When to accept Arrow files from a supplier

Accept Arrow from outside your organization only when both sides control the writer and reader versions and the dataset is a fixed snapshot. Reasonable cases include a supplier whose export pipeline is already Arrow-native (for example Polars write_ipc or pyarrow), columns using an Arrow extension type that the supplier's Parquet path does not preserve, or a second delivery that must be byte-identical to an Arrow artifact you already validated.

Cross-organization Arrow files fail in recognizable ways:

  • Stream vs file confusion: an .arrow file that is really a stream, or the reverse.
  • Feather v1: files that older readers accept but that drop types or compression.
  • Offset overflow: a string or binary column whose single array exceeds the 32-bit offset limit (about 2 GB); the writer should use large_string or large_binary, or split batches.
  • Extension types: custom types identified by ARROW:extension:name metadata that your reader does not register, so they surface as plain storage types and lose meaning.
  • Newer layouts: view types or other features introduced in recent Arrow format versions that your pinned pyarrow cannot read.
  • Time zones: timestamp[us, tz=...] columns written with one zone and naive timestamps in another file of the same delivery.

The last point overlaps with our guide to timestamps and time zones in delivered datasets; the others are specific to Arrow.

Intake checklist for Arrow and Parquet deliveries

A short scripted check catches most format problems before data reaches a training job. Run it on every shard, store the output next to the delivery manifest, and fail the intake rather than "fixing" files silently.

Illustrative example: invented to show structure; it does not describe an available dataset.

StepCheckTool or callReject if
1Magic bytes and footerParquet: PAR1 at both ends [1]; Arrow file: open with pyarrow.ipc.open_fileWrong magic or unreadable footer
2Container flavorTry open_file, then open_streamFlavor differs from the agreed spec
3Schema equalityCompare schema.to_string() across all shards and against the data dictionaryAny field name, type, nullability or order drift
4Writer metadataParquet created_by in file metadata; Arrow custom schema metadataUnpinned or unexpected library version
5Row and batch countsSum num_rows per shard; compare with manifestCount mismatch
6Large-value columnsMax byte length of text columnsSingle values beyond agreed cap
7ChecksumsSHA-256 per file against the manifestAny mismatch
8Round tripLoad into your Datasets cache and sample recordsDecode errors or type coercions

Steps 5 and 7 depend on a manifest; see dataset manifests and checksums for the file layout. Step 3 works best when each record also carries a durable key, covered in stable record IDs and join keys.

A compact delivery spec that keeps Arrow internal might read:

Illustrative example: invented to show structure; it does not describe an available dataset.

delivery_format: parquet
parquet:
  compression: zstd
  target_file_size_mb: 512
  row_group_size_rows: 100000
  writer: "pyarrow (version recorded in created_by)"
schema:
  record_id: string        # stable, never reused
  ticket_created_at: timestamp[us, tz=UTC]
  body_text: large_string  # avoids 2 GB offset overflow
  labels: list<string>
manifest: manifest.jsonl   # path, rows, sha256 per shard
metadata: croissant.json
not_accepted: ["Feather v1", "library cache directories", "pickled DataFrames"]

Pair the spec with machine-readable dataset metadata; our page on Croissant metadata for licensed datasets lists the fields to request.

Where Arrow sits in a licensed-data ingest path

The cleanest design treats Parquet as the contract and Arrow as an implementation detail of your training stack. A typical path is: supplier writes Parquet shards to an agreed location; you verify checksums and schema; you register the shards in a table format or object store prefix; and each training or eval job materializes a memory-mapped Arrow cache local to the compute node.

That separation keeps three things stable. The license and the data dictionary refer to Parquet files with fixed checksums, so audits and deletion requests can point to exact bytes. Arrow caches can be rebuilt or discarded per job, with no compliance meaning. And library upgrades on your side cannot break the supplier contract, because the Parquet spec and its implementation status page are tracked separately from any one reader [2].

If you need table semantics on top of those Parquet files, such as snapshots, schema evolution and time travel, see Iceberg and Delta Lake as a delivery format. For very large multimodal corpora read directly from object storage, streaming dataset formats such as MDS solve a different problem than Arrow caches do. The broader format choice is covered in what file formats AI buyers accept and in the delivery formats hub.

Writing the format into a data request

Put the format decision in the request, before any data moves. When you describe a dataset to a sourcing partner, state the target format (Parquet with ZSTD, or UTF-8 JSONL for nested conversation text [3]), the schema with Arrow types, the large-value policy, and that cache directories are not acceptable deliverables. SourceX works this way: buyers describe the data they need, SourceX looks for US businesses that hold it, and every release is approved by the supplying company.

On the SourceX side, every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect, so keep your own PII scan in the intake checklist.

Sourcing licensed tabular and text data in an Arrow-friendly format

SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents, and finance and legal workflows, and manages the licensing process with the supplier. Data is not held in stock, and a request does not guarantee a match. Tell SourceX what records, schema and format your pipeline needs.

Sources

  1. The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
  2. The Apache Software Foundation (Apache Parquet project), "Apache Parquet Documentation". https://parquet.apache.org/docs
  3. jsonlines.org, "JSON Lines". https://jsonlines.org/
  4. Apache Arrow Project, "Frequently Asked Questions". https://arrow.apache.org/faq/
  5. Hugging Face, "The cache". https://huggingface.co/docs/datasets/about_cache
  6. Apache Arrow Project, "Feather File Format". https://arrow.apache.org/docs/python/feather.html

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data