Schemas, packaging and delivery
WebDataset Tar Shards: When to Request Them for Multimodal Data
Quick answer
Request WebDataset tar shards when you will stream millions of image, audio or video samples with per-sample metadata into a distributed training job. The format packs each sample's files under one shared basename inside numbered POSIX tar archives, so data loaders read shards sequentially from disk or object storage [1]. Specify shard size, key naming, required per-sample fields and a per-shard checksum manifest in the request. Choose Parquet or original files instead when you need random access, column filtering or frequent re-sharding.
By SourceX Editorial · Updated
What the WebDataset format actually contains
A WebDataset delivery is a set of ordinary tar files whose members are grouped into samples by basename. The project repository describes shards numbered like something-000000.tar and addressed as a range with brace notation, such as something-{000000..012345}.tar [1]. Inside a shard, 000123.jpg, 000123.json and 000123.txt share the key 000123 and become one sample with three fields [1]. Because members are plain tar entries, any tool that understands tar (GNU tar, Python's tarfile, bsdtar) can list and extract them without the WebDataset library.
That simplicity is the point and the risk. The tar itself has no embedded schema, no index and no footer: a reader learns what fields exist only by reading samples. Contrast Parquet, where file metadata sits in a footer that records the location of every column chunk, which lets a reader skip to the columns and row groups it needs [3]. For buyers, the absence of a schema means the request, not the file, has to define what a valid sample looks like.
When tar shards beat Parquet, JSONL or loose files
Tar shards win when the workload is sequential, large and binary-heavy. Training jobs that read each sample once per epoch, in roughly shuffled order, across many GPUs benefit from large sequential reads rather than millions of small object GETs. Hugging Face lists WebDataset as a supported Hub format for large-scale datasets and notes that shards are often around 1 GB, with full datasets reaching multiple terabytes [2]. If your loader is already built on webdataset, Hugging Face datasets streaming or a similar tar reader, receiving tar shards removes a conversion step.
They lose when you need to query or edit. Selecting all samples where speaker_age_band = "65+" means scanning every shard's JSON members. Removing one sample after a supplier withdraws consent means rewriting the shard that contains it, then updating the manifest. For tabular metadata with embedded thumbnails, see Parquet vs JSONL for licensed training data; for formats with built-in random access, see streaming training data from object storage.
| Need | Tar shards (WebDataset) | Parquet with binary column | Original files plus sidecars |
|---|---|---|---|
| Sequential streaming at cluster scale | Strong | Good, with tuned row groups | Weak (many small reads) |
| Filter by metadata before reading media | Weak (scan JSON members) | Strong (column pruning via footer) [3] | Moderate (needs separate index) |
| Remove or replace one sample | Rewrite whole shard | Rewrite file or row group | Delete one file |
| Inspect with standard tools | tar -tvf | Needs Parquet reader | File browser |
| Suits audit and rights review | Needs manifest | Needs manifest | Closest to source |
A common compromise is to request both: original files with sidecar metadata as the archival, auditable copy, and tar shards as the training copy built from it. The original files and sidecar metadata guide covers the archival layout.
Choosing a shard size and shard count
Size shards so every data-loading worker gets many shards per epoch, not just a few. A 1 GB target is a widely used starting point [2], but the real constraint is count: with 64 GPUs and 8 loader workers each, you have 512 readers, and a 200 GB dataset in 1 GB shards gives each reader fewer than one shard. That starves shard-level shuffling and leaves some workers idle at the end of an epoch. For small datasets, drop shard size to 100 to 500 MB so shard count comfortably exceeds total workers.
Very small shards create the opposite problem: object storage request overhead and long file listings. Very large shards (10 GB and up) make a single corrupt or withdrawn sample expensive to fix and slow down partial downloads for spot checks. State the target as a range (for example, 0.5 to 1.5 GB, last shard exempt) and ask the supplier to report the actual distribution in the manifest.
Key naming and field extensions
Keys must be unique across the whole dataset, stable across releases, and free of personal data. WebDataset groups members by basename and treats the remaining suffix as the field name [1], so dots inside a key will split it in the wrong place; verify the exact splitting rule against the current library documentation before you fix a convention. Use zero-padded opaque IDs (s0004817233) rather than original filenames such as john_smith_kitchen_2024.jpg, which can leak names, locations or dates into training logs.
Fix the field extensions in the request, because the extension tells the decoder what to do. Typical choices are .jpg or .png for images, .flac or .wav for audio, .mp4 for short clips, .json for structured metadata, .txt for captions or transcripts, and .cls for an integer label. If one sample can have several images or segments, define compound fields (front.jpg, rear.jpg) up front rather than letting the supplier improvise. For audio specs such as sample rate and channel layout, see audio file specs for speech datasets; for long video, tar shards usually hold clips while full recordings follow the video packaging guide.
Per-sample JSON metadata buyers should require
Every sample should carry a JSON member with provenance and rights fields, not just training labels. Since the tar has no schema, write one: field names, types, allowed values and which fields are required. Include a stable sample ID that matches the key, a source record ID the supplier can trace, capture or creation timestamp in ISO 8601 with offset, the license or allowed-use tag that applies, the de-identification method applied, and any split assignment (train, validation, held-out eval).
Illustrative example: invented to show structure; it does not describe an available dataset.
shard: inspections-train-000042.tar
s0004817233.jpg
s0004817233.json
s0004817233.txt
s0004817234.jpg
s0004817234.json
s0004817234.txt
s0004817233.json
{
"sample_id": "s0004817233",
"source_record_id": "INSP-7781-03",
"captured_at": "2025-03-14T09:22:05-05:00",
"media": {"width": 3024, "height": 4032, "mime": "image/jpeg", "exif_stripped": true},
"caption_origin": "technician_note_redacted",
"deid_method": "faces_blurred;names_replaced_v2",
"license_tag": "LIC-2026-014/train",
"split": "train",
"sha256": "9f2c...e41a"
}
Dataset-level facts (creator, license, record structure, file resources) belong outside the shards. Croissant expresses this as JSON-LD built on schema.org, describing both the dataset and its file resources [4], and the Croissant metadata guide covers what to ask for.
Shuffling, epochs and worker splitting
Tar shards give you two levels of randomness, and you need both. First, the loader shuffles the order of shard URLs and assigns them to nodes and workers; second, it keeps an in-memory buffer of samples and draws from it at random. If the supplier wrote shards in source order (all samples from one site, one device or one day together), a small sample buffer will still feed the model long correlated runs. Ask the supplier to shuffle samples across shards before writing, and record the seed in the release notes so the order is reproducible.
Budget memory for the buffer: 5,000 images decoded to 1024 by 1024 RGB arrays take roughly 15 GB per worker, so keep buffers of encoded bytes or size them to your RAM. Check how your loader splits shards across nodes and workers (the webdataset library exposes node and worker splitting functions [1]), and confirm that the shard count divides reasonably among them. Uneven counts cause some ranks to finish early, which matters for synchronous data-parallel training.
Streaming tar shards from S3 and other object stores
Tar shards stream well from object storage because each shard is one large sequential read. WebDataset can read from local paths or any pipe, so a URL such as pipe:aws s3 cp s3://bucket/train-{000000..000511}.tar - streams without staging the full dataset [1]. WebDataset repositories hosted on the Hugging Face Hub follow the same tar-shard layout, so they can be read shard by shard in the same way [2].
Decide who pays for egress before delivery. If the supplier hosts shards in a Requester Pays bucket, your account pays request and download costs while the owner pays storage, and every request must be authenticated [5]. For cross-account patterns and the trade-off between copying and reading in place, see cloud bucket cross-account delivery.
Shard specification checklist for a supplier request
A complete request fixes the layout, the metadata contract and the acceptance tests before any shard is written.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to specify | Acceptance check |
|---|---|---|
| Shard naming | {dataset}-{split}-{000000..NNNNNN}.tar, zero-padded, contiguous | No gaps in range; brace pattern resolves [1] |
| Shard size | Target range, e.g. 0.5 to 1.5 GB; last shard exempt | Size histogram in manifest [2] |
| Key format | Opaque, unique, no dots, no personal data | Uniqueness check across all shards |
| Fields | Required extensions per sample, e.g. .jpg, .json, .txt | Every key has every required field |
| JSON schema | Field names, types, enums, required flags | Validate 100% of JSON members |
| Media spec | Codec, resolution or sample rate, EXIF policy | Decode-test a random sample |
| Shuffle | Cross-shard shuffle before write; seed recorded | Source-ID runs per shard below a threshold |
| Splits | Separate shard sets per split, no key overlap | Set intersection is empty |
| Manifest | Per-shard SHA-256, byte size, sample count, key range | Recompute hashes on receipt |
| Change handling | How withdrawn samples are removed and shards reissued | New manifest with version ID |
Checksums deserve their own process; the manifest and checksum verification guide explains byte-level verification, and dataset versioning covers reissued shards. Use this checklist alongside the multimodal dataset specification template, which defines content and coverage rather than packaging.
Failure modes seen in tar-shard deliveries
Most problems come from writing shards without a contract. Watch for keys that collide across shards, which silently merge or drop samples; macOS ._ resource-fork members or PaxHeader entries that appear as orphan fields; JSON members missing on a fraction of samples, which some loaders skip without error; and split leakage when the same source record lands in both train and eval shards under different keys.
Media issues are the second group. Truncated JPEG or FLAC members decode partially and crash a worker hours into a run. EXIF orientation flags that were not applied produce rotated images, and leftover EXIF GPS tags can carry location data that the de-identification plan meant to remove. Run a full decode pass and a metadata strip check on receipt, not only on the supplier's sample. For broader acceptance testing, see the quality assessment hub and the privacy hub.
How SourceX handles multimodal delivery requests
SourceX sources operational datasets from US companies on request, including documents and new recordings of hands-on work, and manages licensing and ongoing purchases; it holds no stock, and a request does not guarantee a match. You describe the data and packaging you need, such as tar shards with the fields above, and SourceX looks for US businesses that hold it. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Start a request on the SourceX buyers page, or see the category pages for images and inspection photos and voice and audio data. More delivery formats are in the delivery hub and the AI data guides.
Request image, audio or video data as WebDataset shards
If your training stack reads tar shards, put the shard specification in the request from the start so packaging is part of the license discussion. SourceX handles the process from finding a supplier through assessment, agreement, transaction and ongoing management, and nothing is contracted until a supplier agrees. Describe the multimodal data you need on the SourceX buyers page.
Sources
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- Hugging Face, "WebDataset (Hub documentation)". https://huggingface.co/docs/hub/datasets-webdataset
- The Apache Software Foundation (Apache Parquet project), "File Format". https://parquet.apache.org/docs/file-format/
- Akhtar et al. (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Amazon Web Services (Amazon S3 User Guide), "Using Requester Pays general purpose buckets for storage transfers and usage". https://docs.aws.amazon.com/AmazonS3/latest/dev/RequesterPaysBuckets.html
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.