Schemas, packaging and delivery
Delivering Original Files with Sidecar Metadata and Extracted Text
Quick answer
Sidecar metadata files for datasets are small, machine-readable files (usually JSON) that sit next to each original PDF, scan or attachment and describe it without modifying it. A sound document delivery ships three layers per file: the untouched original named by its content hash, a sidecar with source, MIME type, page count, language, redaction status and parent record ID, and derived text and layout (plain text, ALTO or hOCR) as separate files, all covered by one checksummed manifest.
By SourceX Editorial · Updated
Why originals and derived text should travel as separate files
Keep the original byte-identical and put every derived artifact in its own file, because any extraction you bake into the original is one you can never redo. OCR engines, PDF text extractors and layout models improve every year; a research engineer who receives only a supplier's 2026 OCR output is locked into that supplier's error profile. Receiving the original plus the derived text lets you re-extract with your own pipeline, compare against the supplier's output and train layout models on the pixels.
The separation also keeps provenance honest. If a sidecar says ocr_engine: tesseract-5.x and the text file sits beside the original, a later audit can tell which text came from the document and which came from a tool. When text is injected back into a PDF as an invisible layer, that distinction blurs and the file's hash changes. For the broader raw-versus-processed trade-off, see whether AI labs want raw or cleaned data.
Content-addressed naming for originals and attachments
Name each original by a cryptographic hash of its bytes (for example, SHA-256 in hex) rather than by its source filename, because hash names deduplicate attachments automatically and make every reference verifiable. Email and ticketing exports are full of repeats: the same signed contract, logo image or terms PDF attached to hundreds of messages. With content addressing, those collapse to one stored object referenced by many parent records.
Deduplication matters beyond storage. Lee et al. found that common language-model training datasets contain many near-duplicate examples, and that deduplicating them reduced how often models emit memorized training text verbatim [7]. Exact-hash naming only removes byte-identical copies, so ask whether the supplier also ran near-duplicate detection (for example, MinHash on extracted text) and whether the sidecar records a near_dup_cluster ID instead of silently dropping files.
Practical rules buyers can put in a delivery spec:
- Shard the object store by hash prefix (
objects/9f/3a/9f3a...e1.pdf) so no directory holds millions of entries. - Keep the original extension on the hashed name so tools that sniff by suffix still work, but treat the sidecar's
mime_typeas authoritative. - Store the original filename only in the sidecar, and only if it has been reviewed, since filenames often carry personal names and case numbers.
- Never re-encode, linearize, "repair" or re-save a PDF before hashing; any of these changes the bytes and breaks the link to the source system.
What a per-file sidecar should contain
A sidecar should answer, without opening the original, where the file came from, what it is, what was done to it and which record it belongs to. A useful minimum is: sha256, byte_size, mime_type, source_system (for example, an email archive, a ticketing tool or a document management system), parent_record_id, attachment_index, page_count, language, created_at and modified_at as source timestamps, redaction_status, redaction_method, derived_files with their own hashes, and license_scope pointing to the governing agreement's dataset identifier.
The parent link is the field buyers most often forget to request. Without parent_record_id, an attachment is orphaned from the email, ticket or matter that gives it meaning, and RAG systems lose the ability to cite the thread or filter by account, product or date. Field-level guidance for text corpora overlaps here; see metadata fields to require with licensed text corpora and metadata a licensed RAG corpus should ship with.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"sha256": "9f3a4c...e1",
"byte_size": 482113,
"mime_type": "application/pdf",
"original_filename_reviewed": "service_agreement_signed.pdf",
"source_system": "email_archive",
"parent_record_id": "msg-000184233",
"attachment_index": 2,
"page_count": 7,
"language": ["en"],
"has_text_layer": false,
"created_at": "2023-04-11T15:02:09Z",
"redaction_status": "redacted_rendering_licensed",
"redaction_method": "pii-detect + manual review; rasterized pages",
"derived_files": [
{"role": "page_images", "format": "image/png", "dpi": 300, "count": 7},
{"role": "ocr_layout", "format": "alto+xml", "sha256": "41b0...7c"},
{"role": "plain_text", "format": "text/plain; charset=utf-8", "sha256": "c2de...09"}
],
"near_dup_cluster": "nd-55120",
"license_scope": "DS-EXAMPLE-01"
}
Deliver sidecars either as one <hash>.json per original or as a single JSON Lines index, one object per line. JSON Lines requires UTF-8 without a byte order mark and one valid JSON value per line, with no blank lines [9]. The per-file form is easier to inspect; the index is easier to load into a dataframe, and many teams ask for both.
Choosing formats for OCR output and layout
Ask for OCR output in a layout-preserving standard alongside plain text, because plain text alone discards the coordinates that document-AI and table-extraction models train on. Three formats cover most deliveries:
| Format | Maintainer | Structure | Good fit |
|---|---|---|---|
| ALTO XML | Library of Congress | Page, text blocks, lines and words with position and size [2] | Archives, digitized records, long-term storage |
| hOCR | Community spec (kba/hocr-spec) | HTML microformat with classes and bounding boxes in markup [3] | Pipelines already emitting HTML, Tesseract users |
| PAGE XML | PRImA | Regions, lines and reading order designed for ground truth [4] | Layout ground truth, historical documents |
| Plain text | n/a | UTF-8 text, one file per page or per document | Embedding, search, LLM training |
Whichever format you pick, require the sidecar to record the coordinate system (pixels at a stated DPI, or PDF points), the page image the coordinates refer to, and the OCR engine and version. A common failure is OCR coordinates computed against a 300 DPI rendering while the delivered page images are 150 DPI, which silently misaligns every bounding box.
Born-digital PDFs need a different note. If has_text_layer is true, ask whether the derived text came from the embedded text layer or from OCR on a rendering, since the two disagree on ligatures, hyphenation, reading order in multi-column layouts and hidden text. Public benchmarks show why layout matters: FUNSD annotates 199 noisy scanned forms for text detection, OCR, layout analysis and entity linking [8], and real enterprise forms are noisier than most public sets.
Redacted renderings versus originals: deciding the licensed artifact
Decide in writing which file is the licensed artifact, the original or a redacted rendering, because the two have different hashes, different training value and different privacy risk. A redacted rendering is typically produced by detecting personal data, burning black boxes into page images and regenerating the text layer. If redaction only draws a box over the PDF while the underlying text, XMP metadata or embedded attachments remain, the personal data is still in the file.
In many purchases of operational documents, the redacted rendering is what the buyer receives and trains on, and the sidecar should mark it explicitly: redaction_status: redacted_rendering_licensed, with the method recorded. SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method and checks a sample, and no method is perfect. Health records additionally require HIPAA de-identification by Safe Harbor or Expert Determination. The pixel, OCR-layer and metadata pitfalls are covered in redacting PII in scanned documents.
Three checks to run on arrival:
- Extract text from each redacted PDF with a different tool than the supplier used and search for patterns the redaction should have removed.
- Strip and inspect XMP and document-info metadata; author and company fields frequently survive redaction.
- List embedded files and annotations; PDF portfolios and email
.msgfiles can carry nested attachments the sidecar does not mention.
Folder structure and packaging for document datasets
Package the delivery so that a single manifest covers every original, sidecar and derived file, and so the layout is predictable enough to script. BagIt (RFC 8493) is a solid default: a data/ payload directory, a bagit.txt declaration, optional bag-info.txt metadata and manifest-<algorithm>.txt files listing a checksum for every payload file, plus tag manifests for the metadata files themselves [1]. BagIt is an informational RFC rather than a Standards Track specification, but it has wide tool support, so both sides can usually validate a bag without custom code.
Illustrative example: invented to show structure; it does not describe an available dataset.
delivery-2026-10-v1/
bagit.txt
bag-info.txt
manifest-sha256.txt
tagmanifest-sha256.txt
data/
objects/9f/3a/9f3a...e1.pdf
sidecars/9f/3a/9f3a...e1.json
derived/9f/3a/9f3a...e1/page-0001.png
derived/9f/3a/9f3a...e1/ocr.alto.xml
derived/9f/3a/9f3a...e1/text.txt
index/files.jsonl
index/records.jsonl
croissant.json
Add a dataset-level description on top of the per-file sidecars. Croissant, a schema.org-based JSON-LD vocabulary from MLCommons, describes the dataset, its file resources and how records are structured, so loaders can read the index without bespoke parsing [5]; see what to ask suppliers for in Croissant metadata. For training throughput, some teams repack into WebDataset tar shards, where files sharing a basename (for example, 9f3a...e1.pdf, 9f3a...e1.json, 9f3a...e1.txt) become one sample [6]. Treat shards as a derived convenience: keep the bag as the system of record, and check the manifest byte for byte as described in dataset manifests and checksums.
Acceptance checklist for a file-plus-sidecar delivery
Run acceptance against the manifest and sidecars before anyone trains on the data, because sidecar errors propagate into every downstream index.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Check | Pass condition |
|---|---|
| Bag validity | Every payload file is listed in manifest-sha256.txt and every hash recomputes |
| Name integrity | Each object's filename equals the SHA-256 of its bytes |
| Sidecar coverage | Exactly one sidecar per original; no orphan sidecars |
| MIME truth | Sniffed type (for example, via libmagic) matches mime_type |
| Page counts | page_count equals pages parsed and page images delivered |
| Parent links | Every parent_record_id resolves in records.jsonl |
| Derived hashes | Every derived_files hash matches the file on disk |
| Coordinate sanity | Sample OCR boxes overlay correctly on page images at stated DPI |
| Redaction | Spot-check text layers, metadata and embedded files on a sample |
| Encoding | Text files are UTF-8; JSON Lines has no BOM or blank lines |
Recurring purchases need the same contract across releases. Version the sidecar schema, and note added fields in release notes, so your loaders do not break; see handling schema changes across recurring deliveries and the dataset delivery formats hub.
How SourceX handles document deliveries
SourceX sources operational datasets, including documents and the attachments in support, sales, finance and legal workflows, from US companies on request; it does not hold stock, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, comes with diligence materials covering source, rights, preparation and allowed use, and is delivered under a license defining the records, uses, term and delivery. Delivery runs through private, access-controlled workflows, never email attachments, and only after an executed agreement and supplier approval. Buyers can describe the files, sidecar fields and OCR formats they need on the SourceX buyer page. For the categories themselves, see enterprise document datasets, whether AI labs buy PDFs and the AI data hub.
Request original documents with sidecar metadata
Describe the document types, the sidecar fields and the derived formats your pipeline expects, and SourceX looks for US businesses that hold that data. Every release is approved by the supplying company, and pricing and allowed uses are agreed in a license before anything is delivered. Start your request at sourcex.si/buyers.
Sources
- RFC Editor (IETF), "RFC 8493: The BagIt File Packaging Format (V1.0)" (2018). https://www.rfc-editor.org/rfc/rfc8493
- Library of Congress, "ALTO Technical Metadata for Layout and Text Objects". https://loc.gov/standards/alto/description.html
- kba/hocr-spec (GitHub), "hOCR format specification". https://github.com/kba/hocr-spec
- Wikipedia, "PAGE (XML)". ) https://Www.Wikipedia.org/wiki/PAGE_(XML
- MLCommons Croissant working group, "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- WebDataset project, "webdataset (GitHub repository)". https://github.com/webdataset/webdataset
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Jaume, Ekenel, Thiran, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.