Retrieval, RAG and grounding data
Delivery formats that survive chunking: structure, headings and layout for RAG
Quick answer
The best delivery for RAG chunking is two layers: the original files (PDF, DOCX, PPTX, HTML) kept as the citation record, plus a normalized structured text layer, usually Markdown or JSONL, that preserves heading hierarchy, lists, tables, reading order and page numbers. Plain PDF alone forces you to re-parse layout; flat extracted text destroys the section boundaries chunkers rely on. Specify the normalized target, the structure rules and the per-block metadata in the license before delivery, not after ingestion fails.
By SourceX Editorial · Updated
Why licensed documents arrive in mixed formats
Expect a licensed document corpus to arrive as a mix of formats, because the supplier exports whatever its systems hold. Research on retrieval-aware chunking works with mixed-format enterprise-style corpora, and its RAG-Multi-Corpus benchmark pairs 236 documents from five fictional organizations with 786 query-answer pairs carrying ground-truth citations [1]. A knowledge base export may be HTML from Confluence or Zendesk Guide, contracts may be scanned PDFs, and sales enablement may be PPTX. If your request only says "documents," you inherit every one of those parsers and their failure modes.
The fix is to name one normalized target format in the request and treat the originals as evidence. That way your chunker sees one structure model, and your citation layer can still link back to page 14 of the source PDF. For the wider picture of what licensed RAG content involves, start at the retrieval and RAG buyer's guide.
Markdown vs PDF for RAG: what each layer is for
Markdown is the better chunking input and PDF is the better citation record, so a well-specified delivery includes both. PDF encodes glyph positions, not document structure: headings are just larger fonts, multi-column pages interleave when extracted naively, and tables become runs of numbers. Markdown makes structure explicit with # heading levels, list markers and pipe tables, which header-aware splitters can use directly.
Market practice points the same way: one vendor reports that returning full readable Markdown instead of a URL can cut an agent workflow from three LLM calls to one [2]. Converting PDF to Markdown is itself an active research area, with dedicated end-to-end conversion models [3]. Practitioner comparisons rate PyMuPDF4LLM for LLM-ready output, Markitdown for mixed-format batches, Unstructured for production pipelines, Docling for research use and GROBID for reference extraction [4]. Those differences matter because two converters run on the same PDF can produce different heading trees and therefore different chunks.
| Format | Structure it carries | Chunking risk | Role in delivery |
|---|---|---|---|
| Native PDF | Visual layout, page numbers | Reading order, merged columns, lost table cells | Citation original |
| Scanned PDF / TIFF | Pixels only | OCR errors, no text layer | Original; needs OCR layer |
| DOCX | Styles (Heading 1-9), lists, tables | Manual formatting instead of styles | Original; convert from styles |
| HTML | DOM headings, lists, tables | Navigation, footers and boilerplate leak into chunks | Original; strip chrome first |
| PPTX | Slide titles, text boxes, notes | Text box order, speaker notes mixed in | Original; one block per slide |
| Markdown | Heading levels, lists, pipe tables | Complex tables flatten; no page numbers | Normalized chunking layer |
| JSONL blocks | Any fields you define | Schema drift between batches | Normalized layer with metadata |
HTML to Markdown for LLM ingestion: strip the chrome, keep the tree
HTML converts well to Markdown only when the delivery removes site chrome before conversion. Help-center and wiki exports often include breadcrumbs, sidebars, cookie banners, "related articles" widgets and footers, and every one of those becomes repeated text that pollutes embeddings and inflates near-duplicate counts. Ask the supplier, or your own pipeline, to extract the main content element (for example <article> or the platform's body field from the export API) before converting.
Keep <h1>-<h6> levels as #-######, keep <ol>/<ul> nesting, and keep <table> as a table rather than flattening rows into sentences. Preserve anchor IDs on headings, since they let a citation deep-link to the exact section in the source system. Internal links are worth keeping as Markdown links with the original URL so you can rebuild cross-references later.
Structure-preserving chunking needs layout signals the delivery must keep
Structure-aware chunking only works if the delivered text still contains headings, reading order, tables and page anchors. Layout analysis research shows how much structure there is to lose: DocLayNet annotates 80,863 pages with bounding boxes in 11 layout classes such as section headers, captions, tables, footnotes and page headers, in COCO format [6]. Each class implies a chunking rule, for example dropping running page headers and footers and attaching captions to their figure or table.
Reading order is the most common silent failure. OmniDocBench, a CVPR 2025 parsing benchmark, scores reading order with normalized edit distance across diverse PDF types [5]; multi-column and mixed-layout pages are where naive extraction most often scrambles order. Tables are a separate problem with their own datasets and metrics, as PubTables-1M shows for table structure recognition [7]; see table and spreadsheet retrieval data for how to request them. For slide decks and image-heavy scans, where text extraction is not enough, the visual document retrieval guide covers page-image delivery.
Concretely, ask that the normalized layer:
- Keeps the full heading hierarchy, with no skipped levels and no headings rendered as bold body text.
- Emits text in human reading order, with columns serialized left to right and sidebars as separate blocks.
- Keeps tables as tables (Markdown, HTML or cell-level JSON), with header rows marked and merged cells expanded or flagged.
- Records the source page number on every block, so a chunk spanning pages 14-15 can cite both.
- Removes running headers, footers and page numbers from the text body while keeping them in metadata.
- Keeps list items together with their lead-in sentence.
Chunking metadata: section headers and page anchors on every block
Every delivered block should carry its section path and page anchor as metadata, because that is what lets you prepend context to a chunk and cite it precisely. A chunk that reads "The limit is 30 days" is useless for retrieval without "Returns policy > International orders" attached. Storing section_path as an ordered array lets you prefix chunks with breadcrumbs at index time and choose your own chunk size later.
For corpus-level documentation, Croissant offers a schema.org-based JSON-LD vocabulary for describing dataset metadata, file resources and record structure [8], which makes a manifest machine-readable. Block-level fields are covered in more depth in metadata a licensed RAG corpus should ship with. If the corpus is large, columnar Parquet [9] is a practical container for the block table, while JSONL stays easier to inspect.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "kb-00412",
"doc_version": "2025-11-03",
"source_file": "originals/kb-00412.pdf",
"source_sha256": "9f2c...e1",
"block_id": "kb-00412#b0037",
"block_type": "table",
"section_path": ["Returns policy", "International orders", "Refund timelines"],
"heading_level": 3,
"page_start": 14,
"page_end": 15,
"reading_order": 37,
"text_md": "| Region | Window |\n|---|---|\n| EU | 30 days |\n| APAC | 21 days |",
"converter": "docling",
"converter_version": "recorded at export",
"redaction_applied": true
}
Pre-chunked content delivery: when to accept it and when to refuse it
Accept pre-chunked content only alongside the unchunked normalized text, because chunk size and overlap are retrieval decisions that you will want to change. A supplier's 512-token chunks with 50-token overlap may suit their embedding model and not yours, and once boundaries are fixed you cannot re-split by section without the source text. Pre-chunked files are still useful as a reference for how the supplier interprets the structure.
If chunks are delivered, require stable chunk_id values derived from block_id, the chunker name and parameters, and offsets back into the normalized text. That keeps evaluation reproducible, which matters when you later compare runs on pinned corpus snapshots. Never accept embeddings as the only delivery: vectors without text cannot be re-indexed, audited or cited.
A delivery specification to put in the request
Write the format requirements into the request and the license schedule so acceptance testing has something to check against. Rights terms, such as whether the content may be cached or sent to a hosted model, belong alongside the format spec; see RAG content license terms. General file-format expectations are covered on what file formats AI buyers accept, and transfer mechanics on how licensed data is delivered.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | Specification | Acceptance check |
|---|---|---|
| Originals | All source files as exported, with SHA-256 hashes | Hash manifest matches files |
| Normalized layer | CommonMark-style Markdown per document, or JSONL blocks | Parses without errors |
| Headings | Hierarchy preserved from styles or DOM; no skipped levels | Sample of 50 documents reviewed |
| Tables | Kept as tables; header rows marked | Spot-check cell alignment |
| Page anchors | page_start/page_end on every block | No nulls for paginated sources |
| Boilerplate | Running headers, footers and site chrome removed | Near-duplicate block rate reviewed |
| Versioning | doc_version and superseded flags | Older versions identifiable |
| Converter log | Tool name and version per document | Present in manifest |
| Manifest | Corpus-level JSON (Croissant-style) | Counts match delivered files |
Run your own parser over a sample before accepting the full delivery, and score retrieval on a small query set; the retrieval-lift pilot guide describes how. Superseded drafts and near-duplicates deserve their own handling, covered under distractor and near-duplicate documents. If the documents contain personal data, check that redaction happened before normalization so placeholders appear consistently in both layers.
How SourceX handles document requests for RAG
SourceX sources operational datasets from US companies, including documents, support histories and engineering records, and manages the commercial process from licensing through ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement. Personal details such as names, emails and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your format and structure requirements when you submit a buyer request.
Request licensed documents that chunk and cite cleanly
SourceX sources operational documents from US companies on request, rights-reviews each dataset and delivers it under a license that defines records, uses, term and delivery. Describe the documents you need and the delivery structure you expect, and nothing is contracted until a supplier agrees. Start a buyer request.
Frequently asked questions
Is Markdown or JSON better for RAG chunking?
Use Markdown when a human-readable layer and header-aware splitters are the priority, and JSONL blocks when you need typed metadata per block such as page anchors and section paths. Many teams want both: JSONL records whose text field holds Markdown, as in the example above.
Should I chunk by tokens or by section headers?
Split on section boundaries first, then cap by tokens within long sections. That only works if the delivery preserves headings, which is why heading hierarchy belongs in the acceptance criteria.
Do I still need the original PDFs if I get clean Markdown?
Yes. Originals let you re-run conversion with a better parser later, verify a disputed citation against the source page and show reviewers exactly what was licensed.
Sources
- arXiv, "Web Retrieval-Aware Chunking (W-RAC) for Efficient and Cost-Effective Retrieval-Augmented Generation Systems" (2026). https://arxiv.org/pdf/2604.04936
- Firecrawl, "Best news API". https://firecrawl.dev/blog/best-news-api
- arXiv, "Accelerating End-to-End PDF to Markdown Conversion Through Assisted Generation" (2025). https://arxiv.org/pdf/2512.18122
- Tools in Data Science (s-anand.net), "Converting PDFs to Markdown" (2025). https://tds.s-anand.net/2025-05/convert-pdfs-to-markdown/
- arXiv / CVPR 2025, "OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations" (2024). https://arxiv.org/pdf/2412.07626
- arXiv (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- arXiv, "PubTables-1M: Towards comprehensive table extraction from unstructured documents" (2021). https://arxiv.org/pdf/2110.00061
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- The Apache Software Foundation, "Apache Parquet Documentation". https://parquet.apache.org/docs
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.