Multimodal and embodied data
Interleaved Image-Text Data from Business Documents for Multimodal Pre-Training
Quick answer
An interleaved image-text dataset keeps each image at its original position inside a document's text, so a model learns from sequences like "paragraph, diagram, paragraph, photo" rather than isolated caption pairs. Widely used open interleaved corpora, such as OBELICS and MINT-1T, are built from crawled web HTML and PDFs. Licensed business documents (equipment manuals, SOPs with screenshots, inspection reports with photos, training decks) add domain depth and clearer rights, provided you specify serialization, image ownership, metadata handling and deduplication up front.
By SourceX Editorial · Updated
What interleaved data is and how it differs from caption pairs
Interleaved data preserves document reading order with images in place, while caption pairs reduce each image to one short string. In a caption pair, the alt text or caption is the only supervision. In an interleaved document, the image is surrounded by the procedure step that refers to it, the warning that follows it and the table it explains, which is what in-context multimodal learning and long-document VLMs need.
The public reference points are web-derived. OBELICS was assembled from Common Crawl web pages, and MINT-1T extended the approach beyond HTML to web PDFs and arXiv papers to widen coverage of scientific documents. Both inherit the strengths and blind spots of the open web: broad genre coverage, uneven rights information and little operational content.
Those corpora are strong on general web genres and thin on operational documents: torque tables next to exploded-view diagrams, screenshots of a specific ERP screen beside the click path, or a corrosion photo next to the inspector's severity rating. That gap is where licensed business documents earn their place. For pairs rather than sequences, see domain captions from work records; for checking whether text actually describes the adjacent image, see image-text alignment quality checks.
Which business documents produce useful interleaved sequences
The most useful business sources are documents where images carry information the text depends on. Prioritize these:
- Equipment and service manuals. Exploded views, wiring diagrams, part callouts and step photos, with numbered references such as "see Figure 4-12".
- Standard operating procedures with screenshots. Software click paths where every step has a screen capture; dense in UI text and arrows.
- Inspection and field reports. Photos inline with findings, measurements and condition codes, as in inspection reports and inspection photos.
- Training decks and LMS modules. Slides where diagrams and bullet text alternate; see training materials and LMS content.
- Knowledge-base articles. Troubleshooting pages that mix screenshots, error dialogs and resolution text; see internal documentation.
Each source type has a characteristic failure mode. Manuals are often near-duplicated across product revisions. SOP screenshots can expose customer records on screen. Inspection photos carry GPS Exif tags and sometimes faces or license plates. Training decks frequently embed stock images and vendor diagrams the company does not own. The broader catalog of document types is on enterprise document archives.
How to specify serialization so the data is trainable
Serialization is the part of an interleaved request most often left vague, and it determines whether the delivery is trainable without re-parsing. Ask for one record per document (or per logical section for very long manuals), with an ordered list of text and image elements, rather than loose page images plus a separate OCR dump.
Reading order matters most. Multi-column layouts, sidebars, figure captions and footnotes break naive PDF extraction, and parsing benchmarks such as OmniDocBench score reading order with normalized edit distance precisely because it fails so often [3]. Ask the supplier to state the extractor used (for example native DOCX or PPTX XML, HTML DOM, or a PDF layout model) and to keep the original file so you can re-parse later.
Specify these fields at minimum: a stable document ID, an ordered element list with element type (text, image, table, heading), a per-image source ID and file hash, the caption and alt text if present, figure references resolved to image IDs, page and bounding box for PDF sources, and an image placeholder token convention that matches your tokenizer. Tables should come as HTML or Markdown, not as screenshots, unless you want both.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"doc_id": "mfr-manual-0193",
"source_format": "pdf",
"extractor": "layout-model-v2 + native text layer",
"doc_type": "equipment_service_manual",
"language": "en",
"elements": [
{"seq": 1, "type": "heading", "text": "4.3 Replacing the drive belt"},
{"seq": 2, "type": "text", "text": "Isolate power and remove the rear guard (Figure 4-12)."},
{"seq": 3, "type": "image", "image_id": "img-0193-041", "sha256": "...",
"page": 37, "bbox": [72, 210, 540, 498], "caption": "Figure 4-12 Rear guard fasteners",
"alt_text": null, "image_origin": "company_owned", "exif_stripped": true},
{"seq": 4, "type": "table", "format": "html", "text": "<table>...</table>"},
{"seq": 5, "type": "text", "text": "Torque fasteners to the values in Table 4-3."}
],
"figure_refs": {"Figure 4-12": "img-0193-041"},
"redaction_log_id": "rl-0193",
"dedup": {"public_web_match": false, "near_dup_cluster": "c-0193"}
}
Request checklist for a licensed interleaved corpus
A request that a supplier can actually answer names document types, structure, rights and handling, not just token counts. Use this checklist when writing it, and adapt the multimodal dataset specification template for the full version.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Area | What to specify | Why it matters |
|---|---|---|
| Document types | Manuals, SOPs, inspection reports, decks, KB articles; target mix by type | Prevents a corpus dominated by one template |
| Image density | Minimum images per document; exclude logo-only and decorative images | Logos and dividers inflate image counts with no signal |
| Serialization | Ordered elements, figure refs, alt text, placeholders, bounding boxes | Avoids re-parsing and reading-order errors |
| Originals | Source PDF, DOCX, PPTX or HTML retained | Lets you re-extract with a better parser later |
| Image provenance | Flag per image: company-owned, vendor-supplied, stock, unknown | Text ownership does not imply image ownership |
| Metadata | Exif, XMP and document properties stripped or curated; method recorded | Removes GPS, author names and device IDs |
| Personal data | Names, emails, phones, account numbers in text and on screenshots | Screenshots are where redaction is usually missed |
| Deduplication | Exact and near-duplicate removal; overlap check against public web corpora | Public manuals may already be in your crawl |
| Versions | One version per document, or explicit revision chains | Product revisions create near-duplicate floods |
| Allowed uses | Pre-training, fine-tuning, evaluation, and output use stated in the license | Pre-training rights must be explicit |
Image rights inside company documents
Images in business documents often have different owners than the surrounding text, and the license must say which images are covered. A manufacturer's service bulletin reproduced inside a dealer's SOP, a stock photo on a training slide, or a screenshot of third-party software can all sit inside a document the supplier otherwise owns. Ask the supplier to flag image origin per image and to exclude or separately clear anything vendor-supplied or stock.
This matters for downstream disclosure as well as infringement risk. As of October 2026, providers placing general-purpose AI models on the EU market must maintain a copyright policy that identifies and honors rights reservations under Article 4(3) of the DSM Directive, and must publish a training-content summary using the template the Commission issued on 24 July 2025 [5][6]. California AB 2013 requires developers of generative AI made available to Californians to post documentation about training data [7]. Per-image provenance flags make those disclosures much easier to complete accurately.
Open datasets are not a clean substitute here. The Data Provenance Initiative reported license omission rates above 70% and license error rates above 50% on popular dataset hosting sites [4]. For the rights grant itself, see pre-training data license rights and licensing multimodal records from several rightsholders.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Metadata, screenshots and personal data
The privacy risk in interleaved business documents sits mostly in image metadata and screenshot pixels, not in body text. An audit of a large web-scraped image-text pool found non-empty Exif tags on timestamps, geolocation and individuals, and noted that the download tooling stored those tags per sample at download time [1]. Business photos from phones and inspection tablets carry the same GPS and device fields, and Office files carry author and last-modified-by properties.
Require that Exif, XMP and document properties be stripped or reduced to an allow-list (for example orientation only), and that the method be logged per file. Screenshots need visual redaction of customer names, account numbers and email addresses shown in UI fields, which text-only PII scrubbers will not catch. Photos may show faces, badges or license plates.
Indirect identifiers also survive: a site name on a nameplate, a rare equipment serial number, or a small-town address in an inspection header. See de-identifying multimodal records and indirect identifiers in business text.
Deduplication and contamination checks
Deduplicate licensed documents against your existing crawl before you count their value, because many manuals and help-center articles are already public. Run exact hashing on files and images, then near-duplicate detection on text (MinHash or similar) and on images (perceptual hashes), and report the overlap rate with your web corpus separately from internal duplicates.
Revision chains are the second source of inflation. A manual issued in eight revisions can contribute eight near-identical documents; decide whether you want the latest revision only or the chain for change-aware training. Also hold out documents that resemble your evaluation sets. Long-document benchmarks such as MMLongBench-Doc are built on lengthy PDFs with expert-annotated questions [2], and public reports or manuals in your training mix could overlap with similar test material. For held-out tests, see private multimodal evaluation sets and multimodal RAG evaluation data.
How SourceX handles interleaved document requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. The data it looks for includes documents, engineering records, support histories and finance and legal workflows; it does not source scraped web content or generic CCTV or photos. You describe the documents you need, not the businesses that might hold them, and nothing is held in stock, so a request does not guarantee a match.
The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. No de-identification method is perfect, so keep your own checks on screenshots and metadata. You can submit a buyer request with the serialization and rights fields above.
Sourcing interleaved image-text documents for pre-training
SourceX looks for US businesses that hold the manuals, SOPs, inspection reports and training materials you describe, and every release is approved by the supplying company. Datasets are delivered under a license defining records, uses, term and delivery, and terms are agreed per deal. Describe your interleaved document requirements at sourcex.si/buyers, and see the multimodal data hub or pre-training team guide for related sourcing guides.
Frequently asked questions
Can I convert PDFs into interleaved data myself?
Yes, if you receive the original files and the license allows derivatives. Expect reading-order errors on multi-column layouts and figure captions [3], so keep originals and record the extractor version. Whether labs buy PDFs at all is covered in do AI labs buy PDFs.
How is this different from a PDF corpus for VLM pre-training?
A PDF corpus may deliver page images with OCR. An interleaved corpus delivers an ordered text-and-image sequence with figure references resolved, ready to tokenize with image placeholders.
Should I keep alt text and captions?
Keep both as separate fields. Alt text in business documents is often missing or auto-generated, so treat it as optional signal rather than ground truth.
Sources
- arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
- arXiv (Ma et al.), "MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations" (2024). https://arxiv.org/abs/2407.01523
- arXiv / CVPR 2025 (Ouyang et al.), "OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations" (2024). https://arxiv.org/pdf/2412.07626
- arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
- European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
- California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.