Document AI data
Document Splitting Data: Page-Stream Segmentation for Batch-Scanned Packets
Quick answer
Document splitting training data is a set of whole, unsplit packets (loan files, claim submissions, records-request returns, AP scan batches) kept in their original page order, with every page labeled as the first page of a new document or a continuation, and every resulting segment labeled with a document type. Single-document classification sets cannot teach this task, because the boundary signal lives in the transition between neighboring pages. Buy packets, not pages, and evaluate boundary F1 per packet type.
By SourceX Editorial · Updated
What a page-stream segmentation model actually learns
A splitting model learns to decide, for each page in a stream, whether it starts a new document, and then what that document is. In the literature this is called page stream segmentation (PSS); in IDP products it ships as a "splitter" that sits in front of classification and extraction [2][3][4]. The input is an ordered sequence of page images plus OCR text; the output is a list of page ranges, each with a type such as closing_disclosure, w2, police_report or invoice.
The decision depends on context that only whole packets carry. A page that says "Page 3 of 7" in the footer, repeats a header from the previous page, or continues a table is a continuation; a page with a new letterhead, a new form number, a fresh "Page 1" or a barcode cover sheet is likely a boundary. UiPath's guidance on its trainable splitter makes the data consequence explicit: train on original production packets, because the model learns bundling patterns from the pages around each document type, and pre-split files reduce splitting accuracy [1].
That is why this page is distinct from document classification training data. Classic classification benchmarks such as RVL-CDIP label individual document images by class [5]; they contain no page order, no neighbors and no boundaries. A splitter trained only on isolated pages has never seen the hard cases: a two-page form whose second page looks like a cover letter, or three consecutive invoices from the same vendor.
Which packets carry the boundary signal you need
The most useful packets are the ones your production pipeline will receive: real, mixed, unpredictably ordered bundles from operational intake. Typical sources include:
- Mortgage and closing packets: applications, pay stubs, bank statements, appraisals, title documents and disclosures concatenated in one PDF; see the mortgage loan file document data guide for the type taxonomy side.
- Insurance claim submissions: first notice of loss, photos printed to paper, repair estimates, medical bills and correspondence, often faxed in several batches.
- Records-request returns: medical or employment records produced as one long scan with inconsistent separators.
- Accounts payable batch scans: dozens of invoices, credit memos and remittance advices fed through a sheet-fed scanner in one job.
Each source has a different boundary profile. AP batches have many short documents with similar layouts, so false merges dominate; claim and records packets have long documents with heterogeneous pages, so false splits dominate. A buyer who trains on one profile and deploys on the other will see boundary F1 drop sharply, which is why packet-type metadata matters as much as the labels.
Preserve the stream exactly as it was captured
The stream must arrive in the order the scanner, fax server or email ingestion produced it, including the artifacts operators would rather delete. Blank separator sheets, patch-code or barcode cover sheets, fax header lines, rotated pages, duplicate pages and scanner-inserted blank backs are all signal: some splitters learn to use them, and production traffic always contains them. Ask suppliers to keep them in place and label them as their own class (for example separator or blank) rather than dropping them.
Equally, do not let the supplier "clean up" packets by re-ordering pages into a logical sequence. If a closing disclosure's page 4 arrived before page 1 in the original stream, that disorder is part of the distribution; record it with a flag instead of fixing it. Capture conditions also matter, since fax and photocopy degradation changes which cues survive; the degraded document images guide covers how to specify capture channel and resolution.
Ask for these capture fields on every packet:
capture_channel: sheet-fed scan, flatbed, fax, email attachment, portal upload, phone photo.source_container: one PDF, multiple PDFs from one email, multi-page TIFF, or individual images.dpiand color mode per page, because mixed 200 dpi bitonal and 300 dpi color pages within one packet are common.original_page_index, never re-numbered after de-identification or filtering.
Label schema: boundaries, types and page roles
A usable label set marks each page's role and each segment's type, with enough structure to compute both boundary and segment metrics. Google's custom splitter workflow follows the same shape at the product level: split documents into train and test sets, annotate boundaries and types, train, then evaluate before deploying [2]. The minimum is a per-page is_first_page flag and a per-segment doc_type; the useful version adds page roles and ambiguity notes.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"packet_id":"pkt_000417","packet_type":"auto_claim_submission","capture_channel":"fax","page_count":14,"pages":[{"original_page_index":0,"role":"cover_sheet","is_first_page":true},{"original_page_index":1,"role":"body","is_first_page":true},{"original_page_index":2,"role":"body","is_first_page":false},{"original_page_index":3,"role":"blank","is_first_page":false}],"segments":[{"segment_id":"s0","page_range":[0,0],"doc_type":"fax_cover"},{"segment_id":"s1","page_range":[1,3],"doc_type":"repair_estimate","ambiguous_boundary":false,"interleaved":false}],"taxonomy_version":"claims_v3","annotator_ids":["a12","a31"],"adjudicated":true}
Store one packet per line as JSON Lines so packets stream through training jobs without loading whole files; the format requires UTF-8, one valid JSON value per line and no blank lines [8]. Keep page images as separate files referenced by packet_id and original_page_index, and keep OCR output (with coordinates) alongside, so the same labels work for vision-only, text-only and multimodal splitters.
Three labeling rules prevent most downstream disputes:
- Define "document" for each type in the taxonomy. Is a check stapled to a remittance advice one document or two? Is an email cover note plus its attachment one segment? Write the rule down and version it as
taxonomy_version. - Mark interleaving explicitly. When pages of two documents alternate (a common double-feed or re-scan artifact), contiguous page ranges cannot represent it; allow non-contiguous segments or an
interleavedflag. - Record ambiguity instead of forcing a guess. An
ambiguous_boundaryflag lets you exclude or down-weight contested boundaries; measure agreement on boundaries with the methods in inter-annotator agreement metrics.
How to evaluate a splitter on purchased packets
Report boundary-level and segment-level quality separately, per packet type, on held-out packets. Boundary precision, recall and F1 on first-page predictions tell you how often the model splits or merges incorrectly; segment-level exact match (a predicted segment counts only if both its page range and its type match the gold segment) tells you how often downstream extraction receives a clean document.
The two diverge in practice. A model can score well on boundary F1 while losing many segments, because one missed boundary corrupts two documents and one false split corrupts one. For a 40-page loan packet, a single false merge between a pay stub and a bank statement may send an extraction model the wrong page set for both, which is why many teams also track "packets with zero errors" as a straight-through-processing proxy.
Split train and test sets by packet and by source, never by page. Pages from one packet in both sets leak layout and header cues, and packets from the same originating company or branch in both sets leak templates. ISO/IEC 5259-3 does not prescribe these metrics, but it gives a frame for documenting the data quality process behind them, including how splits, label rules and checks were managed [7].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Packet type | Packets in test | Boundary P | Boundary R | Boundary F1 | Segment exact match | Zero-error packets |
|---|---|---|---|---|---|---|
| AP scan batch | 120 | 0.97 | 0.91 | 0.94 | 0.88 | 61% |
| Auto claim submission | 80 | 0.93 | 0.95 | 0.94 | 0.84 | 48% |
| Mortgage closing packet | 60 | 0.95 | 0.89 | 0.92 | 0.81 | 35% |
Read the table across, not down: the pattern of low recall in AP batches (missed boundaries between look-alike invoices) and a low zero-error rate in long closing packets is more actionable than any single average.
De-identify without destroying page cues
Packets concentrate personal data, so de-identification must remove identifiers while keeping the visual and positional cues a splitter relies on. A loan or claim packet can repeat a name, account number and address on dozens of pages, across printed text, handwriting, stamps, fax headers and barcodes. Redaction that misses one fax header line leaks the identity of the whole packet.
The splitter-specific failure mode is over-aggressive redaction: black boxes that cover letterheads, form numbers or "Page x of y" footers erase the boundary signal, and replacing whole pages removes the transitions you are paying for. Prefer surrogate replacement of identifiers in both the image and OCR layer, keep layout positions, and never drop or re-order pages during the process. The PII redaction guide for scanned documents covers pixel, OCR-layer and metadata handling in detail.
Medical records inside claim or records-request packets fall under HIPAA when they come from a covered entity or business associate, which means Safe Harbor removal of the 18 listed identifiers (with no actual knowledge that the remainder could identify the person) or an Expert Determination [6]. Barcode cover sheets deserve a separate check, because patch codes and 2D barcodes can encode a claim or loan number that no OCR-based redaction will catch.
Buyer checklist for a splitting data request
Specify the stream, the labels and the evaluation before you talk to suppliers; vague requests for "multi-page PDFs" return pre-split archives. The document dataset requirements spec has the general template; the items below are the splitter-specific additions.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to ask for | Failure if missing |
|---|---|---|
| Unit of delivery | Whole original packets, not per-document files | Model never sees transitions [1] |
| Order | original_page_index preserved through all processing | Learned order cues are fiction |
| Artifacts | Blank, separator and cover pages kept and labeled | Production separators cause false splits |
| Packet mix | Real distribution of document counts and orderings per packet type [1] | Overfits to tidy, short packets |
| Labels | is_first_page, doc_type, page role, interleaving and ambiguity flags | Cannot compute segment metrics |
| Taxonomy | Versioned definitions of each type and of "one document" | Label drift between batches |
| QA | Double annotation on a sample, adjudication record | Unknown boundary noise |
| Splits | Packet-level and source-level held-out test set | Leakage inflates F1 |
| Privacy | Method recorded, image and OCR layers both treated, barcodes checked [6] | Identifier leakage across pages |
For long-context understanding tasks that consume already-split documents, see multi-page and long document data; for pairing segments with extraction ground truth, see documents paired with system-of-record entries.
How SourceX handles document splitting data requests
SourceX sources operational datasets, including documents and finance and legal workflow records, from US companies on request, and manages the licensing process; data is not held in stock and a request does not guarantee a match. Buyers describe the packets they need (packet types, capture channels, label schema), and SourceX looks for US businesses that hold that data, with every release approved by the supplying company. You can start by describing your packet requirements on the SourceX buyer page, and browse the broader enterprise document datasets overview or the document AI data hub.
Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect; health records require HIPAA de-identification by Safe Harbor or Expert Determination. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval.
Request packet-level document splitting data
If your splitter needs real loan, claim, records or AP packets in original page order, describe the packet types, label schema and volume you need. SourceX will look for US companies that hold matching data and, if a supplier agrees, manage the license from assessment through delivery. Describe your document splitting data needs.
Sources
- UiPath, "Trainable Splitter (Document Understanding user guide)". https://docs.uipath.com/zh-CN/document-understanding/AUTOMATION-CLOUD/LATEST/USER-GUIDE/trainable-splitter
- Google Cloud Document AI, "Custom splitter". https://docs.cloud.google.com/document-ai/docs/custom-splitter
- DocuWare Knowledge Center, "Custom Splitting: training a splitting model". https://knowledgecenter.docuware.com/docs/docuware-idp-custom-splitting
- Konfuzio, "Splitting AI". https://help.konfuzio.com/modules/splitting/index.html
- Harley, Ufkes, Derpanis (arXiv:1502.07058), "Evaluation of Deep Convolutional Nets for Document Image Classification and Retrieval" (2015). https://arxiv.org/pdf/1502.07058
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-3:2024 Data quality for analytics and machine learning, Part 3: Data quality management requirements and guidelines" (2024). https://www.iso.org/standard/81092.html
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.