Skip to content

Document AI data

Signature, Stamp and Seal Detection Data for Executed Documents

Quick answer

A usable signature detection dataset pairs real executed business documents with bounding boxes for every signature, initials, stamp, seal and date stamp, plus unsigned signature blocks as labeled negatives. Public sets are small, single-class and often gated or research-oriented, so production execution checks for contracts, loan files and claims usually need licensed documents from the companies that hold them, labeled with a typed region schema and kept separate from any signer identity.

By SourceX Editorial · Updated

This page covers detection: is a document executed, and where are the marks. Verifying who signed is a different task with different data and legal constraints. For the wider task map, start at the document AI data hub.

What a signature and stamp detection dataset must contain

The core unit is a page image with typed, axis-aligned regions for each execution mark, linked to the page's expected execution slots. A detector that only knows "signature" cannot tell a contract with initials missing on page 7 from a fully executed one, so the label set should be at least five classes: handwritten signature, initials, ink stamp, embossed or raised seal, and date or received stamp. Notary seals deserve their own subtype because notary blocks combine a seal, a commission expiry line and a signature in a small area.

Pair each mark with the slot it fills. A slot is the printed signature line, initials box or "affix seal here" area that the template expects; an empty slot is a labeled negative, not an absence of labels. That pairing is what turns object detection into a completeness check ("3 of 4 required signatures present") rather than a heatmap.

Export in a format your stack already reads. COCO JSON is the common interchange for document layout work, as DocLayNet's 80,863 annotated pages in 11 classes show [5]; YOLO text files work for single-stage detectors such as the public Ultralytics set [1]. Ask for page-level coordinates in pixels with the rendering DPI recorded, because a 150 DPI fax and a 300 DPI scan of the same contract produce different boxes. For delivery formats generally, see dataset delivery formats and schemas.

Why public signature datasets fall short for production

Public signature detection sets are useful for a baseline but are too small and too narrow for execution checks. The Ultralytics signature set has 178 images split 143 for training and 35 for validation, with one signature class and no test split [1]. The Tech4Humans set reaches about 2,819 images by combining Tobacco800 with a Roboflow collection, and access requires sharing contact details [2]. The Roboflow Universe office-documents set is published under CC BY 4.0, but community sets differ in license, labeling consistency and document mix [3].

Tobacco800 descends from IIT-CDIP, the scanned Legacy Tobacco litigation archive that also underlies RVL-CDIP. Researchers have questioned how well that corpus represents current business document processing [4]. It is decades-old, mostly monochrome, and light on modern artifacts such as e-signature certificates, colored notary stamps, phone photos of signed pages and e-signature audit trails.

The sets above center on a signature class; none is described as labeling initials, seals and date stamps separately or as including unsigned signature blocks as negatives. Check commercial terms before training on any of them; public document datasets and commercial use covers how to read those licenses.

Hard cases to request explicitly

The hard cases are overlaps and look-alikes, and a dataset without enough of them will report high mAP on clean pages and fail in production. Request them by name, with target shares, rather than hoping they appear in a random sample.

  • Stamps over printed text. "RECEIVED" and "PAID" stamps often land on body text or totals, so the box overlaps OCR tokens. Labels should allow overlapping regions and mark whether the stamp obscures text.
  • Signatures crossing printed lines. Signatures overrun the signature line, the printed name below it and sometimes a neighboring field. Ask annotators to box the ink extent, not the line.
  • Initials versus marginal notes. Two-letter initials in page footers resemble checkmarks, page numbers and reviewer annotations. Include annotated marginalia as distractors.
  • Faint, partial and color-dropped marks. Blue ink and red stamps can disappear in bitonal scans and faxes. The degraded document images guide covers capture conditions to specify.
  • Electronic signatures. Typed-name signatures, drawn e-signatures and certificate panels from e-signature platforms are executed but look nothing like wet ink. Decide whether they are a class or a separate attribute.
  • Embossed seals. Corporate and notary embossing may show only as a faint relief pattern in a scan. Label them even when barely visible, and record visibility.
  • Look-alike negatives. Logos, letterhead crests, barcodes and QR codes near signature blocks trigger false positives.

If your downstream task includes spotting altered or pasted signatures, keep those labels apart; document tampering and fraud detection data treats that as its own problem.

Illustrative label schema for execution detection

A typed region record that links each mark to an expected slot is the most reusable structure, because it supports detection, completeness scoring and redaction pre-processing from one annotation pass.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "doc_id": "D-000417",
  "doc_type": "promissory_note",
  "page": 4,
  "page_count": 5,
  "render_dpi": 300,
  "capture": "flatbed_scan",
  "slots": [
    {"slot_id": "s1", "slot_type": "signature", "party_role": "borrower", "bbox": [212, 1880, 980, 1990], "filled": true},
    {"slot_id": "s2", "slot_type": "signature", "party_role": "co_borrower", "bbox": [1120, 1880, 1890, 1990], "filled": false},
    {"slot_id": "s3", "slot_type": "notary_seal", "party_role": "notary", "bbox": [1300, 2150, 1700, 2520], "filled": true}
  ],
  "marks": [
    {"mark_id": "m1", "class": "signature", "subtype": "wet_ink", "bbox": [240, 1835, 905, 1975], "slot_id": "s1", "overlaps_text": true, "visibility": "clear"},
    {"mark_id": "m2", "class": "seal", "subtype": "notary_ink_stamp", "bbox": [1315, 2170, 1690, 2505], "slot_id": "s3", "overlaps_text": false, "visibility": "clear"},
    {"mark_id": "m3", "class": "date_stamp", "subtype": "received", "bbox": [1500, 140, 1890, 260], "slot_id": null, "overlaps_text": true, "visibility": "partial"}
  ],
  "execution_status": "partially_executed",
  "identity_fields": "none"
}

Note what the record omits: no signer names, no account numbers and no link from a mark to a person. party_role describes the slot, not the individual. The class list (signature, initials, stamp, seal, date stamp) and the slot list together let you compute signed versus unsigned classification at the page and packet level without a separate labeling pass.

Keeping detection labels apart from signer identity

Detection data should carry no identity labels, and the content inside each box should be handled as personal data until counsel says otherwise. Executed documents name the signer beside the signature, and notary stamps print a name and commission number, so the page around a box is often more identifying than the ink.

As of October 2026, Illinois BIPA's definition of "biometric identifier" excludes writing samples and written signatures [6], and Texas defines biometric identifiers as retina or iris scans, fingerprints, voiceprints and hand or face geometry [8]. That does not settle every case. Signature verification models that extract stroke dynamics, documents from health contexts and other state laws can change the analysis; signed forms that are protected health information held by a HIPAA covered entity or business associate need de-identification under Safe Harbor or Expert Determination before release [9]. Under the CCPA, data counts as "deidentified" only when the holding business also meets the statute's conditions on reasonable measures, public commitments and contractual controls [10]. Where BIPA does apply, the 2024 amendment treats repeated collection of the same identifier from the same person by the same method as one violation [7], which limits exposure but does not remove it.

Two practical tests help. First, train one detector on original boxes and one on boxes whose interior is masked, scrambled or replaced with synthetic strokes, then compare recall per class; if recall holds, you can deliver masked interiors and reduce what you hold. Second, redact printed names, notary commission numbers and account numbers outside the boxes before delivery, and record the method. See the de-identified data buyer's guide for how to specify that.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Sourcing executed documents from the companies that hold them

The realistic source of production-grade execution data is operational archives: signed loan packets, executed vendor contracts, claims forms, delivery receipts and notarized affidavits held by lenders, insurers, law firms and operations teams. These archives contain the real mix of wet ink, e-signatures, stamps and blank slots that public sets lack. They also carry confidentiality obligations to counterparties, so the holder's approval and a clear license matter as much as the labels.

SourceX sources operational datasets, including documents and finance and legal workflows, from US companies on request and manages the licensing process; it does not hold inventory, and a request does not guarantee a match. Every release is approved by the supplying company, each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and personal details are removed or replaced before delivery with the method recorded and a sample checked. No de-identification method is perfect. If you are scoping a request, the SourceX buyer intake asks you to describe the documents and labels you need, not which businesses hold them.

Related owner pages: licensed contracts and legal documents and document AI and RAG use cases. Lending teams should also read mortgage loan file document data.

Request checklist for signature and stamp detection data

A good request names the classes, the document mix, the negatives and the identity handling before any price discussion. Use the checklist below with the fuller document dataset requirements spec.

Illustrative example: invented to show structure; it does not describe an available dataset.

RequirementWhat to specifyWhy it matters
Classessignature, initials, stamp, seal, date stamp; e-signature as class or attributeSingle-class sets cannot drive completeness checks
SlotsExpected execution slots per template, with filled/empty flagUnsigned blocks become labeled negatives
Document mixContracts, loan packets, claims forms, receipts; share per typeExecution conventions differ by document type
CaptureScan, fax, phone photo; DPI recorded per pageColor-dropped and low-DPI marks fail differently
Hard casesTarget share of stamps over text, overrun signatures, embossed seals, logos near blocksClean pages inflate mAP
Annotation QADouble-annotated subset with agreement reported per classDocLayNet-style agreement exposes ambiguous classes [5]
FormatCOCO JSON or YOLO, page-level pixel coordinatesAvoids lossy conversion
IdentityNo signer labels; names, commission numbers, account numbers redacted outside boxes; interior masking optionReduces personal data held
Packet contextPage order and document boundaries for multi-page filesInitials on every page need packet-level scoring
LicenseAllowed uses, term, delivery method, deletion obligationsConfirms training and evaluation are both covered

Batch-scanned packets need boundaries before per-document scoring; page-stream segmentation data covers that step.

Request signature and stamp detection data

SourceX looks for US businesses that hold the executed documents you describe, assesses the data and its licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered through private, access-controlled workflows. Nothing is contracted until a supplier agrees. Describe your classes, document mix and identity requirements at sourcex.si/buyers.

Frequently asked questions

Is signature detection the same as signature verification?

No. Detection finds and classifies marks on a page; verification compares a signature with reference samples to judge whether a specific person made it. Verification needs identity-linked reference sets and raises different consent and biometric questions, so scope and license the two separately.

Can I train a stamp detector on synthetic stamps?

Synthetic stamps pasted onto clean pages help with class balance, but they rarely reproduce ink bleed, partial impressions, rotation and overlap with dense text. Use them to augment real executed pages, and evaluate only on real ones. The trade-offs are covered in synthetic vs real documents.

How many unsigned examples should a dataset include?

Enough that every slot type has empty instances in every document type you score. Without them, a detector learns that a signature line implies a signature, which is the exact failure an execution check exists to catch.

Sources

  1. Ultralytics, "Signature Detection Dataset". https://docs.ultralytics.com/datasets/detect/signature
  2. Hugging Face (Tech4Humans), "tech4humans/signature-detection". https://huggingface.co/datasets/tech4humans/signature-detection
  3. Roboflow Universe, "office-documents_signature-detection". https://universe.roboflow.com/esmail/office-documents_signature-detection
  4. arXiv, "Beyond Document Page Classification: Design, Datasets, and Challenges" (2023). https://arxiv.org/pdf/2308.12896
  5. arXiv (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  6. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  7. Illinois General Assembly, "Public Act 103-0769 (SB 2979), BIPA amendment" (2024). https://www.ilga.gov/documents/legislation/publicacts/103/103-0769.htm
  8. Texas Legislature, "Texas Business and Commerce Code Section 503.001". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  9. U.S. Department of Health and Human Services, OCR, "Guidance Regarding Methods for De-identification of Protected Health Information" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  10. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data