Skip to content

Document AI data

Engineering Drawing Annotation Data: Title Blocks, Dimensions and GD&T

Quick answer

An engineering drawing OCR dataset that works for quoting and inspection is not a pile of page images with transcribed text. It needs layered labels: title block fields linked to their values, each dimension parsed into nominal value, tolerance and units, GD&T feature control frames decomposed into symbol, tolerance zone, modifiers and datum references, plus notes, views, revision blocks and BOM tables. Specify the drafting standard and edition, separate scanned legacy sheets from vector exports, and screen out customer-owned and export-controlled drawings before anything ships.

By SourceX Editorial · Updated

Why generic document OCR fails on mechanical drawings

Generic OCR models fail on drawings because the text is sparse, rotated, interleaved with geometry, and meaningful only through symbols and leader lines. A dimension such as "Ø12.00 +0.05/-0.00" reads as a short string, yet a downstream quoting agent needs to know it is a diameter, which feature it points to, and that the tolerance is unilateral. Stacked tolerances, vertical dimension text, and the ⌀ (diameter), ⌴ (counterbore) and ⌵ (countersink) symbols routinely break line-level recognizers trained on forms and receipts.

Public evidence confirms both the demand and the thin supply. The MechVQA benchmark tests multimodal LLMs on orthographic sheets, isometric views, and part and assembly drawings curated through a semi-automated pipeline [1], which is useful for evaluation but not representative of a shop's legacy archive. A University of Turku thesis by Iiro Partanen that evaluates OCR engines on mixed engineering documents illustrates why drawing-specific evaluation is needed [2]. Commercially, what exists openly tends to be small, discipline-specific sets, such as anonymized reinforced-concrete drawing packs sold for AI and OCR training under no-redistribution terms [3]. For page-level transcription conventions, start with our guide to OCR ground truth data.

The annotation layers a drawing-understanding model needs

A usable dataset defines six or seven label layers, each with its own geometry and schema, rather than one flat transcription. Buyers who ask only for "OCR text" get strings without structure and have to re-annotate later.

  • Title block: part number, part name, drawing number, sheet x of y, scale, material and finish callouts, general tolerance block (for example ".XX ±0.01, .XXX ±0.005, angles ±0.5°"), projection symbol (first or third angle), units, drawn/checked/approved names and dates, company name and CAGE code where present.
  • Revision block: revision letter, ECO or ECN number, description, date and approver per row, plus the zone references that link a revision to changed geometry.
  • Dimensions: bounding polygon of the text, dimension type (linear, diameter, radius, angular, chamfer, thread), nominal value, upper and lower deviation or limit values, units, quantity prefix ("4X"), and a pointer to the leader or extension-line endpoints.
  • GD&T feature control frames: each frame split into characteristic symbol, diameter modifier, tolerance value, material condition modifiers, and ordered primary, secondary and tertiary datum references; datum feature symbols labeled separately.
  • Notes and flags: general notes, flag notes referenced by number, surface finish symbols, weld symbols where present.
  • Views and zones: view boundaries and labels (section A-A, detail B), sheet zone grid, and title block region.
  • BOM or parts list: table structure with item number, part number, description and quantity, linked to balloon callouts on the assembly view.

The BOM layer is a table-recognition problem with drawing-specific quirks such as tables that grow upward from the title block; see table structure recognition data for cell-level conventions. Title block and revision block extraction is closer to key-value linking, the kind of entity-linking labels FUNSD popularized on 199 scanned forms [6]. To keep all of these layers consistent across suppliers, write them into one spec using our document annotation schema guide.

Name the GD&T standard and edition in the spec

Every GD&T label set should record which standard and edition the drawing follows, because symbol semantics and defaults differ between ASME and ISO practice. In US shops that usually means an edition of ASME Y14.5; in European and many Asian supply chains it means ISO GPS standards such as ISO 1101 for form, orientation, location and run-out tolerances. Write the edition into the spec rather than "GD&T" alone, and confirm the current edition with the standards body before a contract cites it.

Legacy archives are messier: a 1985-era sheet may carry ANSI Y14.5M conventions, and a European supplier's drawings may mix ISO GPS with company standards. Model-based definition adds another wrinkle, because the same tolerances may live as annotations on a 3D model instead of the sheet. Add an edition field per sheet and let annotators mark "unknown" rather than guess.

Failure modes to label explicitly rather than drop: frames with illegible datum letters, composite position frames (two rows sharing one symbol), all-around and all-over profile indicators, and projected tolerance zones. These rare cases are exactly where inspection-planning agents go wrong.

Scanned legacy sheets versus vector exports

Treat scanned legacy drawings and vector PDF or DXF exports as separate conditions with separate evaluation splits. Vector exports from CAD can carry exact text and coordinates, so labels can be generated programmatically and checked; scanned blueprints, sepias and aperture-card images carry skew, speckle, faded diazo lines and hand-lettered revisions, so labels must be drawn by people.

The strongest training sets pair the two. Where a supplier still holds the native CAD, the model-based definition can supply exact dimension values and tolerances that become labels for the rendered drawing; our page on 2D drawings paired with 3D CAD models covers that pairing intent. For raw native CAD and PCB files as a data category, see CAD and PCB design datasets. Architectural and civil sheets use different conventions and belong in a separate set; see floor plan and architectural drawing data.

Ask for scan metadata per image: source medium (paper, mylar, aperture card, microfilm), DPI, color depth, compression, and whether the scan was cleaned. A model trained on 400 DPI clean TIFFs is likely to degrade on 200 DPI bitonal microfilm conversions, and you can only diagnose that if the field exists.

Screening rights, customer designs and export control

The biggest risk in drawing data is not label quality but who owns the design and whether it may leave the building. A contract manufacturer's archive is full of customer-owned part drawings marked "proprietary" in the title block; the manufacturer may hold the files without any right to license them. Require the supplier to separate its own designs from customer designs, and to show the basis for any customer drawing it includes.

Export control is the second screen. The ITAR defines technical data to include blueprints, drawings, plans and instructions for defense articles [4], and the EAR defines when release of controlled technology counts as an export, including to foreign persons [5]. Drawings for aerospace, defense or dual-use parts should be excluded or reviewed by an export compliance officer before any transfer; our guide to ITAR and EAR checks for engineering records walks through the review.

De-identification is narrower than it looks. Title blocks carry engineer names, company names, CAGE codes and customer part numbers; removing them can also remove labels you wanted. Decide which fields are masked, which are replaced with consistent pseudonyms, and which are kept, and see how to de-identify CAD and engineering drawings for practical methods.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request template for an engineering drawing annotation dataset

A precise request lets a supplier say yes or no quickly and makes acceptance testing possible. Use the template below as a starting point.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample specification
Drawing typesMachined part drawings, sheet-metal flat patterns, weldment and assembly drawings with parts lists
Conditions60% vector PDF exports, 40% scans of paper and aperture cards; scan DPI recorded per image
StandardsASME Y14.5 (edition per sheet) or ISO GPS; "unknown" allowed
Label layersTitle block key-values, revision rows, dimensions with parsed tolerance, feature control frames, datum symbols, notes, view boundaries, BOM table cells
GeometryPolygon per text element; leader endpoints for dimensions; table cell grid for BOM
ExclusionsCustomer-owned designs without license basis; ITAR or EAR-controlled drawings; architectural and civil sheets
De-identificationPersonal names masked; company and customer identifiers pseudonymized consistently
FormatPage images (PNG or TIFF) plus one JSON per page; optional source CAD reference ID
QA evidenceDouble-annotated subset, agreement by layer, adjudication log

A single dimension record in the delivered JSON might look like this:

{
  "sheet_id": "DWG-000412-S1",
  "element_id": "dim_037",
  "layer": "dimension",
  "polygon": [[1822, 940], [1990, 940], [1990, 972], [1822, 972]],
  "dim_type": "diameter",
  "quantity": 4,
  "nominal": 6.6,
  "upper_dev": 0.1,
  "lower_dev": 0.0,
  "units": "mm",
  "linked_fcf": "fcf_012",
  "standard": "ISO GPS",
  "source_condition": "scan_300dpi_bitonal"
}

Acceptance checks before you sign off on delivery

Acceptance for drawing data should be measured per layer, not with one character error rate. Character accuracy on title block text can be high while tolerance parsing is wrong on a meaningful share of dimensions, and only the second number matters to a quoting model.

Run these checks on a held-out sample: exact-match on parsed nominal and deviation fields; datum order correctness in feature control frames; link accuracy between balloons and BOM rows; and field-level recall on revision blocks. Use the methods in how to audit annotation quality to set sample sizes and reject thresholds. Ask the supplier for a datasheet that records sources, collection and annotation methods and intended use, along the lines of Data Cards [7], so your own model documentation can trace each split.

Sourcing annotated engineering drawing data from US companies

SourceX sources operational datasets, including engineering records and documents, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery, and every release is approved by the supplying company. You describe the drawings and labels you need, not the businesses; tell us what your drawing model needs. For other document tasks, browse the document AI data hub or the full AI data guide.

Request engineering drawing annotation data

Describe the drawing types, conditions, standards and label layers you need, and SourceX will look for US businesses that hold matching data. Pricing and allowed uses are agreed in a license before any delivery, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. arXiv, "MechVQA: Benchmarking and Enhancing Multimodal LLMs on Comprehensive Mechanical Drawing Understanding" (2026). https://arxiv.org/pdf/2605.30794
  2. Iiro Partanen (University of Turku), "AI-powered text extraction from engineering drawings". https://www.utupub.fi/bitstreams/85f13f4b-276a-41eb-b933-be3134af9788/download
  3. Hugging Face (PNEngineeringDatasets), "RC_Foundations_Dataset_V1 dataset card". https://huggingface.co/datasets/PNEngineeringDatasets/RC_Foundations_Dataset_V1/blob/main/README.md
  4. eCFR, "22 CFR 120.33 - Technical data". https://www.ecfr.gov/current/title-22/chapter-I/subchapter-M/part-120/subpart-C/section-120.33
  5. Legal Information Institute, Cornell Law School, "15 CFR Part 734 - Scope of the Export Administration Regulations". https://www.law.cornell.edu/cfr/text/15/part-734
  6. arXiv (Jaume, Ekenel, Thiran), "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  7. arXiv (Pushkarna, Zaldivar, Kjartansson), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data