Skip to content

Multimodal and embodied data

Text-to-CAD Training Data: CAD Models Paired with Design Intent

Quick answer

A useful text-to-CAD dataset pairs parametric CAD, ideally with its full feature history (sketches, constraints, extrudes, fillets), with text engineers actually wrote about why the part exists: requirements, engineering change order (ECO) descriptions, drawing notes and design-review comments. Public corpora give you geometry and synthetic captions that describe shape. Real intent text adds function, tolerance and change rationale. Licensing it means clearing customer ownership, export controls and supplier approval part by part.

By SourceX Editorial · Updated

What public text-to-CAD corpora give you, and where they stop

Public text-to-CAD data is mostly geometry with machine-written captions, which is enough to learn shape grammar but not engineering intent. Text2CAD, the reference text-to-sequence work, annotated the DeepCAD collection of roughly 170K models with about 660K prompts at four levels (abstract, beginner, intermediate and expert), generated by LLM and VLM models (Mistral and LLaVA-NeXT) from rendered views and sequence metadata [1]. Fusion 360 Gallery offers 8,625 human-designed sketch-and-extrude sequences plus a Gym environment for reconstruction [2]. ABC provides one million B-rep models with parametrized curves and surfaces from Onshape, but no construction history and with copyright staying with the original creators [3].

The gap is consistent across these sets. Captions are derived from the shape, so they say "a rectangular plate with four holes" rather than "mounting bracket for the pump, M6 clearance, must clear the 40 mm hose". Follow-up work such as CADmium points out that some text sets, like OmniCAD, contain only beginner-level prompts, and newer systems condition on point clouds and images as well as text [4]. Operational engineering data is the obvious way to close that gap, and it is also the hardest to license.

Which engineering text actually carries design intent

The richest intent text sits in change and review records, not in the CAD file itself. Treat each source below as a candidate pairing and verify coverage during assessment, because many organizations keep only some of them linked to part numbers.

  • Requirements and specifications: system or customer requirements traced to a part number (often in a PLM requirements module or a spreadsheet trace matrix). These give the "must" constraints a prompt would state.
  • ECO / ECR descriptions: the "reason for change" and "description of change" fields on an engineering change order, linked to the before and after revisions. This is the closest real analogue to an edit instruction ("increase wall thickness to 3 mm to pass drop test").
  • Drawing notes and title blocks: general notes, GD&T callouts under ASME Y14.5, material and finish notes. For drawing-image pairs specifically, see the sibling guide on 2D drawings paired with 3D CAD models.
  • Design-review comments: markups and action items from design reviews, often stored as PDF annotations or PLM workflow comments.
  • Feature names and model-tree annotations: engineers who name features ("Boss_Sensor_Mount") leave weak but free intent labels inside the history.

Why feature history beats mesh exports for text-to-CAD

Feature history is the target representation for most text-to-CAD models, so a mesh-only or STEP-only delivery removes the thing you are trying to predict. Text2CAD and Fusion 360 Gallery both train on ordered operation sequences (sketch, profile, extrude with extent and boolean type) rather than on surfaces [1][2]. A neutral STEP AP242 or Parasolid export preserves exact B-rep geometry, which is valuable for evaluation, but it flattens the history into a single solid. STL and OBJ meshes lose exact geometry as well.

Ask whether the supplier can export native files (SolidWorks .sldprt, Creo .prt, NX .prt, Inventor .ipt, Onshape or Fusion documents) and, separately, whether history can be extracted programmatically through the vendor API into a JSON operation list. Native files without a parser are a liability: you will need the same CAD kernel and version to regenerate them, and rebuild failures from missing references or unsupported features are common in older parts. Request a sample that you replay end to end before you commit to a schema.

A record schema for CAD-plus-intent pairs

A text-to-CAD record should keep geometry, history, text and lineage as separate, joinable fields so you can train on any pairing and audit every one. The structure below is a starting point for your request; adjust field names to your pipeline and to what the multimodal specification template already asks for.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "rec-000187",
  "part_ref": "pseudonymized-part-7f3a",
  "revision": {"from": "B", "to": "C"},
  "cad_system": "SolidWorks 2023",
  "files": {
    "native": "rec-000187_revC.sldprt",
    "neutral": "rec-000187_revC.step",
    "history_json": "rec-000187_revC_ops.json"
  },
  "history_ops": [
    {"op": "sketch", "plane": "Top", "entities": 6, "constraints": 11},
    {"op": "extrude", "depth_mm": 12.0, "type": "new_body"},
    {"op": "hole", "standard": "ISO", "size": "M6", "count": 4},
    {"op": "fillet", "radius_mm": 2.0, "edges": 8}
  ],
  "intent_text": [
    {"source": "eco", "field": "reason_for_change", "text": "Cracking at bracket corner in vibration test; add 2 mm fillets."},
    {"source": "drawing_note", "text": "Break all sharp edges 0.5 max. Material 6061-T6, anodize clear."},
    {"source": "requirement", "text": "Bracket shall support 15 kg static load with SF >= 2."}
  ],
  "text_redaction": {"method": "named-entity replacement", "fields": ["customer", "program", "person"]},
  "rights": {"owner": "supplier", "customer_owned": false, "export_review": "EAR99 per supplier classification"},
  "supplier_approved": true
}

The revision pair is the detail most teams forget. An ECO linked to both the "from" and "to" revisions gives you edit-instruction training data (text plus delta), which is more valuable for CAD copilots than single-state captions.

Who owns the geometry, and what export controls gate

Ownership and export classification decide which parts can be licensed at all, and both are settled per part, not per company. Contract manufacturers and engineering service firms often hold large CAD libraries that belong to their customers under build-to-print or design-services agreements. A supplier can only license designs it owns or has the right to sublicense, so ask for a field-level owner flag rather than a blanket statement.

Defense-related parts are the sharper constraint. Under ITAR, technical data includes information required for the design, development, production, manufacture or modification of defense articles, explicitly including drawings and blueprints [5]. Commercial parts may also carry Export Administration Regulations (EAR) classifications other than EAR99. Expect suppliers to exclude anything with a USML or controlled ECCN classification, and ask how the classification was determined.

Public corpora do not remove the rights question either. ABC's own release notes that model copyright remains with the creators, with platform terms also applying [3], and a broad audit of AI datasets found license information missing or wrong on a large share of hosted datasets [6]. Track rights per record, whatever the source; the guide to licensing multimodal records from several rightsholders covers how to structure that.

De-identifying CAD files and engineering text together

CAD files leak identity in places most de-identification passes miss. Native files carry author and last-saved-by properties, file paths with customer or program names, title-block text, embedded part numbers and sometimes logos in sketches or decals. The text side carries names in ECO approver fields, customer names in requirements and program codenames in review comments.

Ask suppliers to cover both sides in one recorded method: strip or replace custom properties and PDM metadata, pseudonymize part numbers consistently across geometry and text so pairs stay joinable, and run entity replacement on the free text. SourceX's existing guide on de-identifying CAD and engineering drawings walks through the file-level steps, and the cluster guide on de-identifying multimodal records covers cross-modal leakage. Geometry itself can be identifying for distinctive products, so treat that as residual risk to discuss, not a solved problem.

Request checklist for a text-to-CAD dataset

Write your request around representation, pairing and rights, since those decide fit more than raw part count.

Illustrative example: invented to show structure; it does not describe an available dataset.

AreaWhat to specifyWhy it matters
CAD system and versionNative format, kernel, oldest acceptable versionRebuild failures and parser coverage
HistoryFull feature tree, or neutral B-rep onlySequence models need operation order [1][2]
Part domainSheet metal, machined, molded, assembliesOperation vocabulary differs sharply
Intent text sourcesECO fields, requirements, drawing notes, reviewsReal "why" versus shape captions
Pair linkageHow text joins to part and revisionUnlinked text is unusable
Revision pairsBefore/after per ECOEdit-instruction training
RightsOwner flag, customer-owned exclusionsOnly owned designs can be licensed
ExportClassification method, controlled exclusionsITAR and EAR scope [5]
De-identificationMetadata, title blocks, text entitiesLeakage across modalities
Eval splitHeld-out parts and familiesAvoid near-duplicate leakage

Near-duplicate leakage deserves its own check: many parts in a company library are configurations or minor revisions of one family, so split by family, not by file. For held-out sets built from private parts, see private evaluation sets for multimodal models. Broader category context sits on the multimodal training data hub and SourceX's overview of CAD and PCB design datasets.

How SourceX sources CAD with design intent

SourceX sources operational datasets, including engineering records and documents, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, such as feature-history CAD linked to ECO text, and SourceX looks for US businesses that hold it. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and every release is approved by the supplying company. You can start from the license CAD and engineering drawings page or go directly to the SourceX buyer intake.

The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Delivery happens through private, access-controlled workflows only after an executed agreement and supplier approval. SourceX does not train models and does not publish prices; terms are agreed per deal.

Source text-to-CAD training data with SourceX

If your text-to-CAD or CAD copilot model needs real engineering intent paired with parametric geometry, describe the CAD systems, history depth, text sources and part domains you need. SourceX will look for US companies that hold that data and manage the license and ongoing purchases with them. Describe your text-to-CAD data request.

Sources

  1. Khan et al., arXiv (NeurIPS 2024), "Text2CAD: Generating Sequential CAD Models from Beginner-to-Expert Level Text Prompts" (2024). https://arxiv.org/pdf/2409.17106
  2. Willis et al., Autodesk Research and MIT, arXiv (SIGGRAPH 2021), "Fusion 360 Gallery: A Dataset and Environment for Programmatic CAD Reconstruction" (2021). https://arxiv.org/abs/2010.02392v2
  3. Koch et al., arXiv (CVPR 2019), "ABC: A Big CAD Model Dataset For Geometric Deep Learning" (2019). https://arxiv.org/pdf/1812.06216
  4. arXiv, "CADmium: Fine-Tuning Code Language Models for Text-Driven Sequential CAD Design" (2025). https://arxiv.org/pdf/2507.09792
  5. eCFR (US Government Publishing Office), "22 CFR 120.33 -- Technical data". https://www.ecfr.gov/current/title-22/chapter-I/subchapter-M/part-120/subpart-C/section-120.33
  6. Longpre et al., arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data