Code and software engineering data
Verilog, SystemVerilog and VHDL Datasets for Chip Design AI
Quick answer
A useful Verilog dataset for LLM training is not a pile of .v files. It is RTL from real design teams paired with the artifacts that grade it: testbenches, SystemVerilog Assertions, functional coverage, lint and synthesis reports, plus the Tcl scripts and design notes engineers wrote around it. Public HDL is small and skewed toward tutorials, so teams building production assistants often look to proprietary corpora. The two things to settle first are rights to third-party IP and export-control screening.
By SourceX Editorial · Updated
Why public HDL data runs out quickly
Public HDL corpora are thin because almost all production RTL sits behind company firewalls, IP licenses and foundry NDAs. Public benchmarks lean on short, self-contained problems such as teaching-site exercises and single modules. These are good tests of small-module generation. They say little about a 400-file SoC subsystem with clock-domain crossings, parameterized generate blocks, vendor macros and a UVM environment, which is what production assistants face.
Open GitHub HDL also carries the usual provenance problems. Audits of popular dataset hosts found license information missing for more than 70% of datasets and wrong in more than half of the cases checked [3]. For HDL that matters doubly, because vendor IP and course solutions are commonly re-uploaded without permission. See open code datasets vs licensed private code for the general tradeoff.
Verification artifacts are the labels
Testbenches, assertions and tool reports turn RTL from pre-training text into graded supervision. Without them, a corpus supports continued pre-training and style learning, but you cannot build pass/fail SFT targets, reward signals or regression evals. Recent reinforcement-learning work on Verilog generation depends on automated verification of each generated design to compute rewards [2], and that verification needs a reference or a testbench. Ask for each design unit to ship with whatever the supplier's flow already produces:
- Testbenches: SystemVerilog/UVM environments (sequences, scoreboards, monitors), cocotb Python tests, or VHDL testbenches with self-checking processes.
- Assertions: SVA properties and
bindfiles, plus formal property results (proven, falsified with counterexample trace, inconclusive). - Coverage: functional covergroups and code coverage exports (line, toggle, FSM, branch), ideally in UCIS or the simulator's native database with a text report.
- Static and implementation reports: lint waivers and violations, CDC/RDC reports, synthesis logs with area/timing summaries and inferred-latch warnings.
- History: commits that fix a failing regression, linked to the bug ticket. These are the hardware equivalent of issue-to-fix pairs.
The unit-test logic carries over from software; unit test generation data covers focal-code and coverage pairing in more depth.
EDA scripts and design documentation belong in the same request
EDA domain adaptation works best when code arrives alongside the text engineers use to operate tools. A recent EDA-domain training corpus drew on vendor tool manuals, engineer Q&A records, papers and script documentation [1], and internal bug reports and design reviews add the reasoning that RTL alone never states. Request Tcl run scripts for synthesis, place-and-route and static timing, SDC constraint files, Makefile or Python regression drivers, and the internal wiki pages or design specs that explain them.
Treat script tasks as their own eval track. One-to-five-line queries ("report all flops in this hierarchy") can be scored by executing the script and diffing output against a golden run. Longer flows, such as an ECO script or a timing-closure loop, usually need engineers to grade them. If your target is an EDA scripting assistant, you will need both kinds of labels, and the second kind requires domain experts; designing evaluation rubrics with domain experts covers that workflow.
What to exclude before anything ships
The main rights trap in RTL is code the supplier uses but does not own. A chip team's repository typically mixes in-house RTL with licensed third-party IP cores (interconnect, SerDes PHY wrappers, memory controllers), encrypted IEEE 1735 IP blocks, foundry PDK files, standard-cell libraries (Liberty .lib, LEF) and memory-compiler outputs, many of them under NDA. None of that is the supplier's to license onward.
Ask for a component inventory per design: each directory or file tagged as in-house, third-party licensed, open source (with SPDX identifier) or generated. Open-source HDL brings its own license mix, including copyleft; see copyleft contamination in licensed code. For ownership questions such as contractor-written blocks or acquired IP, use the code ownership due diligence checklist.
Export controls are the second screen. Some designs, such as cryptographic accelerators or defense-related hardware, may be controlled technology, and sharing controlled source with foreign persons can count as an export. Have counsel classify the material before transfer; the SourceX note on data licensing and export controls outlines the questions.
Matching the corpus to the training stage
Each training stage needs a different slice of an RTL corpus, so specify the stage before you specify volume.
| Stage | What to require | Grading signal |
|---|---|---|
| Continued pre-training | Deduplicated RTL across Verilog, SystemVerilog and VHDL, specs, Tcl, SDC, manuals | Perplexity on held-out internal HDL |
| SFT for RTL generation | Spec or docstring paired with module, plus testbench | Simulation pass against reference |
| Verification assistant | RTL plus SVA, UVM components, coverage holes and closure commits | Assertion proves or fails; coverage delta |
| EDA scripting assistant | Tcl/Python scripts with the run logs they produced | Execution success; expert score |
| Held-out eval | Never-published modules and regressions with tool versions pinned | Pass@k with failure categories |
Pass rates on HDL tasks can swing with prompt wording and simulator version, so fix both when comparing models. Keep evaluation designs out of any training split and tag them with a canary string, as BIG-bench did for its task files [4]; held-out eval sets from private repositories and code benchmark contamination explain why.
A request template for RTL data
The fastest way to get a usable answer from a supplier is a request that names languages, artifacts, tools and exclusions.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: rtl_corpus_v1
languages: [SystemVerilog-2017, Verilog-2005, VHDL-2008]
design_domains: [bus interconnect, DSP datapaths, FSM-heavy control]
per_design_unit:
rtl: required
testbench: required # UVM or cocotb, self-checking
sva_properties: preferred
coverage_report: preferred # functional + code, text export
lint_cdc_reports: preferred
synthesis_log: optional # area, timing, latch warnings
git_history: required # with linked bug IDs
eda_context: [tcl_run_scripts, sdc_constraints, design_specs, wiki_pages]
tools_assumed:
simulators: [Verilator, Icarus, "commercial: state version"]
formal: "state tool and version"
exclusions: [third-party IP cores, IEEE 1735 encrypted blocks,
foundry PDK, std-cell and memory-compiler views]
deliverables: [component_inventory.csv, export_classification_memo]
State which tools the eval tasks assume. A testbench that only runs under a commercial simulator is useless to a team without that license, and Verilator rejects some constructs other simulators accept. The code dataset request specification guide covers the general fields, and evaluating a code dataset sample covers what to check when a sample arrives: does it compile, do testbenches pass, and do coverage numbers reproduce.
How SourceX handles RTL requests
SourceX sources operational datasets, including engineering records, from US companies and manages the licensing process, serving AI teams wherever they are based. Data is sourced on request rather than held in stock, so a request for RTL with testbenches does not guarantee a match. Buyers describe the data, SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, ships with per-dataset diligence materials, and is delivered under a license that defines records, uses, term and delivery, through private access-controlled workflows after an executed agreement. You can start a buyer request once you know which artifacts you need.
Related pages: proprietary code datasets with full git history, CAD and PCB design datasets for board-level files, manufacturing buyers, the code and software engineering data hub, and training data quality assessment.
Request RTL and verification data for your model
If you are training or evaluating an RTL generation, verification or EDA scripting model, describe the languages, verification artifacts and exclusions you need. SourceX runs Find, Assess, Agree, Transact and Manage with US suppliers, and nothing is contracted until a supplier agrees. Describe your RTL data request.
Sources
- arXiv:2604.27415, "ChipLingo: A Systematic Training Framework for Large Language Models in EDA" (2026). https://arxiv.org/pdf/2604.27415
- arXiv:2505.24183, "QiMeng-CodeV-R1: Reasoning-Enhanced Verilog Generation" (2025). https://arxiv.org/pdf/2505.24183
- Longpre et al. (arXiv:2310.16787), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
- Srivastava et al. (arXiv:2206.04615), "Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models" (2022). https://arxiv.org/pdf/2206.04615
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.