Code and software engineering data
Infrastructure-as-Code Datasets: Terraform, Kubernetes and Pipeline Configs for AI
Quick answer
An infrastructure-as-code dataset for AI is production configuration (Terraform or OpenTofu HCL, CloudFormation, Kubernetes manifests, Helm charts, Ansible playbooks and CI pipeline YAML) paired with whatever verified it: terraform validate and plan output, policy-as-code findings, and the commits that fixed them. Public corpora give you syntax and module patterns. Private, licensed IaC adds real change history and remediation pairs, which turn configuration into graded tasks for generation, review and repair agents.
By SourceX Editorial · Updated
What public IaC corpora cover, and where they stop
Public IaC corpora are good for pretraining and syntax coverage but rarely carry the verification signal an agent needs. TerraDS was published to fill the gap in large-scale Terraform corpora [1], and PIPr collects Pulumi and CDK programs written in general-purpose languages from public GitHub repositories, with a permissively licensed subset [2]. Multi-IaC-Eval goes a step further and frames the task as template, modification request and updated template triplets across CloudFormation, Terraform and CDK [3].
What these sources lack is the operational context of a real platform team. Public repos skew toward examples, starter modules and personal projects, not the drift fixes, policy exceptions and incident-driven changes that dominate production history. They also tend to be heavily mirrored, which matters for evaluation; see code benchmark contamination before using any public IaC task as a held-out test.
For the broader trade-off between permissive corpora and licensed private code, the open versus licensed code comparison covers rights, opt-outs and coverage in detail.
Which IaC artifacts make the dataset trainable
The useful unit is a change, not a file: a diff to configuration plus the machine-checkable evidence before and after it. Ask suppliers to package these artifact types together, keyed to a commit SHA.
- Configuration source:
.tfand.tfvars(with values redacted), CloudFormation YAML or JSON templates, Kubernetes manifests, HelmChart.yaml,values.yamland templates, Kustomize overlays, Ansible playbooks and roles, plus pipeline definitions such as.github/workflows/*.yml,.gitlab-ci.ymland Jenkinsfiles. - Validation output: results of
terraform validate,helm lint,kubeconformoransible-lint, recorded with tool versions. - Plan output:
terraform show -jsonrenders a plan as machine-readable JSON with resource changes and a configuration representation, which is far easier to grade than human-readable plan text. - Policy findings: Checkov, Trivy (which absorbed tfsec), OPA/Conftest or Kyverno results, with rule IDs, severity and the file and line referenced.
- Fix commits: the commit that cleared each finding, with its message, linked pull request and review comments.
Pipeline definitions belong here as configuration code. Their run logs and failure traces are a separate data type, closer to production error and stack-trace-to-fix data. Data-platform code such as Airflow DAGs and dbt models has its own page on data pipeline code datasets.
Why state files and saved plans are the main exposure
Terraform state and saved plan files should be excluded from any IaC dataset unless they have been regenerated against a sandbox. State records resource attributes as JSON, and depending on the resources it can hold values such as initial database passwords; marking a variable sensitive = true hides it from CLI output but does not, by itself, keep the value out of state. Confirm the behavior of the supplier's Terraform or OpenTofu version against HashiCorp and OpenTofu documentation before scoping. HashiCorp's own guidance on AI pipelines is to scan repositories, storage and logs and redact hard-coded credentials before data is ingested, because a secret learned by a model is hard to remove afterward [4].
Even without secrets, IaC is a map of the supplier's attack surface. Security groups, IAM policies, VPC peering, public load balancers and bucket policies describe exactly what is exposed and how. Expect suppliers' security teams to review IaC more strictly than application code, and plan for redaction to be part of the scope from the first conversation.
Run full-history secret scanning before anything leaves the supplier, since a credential deleted in a later commit is still in the history. The methods and verification steps are covered in secrets in code datasets.
Redaction that keeps the code valid
Redaction for IaC must preserve parseability and internal consistency, or plan and validate stop working. Replace each sensitive literal with a stable placeholder that is consistent across every file and commit in the delivery.
- AWS account IDs: map to a fixed set of synthetic 12-digit IDs so ARNs still parse.
- ARNs and resource names: rewrite partition, account and name segments while keeping the service and resource-type segments.
- CIDR ranges: remap into reserved documentation or private ranges while keeping prefix lengths and non-overlap, so subnet math and peering rules still hold.
- Hostnames and domains: map to
example.internal-style names consistently across DNS records, ingress hosts and TLS certificates. - Backend blocks: strip bucket names, lock tables and workspace names from
backendandcloudblocks.
After redaction, re-run terraform validate, helm template and the policy scanners and compare finding counts with the pre-redaction run. A redaction that changes which rules fire has changed the label, not just the data.
Licenses inside an IaC repository
A private infrastructure repository usually mixes internal modules with vendored public ones, and each carries its own license. Modules pulled from the Terraform or OpenTofu registry, community Helm charts and Ansible Galaxy roles are often committed into the tree, sometimes modified. Ask for a manifest that separates first-party paths from vendored ones, with each vendored module's source and license.
Copyleft components need the same review as in application code; see copyleft contamination in licensed code. Authorship questions for contractor-written modules are covered in code ownership due diligence.
Request template for an IaC dataset
A precise request lists tools, versions, artifacts and the verification you need, so a supplier can say quickly whether they hold it. Adapt the general guidance in how to specify a code dataset request with these IaC-specific fields.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field | Example value |
|---|---|
| Tools and versions | Terraform 1.5+ or OpenTofu; Helm 3; Kubernetes manifests; Ansible 2.14+; GitHub Actions and GitLab CI YAML |
| Clouds and providers | AWS and GCP providers; Kubernetes provider; no on-prem vSphere |
| History | Full git history for 24 months; merge commits and PR metadata |
| Paired evidence | terraform show -json plan per commit (regenerated in sandbox); Checkov and Conftest findings with rule IDs |
| Remediation pairs | Commits that clear a policy finding, linked to the finding and PR review comments |
| Excluded | State files, saved binary plans, .tfvars values, CI run logs |
| Redaction | Consistent placeholders for account IDs, ARNs, CIDRs, hostnames; method documented |
| Format | JSON Lines, one record per change [5]; dataset card with provenance fields [6] |
| Use | SFT and evaluation of an IaC remediation agent |
An individual record might look like this:
Illustrative example: invented to show structure; it does not describe an available dataset.
{"change_id": "c-000412", "tool": "terraform", "files": ["modules/storage/s3.tf"], "finding_before": {"scanner": "checkov", "rule": "CKV_AWS_19", "severity": "HIGH", "line": 14}, "fix_diff": "...", "validate_after": "success", "findings_after": [], "plan_json_ref": "plans/c-000412.json", "redaction_method": "placeholder-map-v2"}
Evaluating IaC agents without touching live accounts
IaC evaluation should run plan and policy checks against mocked or sandbox providers, never the supplier's live cloud accounts. Graders can check that the output parses, that validate passes, that the targeted policy finding clears, that no new findings appear, and that the plan JSON shows no unintended destroy or replace actions on resources outside the task.
Watch for two failure modes. Agents learn to clear findings by suppressing them, for example adding #checkov:skip comments or broadening exceptions, so graders must treat inline suppressions as failures unless the task allows them. And provider version drift changes plan output, so pin provider versions in the evaluation harness and record them per task. For sample review before you license, use evaluating a code dataset sample, and for security-focused fixes see vulnerability fix pairs.
How SourceX handles IaC requests
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Engineering records, including codebases with history, are among the kinds of data it looks for; see proprietary code datasets and, for the incident side of platform work, ITSM ticket datasets. Nothing is held in stock, and a request does not guarantee a match.
The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your IaC data requirement to SourceX as a buyer, and IT services teams can also review buyers in IT managed services.
More code data types are mapped in the code and software engineering data hub and the main AI data guide.
Sourcing infrastructure-as-code data for your agent
Describe the tools, history and verification evidence your IaC agent needs, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and terms are agreed per deal. Start an infrastructure-as-code data request.
Sources
- MSR 2025 Data and Tool Showcase Track, "TerraDS: A Dataset for Terraform HCL Programs" (2025). https://2025.msrconf.org/details/msr-2025-data-and-tool-showcase-track/12/TerraDS-A-Dataset-for-Terraform-HCL-Programs
- MSR 2024 Data and Tool Showcase Track, "The PIPr Dataset of Public Infrastructure as Code Programs" (2024). https://2024.msrconf.org/details/msr-2024-data-and-tool-showcase-track/26/The-PIPr-Dataset-of-Public-Infrastructure-as-Code-Programs
- arXiv, "Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats" (2025). https://arxiv.org/html/2509.05303v1
- HashiCorp, "Integrating secret hygiene into AI and ML workflows". https://www.hashicorp.com/blog/integrating-secret-hygiene-into-ai-and-ml-workflows
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- Jain et al. (MLCommons Croissant RAI task force), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.