Skip to content

Code and software engineering data

Infrastructure-as-Code Datasets: Terraform, Kubernetes and Pipeline Configs for AI

Quick answer

An infrastructure-as-code dataset for AI is production configuration (Terraform or OpenTofu HCL, CloudFormation, Kubernetes manifests, Helm charts, Ansible playbooks and CI pipeline YAML) paired with whatever verified it: terraform validate and plan output, policy-as-code findings, and the commits that fixed them. Public corpora give you syntax and module patterns. Private, licensed IaC adds real change history and remediation pairs, which turn configuration into graded tasks for generation, review and repair agents.

By SourceX Editorial · Updated

What public IaC corpora cover, and where they stop

Public IaC corpora are good for pretraining and syntax coverage but rarely carry the verification signal an agent needs. TerraDS was published to fill the gap in large-scale Terraform corpora [1], and PIPr collects Pulumi and CDK programs written in general-purpose languages from public GitHub repositories, with a permissively licensed subset [2]. Multi-IaC-Eval goes a step further and frames the task as template, modification request and updated template triplets across CloudFormation, Terraform and CDK [3].

What these sources lack is the operational context of a real platform team. Public repos skew toward examples, starter modules and personal projects, not the drift fixes, policy exceptions and incident-driven changes that dominate production history. They also tend to be heavily mirrored, which matters for evaluation; see code benchmark contamination before using any public IaC task as a held-out test.

For the broader trade-off between permissive corpora and licensed private code, the open versus licensed code comparison covers rights, opt-outs and coverage in detail.

Which IaC artifacts make the dataset trainable

The useful unit is a change, not a file: a diff to configuration plus the machine-checkable evidence before and after it. Ask suppliers to package these artifact types together, keyed to a commit SHA.

  • Configuration source: .tf and .tfvars (with values redacted), CloudFormation YAML or JSON templates, Kubernetes manifests, Helm Chart.yaml, values.yaml and templates, Kustomize overlays, Ansible playbooks and roles, plus pipeline definitions such as .github/workflows/*.yml, .gitlab-ci.yml and Jenkinsfiles.
  • Validation output: results of terraform validate, helm lint, kubeconform or ansible-lint, recorded with tool versions.
  • Plan output: terraform show -json renders a plan as machine-readable JSON with resource changes and a configuration representation, which is far easier to grade than human-readable plan text.
  • Policy findings: Checkov, Trivy (which absorbed tfsec), OPA/Conftest or Kyverno results, with rule IDs, severity and the file and line referenced.
  • Fix commits: the commit that cleared each finding, with its message, linked pull request and review comments.

Pipeline definitions belong here as configuration code. Their run logs and failure traces are a separate data type, closer to production error and stack-trace-to-fix data. Data-platform code such as Airflow DAGs and dbt models has its own page on data pipeline code datasets.

Why state files and saved plans are the main exposure

Terraform state and saved plan files should be excluded from any IaC dataset unless they have been regenerated against a sandbox. State records resource attributes as JSON, and depending on the resources it can hold values such as initial database passwords; marking a variable sensitive = true hides it from CLI output but does not, by itself, keep the value out of state. Confirm the behavior of the supplier's Terraform or OpenTofu version against HashiCorp and OpenTofu documentation before scoping. HashiCorp's own guidance on AI pipelines is to scan repositories, storage and logs and redact hard-coded credentials before data is ingested, because a secret learned by a model is hard to remove afterward [4].

Even without secrets, IaC is a map of the supplier's attack surface. Security groups, IAM policies, VPC peering, public load balancers and bucket policies describe exactly what is exposed and how. Expect suppliers' security teams to review IaC more strictly than application code, and plan for redaction to be part of the scope from the first conversation.

Run full-history secret scanning before anything leaves the supplier, since a credential deleted in a later commit is still in the history. The methods and verification steps are covered in secrets in code datasets.

Redaction that keeps the code valid

Redaction for IaC must preserve parseability and internal consistency, or plan and validate stop working. Replace each sensitive literal with a stable placeholder that is consistent across every file and commit in the delivery.

  • AWS account IDs: map to a fixed set of synthetic 12-digit IDs so ARNs still parse.
  • ARNs and resource names: rewrite partition, account and name segments while keeping the service and resource-type segments.
  • CIDR ranges: remap into reserved documentation or private ranges while keeping prefix lengths and non-overlap, so subnet math and peering rules still hold.
  • Hostnames and domains: map to example.internal-style names consistently across DNS records, ingress hosts and TLS certificates.
  • Backend blocks: strip bucket names, lock tables and workspace names from backend and cloud blocks.

After redaction, re-run terraform validate, helm template and the policy scanners and compare finding counts with the pre-redaction run. A redaction that changes which rules fire has changed the label, not just the data.

Licenses inside an IaC repository

A private infrastructure repository usually mixes internal modules with vendored public ones, and each carries its own license. Modules pulled from the Terraform or OpenTofu registry, community Helm charts and Ansible Galaxy roles are often committed into the tree, sometimes modified. Ask for a manifest that separates first-party paths from vendored ones, with each vendored module's source and license.

Copyleft components need the same review as in application code; see copyleft contamination in licensed code. Authorship questions for contractor-written modules are covered in code ownership due diligence.

Request template for an IaC dataset

A precise request lists tools, versions, artifacts and the verification you need, so a supplier can say quickly whether they hold it. Adapt the general guidance in how to specify a code dataset request with these IaC-specific fields.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample value
Tools and versionsTerraform 1.5+ or OpenTofu; Helm 3; Kubernetes manifests; Ansible 2.14+; GitHub Actions and GitLab CI YAML
Clouds and providersAWS and GCP providers; Kubernetes provider; no on-prem vSphere
HistoryFull git history for 24 months; merge commits and PR metadata
Paired evidenceterraform show -json plan per commit (regenerated in sandbox); Checkov and Conftest findings with rule IDs
Remediation pairsCommits that clear a policy finding, linked to the finding and PR review comments
ExcludedState files, saved binary plans, .tfvars values, CI run logs
RedactionConsistent placeholders for account IDs, ARNs, CIDRs, hostnames; method documented
FormatJSON Lines, one record per change [5]; dataset card with provenance fields [6]
UseSFT and evaluation of an IaC remediation agent

An individual record might look like this:

Illustrative example: invented to show structure; it does not describe an available dataset.

{"change_id": "c-000412", "tool": "terraform", "files": ["modules/storage/s3.tf"], "finding_before": {"scanner": "checkov", "rule": "CKV_AWS_19", "severity": "HIGH", "line": 14}, "fix_diff": "...", "validate_after": "success", "findings_after": [], "plan_json_ref": "plans/c-000412.json", "redaction_method": "placeholder-map-v2"}

Evaluating IaC agents without touching live accounts

IaC evaluation should run plan and policy checks against mocked or sandbox providers, never the supplier's live cloud accounts. Graders can check that the output parses, that validate passes, that the targeted policy finding clears, that no new findings appear, and that the plan JSON shows no unintended destroy or replace actions on resources outside the task.

Watch for two failure modes. Agents learn to clear findings by suppressing them, for example adding #checkov:skip comments or broadening exceptions, so graders must treat inline suppressions as failures unless the task allows them. And provider version drift changes plan output, so pin provider versions in the evaluation harness and record them per task. For sample review before you license, use evaluating a code dataset sample, and for security-focused fixes see vulnerability fix pairs.

How SourceX handles IaC requests

SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Engineering records, including codebases with history, are among the kinds of data it looks for; see proprietary code datasets and, for the incident side of platform work, ITSM ticket datasets. Nothing is held in stock, and a request does not guarantee a match.

The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, through private, access-controlled workflows after an executed agreement and supplier approval. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. You can describe your IaC data requirement to SourceX as a buyer, and IT services teams can also review buyers in IT managed services.

More code data types are mapped in the code and software engineering data hub and the main AI data guide.

Sourcing infrastructure-as-code data for your agent

Describe the tools, history and verification evidence your IaC agent needs, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and terms are agreed per deal. Start an infrastructure-as-code data request.

Sources

  1. MSR 2025 Data and Tool Showcase Track, "TerraDS: A Dataset for Terraform HCL Programs" (2025). https://2025.msrconf.org/details/msr-2025-data-and-tool-showcase-track/12/TerraDS-A-Dataset-for-Terraform-HCL-Programs
  2. MSR 2024 Data and Tool Showcase Track, "The PIPr Dataset of Public Infrastructure as Code Programs" (2024). https://2024.msrconf.org/details/msr-2024-data-and-tool-showcase-track/26/The-PIPr-Dataset-of-Public-Infrastructure-as-Code-Programs
  3. arXiv, "Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats" (2025). https://arxiv.org/html/2509.05303v1
  4. HashiCorp, "Integrating secret hygiene into AI and ML workflows". https://www.hashicorp.com/blog/integrating-secret-hygiene-into-ai-and-ml-workflows
  5. jsonlines.org, "JSON Lines". https://jsonlines.org/
  6. Jain et al. (MLCommons Croissant RAI task force), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data