Code and software engineering data
Data Pipeline Code Datasets: dbt Models, Airflow DAGs and Spark Jobs for AI
Quick answer
A useful data engineering code dataset is a set of whole, real transformation projects, not loose SQL snippets. Each unit should be a dbt project with its compiled manifest.json and data tests, the Airflow or Dagster DAGs that schedule it, Spark or SQL jobs, and run history that links failing runs to the commits that fixed them. Buyers should require lineage, scrubbed credentials, and a stated plan for evaluating agents without access to production warehouses.
By SourceX Editorial · Updated
Why pipeline projects are a different training unit from SQL queries
Agents that build and repair pipelines must reason across files, model dependencies and metadata at once, so the unit of data is the project. Spider 2.0, an enterprise text-to-SQL benchmark of 632 problems, is built on tasks that require navigating project codebases, warehouse schemas and documentation rather than answering one question with one query [1]. A dbt model that compiles in isolation can still break three downstream marts when a column is renamed. Only project-level data, with ref() and source() edges intact, shows a model that dependency.
That distinction sets this page apart from neighboring categories. Query-level data such as analyst SQL logs and stored procedures belongs in production SQL query corpora. Build-level failures in application CI belong in CI build failure logs. Terraform and Kubernetes configs that provision the warehouse belong in infrastructure-as-code datasets.
What a complete dbt, Airflow or Spark project should contain
A complete record includes source code, compiled artifacts, tests, run results and version history for the same project. Ask for each of these layers explicitly, because suppliers often export only the models/ folder.
- dbt project files:
dbt_project.yml,models/(staging, intermediate, marts),schema.ymlwith column descriptions and tests,macros/,seeds/,snapshots/,packages.ymlwith pinned versions, andprofiles.ymlwith credentials stripped. - dbt artifacts: manifest.json, the compiled representation of every node and relationship; run_results.json, which records status and timing for each executed node by
unique_id; and catalog.json fromdbt docs generate. Run results reference nodes only byunique_id, so without the matching manifest they cannot be resolved. - Orchestration: Airflow DAG files with operators, sensors,
schedule, retries and SLAs; task instance states and logs; or the Dagster or Prefect equivalents. - Spark and SQL jobs: PySpark or Scala job code,
spark-submitconfigurations, partitioning and file layout choices, plus the Spark UI or event logs where available. - Lineage: OpenLineage run events, if the supplier emits them. Each event carries an event type, timestamp, run ID and facets, and Airflow events add facets such as nominal start time and parent run.
- Git history: full commit history with messages, pull request descriptions and review comments, so changes map to intent. See delivering repositories with full history.
Turning failing runs and fixing commits into repair tasks
Failure-fix history is the highest-value slice for repair agents, because it supplies an observed break, its error and a human fix. Aggregated run_results.json files can show test failure rates over time, and each failure can be paired with the next commit that turned the node green. Typical breaks include schema drift from an upstream source, a unique or not_null test failing after a join fan-out, an incremental model with a wrong unique_key, a Spark job hitting skew or out-of-memory errors, and an Airflow sensor timing out on a late upstream file.
Treat these pairs as hypotheses until validated. A commit after a failure is not always the fix, and some fixes are config changes made outside Git, such as a warehouse grant. Require the supplier or your own pipeline to confirm that the fix commit changes the failing node, or one of its parents, and that a later run passed.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"task_id": "pipe-repair-00412",
"project": "retail_analytics_dbt",
"base_commit": "a91f3c2",
"fix_commit": "d07e5b8",
"failing_node": "model.retail_analytics.fct_orders",
"failure": {
"artifact": "run_results.json",
"status": "fail",
"test": "test.retail_analytics.unique_fct_orders_order_id",
"message": "Got 214 results, configured to fail if != 0"
},
"root_cause_label": "join_fan_out",
"changed_files": ["models/marts/fct_orders.sql"],
"upstream_nodes": ["model.retail_analytics.stg_order_items"],
"verification": ["dbt compile", "dbt test --select fct_orders+ on masked sample"],
"secrets_scan": "passed",
"pii_handling": "sample rows masked; method recorded"
}
Evaluating pipeline agents when production data cannot ship
Most suppliers will not ship warehouse contents, so decide the evaluation method before you sign. Executable environments are the gold standard for software agents: SWE-Gym pairs each task with a codebase, a runtime and unit tests [2]. For pipelines, the equivalent needs data, which forces a choice among three options.
| Evaluation mode | What it needs | What it can verify | Main gap |
|---|---|---|---|
| Static: compile and lineage | manifest.json, project files | Model compiles; DAG edges and ref() graph are correct | Cannot confirm row-level correctness |
| Test definitions on synthetic data | Schema, test YAML, generated seeds | Tests pass on data matching declared types and constraints | Synthetic data misses real skew and nulls |
| Masked sample warehouse | De-identified sample tables, DuckDB or Spark local | Runs end to end; reproduces many failures | Masking can remove the pattern that caused the break |
Hold out whole projects, not individual tasks, for evaluation. Tasks from one project share macros and naming conventions, and splitting them across train and test inflates scores. For public-repo overlap, apply the checks in code benchmark contamination.
Secrets, personal data and ownership in pipeline repositories
Pipeline code often embeds connection strings and service account credentials, so scan the full history, not just HEAD. Look for warehouse passwords in profiles.yml, Airflow connection URIs in DAG files or airflow.cfg, Snowflake and BigQuery service account JSON, S3 keys in Spark configs, and Fernet keys. Apply the standard in secrets removal for code datasets, and require rotation evidence, not only redaction.
Personal data leaks through seeds, snapshots, test fixtures, sample rows in logs and column names. SourceX removes or replaces personal details such as names, emails, phones and account numbers before delivery, records the method and checks a sample, but no de-identification method is perfect. Confirm ownership too: pipelines often include contractor work and vendored packages, covered in code ownership due diligence.
Documentation and acceptance checks before you license
Ask for dataset documentation that a reviewer can parse, then test a sample against it. Croissant-RAI gives a machine-readable format for provenance and responsible-AI metadata [3], and NIST AI RMF 1.0, a voluntary framework organized around GOVERN, MAP, MEASURE and MANAGE, gives reviewers a structure for recording data risks [4]. Use evaluating a code dataset sample for the general method and add these pipeline checks.
Illustrative example: invented to show structure; it does not describe an available dataset.
Acceptance checklist for a pipeline project sample
dbt parseanddbt compilesucceed with the pinned dbt and adapter versions.- Every
unique_idin run_results.json resolves in the delivered manifest.json. - The DAG files import cleanly in the stated Airflow version with no missing connections beyond documented placeholders.
- A secrets scan across all commits returns no live credentials.
- Seeds, snapshots and logs pass a PII scan with the masking method recorded.
- At least a defined share of failure-fix pairs pass re-verification on the agreed evaluation mode.
- License terms state allowed uses, term and delivery for code, artifacts and logs separately.
For writing the request itself, see how to specify a code dataset request. For broader proprietary code, SourceX's pages on proprietary codebases and software engineering histories cover adjacent sources.
How SourceX sources data engineering code
SourceX sources operational datasets, including engineering records, from US companies on request, and manages licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. You describe the data you need, such as dbt projects with manifests and two years of run history, and SourceX looks for US businesses that hold it. Every release is approved by the supplying company.
Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Delivery uses private, access-controlled workflows after an executed agreement. You can describe your pipeline data needs to SourceX at any stage. The code data buyer's map and the AI data hub cover related categories.
Request data engineering code datasets
If your agents need real dbt projects, orchestration DAGs or Spark jobs with tests and failure-fix history, describe the data and SourceX will look for US companies that hold it. Pricing and allowed uses are agreed per deal in a license, and diligence materials are prepared per dataset. Start a buyer request.
Sources
- arXiv (Lei et al.), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
- arXiv (Pan et al.), "Training Software Engineering Agents and Verifiers with SWE-Gym" (2024). https://arxiv.org/abs/2412.21139v1
- arXiv (Jain et al., MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.