Skip to content

Code and software engineering data

Data Pipeline Code Datasets: dbt Models, Airflow DAGs and Spark Jobs for AI

Quick answer

A useful data engineering code dataset is a set of whole, real transformation projects, not loose SQL snippets. Each unit should be a dbt project with its compiled manifest.json and data tests, the Airflow or Dagster DAGs that schedule it, Spark or SQL jobs, and run history that links failing runs to the commits that fixed them. Buyers should require lineage, scrubbed credentials, and a stated plan for evaluating agents without access to production warehouses.

By SourceX Editorial · Updated

Why pipeline projects are a different training unit from SQL queries

Agents that build and repair pipelines must reason across files, model dependencies and metadata at once, so the unit of data is the project. Spider 2.0, an enterprise text-to-SQL benchmark of 632 problems, is built on tasks that require navigating project codebases, warehouse schemas and documentation rather than answering one question with one query [1]. A dbt model that compiles in isolation can still break three downstream marts when a column is renamed. Only project-level data, with ref() and source() edges intact, shows a model that dependency.

That distinction sets this page apart from neighboring categories. Query-level data such as analyst SQL logs and stored procedures belongs in production SQL query corpora. Build-level failures in application CI belong in CI build failure logs. Terraform and Kubernetes configs that provision the warehouse belong in infrastructure-as-code datasets.

What a complete dbt, Airflow or Spark project should contain

A complete record includes source code, compiled artifacts, tests, run results and version history for the same project. Ask for each of these layers explicitly, because suppliers often export only the models/ folder.

  • dbt project files: dbt_project.yml, models/ (staging, intermediate, marts), schema.yml with column descriptions and tests, macros/, seeds/, snapshots/, packages.yml with pinned versions, and profiles.yml with credentials stripped.
  • dbt artifacts: manifest.json, the compiled representation of every node and relationship; run_results.json, which records status and timing for each executed node by unique_id; and catalog.json from dbt docs generate. Run results reference nodes only by unique_id, so without the matching manifest they cannot be resolved.
  • Orchestration: Airflow DAG files with operators, sensors, schedule, retries and SLAs; task instance states and logs; or the Dagster or Prefect equivalents.
  • Spark and SQL jobs: PySpark or Scala job code, spark-submit configurations, partitioning and file layout choices, plus the Spark UI or event logs where available.
  • Lineage: OpenLineage run events, if the supplier emits them. Each event carries an event type, timestamp, run ID and facets, and Airflow events add facets such as nominal start time and parent run.
  • Git history: full commit history with messages, pull request descriptions and review comments, so changes map to intent. See delivering repositories with full history.

Turning failing runs and fixing commits into repair tasks

Failure-fix history is the highest-value slice for repair agents, because it supplies an observed break, its error and a human fix. Aggregated run_results.json files can show test failure rates over time, and each failure can be paired with the next commit that turned the node green. Typical breaks include schema drift from an upstream source, a unique or not_null test failing after a join fan-out, an incremental model with a wrong unique_key, a Spark job hitting skew or out-of-memory errors, and an Airflow sensor timing out on a late upstream file.

Treat these pairs as hypotheses until validated. A commit after a failure is not always the fix, and some fixes are config changes made outside Git, such as a warehouse grant. Require the supplier or your own pipeline to confirm that the fix commit changes the failing node, or one of its parents, and that a later run passed.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "task_id": "pipe-repair-00412",
  "project": "retail_analytics_dbt",
  "base_commit": "a91f3c2",
  "fix_commit": "d07e5b8",
  "failing_node": "model.retail_analytics.fct_orders",
  "failure": {
    "artifact": "run_results.json",
    "status": "fail",
    "test": "test.retail_analytics.unique_fct_orders_order_id",
    "message": "Got 214 results, configured to fail if != 0"
  },
  "root_cause_label": "join_fan_out",
  "changed_files": ["models/marts/fct_orders.sql"],
  "upstream_nodes": ["model.retail_analytics.stg_order_items"],
  "verification": ["dbt compile", "dbt test --select fct_orders+ on masked sample"],
  "secrets_scan": "passed",
  "pii_handling": "sample rows masked; method recorded"
}

Evaluating pipeline agents when production data cannot ship

Most suppliers will not ship warehouse contents, so decide the evaluation method before you sign. Executable environments are the gold standard for software agents: SWE-Gym pairs each task with a codebase, a runtime and unit tests [2]. For pipelines, the equivalent needs data, which forces a choice among three options.

Evaluation modeWhat it needsWhat it can verifyMain gap
Static: compile and lineagemanifest.json, project filesModel compiles; DAG edges and ref() graph are correctCannot confirm row-level correctness
Test definitions on synthetic dataSchema, test YAML, generated seedsTests pass on data matching declared types and constraintsSynthetic data misses real skew and nulls
Masked sample warehouseDe-identified sample tables, DuckDB or Spark localRuns end to end; reproduces many failuresMasking can remove the pattern that caused the break

Hold out whole projects, not individual tasks, for evaluation. Tasks from one project share macros and naming conventions, and splitting them across train and test inflates scores. For public-repo overlap, apply the checks in code benchmark contamination.

Secrets, personal data and ownership in pipeline repositories

Pipeline code often embeds connection strings and service account credentials, so scan the full history, not just HEAD. Look for warehouse passwords in profiles.yml, Airflow connection URIs in DAG files or airflow.cfg, Snowflake and BigQuery service account JSON, S3 keys in Spark configs, and Fernet keys. Apply the standard in secrets removal for code datasets, and require rotation evidence, not only redaction.

Personal data leaks through seeds, snapshots, test fixtures, sample rows in logs and column names. SourceX removes or replaces personal details such as names, emails, phones and account numbers before delivery, records the method and checks a sample, but no de-identification method is perfect. Confirm ownership too: pipelines often include contractor work and vendored packages, covered in code ownership due diligence.

Documentation and acceptance checks before you license

Ask for dataset documentation that a reviewer can parse, then test a sample against it. Croissant-RAI gives a machine-readable format for provenance and responsible-AI metadata [3], and NIST AI RMF 1.0, a voluntary framework organized around GOVERN, MAP, MEASURE and MANAGE, gives reviewers a structure for recording data risks [4]. Use evaluating a code dataset sample for the general method and add these pipeline checks.

Illustrative example: invented to show structure; it does not describe an available dataset.

Acceptance checklist for a pipeline project sample

  1. dbt parse and dbt compile succeed with the pinned dbt and adapter versions.
  2. Every unique_id in run_results.json resolves in the delivered manifest.json.
  3. The DAG files import cleanly in the stated Airflow version with no missing connections beyond documented placeholders.
  4. A secrets scan across all commits returns no live credentials.
  5. Seeds, snapshots and logs pass a PII scan with the masking method recorded.
  6. At least a defined share of failure-fix pairs pass re-verification on the agreed evaluation mode.
  7. License terms state allowed uses, term and delivery for code, artifacts and logs separately.

For writing the request itself, see how to specify a code dataset request. For broader proprietary code, SourceX's pages on proprietary codebases and software engineering histories cover adjacent sources.

How SourceX sources data engineering code

SourceX sources operational datasets, including engineering records, from US companies on request, and manages licensing and ongoing purchases. Nothing is held in stock, and a request does not guarantee a match. You describe the data you need, such as dbt projects with manifests and two years of run history, and SourceX looks for US businesses that hold it. Every release is approved by the supplying company.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Delivery uses private, access-controlled workflows after an executed agreement. You can describe your pipeline data needs to SourceX at any stage. The code data buyer's map and the AI data hub cover related categories.

Request data engineering code datasets

If your agents need real dbt projects, orchestration DAGs or Spark jobs with tests and failure-fix history, describe the data and SourceX will look for US companies that hold it. Pricing and allowed uses are agreed per deal in a license, and diligence materials are prepared per dataset. Start a buyer request.

Sources

  1. arXiv (Lei et al.), "Spider 2.0: Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows" (2024). https://www.arxiv.org/pdf/2411.07763
  2. arXiv (Pan et al.), "Training Software Engineering Agents and Verifiers with SWE-Gym" (2024). https://arxiv.org/abs/2412.21139v1
  3. arXiv (Jain et al., MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  4. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1" (2023). https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data