Skip to content

Code and software engineering data

CI Build Logs and Build Failure Datasets for Coding Agents

Quick answer

A CI build failure dataset for coding agents links each failing pipeline run to its raw log, a labeled failure span, a root-cause class, the commit that fixed it and the next passing run, plus rerun history that separates flaky tests from real regressions. Public log corpora rarely provide that chain, so teams training build-repair or triage agents usually license CI histories from companies that kept them. Specify the linkage, log limits, retention window and secrets scanning before you buy.

By SourceX Editorial · Updated

Why public log datasets do not cover build repair

Public log datasets mostly capture system and operations logs, not CI runs tied to code changes. Loghub, the standard collection, gathers logs from distributed systems, supercomputers, operating systems and server applications for anomaly detection and parsing research [1]. That is useful for AIOps, but it has no commit, no diff and no notion of "the change that made the build green again."

Academic build-log datasets do exist, mostly mined from public open-source projects on hosted CI. They tend to be small, annotated for log-chunk retrieval or failure classification, and rarely pair a failure with the commit that fixed it. They also inherit the core difficulty of the domain: logs are long, and generic diff tools do a poor job of isolating what changed between a failing and passing run.

That gap is why buyers look to licensed operational data. Companies running GitHub Actions, GitLab CI, Jenkins, Buildkite or CircleCI accumulate exactly the failing-run-to-fix chain that public corpora lack. For the wider landscape of code data types, see the code and software engineering data buyer's map.

The record unit: failing run, log, diagnosis, fix, passing run

The useful unit is a linked chain, and its value depends on how reliably each link was established. A failing run on commit A produces a log; someone diagnoses the failure; commit B fixes it; a run on B passes. If any link is guessed rather than recorded, your reward signal inherits the noise.

Ask every supplier two questions: what share of failing runs are linked to a fixing commit, and how were links inferred? Strong links come from the same branch or pull request with an explicit revert, a "fix CI" commit referencing the run, or a PR review thread. Weak links come from "next green run on the default branch," which can bundle unrelated changes.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueWhy it matters
run_id / attempt88213 / 1Distinguishes reruns of the same run
ci_systemGitHub ActionsLog format and step structure differ by system
head_shaa41f9c2Commit under test
job / steptest-linux / pytestLocalizes the failure
conclusionfailureOutcome label
log_urilogs/88213-1.txt.zstRaw log, compressed
failure_spanlines 4,812-4,866Supervision target for triage
failure_classdependency_resolutionRoot-cause label for classifiers
fix_sha7c0de11Repair target
link_methodsame_pr_followup_commitConfidence of the fix link
verify_run_id88240 (success)Proof the fix turned the build green
rerun_outcomes[failure, success] on same shaFlakiness evidence

This chain is narrower than a full issue-to-PR record. If you need issues, reviews and CI together, the issue-to-fix pairs guide and the software engineering histories page cover that broader record.

Flaky test labels come from reruns, not from guesses

Flakiness is best labeled from reruns of the same commit that produced different outcomes. A test that fails then passes on an identical SHA and environment is a flaky candidate; a test that fails consistently until a code change is a real regression. Without rerun history, an agent trained to "fix" every failure will learn to edit code that was never broken.

Ask for attempt numbers, rerun triggers (manual, automatic retry, scheduled) and any quarantine or retry annotations from the test framework. Ask whether the runner image, cache keys and dependency lockfile were identical across attempts, because an environment change can masquerade as flakiness. Labels such as flaky, infra, real_regression and unknown should carry a stated rule, not an annotator's impression.

Retention windows limit usable CI history

Usable CI history is often shorter than code history because hosted CI services delete logs, artifacts and sometimes run metadata after a configurable retention period. Retention defaults and maximums differ by provider and plan and change over time, so check each provider's current documentation as of the period the supplier's data covers. Lengthening retention generally does not bring back logs that were already evicted.

So a repository with ten years of commits may carry only months of CI logs. Ask each supplier which CI providers they used, what retention applied in each period, and whether logs were exported to their own storage (S3 buckets, Elasticsearch, Datadog, Splunk). Self-hosted Jenkins or GitLab installs often hold longer history, but check their own log-rotation settings.

Logs are long and noisy: specify truncation and failure spans

Build logs routinely run to tens of thousands of lines, so you must specify how they are cut before training. Decide the maximum log length your context window and training budget support, then require the supplier to state their truncation rule: head and tail, failure-span-centered window, or step-level extraction.

Request these preparation details in writing:

  • Raw log retained alongside any truncated version, so you can re-derive spans.
  • Failure-span markers with the rule used (first ERROR, test-framework summary, exit-code step).
  • Timestamps, ANSI color codes and progress bars stripped or preserved, stated explicitly.
  • Step boundaries from the CI system's own structure (job, step, group markers).
  • Paired passing log for diffing, since failing-versus-passing comparison is a strong diagnostic signal.
  • Deduplication of near-identical logs from matrix builds; research on text corpora found that deduplicated training data improves language models and reduces verbatim copying [2].

Secrets in CI output need their own scan

CI logs leak credentials in ways source code does not, so run your secrets standard over logs as well as code. CI masking typically applies only to values registered as secrets; tokens echoed by scripts, printed in stack traces, embedded in URLs or base64-encoded can pass through. Enterprise research shows secrets spread across code and shared documents and need dedicated detection and remediation [3].

Ask for the scanner used (for example Gitleaks or TruffleHog rule sets), whether it ran over logs and artifacts, and how hits were replaced. Require stable placeholder tokens so the agent learns the shape of an auth failure without seeing the key. The secrets removal guide for code datasets gives a full-history scanning checklist.

Matching the data to SFT, RL and evaluation

Each training use needs a different slice of the same CI history. Supervised fine-tuning uses log-plus-diagnosis-plus-fix triples. Reinforcement learning with build or test rewards needs a reproducible environment: lockfiles, container image digests and the test command, so a candidate patch can actually be executed. Triage classifiers need balanced root-cause labels across compiler errors, dependency resolution, test assertions, timeouts and infrastructure failures.

For evaluation, hold out repositories rather than random runs, because runs from the same repository leak fixes across splits. Public coding benchmarks face contamination problems; as of October 2026, OpenAI no longer reports SWE-bench Verified scores, citing contamination [4]. The held-out evaluation sets guide and the SWE task environments guide explain how to build executable splits. Related failure data appears in stack trace to fix data and dependency upgrade pairs.

Buyer checklist for a CI failure data request

Use this checklist to turn a vague "CI logs" request into a specification suppliers can answer.

Illustrative example: invented to show structure; it does not describe an available dataset.

RequirementWhat to state
CI systemsGitHub Actions, GitLab CI, Jenkins, Buildkite, CircleCI
Languages and build toolse.g. Java/Gradle, TypeScript/pnpm, Python/pytest
UnitFailing run linked to fix commit and verifying passing run
Link evidenceMethod per link and target linkage share
FlakinessFull rerun history per SHA with attempt numbers
LogsRaw plus truncated, max length, span-marking rule
EnvironmentLockfiles, image digests, test commands for replay
RetentionPeriod covered per CI provider
SecretsScanner, scope (code, logs, artifacts), placeholder scheme
FormatJSONL metadata, compressed text logs, Parquet tables
SplitRepository-level holdout for evaluation

For more on structuring requests, see how to specify a code dataset request and the coding agent training data overview.

How SourceX handles CI failure data requests

SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data they need, and SourceX looks for businesses that hold it. The process runs Find, Assess (data and licensing permissions), Agree (pricing and allowed uses in a license), Transact and Manage, and nothing is contracted until a supplier agrees. You can describe your CI dataset needs on the buyers page.

Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Personal details such as names and emails are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. The software bug fixing workflow page shows what these engineering records contain.

Request CI build failure data for your agents

Describe the CI systems, languages, linkage and rerun history you need, and SourceX will look for US companies that hold matching engineering records and can license them. Every release is approved by the supplying company, and terms are agreed per deal. Start a CI build failure data request.

Sources

  1. arXiv (He, Zhu et al.), "Loghub: A Large Collection of System Log Datasets towards Automated Log Analytics (also titled 'for AI-driven Log Analytics')" (2020). https://arxiv.org/pdf/2008.06448
  2. arXiv / ACL 2022 (Lee et al.), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  3. arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
  4. OpenAI, "Why we no longer evaluate SWE-bench Verified" (2026). https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data