Skip to content

Software companies

CI build logs and test histories: are they licensable data?

By SourceX Editorial · Updated

Short answer

CI build logs and test histories can be licensable data when the company owns them, they link failures to the commits and reviews that fixed them, and secrets, hostnames and customer details have been stripped. Unlinked logs add little. A failing build tied to its fix, its discussion and a passing rerun is the unit AI developers look for.

Key takeaways

  • Linkage is the value test: a log tied to a commit, pull request and passing rerun teaches far more than a log alone.
  • Logs leak secrets, internal hostnames and customer identifiers in ways source code does not, so scan exported logs separately.
  • Hosted CI services expire old logs and artifacts, so export history before a migration or plan change.
  • Customer-supplied test fixtures and third-party tool output need a rights check before inclusion.

Are CI build logs and test histories licensable?#

CI build logs and test histories are often licensable when three conditions hold: the company owns the records, they connect to the code changes that caused and fixed failures, and they can be cleaned of secrets and of personal or customer details. The logs themselves are machine output; their value comes from the human engineering work around them.

AI developers building coding agents need tasks with a verifiable outcome. A build that failed, the diff that fixed it, the review discussion and the rerun that passed form exactly that kind of task. A folder of logs with no commit hashes, no test names and no link to a fix is far harder to use.

What a CI history actually contains#

A CI history spans more artifacts than most teams expect, and each carries a different mix of value and risk. Inventory them before deciding scope.

Systems differ in what they keep. GitHub Actions, GitLab CI, Jenkins, CircleCI and Buildkite each store logs and artifacts in their own way, and hosted services usually expire old runs on a schedule that can be changed but that few teams have ever reviewed.

Before a CI migration or plan change, export runs together with their commit hashes, branch names, pull request numbers and test report files, and store them next to a mirror of the repository. That way the links survive even if the CI service does not.

  • Build logs: compiler output, dependency resolution, environment setup and step-by-step command output.
  • Test reports: structured results per test case, often JUnit-style XML, with timings and failure messages.
  • Flaky-test records: retries, quarantine lists and the tickets opened to fix unstable tests.
  • Pipeline definitions: the YAML files or scripts describing how builds run, versioned alongside code.
  • Coverage and static analysis output: which lines were exercised and which warnings were raised.
  • Deployment logs: what was released, when, and whether it was rolled back.
  • Artifacts: built packages, container images and screenshots from end-to-end tests.

Why logs linked to fixes matter more than volume#

Logs linked to fixes matter because the link turns output into a lesson. The failure message shows the symptom, the commit shows the change, the pull request shows the reasoning, and the passing rerun confirms the result.

Flaky-test history adds a second layer. A record of a test that failed intermittently, the investigation into timing or shared state, and the eventual fix captures diagnostic judgment that is hard to generate synthetically.

The SourceX Enterprise Data Value Framework describes the same pattern in its own terms. Linkage to human fixes raises human-generated signal and AI utility, while logs anyone could regenerate by rerunning a public project add little, because reproducibility reduces value.

What to strip before anything leaves#

Stripping CI history takes more than a code scan, because logs echo environment variables, print connection strings and record every hostname a build touched. TruffleHog, an AGPL-3.0 open-source scanner, says it scans logs as well as Git, chats and wikis; whatever tool you choose, run it over the exported logs, not only over repositories.

Automated scanning is a first pass. Presidio's documentation, for example, warns that because it relies on automated detection there is no guarantee it will find all sensitive information, and that additional protections should be used. Plan a human review of a sample from every pipeline.

What to strip before anything leaves
ItemWhere it hidesTreatment
API keys, tokens and passwordsDebug output, failed authentication, echoed environment variablesRotate if live, then redact
Internal hostnames and URLsNetwork errors, deploy steps, health checksReplace with consistent placeholders
IP addresses and cloud account IDsProvisioning and test setupMask or generalize
Signed or pre-authorized URLsArtifact uploads and downloadsRemove entirely
Customer names and identifiersFixtures, seeded databases, end-to-end screenshotsExclude or replace with synthetic values
Employee names and emailsCommit metadata, approvers, notificationsPseudonymize consistently
License keys for build toolsTool setup outputRedact

Rights questions specific to CI data#

Rights questions for CI data center on material produced by someone other than your engineers. Most build logs record your own code and processes, but several inputs need a check before they go into a package.

Test fixtures copied from customer environments, or seeded with customer records, belong on the customer side of the ledger and should be excluded. Output from commercial scanners and analysis tools may be governed by those vendors' terms. Open-source dependency output can usually stay as context, but the dependencies themselves remain under their own licenses and are not yours to license.

Hosted CI providers' terms typically treat logs as your content, but read them before relying on that. Work by contractors is usually covered by assignment or work-for-hire clauses; confirm those exist for the period you plan to include.

Illustrative: an inspection software company reviews its pipelines#

Illustrative: a fictional software company that builds inspection apps for commercial elevator contractors runs Jenkins for its older services and GitHub Actions for newer ones. A planned move off Jenkins prompts the CTO to ask which build history is worth keeping.

The team finds that Jenkins kept years of console logs, but only some include commit hashes. GitHub Actions runs link cleanly to pull requests, yet older runs have already expired. A scan of a log sample turns up a database password echoed by a misconfigured step and a long list of internal hostnames.

The company rotates the password, exports the Jenkins logs that carry commit references, keeps the test report XML and drops logs that cannot be tied to any change. It lengthens artifact retention on the newer system and documents the scan. The resulting archive pairs failures with fixes for the services it still maintains.

How SourceX evaluates CI and test history#

SourceX evaluates CI and test history as part of a software engineering package rather than as a standalone dataset. In the Supply step of the SourceX five-step transaction, the fit check asks which CI systems are in use, how far back logs and test reports reach and whether runs link to commits and pull requests.

Preparation covers secret scanning, hostname and identifier replacement and removal of customer fixtures. The scan results go into the privacy record of the SourceX Evidence Packet alongside provenance, licensing rights, permitted use and release authorization, and the supplier approves each release before Delivery.

Frequently asked questions

Are machine-generated logs less valuable than human-written records?

On their own, usually yes. Logs are output a build system produced, and similar logs can often be regenerated. They gain value when tied to human work: the commit that fixed the failure, the review discussion and the issue describing the problem. Package them with that context rather than alone.

Do we have to license source code as well?

Not necessarily. A failure-and-fix pair is most useful with the diff, but some packages include only the changed files, the relevant tests and the surrounding discussion. Scope depends on what the company is comfortable releasing and what the buyer needs, and it is settled during rights review rather than assumed.

What about flaky tests we never fixed?

Unresolved flaky tests have less value as training tasks because there is no confirmed fix. They can still serve as evaluation cases or as context showing how instability affected the team. Keep the retry records and related tickets, and label them unresolved so nobody mistakes them for solved problems.

Can we include CI history from a discontinued product?

Often yes, and a discontinued product can be a cleaner candidate because there is no ongoing release risk. Check that the history was exported before the CI service expired it, that the code and tests are still company-owned and that customer fixtures have been removed.

Should we scan logs with the same tool we use for code?

You can, but configure it for logs. Code scanners are tuned to committed secrets, while logs leak through echoed variables, stack traces and URLs with embedded tokens. Run the scan on the exported copy you would deliver and add a manual review of samples from each pipeline.

Sources

  • TruffleHog, an AGPL-3.0 open-source secret scanner, says it classifies over 800 secret types and scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify