Code and software engineering data
Production Error and Stack Trace to Fix Data for Debugging Agents
Quick answer
A stack-trace-to-fix dataset pairs a production exception (type, message, symbolicated frames, release version and grouping fingerprint) with the code change that resolved it, ideally plus a regression test. Buyers training debugging agents should require exact release tags so each trace replays against its base commit, one representative event per error group with counts, scrubbed request data and breadcrumbs, and a reported link rate from error group to fix. Public crash dumps rarely include the fix, so most usable data sits inside companies.
By SourceX Editorial · Updated
What a stack-trace-to-fix record must contain
A usable record joins three systems that rarely share keys: the error tracker, the version-control host and the issue tracker. The error tracker (Sentry, Bugsnag, Rollbar, Datadog Error Tracking, Crashlytics) supplies the exception and runtime context; Git supplies the base commit and the fixing diff; Jira, Linear or GitHub Issues often supply the human link between the two. If any join is missing, you have either a crash log or a commit, not a training example.
Research has long shown that developers look for stack traces when fixing bugs, which is why traces are a strong conditioning signal for agents [1][2]. Public crash corpora are typically fix-free: the Eclipse AERI dump, for example, publishes raw error-report traces with no linked fixes [3]. For agents that must go from runtime evidence to a patch, you need the fix side too.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"error_group_id": "grp_7f3a",
"fingerprint": ["{{ default }}", "PaymentRetryTimeout"],
"event_count_30d": 4812,
"first_seen_release": "billing-api@2025.11.3",
"representative_event": {
"exception_type": "java.lang.IllegalStateException",
"message": "Retry budget exhausted for [REDACTED_ACCOUNT]",
"frames": [
{"file": "RetryPolicy.java", "function": "nextDelay", "line": 88, "in_app": true},
{"file": "ChargeService.java", "function": "capture", "line": 141, "in_app": true}
],
"breadcrumbs": [{"category": "http", "data": {"method": "POST", "url": "/v1/charges/[ID]", "status_code": 503}}],
"environment": "production"
},
"base_commit": "a1c9e04",
"fix": {"commit": "d44b2f7", "link_method": "issue_key_in_commit_message", "files_changed": 2},
"regression_test": "ChargeServiceTest.capture_retriesWithinBudget",
"resolved_in_release": "billing-api@2025.11.4"
}
Why release tags decide whether a trace is replayable
A trace without a release identifier cannot be pinned to the code that produced it, so an agent cannot be scored on whether its patch fixes the real failure. Line numbers in frames drift with every deploy; ChargeService.java:141 in one build is a different statement in the next. Ask for the release string the SDK attached to each event, plus a mapping from release to commit SHA, which teams usually keep in their deploy pipeline or in error-tracker release metadata.
Then check the mapping on a sample. Pick 20 error groups, check out the base commit, and confirm the top in-app frame lands on plausible code. Failures here usually mean monorepo releases that bundle several services, hotfix branches never merged back, or releases tagged after deploy rather than at build. For the downstream environment work, see SWE task environments with tests.
How error grouping should shape the delivered unit
Deliver one representative event per error group plus volume statistics, not the raw event stream. Error trackers assign every event a fingerprint and merge events with the same fingerprint into one issue, and teams can override that fingerprint from the SDK or with project rules. A single regression can emit millions of identical events, so a raw export over-weights noisy bugs and leaks more personal data per useful example.
Ask the supplier which grouping logic produced the groups. Custom fingerprints can merge unrelated root causes or split one bug across many issues, and minified frontend code without source maps commonly splits a single bug across many groups. Request event_count, first_seen, last_seen, affected release range and the grouping basis (stack trace, exception type, message or custom fingerprint) for each group.
Linking an error group to the fixing change and measuring link rate
The fix link is the expensive, valuable part, and the supplier should report what share of groups have one. Common link methods, in descending reliability, are: the tracker's own resolve-in-commit or resolve-in-release action, an issue key (for example PAY-2291) present in both the error-tracker issue and the commit message, a pull request that references the error-group URL, and heuristic matching on changed files versus in-app frames. Each method should be recorded per record, because heuristic links carry false positives.
Expect a large fraction of groups to have no fix at all: ignored, auto-resolved after traffic dropped, or fixed by a config change outside the repository. Those records still have value for triage and localization training, but label them. The adjacent page on issue-to-fix pairs from private repositories covers validation of the commit side in depth.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field to request | Why it matters | Red flag |
|---|---|---|
| Link method per record | Lets you filter heuristic links | One undifferentiated "linked" flag |
| Link rate by service and year | Shows coverage bias | Only an overall percentage |
| Base commit and fix commit SHAs | Enables replay and diff scoring | Release names with no SHA map |
| Regression test name or diff | Gives a fail-to-pass oracle | Fixes with no test changes anywhere |
| Resolution type | Separates code fixes from config or rollbacks | Rollbacks counted as fixes |
Scrubbing payloads without destroying debugging signal
Remove identifiers from messages, breadcrumbs, request bodies, headers, user context and local variables, while preserving exception types, frame function names, file paths and line numbers. Error events routinely capture emails, account numbers, IP addresses, session cookies and authorization headers, and frame-local variables can hold full customer records. Token replacement that keeps type ([REDACTED_ACCOUNT], [EMAIL]) preserves the shape an agent needs to reason about which value was null or malformed.
For California consumers, "deidentified" under the CCPA requires more than field masking; the business holding the data must also take measures and make commitments about reidentification [4]. Separately, FTC staff have warned AI companies that using customer data for training in ways that break earlier privacy or confidentiality commitments can create liability under laws the FTC enforces, so ask whether the supplier's customer terms allow this use [5]. When you brief a sourcing partner such as SourceX for AI data buyers, state which identifier classes must be removed. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Also scan for secrets: stack-trace payloads often contain connection strings and API keys. The secrets-removal guide for code datasets describes verification steps.
Symbolication maps, minified traces and their own sensitivity
Minified JavaScript, ProGuard or R8-obfuscated Android code and stripped native binaries produce frames that are useless until symbolicated with source maps, mapping files or debug symbols. Ask whether traces were symbolicated at ingest (preferred) or whether raw maps would ship. Source maps frequently embed original source via sourcesContent, which means they are a code disclosure subject to the same ownership review as the repository itself; see code ownership due diligence.
Splitting for training versus evaluation
Hold out by error group and by time, never by event. Events from one group share the same fix, so random event splits leak answers into test. A time-based cutoff after your base model's training data, combined with a group-level holdout, gives a cleaner eval; the approach mirrors held-out SWE eval sets from private repos. Difficulty also varies by exception class (a NullPointerException is a different task from a ClassCastException), so stratify splits by exception type.
Illustrative example: invented to show structure; it does not describe an available dataset.
Request checklist for a stack-trace-to-fix dataset
- Languages, runtimes and error trackers in scope; date range; production environment only.
- One representative event per group, with counts, release range and grouping basis.
- Release-to-commit SHA mapping, verified on a sample.
- Fix commit, link method, resolution type and regression test per linked group.
- Scrubbing method for messages, breadcrumbs, request data and local variables, with a sample check.
- Symbolicated frames; no raw source maps unless the code itself is licensed.
- A dataset card documenting coverage, link rate and known gaps [6].
How this differs from ITSM, postmortem and CI data
This dataset is code-level: one exception, one release, one diff. Operational incident records, such as ITSM ticket datasets and licensed incident postmortems, describe service impact and response, not the frame-to-line fix. CI build failure logs capture failures before release, and software engineering histories supply the broader repository context. The code data buyer's map shows how these fit together.
Sourcing stack-trace-to-fix data through SourceX
SourceX sources operational datasets, including engineering records, from US companies on request; nothing is held in stock and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced with the method recorded and a sample checked, and delivery runs under a license defining records, uses, term and delivery. Describe the error-to-fix data you need at SourceX for AI data buyers.
Sources
- Università della Svizzera italiana (USI), "Misery Loves Company: CrowdStacking Traces to Aid Problem Detection" (2015). https://www.inf.usi.ch/faculty/lanza/PUBS/P/DalS2015a.pdf
- FLOSSHub, "Do stack traces help developers fix bugs?" (2010). https://flosshub.org/node/961
- Eclipse Foundation (University of Maryland mirror), "AERI stack traces mirror". https://mirror.umd.edu/eclipse/scava/aeri_stacktraces/
- California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Hugging Face, "Create a dataset card". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.