Agent, workflow and domain-reasoning data
Rework, reversals and reopened cases: human error-recovery data for agents
Quick answer
The most useful error-recovery examples for AI agents come from business records where a person made or caught a mistake and then fixed it: journal reversals, credit memos, reopened support tickets, re-shipments, voided invoices and ACH returns. To be trainable, each example must link the original record, the detection signal, the corrective action and the final state, with timestamps and reason codes. Synthetic failures teach syntax recovery; human rework records teach which fix a competent operator actually chose.
By SourceX Editorial · Updated
Why human rework records complement agent failure logs
Human rework records add the half of the recovery problem that agent logs rarely contain: a correct fix chosen by someone who understood the business consequence. Research on tool-using agents already shows the demand. PALADIN trains on recovery-annotated trajectories and retrieves matching recovery exemplars at inference time [1], while ParaRecover reports that early errors often go undetected and propagate through later steps [2]. Fission-GRPO observes that smaller models tend to repeat a failed call instead of reading the error, because a bare negative reward says nothing about what to do next [3].
Those papers mostly cover tool-level faults such as malformed arguments, hallucinated tools or HTTP errors. For that layer, see tool-call errors and recoveries. Business-process errors are different: the call succeeded, but the outcome was wrong. A posted invoice had the wrong vendor, a ticket was closed before the customer's issue was solved, or an order shipped to the old address. The record of how a person noticed and unwound that error is the signal an operations agent needs.
Which correction records carry a usable recovery signal
The best sources are systems that write the correction as a new, linked record rather than overwriting the original. Overwrite-in-place systems destroy the "before" state, so the error itself disappears.
| Process | Original record | Correcting record | Link field to request | Detection signal |
|---|---|---|---|---|
| General ledger | Journal entry | Reversing entry | Reversal reference / original document number | Reconciliation break, auditor note, close review |
| Accounts receivable | Invoice | Credit memo, rebill | Reference invoice ID, reason code | Customer dispute, short payment |
| Accounts payable | Vendor invoice, payment | Void, debit memo, re-issued payment | Original payment ID, void reason | Three-way-match exception, vendor statement |
| Payments | ACH entry | Return or reversal entry | Original trace number, return reason code | Bank return file [7] |
| Support / ITSM | Resolved ticket | Reopen event, follow-up ticket | Parent ticket ID, reopen reason | Customer reply, failed verification |
| Fulfillment | Shipment | Re-shipment, RMA, replacement order | Original order and line IDs | Carrier exception, customer claim |
| Field service / repair | Work order | Comeback or rework order | Prior repair order, technician note | Repeat failure within warranty window |
| Data entry | Master-data record | Change log entry | Record ID with field-level old and new values | Validation rule, downstream rejection |
Payment returns are a good model of what a clean signal looks like: Nacha return codes such as R01 (insufficient funds) or R03 (no account or unable to locate) give a standardized cause on every reversal [7]. Most internal processes are messier, with free-text reasons or none at all. For dispute-driven reversals in consumer banking, see Reg E and Reg Z dispute investigation records; for repair comebacks, dealer repair orders with complaint, cause and correction.
How to link the original, the detection and the fix into one sequence
A recovery example is only trainable when the original record, the detection event and the corrective action are joined into one ordered sequence. In most companies these live in different tables or different systems, so linkage is the main preparation cost.
Ask suppliers which join key exists. Common ones are a reversal reference on the correcting journal entry, a reference invoice on the credit memo, a parent ticket ID on a reopened case, and the original order line on an RMA. When no explicit key exists, matching on amount, counterparty and date window produces candidate links that need a confidence score and a manual sample check.
Event-log formats help here. OCEL 2.0 lets one event relate to several business objects, such as an invoice, a payment and a vendor, which is how a correction actually spans records [5]. The public BPI Challenge 2019 purchase-order log shows what procurement rework looks like in event data, including invoice-receipt and goods-receipt mismatches [6]. For broader linkage across ERP, CRM and ticketing, see cross-system workflow records and process mining event logs.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"recovery_id": "rc-000184",
"process": "accounts_receivable",
"original": {
"record_type": "invoice",
"record_id": "INV-55120",
"created_at": "2025-03-04T10:12:00Z",
"actor_role": "billing_specialist",
"fields": {"customer_id": "C-0931", "amount": 4820.00, "price_list": "2024-Q4"}
},
"detection": {
"detected_at": "2025-03-19T15:40:00Z",
"channel": "customer_short_payment",
"detector_role": "collections_analyst",
"evidence": "Payment of 4,338.00 referencing INV-55120; remittance note cites contract price"
},
"diagnosis": {
"root_cause_code": "STALE_PRICE_LIST",
"root_cause_note": "Contract renewal moved customer to 2025 pricing on Mar 1",
"root_cause_source": "free_text_mapped"
},
"correction": [
{"step": 1, "action": "issue_credit_memo", "record_id": "CM-08812", "ref": "INV-55120", "amount": -4820.00, "at": "2025-03-20T09:05:00Z"},
{"step": 2, "action": "rebill", "record_id": "INV-55871", "amount": 4338.00, "price_list": "2025-Q1", "at": "2025-03-20T09:11:00Z"},
{"step": 3, "action": "apply_payment", "payment_id": "PMT-22019", "invoice_id": "INV-55871", "at": "2025-03-20T09:14:00Z"}
],
"outcome": {"status": "closed", "recurrence_within_90d": false},
"link_method": "explicit_reference",
"pii_treatment": "customer names and contacts replaced with stable pseudonyms"
}
What to capture about how the error was detected
Detection is the part buyers most often forget to specify, and it is the part an agent most needs to learn. A model that only sees "credit memo issued" learns to apply fixes; a model that sees why someone looked again learns when to doubt its own output.
Request a detection channel field with a controlled vocabulary. Typical values are customer complaint, reconciliation break, validation rule, peer or supervisor review, downstream system rejection, and audit sample. Record the detection lag (time from original action to detection) and who detected it, since self-detected errors and customer-detected errors teach different behaviors. ParaRecover's finding that undetected early errors compound [2] argues for keeping the intermediate records between the error and its detection, not just the endpoints.
Root-cause notes: valuable, sparse and inconsistent
Root-cause fields are the richest part of a recovery record and the least reliable, so treat them as optional enrichment with measured coverage. Many processes have a reason code dropdown that operators default to "other", plus a free-text comment that may be empty, terse or written for a colleague.
Ask suppliers for the reason-code list as configured in the source system, coverage per field (share of correction records with a non-default code and a non-empty note), and whether codes changed during the period. If you map free text to a taxonomy, record the method in a field such as root_cause_source so evaluation can separate system-coded causes from inferred ones. Structured reflection methods for tool agents depend on explicit error, diagnosis and fix structure [4]; business data rarely arrives that clean, so the mapping step is real work. Rationale capture is covered further in decision records with rationale.
Report base rates of rework by process
Every recovery dataset should ship with the base rate of correction for each process it covers, because rework records are a filtered view of operations. For example, if 3% of invoices got credit memos, a dataset made only of corrected invoices tells an agent nothing about the 97% that were right.
Ask for per-process counts of original records and correcting records over the same period, ideally by month. With both, you can build balanced training mixes, estimate how often an agent should flag work for review, and construct evaluation sets with realistic prevalence. Pair corrected cases with uncorrected look-alikes (same process, same period, similar fields) so a model can learn what distinguishes a record that needed fixing. For allocation methods, see stratified evaluation sets for rare and high-risk cases.
Failure modes that make rework data misleading
The common failures in rework datasets come from accounting conventions, workflow hygiene and duplication rather than from the errors themselves.
- Non-error reversals. Accrual reversals, period-end reclassifications and scheduled auto-reversing entries look like corrections but are routine. Filter by reversal type or source module.
- Reopen noise. Tickets reopen when customers reply "thanks" to a closed case. Require a reopen reason or a substantive follow-up action before counting it as a recovery.
- Commercial concessions. Goodwill credits and price-match credits are not error fixes. Separate them by reason code.
- Silent fixes. Some corrections happen by editing a field in place with no log. If the system lacks field-level change history, those recoveries are invisible, which skews the dataset toward errors serious enough to need a formal document.
- Templated near-duplicates. Batch corrections, such as one pricing mistake fixed across 400 invoices, produce hundreds of near-identical examples. Deduplicate or cap per root cause; research on language-model training data found that near-duplicates increase verbatim memorization [9].
- Survivor bias in outcomes. A first correction that itself failed may be followed by a second fix. Keep chains intact and label the final outcome.
Data quality checks for these issues fit the framework in training data quality metrics.
Using recovery records for SFT and recovery evaluation
Rework sequences work as supervised fine-tuning targets and as held-out evaluation cases, but the format differs. For SFT, render each sequence as context (the original record and detection evidence), a diagnosis turn and the ordered corrective actions as tool calls; function-calling fine-tuning data formats covers the call and result schema.
For evaluation, hold out whole processes or time periods, present the agent with the original record plus the detection signal, and score whether it chooses the same corrective path and end state. Policy-bound customer-service benchmarks such as tau-bench check database end states after an agent interacts with users and APIs [8]; recovery evaluation can use the same end-state approach with the correcting records as the reference. Define success labels explicitly, as described in task success labels for agent trajectories.
Request checklist for error-recovery records
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to specify | Why it matters |
|---|---|---|
| Processes | e.g., AR credit memos, ITSM reopens, RMAs | Scopes systems and linkage effort |
| Linkage | Explicit keys required, or heuristic matching allowed with confidence scores | Determines label noise |
| Detection fields | Channel, detector role, detection timestamp | Teaches when to re-check |
| Root cause | Source reason codes, free-text notes, coverage per field | Sets expectations for diagnosis training |
| Base rates | Original and correction counts per process per month | Enables balanced mixes and realistic evals |
| Exclusions | Accrual reversals, goodwill credits, courtesy reopens | Removes non-error noise |
| Negative pairs | Uncorrected look-alike records | Lets models learn what needed fixing |
| Personal data | Pseudonymization of names, contacts and account numbers | Privacy and license compliance |
| Documentation | Data card covering sources, preparation and intended use [10] | Supports internal review |
| Rights | Written license covering training use | Many public datasets have missing or wrong licenses [11] |
Buyers who want this kind of data from real operations can describe the processes and fields on the SourceX buyer page. SourceX sources operational datasets from US companies on request, including support histories and finance and legal workflows; categories are not inventory, and a request does not guarantee a match. Related owner pages cover workflow task histories, accounting reconciliations, customer support agent data and finance and accounting agent data. For the wider cluster, start at the AI agent training data hub, and compare this with human override and correction logs and exception handling records.
Sourcing error-recovery examples for your agents
SourceX looks for US businesses that hold the correction and rework records you describe, rights-reviews each dataset for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, and every release is approved by the supplying company. Describe the processes, linkage and fields you need at https://sourcex.si/buyers.
Sources
- arXiv, "PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases" (2025). https://arxiv.org/pdf/2509.25238
- arXiv, "ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents" (2026). https://arxiv.org/pdf/2609.12345
- arXiv, "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1
- arXiv, "Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions" (2025). https://arxiv.org/pdf/2509.18847
- arXiv, "OCEL (Object-Centric Event Log) 2.0 Specification" (2024). https://arxiv.org/pdf/2403.01975
- 4TU.ResearchData, "BPI Challenge 2019 (OCEL)". https://data.4tu.nl/datasets/46a7e15b-10c7-4ab2-988d-ee67d8ea515a
- Nacha, "Nacha ACH Return Codes". https://www.nacha.org/rules/ach-return-reason-codes
- arXiv, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
- arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.