Skip to content

Agent, workflow and domain-reasoning data

Rework, reversals and reopened cases: human error-recovery data for agents

Quick answer

The most useful error-recovery examples for AI agents come from business records where a person made or caught a mistake and then fixed it: journal reversals, credit memos, reopened support tickets, re-shipments, voided invoices and ACH returns. To be trainable, each example must link the original record, the detection signal, the corrective action and the final state, with timestamps and reason codes. Synthetic failures teach syntax recovery; human rework records teach which fix a competent operator actually chose.

By SourceX Editorial · Updated

Why human rework records complement agent failure logs

Human rework records add the half of the recovery problem that agent logs rarely contain: a correct fix chosen by someone who understood the business consequence. Research on tool-using agents already shows the demand. PALADIN trains on recovery-annotated trajectories and retrieves matching recovery exemplars at inference time [1], while ParaRecover reports that early errors often go undetected and propagate through later steps [2]. Fission-GRPO observes that smaller models tend to repeat a failed call instead of reading the error, because a bare negative reward says nothing about what to do next [3].

Those papers mostly cover tool-level faults such as malformed arguments, hallucinated tools or HTTP errors. For that layer, see tool-call errors and recoveries. Business-process errors are different: the call succeeded, but the outcome was wrong. A posted invoice had the wrong vendor, a ticket was closed before the customer's issue was solved, or an order shipped to the old address. The record of how a person noticed and unwound that error is the signal an operations agent needs.

Which correction records carry a usable recovery signal

The best sources are systems that write the correction as a new, linked record rather than overwriting the original. Overwrite-in-place systems destroy the "before" state, so the error itself disappears.

ProcessOriginal recordCorrecting recordLink field to requestDetection signal
General ledgerJournal entryReversing entryReversal reference / original document numberReconciliation break, auditor note, close review
Accounts receivableInvoiceCredit memo, rebillReference invoice ID, reason codeCustomer dispute, short payment
Accounts payableVendor invoice, paymentVoid, debit memo, re-issued paymentOriginal payment ID, void reasonThree-way-match exception, vendor statement
PaymentsACH entryReturn or reversal entryOriginal trace number, return reason codeBank return file [7]
Support / ITSMResolved ticketReopen event, follow-up ticketParent ticket ID, reopen reasonCustomer reply, failed verification
FulfillmentShipmentRe-shipment, RMA, replacement orderOriginal order and line IDsCarrier exception, customer claim
Field service / repairWork orderComeback or rework orderPrior repair order, technician noteRepeat failure within warranty window
Data entryMaster-data recordChange log entryRecord ID with field-level old and new valuesValidation rule, downstream rejection

Payment returns are a good model of what a clean signal looks like: Nacha return codes such as R01 (insufficient funds) or R03 (no account or unable to locate) give a standardized cause on every reversal [7]. Most internal processes are messier, with free-text reasons or none at all. For dispute-driven reversals in consumer banking, see Reg E and Reg Z dispute investigation records; for repair comebacks, dealer repair orders with complaint, cause and correction.

A recovery example is only trainable when the original record, the detection event and the corrective action are joined into one ordered sequence. In most companies these live in different tables or different systems, so linkage is the main preparation cost.

Ask suppliers which join key exists. Common ones are a reversal reference on the correcting journal entry, a reference invoice on the credit memo, a parent ticket ID on a reopened case, and the original order line on an RMA. When no explicit key exists, matching on amount, counterparty and date window produces candidate links that need a confidence score and a manual sample check.

Event-log formats help here. OCEL 2.0 lets one event relate to several business objects, such as an invoice, a payment and a vendor, which is how a correction actually spans records [5]. The public BPI Challenge 2019 purchase-order log shows what procurement rework looks like in event data, including invoice-receipt and goods-receipt mismatches [6]. For broader linkage across ERP, CRM and ticketing, see cross-system workflow records and process mining event logs.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "recovery_id": "rc-000184",
  "process": "accounts_receivable",
  "original": {
    "record_type": "invoice",
    "record_id": "INV-55120",
    "created_at": "2025-03-04T10:12:00Z",
    "actor_role": "billing_specialist",
    "fields": {"customer_id": "C-0931", "amount": 4820.00, "price_list": "2024-Q4"}
  },
  "detection": {
    "detected_at": "2025-03-19T15:40:00Z",
    "channel": "customer_short_payment",
    "detector_role": "collections_analyst",
    "evidence": "Payment of 4,338.00 referencing INV-55120; remittance note cites contract price"
  },
  "diagnosis": {
    "root_cause_code": "STALE_PRICE_LIST",
    "root_cause_note": "Contract renewal moved customer to 2025 pricing on Mar 1",
    "root_cause_source": "free_text_mapped"
  },
  "correction": [
    {"step": 1, "action": "issue_credit_memo", "record_id": "CM-08812", "ref": "INV-55120", "amount": -4820.00, "at": "2025-03-20T09:05:00Z"},
    {"step": 2, "action": "rebill", "record_id": "INV-55871", "amount": 4338.00, "price_list": "2025-Q1", "at": "2025-03-20T09:11:00Z"},
    {"step": 3, "action": "apply_payment", "payment_id": "PMT-22019", "invoice_id": "INV-55871", "at": "2025-03-20T09:14:00Z"}
  ],
  "outcome": {"status": "closed", "recurrence_within_90d": false},
  "link_method": "explicit_reference",
  "pii_treatment": "customer names and contacts replaced with stable pseudonyms"
}

What to capture about how the error was detected

Detection is the part buyers most often forget to specify, and it is the part an agent most needs to learn. A model that only sees "credit memo issued" learns to apply fixes; a model that sees why someone looked again learns when to doubt its own output.

Request a detection channel field with a controlled vocabulary. Typical values are customer complaint, reconciliation break, validation rule, peer or supervisor review, downstream system rejection, and audit sample. Record the detection lag (time from original action to detection) and who detected it, since self-detected errors and customer-detected errors teach different behaviors. ParaRecover's finding that undetected early errors compound [2] argues for keeping the intermediate records between the error and its detection, not just the endpoints.

Root-cause notes: valuable, sparse and inconsistent

Root-cause fields are the richest part of a recovery record and the least reliable, so treat them as optional enrichment with measured coverage. Many processes have a reason code dropdown that operators default to "other", plus a free-text comment that may be empty, terse or written for a colleague.

Ask suppliers for the reason-code list as configured in the source system, coverage per field (share of correction records with a non-default code and a non-empty note), and whether codes changed during the period. If you map free text to a taxonomy, record the method in a field such as root_cause_source so evaluation can separate system-coded causes from inferred ones. Structured reflection methods for tool agents depend on explicit error, diagnosis and fix structure [4]; business data rarely arrives that clean, so the mapping step is real work. Rationale capture is covered further in decision records with rationale.

Report base rates of rework by process

Every recovery dataset should ship with the base rate of correction for each process it covers, because rework records are a filtered view of operations. For example, if 3% of invoices got credit memos, a dataset made only of corrected invoices tells an agent nothing about the 97% that were right.

Ask for per-process counts of original records and correcting records over the same period, ideally by month. With both, you can build balanced training mixes, estimate how often an agent should flag work for review, and construct evaluation sets with realistic prevalence. Pair corrected cases with uncorrected look-alikes (same process, same period, similar fields) so a model can learn what distinguishes a record that needed fixing. For allocation methods, see stratified evaluation sets for rare and high-risk cases.

Failure modes that make rework data misleading

The common failures in rework datasets come from accounting conventions, workflow hygiene and duplication rather than from the errors themselves.

  • Non-error reversals. Accrual reversals, period-end reclassifications and scheduled auto-reversing entries look like corrections but are routine. Filter by reversal type or source module.
  • Reopen noise. Tickets reopen when customers reply "thanks" to a closed case. Require a reopen reason or a substantive follow-up action before counting it as a recovery.
  • Commercial concessions. Goodwill credits and price-match credits are not error fixes. Separate them by reason code.
  • Silent fixes. Some corrections happen by editing a field in place with no log. If the system lacks field-level change history, those recoveries are invisible, which skews the dataset toward errors serious enough to need a formal document.
  • Templated near-duplicates. Batch corrections, such as one pricing mistake fixed across 400 invoices, produce hundreds of near-identical examples. Deduplicate or cap per root cause; research on language-model training data found that near-duplicates increase verbatim memorization [9].
  • Survivor bias in outcomes. A first correction that itself failed may be followed by a second fix. Keep chains intact and label the final outcome.

Data quality checks for these issues fit the framework in training data quality metrics.

Using recovery records for SFT and recovery evaluation

Rework sequences work as supervised fine-tuning targets and as held-out evaluation cases, but the format differs. For SFT, render each sequence as context (the original record and detection evidence), a diagnosis turn and the ordered corrective actions as tool calls; function-calling fine-tuning data formats covers the call and result schema.

For evaluation, hold out whole processes or time periods, present the agent with the original record plus the detection signal, and score whether it chooses the same corrective path and end state. Policy-bound customer-service benchmarks such as tau-bench check database end states after an agent interacts with users and APIs [8]; recovery evaluation can use the same end-state approach with the correcting records as the reference. Define success labels explicitly, as described in task success labels for agent trajectories.

Request checklist for error-recovery records

Illustrative example: invented to show structure; it does not describe an available dataset.

ItemWhat to specifyWhy it matters
Processese.g., AR credit memos, ITSM reopens, RMAsScopes systems and linkage effort
LinkageExplicit keys required, or heuristic matching allowed with confidence scoresDetermines label noise
Detection fieldsChannel, detector role, detection timestampTeaches when to re-check
Root causeSource reason codes, free-text notes, coverage per fieldSets expectations for diagnosis training
Base ratesOriginal and correction counts per process per monthEnables balanced mixes and realistic evals
ExclusionsAccrual reversals, goodwill credits, courtesy reopensRemoves non-error noise
Negative pairsUncorrected look-alike recordsLets models learn what needed fixing
Personal dataPseudonymization of names, contacts and account numbersPrivacy and license compliance
DocumentationData card covering sources, preparation and intended use [10]Supports internal review
RightsWritten license covering training useMany public datasets have missing or wrong licenses [11]

Buyers who want this kind of data from real operations can describe the processes and fields on the SourceX buyer page. SourceX sources operational datasets from US companies on request, including support histories and finance and legal workflows; categories are not inventory, and a request does not guarantee a match. Related owner pages cover workflow task histories, accounting reconciliations, customer support agent data and finance and accounting agent data. For the wider cluster, start at the AI agent training data hub, and compare this with human override and correction logs and exception handling records.

Sourcing error-recovery examples for your agents

SourceX looks for US businesses that hold the correction and rework records you describe, rights-reviews each dataset for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, and every release is approved by the supplying company. Describe the processes, linkage and fields you need at https://sourcex.si/buyers.

Sources

  1. arXiv, "PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases" (2025). https://arxiv.org/pdf/2509.25238
  2. arXiv, "ParaRecover: A Process-Level Benchmark for Error Localization and Recovery in Parallel Tool-Use Agents" (2026). https://arxiv.org/pdf/2609.12345
  3. arXiv, "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1
  4. arXiv, "Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions" (2025). https://arxiv.org/pdf/2509.18847
  5. arXiv, "OCEL (Object-Centric Event Log) 2.0 Specification" (2024). https://arxiv.org/pdf/2403.01975
  6. 4TU.ResearchData, "BPI Challenge 2019 (OCEL)". https://data.4tu.nl/datasets/46a7e15b-10c7-4ab2-988d-ee67d8ea515a
  7. Nacha, "Nacha ACH Return Codes". https://www.nacha.org/rules/ach-return-reason-codes
  8. arXiv, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  9. arXiv (ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  10. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  11. arXiv, "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data