Skip to content

Agent, workflow and domain-reasoning data

Tool-call errors and recoveries: real failure data for robust tool use

Quick answer

Tool-call error recovery data is a set of real sequences in which an API or tool call failed (schema validation, authorization, not found, conflict, rate limit, timeout or partial success) paired with what the operator or integration did next and whether it worked. The best raw material is integration run histories with retries, tickets about failed syncs and operator notes. Each sequence should carry the original arguments, the raw error payload, the recovery action, the attempt count and a verified final outcome, so a model learns to read errors instead of repeating them.

By SourceX Editorial · Updated

Why happy-path tool data does not teach recovery

Happy-path function-calling data teaches a model to choose a tool and fill arguments, but it says nothing about what to do when the call returns an error. Recent research makes the gap concrete: smaller models often loop on a failed call rather than reading the error message, and standard reinforcement learning reduces a failure to a sparse negative reward with no guidance on the fix [1]. The ACL 2026 version of that work reports that turning execution errors into corrective training examples lifted Qwen3-8B error recovery by 5.7 points [2].

A second line of work builds explicit error, reflection and correction examples and evaluates them on BFCL v3 multi-turn tasks through a new Tool-Reflection-Bench; the v2 preprint says its datasets will be released after acceptance [3]. PALADIN frames tool-failure cases as a direct training target for self-correcting agents [4]. All three rely mostly on synthesized or injected failures, which is exactly where real operational records add value: real errors have messy payloads, vendor-specific codes and recoveries that took human judgment.

Benchmarks such as the Berkeley Function Calling Leaderboard measure whether calls are well formed and whether multi-turn agentic tasks complete [6]. They do not by themselves supply a distribution of production failures, so teams training for robustness usually need both a benchmark and a corpus of real failure-recovery sequences.

Error classes worth labeling separately

Each error class implies a different correct recovery, so a single "failed" flag is not enough. SHIELDA usefully separates exceptions that start in planning, such as a hallucinated step that only breaks later during execution, from failures inside the tool call itself [5]. The table below is the tool-call half of that picture, with the recovery a well-behaved integration or operator typically applied.

Illustrative example: invented to show structure; it does not describe an available dataset.

Error classTypical signal in real recordsCorrect recovery usually seenCommon wrong recovery to label as negative
Schema or validationHTTP 400 or 422, field-level message such as "invalid date format" or "required: customer_id"Correct the named argument and resend onceResend identical arguments
AuthenticationHTTP 401, expired OAuth token, revoked API keyRefresh token, then retry; escalate if refresh failsRetry in a loop with the dead token
PermissionHTTP 403, missing scope, row-level access deniedStop and report the missing permission to a personSwitch to a broader-privilege tool silently
Not foundHTTP 404, stale ID after a merge or deletionLook up the current ID with a search or list toolCreate a duplicate record
ConflictHTTP 409 or 412, version or ETag mismatch, duplicate keyRe-read the current state, then reapply the changeOverwrite with the stale payload
Rate limitHTTP 429, Retry-After header, quota-exceeded bodyWait for the stated interval, back off, retryImmediate retry or parallel fan-out
TimeoutClient timeout, HTTP 504, gateway resetCheck whether the write landed before retrying; use an idempotency keyBlind retry that double-posts a payment or ticket
Partial successHTTP 207, batch response with per-item errorsRetry only the failed itemsRetry the whole batch
Upstream outageHTTP 500 or 503, maintenance pageFall back to an alternate tool or queue for laterReport success to the user

The timeout and partial-success rows matter most for agents acting on real systems, because the naive recovery creates side effects. Records that show an operator checking state before retrying are rare and valuable.

Where real failure and recovery sequences live

The richest sources are operational systems that already log calls, errors and the next action. Integration platforms and iPaaS run histories (for example connector logs from data sync tools, Zapier-style task histories or MuleSoft and Boomi execution logs) show each attempt with its status code, payload and retry schedule. Our guide to API call logs and integration run histories as tool-use training data covers the success side of those logs; this page is about the failure branches.

Support and engineering tickets about failed syncs add the part logs miss: why the failure happened and what a person decided. A hypothetical ticket that says "CRM sync failing with INVALID_FIELD after admin renamed a picklist; remapped field and replayed 312 records" is a complete error, diagnosis and recovery triple. Reconstructing agent trajectories from ticket and case histories explains how to stitch those into step sequences, and human error-recovery records covers rework and reversals at the task level rather than the call level.

Observability traces are a third source. OpenTelemetry's generative AI semantic conventions define spans for model calls and tool execution that can carry tool-call arguments and results when teams opt in to capturing them [7]. Teams that already run agents in production often have these traces, and so do suppliers that instrumented their own integrations; for Model Context Protocol servers, MCP tool-use data describes the trace shape.

Runbooks and operator notes close the loop. They encode the recovery policy (for instance, "never retry a 409 on invoices; re-fetch and reconcile"), which lets you label whether an observed recovery followed policy.

The record each failure sequence should carry

A usable sequence links the failed call, the evidence the actor saw, every subsequent attempt and a verified outcome. Most function-calling fine-tuning formats are JSONL with a tools field, assistant tool calls and tool outputs returned as messages [8], so the error payload can be placed directly in the tool message the model would see at inference time.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "sequence_id": "seq_00417",
  "source_system": "crm_sync_connector",
  "tool_schema_ref": "crm.update_contact@v3",
  "attempts": [
    {
      "attempt": 1,
      "ts": "2026-03-02T14:01:07Z",
      "arguments": {"contact_id": "C-88213", "phone": "[PHONE_1]", "lifecycle": "Customer"},
      "response_status": 422,
      "error_payload": {"code": "INVALID_ENUM", "field": "lifecycle", "allowed": ["lead", "customer", "churned"]},
      "error_class": "schema_validation"
    },
    {
      "attempt": 2,
      "ts": "2026-03-02T14:01:09Z",
      "recovery_action": "corrected_arguments",
      "arguments": {"contact_id": "C-88213", "phone": "[PHONE_1]", "lifecycle": "customer"},
      "response_status": 200
    }
  ],
  "recovered": true,
  "attempts_to_recover": 2,
  "outcome_verified_by": "read_back_call",
  "side_effects": "none",
  "actor": "integration_auto_retry",
  "policy_compliant": true,
  "pii_method": "pseudonymized_tokens"
}

Useful label fields beyond the example: retry_after_seconds observed versus waited, idempotency_key_present, escalated_to_human with the escalation channel, alternate_tool_used, and recovery_quality (correct, acceptable, harmful). The recovered and attempts_to_recover fields are what separate this data from generic tool-use corpora, because they let you train on and evaluate recovery efficiency, not just eventual success.

Labeling recovery outcomes you can trust

The hardest label is whether the recovery actually succeeded, because a 200 response after a retry does not prove the intended state change happened once. Prefer outcomes verified by a read-back call, a reconciliation report or a closed ticket that confirms the fix. Mark sequences where the outcome is inferred only from the final status code so you can down-weight or exclude them.

Error-class labels derived from status codes are cheap but noisy: vendors return 400 for permission errors, 500 for validation failures and 200 with an error body. Run a second pass that classifies from the error payload text, then audit disagreements. Confident learning, which estimates class-conditional label noise and ranks likely mislabeled examples, is a practical way to target that audit [9].

Also label negative recoveries. Retry storms, duplicate records created after a 404 and blind retries after timeouts are the behaviors you want a model to avoid, and they only exist as training signal if you keep and tag them rather than filtering them out as noise.

Sampling for robustness training and evaluation

Real failure distributions are heavily skewed toward a few classes, usually rate limits and validation errors, so raw logs under-represent the rare, high-cost cases. Stratify by error class, tool, vendor API and recovery action, and set floor counts for conflicts, partial successes and timeouts with side effects. Keep the original class frequencies in metadata so evaluation can report both a stratified and a natural-distribution score.

Split by integration or tenant, not by sequence, so that the same vendor quirk does not leak from training into evaluation. For sandbox-based evaluation, pair sequences with the starting state that reproduces the failure; seed data and state snapshots for enterprise agent sandboxes covers how to capture that state. Document collection windows, sources, the error taxonomy version and known gaps in a Data Card-style summary so downstream users know what the set does and does not represent [10].

Privacy and rights checks specific to error payloads

Error payloads leak more personal and secret data than successful calls. Validation errors often echo the rejected value back ("invalid email: jane.doe@..."), auth failures can contain partial tokens, and stack traces include hostnames and file paths. Redact the request arguments, the echoed values inside error bodies, headers (Authorization, cookies, API keys) and free-text operator notes, and replace identifiers with consistent tokens so multi-attempt sequences still line up.

Confirm that the supplying company may license logs that include third-party API responses, since vendor terms and customer contracts can restrict reuse of response data. For screen-level traces rather than API logs, see PII in screen recordings and computer-use trajectories.

When SourceX sources this kind of data, personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Every dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Teams scoping a request can describe the data they need on the SourceX buyer page.

How this fits with broader tool-use data

Tool-call error recovery data is a narrow slice of agent data with its own label set, so most teams combine it with other sources. Use the agent training data hub to see the workflow, action and decision data around it, and the SourceX overview of training data for tool use and function calling for the broader capability. Task-level failures, where the agent picked the wrong plan rather than mishandling a call, belong in a separate set, and physical-world recovery is covered under robot failure and intervention data.

Request tool-call failure and recovery data

SourceX sources operational datasets from US companies on request, including support histories and engineering records, and manages licensing agreements and ongoing purchases. Data is not held in stock, a request does not guarantee a match, and nothing is contracted until a supplier agrees. Describe the error classes, systems and recovery labels you need at https://sourcex.si/buyers.

Sources

  1. arXiv, "Robust Tool Use via Fission-GRPO: Learning to Recover from Execution Errors" (2026). https://arxiv.org/html/2601.15625v1
  2. ACL Anthology, "Robust Tool Use via Fission-GRPO (ACL 2026 long paper, preview)" (2026). https://preview.aclanthology.org/ingest-acl/2026.acl-long.1880/
  3. arXiv, "Failure Makes the Agent Stronger: Enhancing Accuracy through Structured Reflection for Reliable Tool Interactions" (2025). https://arxiv.org/html/2509.18847v2
  4. arXiv, "PALADIN: Self-Correcting Language Model Agents to Cure Tool-Failure Cases" (2025). https://arxiv.org/pdf/2509.25238
  5. arXiv, "SHIELDA: Structured Handling of Exceptions in LLM-Driven Agentic Workflows" (2025). https://arxiv.org/pdf/2508.07935
  6. ICML 2025, "The Berkeley Function Calling Leaderboard (BFCL): From Tool Use to Agentic Evaluation of Large Language Models" (2025). https://icml.cc/virtual/2025/poster/46593
  7. OpenTelemetry, "Inside the LLM Call: GenAI Observability with OpenTelemetry" (2026). https://opentelemetry.io/blog/2026/genai-observability/
  8. Together AI, "Fine-tuning for function calling". https://docs.together.ai/docs/fine-tuning-function-calling.md
  9. arXiv, "Confident Learning: Estimating Uncertainty in Dataset Labels" (2019). https://arxiv.org/pdf/1911.00068
  10. Google Research (FAccT 2022), "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data