Industry-specific operational data
E-commerce order-support conversations with order state for AI agents
Quick answer
Order-management agents need more than chat transcripts. Useful retail service data joins each conversation to the order record before and after contact, the return or refund policy version in force that day, and every action taken: refund, exchange, cancellation, reship or address change. Public benchmarks in this domain are synthetic, so buyers who want real distributions must license operational records from retailers, with consistent pseudonymization and documented permission for AI training.
By SourceX Editorial · Updated
Why transcripts alone fail order-management agent evaluation
A transcript cannot grade a tool-using agent, because the correct outcome is a change to the order system, not a sentence. τ-bench, a widely cited benchmark for this domain, places agents in a retail domain with a database, a policy document and a simulated user, and grades the run by comparing the final database state with a goal state [1]. Real data has to support the same check: what the order looked like at first contact, what the human agent changed, and whether that change was allowed.
The public alternatives show the gap. τ-bench users are LLM-simulated [1], CRMArena-Pro runs on synthetic CRM records [3], and the Action-Based Conversations Dataset (ABCD) was collected with crowdworkers role-playing customers and agents against written guidelines [2]. These are good test harnesses, but they do not contain real WISMO ("where is my order") traffic, carrier exception codes, partial-shipment edge cases or the way customers actually phrase damaged-item claims.
Real conversation corpora that do exist often carry research-only or non-commercial terms [4]. That is the core sourcing problem for teams building production agents: the data lives inside retailers' helpdesk platforms (Zendesk, Gorgias, Salesforce Service Cloud, Kustomer) and order systems (Shopify, Salesforce Commerce Cloud, Manhattan, an in-house OMS), and it must be joined and licensed deliberately. For the generic ticket version of this data, see customer support ticket datasets.
What an order-support record should contain
A usable record is one conversation linked to one or more orders, with state snapshots and an action log on both sides of the contact. The fields below are the minimum an evaluation lead should ask for; most retailers can produce them only by joining helpdesk exports to OMS and payment events.
- Turns with roles: customer, human agent, bot, macro. Macro and bot turns must be labeled, or SFT data will teach canned text as if it were reasoning; see separating bot, macro and human turns.
- Order state snapshots: status, line items with SKU tokens, fulfillment and carrier events, payment capture status, prior returns, captured at first contact and at resolution.
- Actions taken: typed records such as
refund(line, amount_band, reason),exchange(line, new_sku),cancel(order),reship(line),update_address(order), with timestamps and the system that executed them. - Policy version: the return window, restocking fee and final-sale rules effective on the contact date, as text or a versioned identifier.
- Exception flag: whether the action broke policy (goodwill refund outside the window, waived restocking fee) and the approver role.
- Outcome: resolution code, reopen within 7 days, escalation, chargeback or CSAT where collected.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"conversation_id": "conv_8f21",
"channel": "chat",
"policy_version": "returns_v14_2026-03",
"orders": [{
"order_token": "ORD_0193",
"state_at_contact": {"status": "delivered", "delivered_days_ago": 34,
"lines": [{"line": 1, "sku_token": "SKU_77A", "price_band": "50-100", "return_eligible": false}]},
"state_at_resolution": {"status": "partially_refunded"}
}],
"turns": [
{"role": "customer", "text": "The zipper broke on the jacket from order ORD_0193."},
{"role": "agent_human", "text": "I'm sorry about that. I can offer a refund for the jacket as a one-time exception."}
],
"actions": [{"type": "refund", "line": 1, "amount_band": "50-100", "reason": "defect",
"policy_exception": true, "approved_by_role": "team_lead"}],
"outcome": {"resolution": "refunded", "reopened_7d": false, "csat": 5}
}
The record pairs a policy that says "not eligible" with an action that refunds anyway. That disagreement is the most valuable signal in the file, as long as it is labeled.
Policy exceptions, versioning and grading
Policy adherence can be scored only if the policy text is versioned and exceptions are labeled, since human agents routinely bend rules to keep customers. ABCD formalized the idea that service agents follow guidelines that constrain which actions are allowed [2], and τ-bench tests whether agents follow a domain policy, including declining requests it forbids [1]. Real retail logs complicate both: the same request can be correct under the March policy and wrong under the July one, and a goodwill exception by a supervisor is not a mistake.
Ask for three labels per action: policy-compliant, approved exception, or unapproved deviation. Without them, a model trained by SFT or a reward model will learn that refunds outside the window are normal. For eval sets, decide whether exceptions count as correct agent behavior or are excluded; the companion guide on policy-following evaluation for support agents covers scoring design, and policy-following service agent data covers the cross-industry pattern.
Reliability matters as much as single-run accuracy. τ-bench introduced pass^k, the chance an agent succeeds on all k trials of the same task [1]. Real orders let you build repeated-trial sets on the long tail (split shipments, gift orders, marketplace sellers, BNPL refunds) that synthetic generators rarely produce.
Pseudonymization that keeps orders traceable
Direct identifiers must be replaced consistently, so the same customer, order and tracking number map to the same token in every turn and table. Random redaction per message breaks the join between "my order 4471" in turn 2 and the OMS record, which destroys the state-based grading the data was bought for.
Practical rules buyers should specify:
- Names, emails, phone numbers and street addresses replaced with realistic surrogates; ZIP truncated to three digits or banded region.
- Order numbers, RMA numbers and carrier tracking numbers tokenized with a keyed mapping held by the supplier, never shipped to the buyer.
- Payment fragments (last four digits, card brand plus expiry) removed, not tokenized; amounts banded unless exact amounts are needed for refund math and the license allows them.
- Free text scanned for identifiers customers type unprompted: gift messages, delivery instructions, photos of packing slips.
Request the de-identification method in writing and a reviewed sample before full delivery. Customer contracts and helpdesk privacy notices also determine whether this data may be used for training at all; the FTC has warned that quietly expanding data use to AI training can be unfair or deceptive [5]. The customer contracts and DPAs guide explains what to check.
Buyer checklist for order-support data requests
A precise request names the joins, the time window and the labels, not just "support chats." Use this as a starting template.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Item | What to specify | Why it matters |
|---|---|---|
| Contact types | WISMO, return, exchange, cancel, address change, damaged item | Controls intent mix and long-tail coverage |
| Channels | Chat, email, voice transcripts, SMS | Voice adds ASR noise and different repair patterns |
| Order joins | State at contact and resolution, carrier events | Needed for state-based grading [1] |
| Action log | Typed actions with system of record | Becomes tool-call targets for function-calling training |
| Policy history | Versioned return and refund policies | Enables adherence scoring by date |
| Exception labels | Compliant, approved exception, deviation | Prevents reward hacking toward over-refunding |
| Turn source | Human, bot, macro | Keeps SFT targets clean |
| Privacy | Consistent tokenization, documented method, sample review | Keeps orders traceable without identifiers |
| Rights | Training and eval use, term, delivery format | Confirms the retailer may license the records |
Returns-specific reason codes and RMA dispositions are covered in returns and RMA reason data, and the wider cluster is in industry-specific operational data. For tool schemas built from action logs, see tool use and function calling training data.
How SourceX handles requests for retail service data
SourceX sources operational datasets, including support and sales histories, from US companies on request and manages the licensing process; nothing is held in stock, and a request does not guarantee a match. Buyers describe the data they need, such as order-support conversations joined to OMS events, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and delivery runs through private, access-controlled workflows after an executed agreement. Retail buyers can start from e-commerce buyer requirements or the buyers page.
Request ecommerce customer service data for AI agents
If your order-management agent needs real conversations tied to order state, policy versions and actions, describe the fields, channels and joins you need. SourceX looks for US retailers holding that data and runs the assessment, licensing and delivery process, and nothing is contracted until a supplier agrees. Describe your data request on the SourceX buyers page.
Sources
- Sierra Research (arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- Chen et al., NAACL 2021 (arXiv:2104.00783), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
- Salesforce AI Research (arXiv:2505.18878), "CRMArena-Pro: Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions" (2025). https://arxiv.org/pdf/2505.18878
- LLM Configurator, "ShareGPT dataset entry". https://llmconfigurator.com/en/datasets/sharegpt
- Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.