Skip to content

Agent, workflow and domain-reasoning data

Policy-following service agent data: conversations, policies and back-office actions

Quick answer

Policy-following agent training data links three records for every service case: the policy version in force when the customer made contact, the conversation itself, and the back-office actions staff took, with account or order state captured before and after. That linkage lets you train an agent to call a refund or order-change tool only when a clause allows it, and to check its work against the end state. Transcripts alone teach tone, not rules. The costly work is reconstructing state and labeling which historical actions complied.

By SourceX Editorial · Updated

This page covers sourcing those records for training. For building test tasks from them, see policy-following evaluation for customer support agents; for tickets and transcripts generally, see customer support AI training data and eval sets. The agent training data hub maps the cluster.

Why transcripts without policies and state teach the wrong lesson

A transcript shows what staff said, not which rule permitted an action or what changed in the system, so an agent trained on transcripts alone learns to sound compliant without being compliant. "Policy" here means written business rules, such as a 30-day refund window or an approval limit, not the reinforcement-learning sense of a model's action-selection function.

τ-bench (Tool-Agent-User interaction benchmark) is built from three parts: a database with programmatic APIs, a domain policy document, and user-scenario instructions with ground-truth annotations, graded by comparing the end-of-conversation database state with an annotated goal state [1]. Policy text alone is not enough either. The Alignment Studio authors argue that fine-tuning directly on policy documents gives a model policy vocabulary, not the ability to apply the policy or judge compliance [2]. The model has to see the policy, the situation and the correct action together.

The three linked records, field by field

A usable case joins a versioned policy, a turn-level conversation and a timestamped action log on shared keys. If any join is missing, the case shrinks to a transcript. The table lists the fields to require and where each record usually lives.

RecordFields to requireTypical source systemsDefect to check for
Policy, versionedDocument ID, version or effective-date range, clause IDs, approval limits by role, superseded-by linkKnowledge-base or wiki revision history, macro (saved reply) change logs, approval matricesOnly the current text survives; edits overwrite history
ConversationCase ID, channel, speaker role per turn, timestamps, customer-visible replies kept apart from internal notes, transfer eventsHelp-desk platforms such as Zendesk, Salesforce Service Cloud or Freshdesk; chat and contact-center systemsBot and human turns not distinguished; internal notes merged into the dialogue
Actions and stateAction type, actor (staff, automation or customer), system, arguments (amount, item, reason code), timestamp, approver, before and after values of each changed fieldOrder management, billing and payment systems, CRM and ERP audit tablesOnly current state stored; automated triggers indistinguishable from staff writes

The join keys (case, customer, order and invoice IDs) must survive de-identification unchanged in all three records. See packaging linked records from multiple business systems, and SOP and knowledge base datasets for AI agents for policy documents as standalone data.

Where the policy lives at run time decides what you must buy

Decide first whether your agent will read the policy in its prompt, carry it in its weights, or defer to an external rule checker, because each design needs different records.

DesignPublished exampleWhat each training case must carryWhen the policy changes
Policy in the promptτ-bench supplies the domain policy document to the agent [1]The exact clause text in force for that case, so the model learns to apply whatever text it is givenSwap the document; old cases stay valid if each one carries its own version
Policy in the weightsPA3 trains on domain tool-calling data that follows τ-bench's business policies, and withholds the policies at evaluation [3]Enough cases per clause and outcome for the rule to be learned implicitlyRetrain; cases from superseded versions become wrong labels
Policy enforced outside the modelPolicyLLM compiles policy text into decision graphs checked with the Z3 solver and logs each decision [4]Clause-level decisions to validate the compiled rules, and dialogues showing when to call the checkerRecompile the rules; dialogue data ages more slowly

The PA3 authors note that their training and test data share the same policy descriptions [3], so a benchmark score after that kind of training shows how well the model applies a policy it was trained on, not whether it can apply a new one. For the prompt design, policy revisions inside your case window are useful counterfactuals: refund requests handled under a 30-day window in March and a 14-day window in June teach the model to read the clause instead of memorizing a number. Ask suppliers how many policy versions their history spans. The PolicyLLM authors call fine-tuning brittle under policy change [4], another reason to keep a version on every case.

Turning staff actions into tool calls an agent can make

Each write a staff member made in a back-office system is a candidate tool call. It becomes training data only after it is mapped to your agent's tool schema, attached to the right conversation turn, and separated from automation. The mapping below covers typical refund and order-change workflow data; the tool names are examples, so substitute your agent's own schema.

Recorded eventAgent tool callArguments to recoverReconstruction problem
Refund created against an order paymentissue_refundOrder ID, line items, amount, reason codePartial refunds split across events; refund issued after the chat closed
Item size or color changed on an unfulfilled ordermodify_order_itemsOrder ID, item, new variantChange made by a warehouse substitution rule, not staff
Shipping address editedupdate_shipping_addressOrder ID, pseudonymized addressAddress-validation service rewrites the stored value
Case reassigned to a tier-2 or supervisor queuetransfer_to_humanReason, summary passed onRouting rules produce the same event as a staff decision
Supervisor approval recordedrequest_approvalApprover role, amount, clauseApproval given in a side chat and never logged
Customer and order looked up in the admin consoleget_customer_orders (read)Customer IDReads are rarely logged and must be inferred

The last row matters most. Back-office systems usually log writes, not reads, yet a well-behaved agent looks up the order before changing it, so reconstructed reads should carry a flag such as inferred_read that lets you weight or drop them. Staff also hold powers your agent will not, such as manual ledger adjustments; exclude those cases or relabel the target as an escalation. For the general mechanics, see reconstructing agent trajectories from ticket and case histories and API call logs as tool-use training data.

Some workflows split actions between agent and customer: τ²-bench added a dual-control telecom domain in which both the agent and the simulated user operate tools [5]. Real troubleshooting cases need customer-side steps, such as restarting a device, recorded as user actions.

Common fine-tuning formats are JSONL, one example per line, with a tools list, assistant tool calls carrying an id, and results returned in tool-role messages [6]. Tool definitions are typically function objects with a name, a description and JSON Schema parameters [7], and JSON Lines requires UTF-8 with one valid JSON value per line [8]. See also training data for tool use and function calling.

Illustrative record: an order change granted, a refund declined

The record below shows one case with all three parts linked. In delivery it would sit on a single JSONL line; it is expanded here for reading.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "case_id": "c_5e91a2",
  "channel": "chat",
  "opened_at": "2026-06-03T14:08:00Z",
  "date_offset_applied": true,
  "policy": {
    "doc_id": "returns-and-order-changes",
    "version": "2026-05-12",
    "clauses": {
      "2.1": "Unfulfilled orders: an item may be changed to another size or color once.",
      "4.3": "Refunds to the original payment method within 30 days of delivery; after that, decline and explain."
    }
  },
  "state_before": {
    "orders": [
      {"order_id": "o_71c4", "status": "unfulfilled", "items": [{"sku": "JKT-M-NVY", "qty": 1}]},
      {"order_id": "o_33b0", "status": "delivered", "delivered_at": "2026-04-22", "total": 64.00}
    ]
  },
  "tools": ["get_customer_orders", "modify_order_items", "issue_refund", "transfer_to_human"],
  "messages": [
    {"role": "user", "content": "Can I switch my jacket to a large? Also I want a refund for the boots from April."},
    {"role": "assistant", "provenance": "inferred_read",
     "tool_calls": [{"id": "call_1", "type": "function",
       "function": {"name": "get_customer_orders", "arguments": "{\"customer_id\": \"u_9d20\"}"}}]},
    {"role": "tool", "tool_call_id": "call_1", "content": "{\"orders\": [\"o_71c4 unfulfilled\", \"o_33b0 delivered 2026-04-22\"]}"},
    {"role": "assistant", "provenance": "logged_write", "clause": "2.1",
     "tool_calls": [{"id": "call_2", "type": "function",
       "function": {"name": "modify_order_items", "arguments": "{\"order_id\": \"o_71c4\", \"from_sku\": \"JKT-M-NVY\", \"to_sku\": \"JKT-L-NVY\"}"}}]},
    {"role": "tool", "tool_call_id": "call_2", "content": "{\"status\": \"ok\"}"},
    {"role": "assistant", "clause": "4.3",
     "content": "Your jacket is now a large. The boots were delivered on April 22, more than 30 days ago, so they are outside our refund window."}
  ],
  "state_after_diff": [
    {"table": "order_items", "order_id": "o_71c4", "field": "sku", "before": "JKT-M-NVY", "after": "JKT-L-NVY"}
  ],
  "labels": {"compliance": "compliant", "outcomes": ["granted:2.1", "declined:4.3"], "qa_result": "pass", "reopened_within_14d": false}
}

Three details carry the value. The date offset shifts every date in the case equally, so the 42-day gap that triggers clause 4.3 survives de-identification. The declined refund has no write action, which teaches the agent that a correct conversation can end without one. The provenance and clause keys are not part of standard provider training formats, so keep them in the delivered file and strip them when converting for your trainer.

Labeling compliance, exceptions and refusals

Historical staff behavior is evidence, not ground truth. Label every action against the clause in force before it becomes a training target, or the agent will learn your staff's goodwill habits along with your rules.

LabelHow to detect it in the recordsTraining use
CompliantAction matches the clause for that version; QA pass where scored; not reversedSupervised target
Approved exceptionApproval record from a role authorized for that amount, with a reasonTarget becomes request_approval or a transfer, not the exception itself
Unapproved deviationCredit or refund outside any clause, no approval on recordExclude as a target; use as a negative or correction example
RefusalRequest declined with the clause explained; no write actionSupervised target
Corrected laterReversal, QA failure or supervisor override after the caseRevise-or-keep example

The last row maps onto published methods: CoMAP's agent data includes expert-draft supervision and reflection samples, some showing when a draft action should be revised and when it should be kept [9]. QA-failed and reversed cases supply real drafts and real corrections. Exceptions need rationale too. One study reports that LLMs apply policies rigidly even when that is impractical, and that supervised fine-tuning with human explanations did better than prompting approaches [10], so ask whether approvers recorded a reason.

Related record types: exception handling records, approval and rejection records, human override and correction logs, agent-to-human handoff data and QA scores and corrections.

Coverage: specify a grid, not a ticket count

Specify coverage as intent by clause by outcome by policy version, and ask for counts in each cell. Real service histories are usually dominated by routine grants, while the cases that teach rules sit at the boundaries: day 29 versus day 31, or an amount just above an approval limit.

Licensed records supply the real distribution of requests, phrasing and messy context. Programmatic generation can fill sparse boundary cells: SOPBench represents procedures as graphs of prerequisite checks that must run before service actions, with rule-based verifiers, and notes that its instance generation can scale to produce training data [11].

Expect a blend. Cognitive Kernel mixed open instruction-following, function-calling and agent-trajectory data with a small manually annotated set [12], and PA3 paired general function-calling data with domain tool-calling data [3]. Compare licensed records, commissioned demonstrations and synthetic trajectories, and see data for building user simulators if you will expand dialogues with simulated customers.

De-identifying transcripts and account tables with one key map

Pseudonymize every identifier through one key map shared by the conversation text and the structured tables. Otherwise the trajectory breaks: the order number a customer types must match the order ID in the tool call and in the state snapshot.

  • Deterministic surrogates. Each customer, order or invoice ID maps to one surrogate everywhere, including inside free text and tool-call arguments.
  • One date offset per case or customer. Shift all dates together so refund windows and SLA clocks still compute, and keep amounts exact, because approval limits depend on them.
  • Free text. Customers paste card numbers and addresses into chat. The open-source Presidio project cautions that its ML-based detection cannot guarantee finding all sensitive information [13], so sample-check output by hand.
  • Key custody. The key map stays with the supplier. UK ICO guidance (March 2025, under review as of October 2026) says pseudonymized data remains personal data and that attempting to reverse pseudonymization can be a criminal offense [14]. A supplier relying on California's definition of deidentified information must contractually bind recipients to its conditions [15], so expect that clause in the license.

SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery. The method is recorded for each dataset and a sample is checked after processing. No de-identification method is perfect.

Rights questions for policies, conversations and staff actions

The three records raise different rights questions, so review each one rather than treating the dataset as a single grant.

  • Policies. Internal policies belong to the supplier, but some embed third-party terms, such as carrier or marketplace rules. Confirm the license covers policy text in training prompts and in a deployed agent's context.
  • Conversations. Check the privacy notice in force when each conversation happened. A 2024 FTC staff blog post warned that adopting AI-training uses through a surreptitious, retroactive change to terms or privacy policy may be unfair or deceptive [16]. Outsourced queues add the client's contract terms; see sourcing AI training data through BPOs.
  • Staff notes and actions. Internal notes and approval rationales are employee-authored records.
  • Downstream documentation. As of October 2026, Colorado's SB26-189, signed 14 May 2026, will require developers of automated decision-making technology that materially influences consequential decisions, including lending and insurance, to give deployers documentation covering training data categories from 1 January 2027 [17]. Service agents that grant or deny lending or insurance requests may fall in scope.
  • Scope of use. Name training, evaluation, sandbox replay of account states and synthetic variants explicitly; see license terms for agent workflow data.

SourceX sources operational datasets from US companies, including support histories, documents and finance workflows, and manages the licensing process. Datasets are sourced on request, not held in stock, so a request does not guarantee a match, and every release is approved by the supplying company. Each dataset goes through rights review (does the business own or may it share the records, and are required consents in place) and is delivered under a license defining the included records, allowed uses, term and delivery. To start, describe the linked policy, conversation and action records your agent needs.

What to specify, and red flags in an offer

A good request names the policy versions, the action systems and the labels, not only a conversation count.

Specify:

  • Workflows and intents, such as order changes, refunds, cancellations and escalations
  • Policy documents and macros with revision dates across the case window, plus approval limits by role
  • Conversation turns with roles and timestamps, internal notes kept separate
  • Action logs with actor, arguments, approver and before and after values, joined on shared keys
  • State snapshots at case open, or event logs to rebuild them
  • Per-action compliance labels, approval reasons, QA results and reversal flags
  • Minimum counts per coverage cell, including refusals and approved exceptions
  • One key map, a per-case date offset and exact amounts
  • JSONL with tool schemas, or source tables plus a mapping
  • License scope: training, evaluation, sandbox replay, synthetic expansion

Red flags:

  • Transcripts sold as "agent data" with no action log or state
  • One current policy applied to years of cases
  • Automation events indistinguishable from staff decisions
  • No refusals or escalations in the sample
  • Irreversible actions, such as payouts, with no separate slice for irreversible-action evaluation
  • Policies that read like a public benchmark domain, presented as real business rules

Need policies, conversations and actions linked by case?

Describe the workflows, policy versions, action systems and labels your service agent needs. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery. Nothing is contracted until a supplier agrees. Specify your workflow dataset.

Sources

  1. Yao, Shinn, Razavi, Narasimhan (Sierra; arXiv:2406.12045), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  2. IBM Research authors (arXiv:2403.09704), "Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations" (2024). https://arxiv.org/pdf/2403.09704
  3. arXiv:2603.14602, "PA3: Policy-Aware Agent Alignment through Chain-of-Thought" (2026). https://arxiv.org/pdf/2603.14602
  4. ICML 2026 (conference listing), "PolicyLLM: Neuro-Symbolic Policy Extraction and Enforcement for Runtime AI Governance" (2026). https://icml.cc/virtual/2026/78632
  5. Sierra Research tau2-bench repository (indexed by Algolia DocSearch), "τ²-bench / τ³-bench repository documentation" (2026). https://docsearch.algolia.com/mcp/docs/repo/sierra-research/tau2-bench
  6. Together AI documentation, "Fine-tuning for function calling". https://docs.together.ai/docs/fine-tuning-function-calling
  7. Predibase documentation, "Function Calling (fine-tuning user guide)". https://docs.predibase.com/user-guide/fine-tuning/function_calling
  8. jsonlines.org, "JSON Lines". https://jsonlines.org/
  9. arXiv:2606.02372, "CoMAP: Co-Evolving World Models and Agent Policies for LLM Agents" (2026). https://arxiv.org/pdf/2606.02372
  10. arXiv:2503.02976, "Teaching AI to Handle Exceptions: Supervised Fine-tuning with Human-aligned Judgment" (2025). https://arxiv.org/html/2503.02976v2
  11. alphaXiv listing of arXiv:2503.08669, "SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints" (2025). https://alphaxiv.org/abs/2503.08669
  12. arXiv:2409.10277, "Cognitive Kernel: An Open-source Agent System towards Generalist Autopilots" (2024). https://arxiv.org/pdf/2409.10277
  13. Microsoft presidio project (indexed on pkg.go.dev), "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  14. Information Commissioner's Office (UK), "Pseudonymisation (anonymisation guidance)" (2025). https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/pseudonymisation/
  15. California Legislature (California Legislative Information), "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  16. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  17. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data