Agent, workflow and domain-reasoning data
SOP-to-execution pairs: procedures linked to how work was actually done
Quick answer
SOP execution pairs for agent training link one revision of a standard operating procedure (SOP) to the cases that were actually worked under it: the step-level record of what staff did, in which systems, and whether each step followed, skipped or departed from the procedure. The link is what an agent learns from. Source pairs from systems that already store a procedure ID, checklist completion or process code on each case, match every case to the revision in force at the time, and require labeled deviations with their recorded reasons.
By SourceX Editorial · Updated
Why an SOP library alone cannot teach procedure-following
A procedure corpus shows what should happen; only execution records show what did happen, where people branched and which departures were accepted. An agent that must follow procedures needs both halves plus the label that connects them. SOP corpora alone are covered on the SOP and knowledge base datasets page and the SOP and playbook licensing page; this guide, part of the agent training data hub, covers the link between a procedure version and its executions.
Research benchmarks show the shape of the data. SOP-Bench pairs more than 2,000 tasks drawn from expert-authored SOPs in 12 business domains with executable interfaces and ground-truth outputs [1]. SOPBench models procedures as graphs of prerequisite checks that must pass before a service action, scored by rule-based verifiers [2]. SOP-Agent represents an SOP as a decision graph and restricts the agent to the tools the procedure permits [3], and τ-bench combines policy documents, databases and APIs, and annotated user scenarios in retail and airline domains [5].
What these designs cannot supply is what operating companies hold: procedures revised over years, staff who added steps the document never mentioned, and departures that a supervisor approved or that later caused rework. Those are the cases where an agent must decide whether to follow the text, take a permitted branch or escalate.
Where procedures and executions are already linked
The most reliable pairs come from systems that store a procedure reference on the case while the work happens, so the link need not be inferred later. The table lists common sources and the gap to probe in each.
| Where the work happens | What links a case to a procedure | Execution evidence | Gap to probe |
|---|---|---|---|
| IT service management (incidents, requests, changes) | Knowledge-article or runbook ID on the ticket; change templates | Work notes, state and assignment changes, linked tasks | Article attached at closure, not used during the work |
| Regulated manufacturing (MES, electronic batch records) | Master batch record or recipe version the batch ran against | Step sign-offs, recorded values, deviation report numbers | Deviations and investigations kept in a separate quality system |
| Maintenance (CMMS) | Job plan or preventive-maintenance procedure on the work order | Task completions, readings, parts used, technician notes | Tasks bulk-closed at the end of the job |
| Back office (accounts payable, claims, KYC, underwriting) | Process code, checklist or queue rule on the case | Checklist ticks, field edits in ERP or core systems, approvals | Only the current SOP kept; old revisions overwritten |
| Contact centers | Script or call-flow ID; QA scorecard criteria | Transcripts, desktop events, disposition codes, QA scores | QA scores exist for a small share of calls |
| Workflow engines and RPA | Deployed process-definition version | Instance history, task completions, bot run logs | Human work outside the engine is not recorded |
Process-oriented sources usually export as event logs. IEEE 1849-2023 (XES) defines an XML format for event logs and event streams and supersedes the 2016 edition [6]. OCEL 2.0 adds object-centric logs in SQLite, XML or JSON that record attribute changes and qualified links between events and objects [7]. That suits procedures touching several records at once, such as a vendor, its open invoices and a payment run.
Public logs, such as the 12 IEEE Task Force on Process Mining logs bundled on 4TU.ResearchData under CC BY 4.0 for a process discovery benchmark [8], were not built as SOP pairs, so confirm that procedure text exists before treating any log as a pair. See process mining event logs, OCEL 2.0 object-centric logs and ticket histories as trajectories. For maintenance job plans, see work order datasets; for incident response, runbook execution records.
Matching each case to the procedure revision in force
A pair is valid only if the case is joined to the SOP revision effective when the work was done, not to today's version. Joining every historical case to the current procedure is an easy construction error to make. It mislabels correct work under revision 3 as deviations from revision 4 and teaches an agent the wrong rule.
Ask for document-control fields, not just page text: SOP ID, revision number, effective date, superseded date, approver, and a change summary or redline naming the affected steps. Wiki page history is a weak substitute, because a typo fix and a substantive change both create a new page version. Where the quality system records read-and-understood training, the performer's training date on each revision is a useful cross-check on the effective date.
Specify the join rule in writing. Use case open time for short cases, step timestamps for long ones, and a flag for cases that straddle a revision boundary. Keep revisions with cases on both sides of a change, because those before-and-after sets are what test whether an agent updates its behavior when a procedure changes. The same effective-dating logic applies to matching records to the notice in force at collection.
Adherence and deviation labels that hold up in training
Every executed step needs a label relating it to an SOP step, and every departure needs a type, any recorded reason and an outcome. Without the outcome, nobody can tell a sensible workaround from an error.
| Label | Meaning | Usual evidence | Training use |
|---|---|---|---|
| Followed | Step done as written, in order | Checklist completion, matching system event | Positive demonstration |
| Permitted branch | SOP decision point taken | Field value meeting the branch condition | Conditional logic |
| Skipped | Required step has no evidence | Missing event or unchecked item | Negative example; verifier training |
| Reordered | Steps done out of sequence | Timestamps | Test whether order mattered |
| Added (tacit) step | Work the SOP does not describe | Event with no SOP step reference | Undocumented practice |
| Justified deviation | Departure with recorded reason and approval | Deviation report, waiver, escalation note | When and how to escalate |
| Unjustified deviation | Departure without reason, or later reversed | Rework, reopen, QA or audit finding | Negative example |
| Not applicable | Case routed to the wrong procedure | Reassignment, procedure switch mid-case | Procedure selection |
Labels come from three places, in decreasing order of reliability. Explicit records (checklist ticks, deviation reports, waivers) are strongest. Conformance checking, the process-mining technique that aligns an event log with a model of the procedure, reports moves the log has but the model lacks (candidate tacit steps) and moves the model requires but the log lacks (candidate skips). Annotation of free-text work notes, by people or a model, fills gaps but needs an adjudicated sample and agreement figures.
Treat recurring tacit steps as signal, not noise. When many performers add the same step, such as checking a second system before acting, it often reflects a requirement the document omits. Runtime research also treats departures as events to flag rather than forbid: Compile, Then Page compiles an SOP into an executable program whose soft enforcement lets deviations happen but flags them [4], a behavior that recorded, justified deviations can help teach. Related: exception handling records and decision records with rationale.
What one pair looks like
A usable pair keeps the step text with its revision, references steps by ID, records how the link was made, and stores added steps and substitutions instead of smoothing them away.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"pair_id": "pair-000412",
"sop": {
"sop_id": "AP-017",
"title": "Vendor bank detail change",
"revision": "4",
"effective_from": "2025-03-01",
"superseded_on": "2025-11-14",
"approved_by_role": "finance_controller",
"steps": [
{"step_id": "S1", "text": "Log the request and attach the source document"},
{"step_id": "S2", "text": "Call the vendor on the phone number already on file, not the number in the request"},
{"step_id": "S3", "text": "Update the bank account in the ERP vendor record"},
{"step_id": "S4", "text": "Obtain second approval from an AP supervisor"}
]
},
"case": {
"case_id": "c-88213",
"opened_at": "2025-06-03T14:02:11Z",
"closed_at": "2025-06-04T09:40:52Z",
"linkage": {"method": "sop_id_field_on_ticket", "confidence": "exact"}
},
"executed_steps": [
{"seq": 1, "sop_step": "S1", "system": "itsm", "action": "ticket_created", "actor": "ap_clerk#p17", "ts": "2025-06-03T14:02:11Z", "label": "followed"},
{"seq": 2, "sop_step": null, "system": "erp", "action": "viewed_open_invoices", "actor": "ap_clerk#p17", "ts": "2025-06-03T14:05:40Z", "label": "added_step"},
{"seq": 3, "sop_step": "S2", "system": "telephony", "action": "outbound_call_number_on_file", "actor": "ap_clerk#p17", "ts": "2025-06-03T14:11:02Z", "label": "followed"},
{"seq": 4, "sop_step": "S3", "system": "erp", "action": "vendor_bank_account_updated", "actor": "ap_clerk#p17", "ts": "2025-06-03T14:20:37Z", "label": "followed"},
{"seq": 5, "sop_step": "S4", "system": "erp", "action": "change_approved", "actor": "controller#p03", "ts": "2025-06-04T09:38:15Z", "label": "justified_deviation"}
],
"deviations": [
{"sop_step": "S4", "type": "substituted_approver", "reason": "AP supervisor on leave; controller approved under delegation", "recorded_in": "ticket_comment"}
],
"outcome": {"qa_review": "pass", "reopened": false, "payment_returned": false}
}
Checking a sample before you license
A sample should prove that the linkage is real, that deviations appear in useful numbers, and that outcomes exist to judge them. Ask for cases spanning at least two revisions of each SOP, including labeled deviations, then check:
- Linkage by method: the share of cases whose revision resolves from a stored field versus inference from text similarity or dates alone.
- Step resolution: events map to individual SOP steps, not just case open and close.
- System coverage: if the SOP says "verify in the vendor portal," portal events must be present, or that step cannot be scored.
- Label mix per SOP: counts of followed, branch, skip, added and deviation labels. A sample where nearly every case follows the happy path teaches little about judgment.
- Deviation records: reason text, approver role and later outcome (QA result, reopen, rework, audit finding).
- Timestamps: whether they record when work happened or when it was keyed in.
- Duplication: happy-path cases are often near-identical. Lee et al. found near-duplicates common in language-model training sets and that models trained on deduplicated data emitted memorized text about ten times less often [9]; ask how many distinct execution patterns each SOP has.
- Documentation: linkage method, labeling guidelines and known gaps, ideally as machine-readable metadata such as Croissant-RAI, which extends Croissant with responsible-AI fields for data life cycle and labeling [10].
The general sample procedure is in evaluating an agent data sample before you buy.
Using pairs for fine-tuning versus procedure-following evaluation
For supervised fine-tuning (SFT), a pair becomes a conditioned demonstration: the SOP revision and case state go in, and the next action or remaining action sequence comes out. For evaluation, a pair becomes a task with a checkable end state, and its whole SOP family stays out of training.
In SFT data, include justified deviations with the justification in the target, so the model learns to escalate and document instead of improvising. Use unjustified deviations as negatives for preference data or step-verifier training, not as demonstrations. Policy-following service agent data covers the conversational variant, where customer dialogues are paired with policies and back-office actions.
For evaluation, split by SOP family rather than by case, or near-identical cases leak across the split. Score end states the way τ-bench compares the final database state with an annotated goal state, and report reliability across repeated runs with its pass^k measure [5]. SOPBench's verifiers show how to check that prerequisites ran before an action [2].
Real revision histories add a test that is hard to build otherwise: train or prompt on revision N, give the agent revision N+1, and check that it adopts the change. This is the one split where a family deliberately spans both sides. Suite design is covered in agent evaluation task suites.
Rights, confidentiality and privacy questions specific to SOP pairs
SOP pairs combine two sensitive assets: procedures, which are often confidential know-how, and case records containing customer and employee personal data. Both need rights review before anything is licensed.
On the procedure side, ask whether any SOP embeds third-party material, such as an equipment maker's manual, a franchisor's operations manual or a licensed quality-system template, which the supplier may not be able to license onward. Fraud and security controls (callback rules, approval thresholds, access steps) may need redaction. The license should state whether SOP text may appear in model outputs or published benchmarks; license terms for agent workflow data covers replay, derived tasks and benchmark rights.
On the execution side, performer-level records can read like performance evidence, and deviation logs can overlap with disciplinary or audit files. Request role-level actor IDs with stable pseudonyms, names removed from free-text notes, and HR and investigation records excluded. Monitoring and consent questions for desktop capture are covered in collecting computer-use demonstrations at work.
When SourceX sources this kind of data, each dataset goes through rights review, which checks that the business owns or may share the records and that required consents are in place. Personal details such as names and account numbers are removed or replaced before delivery and a processed sample is checked, but no de-identification method is perfect, so plan your own review of free-text notes. Buyers who would rather describe the pairs than find the companies holding them can submit an SOP-pair data request to SourceX; the categories SourceX describes are kinds of data it sources, not inventory under contract.
Red flags in SOP-pair offers
- Every case links "exactly" although the source system has no procedure field.
- The deviation rate is zero, or every deviation carries the same boilerplate reason.
- SOP text was written or rewritten for the dataset instead of exported from document control.
- All executions come from one team or site, so a local habit looks like the procedure.
- Interview reconstructions or synthetic steps are presented as system records (compare licensed records, commissioned demonstrations and synthetic trajectories).
Writing the request
A useful request names the procedures, linkage and labels you need, so a supplier can tell quickly whether its systems hold them. Adapt this outline; the agent data specification guide covers the rest.
- Procedure families: for example vendor master changes, warranty claims or preventive maintenance.
- Revisions: minimum revisions per SOP and the date range, with cases on both sides of each revision.
- Linkage: stored SOP or process ID preferred; inferred linkage accepted only with a confidence field.
- Granularity: step-level events with timestamps from every system the SOP names.
- Labels: the adherence taxonomy above, deviation reasons, approver roles and outcome fields.
- Volume: cases per SOP family and the minimum share of deviation cases, stated as your own targets.
- Privacy: pseudonymous role-level actors, free-text redaction, exclusions.
- Uses: SFT, evaluation, verifier training, and whether you intend to publish a benchmark.
- Format: pairs as JSON Lines, or XES or OCEL 2.0 logs, plus SOP revisions as Markdown or PDF.
Building agents that must follow real procedures?
Describe the procedures, linkage and labels you need. SourceX looks for US businesses whose systems hold them, checks the data and the supplier's licensing permissions, agrees allowed uses in a license, and coordinates delivery and future purchases. Nothing is contracted until a supplier agrees, and a request does not guarantee a matching dataset. Specify your SOP and execution data.
Sources
- arXiv, "SOP-Bench: Complex Industrial SOPs for Evaluating LLM Agents" (2025). https://arxiv.org/html/2506.08119v2
- arXiv, "SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints" (2025). https://arxiv.org/abs/2503.08669
- arXiv, "SOP-Agent: Empower General Purpose AI Agent with Domain-Specific SOPs" (2025). https://arxiv.org/abs/2501.09316
- arXiv, "Compile, Then Page: Executable SOP Programs and a Capability-Gated Runtime for Procedural LLM Agents" (2026). https://arxiv.org/pdf/2607.11346
- Sierra Research (arXiv), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
- IEEE Standards Association, "IEEE 1849-2023: IEEE Standard for eXtensible Event Stream (XES) for Achieving Interoperability in Event Logs and Event Streams" (2023). https://standards.ieee.org/ieee/1849/10907
- OCEL standard authors (ocel-standard.org), "OCEL (Object-Centric Event Log) 2.0 Specification" (2023). https://arxiv.org/pdf/2403.01975
- 4TU.ResearchData, "Data underlying the paper: Automated Discovery of Process Models from Event Logs: Review and Benchmark". https://data.4tu.nl/datasets/a24f253c-722d-4a3e-9e92-f1a46dbb7473
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Jain et al. (MLCommons), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.