Skip to content

Agent, workflow and domain-reasoning data

Decision records with rationale: what agent builders need from business decisions

Quick answer

Decision rationale data for AI agents is a set of records in which a person made a business decision, each linked to the inputs that person could see at the time, the option chosen, the reason given (a reason code, a note or a cited policy clause) and what happened afterward. Agent teams use it for supervised fine-tuning, reward modeling and evaluation. Its value depends on three things suppliers rarely check: hindsight leakage, outcomes observed for only one branch, and inconsistent deciders.

By SourceX Editorial · Updated

This page covers record design for decisions in general. Domain versions have their own pages: underwriting referrals and exceptions, claims adjudication with coverage reasoning, approval and rejection records and exception handling records. The agent training data hub maps the cluster, and SourceX's page on product decision records covers specs, debates and launch reviews from product teams.

Six parts of a decision record, and the two suppliers usually lack

A usable decision record has six parts: the decision point, an input snapshot as of decision time, the decision, the stated rationale, the decider's context and a linked outcome. Suppliers usually hold the decision and some rationale. The snapshot and the outcome link are what buyers most often have to ask for.

PartWhat it holdsWhere it usually livesTypical gap
Decision pointCase ID, decision type, every option available (approve, approve with conditions, deny, request information, escalate), deadlineCase, ticket or approval workflow in a CRM, ITSM or ERP systemOnly the chosen option is stored
Input snapshotField values, documents and scores the decider could seeChange logs, audit tables, document stores with received timestampsRebuilt from current-state tables, so later values leak in
DecisionChosen option and its parameters (amount, terms, conditions), timestampTransaction or approval recordParameters overwritten by later edits
RationaleReason codes, free-text note, policy clause citedCode fields, work notes, approval comments, memo attachments"Other" codes, boilerplate, notes written after the fact
Decider contextRole, authority limit, queue, stable pseudonymRole tables, authority matrices, assignment historyReal names in the clear, or no ID at all
OutcomeReversal, appeal, reopen, dispute, business result at a fixed horizonDownstream systems such as collections, claims, CRM or ticketingMeasured too early, or only for one branch

A useful mental model comes from a US patent document that describes rationale data structures linking the sequence of human-in-the-loop decisions behind an outcome, so a reviewer can navigate backward in time from the outcome [1]. Every outcome in a licensed dataset should be traversable the same way. If you need the full step-by-step trajectory around a decision rather than the decision itself, see enterprise workflow task histories and reconstructing trajectories from ticket histories.

Rebuilding what the decider could see, not what the system knows now

The input snapshot must show each record as it stood when the decision was made. Joining decisions to today's tables leaks hindsight: a later default, a corrected address or a revised risk tier lets a model "predict" what the decider could not have known.

Most operational systems hold current state and keep history only in change logs. SAP, for example, records changes as change documents: header rows in CDHDR and item rows in CDPOS, whose VALUE_OLD and VALUE_NEW fields hold the before and after values, according to a third-party table reference [2]. ServiceNow keeps work notes and comments as rows in a separate journal table, sys_journal_field, keyed to the parent record, per third-party connector documentation [3]. A supplier who rebuilds as-of values from such logs can deliver a true snapshot; one who exports current tables cannot.

A related failure is documented in a neighboring field. Researchers in predictive process monitoring note that training and test sets built from event logs are often not completely separated, a data leakage problem particular to that field [4]. Point-in-time correct training data explains as-of joins, and target leakage audits show how to test a sample for it. Ask the supplier for:

  • An as_of timestamp on every input field equal to the decision time, with its source (change log, nightly snapshot or document receipt).
  • A visible_to_decider flag. A risk score computed in the background but never shown on screen was not an input to the human's judgment, and a model trained to imitate that human should know it.
  • Document received and opened timestamps, so an attachment that arrived after the decision is excluded.
  • Late corrections tagged as corrections, not merged into the snapshot.

Reason codes, notes and cited policy: three forms of rationale

Rationale comes in three forms that fail differently. Reason codes are consistent but coarse, free-text notes are rich but often written for defensibility after the fact, and policy citations tie a decision to a rule only when the policy text and version come with them.

FormExampleStrengthTypical failureBest training use
Reason codeHold-release code, denial code, disposition code from a fixed listCountable and comparable across deciders"Other" overused; first item in the drop-down over-selected; list revised mid-periodClass labels, stratified sampling, evaluation slices
Free-text noteWork note, approval comment, credit memoRecords the tradeoff and compensating factorsCanned text; written after the outcome; personal data typed inExplanation targets for fine-tuning; rationale-aware reward models
Cited policy"Credit policy 4.2, version 2025-03"Links judgment to a rule and exposes exceptionsPolicy corpus missing or the wrong version deliveredPolicy-following training and evaluation

Some domains standardize codes. Health-plan claim adjustments carry Claim Adjustment Reason Codes (CARCs) and Remittance Advice Remark Codes (RARCs) in the X12 835 remittance, and CAQH CORE publishes an operating rule for their uniform use [5]. Most business decisions use internal code lists, so ask for the code dictionary with effective dates.

Rationale is what makes decision data teach judgment rather than imitation. Research on policy-following language models reports that they apply policies rigidly even when that is impractical, and that supervised fine-tuning on human decisions with explanations worked better than ethical-framework or chain-of-thought prompting [6]. Stated reasons are still justifications, not a view into cognition: a study of automated rationale generation trained on human think-aloud data cautions that its rationales do not necessarily reveal the true decision process [7].

Human notes share that limit, so test them. Compare note timestamps with decision timestamps, and measure how often a note cites a fact present in the input snapshot. Templates and boilerplate in business records covers detecting canned notes.

Linking outcomes without inheriting selective labels

An outcome turns a decision record into a label, but outcomes are usually observed only for the branch the decider chose. Outcome-linked data therefore describes the deciders' policy as much as the truth, and buyers need both per-branch outcomes and decider variation to correct for it.

Lakkaraju and colleagues call this the selective labels problem: in their bail example, whether a defendant fails to appear is observed only if the judge released them [8]. Business decisions have the same shape. Repayment is observed only for approved credit, confirmed fraud only for investigated cases, and claim leakage only for paid claims.

Their "contraction" method compares models with human deciders by exploiting variation in leniency among decision-makers who are assigned comparable cases [8]. De-Arteaga and colleagues instead use consistency among experts as a signal when learning under selective labels [9]. Methods like these depend on decider-level data, so ask for a stable pseudonymous decider ID, enough cases per decider and a description of how cases were routed to deciders.

Specify which outcomes count and when they are read:

  • Reversal: supervisor override, second-level review or upheld appeal, with the original and corrected decisions linked by ID. Human override and correction logs covers this record type.
  • Contest: appeal, dispute, complaint or chargeback filed, even if not upheld.
  • Rework: case reopened or decision re-made. See error-recovery records.
  • Business result at a horizon: paid within terms at 90 days, churned within 12 months, no recontact within 7 days. State the horizon and flag cases still open as censored.

Verifying outcome labels describes the checks, and historical decision bias in operational labels covers what selection does to fairness.

Who decided: role, authority and consistency across deciders

Record each decider's role, authority level and a stable pseudonym, not their identity. Then measure how consistently different deciders treat comparable cases before you treat any single decision as a training label.

The variation can be large. In an insurance-company noise audit described by Kahneman and colleagues, the median difference between underwriters' prices for identical policies was 55 percent, and between claims adjusters' payouts for identical claims 43 percent, according to strategy+business [10]. When deciders in one firm disagree that much, a decision is one professional's opinion under a policy, not ground truth.

  • Double-reviewed cases. Quality reviews, second-level approvals and calibration sessions give direct agreement data. Krippendorff's alpha measures agreement among judges assigning values to the same units, where 1 is perfect reliability and 0 is agreement no better than chance [11]. See choosing an inter-annotator agreement metric.
  • Decider-level rates. Approval rate, override rate and note length per pseudonym show whether one prolific decider would define the agent's judgment.
  • Authority context. Decisions near an authority limit behave differently from routine ones, so record the limit in force on the decision date.
  • Disagreement handling. Condition on policy version and role, use high-agreement subsets as fine-tuning targets, and keep disagreement as soft labels or preference signal; resolving annotator disagreement compares methods.

Illustrative record: releasing a sales order held for credit

The record below shows all six parts for one order-to-cash decision: a credit analyst deciding whether to release an order held because it would push the customer over its credit limit.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "decision_id": "DEC-0193877",
  "decision_type": "credit_hold_release",
  "case_ref": "SO-4471902",
  "decided_at": "2025-06-12T14:22:05Z",
  "options_available": ["release", "partial_release", "keep_hold", "require_prepayment", "escalate"],
  "inputs_as_of_decision": {
    "as_of": "2025-06-12T14:22:05Z",
    "source": "change_log_reconstruction",
    "fields": {
      "order_value_usd": 48200,
      "credit_limit_usd": 150000,
      "open_ar_usd": 131500,
      "ar_over_60_days_usd": 12400,
      "avg_days_to_pay_12m": 41,
      "internal_risk_score": 0.31
    },
    "visible_to_decider": {"internal_risk_score": false},
    "documents": [{"type": "remittance_advice", "received_at": "2025-06-12T09:10:00Z", "opened_by_decider": true}]
  },
  "decision": {"option": "release", "conditions": ["hold next order if 60+ day balance not cleared by 2025-06-19"]},
  "rationale": {
    "reason_codes": ["PAYMENT_IN_TRANSIT"],
    "code_list_version": "2025-01",
    "note": "Remittance advice for ~$14k received today covers the 60+ balance. Release; re-hold next order if not posted by 6/19.",
    "note_written_at": "2025-06-12T14:20:41Z",
    "policy_ref": {"doc": "credit_policy", "section": "4.2", "version": "2025-03"}
  },
  "decider": {"pseudonym": "AN-17", "role": "credit_analyst_2", "authority_limit_usd": 50000},
  "outcome": {"horizon_days": 90, "paid_within_terms": true, "days_to_pay": 38, "disputed": false, "censored": false},
  "related_decisions": [],
  "privacy": {"customer_name": "removed", "account_number": "replaced_with_token", "method_ref": "DM-2"}
}

Four details carry most of the value. The note was written less than two minutes before the decision, so it is contemporaneous, and the risk score existed but was hidden, so an imitation model should not use it. The order sits just under the analyst's authority limit, which explains why no escalation occurred. The outcome names its horizon, so a 30-day reading cannot be confused with a 90-day one.

Matching records to fine-tuning, reward modeling and evaluation

The same decision records serve three training uses, but each needs different parts and different filters. Decide the use before you sample, because the filter for one can destroy the signal for another.

UseModel inputTarget or signalFilterWatch for
Supervised fine-tuningSnapshot fields marked visible to the deciderDecision plus rationale textNot reversed; high-agreement deciders; current policy versionImitating noise; post-decision fields leaking in
Reward modelingSnapshot plus candidate decisionsCorrected over reversed decision; better over worse outcome on comparable casesPairs drawn from comparable cases with outcomes on both branchesSelective labels turning past approval policy into reward
EvaluationSnapshotDecision, matured outcome, or rubric score on the rationaleDecisions after the training cutoff; held-out decidersHuman agreement ceiling; label errors

Treat operational outcomes as noisy labels in evaluation too. Northcutt and colleagues estimate an average label error rate of at least 3.3% across the test sets of 10 widely used datasets and show that such errors can change model rankings [12]. For test-set design, see outcome-labeled evaluation data and task success labels for agent trajectories; for preference pairs, see buying RLHF comparison data.

Rights, privacy and regulatory questions specific to decision data

Decision records carry three kinds of sensitive content at once: personal data about the customer or applicant, employee-written notes, and, for decisions about people, regulatory duties that can attach to the model you train. Settle each before licensing. As of October 2026:

  • Personal data in notes. Names, account numbers and health details are often typed into free text, so ask how notes were scanned and what residual rate a sample showed. SourceX removes or replaces personal details such as names, emails, phone numbers and account numbers before delivery, records the method for each dataset and checks a sample after processing; no de-identification method is perfect.
  • Decider identity. Ask for consistent pseudonyms (the same analyst always maps to the same token) so consistency can be measured, and ask who holds the key. Notes are employee-authored work product; see employee-authored records in training data.
  • Colorado. SB26-189, signed 14 May 2026, covers automated decision-making technology that materially influences consequential decisions in areas including employment, lending, insurance, housing and health care. From 1 January 2027, developers must give deployers documentation that includes training data categories [13]. The Attorney General released interim draft rules on 6 October 2026, with comments due 26 October [14]. See Colorado SB 26-189 training data documentation.
  • EU. If an agent trained on the records will be a high-risk AI system under the AI Act, Article 10 requires governance of training, validation and testing data, including examination of possible biases and identification of data gaps [15]. As amended by Regulation (EU) 2026/1744, Annex III high-risk obligations reportedly apply from 2 December 2027 [16].
  • License scope. Name training on rationale text, deriving reward models and building internal benchmarks as permitted uses, and ask whether outcome refreshes are delivered as outcomes mature. See license terms for agent data.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request checklist and red flags for a decision-record sample

A precise request names the decision type, the six record parts and the checks you will run on a sample, because those choices decide which suppliers can meet it. The agent data specification guide covers task and outcome fields in more depth.

  • Decision and options: the decision type, every option available, and minimum counts per option, including holds, declines and escalations.
  • Period and versions: the decision window, plus policy documents and code lists with effective dates.
  • Snapshot method: as-of reconstruction from change logs, visible_to_decider flags and document timestamps.
  • Rationale fields: the code dictionary, notes, policy references, the share of "Other" codes and the share of notes written after the decision.
  • Decider fields: role, authority limit, stable pseudonym, minimum cases per decider and any double-reviewed subset.
  • Outcomes: types, horizon, censoring flags, availability per branch and links between original and overriding decisions.
  • Handling and format: personal-data removal method, free-text scan results, one JSONL record per decision, and the policy corpus as versioned documents.

Reject or renegotiate a sample when notes are timestamped after case closure, one code covers most decisions, outcomes exist only for approvals, old decisions show input values identical to today's records, or employee and customer names appear in note text.

Decision records like these are a kind of data SourceX sources on request from US companies, not inventory under contract, and a request does not guarantee a match. After you send SourceX a decision-record specification, any release of matching records needs the supplying company's approval. Each dataset goes through rights review and is delivered under a license that defines the records included, permitted uses, term and delivery.

Building agents that need real decision histories?

Describe the decision types, record parts and outcome horizons you need. SourceX looks for US companies that hold matching records, checks each supplier's licensing permissions, and manages the license and delivery through private, access-controlled workflows. Specify your decision-record dataset.

Sources

  1. United States Patent and Trademark Office, "Machine learning traceback-enabled decision rationales as models for explainability". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/12530616
  2. LeanX (third-party SAP table reference), "SAP table CDPOS (change document items) reference". https://leanx.eu/sap/table/cdpos/
  3. CData Software, "CData Python Connector for ServiceNow: sys_journal_field table". https://cdn.cdata.com/help/BNM/py/pg_table-systemjournalfield.htm
  4. Weytjens and De Weerdt, arXiv, "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
  5. CAQH CORE, "CARCs and RARCs 835 Rule (Phase III CORE operating rule)". https://www.caqh.org/sites/default/files/core/phase-iii/policy-rules/CARCsRARCs_835_Rule.pdf
  6. arXiv, "Teaching AI to Handle Exceptions: Supervised Fine-tuning with Human-aligned Judgment" (2025). https://arxiv.org/html/2503.02976v2
  7. Ehsan et al., arXiv, "Automated Rationale Generation: A Technique for Explainable AI and its Effects on Human Perceptions" (2019). https://arxiv.org/pdf/1901.03729
  8. Lakkaraju, Kleinberg, Leskovec, Ludwig and Mullainathan, KDD, "The Selective Labels Problem: Evaluating Algorithmic Predictions in the Presence of Unobservables" (2017). https://www.cs.cornell.edu/home/kleinber/kdd17-selective.pdf
  9. De-Arteaga, Dubrawski and Chouldechova, arXiv, "Learning under selective labels in the presence of expert consistency" (2018). https://arxiv.org/pdf/1807.00905
  10. strategy+business, "How noisy is your company". https://strategy-business.com/article/How-noisy-is-your-company
  11. Klaus Krippendorff, University of Pennsylvania, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  12. Northcutt, Athalye and Mueller, arXiv, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  13. Colorado General Assembly, "SB26-189 Automated Decision-Making Technology" (2026). https://leg.colorado.gov/bills/sb26-189
  14. Colorado Attorney General, "Colorado Automated Decision-Making Technology & Chatbot Safety Rulemaking" (2026). https://coag.gov/ai/
  15. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  16. European Parliament and Council, "Regulation (EU) 2026/1744 (Digital Omnibus on AI)" (2026). https://eur-lex.europa.eu/eli/reg/2026/1744/oj?locale=en

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data