Skip to content

Agent, workflow and domain-reasoning data

Field-level audit trails: reconstructing agent actions from business systems

Quick answer

Field-level audit trails record who changed which field on which record, from what value to what value, and when. That is enough to rebuild state transitions and API-level action sequences for agent training, even when nobody captured a screen. It is not enough for GUI agents, because clicks, navigation and reading time never reach the log. The practical work is confirming field coverage, separating human from system actors, grouping field edits into actions, and naming those actions consistently.

By SourceX Editorial · Updated

What a field-level change log actually records

A field-level change log records state deltas on business objects, not user intent or interface steps. In SAP ERP, a change document has a header in CDHDR (object class, object ID, change number, user, date, time, transaction code) and items in CDPOS that hold the table, field, change flag, old value and new value [1]. Salesforce field history writes one row per tracked field change to objects such as OpportunityFieldHistory, CaseHistory or Custom_Object__History, with Field, OldValue, NewValue, CreatedById and CreatedDate. Database change data capture tools such as Debezium emit before and after row images per committed change, from which the changed fields can be derived [2].

These sources share one shape: object, field, old value, new value, actor, timestamp and sometimes a transaction or session identifier. Ticketing platforms add a twist. ServiceNow stores comments and work notes as entries in sys_journal_field rather than on the incident row, so free text and field changes arrive from different tables [5]. For ticket-specific reconstruction, see reconstructing trajectories from ticket and case histories.

Which agents can learn from before-and-after values

Before-and-after values train agents that act through APIs or record updates, and they mostly fail for agents that act through screens. An agent that calls update_purchase_order(po, {payment_terms, delivery_date}) needs exactly what a change log holds: the prior state, the fields touched and the resulting state. A computer-use agent needs observations and low-level actions (screenshots, accessibility trees, clicks, keystrokes) that change logs never contain; computer-use trajectory records cover that format.

Reconstructed sequences fit three uses well:

Change logs cannot tell you why a field changed. Pair them with notes, approvals or decision records with rationale when the reasoning matters.

Coverage gaps to confirm before anything else

Coverage is the first failure point, because most systems log only configured fields for a limited period. Standard Salesforce field history tracks a capped number of fields per object for a limited retention window, and the paid Field Audit Trail add-on extends both through the FieldHistoryArchive big object; confirm the current limits in the holder's org. SAP logs only fields flagged as change-document-relevant in the ABAP Dictionary, so a custom field without that flag is invisible [1]. Tracking also starts when it is enabled; earlier history does not exist.

Ask the data holder for a field coverage map before discussing volume:

  • Which objects and fields are tracked, and since what date for each.
  • Whether long-text, rich-text and encrypted fields log values or only the fact of a change (some platforms log certain field types without old and new values).
  • Whether archiving or purging jobs have removed history for older records.
  • Whether record creation and deletion are logged (SAP uses change flags for insert, update and delete [1]).
  • Whether child objects, such as line items, have their own history.

A log that covers Status but not Assigned_To produces trajectories with silent gaps. Score coverage per action type, not per system.

Separating human, integration and batch actors

Many log rows are written by integrations, scheduled jobs and workflow rules, and mixing them with human edits teaches an agent to imitate plumbing. Typical signals are technical user IDs (integration or batch users), SAP transaction codes for background jobs, Salesforce rows created by an automated process user, and bursts of identical changes across thousands of records within seconds. CDC streams add another layer, because they capture every committed write regardless of the application path [2].

Tag each row with an actor class: human, rule-driven automation, integration, batch migration. Keep the automation rows, because they explain state the human then reacted to, but exclude them as training targets. If an RPA bot did the work, its run logs are a richer source than the field log; see RPA bot definitions and run logs.

Grouping field edits into one action and naming it

One user action usually writes several fields, so reconstruction groups rows into actions using shared transaction identifiers first and timestamps second. In SAP, rows sharing a CDHDR change number belong to one save [1]; Salesforce history rows written in the same save typically share a CreatedDate and user; CDC events can be grouped by the source transaction ID carried in event metadata. Where no identifier exists, use a short window (for example, same user, same record, within a few seconds) and validate the window on a hand-labeled sample.

Naming is the second step. Map transaction or screen codes plus the set of changed fields to semantic action names: an SAP change on a purchase order where only the delivery date changed becomes reschedule_po_line; a Case update that sets Status=Escalated and changes OwnerId becomes escalate_case. Keep the mapping table versioned, because it is the action vocabulary your agent will learn. When several objects change together (order, delivery, invoice), an object-centric representation such as OCEL 2.0, which links one event to many objects and tracks attribute changes over time, preserves that structure better than a flat case log [3]. The public BPI Challenge 2019 purchase-order log, published in OCEL form, shows how ERP activity becomes object-centric process data [4]; OCEL 2.0 for multi-object business processes goes deeper.

Worked example: from change rows to an action record

The example below shows raw rows from a procurement system and the action record a reconstruction pipeline would emit.

Illustrative example: invented to show structure; it does not describe an available dataset.

Raw change rows (CDHDR/CDPOS-style):

change_noobjectobject_idusertcodetimestamptable.fieldflagoldnew
0004417EINKBELEG4500018821USR_P07ME22N2025-03-04 10:12:31EKET.EINDTU2025-03-102025-03-24
0004417EINKBELEG4500018821USR_P07ME22N2025-03-04 10:12:31EKKO.ZTERMUNT30NT45
0004418EINKBELEG4500018821BATCH_EDI(job)2025-03-04 10:13:02EKES.MENGEI500

Reconstructed action record (JSON lines):

{
  "trajectory_id": "po-4500018821",
  "step": 3,
  "actor_class": "human",
  "actor_role": "buyer",
  "action": "reschedule_po_and_change_terms",
  "source_keys": {"change_no": "0004417", "tcode": "ME22N"},
  "args": {"po": "4500018821", "line": 10, "delivery_date": "2025-03-24", "payment_terms": "NT45"},
  "state_before": {"EKET.EINDT": "2025-03-10", "EKKO.ZTERM": "NT30"},
  "state_after": {"EKET.EINDT": "2025-03-24", "EKKO.ZTERM": "NT45"},
  "grouping_method": "change_number",
  "next_event": {"actor_class": "integration", "action": "supplier_confirmation_received"}
}

Note what the record cannot show: the buyer's reason, the screens viewed and any edits made and abandoned before save. The batch EDI row stays as context for step 4 but is not a training target.

Privacy, regulated logs and rights questions

Audit logs are dense with personal data, because every row carries a user ID and often a customer identifier. Treat user IDs as personal data, replace them with stable pseudonyms that preserve role, and check free-text values in CDPOS or journal tables, which often hold names, phone numbers and notes. Healthcare systems must keep audit controls under the HIPAA Security Rule (45 CFR 164.312(b)) [7], and sharing any PHI in those logs generally requires de-identification under 45 CFR 164.514(a)-(b) or, for a limited data set, a data use agreement under 164.514(e) [9].

Rights matter as much as privacy. The holder owns the log but may be bound by vendor terms, customer contracts or employee notices; confirm these during diligence, and set out allowed uses (training, environment replay, benchmarks) in the license, as described in license terms for agent workflow data.

Buyer checklist for audit-trail datasets

Use this checklist when scoping a request or reviewing a sample. It complements the broader agent data specification guide.

Illustrative example: invented to show structure; it does not describe an available dataset.

CheckWhat good looks likeCommon failure
Field coverage mapTracked fields and start dates per objectHistory only on Status
RetentionFull history for the requested period, including archiveRetention window silently truncates older cases
Old and new valuesPresent for every tracked typeLong text logged as "changed" only
Grouping keyChange number or transaction IDTimestamp-only grouping, unvalidated
Actor classificationHuman, automation, integration and batch taggedIntegration users mixed into targets
Action vocabularyVersioned mapping from tcode plus fields to namesRaw field names used as actions
Outcome linkFinal state or business result per objectNo terminal state; see task success labels
Linked contextComments, approvals, documents joined by object IDSeparate exports with no join keys
PseudonymizationStable role-preserving user tokens, method recordedRaw user IDs or names in free text

If a log fails coverage, compare alternatives: commissioned demonstrations, task mining or synthetic trajectories. Synthetic approaches such as Synatra convert indirect knowledge into demonstrations cheaply [8], but they do not reflect a real company's field-level state. See licensed, commissioned or synthetic trajectories.

How this connects to agent logging standards

Field-level reconstruction gives you the "final external action" layer of an agent log, not the planning layer. Agent observability guidance lists planning steps, tool selection, resource access and final external actions as the elements to record for deployed agents [6]. Change logs can supply the last two well; planning and tool choice have to come from elsewhere, or be left for the model to learn from outcomes.

For finished, packaged workflow datasets rather than the reconstruction method, see SourceX's enterprise workflow datasets and agent trajectories and task execution histories. For multi-system joins, see packaging linked records from multiple business systems, and for the wider cluster, the AI agent training data hub.

SourceX sources operational datasets, including engineering, finance and workflow records, from US companies on request; categories are not inventory, and a request does not guarantee a match. Buyers can describe the objects, fields and history depth they need on the SourceX buyer page.

Sourcing field-level change history data

SourceX looks for US businesses that hold the change histories you describe, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, with personal details such as names, emails and account numbers removed or replaced before delivery. Describe the systems, objects and fields you need at sourcex.si/buyers.

Sources

  1. SAP Community (answers.sap.com), "What is the use of CDHDR and CDPOS change document tables?". https://answers.sap.com/questions/1072996/re-what-is-the-use-of-cdhdr-and-cdpos-change-docum.html
  2. Debezium, "Debezium documentation: Event changes transformation". https://debezium.io/documentation/reference/transformations/event-changes.html
  3. Berti, Koren, Adams, van der Aalst et al. (arXiv), "OCEL (Object-Centric Event Log) 2.0 Specification" (2024). https://arxiv.org/pdf/2403.01975
  4. 4TU.ResearchData, "BPI Challenge 2019 (OCEL)". https://data.4tu.nl/datasets/46a7e15b-10c7-4ab2-988d-ee67d8ea515a
  5. ServiceNow, "Journal fields (ServiceNow Documentation)". https://docs.servicenow.com/bundle/vancouver-platform-administration/page/administer/field-administration/concept/c_JournalFields.html
  6. Collibra, "AI audit trails: what to log for models and agents". https://collibra.com/blog/ai-audit-trails-what-to-log-for-models-and-agents-and-how-a-command-center-captures-it
  7. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.312(b) - Standard: Audit controls". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.312
  8. Ou, Xu, Madaan et al. (NeurIPS 2024, arXiv), "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale" (2024). https://arxiv.org/pdf/2409.15637
  9. Electronic Code of Federal Regulations (eCFR), "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data