Skip to content

Agent, workflow and domain-reasoning data

RPA bot definitions and run logs as agent training and evaluation data

Quick answer

RPA estates hold three assets that agent teams can use: bot definitions (scripted step sequences and UI selectors for named applications), run histories (job and transaction outcomes per item), and exceptions that bots handed to people. Together they supply step-level demonstrations, labeled hard cases and a measured bot baseline for the same tasks. Request all three with stable IDs that join them, credentials stripped, and written confirmation that the platform license and the company's contracts permit export for model training.

By SourceX Editorial · Updated

What an RPA estate contains that agents can learn from

An RPA estate is a structured record of how a company automated specific screen tasks and where that automation broke. Agent builders replacing brittle bots usually want it for two jobs: supervised fine-tuning on the intended step sequence, and evaluation against the bot that currently does the work. Unlike general task mining data from desktop interaction logs, RPA artifacts are already scripted, so the intent of each step is explicit rather than inferred.

The three layers differ by platform but map cleanly:

  • Definitions. UiPath projects are XAML workflow files plus a project.json manifest; Blue Prism stores processes and business objects as XML exported in release packages; Power Automate Desktop flows are written in its own scripting language and stored in Dataverse. Each contains activity order, branching, retry logic, and selectors that target specific fields in SAP GUI, Oracle EBS, mainframe emulators, Citrix sessions or web apps.
  • Run histories. UiPath Orchestrator exposes job, queue item and robot log events, and each transaction carries its queue item reference, status, timestamps, exception details and the robot that ran it. Power Automate keeps desktop flow runs in the Dataverse flowsession table with start, duration, status, machine, robot account and parent flow context.
  • Exceptions and handoffs. Failed queue items record a processing exception type (application or business exception) with a reason and details text. Items routed to a person, plus what that person did next, are the most valuable records in the estate.

Why bot exceptions are the hard cases agents need

Exceptions that a bot handed to a person are pre-labeled failure cases, which is exactly what synthetic data and public benchmarks lack. RPA practice separates business exceptions (the input violates a rule, such as an invoice without a PO number) from application exceptions (the screen, selector or system failed), and UiPath monitoring reports the two separately while accounting for retries. That split gives you two different training signals.

Business exceptions teach policy reasoning: when to stop, escalate, or request missing data. Application exceptions teach robustness: what happens when a modal dialog appears, a field moves, or a session times out. A patent filing on human-in-the-loop RPA training describes logging people's interactions with robots and centralizing them for training [3], which is the same pattern a buyer wants: the bot's last state, the person's resolution, and the final outcome on one record.

The failure mode to avoid is receiving exception counts without resolution data. A row reading "BusinessException: Vendor not found" is a label; it becomes a training example only when joined to what the person did in the vendor master and whether the item then completed. For the handoff side of this design, see agent-to-human handoff data.

Using bot success rates as an evaluation baseline

A bot's historical success rate on a defined task is the most direct baseline an agent has to beat in an RPA migration. If a queue processed thousands of claims with a known split of successes, business exceptions and application exceptions, an agent replaying the same items in a sandbox can be scored on the same three outcomes. This mirrors how execution-based benchmarks such as OSWorld grade agents on final system state in real desktop and web applications [6].

Baselines need care to be fair. Bots only run on items that passed upstream filters, so the agent must face the same input distribution, and retries inflate apparent success unless you count first-attempt outcomes. Pair the bot baseline with a human baseline under matched conditions, as HCAST does for software tasks [7], and see human baseline data for agent evaluation for time, cost and quality fields. For grading design, see agent evaluation task suites.

Rights, platform terms and credential hygiene

Before you design a pipeline, confirm the supplier may export and license these artifacts for training. Definitions built in-house are usually the company's work product, but implementation partners, vendor templates and marketplace components can carry their own terms, and RPA platform agreements may restrict how exported definitions and logs are used. Ask counsel on both sides to read the platform subscription, any systems integrator statement of work, and the customer contracts covering the data the bots touched. Our guide to license terms for agent workflow data covers replay, derived tasks and benchmark rights.

Credentials are the most common contamination. Bots read secrets from Orchestrator assets, credential vaults or connection references, but older projects often hardcode usernames, passwords, API keys or file share paths in arguments, config files and log messages. Run a secrets scanner over definitions and logs, rotate anything found regardless, and treat robot log message text as free text that can contain customer names, account numbers and claim details.

Retention determines how much history exists

How much run history you can buy is set by each platform's logging and retention configuration, not by how long the bots have run. Power Automate desktop flow action logs can be fully enabled, captured only on run failure, or disabled; V1 logs sit in the flowsession AdditionalContext field, and, as of October 2026, Microsoft documents a newer FlowLogs table with built-in time-to-live retention. Microsoft documents bulk deletion of historical run data to manage capacity, so many estates hold only months of step-level detail even when job summaries go back years.

Ask early which log level was configured (UiPath distinguishes levels from verbose to critical), when it changed, and whether logs were forwarded to Elasticsearch, Splunk or a data lake. A forwarded store often holds more history than the platform itself.

Normalizing RPA artifacts into agent trajectories

The practical target is a step record that joins a definition step, the run event it produced, and the outcome. Vendor-native logs are platform-locked, which is why research on UI logging proposes standard formats for ML use [2], and Sapienza research on generating executable RPA scripts shows the reverse mapping from UI logs to executable scripts [1]. If you also need process-level analysis, exporting events to IEEE 1849 XES keeps case, activity and timestamp semantics interoperable with process mining tools [5]; see process mining event logs.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldSource in the RPA estateUse for agents
task_idProcess name plus definition versionGroups runs of one task
definition_versionPackage version or release IDTies behavior to the script in force
item_refQueue item ReferenceJoins bot run, exception and human resolution
step_index, activity_typeWorkflow activity orderAction sequence for SFT
target_app, selectorActivity selector or object elementGrounds actions to UI elements
step_status, timestampRobot log eventTiming and failure point
exception_typeQueue item processing exception typeBusiness vs application label
exception_reasonProcessing exception reason and detailsNatural-language failure cause
human_resolutionTicket, case or audit trail after handoffTarget behavior on hard cases
final_outcomeItem status after human or retryGround truth for evaluation

Two failure modes recur. Selectors drift as applications are upgraded, and a patent filing on updating RPA control code addresses this maintenance problem [4], so always pair a run with the definition version in force that day. Second, the human resolution usually lives outside the RPA platform, in ServiceNow, Jira or the target system's audit log, so plan the join; field-level audit trails explains how to reconstruct those actions.

Request checklist for RPA data

Use this checklist when scoping a purchase or an internal export:

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Platforms and versions in scope (UiPath, Blue Prism, Automation Anywhere, Power Automate) and target applications.
  2. Definition exports for each process, with version history and change dates.
  3. Job and transaction history with first-attempt and post-retry outcomes.
  4. Exception records with type, reason, details text and the item reference.
  5. Human resolution records for handed-off items, joined by item reference.
  6. Log level and retention settings over the requested period, with gaps listed.
  7. Credential and secret scan results, plus the method used to remove personal data from log text.
  8. Written confirmation of export rights under the platform license and the supplier's own customer contracts.
  9. Delivery format (Parquet or JSONL for step records, XES for process mining) and a data dictionary.

Write the request as a specification of tasks, actions and outcomes, following how to write an agent data specification. If screen recordings accompany the logs, apply the redaction guidance for PII in computer-use trajectories.

How SourceX handles requests for RPA and workflow records

SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases; it does not hold RPA logs in stock, and a request does not guarantee a match. Buyers describe the data they need, such as exception histories with human resolutions for a named application, and SourceX looks for US businesses that hold it, with every release approved by the supplying company. You can describe your RPA data requirement to SourceX or browse related workflow and screen activity data and enterprise agent use cases. For the wider cluster, start at the AI agent training data hub.

Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.

Source RPA run logs and exceptions for your agent program

SourceX works with AI teams wherever they are based, moving from finding a supplier through assessing data and licensing permissions to agreeing pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Start an RPA data request at SourceX for buyers.

Sources

  1. Sapienza University of Rome, "Automated Generation of Executable RPA Scripts from User Interface Logs". https://research.uniroma1.it/node/48233
  2. Universidad de Sevilla (idUS), "Towards an OpenSource Logger for the Analysis of RPA Projects". https://idus.us.es/handle/11441/134238
  3. USPTO, "Human-in-the-loop robot training for robotic process automation". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/11815880
  4. USPTO, "Intelligent control code update for robotic process automation". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/10710239
  5. IEEE Standards Association, "IEEE 1849-2023 Standard for eXtensible Event Stream (XES)" (2023). https://standards.ieee.org/ieee/1849/10907
  6. Xie et al. (XLANG Lab, HKU), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  7. METR, "HCAST: Human-Calibrated Autonomy Software Tasks" (2025). https://arxiv.org/pdf/2503.17354

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data