Agent, workflow and domain-reasoning data
Interaction data from legacy desktop, terminal and virtual-desktop applications
Quick answer
Agents that operate mainframe green screens, thick-client Windows apps and remote desktops need interaction data captured at the level those systems actually expose: a character grid and field map for 3270/5250 terminals, an accessibility tree where a thick client provides one, and pixels plus input events for Citrix or RDP sessions. The most useful datasets pair each observation with keyboard-first actions (function keys, field tabs, screen submits), system-ready timing and a semantic step name from the application's transaction catalog.
By SourceX Editorial · Updated
Why legacy-system trajectories need their own data spec
Legacy-system trajectories differ from browser or generic desktop data because the observation and action spaces are different, not just older. A browser agent reads a DOM; a 3270 agent reads a 24x80 (or larger) character buffer divided into protected and unprotected fields, and it moves through the system by pressing an attention key such as Enter, PF3 or Clear rather than clicking links. TN3270E, the Telnet-based protocol most emulators use today, carries that data stream and adds functions such as ATTN, SYSREQ and SNA response handling that an agent may need to reproduce.
Thick clients built on Win32, WinForms, Delphi, PowerBuilder or Java Swing sit in the middle. Some expose a usable tree through Microsoft UI Automation, with control types and patterns such as Invoke and Scroll that a recorder can log alongside each action. Many custom-drawn grids expose little or nothing, and once the app is published through a virtual desktop, the agent and the recorder see only a bitmap stream.
If your general step-record schema is already settled, start from computer-use trajectory data: what every step record must contain and treat this page as the legacy-specific extension. For the wider category map, see the agent training data hub.
Three capture surfaces and what each one gives you
Choose the capture surface by what the target system exposes, because it determines label precision, storage cost and de-identification effort. Terminal emulators can log the screen buffer and keystrokes as text, which is typically far cheaper to store and more precise to label than video. Remote desktops force a pixel-and-coordinates approach with all of its grounding ambiguity.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Surface | Typical systems | Observation you can log | Action space | Main failure mode in data |
|---|---|---|---|---|
| Block-mode terminal (3270/5250) | CICS and IMS transactions, IBM i (AS/400) apps | Screen buffer text, field attributes (protected, numeric, hidden), cursor position | Typed text, Tab/Backtab, Field Exit, PF1-PF24, PA1-PA3, Enter, Clear | Missing keyboard-lock state, so agents learn to type before the host is ready |
| Character terminal (VT100/VT220) | Unix line-of-business apps, telnet/SSH menus | Terminal output stream or rendered grid | Keystrokes, control and escape sequences | Escape-sequence noise; screen state must be reconstructed by an emulator |
| Thick client, local | Win32/.NET/Java desktop apps | Screenshot plus UI Automation tree where exposed | Clicks, keys, shortcuts, menu accelerators | Owner-drawn controls with no tree; coordinates tied to one resolution and DPI |
| Virtual desktop (Citrix, RDP, VDI) | Any app published remotely | Pixels only, plus client-side input events | Mouse at coordinates, keys | Compression artifacts, lag between input and repaint, scaling mismatch |
Practitioners have recorded legacy host sessions for decades. One patent models a legacy host as a finite state machine built from recorded session traces enriched with usage statistics [2], and another describes capturing low-level keyboard and mouse events from legacy applications without modifying them [3]. That prior art is a useful reminder that a screen-to-screen transition graph is often recoverable from terminal logs and can serve as both a training signal and an evaluation oracle.
Designing the action space around keys, fields and screen submits
The action vocabulary must include function keys, field navigation and submit keys as first-class actions, not as generic "key press" strings. A 3270 transaction is a sequence of edits to unprotected fields followed by one attention-identifier key that sends the modified fields to the host; the host then returns a new screen. If the dataset flattens this into "type text" and "press key", models cannot learn which edits were committed together.
Ask suppliers to log, for each step: the AID key pressed, the list of modified fields with row, column and length, and the cursor address at submit. For thick clients, log the accelerator or menu path when one was used, because keyboard-heavy power users rarely click. For virtual desktops, record the client resolution, DPI scaling and any coordinate transform so that pixel actions can be replayed or normalized.
Wait states, refresh timing and the "system ready" signal
Capturing wait states is what teaches an agent when the system is ready, and it is the field most often missing from commodity recordings. On a 3270 host the keyboard is locked (the "X SYSTEM" or input-inhibited indicator) between submit and the next screen; on a remote desktop the equivalent is the interval before the expected region repaints. Agents trained without this signal tend to fire inputs into a locked keyboard or a half-drawn screen.
Request three timestamps per step: action sent, first screen update received, and keyboard unlocked or screen stable. Keep sessions that include host errors, timeouts and "transaction not available" messages rather than filtering them out. Those are the long-tail states covered on our page on exception handling records, and they are also the states a regression suite most needs.
Labeling trajectories with the transaction catalog
Ask the supplier for its screen or transaction catalog so every step can carry a semantic name instead of only raw coordinates or keystrokes. Mainframe shops usually know their screens by transaction code and map name (for example, a CICS transaction ID and BMS map), and IBM i shops by program and display file. Joining those identifiers to the trajectory turns "PF8 on screen 14" into "page forward in the policy coverage list".
Where the business already runs RPA, the bot's run history is a second labeling source. As of October 2026, Power Automate desktop flows, for instance, store run action logs in Dataverse, in the flowsession table or a newer FlowLogs table depending on the environment's log version [1]. Bot logs describe the scripted happy path, so treat them as a step-name dictionary and a source of exceptions, not as human demonstrations. For process-level context, combine them with process mining event logs.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"episode_id": "ep-000213",
"surface": "tn3270e",
"terminal_model": "IBM-3278-2-E",
"task_label": "update policy mailing address",
"step": 4,
"screen_id": "POLM02",
"transaction_code": "PMNT",
"observation": {
"buffer_text_ref": "screens/ep-000213/004.txt",
"fields": [
{"row": 6, "col": 20, "len": 30, "protected": false, "name": "ADDR_LINE_1", "value": "[STREET_1]"},
{"row": 8, "col": 20, "len": 5, "protected": false, "numeric": true, "name": "ZIP", "value": "[ZIP_5]"}
],
"cursor": {"row": 6, "col": 20}
},
"action": {"modified_fields": ["ADDR_LINE_1", "ZIP"], "aid_key": "ENTER"},
"timing_ms": {"sent": 0, "first_update": 412, "keyboard_unlocked": 655},
"result_screen_id": "POLM03",
"host_message": "RECORD UPDATED",
"deidentification": {"method": "field-level token replacement", "version": "v2"}
}
De-identifying screen content before it leaves the supplier
Screen content in these systems routinely contains regulated records, so de-identification of buffers, screenshots and typed values is a precondition of delivery, not a cleanup step. Terminal buffers are easier than video: when the field map is known, a supplier can replace values by field name with consistent tokens, as in the record above. Pixel streams need OCR-based detection and redaction, which misses text in unusual fonts, low-contrast fields and partially repainted regions.
Health plan and provider systems add HIPAA. De-identification there means Expert Determination or Safe Harbor's removal of 18 identifier types, and HHS notes neither method removes all re-identification risk [6]. Keystroke logs are a common leak: a redacted screenshot is useless if the raw typed value of a member ID survives in the action stream. Our guide to PII in screen recordings and computer-use trajectories covers redaction methods in more depth, and collecting computer-use demonstrations at work covers capture consent and monitoring law.
Using legacy trajectories for evaluation, not only training
Legacy interaction data is often more valuable as an evaluation set than as bulk training data, because outcomes on these systems are checkable. A terminal session ends on a known screen with a host message and changed field values, which supports execution-based scoring of the kind OSWorld uses for real desktop environments [4]. As of October 2026, public benchmarks rarely include green screens or published Citrix apps, so a held-out set of real legacy tasks closes a gap rather than duplicating one.
Small, well-labeled sets still matter for training. One recent study started from 312 human-annotated computer-use trajectories and expanded each step with model-generated alternative actions [5]. A few hundred carefully captured legacy episodes with clean field maps and timing can therefore seed augmentation, provided the license permits derived data.
Supplier request checklist for legacy interaction data
Use a written request so suppliers can say quickly whether they hold the data in a usable form.
Illustrative example: invented to show structure; it does not describe an available dataset.
- Systems and surfaces: host type (z/OS CICS, IMS, IBM i, Unix), emulator in use, and whether apps are reached locally or through Citrix, RDP or another VDI.
- Observation format: buffer text with field attributes, UI Automation tree snapshots, or pixels; frame rate and resolution for video.
- Action log: AID or function keys, modified fields, cursor address, accelerators, coordinates with DPI and scaling.
- Timing: sent, first update and ready timestamps per step; network latency notes for remote sessions.
- Labels: transaction codes, screen or map names, task labels, success and error outcomes.
- Scope: human operator sessions versus RPA runs, date range, number of distinct screens, share of exception paths.
- Privacy: fields that hold personal or regulated data, de-identification method and version, sample review results.
- Rights: confirmation that the business owns the session data and that operator monitoring notices covered capture.
For the business-process side of the same work, see document-to-system entry pairs and the distinction from task mining data, which covers broad desktop activity logs rather than legacy-specific capture.
How SourceX sources legacy-system interaction data
SourceX sources operational datasets from US companies on request; it does not hold this data in stock, and a request does not guarantee a match. You describe the systems, surfaces and labels you need, and SourceX looks for US businesses that hold that data, with every release approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and health records require HIPAA de-identification. Related SourceX pages cover training data for computer-use agents, workflow and screen activity data and screen recordings. You can submit a buyer request to SourceX with the checklist above.
Request legacy application interaction data for computer-use agents
SourceX finds US companies that hold the described data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. Describe the legacy-system interaction data you need.
Sources
- Microsoft Learn (Power Automate), "Desktop flow action logs configuration". https://learn.microsoft.com/en-US/power-automate/desktop-flows/configure-desktop-flow-logs
- USPTO, "Modeling interactions with a computer system (US 9047269)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/9047269
- USPTO, "Application instrumentation and monitoring (US 7496575)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/7496575
- Xie et al., arXiv / NeurIPS 2024, "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- He, Jin and Liu, arXiv 2505.13909, "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/pdf/2505.13909
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.