Agent, workflow and domain-reasoning data
Turning screen recordings into action-labeled agent trajectories
Quick answer
To label actions in screen recordings for agent training, use one of three routes: capture input events concurrently with the video (most accurate, but only for new recordings), have humans annotate frames (accurate but slow), or train an inverse dynamics model (IDM) on a small hand-labeled subset and pseudo-label the rest. A practical plan for a large archive combines routes: a gold hand-labeled set, an IDM for scale, and per-action-type agreement checks before training.
By SourceX Editorial · Updated
Video alone is not a trajectory. A computer-use step record needs an observation, an action with typed arguments, a timestamp and, eventually, a task instruction and outcome; see what every step record must contain. The labeling choice decides which of those fields you can trust.
Three routes from pixels to step-level actions
Concurrent event capture records ground truth, human annotation reconstructs it, and IDM pseudo-labeling infers it; the right mix depends on whether you can still record and how much video you already hold.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Route | How actions are obtained | Typical failure modes | Best fit |
|---|---|---|---|
| Concurrent event capture | OS-level hooks log mouse, keyboard and scroll events with timestamps alongside the video; optionally accessibility-tree snapshots | Clock drift between log and video; IME and remote-desktop input lost; password fields captured unless masked | New commissioned recordings |
| Human frame annotation | Annotators step through video and mark action type, screen coordinates, target element and typed text | Missed fast keystrokes, ambiguous drags, inconsistent step boundaries, fatigue on long sessions | Gold set, evaluation, hard actions |
| IDM pseudo-labels | A model trained on a labeled subset predicts the action between frames t and t+1 using past and future frames | Confuses hover vs. click, misses actions with no visible effect, drifts on unseen applications | Large existing archives |
Event capture is the only route that yields exact keystrokes and coordinates, and some products now pair video with structured logs: as of October 2026, Loom's agent recording mode (in open beta) attaches visited links and console and network logs to narrated screen video [8]. Many archives, though, were recorded for quality assurance or compliance; contact centers, for example, can record agent desktops in Amazon Connect [9], and such footage typically comes without an input-event log.
Why inverse dynamics models make archives usable
An IDM is easier to train than a policy because it sees the effect of an action before predicting it, so a modest labeled set can label a much larger unlabeled one. The model takes frames before and after a candidate step and predicts the action between them; because it is non-causal, it can use the menu that opened or the text that appeared to infer the click or keystroke. Game-playing agents (for example, OpenAI's Video PreTraining work on Minecraft [10]) were among the first published uses of this pattern for labeling large volumes of unlabeled video before training a behavioral prior.
Treat any published labeled-hours figure as a reference point from one domain, not a budget for enterprise software. A game has a fixed action space and a single rendering engine; your archive may span dozens of SaaS tools, legacy terminal emulators, and Citrix sessions with compression artifacts. Expect to label a stratified subset per application family, not one global sample.
Computer-use research shows the same leverage pattern on the policy side. PC Agent-E started from 312 human-annotated computer-use trajectories and used a frontier model to synthesize alternative actions per step [1]. Synatra reports synthetic web demonstrations at about 3% of the cost of human-annotated ones [5], which is a useful comparison when you price human annotation of video.
Capture specifications that drive label accuracy
Frame rate, cursor visibility and timing metadata decide most of your achievable accuracy, so fix them in the spec before anyone records or annotates. The common failure is a 1 fps compliance recording where a double-click, a Ctrl+C and a menu open all fall between two frames.
- Frame rate: set a minimum, and require variable-frame-rate encodes to keep presentation timestamps. Low fps destroys keystroke and drag recovery first.
- Cursor overlay: require the cursor to be rendered into the video, ideally with a click indicator. Some capture tools can hide it, and without it click coordinates become guesses.
- Resolution and DPI scaling: record native resolution and the OS scaling factor; coordinates labeled at 150% scaling are wrong when replayed at 100%.
- Clock source: if events are logged, both streams need a shared monotonic clock and a recorded offset.
- Structured context: where possible, capture window titles, process names and accessibility-tree or DOM snapshots. UGround's authors note accessibility trees are often noisy or incomplete [7], so keep pixels as the primary observation and structure as a hint.
- Sensitive fields: mask password inputs at capture; redaction after the fact is harder, as covered in PII redaction for screen recordings.
Segmenting sessions and writing task instructions
Labeling actions does not give you tasks; segmenting a multi-hour session into episodes and writing the instruction for each is a separate, often larger, cost. A support rep's four-hour recording may contain forty tickets interleaved with chat, email and idle time.
Segmentation cues include application switches, ticket or record IDs visible on screen, form submissions and idle gaps. Instructions can then be written post hoc from what was accomplished; OS-Genesis builds GUI trajectories this way, deriving the task after the interaction rather than before [4]. Post-hoc instructions are cheaper but tend to describe exactly what happened, so they under-represent ambiguous or failed requests.
Outcome labels are a third layer. Linking an episode to the ticket resolution, record change or approval it produced is covered in task success labels for agent trajectories.
Measuring pseudo-label accuracy before you train
Measure agreement per action type on a held-out hand-labeled subset, because a single aggregate accuracy hides the actions agents most often get wrong. Clicks on large buttons are easy; text entry, drag-and-drop, scroll distance and keyboard shortcuts are where IDMs and annotators both fail.
Do not treat human labels as perfect ground truth either. Northcutt and colleagues estimated an average label error rate of at least 3.3% across widely used test sets [6], so double-annotate the gold set and adjudicate disagreements before scoring the IDM against it.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Action type | Metric | Tolerance to define | Report |
|---|---|---|---|
| Left/right/double click | Type match plus coordinate within N px of gold | N px at native resolution | Precision, recall, coordinate error distribution |
| Text entry | Character error rate on typed string | Max CER | CER by application |
| Key combination | Exact chord match (e.g., Ctrl+Shift+T) | Exact | Recall; missed-action rate |
| Scroll | Direction match; magnitude within tolerance | Direction exact | Direction accuracy |
| Drag | Start and end within tolerance | N px | Both endpoints |
| No-op / wait | Correctly predicts no action | Exact | False-action rate |
| Step boundary | Timestamp within K ms | K ms | Boundary F1 |
Stratify the held-out set by application, resolution and annotator so you can see where drift comes from. Then validate downstream: train a small policy on pseudo-labeled data and evaluate in an execution-based environment such as OSWorld, which checks task completion in real operating systems rather than step-matching [3]. Offline step-match metrics in the style of Mind2Web, which uses crowdsourced action sequences on real websites [2], are useful for quick iteration but reward copying one valid path.
Illustrative output record for a pseudo-labeled step
Each derived step should carry its label provenance so you can filter, reweight or relabel later. The schema below extends a standard step record with fields that record how the action was obtained.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"episode_id": "ep_000123",
"step": 14,
"t_video_ms": 418230,
"observation": {"frame_ref": "s3://bucket/ep_000123/f_418230.png", "window_title": "Invoice 4471 - ERP", "resolution": [1920, 1080], "dpi_scale": 1.25},
"action": {"type": "click", "button": "left", "x": 1342, "y": 611},
"label_source": "idm",
"label_model": "idm-v3-erp",
"label_confidence": 0.87,
"gold_checked": false,
"instruction_source": "post_hoc_annotator",
"redaction": {"method": "ocr_mask", "version": "2026-09"}
}
Keep label_source values distinct (event_log, human, idm) so training code can weight them and so any later dispute about where a label came from can be answered.
Rights in derived labels and labeling models
Before labeling, confirm that the license for the recordings covers derived action labels, any IDM trained on them, and reuse of that IDM on other video. These are three different permissions, and a license written for "training models on the recordings" may not clearly cover a labeling model you keep using afterward. On-screen third-party software and customer data raise further questions, covered in rights in third-party content captured on screen and license terms for agent data.
Choosing a route for your archive
Choose by what you can still control: if recording is ongoing, add event capture now; if the archive is closed, budget for a gold set and an IDM. A practical sequence for a closed archive is: audit capture specs, discard unrecoverable footage, double-annotate a stratified gold set, train and score an IDM per action type, then segment and write instructions only for episodes that pass.
If you need more or different footage, the trade-offs between licensed records, commissioned demonstrations and synthetic data are compared in licensed vs. commissioned vs. synthetic trajectories, and the wider cluster is at AI agent training data. For licensing recordings themselves, see license screen recordings for AI training and training data for computer-use agents; what makes screen recordings valuable for AI covers the buyer view. Teams that want new recordings of hands-on work captured to a defined spec can describe the data on the SourceX buyers page.
Source screen recordings and trajectories for agent training
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; a request does not guarantee a match, and every release is approved by the supplying company. Each dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery. Describe the recordings or trajectories you need.
Frequently asked questions
Can an IDM trained on one application label another?
Usually only partly. Expect accuracy to drop on unseen applications, especially for text entry and shortcuts, so add labeled samples from each new application family and re-measure per action type.
Is OCR on frames enough to recover typed text?
It recovers the final visible string, not the keystrokes, corrections or shortcuts that produced it. For text-heavy workflows, event capture or human annotation of typing remains necessary.
Should pseudo-labeled and event-logged steps be mixed?
Yes, if each step carries its label source and confidence. A common approach is to filter low-confidence IDM labels or down-weight them relative to logged and human-labeled steps.
Sources
- arXiv (He, Jin, Liu; ICLR 2026), "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/pdf/2505.13909
- arXiv (Deng, Su et al.; NeurIPS 2023), "Mind2Web: Towards a Generalist Agent for the Web" (2023). https://arxiv.org/abs/2306.06070v1
- arXiv (Xie et al.; NeurIPS 2024), "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
- arXiv, "OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis" (2024). https://arxiv.org/pdf/2412.19723
- arXiv (Ou, Xu, Madaan et al.; NeurIPS 2024), "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale" (2024). https://arxiv.org/pdf/2409.15637
- arXiv (Northcutt, Athalye, Mueller; NeurIPS 2021), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- arXiv (UGround), "Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents" (2024). https://arxiv.org/abs/2410.05243v2
- Atlassian (Loom documentation), "Record for AI agents". https://support.atlassian.com/loom/docs/record-for-agent/
- Amazon Web Services (Amazon Connect Administrator Guide), "Agent screen recording". https://docs.aws.amazon.com/connect/latest/adminguide/agent-screen-recording.md
- arXiv, "Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos" (2022). https://arxiv.org/abs/2206.11795
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.