Privacy, de-identification and sensitive data
PII in screen recordings and computer-use trajectories: redaction for agent training data
Quick answer
A computer-use trajectory stores the same personal data in up to six places at once: rendered pixels, the DOM or accessibility tree, typed keystrokes, URLs, clipboard events and tool-call arguments. Redacting only the video, or only the text log, leaves the other copies intact. Treat each step as a multi-channel record, detect entities once, propagate one surrogate to every channel, remove credentials outright, use solid masks rather than blur, and measure residual PII on your own trajectories before training.
By SourceX Editorial · Updated
Where personal data hides in a computer-use step
Personal data in agent trajectories sits in at least six parallel channels, and each needs its own detector. Most failures come from a team securing the channel it looks at most (usually video) and forgetting the others. The step-level record format in computer-use trajectory data shows why: a single click event can carry a screenshot, an accessibility subtree, a target element label and a typed value.
- Pixels. Screenshots and video frames show names in CRM headers, email previews, avatars, address bars, notification toasts, taskbar previews and second-monitor content. OCR is the only way to find text here, and it misses small, rotated or low-contrast type.
- DOM and accessibility tree.
valueattributes on<input>elements,aria-labelstrings,titletooltips, hidden form fields,data-*attributes and off-screen rows of virtualized tables. These often contain more PII than the visible frame, because the tree includes content that is scrolled away. - Keystrokes and typed values.
typeactions record exactly what a demonstrator entered, including passwords typed into fields whose screenshot showed only dots. - URLs and navigation logs. Query strings carry emails, order IDs, search terms and OAuth
codeorstateparameters; fragments can carry access tokens. - Clipboard and file events. Paste payloads, downloaded file names (
Smith_J_W2_2025.pdf) and file-picker paths that include OS usernames (/Users/jsmith/). - Tool-call arguments and results. API calls made by an agent or captured from an integration (
send_email(to=...), SQLWHEREclauses, JSON response bodies). Developers have documented PII redactors that scrub chat messages but pass tool-call arguments through untouched [1].
Window titles, system clocks and timezone strings are lower-risk but still quasi-identifiers when combined with an employer name visible in a logo.
Why single-channel redaction fails for agent data
Single-channel redaction fails because the channels are aligned by design, so any unredacted copy re-identifies the redacted one. If a screenshot masks "Dana Ortiz" but the paired accessibility node still reads aria-label="Dana Ortiz, Account owner", the model learns the name and so does anyone who reads the record. The reverse also happens: a text pipeline replaces the typed value with a surrogate while the next frame still shows the original string echoed in a confirmation banner.
Alignment also breaks the training signal when channels disagree. If the keystroke log says the agent typed [EMAIL_1] but the post-action screenshot shows maria.chen@example.com, the model receives contradictory supervision about what typing does. Annotation pipelines such as the tooling released with OpenCUA capture screens, actions and structured state together [5], which is exactly why redaction has to be joined at the step level rather than run as separate jobs over separate files.
Research on GUI grounding notes that accessibility trees can be noisy or incomplete relative to the rendered screen [8]. For redaction this cuts both ways: text in the frame may have no tree node (canvas-rendered apps, remote desktop sessions, PDFs in a viewer), and the tree may hold text that never appears in the frame.
A consistent redaction pipeline across pixels, DOM, keystrokes and tool calls
The workable pattern is detect-once, replace-everywhere: build an entity table for each episode, then render every channel from that table. Running independent redactors per channel produces different surrogates for the same person and unrecoverable inconsistencies.
- Extract all text per step. Pull DOM/accessibility text, typed values, URLs (decoded), clipboard payloads, tool-call arguments and results, plus OCR output with bounding boxes for each frame.
- Detect entities on the union. Run NER plus pattern recognizers (emails, phone numbers, card numbers with Luhn check, IBANs, SSN formats) over the combined text. Presidio is a common open-source base, but the project itself warns it cannot guarantee finding all sensitive information [4]. For the trade-offs between model-based and rule-based detectors, see LLM-based PII redaction vs NER and regex.
- Build the episode entity table. Map each detected value, and its variants ("Dana Ortiz", "D. Ortiz", "dortiz@"), to one stable surrogate key per episode.
- Separate secrets from PII. Passwords, API keys, session cookies, bearer tokens, OAuth codes, one-time passcodes and private keys go to a removal path, not a surrogate path. Secrets in code and shared documents are common enough to need dedicated scanners [7].
- Rewrite text channels. Replace values in DOM snapshots, typed values, URLs, clipboard and tool-call JSON with surrogates of the same type and format.
- Mask pixels. For every OCR box or DOM element bounding box that matches the entity table, paint an opaque fill or render the surrogate text in place. Do not blur or pixelate: research on image data release finds that common blurring can be partially reversed [2].
- Re-OCR the output and diff. Run OCR again on the redacted frames and search for any original entity string. Any hit fails the episode.
- Log the method. Store detector versions, entity types covered and the masking style per episode alongside the data.
Surrogate replacement keeps the action semantics a model needs: the agent still learns "type an email address into the To field and press Enter." Masking with a [REDACTED] token removes that structure. The masking vs surrogate replacement page covers how each style changes what the model learns.
Credentials and session tokens in agent traces
Credentials must be removed and the episode re-validated, never replaced with a realistic-looking fake. A synthetic password that looks valid teaches the model to type plausible secrets, and a surrogate cookie in a replayable trace can mislead an eval harness into thinking authentication succeeded.
Typical leak points include Authorization: Bearer headers inside captured network logs, Set-Cookie values in HAR files, ?token= and #access_token= in URLs, .env files opened in an editor, terminal history showing export AWS_SECRET_ACCESS_KEY=, and password managers auto-filling fields during the recording. Use a fixed placeholder such as <SECRET:password> so the model learns the action ("enter the password here") without any value. If a secret appears in a frame, mask it and treat the source credential as exposed: the supplying company should rotate it regardless of redaction.
Observability exports deserve the same scrutiny. The OpenTelemetry GenAI semantic conventions define execute_tool spans whose gen_ai.tool.call.arguments and result attributes are captured on an opt-in basis [6]. If a supplier enabled that capture, their trace exports may contain full arguments, including credentials passed to tools.
Illustrative redacted step record
The record below shows one step after redaction, with every channel rewritten from one entity table. Store episodes as JSON Lines, one step per line in UTF-8 [11], so a redaction audit can stream and diff them.
Illustrative example: invented to show structure; it does not describe an available dataset.
{"episode_id":"ep_0412","step":7,
"action":{"type":"type","target":"input#to","typed_value":"{{EMAIL_1}}"},
"url":"https://mail.example.app/compose?draft={{ID_3}}",
"a11y_node":{"role":"textbox","name":"To","value":"{{EMAIL_1}}"},
"screenshot":"frames/ep_0412_007.png",
"pixel_masks":[{"bbox":[412,188,640,206],"entity":"EMAIL_1","style":"surrogate_render"},
{"bbox":[18,4,210,22],"entity":"PERSON_1","style":"solid_fill"}],
"clipboard":null,
"tool_call":{"name":"send_email","arguments":{"to":"{{EMAIL_1}}","subject":"Invoice {{ID_3}}"}},
"secrets_removed":[{"channel":"network_log","kind":"bearer_token"}],
"redaction":{"entity_table":"ep_0412.entities.enc","detectors":["ner-v4","regex-v12","ocr-v3"],"reocr_check":"pass"}}
Note that entity_table is encrypted and held by the supplier, not delivered. Delivering the mapping would let anyone reverse the surrogates.
Measuring residual PII before you train
There is no standard benchmark for redaction quality on agent trajectories, so you have to measure on your own data. As of October 2026, even widely used tools such as Presidio publish no official accuracy benchmark [3]. Build a labeled holdout of real episodes and score recall per channel, not overall.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Channel | What to label in the holdout | Common miss | Check |
|---|---|---|---|
| Pixels | Every visible name, email, phone, address, account number | Small UI text, avatars with initials, browser autofill dropdowns | Re-OCR after masking; human review of a frame sample |
| DOM / a11y tree | value, aria-label, title, hidden inputs | Virtualized off-screen rows | String search of entity table against serialized tree |
| Keystrokes | All type payloads | Values typed into password fields | Confirm no raw typed value survives outside the entity table |
| URLs | Query strings and fragments | URL-encoded emails (%40) | Decode before matching |
| Tool calls | Arguments and results JSON | Nested arrays, base64 bodies | Recursive JSON walk; decode common encodings |
| Clipboard / files | Paste payloads, file names, paths | OS usernames in paths | Path component scan |
For sampling plans and acceptance thresholds, see auditing residual PII in a delivered dataset. NIST's Generative AI Profile lists data privacy, including leakage of personal information from training data, among risks generative AI exacerbates [10], which is the reason the holdout result belongs in your model documentation.
Health, finance and third-party data on screen
Regulated records on screen do not lose their status because they arrived as video. A recording of a clinician working in an EHR can contain protected health information; when the recording is held by a HIPAA covered entity or business associate, de-identification requires Safe Harbor removal of 18 identifier types, which include full-face photographs and any other unique identifying number, or an Expert Determination [9]. A support agent's session in a billing tool can expose primary account numbers, which belong on the removal path.
Third parties appear constantly: customers in a CRM, colleagues in a chat sidebar, senders in an inbox preview. Rights in that content are a separate question from redaction, covered in screen recordings and third-party content rights. For the wider video case (webcams, bystanders, audio), see anonymizing video datasets for AI, and for general guidance from SourceX on this source type, see how to de-identify screen recordings and de-identifying workflow and screen activity data.
What to ask a supplier of screen and trajectory data
Ask suppliers for evidence channel by channel, because "the videos are redacted" does not answer the question. A request should cover:
- Which channels are delivered (frames, video, DOM snapshots, a11y trees, keystrokes, network logs, tool calls) and which were redacted.
- Detector list and versions, entity types covered, and the masking style for pixels (solid fill, surrogate render) versus text (surrogate, token).
- How the same person is kept consistent across channels and steps.
- How credentials, cookies and tokens were removed, and whether exposed credentials were rotated.
- Per-channel residual PII results on a sampled holdout, with the sample size and method.
- Whether capture tooling had input masking enabled at record time, which beats redaction after the fact.
Use the de-identification evidence package checklist to turn these into a document request, and the buyer's guide to de-identified AI training data for the wider privacy picture. Buyers comparing sources can also review how SourceX approaches licensing screen recordings for AI training. On SourceX, personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you are scoping a request for this kind of data, you can describe the trajectories you need to SourceX.
Sourcing redacted computer-use trajectories
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request; nothing is held in stock and a request does not guarantee a match. Every dataset is rights-reviewed, personal details are removed or replaced with the method recorded, and delivery runs through private, access-controlled workflows under a license. Describe the screen and trajectory data you need at SourceX for buyers.
Frequently asked questions
Is blurring enough for on-screen text?
No. Blur and pixelation can be partially reversed [2]. Use an opaque fill, or render the surrogate value in place so the frame stays consistent with the text channels.
Should redaction happen at capture time or afterward?
Both. Input masking in the recorder (for example, never logging password-field values) prevents the worst leaks, but it cannot catch names in rendered content, so post-capture detection is still required.
Can I keep the entity mapping to re-link episodes later?
Keep consistent surrogates within an episode, and across episodes only if your use case needs it. The mapping itself should stay with the data holder; a delivered mapping turns pseudonymized data back into identified data.
Sources
- DEV Community, "Your PII redactor probably leaks tool-call arguments". https://dev.to/crp4222/your-pii-redactor-probably-leaks-tool-call-arguments-357e
- arXiv, "Privacy Blur: Quantifying Privacy and Utility for Image Data Release" (2025). https://arxiv.org/pdf/2512.16086
- Grepture, "Presidio is not enough for PII redaction". https://grepture.com/blog/presidio-not-enough-pii-redaction
- Microsoft (microsoft/presidio project), indexed on pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- arXiv (Wang et al.), "OpenCUA: Open Foundations for Computer-Use Agents" (2025). https://arxiv.org/abs/2508.09123
- Greptime, "OpenTelemetry GenAI semantic conventions" (2026). https://greptime.com/blogs/2026-05-09-opentelemetry-genai-semantic-conventions
- arXiv, "Using AI/ML to Find and Remediate Enterprise Secrets in Code & Document Sharing Platforms" (2024). https://arxiv.org/html/2401.01754v1
- arXiv (UGround), "Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents" (2024). https://arxiv.org/pdf/2410.05243
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
- jsonlines.org, "JSON Lines". https://jsonlines.org/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.