Agent, workflow and domain-reasoning data
GUI grounding data from enterprise applications: screens with labeled elements
Quick answer
A GUI grounding dataset pairs screenshots with element-level labels so a model can turn a description such as "the Post button on the journal entry toolbar" into a click point. For enterprise software, specify native-resolution screens from ERP, CRM, ITSM, spreadsheet and admin applications; a box, role, visible text and interactivity flag for each element; several referring expressions per target; synthetic on-screen values that leave the layout intact; and rights covering both the tenant's records and the vendor's interface.
By SourceX Editorial · Updated
Grounding here means locating an interface element, not grounding answers in documents or an enterprise knowledge graph. A grounding set holds still frames and element labels with no action sequence, unlike computer-use trajectory data; the AI agent training data hub maps the rest of the agent data stack.
Public grounding sets: large on the web, small on business software
The largest public grounding corpora are built mostly from web pages or synthesized interfaces, while sets drawn from desktop and professional software are mostly small evaluation benchmarks or cover office suites. Use them as reference points for label design and difficulty, not as evidence that dense business screens are covered.
| Resource | Contents | Where the screens come from | What it tells a buyer |
|---|---|---|---|
| UGround [1] | About 10M elements with referring expressions over about 1.3M screenshots | Largely web-based synthetic data | Scale is reachable from the web; enterprise density is not |
| OSWorld-G and its training corpus [2] | 564 finely annotated test samples; a 4M-example training corpus | Training data built through interface decomposition and synthesis | A test-category template that includes infeasible instructions |
| ScreenSpot-Pro [3] | Authentic, expert-annotated high-resolution screenshots from 23 applications, five industries and three operating systems | Professional software | Small targets on large, complex screens sharply cut accuracy |
| WinSpot [4] | 5,000 manually validated and refined coordinate-instruction pairs | Windows desktop | Its authors note existing datasets mostly cover web elements |
| GUI-360 [5] | Grounding, screen parsing and action prediction; 105,368 training steps across 13,750 trajectories | Word, Excel and PowerPoint on Windows | Office suites have public coverage; customized line-of-business apps are not its focus |
ScreenSpot-Pro's authors report that the best model they tested scored 18.9% at publication, and that ScreenSeekeR, a planner-guided search that narrows the screen area before grounding, reached 48.1% without additional training [3]. These are 2025 paper figures; check current leaderboards before quoting them. None is described as built around the customized, record-filled screens of an accounts payable clerk or service desk analyst, and "open-sourced" does not settle commercial use; see using public benchmarks commercially.
The element record: boxes, roles, text, state and referring expressions
Each screen needs a lossless native-resolution image with its display geometry, and each element needs a tight box, role, text, state and label source, plus referring expressions that each resolve to exactly one element. Name these fields in your agent data specification before labeling starts.
| Field | What to specify | Defect to catch |
|---|---|---|
| Image | Lossless PNG at physical resolution; width, height, OS scale factor, SHA-256 | JPEG or downscaled frames; logical-pixel boxes on a physical-pixel image |
| Box | Image-pixel coordinates, convention stated (corner pairs or origin plus size) | Mixed conventions; boxes around containers, not hit targets |
| Role | Closed list: button, text_field, combo_box, checkbox, tab, menu_item, grid_cell, column_header, link, icon_button | Free-text roles; unmapped roles from different accessibility APIs |
| Text | Visible text and accessible name, separately | Icon buttons treated as text targets via a hidden name |
| State | Interactive, enabled, checked, focused, occluded, offscreen | Disabled or modal-covered elements labeled clickable |
| Hierarchy | Parent ID and container (dialog, panel, grid row) | "Edit" on every row with nothing to tell them apart |
| Label source | a11y_auto, a11y_verified, human_drawn, model_proposed | Machine boxes indistinguishable from checked ones |
| Expressions | Three or more per target, each typed | Only text-match phrasing; two valid targets |
Box tightness changes scores. ScreenSpot-Pro counts a prediction correct when its point falls inside the ground-truth box [3], so a box padded to include a caption and margin quietly raises accuracy. Ask for boxes drawn to the clickable region, plus a visible-extent box where the two differ, such as a checkbox whose caption also toggles it.
Referring expressions need a type mix; OSWorld-G's categories (text matching, element recognition, layout, fine-grained manipulation, infeasibility) are a usable template [2]. Ask for text-match ("click Save"), icon description ("the funnel icon in the Status column header"), spatial ("the Amount field right of Vendor"), functional ("the control that submits this entry for approval"), and infeasible expressions naming something absent from the screen, with a null target. Annotators who can see the box drift toward text-match phrasing, so set a share for each type.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"screen_id": "scr-000812",
"image": {"file": "screens/scr-000812.png", "width_px": 2560, "height_px": 1440,
"os_scale": 1.5, "sha256": "<hash>"},
"app": {"category": "ERP", "client": "browser", "name": "<application>", "build": "<version>",
"tenant": "t-07", "theme": "dark", "locale": "en-US"},
"capture": {"method": "sandbox_render", "synthetic_values": true, "a11y_source": "DOM+ARIA"},
"bbox_format": "x_min,y_min,x_max,y_max in image pixels",
"elements": [
{"el_id": "e-041", "bbox": [1932, 214, 2016, 250], "role": "button", "visible_text": "Post",
"accessible_name": "Post journal entry", "interactive": true, "enabled": true,
"occluded": false, "parent": "e-030", "label_source": "a11y_verified"},
{"el_id": "e-118", "bbox": [1410, 702, 1566, 734], "role": "grid_cell", "visible_text": "1,240.00",
"interactive": true, "enabled": true, "occluded": false, "parent": "row-07",
"label_source": "human_drawn"}
],
"expressions": [
{"expr_id": "x-1", "target": "e-041", "type": "text_match", "text": "Click Post"},
{"expr_id": "x-2", "target": "e-041", "type": "functional", "text": "Submit this journal entry to the ledger"},
{"expr_id": "x-3", "target": "e-118", "type": "spatial", "text": "Debit amount on the line for account 6100"},
{"expr_id": "x-4", "target": null, "type": "infeasible", "text": "Open the Approvals tab"}
]
}
The grid cell was labeled by hand because the grid exposed no per-cell nodes, the case the label-source flag exists to record.
Coverage: the screen conditions where agents miss clicks
Grounding errors cluster in dense grids, stacked dialogs, small icons on high-resolution displays and unusual scaling, so measure coverage per application and per screen condition rather than as a total screenshot count.
| Dimension | Why it matters | What to request |
|---|---|---|
| Application and client: ERP, CRM, ITSM, EHR, spreadsheets, admin portals; browser, Win32 or WPF, Java, Citrix | Each toolkit draws and exposes controls differently | Screens per application, build and client type |
| Resolution and OS scaling: 1920 × 1080 to 3840 × 2160; 100% to 200% | Target pixel size shifts; ScreenSpot-Pro isolates high-resolution difficulty [3] | Target box area as a share of the screen |
| Density: grids with hundreds of cells, long forms, ribbons | Near-identical targets compete | Elements per screen; repeated-label share |
| Overlays: modals, dropdowns, date pickers, stacked windows | Covered elements look clickable | Screens with open overlays; occlusion flags |
| Theme and locale: light, dark, high-contrast; languages | Contrast and phrasing change | Counts per theme and locale |
| Window layout: multiple monitors, partly visible windows | Coordinates span displays; windows clip elements | Monitor layout in metadata |
| Tenant customization and release drift | Fields and labels differ per tenant and build | Distinct tenants per application; build numbers |
Worked example: why native resolution matters
A 24 × 24 px toolbar icon on a 3840 × 2160 display covers 576 of 8,294,400 pixels, under 0.01% of the screen. Resized to 1280 × 720 before labeling, the icon becomes 8 × 8 px and its box can no longer be drawn accurately. Ask for native-resolution images and keep resizing a training-time choice.
In document layout analysis, source diversity mattered: models trained on DocLayNet's varied sources were more robust than those trained on PubLayNet or DocBank, drawn mainly from scientific repositories [6]. The grounding analogue is many applications and tenants rather than more screens of one deployment; see configuration data for enterprise app replicas.
Accessibility trees bootstrap labels but need human verification
Accessibility APIs produce element boxes, roles and names almost for free, but accessibility trees can be noisy and incomplete [1], and enterprise and legacy applications often expose partial or stale ones, so a grounding dataset must say how each label was derived and include a human-verified sample per application.
On Windows, UI Automation exposes for each element a BoundingRectangle (the screen coordinates of the rectangle that encloses it), a ControlType, an AutomationId and an IsOffscreen flag [7]. Microsoft's documentation also describes AutomationId as expected to be unique among siblings but not across the desktop, optional for applications to support, and not guaranteed to be stable across releases or builds [7], so it is not a durable key for deduplication or cross-version splits. Browser apps expose comparable data through the DOM and ARIA roles, macOS through its accessibility API and Linux desktops through AT-SPI, each with a role vocabulary to map onto yours.
Failure modes to test on a sample:
- Canvas-drawn or custom-painted grids exposed as a single node.
- Icon buttons with no accessible name, or a generated one such as "button1".
- Boxes for containers rather than hit targets, or in logical pixels on a high-DPI capture.
- Nodes still listed while hidden by a modal, scrolled away or mid-animation.
- Citrix, RDP and terminal-emulator sessions that expose only pixels; see legacy desktop, terminal and virtual-desktop interaction data.
Per application, ask for a random sample where annotators correct every machine box, with the correction rate, the share of visible elements missing from the tree, and two-annotator agreement on boxes (intersection over union, IoU, above a stated threshold) and roles (a chance-corrected coefficient such as Krippendorff's alpha, where 1 is perfect agreement and 0 is chance [8]); see inter-annotator agreement metrics. Give evaluation splits a full second pass: an audit of 10 widely used test sets estimated an average label error rate of at least 3.3% and showed that such errors can change model rankings [9].
Replacing on-screen customer data without moving the pixels
Business screens show customer names, emails, account numbers and patient identifiers, so the dataset needs synthetic replacements that keep every box aligned with the image: redaction that changes text length shifts truncation, column widths and wrapping, and labels drift from the pixels.
| Method | Effect on labels | When it fits |
|---|---|---|
| Render synthetic values at the data layer (seeded sandbox or replaced database values), then capture | Layout and boxes are native; expressions use the synthetic values | Preferred where the supplier can re-render; see seed data for agent sandboxes |
| Replace pixels with same-width synthetic strings in the same font | Boxes hold only if rendered widths match; check every replaced region | Captured screens that cannot be re-rendered |
| Solid masks or blur | Text and text-match expressions on masked elements are lost; models learn the masks | Last resort; flag masked elements |
Detection tools miss things: Presidio, an open-source PII detection and anonymization SDK, cautions that there is no guarantee it finds all sensitive information [10]. Check window titles, tabs, tooltips and autocomplete lists too, and keep each synthetic value identical across the image, text fields and expressions.
Clinical screens raise the bar: names, medical record numbers and date elements other than the year are among the 18 identifiers HIPAA's Safe Harbor method removes [11], and an EHR patient banner often shows several at once. See PII in screen recordings and computer-use trajectories.
Two rights layers: the tenant's records and the vendor's interface
A screenshot carries the operating company's records and the software vendor's interface, so confirm per application who can grant training rights to each layer before licensing.
- Tenant records. The business using the software usually controls the data on screen, subject to its customer contracts and privacy notices; confirm it may share them for model training.
- Interface imagery. Layouts, icons, illustrations and help text may belong to the software vendor, and subscription or license agreements can limit screenshots, benchmarking or reverse engineering. Ask which vendor agreements govern each application.
- US copyright position. The US Copyright Office's Part 3 report, still a May 2025 pre-publication version as of October 2026, concludes that copying works into training datasets may be prima facie infringing unless an exception such as fair use applies [12]. It is advisory and does not bind courts.
- Third-party content. Open emails, attachments and other companies' documents carry their own rights; see rights in third-party software and data captured on screen.
- Uses. Training, evaluation, publishing results and building replica environments are separate grants; see license terms for agent data and rights in replicating third-party software.
This section is general information, not legal advice; have counsel review the vendor agreements involved.
Sample checks before you license a grounding set
Box geometry, ambiguous expressions and leftover personal data are cheap to find in a sample and expensive after training, so run these checks before signing; evaluating an agent data sample covers sample size.
- Overlay boxes on a random sample of screenshots; flag offset boxes and boxes drawn around containers.
- Confirm image dimensions match the recorded scale factor and that boxes sit in image pixels.
- Resolve sampled expressions by hand: exactly one valid target, or a null target for infeasible ones.
- Compare counts by application, build, tenant, resolution, theme, role, expression type and target size with your specification.
- Check label-source flags and the human-verified share per application.
- Find near-duplicate screens (the same form showing different records) and hold out whole applications or tenants for evaluation; see contamination-resistant evaluation design.
- Scan pixels and text fields for residual personal data, including titles and tabs.
- Load the export in your own tooling; COCO-style annotation JSON carries info, licenses, categories, images and annotations sections [13], and the box convention must match your loader.
- Read the documentation: capture method, accessibility source, redaction method, annotation guidelines and build list.
End-to-end tests on the same applications are covered in computer-use agent evaluation tasks.
Sourcing labeled screens from operating businesses
Screens of customized, record-filled business software sit with the companies that run it, not on the public web. SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the licensing agreements; it does not source scraped public web content. Datasets are sourced on request rather than held in stock, and a request does not guarantee a match. Each dataset goes through rights review and is delivered under a license defining which records are included and their allowed uses; personal details are removed or replaced before delivery with the method recorded, though no de-identification method is perfect.
A workable request names applications, builds and minimum tenant count, required screen conditions, which label layers must arrive with the screens, the redaction method and the uses to license.
Related records are covered in training data for computer-use agents, licensing screen recordings and the computer-use data glossary entry. Teams can describe the screens and labels they need to SourceX instead of approaching software customers themselves.
Need labeled screens from real business software?
Describe the applications, screen conditions, element fields and uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Send SourceX your grounding data specification.
Sources
- Gou et al., "Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents (UGround)" (2024). https://arxiv.org/pdf/2410.05243v1
- arXiv:2505.13227, "Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis" (2025). https://arxiv.org/html/2505.13227v1
- Li et al. (National University of Singapore, East China Normal University, Hong Kong Baptist University), "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use" (2025). https://arxiv.org/abs/2504.07981v1
- ACL Anthology, "WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models" (2025). https://preview.aclanthology.org/setup/2025.acl-short.85
- alphaXiv (benchmark listing), "GUI-360" (accessed 2026). https://www.alphaxiv.org/benchmarks/nanjing-university/gui-360
- Pfitzmann, Auer, Dolfi, Nassar, Staar (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
- Microsoft, "AutomationElement.AutomationElementInformation Struct" (UI Automation documentation, accessed 2026). https://msdn.microsoft.com/en-us/library/ms606803.aspx
- Klaus Krippendorff (University of Pennsylvania), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API" (accessed 2026). https://pkg.go.dev/github.com/microsoft/presidio
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
- CVAT.ai, "COCO (CVAT format documentation)" (accessed 2026). https://docs.cvat.ai/docs/dataset_management/formats/format-coco/
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.