Skip to content

Agent, workflow and domain-reasoning data

GUI grounding data from enterprise applications: screens with labeled elements

Quick answer

A GUI grounding dataset pairs screenshots with element-level labels so a model can turn a description such as "the Post button on the journal entry toolbar" into a click point. For enterprise software, specify native-resolution screens from ERP, CRM, ITSM, spreadsheet and admin applications; a box, role, visible text and interactivity flag for each element; several referring expressions per target; synthetic on-screen values that leave the layout intact; and rights covering both the tenant's records and the vendor's interface.

By SourceX Editorial · Updated

Grounding here means locating an interface element, not grounding answers in documents or an enterprise knowledge graph. A grounding set holds still frames and element labels with no action sequence, unlike computer-use trajectory data; the AI agent training data hub maps the rest of the agent data stack.

Public grounding sets: large on the web, small on business software

The largest public grounding corpora are built mostly from web pages or synthesized interfaces, while sets drawn from desktop and professional software are mostly small evaluation benchmarks or cover office suites. Use them as reference points for label design and difficulty, not as evidence that dense business screens are covered.

ResourceContentsWhere the screens come fromWhat it tells a buyer
UGround [1]About 10M elements with referring expressions over about 1.3M screenshotsLargely web-based synthetic dataScale is reachable from the web; enterprise density is not
OSWorld-G and its training corpus [2]564 finely annotated test samples; a 4M-example training corpusTraining data built through interface decomposition and synthesisA test-category template that includes infeasible instructions
ScreenSpot-Pro [3]Authentic, expert-annotated high-resolution screenshots from 23 applications, five industries and three operating systemsProfessional softwareSmall targets on large, complex screens sharply cut accuracy
WinSpot [4]5,000 manually validated and refined coordinate-instruction pairsWindows desktopIts authors note existing datasets mostly cover web elements
GUI-360 [5]Grounding, screen parsing and action prediction; 105,368 training steps across 13,750 trajectoriesWord, Excel and PowerPoint on WindowsOffice suites have public coverage; customized line-of-business apps are not its focus

ScreenSpot-Pro's authors report that the best model they tested scored 18.9% at publication, and that ScreenSeekeR, a planner-guided search that narrows the screen area before grounding, reached 48.1% without additional training [3]. These are 2025 paper figures; check current leaderboards before quoting them. None is described as built around the customized, record-filled screens of an accounts payable clerk or service desk analyst, and "open-sourced" does not settle commercial use; see using public benchmarks commercially.

The element record: boxes, roles, text, state and referring expressions

Each screen needs a lossless native-resolution image with its display geometry, and each element needs a tight box, role, text, state and label source, plus referring expressions that each resolve to exactly one element. Name these fields in your agent data specification before labeling starts.

FieldWhat to specifyDefect to catch
ImageLossless PNG at physical resolution; width, height, OS scale factor, SHA-256JPEG or downscaled frames; logical-pixel boxes on a physical-pixel image
BoxImage-pixel coordinates, convention stated (corner pairs or origin plus size)Mixed conventions; boxes around containers, not hit targets
RoleClosed list: button, text_field, combo_box, checkbox, tab, menu_item, grid_cell, column_header, link, icon_buttonFree-text roles; unmapped roles from different accessibility APIs
TextVisible text and accessible name, separatelyIcon buttons treated as text targets via a hidden name
StateInteractive, enabled, checked, focused, occluded, offscreenDisabled or modal-covered elements labeled clickable
HierarchyParent ID and container (dialog, panel, grid row)"Edit" on every row with nothing to tell them apart
Label sourcea11y_auto, a11y_verified, human_drawn, model_proposedMachine boxes indistinguishable from checked ones
ExpressionsThree or more per target, each typedOnly text-match phrasing; two valid targets

Box tightness changes scores. ScreenSpot-Pro counts a prediction correct when its point falls inside the ground-truth box [3], so a box padded to include a caption and margin quietly raises accuracy. Ask for boxes drawn to the clickable region, plus a visible-extent box where the two differ, such as a checkbox whose caption also toggles it.

Referring expressions need a type mix; OSWorld-G's categories (text matching, element recognition, layout, fine-grained manipulation, infeasibility) are a usable template [2]. Ask for text-match ("click Save"), icon description ("the funnel icon in the Status column header"), spatial ("the Amount field right of Vendor"), functional ("the control that submits this entry for approval"), and infeasible expressions naming something absent from the screen, with a null target. Annotators who can see the box drift toward text-match phrasing, so set a share for each type.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "screen_id": "scr-000812",
  "image": {"file": "screens/scr-000812.png", "width_px": 2560, "height_px": 1440,
            "os_scale": 1.5, "sha256": "<hash>"},
  "app": {"category": "ERP", "client": "browser", "name": "<application>", "build": "<version>",
          "tenant": "t-07", "theme": "dark", "locale": "en-US"},
  "capture": {"method": "sandbox_render", "synthetic_values": true, "a11y_source": "DOM+ARIA"},
  "bbox_format": "x_min,y_min,x_max,y_max in image pixels",
  "elements": [
    {"el_id": "e-041", "bbox": [1932, 214, 2016, 250], "role": "button", "visible_text": "Post",
     "accessible_name": "Post journal entry", "interactive": true, "enabled": true,
     "occluded": false, "parent": "e-030", "label_source": "a11y_verified"},
    {"el_id": "e-118", "bbox": [1410, 702, 1566, 734], "role": "grid_cell", "visible_text": "1,240.00",
     "interactive": true, "enabled": true, "occluded": false, "parent": "row-07",
     "label_source": "human_drawn"}
  ],
  "expressions": [
    {"expr_id": "x-1", "target": "e-041", "type": "text_match", "text": "Click Post"},
    {"expr_id": "x-2", "target": "e-041", "type": "functional", "text": "Submit this journal entry to the ledger"},
    {"expr_id": "x-3", "target": "e-118", "type": "spatial", "text": "Debit amount on the line for account 6100"},
    {"expr_id": "x-4", "target": null, "type": "infeasible", "text": "Open the Approvals tab"}
  ]
}

The grid cell was labeled by hand because the grid exposed no per-cell nodes, the case the label-source flag exists to record.

Coverage: the screen conditions where agents miss clicks

Grounding errors cluster in dense grids, stacked dialogs, small icons on high-resolution displays and unusual scaling, so measure coverage per application and per screen condition rather than as a total screenshot count.

DimensionWhy it mattersWhat to request
Application and client: ERP, CRM, ITSM, EHR, spreadsheets, admin portals; browser, Win32 or WPF, Java, CitrixEach toolkit draws and exposes controls differentlyScreens per application, build and client type
Resolution and OS scaling: 1920 × 1080 to 3840 × 2160; 100% to 200%Target pixel size shifts; ScreenSpot-Pro isolates high-resolution difficulty [3]Target box area as a share of the screen
Density: grids with hundreds of cells, long forms, ribbonsNear-identical targets competeElements per screen; repeated-label share
Overlays: modals, dropdowns, date pickers, stacked windowsCovered elements look clickableScreens with open overlays; occlusion flags
Theme and locale: light, dark, high-contrast; languagesContrast and phrasing changeCounts per theme and locale
Window layout: multiple monitors, partly visible windowsCoordinates span displays; windows clip elementsMonitor layout in metadata
Tenant customization and release driftFields and labels differ per tenant and buildDistinct tenants per application; build numbers

Worked example: why native resolution matters

A 24 × 24 px toolbar icon on a 3840 × 2160 display covers 576 of 8,294,400 pixels, under 0.01% of the screen. Resized to 1280 × 720 before labeling, the icon becomes 8 × 8 px and its box can no longer be drawn accurately. Ask for native-resolution images and keep resizing a training-time choice.

In document layout analysis, source diversity mattered: models trained on DocLayNet's varied sources were more robust than those trained on PubLayNet or DocBank, drawn mainly from scientific repositories [6]. The grounding analogue is many applications and tenants rather than more screens of one deployment; see configuration data for enterprise app replicas.

Accessibility trees bootstrap labels but need human verification

Accessibility APIs produce element boxes, roles and names almost for free, but accessibility trees can be noisy and incomplete [1], and enterprise and legacy applications often expose partial or stale ones, so a grounding dataset must say how each label was derived and include a human-verified sample per application.

On Windows, UI Automation exposes for each element a BoundingRectangle (the screen coordinates of the rectangle that encloses it), a ControlType, an AutomationId and an IsOffscreen flag [7]. Microsoft's documentation also describes AutomationId as expected to be unique among siblings but not across the desktop, optional for applications to support, and not guaranteed to be stable across releases or builds [7], so it is not a durable key for deduplication or cross-version splits. Browser apps expose comparable data through the DOM and ARIA roles, macOS through its accessibility API and Linux desktops through AT-SPI, each with a role vocabulary to map onto yours.

Failure modes to test on a sample:

  • Canvas-drawn or custom-painted grids exposed as a single node.
  • Icon buttons with no accessible name, or a generated one such as "button1".
  • Boxes for containers rather than hit targets, or in logical pixels on a high-DPI capture.
  • Nodes still listed while hidden by a modal, scrolled away or mid-animation.
  • Citrix, RDP and terminal-emulator sessions that expose only pixels; see legacy desktop, terminal and virtual-desktop interaction data.

Per application, ask for a random sample where annotators correct every machine box, with the correction rate, the share of visible elements missing from the tree, and two-annotator agreement on boxes (intersection over union, IoU, above a stated threshold) and roles (a chance-corrected coefficient such as Krippendorff's alpha, where 1 is perfect agreement and 0 is chance [8]); see inter-annotator agreement metrics. Give evaluation splits a full second pass: an audit of 10 widely used test sets estimated an average label error rate of at least 3.3% and showed that such errors can change model rankings [9].

Replacing on-screen customer data without moving the pixels

Business screens show customer names, emails, account numbers and patient identifiers, so the dataset needs synthetic replacements that keep every box aligned with the image: redaction that changes text length shifts truncation, column widths and wrapping, and labels drift from the pixels.

MethodEffect on labelsWhen it fits
Render synthetic values at the data layer (seeded sandbox or replaced database values), then captureLayout and boxes are native; expressions use the synthetic valuesPreferred where the supplier can re-render; see seed data for agent sandboxes
Replace pixels with same-width synthetic strings in the same fontBoxes hold only if rendered widths match; check every replaced regionCaptured screens that cannot be re-rendered
Solid masks or blurText and text-match expressions on masked elements are lost; models learn the masksLast resort; flag masked elements

Detection tools miss things: Presidio, an open-source PII detection and anonymization SDK, cautions that there is no guarantee it finds all sensitive information [10]. Check window titles, tabs, tooltips and autocomplete lists too, and keep each synthetic value identical across the image, text fields and expressions.

Clinical screens raise the bar: names, medical record numbers and date elements other than the year are among the 18 identifiers HIPAA's Safe Harbor method removes [11], and an EHR patient banner often shows several at once. See PII in screen recordings and computer-use trajectories.

Two rights layers: the tenant's records and the vendor's interface

A screenshot carries the operating company's records and the software vendor's interface, so confirm per application who can grant training rights to each layer before licensing.

  • Tenant records. The business using the software usually controls the data on screen, subject to its customer contracts and privacy notices; confirm it may share them for model training.
  • Interface imagery. Layouts, icons, illustrations and help text may belong to the software vendor, and subscription or license agreements can limit screenshots, benchmarking or reverse engineering. Ask which vendor agreements govern each application.
  • US copyright position. The US Copyright Office's Part 3 report, still a May 2025 pre-publication version as of October 2026, concludes that copying works into training datasets may be prima facie infringing unless an exception such as fair use applies [12]. It is advisory and does not bind courts.
  • Third-party content. Open emails, attachments and other companies' documents carry their own rights; see rights in third-party software and data captured on screen.
  • Uses. Training, evaluation, publishing results and building replica environments are separate grants; see license terms for agent data and rights in replicating third-party software.

This section is general information, not legal advice; have counsel review the vendor agreements involved.

Sample checks before you license a grounding set

Box geometry, ambiguous expressions and leftover personal data are cheap to find in a sample and expensive after training, so run these checks before signing; evaluating an agent data sample covers sample size.

  1. Overlay boxes on a random sample of screenshots; flag offset boxes and boxes drawn around containers.
  2. Confirm image dimensions match the recorded scale factor and that boxes sit in image pixels.
  3. Resolve sampled expressions by hand: exactly one valid target, or a null target for infeasible ones.
  4. Compare counts by application, build, tenant, resolution, theme, role, expression type and target size with your specification.
  5. Check label-source flags and the human-verified share per application.
  6. Find near-duplicate screens (the same form showing different records) and hold out whole applications or tenants for evaluation; see contamination-resistant evaluation design.
  7. Scan pixels and text fields for residual personal data, including titles and tabs.
  8. Load the export in your own tooling; COCO-style annotation JSON carries info, licenses, categories, images and annotations sections [13], and the box convention must match your loader.
  9. Read the documentation: capture method, accessibility source, redaction method, annotation guidelines and build list.

End-to-end tests on the same applications are covered in computer-use agent evaluation tasks.

Sourcing labeled screens from operating businesses

Screens of customized, record-filled business software sit with the companies that run it, not on the public web. SourceX sources operational datasets from US companies, including new recordings of hands-on work, and manages the licensing agreements; it does not source scraped public web content. Datasets are sourced on request rather than held in stock, and a request does not guarantee a match. Each dataset goes through rights review and is delivered under a license defining which records are included and their allowed uses; personal details are removed or replaced before delivery with the method recorded, though no de-identification method is perfect.

A workable request names applications, builds and minimum tenant count, required screen conditions, which label layers must arrive with the screens, the redaction method and the uses to license.

Related records are covered in training data for computer-use agents, licensing screen recordings and the computer-use data glossary entry. Teams can describe the screens and labels they need to SourceX instead of approaching software customers themselves.

Need labeled screens from real business software?

Describe the applications, screen conditions, element fields and uses you need licensed. SourceX looks for US businesses that hold that data, checks the data and each supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Send SourceX your grounding data specification.

Sources

  1. Gou et al., "Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents (UGround)" (2024). https://arxiv.org/pdf/2410.05243v1
  2. arXiv:2505.13227, "Scaling Computer-Use Grounding via User Interface Decomposition and Synthesis" (2025). https://arxiv.org/html/2505.13227v1
  3. Li et al. (National University of Singapore, East China Normal University, Hong Kong Baptist University), "ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use" (2025). https://arxiv.org/abs/2504.07981v1
  4. ACL Anthology, "WinSpot: GUI Grounding Benchmark with Multimodal Large Language Models" (2025). https://preview.aclanthology.org/setup/2025.acl-short.85
  5. alphaXiv (benchmark listing), "GUI-360" (accessed 2026). https://www.alphaxiv.org/benchmarks/nanjing-university/gui-360
  6. Pfitzmann, Auer, Dolfi, Nassar, Staar (IBM Research), "DocLayNet: A Large Human-Annotated Dataset for Document-Layout Analysis" (2022). https://arxiv.org/abs/2206.01062v1
  7. Microsoft, "AutomationElement.AutomationElementInformation Struct" (UI Automation documentation, accessed 2026). https://msdn.microsoft.com/en-us/library/ms606803.aspx
  8. Klaus Krippendorff (University of Pennsylvania), "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  9. Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  10. Microsoft presidio project (pkg.go.dev), "Presidio - Data Protection API" (accessed 2026). https://pkg.go.dev/github.com/microsoft/presidio
  11. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  12. U.S. Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training (Pre-Publication Version)" (2025). https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Version.pdf
  13. CVAT.ai, "COCO (CVAT format documentation)" (accessed 2026). https://docs.cvat.ai/docs/dataset_management/formats/format-coco/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data