Skip to content

Multimodal and embodied data

Support Tickets with Screenshots and Attachments for Multimodal Support Models

Quick answer

A customer support dataset with screenshots is a set of resolved tickets where an image, photo or file is part of the problem statement and is linked to the agent replies and the final resolution. To be useful for training or evaluating a multimodal support agent, each record needs the attachment in its original position in the thread, a resolution code, the product version, and a label saying whether the image was actually needed to solve the issue. Screenshots also need OCR-based redaction plus visual review before release.

By SourceX Editorial · Updated

What a usable screenshot ticket record contains

A usable record keeps the image attached to the exact message where the customer sent it, not as a loose file at ticket level. Text-led ticket exports, such as those described on our customer support ticket datasets and ITSM ticket datasets pages, often drop attachments or flatten them. For multimodal work, the order of events matters: a screenshot sent after the agent asked "can you show me the error?" teaches a different behavior from one sent in the opening message.

In Zendesk, for example, each reply is recorded as a Comment event in the ticket's audit trail, with its author and whether it was public or an internal note [1]. Exports built from those events keep attachments tied to the right message, while exports that read only ticket-level fields, or skip images pasted inline into the comment body, can silently lose them. Ask the supplier which helpdesk produced the data and how inline images, forwarded email attachments and chat uploads were captured.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "ticket_id": "T-PSEUDO-48121",
  "channel": "web_form",
  "product": "billing-portal",
  "product_version": "2026.3.1",
  "messages": [
    {"seq": 1, "role": "customer", "text": "Export to CSV fails with this error", "attachments": [
      {"att_id": "A1", "kind": "ui_screenshot", "mime": "image/png", "inline": true,
       "width": 1920, "height": 1080, "redaction": {"ocr_pass": true, "visual_review": true, "boxes": 4}}]},
    {"seq": 2, "role": "agent", "text": "Thanks. The banner shows error E-417; clearing the date filter should fix it."},
    {"seq": 3, "role": "customer", "text": "That worked."}
  ],
  "resolution_code": "user_config_filter",
  "attachment_necessary": true,
  "attachment_necessity_basis": "error code visible only in screenshot",
  "status": "solved"
}

Screenshots, photos and files need different handling

UI screenshots, photos of physical things and document attachments are three different modalities with different modeling and privacy risks. Treat them as separate strata in your specification and in your evaluation splits, and ask for a kind field on every attachment.

Attachment kindTypical contentMain modeling taskMain de-identification risk
UI screenshotError dialogs, settings pages, dashboardsReading text, locating UI elements, mapping to known errorsNames, emails, account numbers and other customers' rows visible on screen
PhotoHardware, device labels, receipts, damaged goodsRecognizing objects, reading serials and labelsFaces, home interiors, serial numbers, location metadata
Document or filePDFs, invoices, logs, CSV exportsExtracting fields, parsing logsFull records of third parties, payment data, embedded metadata
Screen recordingShort clips of a workflowFollowing steps over timeEvery frame can expose different data

Document attachments resemble form-understanding data more than chat. Public sets such as FUNSD, with 199 annotated scanned forms [6], show how small and narrow open benchmarks are compared with the invoices, logs and exports customers actually attach. For screen recordings, see our guide on how to de-identify screen recordings.

De-identifying screenshots takes OCR plus human review

Screenshots carry personal data in pixels, so text-only redaction of the ticket body is not enough. A common first pass runs OCR, detects entities in the recognized text and paints over the matching boxes; Microsoft's open-source Presidio SDK covers de-identification for both text and images [2]. Treat any such tool as a first pass: OCR misses small fonts, low-contrast text, rotated photos and partially cropped fields, so a visual review pass is still needed.

Watch for these failure modes:

  • Other customers' data. An admin's screenshot of a list view exposes dozens of unrelated people who never opened the ticket.
  • Avatars and faces. Profile pictures and webcam thumbnails are not text and are invisible to OCR.
  • URLs and browser chrome. Address bars hold tenant names, record IDs and session tokens.
  • File metadata. EXIF location on phone photos, and author fields in PDFs and Office files.
  • Health and payment screens. A screenshot of a patient portal can contain protected health information when it comes from a HIPAA covered entity or business associate; Safe Harbor identifiers include names, phone numbers, emails, account numbers, URLs and IP addresses [8].

Memorization raises the stakes: image models have been shown to regenerate individual training images, including photos of people and logos [7]. Our multimodal de-identification guide covers faces, voices and screens together. SourceX removes or replaces personal details such as names, emails, phones and account numbers before delivery, records the method and checks a sample; no method is perfect, and health records require HIPAA de-identification under Safe Harbor or Expert Determination.

Measure whether the image was actually needed

Label each ticket for whether the attachment was necessary to resolve it, because many screenshots repeat what the text already says. Without this label, a model can score well on a "multimodal" set while ignoring the pixels. A practical labeling rule: the attachment is necessary if a trained agent reading only the text could not reach the logged resolution, for example when an error code, version number or device model appears only in the image.

Keep a text-only baseline on the same tickets. If the multimodal model beats the text-only model only on the attachment-necessary slice, you have evidence the images are doing work; if the gap appears everywhere, check for leakage through agent replies that restate the screenshot. The resolution_code field then lets you score outcome accuracy rather than reply fluency.

Evaluation splits that reflect new product releases

Hold out tickets by product release, not by random sampling, so the evaluation tests generalization to user interfaces the model has not seen. Random splits let near-duplicate screenshots of the same dialog land in both train and test, which inflates scores. Group by product_version and by a perceptual hash of each image before splitting.

Public agent benchmarks show why realistic evaluation matters. On WebArena's self-hosted websites, the best GPT-4 agent reached 14.41% end-to-end success against 78.24% for humans in the v4 paper [4], and OSWorld uses execution-based checks across 369 real desktop and web tasks [3]. Policy-bound support benchmarks such as tau-bench test agents against simulated users and domain rules [5]. Real tickets complement these with the messy screenshots customers actually send; for policy-following tests, see customer support agent policy evaluation, and for multimodal holdouts, private multimodal evaluation sets.

Rights questions specific to attachments

Attachments add rights layers that ticket text does not. A customer's photo, a third-party invoice or a screenshot of another vendor's software may each sit under different terms from the ticket itself. Before buying, check that the supplier's customer contracts and DPAs allow AI training use, a topic covered in customer contracts and DPAs for training use and in licensing multimodal records from several rightsholders.

Practical diligence asks:

  1. Which helpdesk and channels produced the data, and were inline images, internal notes and chat uploads exported [1]?
  2. Were any attachments blocked or flagged by the helpdesk's malware scanning, and how were they handled?
  3. What share of tickets contain at least one attachment, broken down by kind?
  4. What redaction pipeline ran, at what OCR settings, and what did the visual review sample find?
  5. Is there machine-readable documentation, for example Croissant-RAI metadata, describing labeling and preparation [9]?

How SourceX approaches screenshot ticket requests

SourceX sources operational datasets, including support histories, from US companies and manages the commercial process, including licensing agreements and ongoing purchases. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Buyers describe the data they need, and every release is approved by the supplying company. You can start from the multimodal data cluster or describe your requirement on the SourceX buyers page.

Each dataset is rights-reviewed for ownership and consents and delivered under a license defining records, uses, term and delivery. Diligence materials covering source, rights, preparation and allowed use are prepared per dataset, and delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. SourceX does not train models and does not publish prices; terms are agreed per deal. For text-led support data, see customer support AI training data and eval sets, and browse the wider AI data hub.

Request support tickets with screenshots

If you need resolved tickets where screenshots, photos or files are part of the problem, describe the attachment kinds, products, channels and labels you need. SourceX looks for US businesses that hold that data and runs Find, Assess, Agree, Transact and Manage, with nothing contracted until a supplier agrees. Describe your dataset request on the SourceX buyers page.

Sources

  1. Zendesk Developer Docs, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
  2. Microsoft, "Getting started with image de-identification with Presidio". https://microsoft.github.io/presidio
  3. Xie et al., arXiv / NeurIPS 2024, "OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments" (2024). https://arxiv.org/abs/2404.07972v2
  4. Zhou, Xu et al., arXiv, "WebArena: A Realistic Web Environment for Building Autonomous Agents" (2024). https://arxiv.org/abs/2307.13854v4
  5. Sierra Research, arXiv, "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  6. Jaume, Ekenel, Thiran, arXiv, "FUNSD: A Dataset for Form Understanding in Noisy Scanned Documents" (2019). https://arxiv.org/pdf/1905.13538
  7. Carlini et al., USENIX Security 2023, "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
  8. eCFR / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  9. Jain et al., MLCommons, arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data