Skip to content

Video data

Software Screencasts and Tutorial Recordings for Video-Language Models

Quick answer

A usable software tutorial video dataset pairs screen video with time-aligned narration or captions, keeps on-screen text legible after encoding, records the application name and version for every clip, and arrives with customer data and credentials redacted. Public research sets such as PsTuts, TutorialVQA and CodeSCAN are small and single-application [1][2][3], so teams training video-language models on software usually license internal screencasts, support recordings and training libraries from the companies that made them.

By SourceX Editorial · Updated

Why narrated screencasts are a distinct video-text source

Narrated screencasts are valuable because the speaker usually describes the action as it happens, which gives a loose, free alignment between speech and pixels. A trainer saying "open Filters, then Gaussian Blur" while the cursor moves through the menu is a weak caption for that span. That alignment is a working hypothesis, not a guarantee: narrators run ahead, backtrack and say "this one here" with no visible referent. Plan to measure it on a sample before relying on it.

Software video also differs from natural video in what the model has to read. Most of the signal is text: menu labels, cell values, code, error dialogs, URL bars. Motion is sparse and small (a cursor, a highlight, a scrolling pane), and long stretches are static.

Research reflects this split. PsTuts labels Photoshop tutorials for retrieval and captioning [1], TutorialVQA builds about 6K question-answer-span triples from narrated photo-editing screencasts [3], and CodeSCAN focuses on IDE element detection and OCR across 24 languages and more than 90 themes [2]. Related work extends to question answering over screencasts [4] and to generating code-centric educational video [5].

This page covers screencasts as video-text pairs for understanding, captioning, retrieval and UI video QA. If you need per-step action labels for computer-use agents, the conversion is different; see turning screen recordings into action-labeled trajectories. For general caption density and alignment formats, see video-text pairs and dense video captions.

Where licensed screencast supply comes from

Most usable screencast video sits inside companies, not on public platforms. The common holders are:

  • Software vendors: product walkthroughs, release demos, onboarding videos and customer-education academies, often with scripts and chapter markers.
  • Support and success teams: recorded troubleshooting sessions and asynchronous screen-recording replies to tickets. These are realistic but carry the heaviest customer-data exposure.
  • Internal enablement and IT: recorded trainings for ERP, CRM, EHR and finance systems, plus how-to clips for internal tools.
  • Engineering teams: recorded code reviews, incident walkthroughs and pairing sessions. See developer session recordings for coding-agent use.
  • Meeting archives: shared-screen segments from recorded calls, covered in multimodal meeting recordings.

Polished course libraries behave more like licensed media and are covered in corporate training video libraries. Raw support and internal recordings are messier but show real errors, dead ends and recovery, which is what models need to explain real software use. Tooling is also starting to capture richer context: Loom's agent recording mode, in open beta as of October 2026, stores visited links and console and network logs alongside the narrated video [7]. Ask whether any such sidecar data exists, because it can anchor captions to concrete events.

Specifying legibility: resolution, frame rate and compression

On-screen text must survive the whole pipeline, so specify capture and delivery quality rather than accepting platform re-encodes. Small UI text at 1080p is often readable in the source and unreadable after a streaming-bitrate H.264 transcode, because chroma subsampling and macroblocking smear thin glyphs. Ask for the original capture where possible, not the copy exported for a video host.

Practical points to put in the request:

  • Native resolution and scaling: record capture resolution and the OS display scale factor; a 4K display at 200% scaling produces very different glyph sizes from 1080p at 100%.
  • Codec and bitrate: prefer the original recorder output, or a high-bitrate or visually lossless intermediate; ask for 4:4:4 chroma where the recorder supports it.
  • Variable frame rate: many screen recorders write VFR files. Require container timestamps to be preserved so caption timings stay valid after any conversion.
  • Cursor and click visualization: note whether the cursor is baked in, hidden or highlighted, because it affects both grounding and redaction.
  • OCR check: run an off-the-shelf OCR pass on sampled frames and set a minimum legibility rate as an acceptance criterion.

Packaging matters too: keep per-clip sidecar metadata and a clip index rather than a single long file, as described in packaging video datasets.

Narration, captions and the alignment you actually get

Treat narration as a noisy caption track and specify which text layers come with each clip. The layers you may receive, from weakest to strongest, are: auto-generated captions, the producer's SRT or WebVTT file, the original script or storyboard, human-corrected transcripts with word timestamps, and step-level chapter markers written by the author. Each adds cost and accuracy.

Ask the supplier to state the source of every text layer, because mixing auto-captions and human scripts without a flag makes evaluation unreliable. Check drift: scripted product videos often have voice-over recorded separately and laid over silent capture, so the narration can lead or lag the action by seconds. Unscripted support recordings align more tightly but contain filler, crosstalk and customer voices.

If you later want action labels, research such as VPT shows that a small labeled set can train an inverse dynamics model to pseudo-label actions in unlabeled video [6]. That is a downstream modeling choice; the dataset request should still capture raw video and text cleanly.

Version metadata: UIs date the data

Every clip needs the application name, version or build, and capture date, because UI redesigns silently invalidate what the video teaches. A tutorial recorded before a ribbon redesign or a settings-page move will describe menus that no longer exist, and a model trained on it will confidently give stale instructions. Version tags also let you hold out new releases for evaluation and detect leakage between training and test.

Ask for OS and version, browser and version for web apps, locale and UI language, theme (light, dark, high contrast), and whether the product was a production, staging or demo tenant. Demo tenants are cleaner for privacy but show idealized data; production captures show real density and edge cases.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "clip_id": "scr-000412",
  "source_type": "support_recording",
  "app": {"name": "ExampleCRM", "version": "24.3.1", "surface": "web"},
  "environment": {"os": "Windows 11", "browser": "Chrome 128", "locale": "en-US", "theme": "light", "tenant": "production"},
  "capture": {"resolution": "2560x1440", "display_scale": 1.25, "fps_mode": "VFR", "codec": "H.264 original", "chroma": "4:2:0", "cursor": "visible"},
  "duration_s": 312.4,
  "text_layers": [
    {"type": "transcript", "source": "human_corrected", "format": "WebVTT", "word_timestamps": true},
    {"type": "chapters", "source": "author", "count": 6}
  ],
  "redaction": {"method": "OCR-detect + blur", "fields": ["customer_name", "email", "account_number", "api_key"], "audio": "customer voice removed", "qa_sample_rate": 0.05},
  "third_party_software_on_screen": ["ExampleCRM", "Chrome"],
  "rights": {"owner": "supplier", "narrator_consent": "employee notice on file", "allowed_uses": "per license"}
}

Redaction and on-screen rights for screencasts

Screencasts leak more than ordinary video, because personal and confidential data appears as readable text in tabs, sidebars, notifications and autocomplete. Redaction has to cover pixels, audio and any sidecar logs, and it must persist across frames as windows scroll and move. Static masks fail when a panel resizes; frame-level OCR detection plus tracking is the more robust approach, with human review of a sample. The privacy-specific mechanics are in PII redaction for screen recordings and video anonymization.

Customer data on screen raises a separate question: whether the supplier's own customer terms allow that use. The FTC has warned that companies can face liability if they break promises not to use customer data for undisclosed purposes such as training [9]. Recordings from health IT, such as EHR walkthroughs on live charts, contain protected health information and need HIPAA de-identification by Safe Harbor or Expert Determination [10].

Third-party software is the other layer. A recording shows another vendor's interface, icons, documentation pop-ups and sometimes licensed content inside the app. Read rights in third-party software captured on screen and rights layers in a video clip before scoping.

Public datasets do not settle this for you: an audit of popular dataset hosts found license omission above 70% and license error rates above 50% [8]. For EU distribution, general-purpose model providers have had a copyright-compliance policy duty under AI Act Article 53(1)(c) since 2 August 2025, so provenance records matter as of October 2026 [11].

Acceptance checklist for a screencast video request

Use this checklist to turn a vague "screencast dataset" request into something a supplier can answer and you can verify.

Illustrative example: invented to show structure; it does not describe an available dataset.

AreaWhat to specifyHow to verify on a sample
ScopeApplications, versions, task types, languages, total hoursVersion tags present on every clip
LegibilityNative resolution, scale factor, codec, minimum bitrateOCR legibility rate on sampled frames
Text layersTranscript source, format (WebVTT/SRT), word timestamps, chaptersMeasured narration-action offset
MetadataOS, browser, locale, theme, tenant type, capture dateSchema validation of sidecars
RedactionFields, method, audio handling, sidecar logsManual review of redacted frames
RightsOwnership, narrator and customer consent basis, third-party softwareDiligence file per dataset
SplitsHold-out by app version and by sourceNo clip or narrator overlap across splits

How SourceX handles screencast video requests

SourceX sources operational datasets from US companies on request, including support histories, engineering records, documents and new recordings of hands-on work, and manages the licensing process. Buyers describe the data, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Nothing is held in stock and a request does not guarantee a match. It does not source scraped web content, such as tutorials pulled from public video platforms.

Each dataset is rights-reviewed for ownership and consents, and personal details such as names, emails, phones and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Health records require HIPAA de-identification. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval.

Related pages: license screen recordings, workflow and screen activity data, training data for computer-use agents, computer-use data and the video data hub. You can describe the screencast data you need using the checklist above.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Request software tutorial video data

Describe the applications, versions, narration and redaction you need, and SourceX will look for US companies that hold matching recordings. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, and nothing is contracted until a supplier agrees. Start a buyer request.

Sources

  1. CVF / CVPR 2020, "Screencast Tutorial Video Understanding" (2020). https://openaccess.thecvf.com/content_CVPR_2020/papers/Li_Screencast_Tutorial_Video_Understanding_CVPR_2020_paper.pdf
  2. arXiv, "CodeSCAN: ScreenCast ANalysis for Video Programming Tutorials" (2024). https://arxiv.org/pdf/2409.18556
  3. Papers with Code, "TutorialVQA". https://cs.paperswithcode.com/dataset/tutorialvqa
  4. IJCAI, "IJCAI 2020 proceedings paper on screencast tutorial question answering" (2020). https://www.ijcai.org/Proceedings/2020/0148.pdf
  5. arXiv, "Code2Video: A Code-centric Paradigm for Educational Video Generation" (2025). https://arxiv.org/pdf/2510.01174
  6. arXiv (NeurIPS 2022), "Video PreTraining (VPT): Learning to Act by Watching Unlabeled Online Videos" (2022). https://arxiv.org/pdf/2206.11795
  7. Atlassian (Loom support), "Record for AI agents". https://support.atlassian.com/loom/docs/record-for-agent/
  8. arXiv (Longpre et al.), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  9. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  10. U.S. Department of Health and Human Services, OCR, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  11. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data