Skip to content

Code and software engineering data

Developer Session Recordings for Training Coding Agents

Quick answer

A coding agent trajectories dataset built from human developer sessions captures how an engineer actually solves a task: editor events, terminal commands and outputs, documentation lookups, screen video, optional think-aloud narration, and the final diff with its tests and review. Buy it when model-generated trajectories leave gaps in your agent's behavior. Specify capture channels, task source, outcome linkage, consent and redaction before you ask anyone to record.

By SourceX Editorial · Updated

Why human sessions fill gaps that synthetic trajectories leave

Human sessions supply the behaviors that rejection-sampled agent runs underrepresent: real exploration, recovery from wrong hypotheses, and work beyond bug fixing. Public trajectory corpora are mostly model-generated. Together AI's CoderForge Preview, for example, holds about 258k test-verified trajectories produced by Qwen3-Coder-480B with rejection sampling [3], and its dataset card notes a skew toward bug fixing [4]. Benchmarks built from GitHub issues, such as SWE-bench with 2,294 problems across 12 Python repositories, inherit the same issue-to-patch framing [5].

Synthetic data is cheap, which is why it dominates. The Synatra authors report a synthetic web-agent demonstration costing about 3% of a human-annotated one [7]. Yet small human sets still carry weight: one computer-use study started from 312 human-annotated trajectories and augmented them with a frontier model [8]. InstructGPT used supervised fine-tuning on human demonstrations before preference training [6], and a similar approach is a reasonable default here: a modest, well-specified human set can anchor the distribution that synthetic data then scales.

For definitions, see agent trajectory and process supervision. For the outcome side (commits, PRs, reviews), see coding agent training data from private repositories.

Capture channels and what each one costs

Each capture channel adds training signal and a different privacy and storage burden, so pick channels by the policy you plan to train. A text-only agent that acts through tools gets most value from structured editor and terminal events; a computer-use agent needs screen video aligned to those events.

ChannelTypical formatTraining valueMain risk
Editor events (open, edit, navigate, LSP queries, test runs)JSONL event stream from an IDE extension (VS Code, JetBrains)Action sequences for tool-using agents; file navigation patternsCaptures unsaved buffers, including pasted secrets
Terminal input and outputasciinema cast or PTY log with exit codesShell command data for DevOps and CLI agents; error recoveryEnvironment variables, tokens, internal hostnames in output
Browser and documentation lookupsURL and timestamp log, optional page snapshotsShows when experts consult docs or issuesThird-party content and internal wiki pages
Screen videoMP4/H.264 at fixed frame rate with monotonic timestampsGrounding for computer-use agentsCustomer data, chat pop-ups, email previews on screen
Think-aloud narrationWAV/FLAC plus time-aligned transcriptReasoning traces for process supervisionVoiceprints can count as biometric identifiers under some state laws; off-topic personal remarks
Outcome bundleFinal diff, test results, CI status, review commentsOutcome labels for scoring the processReviewer names and internal ticket text

Synchronization is the most common failure. If the editor log, PTY log and video use different clocks, action labels drift by seconds and become useless for step-level supervision. Require one session clock, recorded offsets per channel, and a drift check on a sample. Converting raw video into labeled steps is covered in turning screen recordings into action-labeled trajectories.

Task sources: real backlog work or commissioned tasks

The task source decides both realism and rights complexity, and most buyers end up mixing the two. Real backlog tickets in a company's own repositories give the most realistic distribution of ambiguity, legacy code and cross-team dependencies. They also put proprietary code, customer identifiers and third-party software on screen, which triggers a full ownership review; see code ownership due diligence.

Commissioned tasks on licensed repositories are easier to control. You can balance task types (feature work, refactors, migrations, flaky-test triage, incident response) and pin the repository snapshot so sessions are reproducible. The tradeoff is artificiality: engineers solving a staged task often skip the exploration that makes real sessions valuable. If you need executable environments alongside recordings, specify them as in SWE task environments with tests.

A recording without its outcome cannot be scored, so require each session to link to the final diff, the test results and the review decision. The common convention for software agent trajectories is a binary success label based on whether the patch passes tests, with a stricter variant requiring manual review that the change matches the intended fix [1]. Apply the same labels to human sessions so they sit on one scale with your agent's runs.

Keep failed and abandoned sessions. Agent research treats failures as first-class data: one 2026 analysis of CLI coding agents studied 1,794 trajectories, of which 1,184 failed [2]. Human dead ends, reverted edits and abandoned approaches are useful negative and recovery trajectories. Ask suppliers to report time-per-task, abandon reason and the point where the engineer changed strategy, rather than filtering to clean successes.

Recording engineers at work requires documented consent and, in some workplaces, consultation before the first session is captured. Ask whether the employer's monitoring policy covers recording for third-party AI training, whether a works council or union agreement applies to any participant, and whether narration audio triggers state biometric or recording-consent rules. Treat each as a gate, not a cleanup step.

Redaction must cover every channel, not only the transcript. Secrets appear in terminal output and unsaved editor buffers; customer records appear in screen video and log tails. Scan event streams with a secrets detector, blur or mask video regions, and check a sample by hand. The detailed method is on PII redaction for screen recordings and agent trajectories, and rights in third-party software visible on screen are covered in screen recording third-party content rights.

Session record schema to include in your request

A shared schema prevents most delivery disputes, so attach one to the request. Store events as JSON Lines, which requires UTF-8 without a byte order mark and one valid JSON value per line [9]. Document dataset-level fields (consent basis, redaction method, labeling process) in a machine-readable card; Croissant-RAI is one vocabulary built for this [10].

Illustrative example: invented to show structure; it does not describe an available dataset.

{"session_id": "s-000412", "task": {"source": "backlog", "type": "dependency_upgrade", "repo_snapshot": "sha256:9f2c...", "ticket_text_redacted": true},
 "engineer": {"pseudonym": "eng-17", "years_experience_band": "5-10", "consent_ref": "c-2026-0412"},
 "clock": {"base": "monotonic_ms", "offsets_ms": {"editor": 0, "pty": 12, "video": -40}},
 "events_uri": "events/s-000412.jsonl", "pty_uri": "pty/s-000412.cast", "video_uri": "video/s-000412.mp4", "narration_uri": null,
 "outcome": {"diff_uri": "diff/s-000412.patch", "tests_passed": true, "ci_status": "green", "review": "approved_with_changes", "abandoned": false},
 "metrics": {"wall_time_s": 3820, "strategy_changes": 2, "reverted_edits": 5},
 "redaction": {"method": "secrets_scan+video_mask", "sample_checked": true}}

Buyer checklist before requesting a sample:

  • Capture channels listed, with required formats and one shared clock.
  • Task mix by type and source (backlog vs commissioned), with repository snapshot hashes.
  • Outcome linkage: diff, tests, CI, review decision, abandon flag.
  • Consent reference per engineer, and works-council or union status.
  • Redaction method per channel and the size of the manually checked sample.
  • Allowed uses written out: imitation training, process reward models, held-out trajectory eval.

To check whether a delivered sample matches the spec, follow evaluating a code dataset sample. If you plan to hold sessions out for scoring, keep them separate from training data as described in held-out SWE eval sets from private repositories.

How SourceX handles developer session data requests

SourceX sources operational datasets from US companies, including engineering records and new recordings of hands-on work, and manages the licensing process for AI teams wherever they are based. Data is sourced on request rather than held in stock, so a request does not guarantee a match. Every release is approved by the supplying company, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. You can describe the session data you need on the buyers page.

Personal details such as names, emails and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Delivery runs through private, access-controlled workflows only after an executed agreement and supplier approval. For neighboring data types, start from the code and software engineering data hub, enterprise workflow datasets and agent trajectories, or CI build failure logs.

Request human coding session data for your agent

Describe the tasks, capture channels and outcome labels you need, not the companies you hope to hear from. SourceX looks for US businesses that hold matching data, reviews data and licensing permissions, and agrees pricing and allowed uses in a license before anything is contracted. Start a coding agent data request.

Sources

  1. arXiv, "Understanding Software Engineering Agents: A Study of Thought-Action-Result Trajectories" (2025). https://arxiv.org/pdf/2506.18824
  2. arXiv, "Failure as a Process: An Anatomy of CLI Coding Agent Trajectories" (2026). https://arxiv.org/pdf/2607.09510
  3. Together AI, "CoderForge Preview". https://together.ai/blog/coderforge-preview
  4. Hugging Face, "CoderForge-Preview dataset card". https://huggingface.co/datasets/kshitijthakkar/CoderForge-Preview/blob/main/README.md
  5. arXiv (Princeton, UChicago), "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" (2023). https://arxiv.org/pdf/2310.06770
  6. arXiv (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  7. arXiv (CMU, Amazon AWS AI), "Synatra: Turning Indirect Knowledge into Direct Demonstrations for Digital Agents at Scale" (2024). https://arxiv.org/pdf/2409.15637
  8. arXiv (Shanghai Jiao Tong University, GAIR), "Efficient Agent Training for Computer Use" (2025). https://arxiv.org/pdf/2505.13909
  9. jsonlines.org, "JSON Lines". https://jsonlines.org/
  10. arXiv (MLCommons Croissant RAI task force), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data