Private evaluation sets built from real business work
To evaluate models and agents on real work without contamination, you need private, held-out tasks built from business records that were never published: the context available at a decision point, what a skilled person did next, and the outcome or expert grade that followed. SourceX sources such records, including support cases, engineering histories, workflow trajectories, QA-scored work and contract redlines, from established businesses, and licenses them for evaluation use with personal data removed or pseudonymized.
Dataset types to start with
- Customer support ticket datasets
Resolved cases with disposition codes, reopen counts, CSAT and QA scores become graded items: the model sees the case as it arrived and is scored against what actually resolved it and how reviewers rated the handling.
- Software engineering histories (issues, PRs, reviews)
Issues linked to merged pull requests and CI runs in private repositories make SWE-bench-style tasks that have never been public: the repository before the change plus the issue is the prompt, and the tests added with the real fix are the grader.
- Enterprise workflow and task execution histories
Task execution histories record the context, tool actions, handoffs and outcome of each task, so an agent can be scored on its whole trajectory, including whether it ran the checks and approvals the business requires, not only on its final answer.
- Human feedback and QA-scored work
Reviewer scorecards, corrections and approvals supply grades for open-ended work and a reference for calibrating LLM judges: compare a judge's scores with how human reviewers graded the same items before trusting it at scale.
- Contract negotiation and redline histories
Negotiation histories from first draft to executed version give expert tasks with a known end state — which clauses changed, which positions held, how the playbook was applied — for work where no single answer string is correct.
Why this data is hard to get
Public benchmarks leak into training data
Items published on the web end up in pretraining crawls and in fine-tuning sets assembled from public sources, so a high score can reflect memorization rather than ability. Records that have only ever lived inside a company's systems have not been exposed that way.
Real outcomes exist only inside the business
Public benchmarks are written to have clean answers. Real work has outcomes instead — a case reopened, a change reverted, a clause conceded, an invoice rejected — and they are recorded in help desks, repositories and approval systems that are never published.
Expert grading is slow to recreate
Grading open-ended work in law, finance or support needs domain experts and a rubric they agree on. Businesses already did much of that grading through QA reviews, approvals and code review, but the grades sit in internal tools next to the work they judged.
Freshness needs a continuing source
A benchmark is fixed when it is released, while the work it stands for keeps changing: new products, revised policies, refactored code. Keeping an eval set current takes a steady supply of recent records of the same kind, which only operating businesses produce.
From a business record to a graded eval item
An eval item built from real work has four parts: the context as it stood at a decision point, the task posed to the model, a reference outcome, and a grading method. The decision point matters most, because anything written after it leaks the answer.
- Support: the customer's first message, the account context and the policies in force on that date, not the resolution note written days later. Grade the reply or resolution path against the disposition and the QA review.
- Code: the repository at the commit before the fix, plus the linked issue. Grade the model's patch with the tests the real fix added, and check that existing tests still pass.
- Workflow: the task input, replayed in a sandbox with the same tools. Grade the result and the required steps, such as an approval before a refund.
- Contracts: the counterparty's draft and the negotiation playbook. Experts grade the model's redline against the executed version and a rubric, not by string match.
Outcomes need careful reading. A ticket closed because the customer stopped replying was not resolved, and a merged change that was later reverted was not correct. Filter or relabel these cases before they become ground truth, and document how.
How held-out items leak into training
Held-out items rarely leak into training through one dramatic mistake; they leak through near-copies, public relatives, evaluation traffic and the people who work with them. Near-copies are the first route: templated contracts, reused support macros and forked repositories put almost the same text on both sides of a split, so deduplicate across it before freezing the set. Public relatives are the second: a private ticket about a widely reported outage, or a contract built on a published template, has cousins on the open web, so search distinctive strings before trusting an item as unseen.
Evaluation traffic is the third. Sending items to a model served by a third party can expose them if the provider retains or trains on inputs, so check its data terms first. Embedding a unique canary string in eval files makes copies detectable if they later surface in a crawl or a training mix. People are a fourth route: engineers who read failing items can end up writing training examples that resemble them, so log who has access and treat heavily inspected items as candidates for retirement.
Freshness and rotation
A held-out set loses value with use. Each run that informs a design decision fits the model a little more closely to the set, published examples leak it, and the domain moves on as products, policies and code bases change. Date every item and report results on recent work separately: that shows whether a model keeps up with change, and a model that scores far better on older items than on newer ones may have seen the older ones before.
Version the set like code. Keep a frozen core of anchor items that stays in every version, so scores remain comparable across versions; add dated slices of recent work; and retire items once they have been exposed. When a new version ships, score the current production model on both versions, so trend lines survive the change. Each rotation needs recent records, so agree in the license whether later snapshots from the same sources are in scope.
What good data looks like
- Items drawn from records never published online, with provenance per item — source system, date and the license terms that apply.
- Each item frozen at a decision point, with anything created afterwards, such as resolution notes, final drafts or follow-up fixes, kept out of the prompt.
- A reference outcome or grade for every item, such as passing tests, a disposition, an executed clause or a QA score, with the rubric version recorded.
- A double-graded subset with measured agreement between graders, so you know how much of a score difference is noise.
- Items tagged by task type, difficulty, date and source, so results can be broken down and compared with production traffic.
- Evaluation-only permitted use written into the license, and item fingerprints you can filter every training corpus against.
Questions buyers ask
What makes an evaluation set private and held-out?
A private, held-out evaluation set is never published and never used for training, so a model's score reflects ability rather than recall of the test. Built from business records, each item is a real task frozen at a decision point, with the outcome or expert grade that followed. Access is limited to the people running evaluations, and the license can restrict the data to evaluation use.
How do I keep evaluation data out of training?
Separate it by contract, by storage and in the pipeline. The license can restrict the items to evaluation and prohibit training on them. Store them apart from training data, with restricted access. Fingerprint each item, for example with content hashes and n-gram signatures, filter every training corpus against those fingerprints, and never use eval items as seeds or examples for generating synthetic training data.
How often should a private eval set be refreshed?
Refresh it when it stops telling models apart: when scores saturate, when prompts or models have been tuned against it, when items have been exposed, or when the work it describes has changed. A fixed cadence tied to your release cycle also works, with an extra refresh whenever items leak. Agree a refresh path with the data owner in the license, since new records depend on partners continuing to supply them.
Can the same licensed data be used for training and evaluation?
Yes, if the license allows both uses, but never the same items. A common pattern trains on a source's older records and evaluates on its newest, which also tests whether a model copes with change. State in the license which portion is evaluation-only, so it is not folded into a later training run by someone who never saw the original agreement.
How are real business tasks graded when there is no single right answer?
With the outcome the business recorded plus a rubric. Code tasks can be graded by the project's own tests, and support and workflow tasks by disposition, reopen and handoff data. Open-ended outputs such as replies or redlines are scored by domain experts against a rubric, or by an LLM judge whose scores have been checked against those expert grades. QA scorecards already in the record give a starting rubric.
Can I license an evaluation set exclusively?
Sometimes. In some programs, exclusivity can be agreed for a dataset snapshot or a permitted use, and it matters more for evaluation than for training: items licensed to several buyers are not private to any of them, and a leak by one licensee contaminates the set for everyone. Ask whether the items, or the records behind them, have been licensed to others.
Tell us what you are building
Describe the model or agent, the tasks it must handle, and the volume, format and permitted use you need. SourceX will match it to partner data.
Updated 3 October 2026.