Skip to content

Fine-tuning and post-training data

Expert demonstration data for LLMs: commissioned writing or licensed work product

Quick answer

Expert demonstration data for LLMs is a set of prompts paired with complete responses written or approved by credentialed practitioners, used as supervised fine-tuning (SFT) targets. Buyers get it in two ways: commission experts to write new responses to chosen prompts, or license existing work product, such as memos, reports and answered requests, and rebuild the request each one answered. Commissioning controls coverage. Licensing captures real professional judgment with less new expert time, but raises client confidentiality and third-party rights questions.

By SourceX Editorial · Updated

This page covers full expert-written responses. Expert labels on existing data are a different purchase (expert annotations and labels). The general SFT routes are compared in how to source a supervised fine-tuning dataset, part of the fine-tuning and post-training data guide.

What an expert demonstration teaches, and what it does not

An expert demonstration teaches a model how a practitioner frames, sequences and qualifies an answer. It is a poor way to add facts the base model lacks, which decides what you are paying experts for.

InstructGPT's SFT stage fine-tuned GPT-3 on demonstrations written by contracted labelers before reinforcement learning from human feedback, and human evaluators preferred the resulting outputs to those of the far larger GPT-3 [1]. LIMA fine-tuned a 65B-parameter LLaMA model on 1,000 curated prompt-response pairs with no reinforcement learning. Its authors argue that almost all knowledge is learned in pre-training and that a small, high-quality set mainly teaches output format, and note that such curation is labor-intensive [2]. If the model gets domain facts wrong, a domain corpus for continued pre-training or retrieval is the better fix.

What experts add is judgment made visible: which issue comes first, which missing fact to ask for, when to caveat or refer, and how a conclusion is supported. State whether you want that reasoning written out or only the final answer. For the wider category, see human-generated data for AI and the human data glossary entry.

Neighboring data types are specified and priced differently, so order them separately: domain-expert preference data ranks responses, expert rubric data for RL scores them, reasoning trace datasets expose derivations, and domain-expert raters judge outputs for evaluation rather than training.

"Expert" in a dataset name does not prove a human expert wrote each response, since a model prompted with an expert persona writes in the same voice; review such data under due diligence for synthetic fine-tuning data.

Commissioned expert writing vs licensed expert work product

Commission when you need coverage of specific tasks, formats or edge cases that nobody has written up; license when you need the messy inputs and judgment calls of real practice. The table adds the common hybrid, in which an expert edits licensed work product into a clean target.

Commissioned expert writingLicensed expert work productHybrid: expert-edited work product
PromptWritten by you or the vendor; tends to be tidy and fully specifiedRebuilt from the real request: email, intake form, ticket or request for information (RFI)Rebuilt, then checked by the editing expert
ResponseWritten to your style guide and length rulesWritten for a client or colleague, with house style, hedges and signature blocksOriginal reasoning kept; format normalized to your guide
Credential evidenceVendor's verification of each writerAuthor's role and license at the time of writing, from the holder's recordsOriginal author plus editor
Coverage controlHigh: you choose task families and rare casesLimited to what the business actually handledMedium: gaps filled with commissioned items
Main cost driverExpert hours per response plus review hoursSelection, prompt rebuilding, de-identification and rights reviewEditing plus licensing work
Rights paperAssignment or license from each writer and reviewerHolder's authority to license, client terms, third-party contentBoth sets
Typical failurePolished answers to questions no client asks; undisclosed chatbot draftingClient facts left in the text; answers overtaken by later rule changesEditors rewriting the judgment, not just the format

Related decisions: IP assignment vs license for commissioned data, contracting domain experts for post-training and what drives the cost of human SFT and preference data.

Where expert work product already exists

Professional services firms and in-house expert teams already produce request-and-answer pairs as a by-product of their work. The sourcing task is to find the record holding the question and the one holding the answer, then clear each domain's rights issue.

DomainWork product that answers a requestWhat becomes the promptMain issue before licensing
LegalResearch memos, advice letters, contract redlines (legal briefs and memos)The assigning lawyer's question plus the facts and documents providedPrivilege and client confidentiality; quoted third-party research content
Accounting and taxTechnical accounting memos, replies to client queriesThe client's question and the transaction documentsClient financial data; the rule in force on the memo date
Lending and insuranceCredit memos, underwriting rationales, coverage-position lettersA summary of the application or claim fileConsumer financial information under privacy rules
Clinical operationsPrior-authorization letters, utilization-review rationales, consult repliesThe referral or request with relevant historyProtected health information
Engineering and constructionAnswered RFIs, design-review comments, incident postmortemsThe RFI or change request with the referenced drawingsOwner and client confidentiality; trade secrets
Research and advisoryClient deliverables and analyst notes (research deliverables)The engagement question and the data providedClient ownership of deliverables

Industry detail is in legal LLM fine-tuning data, financial services fine-tuning data and healthcare LLM fine-tuning data.

Rebuilding the request an expert was actually answering

Work product rarely contains a clean prompt, so each licensed record needs the original request rebuilt from what the expert had at the time, and only that. Otherwise the model learns to answer questions that already contain their answers.

  1. Find the trigger. Identify the email, intake form, ticket or RFI that started the work, and its date.
  2. Assemble the context as of that date. Include the documents and messages the expert had; exclude later facts, the outcome, and any revision that superseded the answer.
  3. Write the prompt in the requester's voice. Keep the gaps of the real request; a response that asks for a missing fact is a valuable target.
  4. Strip the wrapper. Remove letterhead, signature blocks, boilerplate disclaimers and billing references, but keep caveats that are part of the advice.
  5. Record the rebuild. Mark the prompt as reconstructed, name the role that rebuilt it, and keep a pointer to the source record.
  6. Tag the as-of date and jurisdiction. Professional answers expire when rules, codes or standards change.

The general record-to-pair mapping is in turning business records into instruction-response pairs. The record below shows the metadata to require on each licensed demonstration.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "id": "xd-eng-rfi-00417",
  "messages": [
    {"role": "user", "content": "RFI from [CONTRACTOR_1], level 3: MEP wants to core a sleeve through the web of beam [BEAM_1], near the column, for the condensate line. Sketch attached. OK to proceed?"},
    {"role": "assistant", "content": "Not at the location shown. The core falls inside the end zone where General Note [N_1] prohibits web penetrations. Two options: (1) move the sleeve to the middle third of the span, clear of other openings, and confirm the new location on a marked-up sketch; or (2) keep the location and submit a reinforcement detail for review. Do not core until we return the marked-up sketch."}
  ],
  "context_refs": ["sketch_[RFI_1].pdf", "framing_plan_L3_rev[R].pdf"],
  "meta": {
    "origin": "licensed_work_product",
    "prompt_origin": "reconstructed_from_rfi_log",
    "reconstructed_by_role": "project engineer",
    "author_role_at_writing": "structural engineer of record",
    "author_credential": {"type": "PE", "verified_by": "holder_hr_record_and_state_board_lookup", "active_on_as_of_date": true},
    "author_pseudonym": "auth-112",
    "as_of_date": "2025-04-17",
    "jurisdiction": "US-[STATE]",
    "superseded_later": false,
    "second_review": {"reviewer_credential": "PE", "verdict": "pass", "criteria_failed": []},
    "ai_assistance": "none_declared",
    "confidentiality_screen": "client_owner_project_names_replaced",
    "rights_ref": "LIC-[ID]"
  }
}

Verifying the experts behind each response

Verify, per record, the credential that matches each task family on the date the response was written; a vendor's statement that its team is "PhD-level" is not evidence. A degree is a proxy for expertise and often the wrong one: a practicing claims adjuster may write a better coverage explanation than an economics PhD.

  • Credential matched to task. Map each task family to the license, certification or job role that qualifies someone to answer it in practice.
  • Independent check. Confirm licenses against the issuing body's public lookup, active on the writing date. For work product, use the holder's records of the author's role at the time.
  • Per-record attribution. A pseudonymous author ID on every record lets you measure quality per author and cap any one author's share.
  • AI-assistance disclosure. Undisclosed chatbot drafting turns "expert-written" into distillation, so require a per-record declaration (verifying human authorship).
  • Prior confidentiality obligations. Commissioned writers declare that responses draw on no current or former client's confidential matters.
  • Documented writer pool. Data Statements, a documentation schema for language datasets, includes fields such as speaker and annotator demographic [3]; record credential type, specialty, years in practice and jurisdiction the same way.

Checking accuracy with a second expert

Have an independent expert of equal or higher credential review a stated sample of every task family against a written rubric, and accept the data on those results. A credential qualifies a writer; it does not verify a specific answer.

A workable rubric scores each response on five points: correct as of its date and jurisdiction; issues spotted versus issues missed; appropriate caveats, questions back and referrals; a conclusion the requester could act on; and no residual client or personal information. Measure how consistently reviewers apply it. Krippendorff's alpha is a reliability coefficient for agreement among raters assigning values to the same items, where 1 means perfect reliability and 0 means agreement no better than chance [4]. Low agreement often means the rubric is ambiguous (designing evaluation rubrics with domain experts).

Treat "gold demonstrations" as a claim to verify. Northcutt and colleagues estimated an average label error rate of at least 3.3% across the test sets of 10 widely used datasets [5]. Ask for the review record behind any gold set and run an annotation quality audit on a sample.

Licensed work product allows one check that commissioned writing cannot: compare the response with what happened next, and drop an RFI answer superseded by a revised drawing. Then run a sample ablation (evaluating a fine-tuning dataset before you buy it).

Confidentiality, privilege and third-party rights in work product

Expert work product was written for a client or employer, so licensing it depends on the holder's authority over the record, the client's terms, any privilege, and third-party content inside the answer. Each needs evidence before delivery.

  • Client confidentiality and privilege. Professional advice carries client facts and, for legal work, may be privileged. De-identification must cover matter names, counterparties and distinctive fact patterns, not only personal identifiers.
  • Third-party content. Answers often quote or paraphrase research databases, standards and client documents. On 29 September 2026 the US Court of Appeals for the Third Circuit, in a precedential opinion, held the Westlaw headnotes at issue copyrightable and ROSS's use of Westlaw material to train a non-generative AI legal-research tool not fair use; commentators note a footnote distinguishing generative AI [6]. A licensed answer can therefore carry rights in third-party summaries it reuses.
  • Financial information. Under Regulation P, a recipient of nonpublic personal information from a nonaffiliated financial institution under one of the rule's exceptions may use it only to carry out the purpose for which it was received, whether or not the recipient is itself a financial institution [7].
  • Health information. HIPAA de-identification uses Safe Harbor, which removes 18 listed identifiers with no actual knowledge that the rest could identify the person, or Expert Determination [8] (Safe Harbor vs Expert Determination for AI training).
  • Memorization. SFT trains a model to reproduce its targets, and researchers have extracted hundreds of verbatim training sequences, including contact details, from GPT-2 [9]. A client name left in one memo is a disclosure risk, not a formatting issue.
  • Commissioned output. Collect a written assignment or license from every writer and reviewer, including subcontractors, and confirm that no writer's employment terms claim the output.

Record origin, author credential and rights reference per record, because disclosure duties draw on them. As of October 2026, general-purpose AI model providers covered by the EU AI Act must publish a sufficiently detailed training-content summary using the AI Office template under Article 53(1)(d) [10]. California's AB 2013 requires developers of generative AI systems offered to Californians to post documentation covering items such as whether training datasets include copyrighted material, whether they were purchased or licensed, whether they include personal information and whether synthetic data was used, first due by 1 January 2026 [11].

SourceX sources operational datasets from US companies; the kinds it sources include engineering records, documents, and finance and legal workflows. Datasets are sourced on request, not held in stock, so a request does not guarantee a match, and the supplying company approves every release.

Rights review checks that the business owns or may share the records and that required consents are in place; the license defines the records included, allowed uses, term and delivery. Names, emails, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked, but no method is perfect, so run your own client-name scan. You can describe the expert work product and credentials you need.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Red flags and a request checklist

These signs in an offer or sample usually mean rework or data you cannot use as described:

  • Credentials stated for the team, with no credential field on each record.
  • Uniform response length and structure across task families, which suggests templates or model drafting.
  • No as-of date or jurisdiction where rules change.
  • Reconstructed prompts that already state the response's conclusion.
  • Letterhead, matter numbers or client names in the sample.
  • No answer on whether client engagement terms allow reuse.

A request should state:

  • The task families, with the credential and target share for each.
  • The route per family and the prompt reconstruction rules.
  • Response rules (writing guidelines for SFT demonstration writers).
  • The per-record metadata shown above, and the review protocol.
  • Volume sized by a learning curve (how much data you need to fine-tune an LLM).
  • License scope (fine-tuning-only data licenses).

General structure is in how to write a data request for suppliers.

Sourcing expert demonstrations from real professional work?

Describe the task families, the credentials that should stand behind each answer, the kinds of work product that could supply them, and the license scope you need. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Request expert work product through SourceX.

Sources

  1. Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Zhou et al., Meta AI and collaborators (arXiv:2305.11206; NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  3. University of Washington Tech Policy Lab, "Data Statements" (schema Version 2, 2021). https://techpolicylab.uw.edu/data-statements/
  4. Klaus Krippendorff, University of Pennsylvania Annenberg School for Communication, "Computing Krippendorff's Alpha-Reliability" (2011). https://www.asc.upenn.edu/sites/default/files/2021-03/Computing%20Krippendorff%27s%20Alpha-Reliability.pdf
  5. Northcutt, Athalye, Mueller (arXiv:2103.14749; NeurIPS 2021 Datasets and Benchmarks), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  6. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (precedential opinion)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  7. Consumer Financial Protection Bureau, "12 CFR 1016.11 - Limits on redisclosure and reuse of information (Regulation P)". https://www.consumerfinance.gov/rules-policy/regulations/1016/11/
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  9. Carlini et al. (USENIX Security 2021), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
  10. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  11. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data