Fine-tuning and post-training data
Contracting domain experts for post-training data: models, QA and IP
Quick answer
To hire domain experts for AI training data, pick an engagement model first (direct contractors, a vendor-managed team, an expert network or an on-demand platform), then write the contract around three things vendors gloss over: a present assignment of IP in every output, confidentiality that covers your prompts and model outputs, and a QA design with calibration, peer review and adjudication. Credential checks and worker classification come next. Prefer pricing per accepted item over hours logged.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Which engagement model fits your post-training program
The right model depends on how narrow the specialty is, how long you need it and how much of the QA you want to own. A CleverX comparison frames three primary models: expert networks, specialist annotation vendors, and on-demand verified platforms [2]. Marketplaces such as OpenTrain publish a job-post model with a flat platform fee, AI interviews and skill tests, and state that they do not host the training data itself [3]. That last point matters: if the platform never touches the data, your own tooling, access controls and audit logs carry the whole confidentiality burden.
| Model | You control | Vendor controls | Typical fit | Main failure mode |
|---|---|---|---|---|
| Direct contractors | Guidelines, QA, tooling, IP paper | Nothing | Long-running niche (e.g., tax law SFT) | Classification risk, recruiting load |
| Vendor-managed team | Spec and acceptance criteria | Recruiting, QA, payroll | Volume RLHF comparisons | Opaque rater pool, QA drift |
| Expert network | Questions, call scope | Introductions, compliance screens | Short consultations, rubric design | Not built for sustained labeling |
| On-demand platform | Tasks and tests | Matching, payments | Fast pilots, mixed skills | Thin vetting, churn |
Many teams mix models: an expert network to recruit a few senior reviewers who write the rubric, then a managed team for volume, with the senior reviewers adjudicating. If your deliverable is preference pairs, see domain-expert preference data; if it is written demonstrations, see expert demonstration data for SFT.
How to secure IP in expert-created outputs
Get a written, present-tense assignment of all rights in every output, because paying an expert does not by itself make you the owner. Under 17 U.S.C. 201, copyright vests initially in the author; the employer or commissioning party is treated as the author only for a "work made for hire," and ownership can otherwise move by a signed written transfer [1]. For independent contractors, the definition of a commissioned work made for hire in 17 U.S.C. 101 is limited to nine categories of work and requires a signed written agreement, so counsel usually treats an express assignment as the reliable route.
Most post-training outputs, such as SFT responses, preference rationales and reward rubrics, do not fit those categories neatly. Use "work made for hire, and to the extent not, hereby assigns" language, plus a waiver or license-back for moral rights where the expert sits outside the US. With a vendor, require the vendor to hold signed assignments from every individual and to assign onward to you; ask for the template, not a summary. For the wider trade-off between owning and licensing commissioned data, read IP assignment vs license in data collection contracts.
Scope the assignment to cover:
- Responses, edits, rankings, scores, free-text rationales and rubrics.
- Derived artifacts: guidelines the expert drafts, error taxonomies, adjudication notes.
- Pre-existing material the expert pastes in: either exclude it, or require a warranty that the expert has the right to contribute it.
Why expert confidentiality terms must cover prompts and model outputs
Your prompts, unreleased model outputs and rubrics reveal your roadmap, so the NDA must name them explicitly as confidential information, not just "data." Standard annotation NDAs often protect the client's customer data but are silent on model checkpoints' behavior, eval sets and reward criteria. Add clauses that bar experts from pasting tasks into third-party chat assistants, reusing your prompts for other labs, or publishing examples in portfolios.
The second exposure is the expert's employer. A practicing physician, lawyer or engineer may be bound by an employment agreement that restricts outside work or the use of employer knowledge, or by professional confidentiality toward clients and patients. Require a representation that outside work is permitted and that no employer, client or patient information will be used; reject outputs that contain identifiable case details. If any task touches health records, the source material has to meet HIPAA de-identification by Expert Determination or Safe Harbor before experts see it [6].
Make your own commitments consistent. If you tell customers their data will not train models, do not route customer conversations to an expert vendor for SFT; the FTC has warned that breaking such commitments can create liability [7].
How to verify expert credentials before work starts
Verify the credential at its source and test the skill on your tasks, because a resume screen proves neither. Vendors should be able to describe their verification workflow step by step, which is the question CleverX recommends buyers ask [2]. Primary-source checks include state bar and medical license lookups, CPA board registries and degree verification services.
Then gate on performance: a paid qualification set of 20 to 50 items with known answers from your senior reviewers, scored against a threshold you set in the contract. Keep a hidden gold set in production so that accuracy is measured continuously, not once. The full method is in verifying domain-expert annotators, and the evaluation-side version is in sourcing domain-expert raters.
What QA layers expert post-training data needs
Expert data needs layered QA because experts disagree, and noisy labels distort the reward models trained on them. Research on RLHF reward modeling describes both incorrect and ambiguous preference labels, and one summary reports inter-annotator agreement of roughly 63% to 72% on general preference data [4]. Expertise narrows that gap only when guidelines define what "better" means in the domain.
Build four layers into the statement of work:
- Calibration: every new expert completes the same 10 to 20 items; a lead reviews disagreements and updates the guideline before production.
- Peer review: a sampled share of items (for example 10% to 20%) goes to a second expert blind to the first answer.
- Adjudication: a senior reviewer resolves disagreements and records a reason code, which feeds guideline revisions.
- Acceptance testing: you, not the vendor, sample each batch against written acceptance criteria and reject below threshold.
Small, carefully curated sets can carry real weight: LIMA fine-tuned a 65B model on 1,000 curated prompt-response pairs, while noting that such curation is labor-intensive [5]. That is the argument for paying for adjudication rather than for more raw volume. For deciding how much to buy, see allocating a post-training data budget.
Expert data contract checklist
Use this checklist to compare vendor paper or to draft a direct contractor agreement. Each row is a term you should be able to point to in the signed document.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Term | What good looks like | Red flag |
|---|---|---|
| IP | "Hereby assigns" all outputs and derivatives; vendor holds individual assignments | "Client may use deliverables" |
| Confidentiality | Prompts, model outputs, rubrics and eval items named; no third-party AI tools | Generic mutual NDA only |
| Outside-work rep | Expert confirms employer permits the work; no employer or client data | Silence |
| Credentials | Primary-source verification plus paid qualification test | "Vetted experts" with no method |
| QA | Calibration, blind peer review rate, adjudication with reason codes, buyer acceptance sampling | Vendor self-reports accuracy |
| Pricing unit | Per accepted item or per adjudicated comparison | Hourly with no output floor |
| Rework | Free rework or credit for rejected batches | Rework billed as new work |
| Data handling | Work done in your tool or a logged workspace; deletion certificate at end | Files exchanged by email |
| Rater metadata | Pseudonymous rater ID, credential type, years of practice per item | No rater-level fields |
| Classification | Who is the employer of record; local contractor rules addressed | "Not our concern" |
A per-item record that supports audits might carry fields such as item_id, rater_id (pseudonymous), credential_type, guideline_version, qa_status (accepted, adjudicated, rejected), adjudicator_id and reason_code.
How worker classification and payment terms change by country
Classification follows the actual working relationship, not the contract label, and the tests differ by country and change over time. In the US, the federal FLSA test, IRS common-law rules and state tests such as California's ABC test can reach different answers for the same expert, and the federal rule has been revisited repeatedly, so confirm current status with employment counsel as of October 2026. Detailed guidelines, mandatory hours and exclusive engagement all push toward employee status.
Outside the US, an employer-of-record or the vendor's own entity usually absorbs local employment, tax and payment obligations, which is one practical reason labs use vendors for multi-country pools. Whatever the model, specify currency, payment timing on acceptance rather than submission, and who bears transfer fees. Classification also feeds back into IP: an employee's work within the scope of employment is work made for hire, while a worker who is legally a contractor generally owns their output unless it is assigned in writing [1].
When licensing existing expert records beats commissioning new work
Licensing existing records can be faster when the expertise you need already exists in a company's operational history. Resolved engineering tickets, legal and finance workflows, and support histories contain expert judgment captured in real work, and they can be turned into instruction-response pairs (see turning business records into instruction-response pairs). Commissioning remains the better route for preference rankings on your own model's outputs and for rubrics tied to your policy.
SourceX sources operational datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery; personal details are removed or replaced before delivery, and no method is perfect. Buyers who want to compare licensed records with a commissioned program can describe the data they need. For the broader procurement process, see how to procure enterprise training data and the fine-tuning and post-training data hub.
Hire domain experts or license expert records for AI training data
If your post-training plan needs expert judgment that already exists in business records, describe the data rather than the companies, and SourceX looks for US businesses that hold it. Nothing is contracted until a supplier agrees, every release is approved by the supplying company, and a request does not guarantee a match. Start at SourceX for buyers.
Sources
- U.S. Government Publishing Office (govinfo), "17 U.S.C. 201 - Ownership of copyright (U.S. Code, 2024 edition)" (2024). https://www.govinfo.gov/content/pkg/USCODE-2024-title17/html/USCODE-2024-title17-chap2-sec201.htm
- CleverX, "Domain experts for AI training by industry". https://cleverx.com/blog/domain-experts-for-ai-training-by-industry
- OpenTrain AI, "How it works". https://opentrain.ai/how-it-works/
- alphaXiv, "Secrets of RLHF in Large Language Models Part II: Reward Modeling (arXiv:2401.06080) overview" (2024). https://www.alphaxiv.org/overview/2401.06080
- Zhou et al., arXiv:2305.11206 (NeurIPS 2023), "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.