Recruiting and hiring workflow datasets for AI training
A recruitment dataset is a de-identified history of real hiring processes: requisitions, job postings, applications, screening outcomes, interview scorecards, offers and, where linked, what happened after the hire. SourceX sources these histories from employers, staffing agencies and recruitment process outsourcing providers running applicant tracking systems such as Greenhouse, Workday or Bullhorn. Candidate identities are removed before delivery, and decision labels come with the context a buyer needs to test them for bias.
Dataset manifest
Sourced to your spec- What it is
- Requisitions, applications, screening and interview decisions, offers and post-hire outcomes
- Typical systems
- Greenhouse, Workday Recruiting, iCIMS, Bullhorn, Lever, SAP SuccessFactors
- Typical history
- Varies by partner; ATS migrations often cut older history short
- Modality
- Structured pipeline events, résumé and posting text, scorecards and offer documents
- Delivery formats
- Agreed per order; JSONL or Parquet tables keyed by requisition and candidate
- Preparation
- Candidate and interviewer identities removed; résumés generalized; demographics separated or excluded
- Licensing
- Permitted use agreed per order and may exclude automated candidate screening
- Availability
- Sourced to your spec from employers and staffing firms; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| requisition_id | string | Pseudonymous requisition key linking postings, applications, interviews, offers and placements. |
| requisition | object | Job family, level, employment type, work model, location at metro or region level, openings, open and close dates and close reason. |
| posting | object | The job description as posted, with each revision and the date it went live. |
| candidate | object | Random candidate token, application source, and résumé-derived skills and experience bands, with names and contact details removed. |
| application | object | De-identified résumé text, answers to screening questions and knockout results. |
| stage_history | array | Every move through the pipeline, from applied through screen, interviews and offer to hire or exit, timestamped, with the actor's role. |
| disposition | object | Where and why a candidate left the process — rejected, withdrew or requisition closed — and whether a person or an automated rule made the call. |
| screen_notes | array | Recruiter and sourcer notes from outreach and phone screens, de-identified. |
| interviews | array | Each interview with interviewer pseudonym and role, rubric version, competency scores, written feedback and recommendation, plus the debrief decision. |
| offer | object | Offer and acceptance dates, pay expressed as position in band, counteroffers and the reason for any decline. |
| placement | object | For staffing agencies, the job order, submittal, assignment dates, extensions and how the assignment ended, with bill and pay rates banded. |
| post_hire | object | Retention at agreed checkpoints and, where licensed, a first performance review band. |
| self_id | object | Voluntary self-identified demographics, held in a separate table or aggregated for adverse impact testing, or excluded. |
Example record
{
"requisition_id": "req_3a9d10",
"requisition": { "job_family": "customer_success", "level": "IC2", "employment": "full_time",
"work_model": "hybrid", "location": "[METRO_AREA]", "openings": 2,
"opened": "2023-02-06", "closed": "2023-04-14", "close_reason": "filled" },
"posting": { "versions": [
{ "v": 1, "live": "2023-02-07", "text_ref": "jd_3a9d_v1" },
{ "v": 2, "live": "2023-02-21", "text_ref": "jd_3a9d_v2", "change": "degree requirement removed; pay range added" } ] },
"candidate": { "id": "cand_8f12c4", "source": "employee_referral", "experience_band": "3-5y",
"skills": ["saas_onboarding", "renewals", "crm_admin"], "resume_ref": "res_8f12_deid" },
"application": { "applied": "2023-02-09", "knockouts": { "work_authorization": "pass", "hybrid_ok": "pass" } },
"stage_history": [
{ "t": "2023-02-09T15:20:11Z", "stage": "applied" },
{ "t": "2023-02-13T10:02:45Z", "stage": "recruiter_screen", "by": "recruiter" },
{ "t": "2023-02-20T17:31:09Z", "stage": "hiring_manager_interview", "by": "recruiter" },
{ "t": "2023-03-03T12:10:52Z", "stage": "panel", "by": "hiring_manager" },
{ "t": "2023-03-15T09:44:30Z", "stage": "offer", "by": "hiring_manager" },
{ "t": "2023-03-21T16:05:00Z", "stage": "hired", "by": "recruiter" }
],
"screen_notes": [{ "role": "recruiter",
"text": "Runs renewals for a mid-market book at [EMPLOYER_1]. Wants hybrid. Pay ask near top of band." }],
"interviews": [
{ "stage": "hiring_manager_interview", "interviewer": "intv_r2", "role": "cs_manager", "rubric": "cs_ic2_v3",
"scores": { "customer_judgment": 4, "commercial_acumen": 3, "communication": 4 }, "recommend": "yes",
"feedback": "Strong on the churn-risk role play. Thin on expansion pricing; would need ramp time." },
{ "stage": "panel", "interviewer": "intv_k7", "role": "senior_csm", "rubric": "cs_ic2_v3",
"scores": { "customer_judgment": 5, "commercial_acumen": 3, "communication": 4 }, "recommend": "strong_yes" }
],
"debrief": { "decision": "offer", "rationale": "Panel aligned; commercial gap coachable in first quarter." },
"offer": { "extended": "2023-03-15", "band_position": 0.62, "counter": true, "final_band_position": 0.70,
"accepted": "2023-03-21", "start": "2023-04-10" },
"post_hire": { "retained_90d": true, "retained_12m": true, "first_review_band": "meets" },
"self_id": { "status": "held_separately", "audit_key": "aud_51e0" }
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Train recruiting coordination agents
Stage histories and recruiter notes show the work around a hire — screening calls, scheduling interview loops, chasing feedback, moving candidates on — with the time each step actually took.
Turn interview notes into structured evaluations
Written feedback paired with rubric scores teaches models to summarize interviews against defined competencies, and to flag feedback that gives a score without evidence.
Write and revise job postings
Posting revisions tied to the applicant pool and time to fill that followed show which changes widened or narrowed who applied.
Build fairness evaluation sets for hiring models
Real applicant pools with self-identified demographics held apart let you test whether a screening or matching model produces adverse impact before anyone deploys it.
Model staffing submittals and placements
Job orders, submittals and assignment outcomes show which candidates clients accepted and which placements ran their full term or ended early.
Use-case guides: Enterprise and computer-use agents, Private evaluation sets
What makes this data valuable
The whole funnel
Rejected and withdrawn candidates appear with the stage and reason, not only the people who were hired.
Structured scorecards
Competency scores with the rubric version sit beside the written feedback that justified them.
Later outcomes
Retention or assignment completion adds a label that arrives months after the hiring decision.
Stage timestamps
Dated transitions show where candidates waited, dropped out or were fast-tracked.
Posting history
Each version of a job description is linked to the applications it attracted.
Separable demographics
Self-identification held apart from decision records supports bias testing without training on it.
Which labels in hiring data deserve trust
A hiring record mixes two kinds of information. Process data — what was posted, who applied, how long each stage took, what interviewers wrote and how scorecards turned into decisions — describes the work of recruiting fairly reliably. Decision data — who advanced and who received an offer — records a judgment made under that team's conditions at the time: budget, urgency, the other people in the pool and the interviewers' preferences, including any bias.
That split should decide how the data is used. Tasks built on process data, such as coordinating interview loops, summarizing feedback against a rubric or drafting postings, carry little risk of learning whom to favor. Tasks that treat decisions as targets, such as ranking applicants, inherit past selection patterns and need adverse impact testing before and after training.
Outcomes after the hire do not settle the question. Retention and performance exist only for people who were hired, so a model never learns how a rejected candidate would have done. Researchers call this the selective labels problem: even outcome labels describe a pool the old process had already filtered.
Traps in applicant tracking exports
- Policy changes split the history. Dropping a degree requirement, publishing pay ranges or moving to structured interviews changes who applies and how they are judged. Date these changes and avoid pooling across them blindly.
- The applicant tracking system records clicks, not always decisions. Bulk rejections when a requisition closes, auto-advance rules and stages skipped for internal candidates create events no one actually decided.
- Evergreen and canceled requisitions. Postings kept open to collect résumés, and requisitions canceled without a hire, carry no real decision for most of their applicants.
- Referrals and internal moves follow other paths. These candidates often skip stages or meet different interviewers, so mixing them with outside applicants distorts stage-transition and time-to-hire models.
- Interviewer calibration drifts. Rubrics are revised and some interviewers score high on everything. Keep interviewer pseudonyms so scores can be normalized per interviewer and rubric version.
What to check before licensing
- Establish who controls the records. An employer owns its applicant tracking data; a recruitment process outsourcing provider usually works inside its client's system, so the client must authorize; a staffing agency typically owns its candidate database, but client job orders, feedback and rates are often confidential under client agreements.
- Read the candidate privacy notices and retention policies that applied at collection, including for candidates in the EU, the UK and California, where state privacy law now covers job applicants, and check that the agreed use fits them.
- Run an adverse impact check on the sample before using decisions as labels — selection rates by group at each stage where self-identification allows — and look for proxies such as names, graduation years, addresses, employment gaps and affinity-group memberships.
- Find out which decisions were automated, such as knockout questions, assessment cut scores, ranking tools or auto-reject rules, because those labels record a tool's output, and ask whether the tool was audited.
- Review résumé de-identification on the sample for indirect identifiers as well as direct ones, including unusual titles, small employers, exact dates, publications, portfolio links and photos.
- Confirm exclusions for material that should not travel, including background check reports (consumer reports under the US Fair Credit Reporting Act), medical and accommodation details, immigration documents and recorded video interviews.
- Check how disposition reasons were recorded, since many teams default to generic codes, and measure on the sample how often a reason is specific enough to learn from.
- Deduplicate candidates who applied to several requisitions or arrived through more than one agency, so the same person does not land in both training and test sets.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Can I license real applicant tracking data for AI training?
Yes, provided whoever controls the records consents and candidate identities are removed first. Employers own their applicant tracking data, recruitment process outsourcing providers usually work inside a client's system, and staffing agencies typically own their candidate databases, though client details may be confidential. SourceX confirms which of these applies before a dataset is offered. Whether matching histories exist depends on which partners agree to license.
Does training on historical hiring decisions risk building in bias?
Yes. A past decision records what a hiring team chose, under its own constraints and possible biases, rather than how good a candidate was, and a model can reproduce adverse impact through proxies such as names, schools, addresses or career gaps. Decision labels are safest when you test selection rates by group, remove known proxies and treat recorded outcomes as one signal among several. Self-identified demographics, where they can be licensed, make those tests possible.
Are AI hiring tools regulated?
AI hiring tools are regulated in some jurisdictions, which matters if you train a model for hiring use. New York City requires a bias audit within a year before an automated employment decision tool is used, plus notice to candidates. The EU AI Act treats AI used to recruit or select candidates as high-risk, with requirements that include examining training data for bias. Illinois regulates AI analysis of video interviews, and other states have passed or proposed similar rules.
Can a staffing agency license its candidate and placement data?
Often, within limits. An agency's candidate database and its recruiters' notes are usually its own records, but job orders, client interview feedback, bill rates and client names typically fall under confidentiality terms in client agreements and may need the client's consent or removal. The privacy notices candidates received also apply. Placement histories are scoped to what the agency controls outright, with client identities tokenized.
How are résumés de-identified?
Names, contact details, addresses, profile links and photos are removed, and the details that make a career history unique are generalized: exact dates become durations, small employers become industry and size bands, and rare titles map to standard ones. How far to generalize is agreed in scope, because skills, titles and career paths carry most of the signal. Residual identifiers are checked on the sample before licensing.
Can I get demographic data to test for adverse impact?
Sometimes, in a separate table. Many employers collect voluntary self-identification for equal employment reporting and keep it away from hiring decision-makers. Where the data owner agrees and the law allows, it can be delivered aggregated, or under a separate audit key with strict use limits, so you can measure adverse impact without training on protected characteristics. In the EU and UK, data on racial or ethnic origin is special-category personal data, and sharing it faces stricter conditions.
Do recruiting datasets include what happened after the hire?
Sometimes. Retention at set checkpoints, or completion of a staffing assignment, can be linked when the partner holds HR or assignment records for the same people. Performance ratings are more sensitive employee data and are often excluded or reduced to broad bands. Say which outcomes you need in your request, because they decide which partners can supply the data.
Related datasets
- Enterprise workflow and task execution histories
Linked task trajectories from request to outcome, across every tool the work touched
- Human feedback and QA-scored work
Work items with scores, verdicts and corrections from the people who reviewed them
- Approved workplace email and chat exports
Approved, de-identified exports of team email threads and chat channels
- SOPs, playbooks and internal knowledge bases
Written procedures with page history, ownership and links to execution records
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026. Own data like this? See how companies license it to AI developers.