Skip to content

Industry-specific operational data

Skills-labeled job requisitions for skills extraction and taxonomy mapping

Quick answer

The most useful skills extraction training data is not scraped job ads but internal requisitions where recruiters and hiring managers confirmed which skills are required, which are preferred, and which taxonomy code each one maps to. Buyers should ask for span-level skill labels with label provenance, ONET-SOC or ESCO codes with the taxonomy version recorded, requisition revision history, and tokenized client and manager names. No candidate data is needed to train skills extraction, title normalization or requisition drafting.

By SourceX Editorial · Updated

Why real recruiter labels beat synthetic and weakly supervised data

Recruiter-verified labels fill the gap the research literature keeps naming: real annotated job-posting data is scarce. SkillSpan built a human-annotated English job-posting corpus with span-level guidelines for hard and soft skills, noting that earlier datasets relied on crowd-sourced or document-level labels [1]. Later work on extreme multi-label skill extraction states that scarcity of real-life annotation data is the main problem and turns to LLM-generated sentences built from ESCO skill descriptions, which may not reflect how real ads are written [2]. Weak-supervision approaches map job-ad text to ESCO to avoid manual labeling, at the cost of noisy alignments [3].

Operational requisitions offer a different label source. In an applicant tracking system (ATS) or vendor management system (VMS), a recruiter often edits the required and preferred skill lists after an intake call, and those edits are a human judgment about the role, not a crowd-worker's guess. That makes them strong supervision for span extraction, skill-to-code classification and drafting models, provided you know which field each label came from.

Where the skill labels in a requisition come from

Each label source in a requisition has a different reliability, so the dataset should carry a provenance field per label rather than one merged skill list. Treat the reliability ranking below as a working hypothesis to test on a sample, not a fixed rule.

Label sourceTypical system and fieldStrengthFailure mode
Hiring-manager intake formIntake questionnaire, "must-have skills" free textClosest to real job needWishlists, vague phrasing ("strong communicator")
Recruiter-edited required vs. preferred skillsATS requisition skill fieldsHuman-verified, split by importanceRecruiter shorthand; copied from templates
ATS skill tagsPicklist tags tied to a vendor skills libraryNormalized, machine-readyTag drift when the vendor library changes
Posted job ad textCareer-site or job-board posting bodyNatural language for span extractionMarketing boilerplate, EEO text, benefits noise
Staffing job orderVMS or front-office job order from a clientRate, duration and skills in one recordClient names and bill rates need tokenizing

The gap between the intake form and the final posted ad is itself a signal. A model trained only on posted ads learns marketing language; one trained on intake-plus-final pairs learns which requirements survive review.

Mapping skills and titles to O*NET-SOC and ESCO

Taxonomy mappings are only reusable if every code carries the taxonomy name and version it came from. In the US, the common target is O*NET-SOC, which is built on the federal Standard Occupational Classification; in the EU, it is ESCO, the European skills, competences, qualifications and occupations classification that weak-supervision research already uses as a label space [3]. Both are revised periodically, and ESCO publishes numbered releases with downloadable files and an API. As of October 2026, confirm the current release of each before you fix your label schema, because a code valid in one version can be split, merged or relabeled in the next.

For title normalization, ask for three fields per requisition: the raw internal title ("Sr. SWE II - Payments"), the normalized title the recruiter or ATS chose, and the occupation code. Internal leveling suffixes, team names and requisition prefixes are what title models must learn to strip. If the supplier used a vendor skills library rather than O*NET or ESCO, request the crosswalk table and its date, or plan to map codes yourself and record that the mapping is buyer-created.

Retail catalogs face a similar many-to-one mapping problem with different failure modes; see product categorization and taxonomy-mapping data and product and spend classification training data for how code-version drift is handled there.

Requisition revisions as drafting supervision

Revision history turns a static posting into supervised fine-tuning pairs for requisition drafting. When the ATS keeps versions, each edit from draft to approved posting shows which skills a hiring manager added, which a recruiter softened from required to preferred, and which compliance language legal inserted. Ask whether the export includes version timestamps and the editor's role (manager, recruiter, compensation, legal), with personal names replaced.

Structured compensation fields are worth requesting as well. Where a pay-transparency law applies to the posting, the requisition often carries minimum and maximum pay, currency and pay period as separate fields; ask the supplier which postings were subject to such rules rather than inferring it. For drafting models released to the public, remember that California AB 2013 (documentation due January 1, 2026) requires developers of generative AI systems offered to Californians to post documentation about their training data, so keep the provenance you receive [4]. For output formats, see structured-output fine-tuning data.

Illustrative record schema for a skills-labeled requisition

A practical delivery is one JSON Lines record per requisition version, UTF-8 encoded with one JSON object per line [6], with labels kept separate from text so you can rebuild spans.

Illustrative example: invented to show structure; it does not describe an available dataset.

{"req_id": "REQ-TOKEN-8F21", "version": 3, "version_ts": "2025-03-14", "editor_role": "recruiter",
 "client_token": "CLIENT-0042", "raw_title": "Sr. SWE II (Payments)",
 "normalized_title": "Software Developer", "onet_soc_code": "15-1252.00", "onet_soc_version": "O*NET-SOC 2019",
 "esco_occupation_uri": null, "posting_text": "...", "pay_min": 140000, "pay_max": 175000, "pay_period": "annual",
 "skills": [
   {"span": [212, 224], "text": "Apache Spark", "importance": "required", "label_source": "recruiter_edit",
    "taxonomy": "ESCO", "taxonomy_version": "v1.2", "code": "esco-skill-uri-placeholder"},
   {"span": [301, 318], "text": "stakeholder comms", "importance": "preferred", "label_source": "manager_intake",
    "taxonomy": null, "taxonomy_version": null, "code": null}
 ],
 "flags": [{"span": [402, 419], "type": "age_coded_term", "text": "digital native"}]}

Note three design choices. Unmapped skills keep a null code rather than a forced match. The flags array marks age- or gender-coded terms instead of silently deleting them, so you can train a biased-language detector and audit the drafting model. And the client is a token, not a name.

Removing identities without destroying the labels

You can build this dataset without any candidate data, but the requisition itself still contains identifiers that must be tokenized. Staffing job orders name the client company, hiring managers appear in approval chains, and small-team postings ("report to our CFO in Boise") can re-identify a person even after names go. See indirect identifiers in business text for why job titles plus location are a classic quasi-identifier.

Consistent tokens matter more than deletion: if "CLIENT-0042" appears across 300 requisitions, your model still learns per-client patterns without learning the client. Memorization is the residual risk. Language models can regurgitate verbatim training sequences, including contact details [7], so check that recruiter email signatures and phone numbers in posting footers were removed before training a generative drafting model. More on that in training-data extraction and memorization risk.

Acceptance checklist for skills-labeled requisition data

Before you accept a delivery, test the labels the way you would test any annotation vendor's work. The steps below assume you hold a sample and the label guidelines.

  1. Label provenance: every skill label names its source field (intake, recruiter edit, ATS tag, posting text).
  2. Taxonomy versioning: every code carries taxonomy and version; spot-check 50 codes against the official O*NET or ESCO release files.
  3. Required vs. preferred: the split is explicit, not inferred from word order.
  4. Span integrity: character offsets match the delivered text after de-identification edits.
  5. Inter-annotator signal: where two people touched a requisition, measure agreement on skill sets; see annotation quality audit.
  6. Bias flags: coded terms are flagged with type and span, not removed.
  7. Identity removal: client, manager and recruiter names tokenized consistently; footers and signatures stripped.
  8. Distribution: counts by occupation group, seniority, industry and region, so you can see coverage gaps.
  9. Documentation: a dataset card covering source systems, label process, known biases and gaps [5].

How SourceX approaches requisition data requests

SourceX sources operational datasets, including documents and workflow records, from US companies on request and manages the licensing process; requisition data is not held in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery. Personal details such as names, emails and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Recruiting-specific context is on the recruiting and staffing buyers page, and you can describe the requisition data you need to SourceX.

If your need is broader than skill labels, the owner pages for licensing job descriptions and interview notes, recruitment and hiring workflow datasets and training data for recruiting AI cover those intents. For other sector datasets, start at the industry-specific operational data hub or the AI data buyer guides.

Source recruiter-labeled requisitions for skills extraction

Describe the requisitions, label fields and taxonomies you need rather than naming specific businesses; SourceX looks for US companies that hold that data, and every release is approved by the supplying company. Nothing is contracted until a supplier agrees, and pricing and allowed uses are set per deal in a license. Start a buyer request at SourceX.

Sources

  1. arXiv (Zhang et al.), "SkillSpan: Hard and Soft Skill Extraction from English Job Postings" (2022). https://arxiv.org/pdf/2204.12811
  2. arXiv, "Extreme Multi-Label Skill Extraction Training using Large Language Models" (2023). https://arxiv.org/pdf/2307.10778
  3. arXiv (ar5iv), "Skill extraction from job postings using weak supervision (ESCO)" (2022). https://ar5iv.labs.arxiv.org/html/2209.08071
  4. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  5. Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
  6. jsonlines.org, "JSON Lines". https://jsonlines.org/
  7. USENIX Security 2021 (Carlini et al.), "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data