Software companies
Applicant data in HR software: can the vendor train screening models on it?
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
An HR software vendor can train screening models on applicant data only where the employer's contract allows it, the vendor's privacy role permits it, and the use holds up under bias and employment-law review. Most ATS vendors process applicant data on the employer's behalf, so the default answer is no unless contracts and applicant notices clearly say otherwise.
Key takeaways
- Employers usually control applicant data; the ATS vendor processes it to provide the service.
- Training a cross-customer screening model is a separate purpose that most processor terms do not cover.
- Contract deletion promises reach training copies, so a model built on applicant data has to respect them.
- Screening models used in hiring carry discrimination risk that sits apart from privacy and contract questions.
- Vendor-owned records such as support tickets and engineering history are a safer starting point for any licensing.
Who controls applicant data in an ATS?#
Applicant data in an ATS is usually controlled by the employer that collected it, with the vendor processing it to run the hiring workflow. Resumes, application answers, interview scorecards, assessment results, recruiter notes, disposition codes and offer details all fall in that category.
The employer decides why the data is collected and how it is used. The vendor's rights come from the employer contract and the data processing terms attached to it. Under GDPR that is the processor role; under California law the vendor is often a service provider; other frameworks use similar ideas.
State law adds a patchwork. The Colorado Attorney General states that the Colorado Privacy Act does not cover personal data of people acting in an employment context, such as a job applicant. California went the other way: its privacy agency began preliminary rulemaking on April 20, 2026 on how the CCPA applies to employee, job applicant and contractor information. Which laws may apply to a dataset is assessed deal by deal with counsel.
Three checks before any model training#
Three checks decide whether a vendor can train screening models on applicant data: the processor check, the contract check and the bias-liability check. Failing any one usually ends the analysis, whatever the other two show.
| Check | Question | Where to look | Common blocker |
|---|---|---|---|
| Processor check | Does the vendor's privacy role allow use beyond serving that employer? | DPA, privacy notices, applicable privacy laws | Processing limited to providing the services |
| Contract check | Does the employer contract permit training, aggregation or de-identified use? | MSA, order forms, enterprise amendments, AI addenda | No training right, or a clause struck in negotiation |
| Bias-liability check | Could the model or its training data produce discriminatory outcomes? | Model design, labels, testing plan, employment law advice | Historical hiring decisions used as labels without testing |
The processor check: purpose limits come first#
The processor check asks whether the vendor's role lets it use applicant data for its own purposes at all. A processor or service provider generally uses personal data only to provide services to the customer that supplied it, and training a model that serves every customer is a different purpose.
Some privacy frameworks allow a service provider to use data to build or improve its services within limits. Whether that reaches a screening model, and on what conditions, is a question for counsel in each jurisdiction involved.
De-identification is harder here than it looks. Resumes and recruiter notes are free text, and a combination of employers, schools, job titles and locations can point to one person even after names are removed.
The contract check: read every version#
The contract check reads every version of the employer agreement still in force, because hiring software customers often negotiate their paper. A clause that allows aggregated use in your standard terms may be missing from the contracts of your largest customers.
Deletion terms deserve special attention. Greenhouse's Master Subscription Agreement, for example, says all Customer Data is queued for deletion 90 days after the agreement expires or is terminated. A vendor with a similar promise has to account for applicant data sitting in training sets and feature stores, not only in the production database.
Look at security questionnaires and RFP answers too. A statement your team gave a customer, such as a promise that customer data never trains models, may have been incorporated into the contract or relied on by the customer. Treat those answers as commitments when deciding what the contract allows.
- Permitted use clauses limiting processing to providing the services.
- Aggregated, anonymized or de-identified data clauses, and who owns the output.
- AI or machine learning addenda, including promises not to train on customer data.
- Confidentiality clauses covering applicant information.
- Deletion and return obligations at termination, and whether they reach backups and derived data.
The bias-liability check is a separate risk#
The bias-liability check stands apart from privacy and contract review, because a model trained with every permission in place can still screen people unfairly. Anti-discrimination law applies to hiring outcomes, and litigation over AI hiring tools has put vendors, not only employers, in the frame.
Training data is where much of that risk starts. Past disposition codes and hiring decisions reflect past practices, so using them as labels can teach a model to repeat them. Some jurisdictions also regulate automated employment decision tools with notice or audit requirements, which may apply depending on where employers and applicants are located.
Document the model's purpose, labels, testing plan and who reviews results. That record matters in customer audits, in diligence and in any later dispute.
Alternatives to training on pooled applicant data#
Alternatives exist for vendors whose contracts do not support cross-customer training. Each trades some model reach for a clearer rights position, and several can run together.
Whichever option you choose, keep the rights basis for each customer in a register. A buyer, auditor or large customer will ask for it, and rebuilding it later from old email threads is slow and unreliable.
| Option | What it uses | Main limitation |
|---|---|---|
| Per-customer models | One employer's data, used only for that employer | Smaller training sets and more operational overhead |
| Opt-in program | Data from employers who sign a specific amendment and update applicant notices | Slower to assemble and limited to willing customers |
| Synthetic or public job data | Generated resumes and public job descriptions | Weaker signal for real screening decisions |
| Vendor-owned records | Support tickets, configuration histories, engineering records | Useful for product and support AI, not for ranking applicants |
Illustrative: an ATS vendor weighs a ranking feature#
Illustrative: a fictional ATS vendor serves regional logistics and warehouse employers. Its product team wanted an applicant ranking feature trained on years of applications, scorecards and dispositions pooled across all customers.
The review found that the DPA limited processing to providing the services, that several large customers had struck the aggregated data clause, and that disposition codes were used inconsistently from one employer to the next. Counsel also flagged differing state rules on applicant data.
The vendor dropped pooled training. It ran a per-customer ranking pilot with employers who signed an amendment and updated their applicant notices, with a bias testing plan agreed before launch. Separately, it scoped a license of its own support and engineering records, with candidate details removed, as a lower-risk project.
How SourceX treats applicant data#
SourceX excludes applicant and candidate personal data from packages by default. In the SourceX five-step transaction of Supply, Rights, Preparation, Approval and Delivery, the Rights and Preparation steps confirm that only vendor-owned records, such as support tickets, issue histories and code reviews, move forward.
SourceX does not advise a vendor on training its own models, and its own rights in a deidentified dataset are set out in the signed supplier agreement. What it adds is documentation: a SourceX Evidence Packet for each package, recording provenance, licensing rights, permitted use, the privacy record and release authorization.
Frequently asked questions
Is de-identified applicant data free to use for training?
Not automatically. De-identification standards differ by law, and contracts may restrict de-identified use as well. Resumes and recruiter notes are hard to de-identify reliably because career details can single out a person. Counsel should review both the method and the contractual basis before any use.
Can an employer customer give us permission to train?
An employer can grant rights within its own legal limits, usually through a written amendment. It may also need to update applicant notices, and some uses may need more than notice. Confirm what the employer is able to authorize rather than relying on a signature alone.
Do privacy laws treat applicants differently from consumers?
Some do. Several state consumer privacy laws exclude people acting in an employment context, while California applies its law to applicant data and is working on rules for it. Employment and anti-discrimination laws apply on top. Which rules govern a dataset depends on where employers and applicants are.
Is improving resume parsing the same as training a screening model?
Not in risk terms. Parsing extracts fields from documents; screening ranks or filters people. Contracts may permit one and not the other, and the bias exposure differs sharply. Describe each purpose separately in contracts, notices and internal records.
What should a vendor document if it does train on applicant data?
Record the rights basis by customer, the data and labels used, the de-identification method, the bias testing plan and results, and how deletion requests and contract terminations reach training sets. That record supports customer audits, regulator questions and any future buyer's diligence.
Sources
- The Colorado Attorney General states that the Colorado Privacy Act does not cover personal data of individuals acting in a commercial or employment context, such as a job applicant. Source
- The California Privacy Protection Agency initiated preliminary rulemaking on April 20, 2026 focused on how the CCPA applies to personal information of employees, job applicants and independent contractors. Source
- Greenhouse's Master Subscription Agreement says all Customer Data is queued for deletion 90 days after the Agreement expires or is terminated for any reason. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.