Skip to content

Consulting and recruiting

De-identifying resumes and candidate profiles: what to remove

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

To de-identify a resume, remove direct identifiers, generalize quasi-identifiers and review free text by hand. Names, contact details, profile links, photos and ID numbers go; employers, schools, dates and locations become categories; summaries, cover letters and recruiter notes are scanned and then read. A resume is de-identified only when what remains cannot point to one person.

Key takeaways

  • Direct identifiers are removed; quasi-identifiers such as employer, school, dates and city are generalized.
  • Career history is the hardest part, because a sequence of employers and titles can be unique to one person.
  • Automated detection is a first pass only; a person needs to read the free text.
  • Blind resumes for fair hiring and de-identified resumes for AI data have different goals and different standards.
  • Re-identification risk is tested across the whole dataset, not resume by resume.

Why is a resume harder to de-identify than a form?#

A resume is harder to de-identify than a form because the person wrote it to be recognized. Identity is spread across the whole document: an employer sequence, a niche certification, a conference talk and a city can together describe one individual even after the name is gone.

This is also why a blind resume is not the same thing. A blind resume hides a few fields from a hiring manager to reduce bias in one decision. De-identification for an AI dataset has to hold up against anyone who holds the data, can search the web and has time, so it removes and generalizes far more.

Decide the purpose before choosing treatments. If the dataset is meant to show how skills and roles progress over a career, durations and role families must survive and exact employers can go. If it is meant to show how resumes are matched to job orders, skills and seniority matter most and career detail can be coarser. Fields that serve no stated purpose should be removed rather than generalized.

Field-by-field table: direct identifiers, quasi-identifiers and free text#

The field-by-field table below gives a default treatment for the content found on most resumes. Remove means the content is deleted, generalize means it is replaced with a broader category and review means a person reads it after automated scanning.

Field-by-field table: direct identifiers, quasi-identifiers and free text
FieldTypeTreatmentWhy
Full name and initialsDirect identifierRemoveNames the person
Email, phone, street addressDirect identifierRemoveContact details are unique
LinkedIn, GitHub and portfolio linksDirect identifierRemoveLead straight to a public profile
PhotoDirect identifierRemoveIdentifies and reveals protected traits
Date of birth or ageSensitiveRemoveProtected characteristic
Government ID, license and badge numbersDirect identifierRemoveUnique by design
Employer namesQuasi-identifierGeneralize to industry and size bandEmployer sequences can be unique
Job titlesQuasi-identifierGeneralize rare titles to a role familyUnusual titles are searchable
Employment datesQuasi-identifierConvert to durations or shift consistentlyExact dates narrow the field
School names and graduation yearsQuasi-identifierGeneralize to degree level and field; drop yearsYears reveal age
City and stateQuasi-identifierGeneralize to regionSmall markets narrow the field
Certifications with ID numbersMixedKeep the certification name; remove the numberNumbers can be looked up
Publications, patents, awards, talksQuasi-identifierRemove or reduce to a count by typeSearchable by title
Military service, visa or work authorizationSensitiveRemoveProtected or sensitive status
Salary history and expectationsSensitiveRemovePersonal and restricted in some places
Hobbies, affiliations, volunteer rolesSensitiveRemoveCan reveal religion, politics, health or ethnicity
Summary, objective, cover letterFree textReviewMixes all of the above in prose
ReferencesThird-party dataRemoveOther people's personal data

What else sits in an ATS candidate profile?#

An ATS candidate profile carries far more than the resume, and each part needs its own rule. Profiles usually hold status history, source, tags, recruiter notes, attached files and communications logs, and some hold EEO self-identification answers collected for reporting.

Treat the profile as a set of separate decisions rather than one record:

  • Status history and stage dates: keep, with tokens and consistent date shifting.
  • Source and campaign codes: keep as categories; remove referrer names.
  • Skills tags: keep, after checking that no tag encodes a protected trait.
  • Recruiter notes: review line by line for names, health, family and pay remarks.
  • Attachments such as cover letters, transcripts and ID scans: exclude unless reviewed individually.
  • EEO and voluntary self-identification: exclude entirely.
  • Email and text logs: usually exclude.

Where automated tools miss identifiers#

Automated tools catch the obvious identifiers well and miss the contextual ones. Presidio, an open-source PII detection and anonymization SDK now governed by the community Data Privacy Stack organization, combines named-entity recognition, regular expressions, rule-based logic and checksums. The project is candid about the limit: its documentation says automated detection cannot promise to catch every piece of sensitive information, so other safeguards are needed alongside it.

Resumes produce predictable misses: names embedded in email handles or file names, project and product names that identify an employer, client names in consulting or contract histories, foreign-language text and details inside images or tables. Run automated detection first, then have a trained reviewer read a meaningful sample of every batch and all free text flagged as uncertain.

How do you test whether a de-identified set can be re-identified?#

Testing for re-identification means checking the dataset as a whole, because risk comes from rare combinations of fields shared by few records. The research behind this is old and consistent: Latanya Sweeney estimated from 1990 Census data that 87% of the US population was likely unique on five-digit ZIP code, gender and date of birth, and Philippe Golle's 2006 revisit put the figure at about 63% on 2000 Census data. The exact number depends on method; the lesson for resumes, which carry far more detail than three fields, does not.

The standard test is k-anonymity, defined in Sweeney's 2002 paper: a release is k-anonymous when each person's record cannot be distinguished from at least k-1 others. For resumes, that means every combination of role family, region, degree level and years of experience should be shared by at least a set number of records; l-diversity adds a check that sensitive values inside each group vary. Google's Sensitive Data Protection API, for example, offers k-anonymity, l-diversity, k-map and delta-presence as built-in risk analyses. When a combination is too rare, coarsen a field or remove the record; niche specialists in small markets are the usual cases.

Illustrative: an engineering staffing firm prepares a resume sample#

Illustrative: a fictional engineering staffing firm wants to understand what its candidate archive would look like after de-identification. It pulls an internal sample of resumes and ATS profiles for mechanical and controls engineers.

The privacy lead applies the field table, generalizing employers to industry and size and converting employment dates to durations. Automated detection removes contact details, but the reviewer finds employer names inside project descriptions and a patent title that would identify one engineer. A k-anonymity check shows that several senior controls engineers in one small region are unique, so region becomes a broader area for that group. The firm concludes that resumes need heavier treatment than its workflow records and scopes them as a separate decision.

How SourceX approaches resumes and candidate profiles#

SourceX treats resumes and candidate profiles as high privacy burden records. In the SourceX five-step transaction, the Rights step checks candidate notices and the consent language in use when the records were collected, and Preparation applies field rules the supplier has approved, with a prepared sample reviewed before release.

The privacy record in the SourceX Evidence Packet lists each field treatment, the detection tools and human review used and the results of re-identification testing. Which privacy laws may apply is assessed deal by deal with counsel.

Frequently asked questions

Should employer names ever stay in a de-identified resume?

Rarely. A single large employer may seem harmless, but the sequence of employers is what identifies people. Generalizing to industry and size band keeps the career pattern readable while breaking the link to a searchable profile.

Is a de-identified resume still personal data?

It can be. GDPR Recital 26 treats pseudonymised data that could be attributed to a person using additional information as information on an identifiable person, and other laws take similar views. If your firm keeps a key that links prepared records back to candidates, keep it inside the firm, minimize the fields that travel and have counsel confirm how the applicable laws classify the output.

Can we use an AI model to rewrite resumes into a neutral format?

It can help standardize content, but models can drop or invent details and may not remove identifiers reliably. If you use one, keep it inside your controlled environment, review outputs and record the method in the privacy record, because the result is a derived document.

Do candidates need to be told before their resumes are de-identified for AI use?

The answer turns on the wording of your candidate notices when the records were collected and on the laws covering those candidates. De-identification lowers risk but does not always remove notice or consent obligations, so counsel should review the position before any external use.

What about candidates who asked us to delete their data?

Exclude them before preparation starts. Keep a suppression list of deletion requests, opt-outs and candidates whose records your retention policy says should already be gone, run it against every export and record that check in the privacy record.

Are non-English resumes handled the same way?

The rules are the same, but detection tools often perform less well outside English and reviewers need the language. Many firms handle them as a separate batch or leave them out of a first package.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums with context; its documentation warns there is no guarantee it will find all sensitive information and that additional systems and protections should be employed. Source
  • Presidio moved from a Microsoft-owned project to an independent, community-governed open-source project under the GitHub organization Data Privacy Stack, and remains MIT-licensed. Source
  • Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
  • In her 2000 working paper "Simple Demographics Often Identify People Uniquely," Latanya Sweeney estimated from 1990 U.S. Census data that 87% of the U.S. population (216 million of 248 million) were likely unique on {5-digit ZIP, gender, date of birth}. She put the figure at 53% (132 million) for {place, gender, date of birth}. Source
  • Philippe Golle's 2006 paper "Revisiting the Uniqueness of Simple Demographics in the US Population" found that gender, ZIP code and full date of birth uniquely identify about 63% of the U.S. population on 2000 Census data, fewer than the 87% Sweeney had reported from 1990 Census data. Source
  • Latanya Sweeney's paper "k-anonymity: a model for protecting privacy" was published in 2002 in the International Journal on Uncertainty, Fuzziness and Knowledge-Based Systems, vol. 10, no. 5, pp. 557-570. It defines a release as k-anonymous when each person's record cannot be distinguished from at least k-1 other individuals in the same release. Source
  • GDPR Recital 26 states that personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information, should be considered to be information on an identifiable natural person. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify