Skip to content

Privacy and preparation

Spanish-language and multilingual records: why PII tools miss names

By SourceX Editorial · Updated

Short answer

PII tools miss names in Spanish and mixed-language records because name detection relies on models and context words tuned mostly for English. For bilingual service records, detect the language of each message first, run recognizers built for that language, and review a larger human sample of the non-English slice than you would for English-only records.

Key takeaways

  • Detect language per message or per field, not per file, because bilingual records switch languages mid-thread.
  • An English-tuned name model often reads Spanish given names such as Rosa, Luz or Ángel as ordinary words.
  • Double surnames and particles such as de la need rules that capture the whole name, not only its first part.
  • Context words that boost detection need Spanish equivalents, such as me llamo, se llama, señora and teléfono.
  • Measure misses separately for each language slice so a strong English result cannot hide a weak Spanish one.

Why do PII tools miss names in Spanish and bilingual records?#

PII tools miss names in Spanish and bilingual records because name detection usually comes from a model trained mostly on English text, and the supporting rules assume English wording. A model that has learned to expect a surname after Mr. has little to work with when a customer writes habló la señora Pérez, or texts a name in lowercase.

Patterned identifiers hold up better. Email addresses, card numbers and US phone numbers look the same in any language, so regular expressions still catch them. Coverage drops on names, street addresses and free-text descriptions of people, which are exactly what fill technician notes and customer messages in a home services company.

Open-source tools make the design easy to see. Presidio, the open-source toolkit now run as a community project under the Data Privacy Stack organization, combines a named-entity model with pattern rules, checksums and context words. For names, the model does most of the work, so its training language matters most. Context words, the nearby terms that raise confidence in a match, are usually configured per language as well, so an English list adds nothing on a Spanish note. The project itself warns that automated detection cannot guarantee it finds all sensitive information, and that warning weighs heavier outside the language a pipeline was tuned for.

Where names hide in a contractor's job records#

Spanish shows up across the whole job record, not only in a separate Spanish queue. In a contractor running ServiceTitan, Housecall Pro or Jobber, bilingual customer service reps type booking notes, technicians add job notes from a phone, and customers text back in whichever language is easier for them.

The table maps the usual record types to the specific way names slip through. Use it to decide where your bilingual reviewers spend their time.

Where names hide in a contractor's job records
RecordTypical Spanish or mixed contentWhat the scanner tends to miss
Call booking notesCliente dice que su hijo Jesús abre la puertaGiven names that read as words; relatives named in passing
Technician job notesShorthand that switches between English and SpanishTenants, neighbors and property managers named mid-sentence
Customer text messagesLowercase names, missing accents, nicknamesNames without capital letters or titles to anchor them
Call and voicemail transcriptsSpeech-to-text spellings of Spanish namesMisspelled names that match no dictionary
Estimates and invoicesBilling names with two surnamesThe second surname left behind after the first is masked
Complaint emails and reviewsFormal Spanish with don, doña or señorTitles that should have triggered a name match

Five name patterns that English-tuned tools get wrong#

Five name patterns come up again and again when a bilingual reviewer checks an English-tuned scanner's output. Each one has a specific fix, so test for them by name rather than hoping a general model improves.

  • Given names that are also everyday words, such as Rosa, Luz, Blanca, Paz, Dolores and Ángel. A word-list rule over-redacts la luz de la cocina, the kitchen light, while a weak model misses Luz llamó otra vez, Luz called again.
  • Double surnames and particles, as in María Fernanda García López or Juan de la Cruz. Tools often mask the first surname and leave López or de la Cruz, which can still point to one household in a small service area.
  • Missing accents and lowercase. Exports and texting gateways often drop accents, and customers type jose or maria. Rules that depend on capitalization or exact spelling lose them.
  • Titles and kinship words: don Ramón, doña Elena, la señora Ruiz, su esposo Beto. These are strong signals in Spanish, but only when the context word list includes them.
  • Code-switching inside one sentence, such as Customer said que su suegro Arturo will be home after lunch. A detector run on the whole note may label it English and send it to the English model, which skips the Spanish clause.

A checklist for mixed-language records#

The fix for mixed-language records is a pipeline that treats language as a property of each piece of text, not of the dataset. The steps assume records are exported with a stable record ID, so every finding can be traced back to its source.

One step does more than the rest: use your own customer table as a dictionary. The customer master in your field service system already holds the names most likely to appear in notes, in every spelling customers gave you. Route matches on names that are also everyday words to a reviewer instead of masking them automatically, or the kitchen light disappears along with Luz.

  • Split records into units small enough to carry one language: a message, a note, a field, or a sentence in long notes.
  • Detect the language of each unit, store a confidence value, and send low-confidence or mixed units through both pipelines.
  • Run a name and location model trained for Spanish on Spanish units and the English model on English units.
  • Add Spanish context words and titles to your rules: me llamo, mi nombre es, se llama, señor, señora, don, doña, teléfono, dirección, calle and avenida.
  • Extend surname rules so a match expands across particles and a second surname.
  • Match against a copy of the text with accents and case stripped, then map each finding back to the original.
  • Run known customer and employee names from the field service customer list, CRM and payroll as an exact-match deny list against every unit, matching on the accent-stripped copy.
  • Draw a separate human review sample from each language slice and record misses per slice.

How to review the Spanish slice#

The Spanish slice needs its own reviewers and its own pass mark. A blended result can look acceptable while many Spanish notes still carry a name, simply because English notes outnumber them.

Review samples should be larger relative to the slice for languages your pipeline handles less well. You are measuring a process you trust less, so you need more evidence before you sign off. Fix the release threshold before anyone reads a sample, so a weak Spanish result cannot be talked up after the fact.

How to review the Spanish slice
SliceReviewerWhat to countRelease rule
English onlyAny trained reviewerMissed names, addresses and phone numbersRelease when misses meet the agreed threshold
Spanish onlyFluent Spanish readerThe same, plus surname fragments and titlesSame threshold, measured on this slice alone
Mixed languageBilingual readerMisses at the point where language switchesHold until reviewed, even if other slices pass
Low-confidence languageBilingual readerWrong routing and missed namesFix routing, then draw a fresh sample

Illustrative: a bilingual HVAC and plumbing contractor#

Illustrative: a fictional HVAC and plumbing contractor in the Southwest runs ServiceTitan, with a bilingual call center and technicians who write notes in both languages. Before a planned export, its IT lead ran an English-tuned scanner over booking notes, job notes and customer texts.

The English sample looked clean. A bilingual reviewer then read a separate sample of Spanish and mixed notes and found names such as Rosa and Ángel left in place, second surnames remaining after the first was masked, and relatives named after su hijo or su suegra. The team added per-note language detection, a Spanish model, Spanish context words, the customer master as a deny list and surname expansion. A fresh Spanish sample met the same release rule as the English one, and the company kept its Spanish notes in scope instead of dropping them.

How SourceX reviews language slices#

SourceX reviews mixed-language records one language slice at a time, starting before any file moves. At the fit check stage the only question is descriptive: which systems hold Spanish or mixed-language text, and roughly what share of the records it covers.

If the records proceed, Preparation reports review results for each language slice, so a gap in one language is visible rather than averaged away. The privacy record in the SourceX Evidence Packet states which languages were reviewed, by whom and against what release rule, and the supplier approves that result before anything moves to Delivery in the SourceX five-step transaction.

Frequently asked questions

Do AI developers want Spanish-language service records at all?

Some do, depending on the buyer's project, while many buyers prioritize predominantly English records. Bilingual service records show how real customers describe problems and how staff respond across languages, which is hard to reproduce. Confirm buyer interest during scoping before building a full Spanish pipeline, and record the language mix in your inventory so it is visible early.

Is translating everything to English before redaction a good shortcut?

Usually not. Translation can alter or drop names, and a buyer may want the original language. If you translate to help detection, map each finding back to the original text and redact there. Never deliver a translated version without saying so in the dataset documentation.

What about other languages, such as Vietnamese or Portuguese?

The same method applies: detect language per unit, use a model built for that language, add native context words, and review with fluent readers. If a language appears in only a few records and you have no reviewer for it, excluding that slice is often the safer choice.

Can we rely on a cloud provider's PII API for Spanish?

Check the provider's current documentation for which entity types it supports in Spanish, because support often differs by language and by entity type. Then test it on your own notes with a bilingual reviewer. Coverage claims are a starting point for your pipeline, not a release decision.

Should accents be kept in the delivered data?

Yes, in the delivered text. Strip accents only in a working copy used for matching, then apply redactions to the original. Accented spellings carry meaning, and a buyer studying real customer language will want them intact.

Sources

  • Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums with context. Source
  • Presidio's documentation warns that because it is using automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information, so additional systems and protections should be employed. Source
  • Presidio moved from a Microsoft-owned project to an independent, community-governed open-source project under the GitHub organization Data Privacy Stack and remains MIT-licensed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify