Manufacturing
Masking part numbers and customer names without breaking record links
By SourceX Editorial · Updated
Short answer
To mask part numbers and customer names without breaking record links, replace each real identifier with one consistent token everywhere it appears, in every system and every year. Keep the attributes that carry meaning, such as part family, material and revision order, drop what identifies the customer, and prove the joins still return the same counts.
Key takeaways
- One entity, one token: the same customer or part gets the same replacement in every table, file and year.
- Keep the part family and drop the customer, so meaning survives while identity does not.
- Generate tokens with a secret key the manufacturer keeps, and never ship the mapping table with the data.
- Free text and smart part numbers leak identities most often, so they need their own rules and a human review.
- Test masked data by comparing join counts and distinct-entity counts before and after masking.
Why simple redaction breaks manufacturing records#
Simple redaction breaks manufacturing records because part numbers and customer IDs are the keys that join them. Blank out the part number on a job, an NCR and a return, and nobody can tell that all three concern the same valve body. Replace it with a fresh random value on each row, and the result is the same.
Manufacturing makes this harder than most fields. The same part number appears in quotes, BOMs, routings, jobs, inspection plans, NCRs, shipments and warranty claims, usually across an ERP, a QMS and several spreadsheets. A masking approach that works table by table leaves the chain in pieces.
The rule set for consistent masking#
Consistent masking follows a small set of rules applied across every system in scope. Agree on them before anyone writes a script, because changing a rule halfway means redoing every table already processed.
- One entity, one token: each customer, part, drawing and supplier gets a single replacement used everywhere.
- Keyed, not guessable: tokens come from a keyed method such as deterministic encryption or a keyed hash, with the key held by the manufacturer.
- Keep the family, drop the customer: preserve part family, material, process and size class, and remove customer identity and customer part numbers.
- Preserve order where it matters: revisions keep their sequence, so a later revision still follows an earlier one.
- Generalize what would re-identify: ship-to addresses become regions and customer names become industry segments.
- Never share a token between two real items: a part number that was reused for a new item gets one token per real item.
- Keep the mapping out of the delivery: the table linking tokens to real values stays inside the company.
Which fields to tokenize, generalize or drop#
Each field in scope gets one treatment: tokenize, generalize, replace or drop. The table shows typical choices for a manufacturer; the final call depends on customer contracts and the rights review.
| Field | Treatment | Reason |
|---|---|---|
| Customer name and ID | Tokenize | Keeps every order, complaint and return tied to one anonymous customer |
| Customer industry | Generalize to a segment | Preserves context, such as food equipment or trucking, without identity |
| Ship-to address | Generalize to region or drop | Addresses identify customers and their sites |
| Internal part number | Tokenize, keep family attributes | Joins survive and the type of part stays visible |
| Customer part number | Drop, or tokenize separately | Often names the customer's product directly |
| Drawing number and revision | Tokenize number, keep revision order | Shows design changes without exposing the drawing |
| Lot and serial number | Tokenize | Links production, inspection and field returns |
| Operator and technician names | Replace with role codes | Personal data with little value for most uses |
| Prices and costs | Drop, band or scale after review | Often confidential under customer or supplier terms |
Smart part numbers and free text: the two hard cases#
Smart part numbers and free text are where masking most often fails. A smart part number encodes meaning in segments, such as a customer prefix, a product family, a material code and a size. Tokenizing the whole string hides the family; leaving it intact exposes the customer prefix.
The fix for smart numbers is to parse the segments first. Tokenize the customer segment, keep family and material segments that follow your own coding scheme, and store the parsed parts as separate fields so their meaning survives.
Free text is harder. NCR descriptions, job notes and service comments mention customer names, drawing numbers and part numbers in passing. Build a dictionary of every real identifier from master data and replace matches in text with the same tokens used in structured fields. Then add automated detection for names and contact details, plus human review of samples, since Presidio's own documentation warns that automated detection gives no guarantee of finding all sensitive information.
Deterministic encryption, keyed hashes and format-preserving tokens#
The method behind the tokens decides whether masking stays consistent and whether the manufacturer can reverse a token if it ever needs to. Deterministic methods always produce the same output for the same input under the same key, which is what keeps record links intact.
Mainstream tools support these methods. Google's Sensitive Data Protection API offers both deterministic encryption and format-preserving encryption, and its API notes recommend deterministic encryption where the input alphabet need not be preserved, because format-preserving encryption carries significant latency costs.
A keyed hash is a simpler choice when nobody will ever need to reverse a token. Whichever method you pick, record it in the package documentation and protect the key like any other secret. Pseudonymized records can still count as personal data under some privacy laws when people are involved, so masking does not end the privacy review.
Illustrative: before and after masking a nonconformance chain#
Illustrative: a fictional maker of stainless steel process valves masks the records behind one complaint: the sales order, the job, an NCR and a return. Every value shown is invented for the example.
The same tokens appear in the job, the return and every other record about that customer and part, so a reviewer can follow the chain end to end without learning who the customer is. The customer's own part number disappears entirely, while the valve family and material remain readable.
| Field | Before masking | After masking |
|---|---|---|
| Customer | Customer A, a dairy equipment builder | CUST-K7Q2, segment: food equipment |
| Customer part number | CA-44-0917 | Dropped |
| Internal part number | VB-316-2IN-CA | PART-9F3A; family: valve body; material: 316 stainless; size class: 2 inch |
| Drawing and revision | D-11872 rev C | DWG-41BE, third revision |
| Lot | L2318-07 | LOT-77D0 |
| NCR text | Porosity on Customer A order, per D-11872 rev C | Porosity on CUST-K7Q2 order, per DWG-41BE, third revision |
| Inspector | Inspector full name | Role: final inspector |
How to test that the links survived#
Testing masked data means proving that joins behave exactly as they did before masking and that no real identifier slipped through. Run the tests on the full masked set, not on a sample.
- Compare row counts for each join, such as jobs to orders and NCRs to jobs, before and after masking.
- Compare distinct counts of customers, parts and lots; any change means tokens collided or split.
- Rebuild several complete chains by hand and confirm each step still connects.
- Search the output for every real identifier in the master-data dictionary, including partial matches.
- Confirm the mapping table and the key are absent from the delivery folder.
How SourceX handles masking#
Masking happens in the Preparation step of the SourceX five-step transaction, after the Rights step has settled which records and fields are in scope. The manufacturer keeps the key and the mapping; the buyer receives tokens only.
The methods used, the fields treated and the test results go into the privacy record of the SourceX Evidence Packet. That mirrors provenance standards such as the Data & Trust Alliance's Data Provenance Standards, whose list of privacy-enhancing tools includes masking, pseudonymization and tokenization. The manufacturer approves the masked package before delivery.
Frequently asked questions
Can a buyer reverse the tokens?
Not without the key or the mapping table, which stay with the manufacturer. The remaining risk is inference: a rare part family combined with a region and a date might point to one customer. That is why fields are generalized as well as tokenized, and why re-identification risk is reviewed before release.
Should we keep the mapping table at all?
Usually yes, under tight access control. It lets you answer a buyer's question about a record, honor a later deletion request or correct an error. If nobody will ever need to reverse a token, a keyed hash with no stored mapping is a simpler option.
What about drawing numbers inside attached PDFs and images?
Customer drawings attached to jobs and NCRs are usually excluded rather than masked, because the drawing itself is the customer's design. For your own drawings, text inside images and title blocks is hard to mask reliably, so most packages describe the drawing in fields instead of shipping the file.
Do we need to mask our own part numbers?
Often yes, even when they identify only your products. Internal part numbers can reveal customers through prefixes, match public catalogs or expose a product line you treat as confidential. Tokenizing them costs little once the method is in place, and the family attributes keep their meaning.
Is masking enough to meet customer confidentiality terms?
Sometimes, but it depends on the wording. Some contracts restrict any use of information about the customer's parts, masked or not. The rights review reads those terms first and decides whether masking is enough or whether that customer's records stay out.
Sources
- Google's Sensitive Data Protection API supports format-preserving encryption and deterministic encryption, and recommends deterministic encryption where the input alphabet need not be preserved because FPE incurs significant latency costs. Source
- Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, so additional systems and protections should be employed. Source
- The Data Provenance Standards' Privacy Enhancing Tools code list includes masking, pseudonymization and tokenization, among other techniques. Source
Related resources
- QuestionDo AI labs buy images?
- InsightHow do I de-identify images and inspection photos for AI training?
- InsightHow do I de-identify inspection reports for AI training?
- InsightWhat compliance checks do food and beverage manufacturers need before licensing data to AI companies?
- SolutionData partnerships between businesses and AI developers
- IndustryBPO & contact centers data
See if your company qualifies
A short company assessment. No data uploads are needed.