Logistics and distribution
How to remove shipper, consignee and carrier identities from freight records
By SourceX Editorial · Reviewed by Noah Loul ·
Short answer
To anonymize shipment data, remove or replace every field that identifies a shipper, consignee or carrier: names, addresses, contacts, reference numbers, carrier identifiers, free text and documents such as bills of lading. Use consistent tokens so loads still join, generalize locations and dates, then test whether lanes, volumes or commodities still point to one company.
Key takeaways
- Names are the easy part; reference numbers, free text and documents leak most identities.
- Consistent tokens keep orders, stops, invoices and exceptions joined after names are removed.
- Generalize locations and shift dates consistently so a facility cannot be inferred from a lane or schedule.
- Automated detection tools help but do not catch everything, so add custom rules and human review.
- Residual-risk checks look for unique combinations of fields, not just leftover names.
Which fields identify shippers, consignees and carriers?#
The fields that identify shippers, consignees and carriers fall into five groups: names, locations, contacts, reference numbers and free text. Most freight records carry all five, spread across orders, stops, invoices, tracking events, claims and attached documents.
IT teams usually find the named fields quickly. The harder work sits in identifiers that look neutral, such as PRO numbers, MC and USDOT numbers, SCAC codes, seal, container and trailer numbers, and in notes where a dispatcher typed a receiver's name, a dock instruction or a shipper's internal project code.
The field-by-field removal checklist#
Ten field groups in a typical TMS, WMS or brokerage system carry party identities, and each needs a set treatment: replace with tokens, generalize, remove or review by hand. Treatments are starting points; the rights review and the purpose of the dataset decide the final choice for each field.
| Field group | Examples | Treatment |
|---|---|---|
| Party names | Shipper, consignee, bill-to, broker, carrier, notify party | Replace with consistent tokens per party |
| Addresses and locations | Street address, facility name, dock, geocodes | Generalize to region or metro area; drop geocodes |
| Contacts | Names, phones and emails of dispatchers, receivers and drivers | Remove |
| Shipment references | BOL, PRO, PO, load, appointment and pickup numbers | Replace with surrogate IDs that keep joins |
| Carrier and equipment identifiers | MC and USDOT numbers, SCAC, tractor and trailer numbers, plates, VINs | Remove or tokenize; never keep raw |
| Commercial terms | Line-haul rates, accessorials, fuel surcharge, customer-specific charges | Keep only if rights allow; otherwise band or remove |
| Commodity and product | Item descriptions, SKUs, branded product names | Generalize to commodity class where a product names its maker |
| Dates and times | Pickup and delivery appointments and actuals | Shift consistently per party, or round |
| Free text | Special instructions, dispatcher notes, check-call comments, emails | Scan, redact, then sample by hand |
| Documents | BOL and POD images, signatures, stamps, letterheads | Exclude, or redact after text extraction with manual review |
How do you keep loads joined after removing names?#
Loads stay joined after names are removed when every identifier is replaced with a consistent token: the same shipper always becomes the same token, and each load ID maps to exactly one surrogate. That preserves the chain from order to stop to invoice to exception, which is where most of the analytical value of freight records sits.
Deterministic encryption or keyed hashing produces consistent tokens without a lookup table ever leaving the company. Google's Sensitive Data Protection API, for example, supports both deterministic and format-preserving encryption, and recommends deterministic encryption where the original character set need not be preserved, citing latency costs. Whatever the tool, keep the keys inside the company and out of any delivered package.
Dates need the same discipline. Shift all dates for a given party by the same offset so intervals such as dwell time, transit time and detention stay accurate. The same Google API includes a transform that shifts dates by a random number of days; whichever tool you use, confirm it applies one consistent offset per party rather than a new offset per row.
Free text and documents need their own pass#
Free text and documents need their own pass because identities in them do not sit in labeled fields. A dispatcher note may say that a named receiver refused a pallet at a named dock, and a bill of lading image carries letterhead, signatures and handwritten exceptions.
Automated PII detection is a useful first layer. Presidio, an open-source tool now maintained by the community-governed Data Privacy Stack organization, pairs named-entity recognition with regular expressions and rules, and its maintainers say plainly that automated detection can miss sensitive information and should be layered with other safeguards. Add custom rules built from your customer, carrier and vendor master files, then review samples by hand.
For documents, the practical choice is often to exclude images and keep the structured data already extracted from them. Where documents are essential, redact after text extraction and check every page layout manually, since each shipper's bill of lading looks different.
Residual-risk checks before release#
Residual-risk checks test whether a prepared dataset can still point to a specific shipper, consignee or carrier through combinations of the remaining fields. A unique lane, an unusual commodity and a regular pickup schedule can identify a shipper as surely as its name.
- Count records per origin region, destination region and commodity class; small groups may identify a party.
- Look for lanes served by a single shipper or a single carrier in the dataset.
- Check that rate bands cannot be matched to a known contract or tariff.
- Search free text for every name in the customer, carrier and vendor master files.
- Confirm that tokens and shifted dates cannot be reversed without keys held inside the company.
- Ask someone who knows the business to try to identify major customers from a prepared sample.
Which formal methods support the review?#
Formal methods support the review by measuring how many records share each combination of quasi-identifiers such as region, commodity and date. Metrics like k-anonymity and l-diversity are built into mainstream tooling; Google's Sensitive Data Protection API, for instance, includes risk analysis for k-anonymity and l-diversity as well as k-map estimation and delta-presence estimation.
Shared vocabulary helps document the work. The Data Provenance Standards list masking, redaction, pseudonymization, tokenization and other privacy-enhancing tools as metadata values, which gives buyers and suppliers a common way to record which techniques were applied to which fields.
Illustrative: a 3PL prepares exception records#
Illustrative: a fictional 3PL serving both B2B and direct-to-consumer clients wants to assess its delivery exception and damage records. The IT director builds a token map from the client and carrier masters, removes consumer ship-to names and addresses entirely, and generalizes business consignees to region.
Dates are shifted per client, free text is scanned and then sampled, and proof-of-delivery images are excluded. The residual check finds that one client is the only shipper of a specialty commodity in its region, so that commodity is generalized one level further. An account manager then fails to identify any client from the sample, and the result goes into the preparation notes.
How SourceX handles identity removal#
SourceX handles identity removal in the Preparation step of the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. Personal and confidential details are removed before release, and the supplier reviews the result in the Approval step.
The privacy record in the SourceX Evidence Packet lists the fields removed or transformed, the techniques used and the residual-risk checks performed, so both sides can see how identities were handled.
Frequently asked questions
Is tokenized freight data still personal data?
It can be. Where records include people, such as drivers, receivers or consumer ship-to names, tokenized data may still count as personal data under laws such as GDPR while keys exist, and US state privacy laws draw their own lines between de-identified and pseudonymous data. For company names the main concern is usually confidentiality, and counsel assesses both.
Should we remove rates as well as names?
Often, at least in part. Rates are commercially sensitive to shippers, brokers and carriers, and contracts may protect them. Where rates matter to the dataset, banding or indexing them keeps the pattern while removing the exact figure. The rights review, not preparation alone, makes that call.
Can we keep ZIP codes?
Full ZIP codes on pickup and delivery points can identify facilities, especially industrial sites with few neighbors. Most preparations generalize to a coarser area such as a metro or region and drop geocodes. Recheck small groups afterward, because a rural region with one large shipper can still point to it.
Do we need to remove our own company name?
The supplier's identity is usually known to the licensee through the license itself, so removing it is less about secrecy than consistency. Remove it where it would reveal counterparties, such as facility names or email signatures, and follow what the license says about attribution.
Should preparation happen inside our own systems?
Usually, yes. Running tokenization, redaction and residual checks in company-controlled storage means raw records never leave your environment, and only the prepared copy moves. Keep the token keys, the field map and the preparation scripts together so the process can be repeated for later data refreshes.
Sources
- Google's Sensitive Data Protection API supports de-identification transforms including format-preserving encryption, deterministic encryption and date shifting by a random number of days, and recommends deterministic encryption over format-preserving encryption where the input alphabet need not be preserved because FPE incurs significant latency costs. Source
- Google's Sensitive Data Protection API offers four re-identification risk-analysis metrics: k-anonymity, l-diversity, k-map estimation and delta-presence estimation. Source
- Presidio is an open-source, MIT-licensed SDK for PII identification and anonymization that combines named-entity recognition, regular expressions, rule-based logic and checksums; its documentation warns there is no guarantee it will find all sensitive information and that additional systems and protections should be employed. Source
- Presidio moved from a Microsoft-owned project to an independent, community-governed open-source project under the Data Privacy Stack GitHub organization. Source
- The Data Provenance Standards' Privacy Enhancing Tools code list includes anonymization, encryption, masking, minimization, redaction, pseudonymization, tokenization, k-anonymity and other techniques. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.