Privacy, de-identification and sensitive data
Dates, ages and ZIP codes under HIPAA Safe Harbor: what temporal and geographic models lose
Quick answer
HIPAA Safe Harbor keeps the year and strips every more specific date element tied to a patient: day, month, admission, discharge, birth and death [1]. It collapses ages over 89 into one "90 or older" bucket and allows only three-digit ZIP prefixes whose combined population exceeds 20,000 [1][2]. Date shifting and "days since index" offsets are not named as exceptions in the rule. If your model needs intervals, seasonality or sub-state geography, plan for Expert Determination rather than assuming Safe Harbor output will be usable [2][3].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What the Safe Harbor date rule actually removes
The date rule removes all elements of dates, except year, for dates directly related to an individual [1]. The regulation's examples are birth date, admission date, discharge date and date of death, but in operational records the same logic reaches service dates, claim received dates, prescription fill dates, appointment timestamps and the time component of any of them [1][2]. A Safe Harbor dataset therefore arrives with a service_year column where your feature pipeline expected service_ts.
The practical failure modes are predictable:
- Intervals disappear. Length of stay, time to readmission, days between denial and appeal, and claim lag cannot be computed from two year values.
- Ordering becomes ambiguous. Events in the same year lose their sequence unless the supplier provides an ordinal that is itself defensible under the rule.
- Seasonality collapses. Influenza-season volume, month-end billing cycles and quarter-end payer behavior all live below year granularity.
- Leakage controls weaken. Without event dates you cannot build a clean as-of cutoff, so label information can bleed into features.
Dates not related to an individual, such as a fee-schedule effective date or a software release date, are not the target of the rule [1][2]. Ask the supplier to separate those administrative dates from patient-linked dates in the data dictionary so they are not stripped unnecessarily.
Ages over 89 and the hidden year problem
Safe Harbor requires removing all ages over 89 and all elements of dates, including year, that would indicate such an age, though those ages may be aggregated into a single category of 90 or older [1]. The second half of that sentence is the one engineers miss. A retained birth_year of 1931 in a 2024 encounter reveals an age over 89, so the birth year itself must go or be top-coded.
For modeling, the 90-plus bucket is usually tolerable for risk scores, because the population is small and outcomes are often modeled as a single elderly stratum. It is less tolerable for geriatric, end-of-life or long-term care use cases, where that bucket is the target population. If that describes your project, state it in the request so the supplier can evaluate whether Expert Determination is warranted.
The three-digit ZIP rule and what geographic models keep
Safe Harbor removes all geographic subdivisions smaller than a state, including street address, city, county, precinct and five-digit ZIP code [1]. The one carve-out is the initial three digits of a ZIP code, allowed only if the geographic unit formed by combining all ZIP codes with those same three digits contains more than 20,000 people; otherwise the prefix must be changed to 000 [1]. The rule ties that population test to current publicly available Census Bureau data, and OCR's guidance explains how to apply it [1][2].
What a geographic model keeps under Safe Harbor:
- State, which is not a subdivision smaller than a state.
- A three-digit ZIP prefix for most of the country, with sparsely populated prefixes masked as 000.
- No county FIPS, census tract, hospital referral region or drive-time features, unless an expert has evaluated them.
The masked prefixes are not random. They cluster in rural and frontier areas, so a model trained on Safe Harbor geography has the least spatial resolution exactly where access-to-care and rural-outcome questions matter most. Treat "000" as a distinct category, never as missing, and audit performance on it separately. The linkage research behind these rules is sobering: Rocher and colleagues' generative model estimated that 15 demographic attributes, such as ZIP code, birth date and gender, would correctly re-identify 99.98% of Americans in any dataset, a model-based estimate rather than a count of actual re-identifications [4].
Date shifting and relative time: utility versus the rule text
Date shifting moves every date for a patient by a random offset, usually a per-person value within a fixed window, so intervals and sequence inside that patient's record are preserved. Relative-time encoding replaces dates with offsets from an index event, such as days_since_first_encounter. Both are standard engineering techniques, and both preserve the features most temporal models need.
Neither appears as an exception in the Safe Harbor text, which lists permitted retentions narrowly: year, ages aggregated at 90-plus, qualifying three-digit ZIP prefixes and a re-identification code that meets 164.514(c) [1]. A shifted date is still a full date derived from the true date, and an interval can itself be identifying when it is rare or linkable, such as a 400-day inpatient stay. That is why these techniques are commonly documented under Expert Determination, where a qualified expert evaluates the actual transformation and the anticipated recipient [2][3].
Three engineering points matter whichever route the supplier uses:
- Shift per person, not per row. Per-row shifting destroys sequence and gives you noise, not privacy.
- Keep the offset key with the supplier. The offset table is a re-identification key; you should never receive it.
- Check the weekday and holiday artifacts. A shift that is not a multiple of seven moves weekend admissions onto weekdays, which corrupts staffing, day-of-week and holiday features. Ask whether the shift preserves weekday.
Choosing a route: what each method can preserve
The method should follow from the temporal and geographic resolution the model needs, not the other way around. The comparison below summarizes typical outcomes; an Expert Determination is case-specific and may reject a transformation [2][3].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field your model uses | Safe Harbor output [1] | Often retained under Expert Determination [2][3] | Modeling impact if lost |
|---|---|---|---|
| Encounter date | Year only | Month, shifted date or relative day | No intervals, no seasonality |
| Length of stay | Not derivable from year | Integer days, possibly top-coded | Lose a core outcome and feature |
| Age | Integer to 89, then "90+" | Similar, sometimes finer banding | Weak geriatric stratification |
| Birth date | Year only, removed if implies 90+ | Year or age band | Minor for most models |
| Geography | State plus qualifying ZIP3, else 000 | ZIP3 or wider region, sometimes county | Rural signal collapses |
| Event order within year | Not guaranteed | Preserved via shifted or relative time | Sequence models fail |
For a deeper method comparison, see Safe Harbor vs Expert Determination for AI training data, and before relying on an expert's report, read how to review a HIPAA Expert Determination report. If you do not need de-identified status at all, a HIPAA limited data set under a data use agreement can retain dates and five-digit ZIP codes under contractual controls [1].
Specifying temporal granularity in a health data request
The cheapest fix is to ask only for the resolution your model uses. Every extra unit of date or location precision raises re-identification risk and narrows the set of suppliers who can say yes. Work backward from the feature list, then write the request field by field.
Illustrative example: invented to show structure; it does not describe an available dataset.
request: de-identified outpatient encounter and claims records
use_case: 30-day readmission and denial-to-appeal forecasting
deid_route_requested: expert_determination # safe_harbor acceptable if fields below are met
temporal_fields:
encounter_time: relative_days_from_first_encounter # integer, per-person index
seasonality: service_month # 1-12, no day
weekday_preserved: required
claim_lag_days: integer, top-coded at 365
age:
representation: integer_years, 90_plus_bucket
geography:
representation: zip3_or_000
wider_region_acceptable: true
excluded: [free_text_notes, five_digit_zip, exact_dates, offset_keys]
documentation_requested: [data_dictionary, deid_method_statement, expert_report_summary]
When reviewing the delivered data, check that relative day counts start at zero for the index event, that no free-text field such as a note or claim memo carries an exact date, and that the dictionary flags which date columns were transformed and how. Free text is the usual way exact dates leak past a structured-field transformation. Also confirm the temporal coverage, date gaps and seasonality of the file you received match what you scoped.
Where SourceX fits for temporal health data
SourceX sources operational datasets from US companies on request, and buyers can describe needs such as healthcare operational records; categories are not inventory, nothing is held in stock and a request does not guarantee a match. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, personal details are removed or replaced before delivery with the method recorded, and every release is approved by the supplying company. Buyers can describe the temporal and geographic fields they need on the SourceX buyers page. For the wider cluster, see the privacy and de-identification hub and the Safe Harbor glossary entry.
Request health records scoped for temporal models
If your forecasting or process model depends on intervals, seasonality or regional geography, describe those fields and the de-identification route you need. SourceX looks for US businesses that hold the data, rights-reviews each dataset and delivers it under a license that defines records, uses, term and delivery. Start your request on the SourceX buyers page.
Frequently asked questions
Can a Safe Harbor dataset include the patient's age in years?
Yes, up to 89. Ages over 89 must be removed or aggregated into a single 90-or-older category, and any date element, including year, that would reveal such an age must also go [1].
Is a five-digit ZIP code ever allowed under Safe Harbor?
No. Only the first three digits may be kept, and only where the combined population for that prefix exceeds 20,000; otherwise the prefix becomes 000 [1].
Does a random per-patient date shift make data Safe Harbor compliant?
The Safe Harbor text does not list shifted dates as a permitted retention, so treat shifted or relative dates as something to document under Expert Determination unless counsel concludes otherwise [1][2].
Sources
- Electronic Code of Federal Regulations (eCFR), Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- Tonic.ai, "HIPAA AI compliance guide". https://www.tonic.ai/guides/hipaa-ai-compliance
- Rocher, Hendrickx, de Montjoye (Nature Communications), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.