Image data
DICOM De-identification for AI Training Data: Header Attributes, Burned-In Pixel PHI and What Annex E Does Not Cover
Quick answer
For AI training data, require the DICOM PS3.15 Annex E Basic Application Level Confidentiality Profile plus named options, especially Clean Pixel Data, Clean Descriptors and an explicit choice on UIDs, dates and private tags, and require the supplier to record what it applied in each file. Annex E governs header attributes well, but the standard itself says its profiles do not guarantee removal of all identifying information [1]. Burned-in text, faces in volumetric scans and free-text fields need separate controls and a human-checked sample.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What Annex E actually specifies
Annex E is a per-attribute rulebook: Table E.1-1 lists DICOM attributes and assigns each an action under the Basic Profile and under each option [1]. The actions are short codes: D (replace with a non-zero-length dummy value), Z (replace with zero length or a dummy), X (remove), K (keep), C (clean, meaning replace identifying content with similar non-identifying values) and U (replace UIDs with internally consistent new ones), plus conditional combinations such as X/Z or X/Z/D that depend on the attribute's type in the IOD [1]. The Basic Profile is deliberately conservative; it was designed for clinical trials, teaching files, publications and registry submissions, not for training-data utility [2].
That conservatism is the main tension for model builders. The Basic Profile removes or blanks many attributes a model team wants, such as acquisition dates, device serial numbers and institution names. The options exist to keep some of them under stated conditions, so the buyer's job is to choose options deliberately rather than accept "Annex E compliant" as a single setting.
The options you will meet most often in imaging deals are:
- Clean Pixel Data: remove burned-in identifying text from pixel data.
- Clean Recognizable Visual Features: address features such as faces that can identify a patient.
- Clean Graphics and Clean Structured Content: scrub overlays, presentation-state graphics and structured report content.
- Clean Descriptors: clean free-text attributes like Study Description, Series Description and Image Comments rather than keeping them verbatim.
- Retain Longitudinal Temporal Information (full or modified dates), Retain Patient Characteristics, Retain Device Identity, Retain Institution Identity, Retain UIDs and Retain Safe Private: each relaxes the Basic Profile for one class of attributes.
Choosing options for a training dataset
The right option set keeps what the model needs for label validity and stratification while removing what links a study to a person. A segmentation model for chest CT needs Pixel Spacing, Slice Thickness, Image Orientation (Patient), Rescale Slope and Intercept, Convolution Kernel and Manufacturer Model Name; none of those identifies a patient, and most survive the Basic Profile. A longitudinal progression model also needs time between studies, which is where Retain Longitudinal Temporal Information with modified dates (a consistent per-patient shift) matters.
Patient's Age, Patient's Sex, Patient's Size and Patient's Weight are only retained under Retain Patient Characteristics [1]. If your evaluation plan stratifies by age band or sex, ask for that option explicitly and then check whether ages above 89 are aggregated, because HIPAA Safe Harbor treats ages over 89 as an identifier element [6]. Device Serial Number and Station Name usually add little to model quality and add linkage risk, so leave Retain Device Identity off unless you are studying scanner drift at the unit level.
Institution Name and Institution Address are useful for site-level generalization studies but can narrow a patient population. A common compromise is to replace the institution with a stable pseudonymous site code in a sidecar manifest rather than retaining the real name in the header. For dataset-level sourcing questions beyond de-identification, see how to source licensed medical imaging datasets.
Private tags, UIDs and File Meta Information
Private tags, UIDs and the Part 10 file wrapper are where well-intentioned header scrubbing most often leaks. Vendors store data in odd-group private elements, for example (0019,xxxx) or (0029,xxxx) blocks, and some of them carry patient names, accession numbers, protocol notes or embedded CSA headers. The Basic Profile removes private attributes; Retain Safe Private keeps only those known to be safe, ideally documented through the Private Data Element Characteristics Sequence [1]. Ask the supplier which private creators it whitelisted and why, because a blanket keep is a frequent failure mode.
UID remapping must be consistent, not random per file. Study Instance UID, Series Instance UID, SOP Instance UID and Frame of Reference UID need deterministic replacement within the dataset so that series stay grouped, registrations still align and references inside Structured Reports, segmentation objects (SEG), RT Structure Sets and Presentation States still resolve [1][3]. A mapping table that sits with the supplier, never with the buyer, keeps re-linkage possible for corrections without exposing it to you. New UIDs generated under the 2.25 UUID-derived root avoid embedding the original organization's registered root.
Annex E also warns that the File Meta Information group (0002,xxxx) and the 128-byte Part 10 preamble can carry identifying content and should be replaced rather than copied [1]. Source Application Entity Title and Implementation Version Name reveal sending systems, and some preambles contain other file headers. Re-writing files with a fresh meta header is the safe default.
Burned-in pixel PHI: where header profiles stop
Header de-identification does nothing to text rendered into pixels, so burned-in PHI needs its own detection and redaction step. The IHE handbook treats pixel data separately from attribute handling for exactly this reason [3]. High-risk sources include ultrasound, where patient name, ID and date are often rendered in a top banner; secondary capture (SC) images and screenshots of 3D workstations; scanned paper requisitions and consent forms stored as DICOM; dose report screens from CT; and some fluoroscopy, endoscopy and dental captures.
Do not rely on the Burned In Annotation attribute (0028,0301) alone. Many modalities populate it inconsistently or leave it empty, so a filter of "Burned In Annotation = NO" is a weak control. A defensible pipeline combines three elements:
- Modality- and model-specific rules that blank known banner regions (for example a fixed rectangle on a given ultrasound model's output).
- OCR-based text detection across all frames, including multi-frame cines, with a low threshold and human review of hits.
- A stratified human QA sample per modality, manufacturer and SOP Class, reported with the number of images reviewed and the number of residual findings.
If any image is masked, ask how. Pixel redaction in compressed transfer syntaxes such as JPEG Baseline or JPEG 2000 requires decompress, mask and re-encode; a tool that edits only the header leaves the original pixels untouched. For the photographic side of image privacy, such as EXIF in camera images, see EXIF metadata in image training data.
Faces, free text and other residual risk
Some identifiers live in the anatomy or in content Annex E handles only partially. Head CT and MR volumes can be rendered into recognizable face surfaces, so neuroimaging datasets typically need defacing or skull-stripping under Clean Recognizable Visual Features, and the buyer should test that the defacing does not damage the region the model needs. Free-text fields, Structured Report content and embedded PDF reports (Encapsulated PDF) need the same treatment as clinical notes; see de-identifying clinical free text for LLM training.
Re-identification research shows that removing direct identifiers does not reduce risk to zero, especially when rare conditions, small sites or linkable dates remain [8]. Treat the de-identification report as one input to a re-identification risk assessment rather than as a final answer.
Mapping Annex E to HIPAA in the agreement
Annex E is a technical standard, not a HIPAA method, so the agreement should state which HIPAA path the dataset meets and how the DICOM options support it. HIPAA offers two routes in 45 CFR 164.514(b): Safe Harbor, which removes 18 identifier types and requires no actual knowledge that the remainder could identify someone, and Expert Determination, in which a qualified expert finds the risk very small [6][7]. Retaining modified dates, full ages or site codes usually pushes a dataset toward Expert Determination; the trade-offs are covered in Safe Harbor vs Expert Determination for AI training data.
A limited data set under 164.514(e) is a different instrument: it is still PHI and requires a data use agreement [7]. Buyers should not confuse "dates retained" with "de-identified." Glossary background is in Safe Harbor de-identification, expert determination and PHI.
Specifying and verifying the supplier's process
Write the de-identification requirements into the request, then verify them on delivered files rather than on a policy document. Each object should carry Patient Identity Removed (0012,0062) = YES, a De-identification Method (0012,0063) text, and a De-identification Method Code Sequence (0012,0064) listing the profile and options from CID 7050, as the IHE handbook recommends [3]. That makes the applied method machine-checkable across millions of files.
Tool choice matters. A study of free DICOM de-identification tools found they differ in function and in how safely they protect patient privacy [4], and some libraries describe themselves as best-effort implementations rather than reference implementations of the profile [5]. Ask which tool and version ran, what configuration file was used, and whether the supplier validated it against test objects containing known PHI in headers, private tags and pixels.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to ask for | How to verify on delivery |
|---|---|---|
| Base profile | Basic Application Level Confidentiality Profile | 113100 present in (0012,0064) on every object |
| Pixel data | Clean Pixel Data with OCR detection plus human QA | QA report by modality and SOP Class; spot-check ultrasound and SC |
| Faces | Defacing for head CT/MR volumes | Render 3D surface on a sample; confirm target anatomy intact |
| Dates | Modified dates, consistent shift per patient | Intervals preserved; no real dates in any VR=DA field |
| Patient characteristics | Age, sex retained; ages over 89 grouped | Distribution check on Patient's Age |
| UIDs | Consistent remap, new root | Series grouping and SEG/SR references resolve |
| Private tags | Removed, or whitelisted safe private only | Dump of remaining odd-group elements with private creators |
| File wrapper | New File Meta Information and preamble | Inspect group 0002 and first 128 bytes |
| Free text | Clean Descriptors; SR and Encapsulated PDF scrubbed or excluded | Text search for names, MRNs, accession patterns |
| Documentation | Tool, version, config, validation results | Matches (0012,0063) text and evidence package |
The written evidence behind each row belongs in a de-identification evidence package, and the related contract language is covered in re-identification prohibition clauses.
Where SourceX fits in DICOM sourcing
SourceX sources operational datasets from US companies on request and manages licensing and ongoing purchases; it does not hold imaging in stock, and a request does not guarantee a match. Health records require HIPAA de-identification by Safe Harbor or Expert Determination, every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery, and diligence materials on source, rights and preparation are prepared per dataset. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Buyers can describe the imaging data they need, and background on medical data is on licensing medical records for AI training and whether AI labs buy medical data. More image topics are in the image data hub and the AI data guides.
Request de-identified DICOM data for model training
Describe the modalities, SOP Classes, Annex E options and HIPAA method you need, and SourceX looks for US businesses that hold matching data. Every release is approved by the supplying company, and nothing is contracted until a supplier agrees. Start your request at sourcex.si/buyers.
Sources
- NEMA / DICOM Standards Committee, "DICOM PS3.15 Security and System Management Profiles, Annex E: Attribute Confidentiality Profiles (current edition)" (2026). https://dicom.nema.org/medical/dicom/current/output/chtml/part15/chapter_E.html
- NEMA / DICOM Standards Committee, "DICOM PS3.15 2022a, E.2 Basic Application Level Confidentiality Profile" (2022). https://dicom.nema.org/medical/dicom/2022a/output/chtml/part15/sect_E.2.html
- IHE International, "IHE De-Identification Handbook: DICOM example". https://profiles.ihe.net/ITI/DeId/dicom-example.html
- European Radiology, "Free DICOM de-identification tools in clinical research: functioning and safety of patient privacy". https://www.european-radiology.org/?p=2108
- HexDocs (dicom library), "Dicom.DeIdentification module documentation". https://hexdocs.pm/dicom/Dicom.DeIdentification.html
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- eCFR, Office of the Federal Register, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.