Image data
EXIF Metadata in Image Training Data: What to Strip and What to Keep
Quick answer
Strip everything that locates, times or identifies a person or a specific device: the GPS IFD, owner and artist names, body and lens serial numbers, MakerNote blocks, embedded thumbnails, and XMP or IPTC location, creator and face-region fields. Keep, as sanitized sidecar columns, the fields that explain pixels: orientation (applied, then dropped), make, model, lens model, focal length, aperture, exposure, ISO and color space. Then audit the manifests and parquet files, not only the JPEGs, because metadata often survives there.
By SourceX Editorial · Updated
Why image metadata is a training-data risk, not just a privacy footnote
Image metadata is a second dataset riding inside every file, and it routinely carries personal data that the pixels do not. A 2025 audit of a large web-scraped ML image pool found non-empty EXIF tags covering timestamps, geolocation and named individuals, and noted that the dataset's download tool extracted EXIF tags for every sample, so the metadata was copied into side files at download time as well as staying in the originals [1]. Stripping the images alone would not have removed it.
For buyers of operational imagery (field inspections, claims photos, site surveys, retail audits), the exposure is sharper. A GPSLatitude/GPSLongitude pair on a homeowner's roof photo is a street address, and a DateTimeOriginal plus a BodySerialNumber can tie a frame to a named technician through a fleet asset register. Treat these as personally identifiable information under the same review as names in captions.
Metadata also biases models. Fields like Make, Model and LensModel correlate with site, contractor or era, which is useful for domain-shift analysis but becomes a shortcut feature if a pipeline ever feeds raw tags into a multimodal model's text stream.
Where identifying data hides: EXIF, XMP, IPTC and container chunks
Identifying fields live in at least four places, and a policy that only says "remove EXIF GPS tags" misses most of them [5]. Each layer is written by different software, so a phone, an editing tool and a digital asset management (DAM) export can each add their own copy of location or authorship.
- EXIF (TIFF-structured IFDs in JPEG APP1, TIFF, HEIF, the WebP EXIF chunk and the PNG eXIf chunk). The GPS IFD (GPSLatitude, GPSLongitude, GPSAltitude, GPSTimeStamp, GPSDateStamp, GPSImgDirection), CameraOwnerName, BodySerialNumber, LensSerialNumber, ImageUniqueID, Artist, Copyright, and the opaque MakerNote, where vendors store internal serials and sometimes face-detection data.
- IFD1 thumbnail. A small embedded JPEG preview. If a supplier blurred faces or license plates in the main image but did not regenerate the thumbnail, the unredacted original ships inside the file.
- XMP (RDF/XML packet). exif:GPS* duplicates, photoshop:City and photoshop:State, Iptc4xmpCore:Location, dc:creator, xmpMM:History (edit trail with software and timestamps), xmpMM:DocumentID/InstanceID, and MWG or Microsoft face regions (mwg-rs:Regions) that can carry person names with bounding boxes.
- IPTC IIM (Photoshop APP13). By-line, City, Sub-location, Province-State, Country, Caption-Abstract and Keywords, which in enterprise DAM exports often contain customer names, job numbers or addresses.
Two further layers matter for training data. PNG tEXt/iTXt chunks can hold free-text comments, and C2PA content credentials can carry a signed manifest with authorship and a training and data mining (TDM) assertion [6][7]. The TDM assertion is a rights signal, not personal data, so record it before stripping rather than discarding it.
The field policy: strip, transform or keep
The default should be "strip all, re-add an allowlist," because denylists fail open when a new phone firmware writes a tag nobody listed. Selective approaches exist in public releases: one dataset removed sensitive fields such as GPS before publication while keeping originals for trusted parties [2], and a utility drone imagery release removed EXIF entirely as part of anonymization [3]. Open-source tools now support GPS-only, full, or keep-a-safe-set modes that also preserve orientation [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
| Field or block | Layer | Action | Reason |
|---|---|---|---|
| GPSLatitude, GPSLongitude, GPSAltitude, GPSImgDirection | EXIF GPS IFD, XMP exif:GPS* | Strip; optionally replace with coarse region code in sidecar | Precise location identifies homes and workers |
| GPSDateStamp, DateTimeOriginal, OffsetTimeOriginal, SubSecTimeOriginal | EXIF, XMP | Transform: keep month or season and local hour bucket in sidecar | Exact timestamps link to shift rosters; coarse time supports lighting analysis |
| CameraOwnerName, Artist, By-line, dc:creator | EXIF, IPTC, XMP | Strip | Direct identifier |
| BodySerialNumber, LensSerialNumber, ImageUniqueID, MakerNote | EXIF | Strip; optionally replace serial with salted hash in sidecar | Serials link frames to people via asset registers |
| IFD1 thumbnail, XMP thumbnails | EXIF, XMP | Strip | May contain unredacted pixels |
| mwg-rs:Regions, MP:RegionInfo | XMP | Strip | Named face regions |
| City, Sub-location, Caption-Abstract, Keywords | IPTC, XMP | Strip from file; review text separately as a caption source | Free text often contains names and addresses |
| xmpMM:History, CreatorTool | XMP | Strip; record "edited: yes/no" in sidecar | Edit trail is useful for provenance, not for training |
| Orientation | EXIF 0x0112 | Apply to pixels, then drop | Prevents silently rotated training images |
| Make, Model, LensModel | EXIF | Keep in sidecar | Domain-shift and camera diversity analysis |
| FocalLength, FNumber, ExposureTime, ISO, Flash, WhiteBalance | EXIF | Keep in sidecar | Explains blur, noise and color cast |
| ICC profile, ColorSpace | ICC, EXIF | Keep in file or convert to sRGB and record | Color consistency across sources |
| C2PA manifest and TDM assertion | JUMBF | Record assertion in provenance log, then strip | Rights signal [6][7] |
The sidecar approach matters because the useful camera fields are also the ones that drift. Keeping Make, Model and exposure as separate columns lets you stratify evaluation by device, which pairs naturally with the analysis in camera, lens and lighting diversity and with coverage gap analysis.
Orientation: the tag you must apply before you strip
Orientation is the one EXIF field whose removal changes what the model sees. Phones usually store pixels in sensor order and set Orientation to 6 or 8 for portrait shots; a viewer rotates on display, but a training loader that ignores the tag feeds sideways images, and a loader that honors it behaves differently once the tag is gone. Some stripping tools preserve the tag by default for this reason [4].
The safe sequence is decode, apply the rotation to the pixel array (for example Pillow's ImageOps.exif_transpose), re-encode, then strip. Do this before annotation, because bounding boxes and polygons drawn on a rotated display will not match the stored pixel grid. If annotations already exist, ask the supplier which coordinate frame they used and check a handful of portrait images by overlay before accepting the delivery.
Re-encoding a JPEG to bake in rotation adds one generation of compression loss. Lossless JPEG rotation (jpegtran-style transforms on dimensions aligned to the JPEG MCU, 8 or 16 pixels depending on chroma subsampling) avoids it; otherwise record the re-encode quality setting alongside the resolution and compression requirements for the set.
Where metadata survives after you strip the images
Metadata most often leaks through the files around the images, not the images themselves. The web-scale audit is the clearest case: the download tooling wrote EXIF into per-sample records, so a stripped image could still sit next to a parquet row holding its coordinates [1].
Check these locations on every delivery:
- Manifests and shard metadata. Per-sample JSON in WebDataset tar shards, parquet columns, and COCO-style
imagesentries with extra keys. - Filenames and paths.
IMG_20260314_071522.jpgencodes a timestamp; DAM paths often contain job numbers, customer names or street names. - Annotation exports. CVAT, Label Studio and similar exports may copy source filenames and EXIF dates into task metadata.
- Original-format archives. HEIC and RAW (DNG, CR3, NEF) originals carry their own metadata and usually more MakerNote data than derived JPEGs.
- Burned-in overlays. Timestamp and GPS stamps printed on the pixels by inspection apps or dashcams are a redaction problem, not a metadata problem, and need pixel-level handling.
A verification gate for receiving an image delivery
Verification should be scripted, run on every file, and produce a report the privacy reviewer can sign. Sampling is fine for pixel review, but tag checks are cheap enough to run exhaustively.
Illustrative example: invented to show structure; it does not describe an available dataset.
METADATA ACCEPTANCE CHECK (per delivery)
1. Inventory: dump all tags per file (e.g., exiftool -json -G1 -a -u) to a log.
2. Fail if any file contains: GPS:* | XMP-exif:GPS* | *SerialNumber | OwnerName
| Artist | By-line | XMP-mwg-rs:* | IFD1 ThumbnailImage | MakerNotes:*
3. Fail if any IPTC/XMP free-text field is non-empty (City, Caption-Abstract, Keywords).
4. Orientation: assert tag absent or = 1; spot-check 50 portrait-origin files by overlay.
5. Sidecar: confirm columns make, model, lens_model, focal_length_mm, f_number,
exposure_s, iso, capture_month, hour_bucket, serial_hash; no raw serials or coordinates.
6. Side files: grep manifests, parquet and annotation exports for lat/lon patterns,
date-stamped filenames and serial formats.
7. Provenance log: record C2PA/TDM assertions found before stripping.
8. Sign-off: reviewer, tool versions, counts per failure class.
Record the outcome in the dataset card so downstream users know which fields were removed and which were transformed [10]. If you publish machine-readable metadata, a Croissant record can describe the sidecar's fields and file resources alongside the images [9]; see dataset cards for licensed enterprise data for what licensed sets usually document.
Health, minors and other regulated imagery
Regulated imagery needs a de-identification standard, not just a tag policy. For clinical or care-setting photos covered by HIPAA, de-identification follows 45 CFR 164.514 through either Safe Harbor or Expert Determination; Safe Harbor's identifier list reaches dates, geographic detail below the state level, device identifiers and serial numbers, and full-face photographs, so EXIF dates, GPS and serials all fall inside it [8].
Stripping metadata therefore satisfies only part of the standard: faces, tattoos, wristbands and screens in frame still need pixel work, and an Expert Determination report should name the metadata policy it relied on. For imagery that may include children, schools or homes, apply the strictest version of the table above and drop time and location sidecars entirely.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Writing the metadata policy into an image data request
The cheapest place to settle metadata handling is the request, before any supplier exports files. State the allowlist, the sidecar schema, whether orientation must be baked in, which formats you accept, and whether you need originals held back for audit. The image datasets hub covers the wider request; accepted file formats covers container choices.
SourceX sources operational datasets from US companies on request, including new recordings of hands-on work; categories are not inventory and a request does not guarantee a match. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. If you are scoping an image set, you can describe the data you need to SourceX, including your metadata policy.
Get image data delivered with a defined metadata policy
SourceX sources operational datasets from US companies on request, with every dataset rights-reviewed and delivered under a license that defines records, uses, term and delivery. Every release is approved by the supplying company, and delivery runs through private, access-controlled workflows after an executed agreement. Describe the images and metadata policy you need.
Frequently asked questions
Should I keep DateTimeOriginal for lighting and seasonality analysis?
Keep a derived value, not the raw tag. Month or season plus an hour bucket supports day/night and seasonal stratification without producing a timestamp that can be joined to a shift roster or a GPS track.
Is hashing the camera serial number enough?
Only with a secret salt held outside the dataset. Unsalted hashes of serial numbers are reversible by enumeration because serial formats are short and structured, and the hash still allows linking all frames from one device.
Do PNG and WebP files need the same checks as JPEG?
Yes. Both formats can carry EXIF and XMP in dedicated chunks, and PNG adds free-text tEXt/iTXt chunks, so run the same inventory and fail rules regardless of extension.
Sources
- arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
- arXiv, "\"ScatSpotter\" -- A Dog Poop Detection Dataset" (2024). https://arxiv.org/pdf/2412.16473
- IEEE DataPort, "EPRI Distribution Inspection Imagery". https://ieee-dataport.org/open-access/drone-based-distribution-inspection-imagery
- PyPI, "pheser". https://pypi.org/project/pheser/
- BlurMe, "Remove EXIF and GPS metadata from photos". https://www.blur.me/blog/remove-exif-gps-metadata-from-photos/
- IPTC, "CAWG Training and Data Mining Assertion". https://metawatch.iptc.org/ai-policy/cawg-training-mining/
- C2PA, "Clarification to C2PA TDM assertions reference". https://c2pa.org/c2pa-clarification-to-c2pa-tdm-assertions-reference/
- eCFR, Office of the Federal Register / HHS, "45 CFR 164.514 - Other requirements relating to uses and disclosures of protected health information". https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
- arXiv (MLCommons Croissant working group), "Croissant: A Metadata Format for ML-Ready Datasets" (2024). https://arxiv.org/pdf/2403.19546
- Hugging Face, "Create a dataset card (Datasets library documentation)". https://huggingface.co/docs/datasets/v2.19.0/en/dataset_card
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.