Skip to content

Multimodal and embodied data

De-Identifying Multimodal Records: Faces, Voices, Screens and Metadata Together

Quick answer

To de-identify multimodal data, treat each record as one identity surface rather than separate files. Build an inventory of every channel that can carry identity (face pixels, voice timbre, names spoken in audio, on-screen and in-frame text, transcripts, captions, container and Exif metadata), redact each with a method sized to the risk, then propagate every redaction across all time-aligned streams. Finish by auditing a sample with cross-modal attacks and recording the method per modality, because single-modality tools miss leaks that cross channels.

By SourceX Editorial · Updated

This page covers the cross-modal problem. For single-modality depth, see the SourceX insights on de-identifying voice and audio data, de-identifying screen recordings and de-identifying egocentric video of manual work. The broader cluster lives at the multimodal training data hub.

Why single-modality redaction fails on multimodal records

Single-modality redaction fails because identity leaks through whichever channel the pipeline did not process. A face-blurred meeting clip still identifies a participant when someone says "thanks, Priya" on the audio track, when the video-conference tile shows a display name, or when the MP4's com.apple.quicktime.location.ISO6709 atom places the recording at a home address. Each tool reports success on its own stream while the record as a whole stays identifiable.

The failure modes repeat across sources. Common ones include:

  • Spoken names and numbers that ASR transcribes into the caption track or a sidecar .vtt file the video pipeline never opens.
  • Rendered text in frame: badges, lanyards, whiteboards, email clients, CRM screens, Slack sidebars, license plates and shipping labels.
  • Voice as a biometric: even with names bleeped, the raw voice can be matched to a speaker-verification enrollment.
  • Metadata: Exif GPS, device serials, owner names in camera settings, file paths such as /Users/jsmith/, and calendar titles embedded in meeting exports.
  • Linkage across streams: timestamps, room IDs and session IDs that join a de-identified clip back to a calendar entry or badge-swipe log.

Recent audits show the metadata risk is not theoretical. A 2025 analysis of a large web-scraped machine-learning image dataset found non-empty Exif tags relating to timestamps, geolocation and individuals, many disclosing full names, and noted that the download tool extracted those tags for every sample [3]. Operational recordings from companies tend to carry richer metadata than web images, not less.

Building the identity inventory for one record

The identity inventory is a per-record list of every channel, its identifiers and its time alignment, and it should exist before any tool runs. Without it, redaction coverage is whatever the tooling happened to support. Start from the container: list every video, audio, subtitle, data and attachment track (ffprobe -show_streams exposes them), every sidecar file (transcripts, chat logs, telemetry, annotation JSON) and every metadata layer (Exif, XMP, QuickTime atoms, Matroska tags).

For each channel, record three things. First, the direct identifiers it can carry (face, voice, name, account number, email, plate). Second, the quasi-identifiers (location, timestamp, employer logo, distinctive tattoo, accent plus job title). Third, the clock it runs on, because redaction has to be mapped between streams that may drift. If the source cannot document clock alignment, fix that first; the time synchronization checks for multimodal datasets apply directly here, since a redaction offset by 400 ms leaks the first syllable of a name.

Choosing methods per modality without destroying training value

The right method for each modality is the weakest transformation that removes identity for the intended release, because heavier masking costs model quality. Evidence from vision is the clearest. Hukkelås and Lindseth trained detectors on anonymized versions of standard datasets and evaluated on originals; traditional blur and mask anonymization hurt performance noticeably, full-body anonymization cost more than face-only, and realistic generative anonymization narrowed the gap [1].

Strength matters as much as scope. Research quantifying blur for image release reports that weak blur settings can leave faces recoverable or recognizable, so a kernel tuned to look anonymous to a human reviewer may not resist a recognition model [2]. For release outside the supplier's control, prefer solid masking or generative face replacement over light Gaussian blur, and record the kernel or model version.

Voice has its own tradeoff. The VoicePrivacy 2024 Challenge evaluates speaker anonymization against speaker-verification attackers, including attackers that know the anonymization system, while measuring utility with ASR word error rate and emotion recognition [4]. Buyers training speech or paralinguistic models should ask which attacker model the supplier tested against, not only whether voices were "converted."

ModalityTypical identifiersLower-utility-cost optionStronger option for wide releaseUtility to check afterward
Face videoFace, tattoos, gaitFace-only masking or generative replacementFull-body generative replacementDetection, pose, action recognition
VoiceTimbre, prosody, accentSpeaker anonymization tested against ASVRe-synthesis from transcript (TTS)WER, emotion, turn-taking
Spoken contentNames, numbers, addressesBleep or tone over NER-detected spansBleep plus transcript token replacementTranscript alignment, dialogue coherence
Screen and in-frame textNames, emails, account IDsOCR-driven box maskingFull UI region maskingUI grounding, OCR targets
MetadataGPS, serials, owner names, pathsAllowlist of kept fieldsFull strip with re-derived timestampsSync and ordering

The table is a starting point, not a standard. Ego4D shows the pattern at scale: consenting participants, with de-identification applied where needed rather than uniformly [10].

Propagating one redaction across synchronized streams

Consistent redaction means every detected identity event becomes a time-coded span that every stream honors, not a per-stream decision. When NER on the transcript finds a person name at 00:12:04.310 to 00:12:04.870, the pipeline should bleep the audio over that span, replace the token in the transcript and captions with a consistent surrogate ([PERSON_07]), and check the video frames for the same name on screen. When OCR finds an email address in a shared-screen frame, the same surrogate should appear wherever that person is mentioned in audio.

Three engineering rules prevent most cross-modal leaks:

  1. One surrogate map per record (or per project, if linkage is licensed). The same real person gets the same token in transcript, captions, chat and annotations, so dialogue stays coherent and nothing re-links by mismatch.
  2. Pad spans. Extend audio bleeps and video masks by a small margin on both sides to absorb clock drift and ASR boundary error.
  3. Redact derived artifacts last. Thumbnails, keyframe sprites, waveform previews, embeddings and ASR lattices are often generated before redaction and shipped by accident.

Illustrative example: invented to show structure; it does not describe an available dataset.

record_id: rec_000417
streams:
  video_main:   {codec: h264, clock: pts, offset_ms: 0}
  audio_mix:    {codec: aac, clock: pts, offset_ms: 0}
  screen_share: {codec: h264, clock: pts, offset_ms: 212}
  transcript:   {format: webvtt, clock: asr_word_times, offset_ms: 0}
  metadata:     {layers: [quicktime_atoms, xmp]}
redactions:
  - id: r1
    trigger: ner_person_in_transcript
    span_ms: [724310, 724870]
    pad_ms: 150
    surrogate: "[PERSON_07]"
    applied_to: [audio_mix:bleep, transcript:replace, captions:replace]
  - id: r2
    trigger: ocr_email_on_screen
    span_ms: [731000, 742500]
    bbox_source: screen_share
    surrogate: "[EMAIL_03]"
    applied_to: [screen_share:mask, transcript:replace_if_spoken]
methods:
  face: {tool: generative_replacement, version: "x.y", scope: face_only}
  voice: {tool: speaker_anonymization, attacker_tested: semi_informed_asv}
  metadata: {policy: allowlist, kept: [duration, codec, relative_time]}
audit:
  sample_size_records: 50
  checks: [manual_review, asv_attack, ocr_rescan, metadata_diff]
  residual_findings: 1
  action: rerun_ocr_mask_on_screen_share

A manifest like this is what a privacy reviewer and the data team should both be able to read. It also connects to the multimodal dataset specification template, where buyers state the redaction policy they expect before sourcing starts.

Stripping and curating metadata without breaking alignment

Metadata should be curated with an allowlist, not stripped blindly, because a training pipeline still needs duration, codec, frame rate and relative time. Blind stripping (exiftool -all=) can remove the creation timestamps used to align streams. Re-encoding can drop or rewrite edit lists in MP4 files and shift audio relative to video.

A practical policy keeps technical fields, converts absolute timestamps to offsets from record start, drops GPS, device serials, owner and author fields, software paths and free-text comments, and hashes any session or room ID that must survive for linkage under the license. Then diff metadata before and after with the same tool across every layer. Exif alone can carry timestamps, geolocation and personal names [3], and XMP and QuickTime atoms can hold copies of the same fields.

Auditing residual re-identification risk across modalities

Auditing means attacking the processed sample the way a motivated recipient would, one modality at a time and then jointly. NIST SP 800-188 cautions that traditional de-identification has inherent limitations and frames de-identification as a governed process with documented decisions and risk assessment, not a one-time transformation [9]. For multimodal records, a useful audit pass includes:

  • Face: run a face detector and a recognition model on the processed video; count any detections the masking missed.
  • Voice: run speaker verification against enrollment audio of known participants, if available under consent.
  • Text: OCR every frame (or a dense sampling) and run NER over the OCR output and the final transcript.
  • Metadata: dump all layers and compare against the allowlist.
  • Joint: have a reviewer who knows the source context try to identify participants from combined cues (role, accent, office layout, project names).

Record what failed and what was re-run. The concept is covered in the SourceX glossary entry on re-identification, and the full company-data workflow sits in the de-identification playbook. If the processed records will also serve as a held-out test set, read de-identifying evaluation data without breaking the test before choosing surrogates.

Faces and voices are regulated as biometric identifiers in several US states, so de-identification decisions carry legal weight beyond general privacy law. Illinois BIPA defines biometric identifiers to include voiceprints and scans of face geometry while excluding photographs [5]. Texas Business and Commerce Code Section 503.001 covers voiceprints and records of face geometry and requires notice and consent before capture for a commercial purpose [6]. Plaintiffs have filed biometric privacy suits over face and voice data used in AI development, which makes face and voice handling a live litigation question as of October 2026 [7].

Health contexts add HIPAA. The Safe Harbor method's 18 identifiers include full-face photographs and comparable images and biometric identifiers such as voice prints; Expert Determination is the alternative when Safe Harbor would remove too much [8]. For state-by-state detail see biometric data rules for AI training under BIPA and Texas CUBI, and for EU recipients see anonymised versus pseudonymised training data under GDPR. Consistent surrogate maps are pseudonymization, not anonymization, if the supplier keeps the key.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

What to ask a supplier before licensing de-identified multimodal records

Buyers should ask for evidence of method and coverage per modality, not a statement that data is "anonymized." Useful requests:

  • The identity inventory: which streams, sidecars and metadata layers exist per record.
  • The method and version per modality, and whether the face method is face-only or full-body.
  • The attacker model used to test voice anonymization.
  • The surrogate policy: per record or per project, and who holds the key.
  • Audit sample size, checks run and residual findings, with what was re-processed.
  • Which derived artifacts (thumbnails, embeddings, ASR lattices) are excluded or regenerated after redaction.
  • Consent and notice records for people captured on camera or microphone, including bystanders.

Rights questions for records assembled from several parties are covered in licensing multimodal records from several rightsholders, and meeting-specific issues in multimodal meeting recordings for AI. When you are ready to describe the records you need, SourceX takes buyer requests for operational datasets.

Sourcing de-identified multimodal records through SourceX

SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request, and every release is approved by the supplying company. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe the modalities and redaction policy you need at SourceX for buyers.

Sources

  1. Hukkelås and Lindseth, CVPR 2023 Workshops (CVF Open Access), "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
  2. arXiv, "Privacy Blur: Quantifying Privacy and Utility for Image Data Release" (2025). https://arxiv.org/pdf/2512.16086
  3. arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
  4. arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677
  5. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  6. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  7. Perkins Coie, "New Biometrics Lawsuits Signal Potential Legal Risks for AI". https://perkinscoie.com/insights/update/new-biometrics-lawsuits-signal-potential-legal-risks-ai
  8. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  9. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  10. Grauman et al., arXiv, "Ego4D: Around the World in 3,000 Hours of Egocentric Video" (2021). https://arxiv.org/abs/2110.07058v1

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data