Privacy, de-identification and sensitive data
Licensing image and video data that contains faces: consent, releases or anonymization
Quick answer
Decide by what the model must learn. If it needs facial detail (identity, expression, gaze, age cues), license only footage backed by per-person written releases that satisfy biometric statutes such as Illinois BIPA and Texas Section 503.001. If it needs presence, pose, hands or actions, license footage anonymized at the source, preferably with realistic face replacement rather than heavy blur, and treat bystanders and employees as a separate rights question. Most workplace and retail datasets belong in the second lane.
By SourceX Editorial · Updated
The deciding question: does the model need facial detail?
The single most useful scoping step is to write down which labels depend on pixels inside the face region. Person detection, pose estimation, hand-object interaction, action recognition and robot perception of human co-workers mostly do not; face recognition, expression and attention models do. That one answer determines whether you are buying biometric-sensitive data with a consent chain or ordinary operational footage with a de-identification record.
The evidence supports keeping facial detail out when you can. Hukkelas and Lindseth trained detectors on anonymized versions of standard datasets and evaluated on originals: traditional blurring and masking cost noticeable accuracy on some tasks, while realistic face replacement cost little for many detection tasks [5]. The magnitude depends on what is anonymized and how, which is why our companion page on whether face and plate blurring hurts computer vision training goes deeper on the numbers.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Model objective | Needs facial pixels? | Recommended lane | Anonymization method to request |
|---|---|---|---|
| Person and PPE detection on a warehouse floor | No | Anonymize at source | Realistic face replacement or tight-box blur |
| Pose and action recognition (picking, packing) | No, except head orientation | Anonymize at source | Face replacement that preserves head pose keypoints |
| Robot perception of nearby workers | No | Anonymize at source | Face replacement plus badge and screen masking |
| Driver or operator attention and fatigue | Yes (eyes, gaze) | Consented collection | None in training; releases plus biometric consent |
| Customer expression or sentiment | Yes | Consented collection, often commissioned | None in training; strict access controls |
| Face recognition or verification | Yes, identity | Consented collection only | None; identity is the label |
When you need faces: releases plus biometric consent
A photo model release alone is not enough when faces are processed as biometrics. A standard model release grants rights to use a likeness in an image; biometric laws regulate capturing face geometry, which is what face embeddings, landmark extraction and many training pipelines produce. Treat the two documents as layers, and see model and property releases for AI training images for the release layer itself.
Illinois BIPA is the strictest reference point. It defines biometric identifiers to include a scan of face geometry, expressly excludes photographs themselves, and requires a private entity to inform the person in writing, state the purpose and retention period, and obtain a written release before collecting, plus publish a retention and destruction schedule [1]. As of October 2026, the 2024 amendment in SB 2979 makes repeated collection of the same identifier from the same person by the same method a single violation, which lowers but does not remove exposure [2]. Face datasets have drawn BIPA suits before; IBM's Diversity in Faces dataset is a frequently cited example [4].
Texas Section 503.001 covers a record of face geometry and bars capture for a commercial purpose unless the person is informed before capture and consents [3]. Washington's My Health My Data Act counts biometric data within consumer health data, which adds a second consent regime for some Washington-linked footage [8]. Ask counsel to map the states where subjects were recorded, not where the supplier is incorporated.
For commissioned or consented collection, ask for:
- A release per identifiable person, signed before capture, naming AI model training and evaluation as purposes.
- A separate biometric notice and consent that states purpose, retention period and destruction trigger, keyed to BIPA and Texas requirements.
- A participant ID that links every frame track to its release, so a withdrawal can be executed against specific files.
- The withdrawal process and what happens to copies already delivered.
- Minor and guardian handling, or a statement that no minors appear.
Our page on consent language for commissioned AI data collection covers participant wording in detail.
When you do not need faces: anonymize before the files leave the supplier
Anonymizing at the source is the lower-risk lane because the buyer never holds the identifiable original. If faces are replaced or masked in the supplier's environment and the originals stay there, your team is not the party that collected or stored face geometry. Ask for anonymization to happen before transfer, not as a step your pipeline runs after ingestion.
Specify the method, not just "blurred." Weak blur can leave faces recognizable at high resolution, and heavy blur degrades the context models use; pixelation and black boxes destroy head pose cues; realistic synthetic replacement (DeepPrivacy2-style generators) preserves pose and silhouette while removing identity [5]. Also request the detector recall on a held-out, hand-labeled sample, because every missed face in a long video is an identifiable frame. The image-specific mechanics are covered in how to de-identify images and inspection photos.
Faces are not the only identifiers in frame. Name badges, uniform embroidery, monitors showing names or account data, whiteboards, license plates and reflections all leak identity, and the audio track can carry names and voiceprints. File metadata is a separate leak: an audit of a large ML image dataset found Exif tags with timestamps, geolocation and individuals' names, and noted the download tooling stored that metadata too [9]. Require Exif and XMP stripping (or an allowlist of retained fields) and container-level metadata removal for MP4 and MOV files. For the full video checklist see anonymizing video datasets for AI.
Workplace footage: employees and bystanders are separate cases
Workplace video usually needs two different rights analyses, one for employees who were filmed and one for everyone else. Employees may have received a recording notice tied to safety or security, but that notice rarely mentions licensing footage to a third party for AI training. Customers, visitors, delivery drivers and contractors received, at most, a door sign.
Practical handling that holds up in diligence:
- Treat employee footage as consented only if there is a specific consent covering third-party AI training; otherwise anonymize faces and audio.
- Treat bystanders as never consented; anonymize every face not covered by a release, including partial and background faces.
- Exclude restrooms, break rooms, locker areas and medical spaces outright.
- Drop or transcribe-and-redact audio unless you have a clear basis for it; see recording employees on video for AI datasets.
- For storefront footage, read retail store video privacy constraints before scoping camera angles.
Why the model and its outputs are part of the risk
Face data risk does not end at the dataset, because regulators and researchers have both tied it to trained models. In the Everalbum matter the FTC's final order required the company to delete not only photos and videos from deactivated accounts but also models and algorithms developed using users' photos and videos [6]. A buyer that trains on improperly sourced face data can therefore face a remedy that reaches the model weights.
Generative models add an output-side exposure. Carlini and colleagues extracted more than a thousand training examples from diffusion models, including photos of individual people [7]. If you train generative or multimodal models on face-bearing media, keep provenance per file so you can show which images entered which run, and see training-data extraction and memorization risk.
A request template that keeps both lanes honest
A good request names the lane, the method and the evidence you will accept before any files move. Suppliers respond better to a precise specification than to "privacy-safe footage," and counsel can review it before you see samples.
Illustrative example: invented to show structure; it does not describe an available dataset.
request:
modality: video
environment: warehouse pick-and-pack stations, fixed overhead cameras
model_objective: action recognition (pick, scan, pack, place)
facial_detail_required: false
lane: anonymize_at_source
anonymization:
faces: realistic_replacement # not gaussian_blur
badges_and_screens: masked
audio: removed
metadata: exif_xmp_and_container_tags_stripped
detector_recall_report: required # on hand-labeled held-out sample
exclusions: [restrooms, break_rooms, medical_areas]
people_in_frame:
employees: anonymized unless third-party AI training consent exists
bystanders: always anonymized
evidence_requested:
- anonymization method and tool version
- sample QA results and residual-risk statement
- recording notice text shown to staff
- states where recording took place
For a consented-face request, flip facial_detail_required to true, set the lane to consented_collection, and replace the anonymization block with release and biometric-consent artifacts keyed by participant ID. The broader document list is in the de-identification evidence package checklist.
How SourceX approaches face-bearing footage
SourceX sources operational datasets, including new recordings of hands-on work, from US companies on request, and every release is approved by the supplying company. It does not source scraped web content or generic CCTV or photos, so requests should describe a specific operational setting and objective. Each dataset is rights-reviewed for ownership and consents, personal details are removed or replaced before delivery with the method recorded and a sample checked, and no method is perfect. To scope a request in either lane, start at the SourceX buyer intake; for recordings specifically, see licensing video recordings for AI training.
Source face-bearing image or video data with the right paperwork
SourceX looks for US businesses that hold the operational footage you describe, and nothing is contracted until a supplier agrees. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery. Describe your objective and lane at sourcex.si/buyers.
Frequently asked questions
Can you train on images that contain faces without consent?
It depends on whether faces are processed as biometrics, where subjects were recorded, and how the data was collected. If the objective does not need facial detail, the safer route is anonymization before transfer so you never hold identifiable faces. If it does, obtain releases and biometric consent; the privacy cluster hub covers the related legal definitions.
Is a standard model release enough for AI training?
Usually not on its own. A release covers use of a likeness; statutes such as BIPA and Texas Section 503.001 separately regulate capture of face geometry and require specific notice and consent [1][3].
Does blurring faces make footage non-personal data?
Not automatically. Missed detections, badges, screens, audio and metadata can still identify people, and under GDPR-style regimes context and auxiliary data matter; compare de-identified vs. anonymized legal definitions. This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sources
- Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
- Illinois General Assembly, "SB 2979 (103rd General Assembly), BIPA amendment, engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
- Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier". https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
- Bloomberg Law, "IBM Trims Privacy Lawsuit Over Its Diversity in Faces Dataset". https://news.bloomberglaw.com/ip-law/ibm-trims-privacy-lawsuit-over-its-diversity-in-faces-dataset
- Hukkelas and Lindseth, CVPR 2023 Workshops, "Does Image Anonymization Impact Computer Vision Training?" (2023). https://openaccess.thecvf.com/content/CVPR2023W/WAD/papers/Hukkelas_Does_Image_Anonymization_Impact_Computer_Vision_Training_CVPRW_2023_paper.pdf
- Federal Trade Commission, "FTC Finalizes Settlement with Photo App Developer Related to Misuse of Facial Recognition Technology" (2021). https://www.ftc.gov/news-events/news/press-releases/2021/05/ftc-finalizes-settlement-photo-app-developer-related-misuse-facial-recognition-technology
- Carlini et al., USENIX Security 2023, "Extracting Training Data from Diffusion Models" (2023). https://www.usenix.org/conference/usenixsecurity23/presentation/carlini
- Washington State Legislature, "Chapter 19.373 RCW, Washington My Health My Data Act". https://app.leg.wa.gov/RCW/default.aspx?cite=19.373&full=true
- arXiv, "A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset" (2025). https://arxiv.org/pdf/2506.17185
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.