Privacy, de-identification and sensitive data
The DOJ bulk sensitive data rule and licensed AI training data: what buyers outside the US must check
Quick answer
Licensing US operational data to a non-US AI developer is usually "data brokerage" under 28 CFR Part 202, the Justice Department rule implementing Executive Order 14117 [1][2]. If the buyer is not a covered person, the deal is allowed, but the US licensor must contractually bar onward brokerage to countries of concern and covered persons [1]. Bulk thresholds apply to de-identified and pseudonymized data too, so stripping names does not take a training corpus out of scope [2].
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Why a training-data license falls inside 28 CFR Part 202
A dataset license is inside the rule because Part 202 defines data brokerage to include licensing access to data that the recipient did not collect directly from the individuals [1]. That describes nearly every third-party training-data deal: support tickets, claims files, call recordings or transaction logs collected by a US company and then licensed to a model developer.
DOJ was explicit about the AI angle. Its threat model is that countries of concern use bulk US sensitive personal data to build and enhance AI capabilities, and combine unrelated datasets to identify individuals [3][2]. The rule was issued in December 2024, published in the Federal Register in January 2025, and its prohibitions took effect April 8, 2025 [1]. The affirmative due diligence, audit and reporting provisions for US persons applied from October 6, 2025 [1].
The rule's affirmative duties fall on US persons, although its evasion provision reaches any person who causes or conspires in a violation [1]. In practice it reaches you through the contract your US counterparty must sign and through the screening it must run before release [1][4]. As of October 2026 the program has been in force for more than a year, and compliance commentary describes onward-transfer terms becoming routine in data deals [5].
Countries of concern and covered persons: the first screen
The first question is whether your organization, or anyone who will touch the data, is a covered person. As of October 2026, the countries of concern listed in the rule are China (including Hong Kong and Macau), Cuba, Iran, North Korea, Russia and Venezuela [1]; check the current eCFR text, because DOJ can amend the list.
Covered persons broadly include entities that are 50% or more owned by a country of concern or by covered persons, entities organized under the laws of or with a principal place of business in a country of concern, individuals primarily resident in a country of concern, and individuals who are employees or contractors of a country of concern or a covered entity [1]. DOJ can also designate specific persons by name [1]. A data-brokerage transaction with a covered person is prohibited outright, not merely restricted [1].
Non-US developers trip this screen in ways they do not expect. Common failure modes:
- A parent or major investor whose ownership chain aggregates to 50% or more from covered persons.
- An annotation or evaluation vendor whose staff are resident in a country of concern and will see raw records.
- A subsidiary or research lab in Hong Kong that shares the same training cluster or object storage bucket.
- Evasion structures, such as routing the license through a third-country affiliate, which Part 202 separately prohibits [1].
Bulk thresholds and why de-identification does not remove you from scope
Part 202 applies when a dataset meets the "bulk" threshold for a category of sensitive personal data, counted over the preceding 12 months and aggregated across transactions between the same parties [1]. The thresholds below follow law-firm summaries of the final rule; confirm current values against the eCFR text of Part 202 before relying on them [1][2].
| Sensitive data category | Bulk threshold (US persons or devices) | Typical AI training source |
|---|---|---|
| Human genomic data | More than 100 US persons | Lab and research records |
| Other human 'omic data | More than 1,000 US persons | Clinical and research records |
| Biometric identifiers (face, voice, gait and similar) | More than 1,000 US persons | Call-center audio, video of hands-on work |
| Precise geolocation data | More than 1,000 US devices | Field-service, fleet and mobile app logs |
| Personal health data | More than 10,000 US persons | Claims, EHR extracts, wellness apps |
| Personal financial data | More than 10,000 US persons | Card transactions, credit files, collections notes |
| Covered personal identifiers | More than 100,000 US persons | CRM, support and billing systems |
| Combined categories | Lowest applicable threshold | Joined multi-source corpora |
| US government-related data | No threshold | Data on current or recent government personnel or sensitive locations |
The critical point for training data is that bulk US sensitive personal data counts "regardless of whether the data is anonymized, pseudonymized, de-identified, or encrypted" [1][2]. A HIPAA Safe Harbor extract, a tokenized card-transaction table or a voice corpus with names bleeped can all still be covered. De-identification still matters for HIPAA, GLBA and state law, which our guides to financial records and GLBA and biometric data under BIPA and CUBI cover, but it is not a Part 202 exit.
Covered personal identifiers deserve attention because operational data is full of them. The category lists items such as government ID numbers, financial account numbers, device and advertising identifiers, account login data, network identifiers and call-detail data, and generally applies when they appear in combination with another listed identifier or linking data [1]. Demographic or contact fields linked only to each other do not count on their own [1]. A support-ticket export with customer IDs, IP addresses and device fingerprints can therefore cross 100,000 people quickly.
Some content falls outside "sensitive personal data" entirely, including certain lawfully available government public records, personal communications that do not transfer anything of value, and information or informational materials such as published media [1]. Do not assume a corpus of emails or chats is exempt: the carve-outs are narrow and apply per data element, and business records attached to messages may still be identifiers or financial data. Ask counsel to map each field.
The onward-transfer clause a non-US licensee will be asked to sign
If you are a foreign person but not a covered person, the license is permitted only if the US party contractually requires you not to engage in a later data-brokerage transaction of the same data with a country of concern or covered person [1]. The rule also expects the US party to report known or suspected violations of that clause to DOJ, so a breach on your side becomes a regulatory event for your supplier [1].
Treat the clause as a set of operational controls, not boilerplate. A license that bars onward brokerage but lets you sublicense the corpus to an unvetted reseller, or host it in a jurisdiction where a covered-person contractor has admin rights, creates exactly the exposure the rule targets [4].
Illustrative example: invented to show structure; it does not describe an available dataset.
Part 202 clause and control checklist for a foreign-person licensee
| Item | What to check | Evidence to keep |
|---|---|---|
| Onward brokerage bar | Clause prohibits sale, license or similar transfer of the licensed data to any country of concern or covered person | Executed license, clause reference |
| Sublicensing chain | Any permitted sublicense flows down the same bar, in writing | Flow-down template, sublicensee list |
| Covered-person attestation | Your entity and affiliates confirm they are not covered persons; ownership chart attached | Signed attestation, cap table summary |
| Access control | Raw records restricted to named roles; no access by staff primarily resident in countries of concern | IAM group export, HR residency check |
| Vendors and annotators | Labeling, eval and red-team vendors screened against the covered-person definition | Vendor screening log |
| Hosting location | Storage region and cloud tenant documented; no replicas in countries of concern | Bucket policy, region list |
| Derived artifacts | Agreement states whether embeddings, filtered subsets and synthetic data derived from the records are treated as the licensed data | License definitions section |
| Incident notice | You must notify the licensor promptly of a suspected onward transfer so it can meet its reporting duty | Notice clause, contact roster |
The derived-artifact row is where AI deals most often go wrong. Part 202 is written around data, not models, so whether a filtered subset, a sample shipped for evaluation or an embedding index is "the same data" is a contract question you should settle before delivery rather than after an audit [2].
Due diligence your US counterparty will run, and what to have ready
Expect your supplier to run know-your-customer diligence on you, because US persons must avoid prohibited transactions, enforce the onward-transfer clause and, for restricted transactions, maintain a data compliance program with due diligence, recordkeeping and audits [1]. Having the answers ready shortens procurement and avoids the deal stalling at legal review.
A practical diligence pack for a non-US AI buyer includes:
- Legal entity name, jurisdiction of organization and principal place of business for the contracting entity and its parent.
- A simplified ownership chart showing any holder at or above 10%, plus a statement on aggregate ownership by persons connected to countries of concern.
- Where the data will be stored and processed, by cloud provider and region.
- Which teams and vendors will access raw records, with residency confirmation.
- The intended use (pretraining, fine-tuning, evaluation or retrieval) and whether any derived data will be shared with third parties.
- A named compliance contact who will receive and answer onward-transfer notices.
Your own privacy work runs in parallel. If the US data contains EU or UK personal data, Part 202 does not replace transfer mechanics, which our guide to EU personal data in US datasets covers. For de-identification proof under HIPAA and state law, use the de-identification evidence package checklist.
Restricted transactions versus data brokerage: which one applies to you
Data brokerage is the category that governs a training-data license; vendor, employment and investment agreements are "restricted" transactions that are allowed with covered persons only if CISA security requirements are met [1]. The distinction matters because a buyer sometimes plays two roles.
If your company also provides services to the US supplier, such as hosting, labeling or model development under a vendor agreement, that relationship can itself be a restricted transaction if you are a covered person [1]. If you are not a covered person, the vendor-agreement rules generally do not bite, but the data-brokerage onward-transfer clause still does. Map each agreement separately rather than assuming one license covers everything.
Certain exemptions exist, including some for clinical investigations, regulatory submissions, corporate group transactions and transactions required by law [1]. They are fact-specific and rarely fit a commercial training license, so do not build a deal on one without written advice.
How SourceX fits a cross-border training-data request
SourceX sources operational datasets from US companies on request and manages the commercial process, including the license, for AI teams wherever they are based. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, and diligence materials are prepared per dataset. As this page explains, that de-identification does not by itself take a dataset outside Part 202, so the license terms and your own screening still matter. You can describe the data you need on the SourceX buyers page, and see the general guide for buyers outside the US.
Every release is approved by the supplying company, delivery runs through private, access-controlled workflows only after an executed agreement, and nothing is contracted until a supplier agrees. A request does not guarantee a match. For the wider privacy picture, start from the privacy and de-identification hub or the AI data guides.
License US training data with the DOJ bulk data rule in view
SourceX finds US businesses that hold the operational data you describe, assesses the data and licensing permissions, and agrees allowed uses in a license that defines records, uses, term and delivery. Bring your entity details and intended use so the Part 202 questions are answered early. Start a buyer request.
Sources
- White & Case, "DOJ issues final rule prohibiting and restricting transfers of bulk sensitive personal data" (2025). https://www.whitecase.com/insight-alert/bulk-data-final-rule
- White & Case, "DOJ issues final rule prohibiting and restricting transfers of bulk sensitive personal data" (2025, alternate publication). https://www.whitecase.com/insight-alert/doj-issues-final-rule-prohibiting-and-restricting-transfers-bulk-sensitive-personal
- DLA Piper, "DOJ announces proposed rule to mitigate data security risks related to AI" (2024). https://www.dlapiper.com/insights/publications/2024/02/doj-announces-proposed-rule-to-mitigate-data-security-risks-related-to-ai
- Neudata, "Sensitive data, serious stakes: what the DOJ's new data rule means for your data supply chain". https://www.neudata.co/sentry-intelligence/sensitive-data-serious-stakes-what-the-doj-s-new-data-rule-means-for-your-data-supply-chain
- Model Diplomat, "DOJ's Bulk Data Rule Turns One: Compliance". https://modeldiplomat.com/story/dojs-bulk-data-rule-turns-one-compliance
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.