Privacy, de-identification and sensitive data
Student records and FERPA: when de-identified education data can be licensed for AI
Quick answer
Under FERPA, education records stripped of personally identifiable information can be released without consent once the holder reasonably determines no student is identifiable [1]. That makes de-identified student data legally licensable in principle, but the gate is rarely FERPA itself. It is the chain behind the data: the school-vendor contract, state student privacy statutes, COPPA, and re-identification risk in rich LMS and assessment logs. A buyer should treat each link as a separate approval, not assume "de-identified" settles it.
By SourceX Editorial · Updated
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
What FERPA actually permits for de-identified education records
FERPA permits disclosure of de-identified records without parental or eligible-student consent, but only after removal of all personally identifiable information and a "reasonable determination" that a student's identity is not personally identifiable, considering single and multiple releases and other reasonably available information [1]. The test is judgment-based. There is no HIPAA-style list of 18 identifiers and no Expert Determination path to point at, and one 2026 law-journal analysis describes FERPA's treatment of de-identified data as permissive under a minimal standard [2].
Two details matter for AI buyers. First, the "personally identifiable information" definition in 34 CFR 99.3 reaches indirect identifiers and information that would let a reasonable person in the school community identify a student with reasonable certainty, so a small-school cohort with a rare disability code can stay identifiable after names are gone. Second, 99.31(b)(2) separately allows de-identified student-level data with an attached record code for education research, with conditions on how the code is generated and used [1]; that research path is not a general commercial training license. The current eCFR text should be checked before any clause relies on these provisions, since the version cited here is the 2018 annual edition.
For the definitional differences between de-identified, pseudonymized and aggregated data across regimes, see de-identified vs anonymized legal definitions. The supplier-side view of the statute sits on the FERPA licensing page.
Why the school-vendor contract usually decides the question
The contract between the school or district and the edtech vendor usually controls more than FERPA does, because most vendors hold student data as a "school official" under 99.31(a)(1), which ties their use to the school's purposes and its direct control [1]. A vendor that later wants to license derived data to an AI developer must show that its agreement allowed the de-identified data to leave that relationship. Practitioners report that school contracts increasingly prohibit training on student data, including de-identified data [3].
Common blocking clauses look like these, and counsel should search for each in every agreement covering the source population:
- Purpose limitation: data may be used "solely to provide the services" to the district.
- No secondary use of de-identified data: de-identified data may be used only to improve the vendor's own product, not transferred.
- Transfer conditions: de-identified data may go to third parties only with district notice and a recipient covenant not to re-identify.
- Model-training bans: explicit prohibitions on using student data, in any form, to train or fine-tune AI models.
- Return or destruction: obligations at term end that may reach derived datasets and, arguably, model artifacts.
Districts commonly use template agreements such as state or consortium data privacy agreements, and the de-identified-data article differs between versions. A vendor serving many districts may hold dozens of signed variants, so the licensable population is often the subset of districts whose signed terms permit transfer, not the whole customer base.
State student privacy laws and COPPA add separate limits
State student privacy statutes can restrict operators of K-12 services independently of FERPA, so a dataset can be FERPA-compliant and still unlicensable [3]. California's Student Online Personal Information Protection Act (SOPIPA) and the many state laws modeled on it limit operators' use of student information to school purposes, with narrow allowances for de-identified data; whether a third-party training license fits those allowances is a question for counsel in each state where students are located.
COPPA applies when the vendor collected personal information directly from children under 13 on a commercial service. The amended COPPA Rule was published April 22, 2025, became effective June 23, 2025, and carried a compliance date of April 22, 2026 [5]. As of October 2026, buyers should ask whether any child data in scope was collected under the school-authorization model, and whether disclosures to third parties were separately consented to where the amended rule requires it. Practitioner guidance also flags edge cases, such as AI features that generate new data about students, as needing their own FERPA analysis [4].
Promises matter too. The FTC has warned that companies may face liability if they break commitments not to use customer data for undisclosed purposes such as training [8]. A vendor privacy policy saying "we never use student data to train AI" is a constraint on any licensed dataset regardless of de-identification.
Re-identification risk in LMS, assessment and tutoring data
Education datasets are high-risk for re-identification because they are longitudinal, small-cohort and behaviorally distinctive. Re-identification of educational datasets has been demonstrated, and models trained on de-identified behavioral data still learn real student patterns [2]. Research on general populations shows that a handful of quasi-identifiers can make most individuals unique even in sampled data [7].
The quasi-identifiers that survive naive de-identification in this domain are predictable:
- LMS event logs: clickstream timestamps, session IDs, device or IP fields, course section IDs that map to a single teacher and a class of 20.
- Assessment records: state assessment scores with grade, school, year, and subgroup flags (IEP, 504, English learner, free or reduced-price lunch).
- Free text: essays, tutoring chat transcripts and teacher comments that name classmates, towns, teams, medical conditions or family events.
- Rare combinations: small schools, rare languages, gifted or special education placements, and transfer histories.
NIST SP 800-188 describes both traditional de-identification and formal methods such as differential privacy, and cautions about the limits of the traditional approach [6]. For text-heavy tutoring data, the newer threat model is LLM-assisted re-identification; see LLM-assisted re-identification of de-identified text. If you plan to join student data with other licensed sources, review linkage and mosaic risk first, because FERPA's reasonable determination explicitly considers other reasonably available information [1].
Chain-of-rights checklist for student-derived training data
A buyer can license student-derived data with confidence only when every link from student to model is documented. The table below is a working checklist for counsel; each row should produce a document or a recorded answer, not an assurance.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Link in the chain | Question to answer | Evidence to request | Red flag |
|---|---|---|---|
| Collection | Under which authority was the data collected (school official, consent, COPPA school authorization)? | Data map by source system (SIS, LMS, assessment platform, tutoring app) | Vendor cannot say which districts' data is in scope |
| School contract | Does each governing agreement permit transfer of de-identified data to third parties for model training? | Clause extracts per district or template version, with a list of excluded districts | Blanket "our contracts allow it" with no extracts |
| State law | Which states' students are included, and does each state's student privacy law permit this use? | State-by-state counsel memo or matrix | K-12 data from SOPIPA-style states with no analysis |
| Vendor promises | Do privacy policies or marketing statements disclaim AI training use? | Archived policy versions covering the collection period | Policy promised no training at time of collection |
| De-identification | What was removed, generalized or suppressed, and how was the reasonable determination made? | Method write-up, small-cell thresholds, free-text redaction results on a sample | Only direct identifiers removed from longitudinal logs |
| Residual risk | Was re-identification tested against realistic outside data? | Attack test summary, uniqueness metrics | No testing on essays or chat transcripts |
| License | Does the license define records, permitted uses, term, re-identification ban and delivery? | Draft license and diligence pack | Use rights broader than the upstream contracts allow |
The de-identification evidence package checklist covers the generic documents; this table adds the education-specific links that most often fail.
Contract terms a buyer should require
Buyers should mirror the upstream restrictions downstream so the license never grants more than the school contracts allow. A clean license for student-derived data typically includes a warranty-backed representation from the licensor about upstream authority, a covenant not to re-identify or link to student identities, a ban on onward transfer of raw records, and a defined response if a district later objects or a contract is found not to permit the use.
Model outputs deserve their own clause. Tutoring and assessment models can memorize essay passages or rare score patterns, so require training and evaluation controls proportionate to risk; see training-data extraction and memorization risk. Where upstream contracts forbid any transfer, a data clean room or vendor-side training arrangement may be the only workable structure.
Content that is not student records at all, such as course materials, instructional videos and teacher-authored curricula, raises copyright and licensing questions rather than FERPA ones. If that is what your model needs, see training materials and LMS content. The broader privacy picture is in the de-identified data buyer's guide.
How SourceX handles requests that touch student data
SourceX sources operational datasets from US companies on request, rather than holding inventory, so a request for education-adjacent data does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines the records, uses, term and delivery. Personal details such as names, emails, phones and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect.
Diligence materials on source, rights, preparation and allowed use are prepared per dataset, and each release is approved by the supplying company. You can describe the education or LMS data you need without naming suppliers.
Request student-derived training data with the rights chain documented
Describe the records, fields and allowed uses your tutoring or assessment model needs. SourceX looks for US businesses that hold that data, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered through private, access-controlled workflows. Start at SourceX for AI data buyers.
Sources
- U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information? (CFR 2018 edition)" (2018). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
- Michigan Journal of Environmental & Administrative Law (MJEAL Online), "Clarke - Spring 2026" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/
- Promise Legal, "AI in EdTech: FERPA, COPPA, and State Student Privacy Laws When Your App Adds AI Features". https://blog.promise.legal/edtech-ai-compliance-ferpa-coppa/
- Promise Legal, "FERPA Edge Cases for AI Features in K-12 Products". https://blog.promise.legal/ferpa-edge-cases-for-ai-features-in-k-12-products/
- Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
- National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
- Nature Communications (Rocher, Hendrickx, de Montjoye), "Estimating the success of re-identifications in incomplete datasets using generative models" (2019). https://pmc.ncbi.nlm.nih.gov/articles/PMC6650473
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.