Skip to content

Industry-specific operational data

Sourcing learner data for education AI: K-12, higher ed and corporate training compared

Quick answer

The most licensable education data for AI training comes from corporate learning and development (L&D) programs, certification bodies, and content owners such as course and item-bank publishers. K-12 and higher-ed student records are the hardest to license, because FERPA, state student-privacy laws and district contracts restrict or bar vendors from reusing them. Consumer learning apps and test-prep sit in between. Their rights depend on consumer privacy terms and, for users under 13, COPPA. Pick sources by who collected the data, under which contract, and whether secondary use was ever approved.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Which education segments can realistically license learner data?

Licensing odds fall as you move from adult, employer-held records to minors' records held under school contracts. The table below compares the five main source types on the factors that usually decide a deal. Use it to rank sourcing effort before you send a single request. The hub for this cluster, industry-specific operational data for AI, covers the same logic for other verticals.

Illustrative example: invented to show structure; it does not describe an available dataset.

Source segmentTypical dataMain legal regimeWho must approve reuseRealistic licensing odds
K-12 districts and K-12 edtech vendorsLMS activity, assignment submissions, assessment scores, tutoring chatsFERPA, state student-privacy laws, COPPA for under-13 usersDistrict (data owner), often via a data privacy agreement; vendor cannot decide aloneLow for records; moderate for de-identified aggregates if the contract allows
Higher-ed institutions and their vendorsCanvas or Moodle logs, graded essays, discussion posts, advising notesFERPA, institutional IRB and data-governance policyRegistrar or data steward; IRB if research-derivedLow to moderate; better for research partnerships than commercial licenses
Consumer learning apps and test-prepPractice-question responses, timing, hint use, essays, tutoring sessionsConsumer privacy law (for example CCPA), FTC Act, COPPA if under 13The company, bounded by its own privacy policy and terms at collectionModerate if terms disclosed model training or the data is properly de-identified
Corporate L&D and training providersxAPI/SCORM completions, scenario simulations, coaching transcripts, skills assessmentsEmployment and contract terms, state privacy law; generally outside FERPAThe employer client and, for vendors, each client contractModerate to high, client by client
Certification and licensing bodiesItem-level responses, scored constructed responses, rater annotationsCandidate agreements, psychometric security policiesThe certification bodyModerate for retired items and scored responses; low for live item banks

Why are K-12 and higher-ed student records the hardest to license?

Student records are hard to license because the institution, not the edtech vendor, controls them, and modern contracts usually forbid secondary use. FERPA lets a school disclose records to a vendor acting as a "school official" under direct control, and it allows release of de-identified records without consent once the school has made a reasonable determination that students cannot be identified [1]. Neither path automatically lets the vendor turn around and license the data to a third-party model builder.

Contract practice is stricter than the statute. Practitioner guidance recommends that schools bar AI vendors from training models on student data, de-identified or not [3], and district data privacy agreements frequently mirror that language. K-12 products add edge cases, such as AI features that process student work under the school-official exception but then retain prompts for model improvement [4]. Many states also have student-privacy statutes, such as California's SOPIPA, that restrict operator use of student data beyond the school purpose; check each state where the data originated.

De-identification is a weaker shield than it looks. A 2026 law-journal analysis argues that FERPA treats de-identified data permissively while setting a minimal de-identification bar, and that re-identification has been repeatedly demonstrated [2]. NIST's survey of the field documents re-identification of datasets that were believed to be de-identified [10]. For learner data, quasi-identifiers such as school, grade, course section, timestamps and free-text essays make linkage easier than in most tabular data.

How do edtech platform terms change what a vendor can license?

An edtech vendor can license learner data only to the extent the terms under which that data entered its platform allow. Platform terms that reserve rights to "de-identified aggregate data" for product improvement or model training are a known weak point, because the institution may never have agreed to that use in its own contract [5]. The vendor's click-through terms and the district's signed agreement can conflict, and many signed agreements include precedence clauses that override click-through terms.

Consumer-facing promises matter too. The FTC has warned that companies may face liability under FTC-enforced laws if they break commitments not to use customer data for undisclosed purposes, including model training [8]. If a test-prep app's privacy policy said "we never share your answers," a later license to you inherits that problem. Ask for the policy version in force when each record was collected, not the current one.

For users under 13, the amended COPPA Rule was published on April 22, 2025, took effect June 23, 2025, and had a general compliance date of April 22, 2026 [6]. Law-firm summaries describe a new requirement for separate verifiable parental consent before disclosing children's personal information to third parties [7]. As of October 2026, treat any identifiable under-13 data offered for training as a red flag unless the supplier can show that consent trail.

What makes corporate training data easier to source?

Corporate L&D data is easier to source because adult employees' training records generally fall outside FERPA, and the employer or training vendor usually holds the rights under contract. The constraints shift to employment terms, state privacy law, and each enterprise client's agreement with the training provider. For a broader view of this vertical, see the corporate training buyer overview.

The data is also more structured than many buyers expect. Modern learning stacks emit xAPI statements to a Learning Record Store, SCORM 2004 or cmi5 runtime data from the LMS, and 1EdTech Caliper events from some platforms. Useful records include scenario-simulation branches, sales-coaching roleplay transcripts, compliance-quiz item responses, and manager assessments of skills.

Watch for three failure modes. Multi-client training vendors often cannot separate one client's learners from another's, which mirrors the queue problem covered in sourcing AI training data through BPOs. Coaching transcripts may contain customer names, deal values or health disclosures from the employee. Completion data alone (pass/fail, minutes spent) is rarely worth licensing for tutoring or grading models.

Is course content easier to license than learner records?

Yes, course content, item banks and scored response sets are usually easier to license than learner records because a single publisher or certification body typically owns them outright. Content licensing turns on copyright and contributor agreements rather than student privacy. For many tutoring and grading use cases, a corpus of rubrics, worked solutions and rater-scored responses is more useful than raw clickstream.

If your model needs instructional material rather than learner behavior, start with training materials and LMS content. For continued pre-training on curricula and textbooks, the requirements in sourcing domain corpora for continued pre-training apply. Retired assessment items are often licensable when live items are not, because certification bodies protect active forms against exposure.

What should a learner-data record look like when it arrives?

A licensable learner-data record carries a pseudonymous learner key, the content item it refers to, the learner's response, the score and rubric, and provenance fields that tie it back to the collection terms. Without provenance fields, you cannot prove later which records fall under which permission. The schema below shows the minimum structure worth asking for.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "rec_000184",
  "segment": "corporate_ld",
  "learner_key": "lk_7f3a91",
  "cohort": {"role_family": "inside_sales", "tenure_band": "1-3y"},
  "activity": {"standard": "xAPI", "verb": "answered", "object_id": "item:objection-handling-q12"},
  "response_text": "[CUSTOMER_NAME] said the price was too high, so I...",
  "result": {"score_scaled": 0.72, "rubric_id": "rub_objection_v3", "rater_type": "human"},
  "timestamp_bucket": "2025-Q3",
  "provenance": {
    "collected_under": "client_msa_v4 + platform_terms_2025-01",
    "secondary_use_approved_by": "employer_client",
    "deid_method": "named-entity replacement + date bucketing",
    "deid_sample_checked": true
  }
}

Note the replacement token in response_text, the date bucketing, and the absence of free-form employer names. For delivery formats, manifests and schema documentation, see dataset delivery formats and schemas.

Learner-data supplier checklist

Run every candidate source through the same questions before you negotiate price. They surface common deal-breakers early. The broader workflow sits in AI training data procurement.

Illustrative example: invented to show structure; it does not describe an available dataset.

  1. Collector. Who collected the data: the school, the vendor, the employer, or the learner directly through a consumer app?
  2. Terms at collection. Which contract, data privacy agreement or privacy policy version governed each record, and does it mention model training or secondary use?
  3. Approver. Did the institution or employer client approve the secondary use in writing, or is the vendor relying only on its own terms [5]?
  4. Age. Can the supplier show that no under-13 personal data is included, or show verifiable parental consent for disclosure [7]?
  5. De-identification method. Which identifiers were removed or replaced, how free text was handled, and was a sample checked? For CCPA-covered data, can the holder meet the statute's conditions for deidentified information, including public commitments and contractual bans on re-identification [9]?
  6. Re-identification controls. Does the license prohibit re-identification and linkage with other datasets, and who audits that?
  7. Segment separation. Can records be filtered by client, institution and state so you can exclude any source whose terms are unclear?
  8. Quality and contamination. Are assessment items in your eval set also in the training data? See training data quality assessment.

Choosing sources by model use case

Match the source to the model's job rather than collecting every learner signal available. Tutoring models gain most from multi-turn tutoring or coaching transcripts and worked solutions; grading models need scored constructed responses with rubrics and, ideally, double-rated items for agreement statistics. Learning-analytics models need longitudinal event streams with stable pseudonymous keys, which are also the hardest to de-identify safely.

For evaluation sets, small, well-documented samples from certification bodies or corporate assessments often beat large clickstream dumps. For post-training on real learner questions, the guidance on real-world prompt sets for post-training applies, with the added need to strip minors' details. Institutions weighing their side of these deals can read compliance checks for higher education institutions and data licensing rules for education and tutoring companies. Suppliers' FERPA obligations are covered in the FERPA licensing guide.

If your requirement centers on adult learner or training-program data held by US companies, you can describe the dataset to SourceX; data is sourced on request rather than held in stock, and a request does not guarantee a match.

Requesting education and training data for AI

SourceX looks for US businesses that hold the data you describe, rights-reviews each dataset for ownership and consents, and delivers it under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Nothing is contracted until the supplying company agrees, so start by describing the learner or training data you need.

Sources

  1. U.S. Government Publishing Office / U.S. Department of Education, "34 CFR 99.31 - Under what conditions is prior consent not required to disclose information? (CFR 2018 edition)" (2018). https://www.govinfo.gov/content/pkg/CFR-2018-title34-vol1/pdf/CFR-2018-title34-vol1-sec99-31.pdf
  2. Michigan Journal of Environmental & Administrative Law (online), "Clarke, Spring 2026 (FERPA, de-identified data and AI)" (2026). https://www.mjeal-online.org/2026/04/12/clarke-spring-2026/
  3. Promise Legal, "AI in EdTech: FERPA, COPPA, and State Student Privacy Laws When Your App Adds AI Features". https://blog.promise.legal/edtech-ai-compliance-ferpa-coppa/
  4. Promise Legal, "FERPA Edge Cases for AI Features in K-12 Products". https://blog.promise.legal/ferpa-edge-cases-for-ai-features-in-k-12-products/
  5. Everything PR News, "The FERPA problem with AI vendors in edtech". https://everything-pr.com/ferpa-problem-ai-vendors-edtech
  6. Federal Trade Commission, Federal Register, "Children's Online Privacy Protection Rule (Final Rule amendments), 90 FR 16918" (2025). https://www.federalregister.gov/documents/2025/04/22/2025-05904/childrens-online-privacy-protection-rule
  7. Loeb & Loeb, "Children's Online Privacy in 2025: The Amended COPPA Rule" (2025). https://www.loeb.com/en/insights/publications/2025/05/childrens-online-privacy-in-2025-the-amended-coppa-rule
  8. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  9. California Legislature, "California Civil Code section 1798.140 (CCPA definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV&sectionNum=1798.140
  10. National Institute of Standards and Technology, "De-Identification of Personal Information (NISTIR 8053)" (2015). https://nvlpubs.nist.gov/nistpubs/ir/2015/NIST.IR.8053.pdf
  11. California Legislature, "California Business and Professions Code section 22584 (SOPIPA)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=BPC&sectionNum=22584

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data