For AI data buyers
AI training data licensing terms
Quick answer
Plain-language definitions of the words used when AI developers license business data for training: usage rights, exclusivity, permitted use and preparation. Type a word to filter.
99 terms
- Agent trajectoryThe full sequence of steps an AI agent or person takes to complete a task, including actions, tool calls and results.
- AI training dataInformation used to teach an AI model patterns, such as text, conversations, images, audio, code or action records.
- API rate limitA cap on how many requests a program can make to a software service in a given time.
- Benchmark contaminationWhen test examples leak into a model's training data, inflating its scores.
- Call transcriptA written text version of a phone or video call, often produced by speech-to-text software.
- Chain of custodyA documented record of who held, accessed or changed something, and when, from collection to delivery.
- Chain of thoughtA model's intermediate reasoning steps written out before its final answer.
- Computer-use dataRecords of people operating software, such as screen recordings or click and keystroke sequences.
- Consent managementThe tools and processes a company uses to collect, record and honor people's choices about how their data is used.
- Data annotationAdding labels or metadata to raw data so models can learn from it.
- Data anonymizationIrreversibly transforming data so individuals cannot be identified.
- Data as a productTreating a dataset as something designed, documented and maintained for its users, with clear ownership and quality standards.
- Data brokerA company that collects personal information about consumers, often from public and commercial sources, and sells it to others.
- Data catalogAn inventory of a company's data assets with descriptions, owners and metadata so people can find and understand them.
- Data clean roomA secure environment where parties can analyze combined data without seeing each other's raw records.
- Data controllerUnder GDPR, the organization that decides why and how personal data is processed.
- Data cooperativeAn organization where members pool data and govern its use collectively, often sharing the benefits.
- Data exportCopying records out of a software system into files or another system.
- Data governancePolicies and controls that manage data's ownership, quality, access, use, retention and deletion.
- Data intermediaryAn organization that connects data holders with data users and helps arrange access on agreed terms.
- Data inventoryA list of the systems, record types, volumes and date ranges a company holds.
- Data lakehouseA data platform that combines low-cost data lake storage with database-style tables and governance.
- Data licensingGranting specific, limited rights to use a dataset while the owner retains ownership.
- Data lineageA map of where data came from and every step it passed through before reaching its current form.
- Data minimizationCollecting, keeping and sharing only the data needed for a specific purpose.
- Data processing agreementA contract that sets the rules for how a processor handles personal data for a controller.
- Data processorUnder GDPR, an organization that handles personal data on behalf of a controller and only on its instructions.
- Data provenanceThe documented origin and history of a dataset: where it came from, who authorized it and what was done to it.
- Data residencyA requirement or commitment that data is stored and processed in a specific country or region.
- Data sovereigntyThe principle that data is subject to the laws of the country or region where it is collected or stored.
- Data stewardThe person responsible for the quality, documentation and appropriate use of a dataset.
- Data trustA legal structure in which independent trustees manage data on behalf of beneficiaries.
- DatasetA defined collection of data assembled for a purpose, such as training or evaluating a model.
- Dataset cardA short document describing a dataset: contents, sources, collection, preparation, intended uses and limitations.
- Dataset licensingLicensing a specific, defined dataset rather than ongoing access to a system.
- Datasheet for datasetsA structured questionnaire, proposed by Gebru and colleagues in 2018, for documenting a dataset's motivation, composition, collection and uses.
- DiarizationIdentifying who spoke when in an audio recording, separating speakers into labeled turns.
- Differential privacyA mathematical technique that adds carefully measured noise to data or results so individuals cannot be singled out.
- Egocentric videoVideo recorded from a person's point of view, usually with a head-mounted or body-worn camera.
- EmbeddingsLists of numbers that represent the meaning of text, images or other data so similar items sit close together.
- Enterprise AIAI systems built for or deployed within organizations to perform business tasks.
- Enterprise dataData an organization generates while operating: communications, transactions, documents, tickets, code and workflows.
- ETLExtract, transform, load: moving data out of source systems, cleaning and reshaping it, and loading it somewhere new.
- Eval setA collection of examples used to measure model performance, kept separate from training data.
- Exclusive licenseA license where the owner agrees not to license the same data to anyone else within an agreed scope.
- Expert determinationThe HIPAA method where a qualified statistical expert certifies that the risk of identifying individuals is very small.
- Field of useA contract term limiting a license to specific purposes, products or industries.
- Fine-tuningFurther training of an existing model on a smaller, focused dataset to improve it for a task or domain.
- Golden datasetA carefully verified set of examples with trusted correct answers used as a benchmark.
- Held-out dataData deliberately kept out of training so it can test whether a model generalizes.
- Human dataData created by people rather than generated by models, including expert work products and human judgments.
- IndemnificationA contract promise by one party to cover the other's losses if certain things go wrong, such as a third-party claim.
- Instruction tuningTraining a model on instruction and response pairs so it follows requests well.
- K-anonymityA privacy measure where each record is indistinguishable from at least k−1 others on identifying attributes.
- Knowledge baseA collection of articles, guides and answers a company maintains so customers or staff can solve problems themselves.
- Lawful basisOne of the six legal grounds under GDPR that permit processing personal data, such as consent, contract or legitimate interests.
- Legitimate interestA GDPR lawful basis that allows processing when a company's interest is balanced against, and not overridden by, individuals' rights.
- Minimum guaranteeA fixed amount a buyer agrees to pay regardless of how much revenue share or usage fees would otherwise produce.
- Model evaluationMeasuring how well an AI model performs on defined tasks.
- Most-favored-nation clauseA promise that a party will receive terms at least as good as those given to others.
- Multimodal dataData combining more than one type, such as text with images, audio or video.
- Named entity recognitionNER: software that finds and labels names of people, companies, places, dates and similar items in text.
- Non-exclusive licenseA license that lets the owner license the same data to other buyers.
- OCROptical character recognition: software that turns text in scans, photos or PDFs into machine-readable text.
- Operational dataRecords of how a business runs day to day, such as work orders, tickets, approvals and process logs.
- Opt-outA choice that lets a person stop a company from using their data for a particular purpose.
- Personally identifiable informationInformation that can identify a person directly or in combination, such as names, emails, account numbers, faces or voices.
- PHIProtected health information: health details tied to an identifiable person, such as diagnoses, treatment notes, claims or appointment records.
- PII detectionAutomated scanning that finds personal details, such as names, emails, phone and account numbers, in data.
- Post-trainingTraining steps after pretraining that shape a model's behavior, such as instruction tuning and reinforcement learning.
- Preference dataExamples where people compare two or more model responses and pick the better one.
- PretrainingThe first, large-scale stage of training a language model on broad text to learn general patterns.
- Process supervisionTraining that rewards each correct step in a model's reasoning, rather than only the final answer.
- Proprietary dataData a company owns and controls that is not publicly available.
- PseudonymizationReplacing direct identifiers, such as names, with codes so records no longer point to a person without extra information.
- Re-identificationLinking supposedly anonymous data back to a specific person, often by combining it with other information.
- Real-world dataData collected from actual activity rather than simulated or generated.
- Red teamingDeliberately testing an AI system by trying to make it fail, misbehave or reveal information it should not.
- RedactionRemoving or masking sensitive information, such as names, account numbers or addresses, from a record before it is shared.
- Reinforcement learningA training approach in which a model learns by receiving rewards or penalties for its actions.
- Retention policyA company's rules for how long each type of record is kept and when it is deleted.
- Retrieval-augmented generationRAG: an AI technique where a model looks up relevant documents at answer time and uses them to write its response.
- Revenue shareA payment model where the licensor receives a percentage of revenue tied to the licensed asset.
- Reward modelA model trained to score responses by how well they match human preferences or task success.
- RL environmentA simulated setting where an AI agent takes actions, receives feedback and learns through reinforcement learning.
- RLAIFReinforcement learning from AI feedback: using an AI model's judgments in place of, or alongside, human feedback.
- RLHFReinforcement learning from human feedback: training a model using human preference judgments.
- Safe Harbor de-identificationThe HIPAA method that de-identifies health data by removing 18 specified identifiers and confirming no actual knowledge of re-identification.
- SOPStandard operating procedure: a written, step-by-step description of how a team does a recurring task.
- Speech-to-textSoftware that converts spoken audio into written text.
- Standard contractual clausesTemplate contract terms approved by the European Commission for transferring personal data from the EU to other countries.
- SublicensePermission for a licensee to grant some of its rights to third parties.
- Supervised fine-tuningFine-tuning with labeled input and output examples showing the desired response.
- Synthetic dataData generated by models or simulations rather than collected from real activity.
- Ticket dataRecords from help desk or issue tracking tools: requests, conversations, status changes, assignments and resolutions.
- Token countThe number of tokens in a dataset or text, used to estimate size and cost.
- TokensThe small chunks of text, often word pieces, that language models read and produce.
- Tool-use dataExamples of a model or person calling tools, such as search, APIs or calculators, and using the results.
- Warranty of titleA seller's promise that it owns, or has the right to license, what it is providing.