Skip to content

Definitions and comparisons

Training vs inference: when does an AI model actually use your data?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Training uses data to change a model's weights, so what the model learns stays in it; inference runs a finished model on an input to produce an answer without changing the model. The plain rule for contracts: the license grant, permitted use and deletion terms govern training, while service terms, confidentiality and retention govern inference.

Key takeaways

  • Training changes the model; inference uses the model without changing it.
  • Data used in training is hard to remove from a trained model, so licenses settle that question before delivery.
  • Retrieval systems use your documents at inference time but keep copies and embeddings in an index that can be deleted.
  • Whether a vendor trains on your prompts depends on its terms and your plan, not on how inference works.
  • Evaluation is a third use: data tests a model without training it, and a license can grant it on its own.

What is the difference between training and inference?#

Training is the process of adjusting a model's internal parameters, its weights, by exposing it to data, while inference is running the trained model on new input to produce an output. Training happens before a model is put to work; inference happens every time someone asks it something.

The difference matters for control. After training, patterns from the data are spread across the model's weights, and the original files are no longer needed to run it. During inference, the input is processed to produce an answer and the model does not change, although the service around it may log both the input and the output.

Training vs inference side by side#

Training and inference differ in what happens to the data, how long the effect lasts and which part of a contract controls it. The two-column comparison below sets them side by side.

The row about undoing is the one owners tend to miss. Deleting a file after training does not delete what the model learned from it, which is why training rights are negotiated before data is delivered rather than after.

Training vs inference side by side
QuestionTrainingInference
What happens to the dataUsed to adjust model weightsProcessed as input to produce an output
Does the model change?YesNo
How long the effect lastsAs long as the model and models derived from it existAs long as logs, caches or stored inputs are kept
Can it be undone?Hard; usually means retraining without the dataYes, by deleting logs and stored inputs
Typical exampleA developer trains on licensed support ticketsAn agent asks an assistant to summarize one ticket
Clause that governs itLicense grant, permitted use, deletion and model termsService terms, confidentiality, retention and the DPA

Where do fine-tuning, retrieval and evaluation fit?#

Fine-tuning is a form of training, retrieval is a form of inference, and evaluation is a separate use altogether. Each is common in AI products, and each should be named in a license rather than assumed from a general phrase such as AI purposes.

  • Fine-tuning: further training of an existing model on a narrower dataset. It changes weights, so it counts as training.
  • Retrieval-augmented generation: documents are indexed, and relevant passages are fetched and handed to the model at inference time. The model does not change, but copies sit in an index.
  • Embeddings: numerical representations of text stored for search. They are derived from your data and belong in the deletion terms.
  • Evaluation: using records to test how well a model performs without training on them. Some licenses grant evaluation only.
  • Synthetic data generation: using records to produce new examples. A license should say whether this is allowed and who owns the results.

Does AI learn from my prompts?#

Whether an AI service learns from your prompts depends on the provider's terms and your plan, not on the mechanics of inference. Inference by itself does not train the model, but a provider may keep prompts and outputs and later use them for training if its terms allow it.

Vendor terms differ on this point, and plans of the same product can differ from each other. GitHub's Terms of Service, for example, grant GitHub a license to use AI-feature inputs and outputs to train AI models, with an opt-out in account settings, while customers under a GitHub Customer Agreement are excluded from that training license. Asana's Product-Specific Terms, by contrast, say Asana does not use Customer Data to train the generative AI models used to provide Asana AI. Read the data use and retention sections for the plan your company actually pays for, check admin settings for any training opt-out, and record the answer in your vendor inventory.

The weak point is usually behavior rather than contracts. Staff who paste customer records into a personal account bypass whatever the company plan says, so an acceptable use policy matters as much as the vendor's terms.

Which contract clause governs each use?#

The plain rule is that training is governed by what you grant, and inference is governed by what the service may keep. A data license controls training through its grant and permitted use, and it should address deletion, derived data and trained models. A service agreement controls inference through confidentiality, retention, data processing terms and any clause about using inputs for training.

This is general information, not legal advice. Clause wording and applicable law vary, so review agreements with counsel before relying on any of them.

Which contract clause governs each use?
UseClause to readQuestion to ask
Training or fine-tuningGrant and permitted useIs training named, and for which models?
Models after the term endsModel and deletion termsMay models trained during the term stay in use?
RetrievalPermitted use and deletionAre indexes and embeddings deleted at the end?
EvaluationPermitted useIs evaluation allowed separately from training?
Prompts and outputs in a vendor toolData use, retention and DPACan the vendor train on inputs, and can we opt out?

A short checklist for owners#

A short checklist keeps the two uses apart in practice, because the people who buy AI tools and the people who negotiate data licenses are rarely the same people.

  • List every AI tool in use, the plan it runs on and whether its terms allow training on inputs.
  • Record opt-out settings and who in IT controls them.
  • For any data license, name each permitted use: training, fine-tuning, evaluation, retrieval or synthetic generation.
  • Ask how models trained during the term are treated after it ends.
  • Require deletion of copies, indexes and embeddings, with written confirmation.
  • Keep the answers in one place so finance, legal and IT read the same record.

Illustrative: a machine shop separates two AI questions#

Illustrative: a fictional precision machining company uses an AI assistant to summarize nonconformance reports from its QMS, and a model developer has asked about licensing its NCR and CAPA history. Leadership has been treating both as one question about AI and data.

The quality manager splits them. For the assistant, IT confirms the business plan's terms on input retention and training, turns on the available opt-out and bars personal accounts for work data. For the license, counsel scopes training and evaluation on NCR and CAPA records with customer part numbers and customer-owned drawings removed, and asks that models trained during the term be addressed expressly in the deletion clause. Each question now has an owner and a contract.

How SourceX defines permitted use#

SourceX licenses state permitted use explicitly, so training, fine-tuning, evaluation and retrieval are named rather than implied. The supplier approves that scope in the Approval step of the SourceX five-step transaction, and the permitted use is recorded in the SourceX Evidence Packet alongside provenance, licensing rights, the privacy record and release authorization.

SourceX is not a model developer, though the signed agreement does license it to use the deidentified dataset, including for model training. It manages the transaction between the company that holds the records and the developer that uses them.

Frequently asked questions

Can licensed data be removed from a model after training?

Generally not in a simple way. Removing a dataset's influence usually means retraining without it, and research on machine unlearning is still developing. That is why licenses deal with trained models up front: whether models trained during the term may stay in use, and which copies of the data must be deleted.

Is evaluation data the same as training data?

No. Evaluation data tests a model's performance without changing it, and it is usually kept apart from training data so results stay honest. A license can grant evaluation only, which some owners find a comfortable first step because the data does not shape the model.

Does using retrieval on licensed data count as training?

Not technically, because the model's weights do not change. Retrieval does keep copies and embeddings of your data in an index, so a license should say whether retrieval is permitted and require deletion of indexes and embeddings when the term ends.

Can inference data become training data later?

Yes, if the provider's terms allow it. Prompts and outputs may be logged, reviewed and later used for training under some plans. Check the terms for your plan, the admin settings and any opt-out, and treat them as part of your regular vendor review.

Is a model's output a copy of my data?

Usually not in a direct sense, but models can sometimes reproduce passages from their training data, particularly text that appeared many times. Licenses address this through confidentiality duties, limits on outputting identifiable content and the privacy preparation done before delivery, which removes personal and confidential details from the data in the first place.

Sources

  • GitHub's Terms of Service (Section J) grant GitHub a license to use AI-feature inputs and outputs to train AI models, which users can opt out of, and customers under a GitHub Customer Agreement are excluded from this training license. Source
  • Asana's Product-Specific Terms state that Asana does not use or permit third-party AI partners to use Customer Data to train generative AI models used to provide Asana AI. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify