Skip to content

Data licensing for AI training

Fine-tuning hosted models on licensed data: processor and API terms

Quick answer

You can upload licensed data to a hosted fine-tuning API only if your data license permits disclosure to that provider as your processor for that purpose. The provider's terms never grant that right; they only describe how the provider handles your files. Before the first upload, confirm four things: the grant covers training on third-party infrastructure, the provider is a permitted subprocessor, retention and deletion at the provider match the license, and you hold the resulting fine-tuned weights on terms the license allows.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

The data license decides; provider terms only describe handling

The license from the data owner is the only document that can authorize sending its records to OpenAI, Azure OpenAI, Vertex AI or any other hosted tuning service. A provider's promise that it will not train on your uploads limits the provider's use, but it does not expand what your licensor allowed you to do. Developer forums show the confusion: teams ask whether the provider will use their fine-tuning files [1], when the prior question is whether they may upload them at all.

Three license patterns cause most problems. An "internal use only" grant that limits access to "Licensee's employees and contractors" may not reach a model provider's infrastructure staff or automated pipelines; see internal-use-only data licenses. A grant that says "train models on Licensee's systems" can be read to exclude a hosted job unless "systems" is defined to include cloud services. A fine-tuning-only license may cover the use but stay silent on where the computation runs.

For the broader map of rights, terms and pricing, start at the AI training data licensing hub. If you are sending licensed content to a model API for retrieval rather than tuning, the issues differ; see sublicense and processor terms for RAG. Teams still sourcing data can describe the records they need and raise hosted fine-tuning as an intended use when license terms are agreed.

Grant language that actually covers a hosted fine-tuning job

A grant covers hosted fine-tuning when it names the activity, the infrastructure and the provider's role in one place. Look for (or negotiate) a definition of "Authorized Processor" that includes cloud and model providers acting on the licensee's documented instructions, a list or category of named providers, and an express statement that disclosure to an Authorized Processor for training is not a sublicense.

The processor framing matters because a hosted tuning service receives records, transforms them into training batches and stores the outputs. Under GDPR Article 28, a processor acts only on the controller's documented instructions, and a processor that decides purposes and means on its own is treated as a controller [4]. Your data license should mirror that logic for non-personal licensed content too: the provider must act only for you, with no independent right to use the data. The data processor glossary entry explains the role in more depth, and writing the AI training rights grant covers definitions.

Watch the "Licensed Materials" definition as well. Many licenses extend restrictions to "copies, extracts and derivatives," which can sweep in the JSONL training file you build, the provider's cached copy and, depending on drafting, the fine-tuned weights.

What to verify in the model provider's terms

The provider check is a contract review of five terms: training use, retention, deletion, region and model exclusivity. Market practice among the large providers is broadly similar on paper, but the binding version lives in your enterprise agreement, data processing addendum (DPA) and service-specific terms, not in marketing pages.

  • No training on customer data. The large providers publish commitments that customer fine-tuning data is not used to train their base models, and developers routinely ask how far those commitments reach [1]. Confirm the commitment is contractual and covers uploaded files, job metadata and validation sets.
  • Retention of training files. On OpenAI, uploaded fine-tuning files persist until you delete them, according to policy quoted in developer discussions [1]; a job references the uploaded training_file and produces a model identifier in your organization [2]. On Vertex AI, JSONL datasets are staged in your own Cloud Storage bucket [3], so retention follows your bucket policy plus whatever the service copies during the job.
  • Deletion mechanics. Providers typically expose delete calls for files and fine-tuned models. Record which API calls delete the file, the job artifacts and the model, and whether backups are covered.
  • Region of processing. Some providers store training data and fine-tuned models in the resource's region but also offer global or multi-region deployment types that can process data elsewhere. If your license or a privacy law restricts transfers, use regional deployments and document the choice; global deployment types can process data outside the resource region [9].
  • Model exclusivity. Forum answers quoting OpenAI policy say fine-tuned models are reserved for the customer that created them [1]. Confirm the same in writing for any provider you use.

Treat published promises as enforceable expectations, not as fixed. The FTC has warned that model-as-a-service companies may face liability for breaking promises not to use customer data for training [6], and that quietly changing terms to permit more data use can be unfair or deceptive [7]. Even so, your license compliance should rest on signed terms and a change-notice clause, not on regulator posture. The software vendor terms guide covers how platform terms interact with data licenses.

Who holds the fine-tuned model, and what happens when the license ends

A hosted fine-tune creates an asset you control logically but cannot usually export: on most proprietary-model services the adapted weights stay inside the provider's platform under your account. That makes three license questions concrete. Does the license treat the fine-tuned model as a derivative of the licensed data? May you keep serving it after term? And must you delete it, and prove deletion, at the provider?

If the license requires model deletion on termination, you need the provider's model-delete endpoint and a log of its execution, because you cannot "return" weights you never held. If the license permits retention, make sure it names hosted models expressly. See what happens to trained models when a data license ends, derivative and successor model rights and deletion and return clauses.

Personal data, PHI and cross-border processing

When licensed records contain personal data, the hosted provider becomes your processor or subprocessor, and the chain from licensor to you to provider must hold together. Under GDPR, that means a DPA with Article 28 terms with the provider [4], a lawful basis for training that you can document, and transfer mechanisms that match the provider's actual processing region.

Health data adds HIPAA. If records were de-identified under 45 CFR 164.514 (Safe Harbor or Expert Determination), uploading them is outside HIPAA's disclosure rules, but an Expert Determination may be conditioned on the recipient environment, so check whether a hosted provider was in scope [5]. A limited data set requires a data use agreement, and onward disclosure to a provider needs its own analysis [5]. Identified PHI generally requires a business associate agreement with the provider before upload.

Fine-tuning a hosted model usually does not make you a general-purpose AI model provider; Commission guidance reserves that for significant modifications. If your modification is significant and the model is placed on the EU market, Article 53 duties on documentation, copyright policy and a training-content summary may apply, and have applied since 2 August 2025, with AI Office enforcement for new models from 2 August 2026 and models already on the market from 2 August 2027 [8]. Your license should let you describe the licensed source in that summary.

Pre-upload clearance checklist

Run this checklist before any licensed file leaves your environment for a hosted tuning job.

Illustrative example: invented to show structure; it does not describe an available dataset.

#CheckWhere to lookPass condition
1Grant covers model trainingLicense: grant and "Permitted Purpose"Training or fine-tuning named expressly
2Hosted infrastructure allowedLicense: "Authorized Processor," "Licensee Systems"Cloud or model providers permitted, by name or category
3Provider disclosure is not a sublicenseLicense: sublicensing clauseDisclosure to processors expressly carved out
4Provider will not train on uploadsEnterprise agreement, DPA, service termsContractual no-training commitment covering files and outputs
5Retention and deletion alignProvider docs and DPA; license deletion clauseFile, job artifact and model deletion paths documented
6Region matches license and privacy lawDeployment type, DPA transfer termsRegional deployment or approved transfer mechanism
7Fine-tuned model status is clearLicense: derivatives, post-term rightsHosted model retention or deletion rule stated
8Personal data handledLicensor de-identification record; DPA or BAAMethod recorded; BAA in place if PHI remains
9Change noticeProvider termsNotice and exit right if data-use terms change
10Audit trailInternal ticketUpload IDs, job IDs, model IDs and deletion logs stored

Illustrative rider for a hosted fine-tuning use

A short rider can close the most common gap when the base license predates hosted tuning. Adapt the wording with counsel.

Illustrative example: invented to show structure; it does not describe an available dataset.

Authorized Processors. Licensee may disclose Licensed Data to a cloud
or model service provider ("Authorized Processor") solely to perform
fine-tuning, evaluation and hosting of Licensee Models, provided the
Authorized Processor (a) acts only on Licensee's documented instructions,
(b) is contractually prohibited from using Licensed Data to train or
improve its own or third-party models, (c) deletes Licensed Data on
Licensee's instruction, and (d) processes Licensed Data only in the
regions listed in Schedule B. Such disclosure is not a sublicense.
Licensee remains responsible for each Authorized Processor's acts.

Failure modes to plan for in hosted fine-tuning

Most problems surface after the job has run, when deletion or audit requests arrive. Failures to plan for include validation files uploaded but forgotten when training files were deleted; a Global deployment chosen for capacity that breached a regional restriction; a license that allowed "internal models" read as excluding provider-hosted weights; and personal fields left in system prompts or metadata that the de-identification pass never touched. Each one is cheaper to prevent with the checklist above than to remediate.

The same discipline applies to evaluating the data itself; see how to evaluate a fine-tuning dataset before you buy it and the fine-tuning glossary entry.

How SourceX supports licensed data for hosted fine-tuning

SourceX sources operational datasets from US companies on request and manages the licensing process. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery. If you plan to fine-tune a hosted model, describe the data and the provider setup when you submit a buyer request.

Sources

  1. OpenAI Developer Community, "Is the data I upload for fine-tuning used by OpenAI?". https://community.openai.com/t/is-the-data-i-upload-for-fine-tuning-used-by-openai/831241
  2. OpenAI, "Fine-tuning API reference". https://developers.openai.com/api/docs/api-reference/fine-tuning
  3. Google Cloud, "Prepare supervised fine-tuning data for Gemini models". https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini-supervised-tuning-prepare
  4. GDPR-Text.com, "Article 28: Processor" (Regulation (EU) 2016/679). https://gdpr-text.com/en/read/article-28/
  5. eCFR, Office of the Federal Register / HHS, "45 CFR 164.514" (2026). https://www.ecfr.gov/current/title-45/subtitle-A/subchapter-C/part-164/subpart-E/section-164.514
  6. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  7. Federal Trade Commission, Office of Technology, "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  8. Official Journal of the EU, via EUR-Lex, "Regulation (EU) 2024/1689 (AI Act), consolidated text of 27 July 2026". https://eur-lex.europa.eu/legal-content/EN/TXT/HTML/?uri=CELEX:02024R1689-20260727
  9. Microsoft, "Azure OpenAI Service deployment types". https://learn.microsoft.com/en-us/azure/ai-services/openai/how-to/deployment-types

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data