Skip to content

Consulting and recruiting

Can consulting firms use client data to train internal AI tools?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

Consulting firms can use client data to train internal AI tools only where the engagement contract and applicable privacy law allow it, and standard confidentiality clauses usually do not. The working rule: applying AI to a client's material for that client's engagement is delivery, while training a firm-wide tool on it needs explicit contract language or written client consent.

Key takeaways

  • Training a firm tool is a different use from serving the client, and many confidentiality clauses limit client information to performing the services.
  • Four clauses decide the answer: confidentiality, use restriction, aggregation or derived-data rights, and return or deletion.
  • De-identification reduces privacy exposure but does not cure a contractual use restriction.
  • Firm-owned methods, templates and internal reviews are a safer training base than client-supplied data or client-owned deliverables.
  • Record which engagements fed each model version, because a deletion request is hard to honor once material sits in model weights.

Why training is a different use from client work#

Training is a different use from client work because the client's information outlives the engagement and starts serving other clients. When a consultant summarizes a client's board pack in an approved AI tool for that client's project, the material stays inside the purpose the contract describes. When the firm fine-tunes a model on that board pack, the client's facts and reasoning become part of a firm asset.

Most consulting confidentiality clauses limit client information to performing the services for that client. Read that way, a model trained on one client's records and then used on another client's engagement is a use for another purpose, even if no document is ever shown to the second client.

The table compares common AI activities by what happens to the client's material and where contracts usually land. Contracts differ, so treat the right-hand column as a default to test against your own agreements, clause by clause with counsel.

Why training is a different use from client work
AI activityWhat happens to client materialUsual contract position
Analyzing a client document in an approved tool for that engagementUsed for that client's work and not kept by the vendor for trainingGenerally within the services, unless the client bars AI tools
Search index of one client's files, used only by that engagement teamStored and searchable, scoped to one clientUsually acceptable if storage and access terms are met
Firm-wide search over all past deliverablesRetrieved and shown to teams serving other clientsOften restricted; needs an ownership and confidentiality review
Fine-tuning a model on client deliverables or dataAbsorbed into model weights used across clientsTypically needs explicit permission or consent
Benchmarks derived from many clientsReduced to statistics that should not identify any clientAllowed only where an aggregation clause covers it
Shared prompt library with client excerptsClient text persists and circulates inside the firmSame review as firm-wide search

The four clauses that decide the answer#

Four clauses decide whether client data can train a firm tool: confidentiality, use restriction, aggregation or derived-data rights, and return or deletion. Read them together across the MSA and every SOW or engagement letter, and check whether the client's own paper or purchase terms override yours.

Two further terms change the reading. A residuals clause may let consultants keep general know-how retained in memory, but it rarely reaches structured records fed into a model. The deliverables and background IP clauses decide who owns the work product, and if the client owns a report, training on it uses the client's property, not just its confidential information.

  • Confidentiality: how client information is defined, whether it covers material the firm created from client inputs, and which exceptions apply, such as public or independently developed information.
  • Use restriction: whether client information may be used solely to perform the services. A sole-purpose clause is the most common block on training.
  • Aggregation and derived-data rights: whether the firm may create and keep anonymized or aggregated data, for what stated purpose, and whether the clause mentions models, machine learning or improving the firm's services.
  • Return or deletion: what must be returned or destroyed when the engagement ends, whether derived materials are included, and what certification the client can demand.

Does de-identifying client data make training acceptable?#

De-identifying client data reduces privacy exposure but does not cure a contractual use restriction. Removing names from interview transcripts deals with the personal data of the client's employees; it does not change a clause that limits all client information to the engagement.

Consulting material is also hard to de-identify. A pricing study for a regional distributor, a reorganization plan or a carve-out model can point to one client through industry, geography, timing and unusual events even with every name removed. Treat anything a knowledgeable reader could attribute to a client as still identifying.

Privacy laws may apply on top of the contract. Under GDPR, information that is truly anonymous falls outside the regulation, but Recital 26 says pseudonymised data that could be attributed to a person with additional information is still information about an identifiable person. In California, the CCPA treats information as deidentified only if, among other conditions, the business publicly commits to keep it in deidentified form and not to re-identify it. Whether a given training use fits the purposes people were told about is assessed case by case with counsel.

Which firm records are safer to train on?#

Firm-owned records are the safer training base: methods, templates, internal reviews and operating records the firm created for itself. They still need a check for embedded client details, but the starting position is very different from client-supplied data.

Which firm records are safer to train on?
Record typeTypical systemWho usually controls itTraining posture
Methodology decks, frameworks, templatesSharePoint, Confluence, NotionFirm, as background IPUsually usable after client examples are removed
Internal project reviews and lessons learnedShared drives, PSA notesFirmUsable once client names and facts are stripped
Proposals and statements of workCRM, proposal libraryFirm, but may quote client RFPsUsable after RFP content and pricing are reviewed
Staffing, time and project financialsPSA such as Deltek or KantataFirmUsable for internal operations tools; holds employee data
Interview notes and transcriptsTeams, Zoom, OneDriveClient confidential, with personal dataExcluded without consent
Client-supplied data and models built from itExcel, data rooms, client portalsClientExcluded without explicit permission
Final deliverablesSharePoint, client portalsOften assigned to the clientExcluded unless the contract reserves rights

Plan for deletion before you train#

Deletion planning matters because a client's material is hard to remove from a trained model. A retrieval index can drop a client's documents on request, while weights fine-tuned on those documents usually cannot forget them without retraining from a clean set or retiring the model.

Keep lineage from the start: which engagements, document sets and date ranges fed each model version, and under which permission. When a client invokes its return-or-destroy clause, that record tells you whether the request touches a model, an index or nothing at all.

Separate permissions by layer as well. A client may accept its documents in a firm search tool but refuse fine-tuning, so record the two answers as separate fields in the PSA or CRM rather than one AI flag.

Training consent is easiest to obtain as a specific, bounded request rather than a broad license buried in boilerplate. Client procurement and legal teams tend to accept narrow terms they can explain internally and push back on open-ended ones.

  • Name the material: specific record types, such as final reports with client identifiers removed, not all client information.
  • Name the use: internal tools used by firm staff only, never provided to other clients.
  • State the safeguards: de-identification, access controls and no attribution of any output to the client.
  • Offer withdrawal terms: what happens if the client withdraws consent, and whether that applies to future model versions only.
  • Add the clause to the MSA template and to renewals instead of sending retroactive requests to former clients.
  • Record each answer against the engagement in the PSA or CRM, linked to the signed clause.

Illustrative: a pricing consultancy decides what its assistant learns from#

Illustrative: a fictional pricing and commercial strategy consultancy wanted an internal assistant that drafts pricing diagnostics in the firm's house style. The knowledge lead proposed fine-tuning on every final report in SharePoint and on the Excel price-waterfall models built from client transaction data.

The general counsel sampled engagement letters across the archive. Older letters on the firm's paper limited client information to performing the services and assigned final reports to the client. Newer ones allowed anonymized benchmarks but said nothing about models. Several large clients had contracted on their own paper, with return-or-destroy clauses that reached derived materials.

The firm fine-tuned only on its own methodology decks, diagnostic templates and internal project reviews scrubbed of client facts. Client reports went into per-engagement search spaces that close when each engagement ends, a narrow consent clause went into the MSA template, and the PSA gained separate search and training fields. The assistant launched with no client transaction data in its training set.

How SourceX treats client material in a licensing review#

SourceX applies the same distinction when a consulting firm considers licensing records to AI developers, a stricter question than internal training because the licensed records go to an outside party. SourceX's own rights in a deidentified dataset are set out in the signed supplier agreement. In the Rights step of the SourceX five-step transaction (Supply, Rights, Preparation, Approval, Delivery), client-supplied data and client-owned deliverables are carved out unless the client has authorized the use.

Firm-owned records such as proposals, staffing and project reviews, playbooks and internal knowledge are the usual candidates. For each package that proceeds, the SourceX Evidence Packet documents provenance, licensing rights, permitted use, the privacy record and release authorization, the same lineage discipline that internal training needs.

Frequently asked questions

Does an AI vendor's promise not to train on our inputs settle the question?

No. A vendor commitment not to train on your inputs covers what the vendor does with client material, not what the firm does with it. Fine-tuning a firm model on client records through that vendor is still the firm's own use, and the engagement contract still decides whether it is allowed.

What if an old engagement had no written confidentiality clause?

Silence is not permission. The client may rely on a separate NDA, purchase order terms, professional standards or an implied duty of confidence, and privacy law applies whatever the contract says. Treat undocumented engagements as excluded from training until counsel has reviewed what governs them.

Can we train on former clients' data if we anonymize it?

Check the survival and return-or-destroy clauses first. Many confidentiality obligations survive termination, and some clients were entitled to have their material destroyed when the work ended. Asking former clients for consent is possible but rarely practical, so most firms exclude former-client material by default and build from firm-owned records instead.

Are benchmarks built from many clients the same as training data?

No. A benchmark reduces many engagements to statistics that should not reveal any one client, and an aggregation clause may expressly allow it. Training on raw reports keeps each client's facts and reasoning inside the model. If your aggregation clause names benchmarking or service improvement, read its stated purpose before stretching it to model training.

Should we tell clients what our internal tools were trained on?

It helps. Client security questionnaires often ask whether their data trains any firm or vendor model. A short written statement of what each internal tool was trained on, backed by the lineage record, answers that quickly and replaces vague assurances that procurement teams tend to probe.

Sources

  • GDPR Recital 26 states that data protection principles should not apply to anonymous information, and that personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information, should be considered information on an identifiable natural person. Source
  • Post-CPRA, Cal. Civ. Code 1798.140(m) treats information as deidentified only if the business takes reasonable measures to ensure it cannot be associated with a consumer or household, publicly commits to maintain and use it in deidentified form and not attempt to reidentify it, and contractually obligates any recipients to comply. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify