Skip to content

Getting started

Is AI already training on my company's data without permission?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

AI may already be training on your company's public web pages, but probably not on its private records. Public pages are likely crawled. Data held by your software vendors depends on their terms and your settings. Private systems are not reachable without a breach, a leak or a license you approve. Check each branch separately.

Key takeaways

  • Anything on your public website may already have been collected by crawlers, and blocking them now does not recall it.
  • A software vendor's right to use your data for AI is set by its terms, any AI addendum and your admin settings.
  • Records in private systems reach AI developers only through a security failure, staff copying them into AI tools, vendor terms you accepted, or a license you approve.
  • A frequent practical exposure is employees pasting company records into consumer AI tools.

Three places your data can be, and what decides each#

Whether AI trains on company data depends on three branches: what is public, what your vendors hold and what stays in your private systems. Each branch has a different answer and a different set of controls, so a single yes or no is misleading.

Work through the branches in order. The first is mostly about accepting what has already happened, the second about reading contracts and settings, and the third about how your own people use AI tools.

Three places your data can be, and what decides each
Where the data sitsCan AI developers reach it?What decides itWhat you can do
Public website, documentation, blog, job postsLikely, through web crawlersCrawler behavior, robots.txt, site termsReview robots.txt and terms; assume past content was collected
Software vendors' systemsPossibly, through the vendor's own AI featuresVendor terms, AI addenda, your admin settingsRead the AI clauses, apply opt-outs, record your choices
Private systems and archivesNot without access you grant or loseYour security, staff behavior, contractsControl AI tool use and access; license only on your terms

Branch one: your public web pages#

Public web pages are the branch most likely to have been used already. Marketing pages, product documentation, help-center articles, blog posts, press releases and job posts are all reachable by crawlers, and AI developers have collected public web text at large scale.

Draw the line carefully. A public help-center article about how to reset a device is public; the support conversations that led someone to write it are not. Most of what makes a company's records distinctive never appears on its website.

Some AI developers publish the names of their crawlers and explain how to block them in robots.txt. A robots.txt file is a request, not a lock: it works only for crawlers that choose to honor it, and it applies going forward. Content already collected stays collected, and removing data from a trained model is difficult.

Website terms of use are the other lever. They can state that automated collection for AI training is not permitted, which records the company's position even though a crawler may never read them. Whether such terms can be enforced against a particular developer is unsettled, so treat them as a statement of intent and have counsel review the wording.

Branch two: data your software vendors hold#

Data held by your software vendors may be used for AI depending on their terms and your settings. Helpdesks, CRMs, call recording tools, field service platforms and productivity suites keep adding AI features, and their contracts differ on whether customer data is used only to serve you or also to improve models shared across customers.

Read the documents rather than the marketing page. The answer usually sits in a combination of the main terms, a data processing agreement, an AI-specific addendum and an admin setting.

  • Find the clause on using customer data for product improvement or model training.
  • Check whether the vendor treats de-identified or aggregated data differently from your raw records.
  • Look for an opt-out or admin setting, apply it if you choose to, and keep a screenshot.
  • Note whether the vendor can change its terms by notice, and who in your company receives those notices.
  • Check whether your plan tier changes the commitments, since business and consumer plans often differ.

Branch three: your private systems#

Private systems such as your ERP, file servers, internal wikis and archived helpdesks are not reachable by AI developers unless access is granted or lost. No crawler can read a closed system. Apart from software vendors acting under terms you accepted, which is branch two, the legitimate route for an AI developer to obtain internal records is an agreement your company approves.

Exposure in this branch often comes from inside rather than outside. The common paths are employees pasting customer emails, estimates or code into consumer AI tools; drive or wiki links set to anyone with the link; contractors working on personal devices; and, less often, a security breach. Each has an ordinary control: an acceptable-use policy with approved tools, link-sharing settings, contractor agreements and basic security hygiene.

Old systems deserve a look too. A helpdesk or CRM the company stopped using may still hold years of records on a vendor account nobody monitors, under terms nobody has reread since signing. Either bring those accounts into the vendor review or archive the records and close the account properly.

A quick exposure check#

A quick exposure check needs one owner, a list and a few documents. It turns a vague worry into a short written record that management and the board can rely on.

  • Open your robots.txt and website terms, and decide whether they reflect your current position on AI crawlers.
  • List every vendor that holds customer, employee or operational records.
  • Read each vendor's AI and data use clauses and check the admin settings.
  • Ask team leads which AI tools staff use, and with what kinds of records.
  • Search for publicly shared links to company drives and wikis, and close the ones that should not be open.
  • Write down what you found, what you changed and who owns each follow-up.

Illustrative: a roofing company owner checks all three branches#

Illustrative: the owner of a fictional roofing and restoration company hears that a competitor is generating estimates with AI and wonders whether the company's own work helped train the tool. The office manager runs the exposure check.

The public branch shows that the company's project galleries and blog have been online for years and were likely crawled; the owner accepts that and updates robots.txt going forward. The vendor branch turns up a call recording tool whose terms allow it to use de-identified recordings to improve its AI; an admin opt-out exists and is applied. The private branch reveals that estimators had been pasting customer emails into a free chatbot to draft replies.

The owner adopts an approved AI tool on business terms, adds a one-page policy and confirms that the company's estimates, job photos and warranty history have never left its own systems. Those private records remain something the company could choose to license later, on its own terms.

Where licensing fits, and how SourceX approaches it#

Licensing is the route by which private operational records reach AI developers with the company's consent, and it requires an authorized signature. The company keeps ownership and decides the permitted use, because the records are licensed rather than sold outright.

SourceX runs that route as a managed transaction in five steps, Supply, Rights, Preparation, Approval and Delivery, known as the SourceX five-step transaction. Nothing is shared during the initial assessment, the supplier approves every step, and the SourceX Evidence Packet records the permitted use and release authorization for anything that is licensed. That written trail is the opposite of the untracked exposure described in the three branches above.

Frequently asked questions

Can I find out whether a specific AI model trained on my website?

Usually not with certainty. Few developers publish complete lists of training sources, and models rarely reproduce a page word for word. Your server logs may show visits from named AI crawlers, which tells you a page was fetched, not whether it was used in training.

Does blocking AI crawlers now remove what they already collected?

No. Blocking applies to future crawling by crawlers that honor it. Content already collected may remain in datasets, and removing the influence of specific data from a trained model is technically difficult. Blocking is still worth doing if you want to limit future collection.

Is it legal for AI companies to train on public web content?

That question is being argued before courts and regulators in several countries, and the answer may differ by jurisdiction, content type and how the content was used. Treat it as unsettled, and ask counsel if a specific use of your content concerns you.

If AI has already learned from our public content, are our private records still worth anything?

Often yes. Public content shows what you say about your work; private records such as tickets, jobs, orders and quality reports show how the work was done and what happened. That internal detail is exactly what public sources lack.

Should we stop using AI features in our business software?

Not necessarily. Many AI features use your data only to serve you. The decision rests on each vendor's terms and settings, so read the AI clauses, apply the opt-outs you want and record the choice rather than switching features off by default.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify