Skip to content

Getting started

How AI training on business records actually works

By SourceX Editorial · Updated

Short answer

AI training on business records turns prepared records into examples that adjust a model's internal weights, so the model learns patterns rather than storing your database. Licensed records pass through four steps: preparation, training or evaluation, testing, and deletion or retention under the contract. Because rare verbatim memorization is possible, preparation removes personal and confidential details first.

Key takeaways

  • A trained model holds adjusted weights, not files; your records are not stored inside it as a searchable database.
  • Some records are used only for evaluation, as tests of whether a model can do real work, and never for training.
  • Models can occasionally reproduce rare strings they saw, which is why names, account numbers and secrets are removed before delivery.
  • What happens to the delivered records after training is set by the contract: deletion, retention for a term, or limited further use.

Where do AI developers get training data today?#

AI developers get training data from several sources, and each fills a different gap. Seeing the whole mix explains why operating companies' records have started to interest them.

The last row is where your company's records fit. Support tickets linked to fixes, estimates linked to jobs and invoices, and NCRs linked to corrective actions show how real work unfolds over time, which the other sources struggle to supply.

Where do AI developers get training data today?
SourceWhat it providesMain limits
Public web crawlsBroad language and general knowledgeLittle detail on how work is done inside companies
Licensed publisher and media contentEdited, high-quality text and archivesCovers published output, not internal decisions
Data vendors and brokersPackaged datasets on specific topicsProvenance and rights can be hard to verify
Contracted human expertsWritten examples, ratings and correctionsCreated for the purpose, so costly and sometimes artificial
Simulated environmentsPractice tasks where models act and get scoredOnly as realistic as the scenarios their designers build
Enterprise archivesReal workflows with requests, decisions and outcomesNeed a rights review and privacy preparation

Step 1: preparation#

Preparation turns raw exports into a clean, documented dataset before anything leaves the supplier's control. Personal and confidential details are removed or masked, duplicates and system noise are filtered out, and records are shaped into examples, such as a support thread with its resolution or an issue with its code change and review.

Automated tools help but do not finish the job. The open-source Presidio project, which detects and anonymizes personal information in text, warns in its own documentation that there is no guarantee it will find all sensitive information and that additional systems and protections should be employed. Good preparation therefore pairs automated scanning with human review of samples.

Preparation also produces documentation: what the records are, where they came from, what was removed and which uses are permitted. Buyers rely on it to decide how the data can be used.

  • Mask or remove names, email addresses, phone numbers, street addresses and account numbers.
  • Strip secrets such as passwords, API keys and access tokens from tickets, logs and code.
  • Remove confidential customer material, internal pricing and third-party content outside the license.
  • Drop duplicates, auto-replies and system notifications that add no signal.
  • Pair each request with its decision and outcome so a record reads as one complete example.

Step 2: training or evaluation#

Training means showing a model many examples and nudging its internal weights so it gets better at producing the right output. Each example shifts the weights by a tiny amount; no single record is filed away for later lookup.

Business records are used in a few distinct ways. In fine-tuning, a model sees examples of tasks done well, such as a ticket and the reply that solved it. In reinforcement-style training, a model attempts a task and is scored against an outcome, such as whether its proposed fix matches the change that resolved the issue. In evaluation, records are held back entirely and used as a test the model has never seen.

Evaluation use deserves attention because it is often the first use buyers want. Real resolved cases make hard, realistic tests, and evaluation sets are deliberately kept out of training so that scores stay honest.

Step 3: testing#

Testing checks whether training worked and whether it created new problems. Developers measure performance on held-out tasks, compare against earlier model versions and look for regressions in other skills.

Careful developers also probe for leakage by trying to make the model repeat specific strings from its training data, such as names or identifiers. Thorough preparation makes that test far less worrying, because the sensitive strings were never in the delivered data.

Step 4: deletion or retention#

Deletion or retention of the delivered records is decided by the license, not by the technology. Common terms include deleting delivered files after training or at the end of a term, retaining them for evaluation during the term, and confirming deletion in writing on request.

Removing a dataset's influence from a trained model is a different matter. Once weights have been adjusted, extracting one dataset's effect generally means retraining, which is why permitted use, term and deletion are settled before delivery rather than after.

What a model keeps and does not keep#

A model keeps statistical patterns learned across many examples and does not keep a retrievable copy of your files. The practical risks sit at the edges, in rare strings and in how carefully records were prepared.

Read the table as a guide to where preparation effort should go. The rows marked occasionally or should not be are the ones a supplier controls before delivery.

What a model keeps and does not keep
ItemKept in the model?Why
Your raw files and databaseNoTraining adjusts weights; files are not stored for lookup
Patterns in how problems get solvedYesLearning those patterns is the purpose of training on workflow records
Customer names and account numbersShould not beRemoved or masked during preparation
Rare exact strings that slip throughOccasionallyModels can memorize unusual text seen during training
Your company's identityUsually notSupplier names can be removed and attribution restricted by contract
Records used only for evaluationNoEvaluation sets are scored against, not trained on

Illustrative: support tickets through all four steps#

Illustrative: a fictional maker of route-planning software for landscaping businesses licenses several years of resolved support tickets linked to Jira issues. In preparation, customer names, phone numbers, addresses and account IDs are masked, internal pricing notes are removed, and each ticket is paired with its resolution and linked fix.

The buyer holds back a slice of tickets as an evaluation set to test whether its model can diagnose problems in scheduling software, and uses the rest for fine-tuning. Testing shows better diagnosis on the held-out tickets, and leakage probes return no customer details. At the end of the term, the buyer deletes the delivered files and confirms deletion, as the license required.

How SourceX fits into the process#

SourceX is not a model developer, though the signed agreement does license it to use the deidentified dataset, including for model training. Its job is the transaction around the records, run as the SourceX five-step transaction (Supply, Rights, Preparation, Approval, Delivery), and nothing ships until the supplier approves the release.

Before any file moves, the SourceX Evidence Packet sets down provenance, licensing rights, permitted use, the privacy record and release authorization, so the split between training and evaluation use and the retention and deletion terms are agreed in writing.

Frequently asked questions

Can a model repeat my customers' names after training?

It should not, if names were removed before delivery. Models can occasionally reproduce rare strings seen in training, which is why preparation masks names, contact details, account numbers and secrets rather than relying on the model to forget them. Ask to see the preparation approach and samples before approving release.

Does the AI developer get access to our live systems?

No. Licensing works from prepared exports, not system access. Records are exported, prepared and approved before delivery, and large datasets can stay in the supplier's own storage or ship on encrypted drives. The buyer never logs in to your help desk, CRM or ERP.

Is evaluation-only use lower risk than training?

It can be, because evaluation records score a model rather than change its weights, and contracts can limit them to testing. The same preparation still applies, since evaluation files are handled by the buyer's staff and systems. Some suppliers start with evaluation use to get comfortable with the process.

Will a model trained on our records help our competitors?

A model trained on many sources learns general patterns of how work gets done; it does not hand competitors your customer list or pricing, which preparation removes. Consider which record families reveal competitive strategy, keep those out of scope, and use the contract to restrict attribution to your company.

Can we verify that delivered data was deleted?

Verification usually relies on contract terms: deletion obligations, written confirmation and, in some deals, audit rights. Technical proof of deletion is difficult for any data shared with another party, so the quality of the counterparty and the contract matter more than any single tool.

Sources

  • Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify