Skip to content

Getting started

How to prepare company data for AI

By SourceX Editorial · Updated

Short answer

To prepare company data for AI, work through ten steps: inventory systems, name owners, check permissions, confirm retention, measure history, link related records, find sensitive details, choose a removal method, document the result and get approvals. Start with rights and ownership rather than tools; a tidy dataset you are not allowed to use is wasted work.

Key takeaways

  • AI-ready data is defined by ownership, permission and linkage, not by how tidy the spreadsheets look.
  • Links between records, such as a ticket joined to the bug it exposed, are what make operational data useful to AI.
  • Automated PII detection helps but misses things, so pair it with human review of samples.
  • The same preparation serves internal AI tools and external licensing, but licensing needs written rights and a privacy record.
  • Approvals are granted last but should be identified first, so nobody prepares data that cannot be released.

What does AI-ready company data actually mean?#

AI-ready company data is a set of records that is findable, owned, permitted for the intended use, linked across systems, free of details that should not leave, and documented well enough for someone else to rely on. Formatting matters far less than those six properties.

Most companies prepare data for one of two purposes. The first is internal: a support assistant, a search tool over project files, or forecasting from order history. The second is external: licensing records to AI developers who use them to train or evaluate models. The steps overlap, but licensing raises the bar, because another party needs written evidence of rights and of how personal details were handled.

Licensing also changes who relies on the work. Records are licensed, not sold outright, so the company keeps ownership and sets the permitted use, but the preparation has to hold up for a reader outside the building.

The ten-step checklist#

The ten steps run roughly in order. Approvals are granted at the end, but the people who give them should be identified at the start.

  • Step 1, inventory: list each system, the record families it holds and the years still exportable.
  • Step 2, owners: name a business owner for every record family, not just a system administrator.
  • Step 3, permissions: check customer contracts, vendor terms, employee notices and IP assignments.
  • Step 4, retention: confirm what you are allowed or required to keep, and any legal holds.
  • Step 5, history and exports: test how far back each export reaches and what it leaves out.
  • Step 6, linking: keep the IDs that join a request to the work done and the outcome.
  • Step 7, sensitive data: find personal details, credentials, client confidential material and HR content.
  • Step 8, removal method: decide per field whether to exclude, redact, pseudonymize, generalize or keep.
  • Step 9, documentation: write a data dictionary, a provenance note and a list of known gaps.
  • Step 10, approvals: get sign-off from the authorized signer and any board, investor or lender consents.

Steps one to three: inventory, owners and permissions#

The first three steps establish what exists, who decides about it and whether the intended use is allowed. Skipping them is the most common reason AI data projects stall after a lot of cleanup has already been done.

Owners matter because system administrators can export data but cannot judge whether a use is appropriate. The head of support knows which customers have unusual contract terms; the quality manager knows which part numbers belong to a customer's proprietary design.

Steps one to three: inventory, owners and permissions
Record familyTypical systemsUsual ownerPermission question
Support conversationsZendesk, Intercom, FreshdeskHead of supportDo customer contracts restrict reuse of their communications?
CRM historiesSalesforce, HubSpotSales operationsWhat do privacy notices say about contact data?
Engineering workJira, Linear, GitHub, GitLabCTODoes any repository contain customer or third-party code?
Jobs and dispatchServiceTitan, Housecall Pro, JobberCOODo franchise or customer agreements control these records?
Orders and exceptionsNetSuite, Epicor, AcumaticaOperations leadDo client contracts treat order data as client property?
Quality and maintenanceQMS, MES, maintenance systemQuality managerAre customer-owned designs or export-controlled items involved?

Retention, history and linking decide how much of the record set is usable. Retention tells you what you may keep, history tells you how much you can still reach, and linking tells you whether the records explain anything.

Check retention schedules and legal holds before building any dataset, because records that should already have been deleted do not belong in a prepared set. Then test the exports. Many software plans limit how far back an export goes or leave out attachments, internal notes and audit history, so check your plan and the vendor's documentation rather than assuming the export is complete.

Linking is the most underrated step. A ticket on its own is a complaint; a ticket joined to the Jira issue, the code review, the release note and the customer's reply is a worked example of how a problem gets solved. Keep native IDs, ticket numbers, job numbers and order numbers through every export so those joins survive.

Steps seven and eight: find sensitive data and decide how to remove it#

Sensitive data in operational records mostly hides in free text, not in labeled fields. Names, phone numbers and addresses appear inside ticket bodies and dispatch notes; passwords and API keys get pasted into tickets and committed to code; HR matters surface in chat channels and email threads.

Automated tools help with the first pass. Presidio, an open-source PII detection SDK now governed by the community under the Data Privacy Stack organization, combines entity recognition, patterns and rules, yet its own documentation warns that there is no guarantee it will find all sensitive information. Secret scanners such as gitleaks and TruffleHog look for passwords, API keys and tokens in git repositories and other sources. Treat their output as a starting point and review samples by hand.

Pseudonymized data may still count as personal data under some privacy laws, so a consistent token is not a shortcut around the privacy review.

Steps seven and eight: find sensitive data and decide how to remove it
MethodWhat it doesWhen it fits
ExcludeDrops a record family, field or channel entirelyHR channels, health details, client deliverables
RedactReplaces a value with a placeholderNames, emails and phone numbers in free text
PseudonymizeReplaces a value with a consistent tokenRecords that must still link by customer or employee
GeneralizeCoarsens a value, such as a street address to a regionLocations and dates where detail adds risk
KeepLeaves the value as writtenPart numbers, error codes and product names you own

Steps nine and ten: documentation and approvals#

Documentation turns a folder of exports into something another team can rely on. A short data dictionary, a provenance note that says which system and export produced each file, and an honest list of known gaps answer most questions before they are asked.

Approvals close the checklist. Identify the authorized signer early, check whether the board, investors or lenders must consent to licensing company records, and record each approval in writing. Internal AI projects need lighter approvals, but a short sign-off from the data owner still helps when an auditor later asks what a model was fed.

Which preparation mistakes cost the most time?#

The most expensive preparation mistake is cleaning records before anyone has checked whether they may be used. Weeks of redaction on a record family that a customer contract or vendor term rules out cannot be recovered.

The other costly mistakes are quieter and show up only when someone tries to use the result.

  • Exporting flat tables: a CSV of ticket fields without internal notes, attachments or the linked issue ID loses the part that explains the work.
  • Trusting one automated pass: detection tools miss names in odd formats, signatures in scanned files and secrets pasted mid-sentence.
  • Preparing everything at once: one well-documented record family is more useful than ten half-finished ones.
  • Losing the export trail: nobody recorded which system, date range or query produced a file, so it cannot be refreshed or defended.
  • Leaving approvals to the end: a lender or investor consent discovered after preparation can delay or narrow the whole project.

Illustrative: a property management software company prepares one archive for two uses#

Illustrative: a fictional property management software vendor wants an internal support assistant and is also curious whether its records could be licensed. Its CTO inventories Zendesk, Jira, GitHub, Confluence and Slack, and names the head of support, the VP of engineering and the controller as owners.

The review finds tenant names and phone numbers in ticket bodies, several API keys pasted into old tickets, and an HR channel in Slack that has to be excluded. Ticket-to-issue links exist for most product bugs, so the team keeps those IDs through the export. Counsel confirms that customer contracts allow internal use but flags a clause on customer communications that requires de-identification before any outside license.

The company prepares one redacted archive with a data dictionary, uses it for the internal assistant, and runs a separate rights review before considering any external license. One preparation effort serves both decisions.

How SourceX approaches preparation for licensing#

SourceX orders the work so that effort follows rights. In the SourceX five-step transaction (Supply, Rights, Preparation, Approval and Delivery), preparation comes after the rights review, so nothing is cleaned that cannot be licensed. Preparation removes personal and confidential details, and the supplier reviews and approves the prepared result before anything is released.

The privacy record in the SourceX Evidence Packet describes what was removed and how, which is the document an outside licensee will ask for. The initial fit check collects metadata only, and nothing is shared during that assessment.

Frequently asked questions

Do we need a data warehouse before preparing data for AI?

No. A warehouse helps with analytics, but preparation for AI can start from exports of the source systems. What matters is that exports keep their IDs and notes, that owners and permissions are documented, and that sensitive details are handled before anything leaves your control.

How clean does data need to be for AI?

Clean enough to be accurate and safe, not polished. AI developers value records that show how work really happened, including messy notes and corrections. Fix broken encodings, duplicates and missing links, remove sensitive details, and leave the substance of the work as it was written.

Who should lead AI data preparation inside the company?

Usually the CTO or head of IT runs the technical work, with a business owner for each record family and counsel for permissions. The CEO or another authorized signer should know the scope early, because any licensing decision ultimately needs their approval.

Can we prepare data for AI without moving it out of our systems?

Partly. Inventory, ownership and permission checks need only metadata. Exports, redaction and review need copies, which can stay in company-controlled storage with access limited to named people until an approved release.

Is preparing data for licensing different from preparing it for our own AI tools?

The steps are the same, but the evidence bar is higher. An outside licensee needs documented rights, a record of how personal details were removed and a signed release. Internal use still needs permission and privacy checks, with lighter paperwork.

Sources

  • Presidio's own documentation warns that "because it is using automated detection mechanisms, there is no guarantee that Presidio will find all sensitive information. Consequently, additional systems and protections should be employed." Source
  • Presidio has moved from a Microsoft-owned project to an independent, community-governed open-source project under the GitHub organization Data Privacy Stack, and remains MIT-licensed. Source
  • Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin. Source
  • TruffleHog, an AGPL-3.0 open-source secret scanner from Truffle Security, scans sources including Git, chats, wikis, logs, object stores and filesystems. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify