AI data market
Is your company data already contaminated by AI-written text?
By SourceX Editorial · Updated
Short answer
Company data already contains AI-generated text wherever staff used drafting assistants, suggested replies, meeting summaries or chatbots. Records created before those tools arrived are cleaner human-generated signal. The practical fix is not a detector: date each tool's rollout, list the systems and fields it touched, and split archives at those dates so buyers know exactly what they are getting.
Key takeaways
- Human-generated signal is a value driver, so buyers want to know which records were written by people.
- Metadata such as author IDs, integration fields and enablement dates is more reliable than AI text detectors.
- Outcome fields like resolution codes, invoices and fix commits stay valuable even after AI tools arrive.
- Bot-authored records are usually excluded or labeled as a separate set.
- Labeling AI-assisted content from now on is cheaper than reconstructing it later.
Why AI-written text matters to buyers#
AI-written text matters to buyers because much of the value of company records comes from the fact that people wrote them while doing real work. Human-generated signal is one of the drivers in the SourceX Enterprise Data Value Framework, alongside domain expertise and data cleanliness.
Model developers are cautious about training on text that earlier models produced. Research on repeated training with model-generated output has raised concerns about narrower, less accurate models over time, so buyers prefer to know where machine-written text sits in a package and how much of it there is.
AI assistance does not make records worthless. A support reply drafted by a tool, edited by an agent and followed by a confirmed resolution still records a real problem and a real outcome. What buyers need is an honest description, not a perfectly clean archive.
Where AI text has probably entered your systems#
AI text usually enters through features that staff switched on to save time. Each system below may contain built-in or add-on AI features, depending on plan, configuration and what your team enabled; check your own admin settings rather than assuming.
| System | Where AI text may appear | What to check |
|---|---|---|
| Help desk, such as Zendesk or Intercom | Suggested replies, ticket summaries, chatbot conversations | Author or agent IDs, bot users, app and integration fields |
| Email and calendar | Drafting assistants and smart replies | Admin settings and add-on install dates; detection in text is weak |
| CRM, such as Salesforce or HubSpot | Call summaries, generated emails and sequences | Activity source fields and integration logs |
| Meeting and call tools | Automated notes, summaries and action items | Which recordings or notes were machine-produced |
| Code hosting and IDEs, such as GitHub | Completion suggestions, generated pull request descriptions | Tool rollout dates by team; bot commits and comments |
| Wikis, such as Confluence or Notion | Drafted pages and summaries | Page history, editor attribution, feature enablement dates |
Build an AI adoption timeline#
An AI adoption timeline is the single most useful document for this problem. It records when each AI feature or tool became available to which teams, and it lets you sort records by date rather than by guesswork.
- List every AI tool and feature in use, including ones individual teams adopted without a formal rollout.
- Find enablement dates in admin audit logs, billing records, procurement tickets and internal announcements.
- Map each tool to the teams and systems it touched; support may have adopted years before engineering, or the reverse.
- Identify bot accounts and integration users that write records directly, and note their user IDs.
- Record policy changes, such as a rule requiring agents to review every suggested reply before sending.
- Mark known gaps where dates cannot be confirmed, rather than filling them with estimates.
Can you detect AI-written text after the fact?#
AI-written text cannot be detected reliably one record at a time. Text classifiers produce false positives on concise, formulaic or non-native human writing, which describes a lot of support and engineering text. Running a detector across an archive and deleting what it flags risks removing good human records while missing machine-written ones.
Metadata is a better tool. Author IDs reveal bot users, integration fields show which app created a record, edit histories show whether a person changed a draft, and timestamps place records before or after a rollout.
Labeling standards are starting to address this for new content. The C2PA specification version 2.4, released in April 2026, added an AI disclosure assertion and support for embedding provenance manifests in HTML documents and structured text formats. Historical business records almost never carry such labels, so the adoption timeline remains the main evidence for older archives.
Splitting archives by date and author type#
Splitting archives turns an uncomfortable unknown into a clear description. A typical split has four segments, each presented differently to a buyer.
| Segment | What it contains | How to present it |
|---|---|---|
| Before any AI tools | Records written entirely by staff and customers | Label as pre-adoption human-generated records with the cut-off date |
| Transition period | Some teams using AI tools, others not | Label by team and date; flag uncertainty openly |
| After adoption, human-reviewed | Drafts or summaries reviewed and sent by a person | Label as possibly AI-assisted; keep edit history where available |
| Bot-authored | Chatbot replies, auto-generated summaries, bot commits | Exclude, or deliver as a separately labeled set if the buyer wants it |
Label AI-assisted content from now on#
Labeling AI-assisted content going forward costs little and saves a painful reconstruction later. Keep machine-generated summaries in their own fields instead of overwriting human notes, give every bot and integration its own user account, and tag tickets or documents where an AI draft was used.
The same labels help internal teams. Quality reviews, training for new agents and internal AI projects all benefit from knowing which text a person actually wrote.
Write the rule into the AI tool policy, not only into system settings. A short policy that names approved tools, says where their output is stored and asks staff to mark AI-drafted text gives every record after the policy date a fixed reference point, which makes the next adoption timeline far easier to build.
Illustrative: a warehouse software company sorting its support archive#
Illustrative: a fictional B2B software company sells warehouse management software to regional distributors. It holds many years of Zendesk tickets linked to Jira issues and GitHub pull requests. Partway through its history, support switched on AI reply suggestions and ticket summaries and added a chatbot for password resets; engineering adopted a code assistant about a year later.
The CTO built an adoption timeline from admin audit logs, billing history and the original rollout announcements. Chatbot tickets were identified by their bot user ID and excluded. AI summaries lived in a separate field and were dropped from the package. Agent replies after the rollout were kept but labeled as possibly AI-assisted, and pull requests were split at the code assistant's rollout date by team.
The outcome was a package that described itself accurately: a long pre-adoption period of human-written replies and reviews, a labeled later period, and resolution fields and linked fixes throughout. The buyer's review focused on content rather than on doubts about where the text came from.
How SourceX handles AI-assisted records#
SourceX treats AI-assisted text as a Preparation question within the SourceX five-step transaction of Supply, Rights, Preparation, Approval and Delivery. The fit check asks when AI tools were adopted in each system, as metadata, before any files are discussed.
The provenance section of the SourceX Evidence Packet records date ranges, segments and how bot-authored or AI-summarized content was handled, so a buyer can rely on a documented description instead of running its own guesswork.
Frequently asked questions
Should we stop using AI tools to protect our data's value?
No. The productivity gain from AI tools is real, and records created with AI assistance still capture real problems and outcomes. Label AI-assisted content and keep bot output separate instead. That protects value without asking staff to give up useful tools.
Is AI-generated code in our repositories a problem for licensing?
It is a disclosure question. Code completion suggestions accepted by engineers are hard to separate line by line, so date tool rollouts by team and describe them. Review comments, issue discussions and design debates remain human records and are often the most valuable part of an engineering package.
Will buyers require disclosure of AI-written content?
Expect the question. Provenance questions usually cover how records were created, and a contract may include representations about the source of the data. Answering from a documented adoption timeline is safer than estimating, and it avoids promises the company cannot support.
Do customer messages count as human-generated if customers used AI to write them?
Possibly not, and you usually cannot tell. Customer-side AI use is outside your control, so describe it as a general limitation rather than trying to classify individual messages. Your own staff's records and the outcome fields carry the clearer signal.
Does AI-assisted text change how records are de-identified?
Somewhat. AI summaries can restate names, account numbers or addresses from the original ticket in new fields, so de-identification has to cover summary fields as well as the original notes. If summary fields are excluded from the package, that extra exposure leaves with them.
Sources
- C2PA specification version 2.4 (April 2026) added an AI Disclosure Assertion (c2pa.ai-disclosure) and support for embedding manifests in HTML documents and structured text formats. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.