AI data market
AI data licensing trends in 2026 that operating companies should know
By SourceX Editorial · Updated
Short answer
The AI data licensing trends that matter most for operating companies in 2026 are demand for records of real work, agent environments built from company archives, closed vendor API routes and rising documentation duties. For a 50 to 500 person company, start with records that link a request to a decision and an outcome, with clear rights.
Key takeaways
- Buyers increasingly want records that show how work gets done, not more public web text.
- Company archives are being turned into practice environments for AI agents, so permitted use is now a key license term.
- Platform terms such as Slack's bar third-party apps from using API data for model training, so licensing runs through the record owner's own exports.
- Provenance standards and transparency laws lead buyers to ask suppliers for system names, date ranges, consents and preparation notes.
- Prices remain private and deal by deal; headline deals say little about what a mid-sized company's records are worth.
Which AI data licensing trends matter in 2026?#
The AI data licensing trends that matter for operating companies are the ones that change what buyers ask for and what suppliers must document. Much public coverage of AI data deals has centered on publishers and media archives, yet the shifts below reach logistics firms, contractors, manufacturers and software companies just as directly.
The table pairs each trend with a one-line meaning for a company of 50 to 500 employees. The sections that follow explain the trends that carry the most practical work.
| Trend | What it means for a mid-sized operating company |
|---|---|
| Demand is moving toward records of real work | Start your inventory with records that link a request to a decision and an outcome: tickets, jobs, orders, code reviews. |
| Agent environments are being built from company records | Linked task-and-outcome histories can serve evaluation as well as training, so define permitted use precisely. |
| Closed and bankrupt companies' archives have entered the market | Value survived only where records were preserved; a running company keeps far more control over scope. |
| Major software vendors are restricting API routes to training data | Licensing runs through exports the record owner controls, not through third-party apps pulling data. |
| Provenance documentation is becoming standard | Keep system names, date ranges, consents and preparation notes for every package. |
| Transparency laws push questions down to suppliers | Decide early how your company may be named or described in a developer's training data summary. |
| Privacy tooling is maturing but carries no guarantee | Plan for human review on top of automated redaction, and keep a written privacy record. |
Why buyers want records of real work#
Buyers want records of real work because AI systems are being asked to act inside businesses, not just answer questions. A model that will triage service calls or resolve order exceptions needs examples of how people made those decisions, with the context they had and the result that followed.
Public text is thin on that internal reasoning, and its supply is finite: Epoch AI has projected that language models will fully use the stock of public human-generated text sometime between 2026 and 2032. Developers have looked elsewhere. Wired reported in January 2026 that OpenAI and Handshake AI asked contractors to upload real work from past and current jobs, such as documents, spreadsheets and code, and reportedly told them to remove proprietary and personal details first.
For an operating company, that cuts two ways. Records of real work have value, and some of that work may already be leaving through former staff, so check confidentiality agreements. A ServiceTitan job that runs from inquiry to estimate, dispatch, invoice and callback, or a NetSuite order exception with its substitution decision and credit memo, shows steps a general corpus rarely contains. As more workplace text is drafted with AI assistance, buyers may also ask when such tools entered your workflows.
Why are agent environments and closed-company archives in the news?#
Agent environments and closed-company archives are in the news because developers need realistic practice settings for AI agents, and archives of whole companies supply them. TechCrunch reported in September 2025 that leading AI labs wanted more reinforcement learning environments, simulated workspaces where agents train on multistep tasks, and Forbes reported in April 2026 that records bought from defunct companies were being fed into such gyms.
Supply followed. Forbes described cielo24, a closed company whose internal chat, project-tracking tickets and emails became items for sale to AI developers after a wind-down. In August 2026, SiliconANGLE reported that Google agreed to buy Spirit Airlines' internal business data in a bankruptcy auction, and other coverage said passenger profiles, loyalty records and privileged legal material were excluded.
For an operating company the lesson is about control, not urgency. A wind-down often sells under time pressure with no one left to review channels, while a running company can choose record families, exclude sensitive material and state whether records may be used for training, evaluation or environment building.
Why are major software vendors restricting API routes to training data?#
Major software vendors are restricting the API routes that third-party apps could use to pull workplace data for model training, which pushes licensing back to the company that owns the records. Slack updated its App Developer Policy in December 2024 to make explicit that using data to train an LLM is prohibited, and a May 2025 change to its API terms, as summarized by the law firm Hunton Andrews Kurth, barred bulk export through the APIs, persistent archives and use in large language models.
HubSpot's updated developer terms similarly restrict using data accessed through its APIs to train or improve AI models, with a carve-out for single-customer use. These terms govern apps and integrations, not a company's own admin exports. The practical meaning: a buyer cannot simply connect an app and pull your history, and a scoped export that your admins run and review is the route to a license.
Why is documentation becoming part of the deal?#
Documentation is becoming part of every data license because buyers must explain where their training data came from. One example is the Data & Trust Alliance's Data Provenance Standards, which sort dataset metadata into three groups, Source, Provenance and Use, and present them as what a developer needs to choose training datasets properly.
Other standards point the same way. MLCommons published version 1.1 of its Croissant dataset metadata format on January 29, 2026, and C2PA specification 2.4, released in April 2026, added an AI disclosure assertion. Suppliers do not need to adopt these formats, but they should expect buyers to ask for the facts they describe.
Keep these facts for each package from the start:
- Each source system and record family, and the years it spans.
- Who created the records and under which employee, customer and vendor terms.
- Consents, notices and contract clauses that affect permitted use.
- Preparation steps applied, including de-identification and exclusions.
- Who approved release, and when.
How do transparency rules reach suppliers?#
Transparency rules reach suppliers indirectly, through the developers who license their records. California AB 2013, signed in September 2024, required developers of generative AI systems released since January 1, 2022 to post training data documentation by January 1, 2026, including whether datasets were purchased or licensed and whether they include personal information.
In the EU, Article 53(1)(d) of the AI Act requires providers of general-purpose AI models to publish a sufficiently detailed summary of training content, following a template the European Commission released on July 24, 2025. Neither rule asks a supplier to publish anything, but both lead buyers to ask how a dataset may be described and whether the supplier may be named, so settle description and naming in the license.
Privacy preparation: better tools, same responsibility#
Privacy preparation tools kept changing in 2026, but responsibility for what leaves the company did not move. Presidio, an open-source tool for finding and removing personal details in text, moved from Microsoft ownership to independent community governance under the Data Privacy Stack organization, recorded in a release dated June 28, 2026.
The project is candid about limits: because detection is automated, its documentation says it cannot promise to catch every sensitive detail and recommends layering other protections on top. For an operating company, that means sampling prepared records, reviewing free-text fields by hand and keeping a privacy record of what was removed.
Illustrative: a contract manufacturer reads the trends#
Illustrative: the CEO of a fictional contract manufacturer reads 2026 market coverage and asks whether any of it applies to a company with an ERP, an MES and a QMS. The quality manager points to nonconformance reports and CAPAs that link each defect to a root cause, a disposition and a corrective action.
The team describes its systems, years of history and record families in a metadata-only fit check. Customer-owned drawings and export-controlled jobs are excluded at the start, and the CFO asks that any developer summary describe the records as licensed quality records from a US manufacturer without naming the company. Outcome: a defined scope, a documentation file in place and no data moved before the decision to proceed.
How SourceX reads these trends#
SourceX weighs these trends against the SourceX Enterprise Data Value Framework, its own qualitative methodology, which publishes no prices or index values. Several trends map directly onto its drivers: demand for records of real work raises the weight of human-generated signal, domain expertise and AI utility; documentation expectations sit under rights; and privacy tooling changes preparation cost and privacy burden, which reduce net value. Uniqueness, scale, recency and data cleanliness still add value, exclusivity still affects price and reproducibility still reduces value.
What has not changed is the transaction itself. Records are licensed, not sold outright; the company keeps ownership; and each deal moves through the SourceX five-step transaction of Supply, Rights, Preparation, Approval and Delivery, with the supplier approving every step and no files shared during the initial assessment.
Frequently asked questions
Are AI developers still paying for business data in 2026?
Public reporting in 2026 described developers buying workplace records from closed and bankrupt companies, and demand continues for linked operating records that show decisions and outcomes. Prices are negotiated privately, deal by deal, so no published figure tells a company what its own records would earn.
Do headline deals between AI labs and publishers tell us what our data is worth?
Rarely. Publisher deals usually involve large public archives, brand value and display rights that operating records do not carry. A mid-sized company's records are valued on different drivers, such as uniqueness, linkage, rights and preparation cost, and the terms of most deals are private.
Do these trends apply to companies outside software?
Yes. Home services contractors, distributors, 3PLs, engineering firms and manufacturers hold job, order, exception and quality records that show real decisions. Software companies have an early advantage because their records are already structured, but the same trends shape demand across operating industries.
Is it better to wait until the market matures?
Waiting has costs that are easy to miss. Systems get retired, retention settings delete older records and the people who understand an archive move on. A metadata-only inventory costs little and keeps options open, whatever happens to buyer demand in later years.
What is a sensible first step for a mid-sized company?
List the systems that hold support, sales, project, job or quality records, note roughly how many years each can still export, and flag known restrictions such as customer contracts. That list settles most fit questions while every file stays where it is.
Sources
- The Data & Trust Alliance's Data Provenance Standards define dataset metadata in three groups, Source, Provenance and Use, which the specification says is needed to enable proper dataset selection for AI model training. Source
- MLCommons' Croissant metadata format for machine-learning datasets published version 1.0 on 2024-03-01 and version 1.1 on 2026-01-29. Source
- C2PA specification version 2.4 (April 2026) added an AI Disclosure Assertion, among other additions. Source
- Presidio moved from a Microsoft-owned project to an independent, community-governed project under the Data Privacy Stack GitHub organization, recorded in release 2.2.363 dated 2026-06-28. Source
- Presidio's documentation warns that automated detection gives no guarantee of finding all sensitive information and that additional systems and protections should be employed. Source
- Epoch AI projects that language models will fully utilize the stock of public human-generated text between 2026 and 2032, or earlier if models are intensely overtrained. Source
- Wired reported in January 2026 that OpenAI and Handshake AI asked third-party contractors to upload real work from past and current jobs, such as documents, spreadsheets and code repositories. Source
- Per reporting on the OpenAI/Handshake request, contractors were told to remove proprietary or personal information before uploading. Source
- TechCrunch reported on September 16, 2025 that leading AI labs are demanding more reinforcement-learning environments, simulated workspaces where agents train on multistep tasks. Source
- Forbes reported on April 16, 2026 that cielo24's remaining records (internal chat, project-tracking tickets and emails) became items for sale to AI developers after a wind-down, and that records bought from defunct companies are fed into reinforcement learning gyms. Source
- SiliconANGLE reported on August 17, 2026 that Google agreed to buy Spirit Airlines' internal business data in a bankruptcy auction, according to filings in the U.S. Bankruptcy Court for the Southern District of New York. Source
- Reporting on the Spirit Airlines sale states that it excludes passenger profiles, Free Spirit loyalty records and privileged legal materials. Source
- In a changelog entry dated December 10, 2024, Slack said it was updating its App Developer Policy, making explicit that the use of data to train an LLM is prohibited. Source
- According to Hunton Andrews Kurth, the Slack API Terms of Service were modified effective May 29, 2025 to prohibit bulk exporting of data accessible through Slack's APIs, persistent copies or archives, and use of such data in large language models. Source
- HubSpot's updated Developer Terms restrict using data accessed through HubSpot APIs to train, fine-tune or improve AI or machine learning models, with a carve-out for legitimate single-customer use cases. Source
- California AB 2013, signed September 28, 2024, requires developers of generative AI systems released on or after January 1, 2022 to post training-data documentation by January 1, 2026, stating whether datasets were purchased or licensed and whether they include personal information. Source
- Article 53(1)(d) of the EU AI Act requires providers of general-purpose AI models to make publicly available a sufficiently detailed summary of training content following an AI Office template, which the European Commission published on July 24, 2025. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.