Skip to content

Getting started

Why AI data deals are moving from publishers to operating companies

By SourceX Editorial · Updated

Short answer

AI data licensing deals are moving from publishers to operating companies because the first wave licensed published text and media, while developers building agents now need records of how work actually gets done: tickets, orders, jobs and reviews. Those records sit inside private companies, so the next deals depend on documented rights and careful privacy preparation.

Key takeaways

  • The first widely reported AI data deals licensed published content: news, forum posts, stock media and code.
  • Agents that complete business tasks need records that link a request, the decisions made and the outcome.
  • Those workflow records live in private systems such as help desks, field service platforms, ERPs and QMS tools.
  • Public reporting moved from publisher and forum deals in 2023 and 2024 to closed platforms in 2025 and private workplace records by 2026.
  • Operating-company deals turn on customer contracts, vendor terms and privacy preparation rather than copyright alone.

What did the first wave of AI data deals license?#

The first wave of widely reported AI data deals licensed content that was already published: news archives, forum and question-and-answer posts, stock images and video, and code. Those deals drew attention because the rights holders were well known and the content was easy to describe.

Published content suited early buyers for practical reasons. Rights sat with a small number of owners, the material was already organized and searchable, and it covered the broad range of topics that general-purpose language models needed. Many terms were never disclosed, so the public record shows that deals happened more clearly than what was in them.

For an operating company, the lesson from that wave is limited. A distributor or an HVAC contractor does not publish its order exceptions or job histories, so coverage of publisher deals says little about whether those records have a market.

What changed in what AI developers need?#

AI developers now need examples of work being done, not only examples of writing. Models are being built as agents that open tickets, schedule jobs, reconcile orders and draft responses, and training or evaluating them takes records that show a request, the decisions made along the way and the outcome.

Public text rarely contains that chain. A news archive shows how events were reported; it does not show how a dispatcher reassigned a job after a technician called in sick, how a support engineer traced a defect to a specific release, or how a quality manager closed a nonconformance report. Those sequences live in ServiceTitan, Zendesk, Jira, NetSuite and QMS records inside private companies.

That is the core of the shift. When the useful record moves from the public web into company systems, the counterparty moves from publishers to operators.

Deal categories and what drove each#

Each category of AI data deal followed a capability that developers were trying to build. The table summarizes the pattern without dates, because timing and terms varied by deal and many were never made public.

The last row differs from the others in one important way. Workflow records were never meant to be public, so the rights and privacy work falls on the company that holds them, not on a platform's terms of service or a publisher's contributor agreements.

Deal categories and what drove each
CategoryTypical recordsWhat drove demandWho holds the records
News and publishingArticles, archives, reference contentBroad language coverage and current eventsPublishers and media groups
Forums and Q&AThreads, answers, votes, moderation historyConversational, human-written discussionOnline platforms
Stock mediaImages and video with captions and metadataImage, video and multimodal modelsMedia libraries
CodeRepositories, issues, code reviewsCoding assistantsCode platforms and software companies
Enterprise workflow recordsSupport tickets, orders, jobs, RFIs, NCRs, CRM historiesAgents that complete business tasks end to endPrivate operating companies

A dated timeline of the shift, 2023 to 2026#

A dated timeline of public reports shows the shift in stages: publishers and forums licensed archives first, open platforms then tightened access, and by 2026 reporting turned to private workplace records. The entries below are third-party deals and reports, not SourceX transactions, and most financial terms were never disclosed.

Documentation expectations moved in parallel. Datasheets for Datasets was published in Communications of the ACM in December 2021, MLCommons published the Croissant 1.0 dataset metadata format in March 2024, and the Data & Trust Alliance released Data Provenance Standards in July 2024 defining Source, Provenance and Use metadata for AI training datasets. None is a law, but together they show buyers expecting provenance, rights and permitted use to be recorded for each dataset.

  • April 2023 (forums): Reddit announces paid premium access to its Data API for large-scale commercial use, a change its CEO linked to companies training AI models on Reddit data.
  • July 2023 (news): The Associated Press and OpenAI announce a deal under which OpenAI licenses part of AP's text archive; financial terms were not disclosed.
  • December 2023 (news): Axel Springer and OpenAI announce a partnership covering news content from brands including Politico and Business Insider.
  • February 2024 (forums): Reuters reports that Reddit struck a deal to make its content available for training Google's AI models.
  • June 2025 (platforms close off): X updates its Developer Agreement to bar using its API or content to fine-tune or train a foundation or frontier model, as TechCrunch reported.
  • October 2025 (workplace expertise): TechCrunch reports Mercor's CEO describing AI labs hiring former senior staff of banks, consulting firms and law firms because the companies themselves do not want to hand over data.
  • January 2026 (marketplaces): Cloudflare announces it has acquired Human Native, an AI data marketplace where content owners license material for AI training.
  • April 2026 (operating records): Forbes reports that, after cielo24 was wound down, its remaining internal chat, project-tracking tickets and emails became items for sale to AI developers.

Why operating companies hold what developers now want#

Operating companies hold records that are private, written by people doing the work and tied to outcomes, which is the combination public sources lack. In the SourceX Enterprise Data Value Framework those traits map to uniqueness, domain expertise and human-generated signal, which increase value, while reproducibility, the ease of recreating records from public sources, reduces it.

A typical fit is a US company with 50+ full-time employees at peak and several years of operating history, running connected systems such as a CRM, a help desk, a field service platform or an ERP. Software companies, engineering and consulting firms, trades, logistics businesses and manufacturers qualify on their records, not on being in technology.

Why operating companies hold what developers now want
QuestionPublisher dealOperating-company deal
What is licensedPublished articles, posts or mediaInternal workflow records with outcomes
Main rights questionCopyright and contributor termsCustomer contracts, vendor terms and employee notices
Privacy workLimited, since the content was already publicSubstantial: personal and confidential details removed before delivery
Who approvesPublisher leadershipOwner or CEO, with counsel and sometimes investors
VisibilityOften announcedUsually confidential

Illustrative: a property software company checks whether it fits#

Illustrative: the CEO of a fictional property management software company reads coverage of publisher licensing and assumes it has nothing to do with a business that sells to landlords. Its records tell a different story: years of Zendesk tickets linked to Jira issues, GitHub pull requests and release notes, so a customer complaint can be traced to the fix that resolved it.

The company runs a metadata-only fit check, describing systems, years of history and record types without sharing files. The review flags the support-to-fix chain as the strongest package and recommends excluding documents that tenants uploaded to customer portals, which customers control. The CEO approves a rights review of that narrower scope before any export is prepared.

How SourceX approaches the operating-company side#

SourceX works on the operating-company side of this shift, helping established US companies license support conversations, CRM histories, job and dispatch records, orders and exceptions, and quality and maintenance records. Data is licensed, not sold outright, and the company keeps ownership.

Each transaction follows the SourceX five-step transaction of Supply, Rights, Preparation, Approval and Delivery. Each package carries a SourceX Evidence Packet with provenance, licensing rights, permitted use, the privacy record and release authorization, the same categories the documentation standards above ask dataset builders to record.

Frequently asked questions

Are AI companies still licensing publisher content?

Publisher and platform licensing has not stopped; the change described here is additive. As developers build agents for specific business tasks, they look for records that published content does not contain, and those sit with operating companies. Both kinds of deals can continue at the same time.

Why are operating-company deals rarely announced?

Operating companies usually treat a data license as confidential, much like a supplier agreement, and their records involve customers and staff. Competitive sensitivity also makes a public announcement less attractive. Publishers, by contrast, often announced deals because licensing content is part of their public business model.

Does a company have to be in technology to take part?

No. Workflow records from trades, logistics, manufacturing, engineering and consulting firms can be as relevant as software company records. What matters is whether the records are retained, connected to outcomes, mostly in English and covered by rights the company can document.

Will demand grow for every type of business record?

Not evenly. Demand depends on what developers are building at a given time, so some record types draw strong interest while others draw little. A fit check against current buyer requests is more reliable than assuming that all operational data is in demand.

Does licensing to one developer stop us licensing to others?

Only if the license grants exclusivity. A non-exclusive license lets a company license the same records to other buyers, while exclusivity, limited by record type, field of use or time, usually raises the price because the buyer gets something competitors cannot. The choice is made deal by deal.

Sources

  • Datasheets for Datasets by Gebru, Morgenstern, Vecchione, Vaughan, Wallach, Daume III and Crawford was published in Communications of the ACM, vol. 64, no. 12 (December 2021). Source
  • MLCommons Croissant 1.0 was published on 2024-03-01. Source
  • On April 18, 2023, Reddit announced premium paid access for third parties needing large-scale or commercial use of its Data API, a change CEO Steve Huffman linked to companies using Reddit data to train AI models. Source
  • The Associated Press and OpenAI announced on July 13, 2023 a deal under which OpenAI licenses part of AP's text archive; financial terms were not disclosed. Source
  • Axel Springer and OpenAI announced on December 13, 2023 a partnership covering news content from Axel Springer brands including Politico, Business Insider, Bild and Welt. Source
  • Reuters reported on February 22, 2024 that Reddit had struck a deal to make its content available for training Google's AI models. Source
  • In June 2025, X updated its Developer Agreement to bar developers from using the X API or X Content to fine-tune or train a foundation or frontier model, as reported by TechCrunch on June 5, 2025. Source
  • TechCrunch reported on October 29, 2025 that Mercor's CEO described AI labs tapping former senior employees of investment banks, consulting firms and law firms because the companies themselves do not want to hand over data. Source
  • Cloudflare announced in January 2026 that it had acquired Human Native, an AI data marketplace where creators license content for AI training. Source
  • Forbes reported on April 16, 2026 that after cielo24 was wound down, its remaining internal chat, project-tracking tickets and emails became items for sale to AI developers. Source
  • The Data & Trust Alliance Data Provenance Standards (v1.0.0, July 2024) define Source, Provenance and Use metadata needed to enable proper dataset selection for AI model training. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify