Skip to content

Procurement, samples and ongoing supply

Sourcing Data Directly from Operating Companies: A Buyer's Playbook

Quick answer

To buy data directly from companies for AI training, identify businesses by the system of record that holds the workflow you need (Zendesk queues, Jira projects, NetSuite ledgers, contract repositories), not by industry alone. Then approach an executive sponsor with a narrow, de-identified request and expect legal, privacy and security review on their side. Budget for confidentiality carve-outs that shrink usable volume. Close under a license that names records, uses, term and delivery. Direct deals reach data nobody sells, but cycle times are long.

By SourceX Editorial · Updated

Why operating-company data needs a different procurement motion

Operating companies are not data vendors, so the deal runs through people who have never sold a record. A company that runs a support desk, a field-service operation or an accounts-payable team holds operational history as a by-product. Usually no one there owns "data sales," has a price list, or has a standard license. Your counterpart is often a general counsel, a CISO and a COO who each have veto power and no upside quota.

This differs from buying from a catalog vendor or a collection shop.

Market commentary describes bilateral licensing deals as expensive to negotiate and slow to scale [1]. For the economics of doing this yourself versus through a managed channel, see managed data sourcing vs an in-house function. For the broader buying lifecycle, start at the AI training data procurement hub.

Identify holders by system of record, not industry

The fastest way to find a holder is to name the system and workflow that produces the records your model needs. "Insurance data" is not a target. "Closed first-notice-of-loss claims with adjuster notes and final reserve, from a Guidewire ClaimCenter instance, 2019 to 2025" is a target, and it points to a specific kind of company and a specific internal owner.

Map each training objective to a workflow, a system and an internal owner:

  • Agent training on ticket resolution: Zendesk, Salesforce Service Cloud or Freshdesk exports with ticket threads, macros used, status transitions and resolution codes. Owner: VP Support or CX operations.
  • Code and engineering agents: Jira or Linear issues linked to Git commits, pull request reviews and CI results. Owner: VP Engineering, with security sign-off.
  • Finance workflow models: NetSuite, SAP or QuickBooks journal entries, AP invoice images with approval chains, reconciliation exceptions. Owner: Controller or CFO.
  • Legal document models: contract repositories (Ironclad, iManage, SharePoint) with redline histories and clause metadata. Owner: General Counsel, who will screen out privileged material.
  • Sales and voice data: Gong or call-center recordings with CRM outcomes. Owner: RevOps, with privacy review because voiceprints can be biometric identifiers under statutes such as Illinois BIPA [4].

Size the target list for attrition. Many holders run multi-tenant SaaS, so the data sits with them, but their customer contracts and privacy notices decide what they may license. Holders also differ in how clean their exports are: a company still on a self-hosted ticketing system may have richer history but harder extraction.

What a holder must clear internally before saying yes

A supplier's yes is the end of an internal approval chain, and each link can kill or reshape the deal. Expect these reviews, roughly in this order:

  1. Commercial and strategic: does licensing expose competitive information, such as pricing, customer lists or process know-how, to a lab that might later serve competitors?
  2. Contractual rights: do customer contracts, MSAs or DPAs restrict secondary use? B2B holders frequently process data on behalf of their own clients and often cannot license it without those clients' permission.
  3. Privacy promises: FTC staff warned in a January 2024 Office of Technology post that companies may be liable under laws it enforces when they break privacy or confidentiality commitments, including by using customer data for undisclosed purposes such as training [2]. A holder's published privacy notice is therefore a hard boundary on what it can release.
  4. Sector law: health records need HIPAA de-identification by Safe Harbor or Expert Determination [3]; biometric data such as voiceprints and face geometry carries consent and retention duties under statutes like BIPA [4].
  5. Security: the CISO wants to know how extraction happens, who touches raw data, and how delivery is controlled.
  6. Ownership of content: a holder's own written work, such as editorial headnotes, knowledge-base articles or annotated templates, can be copyrightable. The Third Circuit affirmed in September 2026 that Westlaw headnotes were protectable and that their use for a non-generative tool was not fair use [6].

Bring answers to these questions in your first package. A request that already specifies de-identification, access-controlled delivery and narrow allowed uses gives internal reviewers something they can approve.

Confidentiality carve-outs shrink usable volume

Plan on receiving materially less than the holder's raw record count. Typical carve-outs remove client-confidential records, privileged legal material, records under litigation hold, data subject to contractual no-reuse clauses, regulated categories the holder will not de-identify, and anything touching minors. A ten-year ticket archive can lose whole customer segments because those customers' MSAs prohibit reuse.

De-identification also changes what survives. Removing names, emails, phone numbers, account numbers and free-text mentions of them can break entity links across a workflow, so a ticket and its linked refund no longer join. NIST notes that traditional de-identification has inherent limitations compared with formal privacy methods [5]. Specify which join keys must survive as consistent pseudonyms before extraction, not after delivery.

Price and size against post-carve-out volume. Use sizing a data purchase and estimating a dataset's value before purchase to decide whether the residual is still worth the effort.

Outreach that gets past the first meeting

Effective outreach leads with a bounded, low-risk request rather than a partnership pitch. Executives respond to specifics: which records, which fields removed, which uses, how delivery works, and what their company is not being asked to do.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample entry in a first-contact data request
RecordsClosed B2B support tickets, Jan 2021 to Dec 2025, English only
Source systemZendesk export (tickets, comments, audits, ticket_fields)
Required fieldsThread text, status transitions, tags, priority, resolution code, time-to-resolve
Removed or replacedRequester names, emails, phone numbers, account numbers, URLs with tokens, attachments
Consistent pseudonymsOrganization ID and agent ID, so threads and escalations still join
ExcludedTickets from customers whose contracts bar reuse; legal-hold tickets; minors
Allowed useSupervised fine-tuning and held-out evaluation for a support agent; no redistribution
Not requestedCustomer lists, pricing, source code, raw recordings
DeliveryAccess-controlled transfer after signature; no email attachments
Sample200 de-identified tickets for schema and quality review under NDA

Send this as a one-page brief to the business owner, with a separate data-handling annex for counsel and security. Keep your own roadmap confidential where you can; confidential data sourcing covers how much to disclose about intended models.

Terms that decide whether a direct deal closes

Direct deals close when the license is narrow enough for the holder and specific enough for your model team. The points that most often stall negotiation:

  • Use scope: training, fine-tuning, evaluation and retrieval are different rights. Licensors increasingly separate training rights from grounding rights and price them differently [9], so state whether records will sit in a RAG index or only in training runs.
  • Disclosure obligations: your public documentation may need to describe sources. As of October 2026, California AB 2013 has required generative AI developers to post training-data documentation since January 1, 2026 [7], and the EU AI Office template for the Article 53(1)(d) training-content summary was published on 24 July 2025 [8]. Holders often want to know whether they will be named.
  • Competitive restrictions: holders may ask that their data not be used to build products for direct competitors. Define "competitor" by product, not by the holder's market narrative.
  • Refresh: a recurring feed (monthly ticket drops, quarterly ledgers) needs a defined schema, change notice and quality checks per drop.
  • Remedies: agree in advance what happens if a delivery has residual personal data or broken fields; see remedies when a data delivery fails and acceptance criteria for licensed training data.

Run your standard data provider due diligence questionnaire, adapted for a holder that has never answered one. Expect to help them fill it in.

When to use an intermediary instead

An intermediary makes sense when you need many holders, when the holder lacks capacity to prepare data, or when rights review and de-identification would otherwise sit on your critical path. A lab with a large partnerships team may run two or three strategic bilateral deals itself and route long-tail categories through a managed channel.

SourceX is one such channel for US operational data: it sources datasets from US companies on request and manages the commercial process, including licensing agreements and ongoing purchases. Buyers describe the data they need, not the businesses; SourceX looks for US companies that hold it, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails, phones and account numbers are removed or replaced before delivery with the method recorded and a sample checked, though no method is perfect. You can submit a data request to SourceX.

The fit is narrower than all direct sourcing. SourceX does not hold inventory, a request does not guarantee a match, and it does not source scraped web content, standalone contact lists or generic CCTV or photos. Compare channels in the build, buy or synthesize framework and the supplier-side view of data partnerships for AI.

Direct-deal readiness checklist

Use this list before the first outreach call to a holder:

  • Training objective mapped to a workflow, source system and internal owner
  • Field-level request with required fields, removed fields and pseudonymized join keys
  • Allowed-use statement separating training, evaluation and retrieval
  • Expected carve-outs listed, with a post-carve-out volume estimate
  • Sector-law screen (HIPAA, BIPA, contractual no-reuse) completed
  • Security annex covering extraction, access control and delivery
  • Disclosure position on AB 2013 and EU training-content summaries agreed internally
  • Sample review plan, acceptance criteria and remedies drafted

For how labs structure these purchases end to end, see how AI labs purchase training data, and for the broader catalog of buyer guides, the AI data hub.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Source operational data from US companies

If the records you need sit inside US businesses that have never sold them, describe the data, fields and allowed uses, and SourceX will look for companies that hold it. Nothing is contracted until a supplier agrees, and every dataset is delivered under a license defining records, uses, term and delivery. Describe the company data you need.

Sources

  1. Bria, "Content licensing radar". https://bria.ai/content-licensing-radar
  2. Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  3. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
  4. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  5. National Institute of Standards and Technology, "De-Identifying Government Datasets: Techniques and Governance (NIST SP 800-188)" (2023). https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-188.pdf
  6. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  7. California Legislature, "AB-2013 Generative artificial intelligence: training data transparency" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  8. European Commission (AI Office), "Explanatory Notice and Template for the Public Summary of Training Content for general-purpose AI models" (2025). https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training-content-general-purpose-ai-models
  9. Digiday, "WTF is AI 'grounding' licensing, and why do publishers say it matters over training deals?". https://digiday.com/media/wtf-is-ai-grounding-licensing-and-why-do-publishers-say-it-matters-over-training-deals/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data