Approved workplace email and chat datasets for AI training
A workplace email and chat dataset is a collection of real work conversations from a company's own systems: email threads and chat channels with their reply structure, participant roles, timestamps and attachment text. SourceX sources only exports a data owner has approved, scoped to specific teams, mailboxes or channels and a date range, screened for privileged and confidential material, and de-identified before delivery, including the customers and suppliers who appear in the threads.
Dataset manifest
Sourced to your spec- What it is
- Approved, de-identified exports of team email threads and chat channels
- Typical systems
- Microsoft 365 and Exchange, Gmail, Slack, Microsoft Teams, Google Chat
- Typical history
- Varies by partner, retention policy and the approved date range
- Modality
- Message text plus structured thread, participant and channel metadata
- Delivery formats
- Agreed per order; normalized JSONL threads with attachment text are common
- Preparation
- Names, addresses, phone numbers and signatures de-identified; privileged threads removed
- Licensing
- Limited to the approved teams, channels and dates; permitted use set in the license
- Availability
- Only where a data owner approves a scoped export; not guaranteed
What a delivery contains
Fields vary by source system and are fixed per order. A typical delivery includes:
| Field | Type | What it holds |
|---|---|---|
| thread_id | string | Pseudonymous thread ID, stable across every file in the delivery. |
| source | enum | Email, channel message or thread reply. Direct messages only where explicitly approved. |
| export_scope | object | The approved team, mailbox set or channel the item came from, and the export's date window. |
| channel | object | For chat, a pseudonymous channel ID, its type (public, private or shared with another company) and a purpose label. |
| participants | array | Pseudonymous IDs with job role, department and an internal or external flag. |
| messages | array | Messages in order, each with timestamp, sender, recipients and de-identified text. |
| messages[].reply_to | string | The parent message, rebuilt from email reply headers or chat thread references. |
| messages[].body_text | string | De-identified text with quoted history and signature blocks split out, so earlier messages are not repeated. |
| messages[].mentions | array | People mentioned in chat or newly copied on email, marking where a request changes hands. |
| attachments | array | File type with de-identified extracted text, or a withheld marker and reason code. |
| exclusions | array | Reason codes for messages removed from a thread, such as a privilege screen or an HR matter. |
| redaction | object | Entity types replaced and the placeholder scheme, so consistent pseudonyms can be told apart from blanket redaction. |
| language | string | Detected language, set per message where a thread switches languages. |
Example record
{
"thread_id": "th_2c81e0",
"source": "email",
"export_scope": { "team": "procurement", "unit": "team_mailboxes",
"window": ["2022-01-01", "2023-12-31"] },
"participants": [
{ "id": "p_0412", "role": "buyer", "dept": "procurement", "external": false },
{ "id": "p_1187", "role": "category_manager", "dept": "procurement", "external": false },
{ "id": "p_9d03", "role": "account_manager", "org": "[SUPPLIER_1]", "external": true }
],
"messages": [
{ "id": "m1", "ts": "2023-05-02T08:14:09Z", "from": "p_9d03", "to": ["p_0412"],
"body_text": "Hi [PERSON_1], due to resin costs our list prices rise 9% from [DATE_1].",
"signature": "removed" },
{ "id": "m2", "ts": "2023-05-02T10:31:55Z", "from": "p_0412", "to": ["p_1187"], "reply_to": "m1",
"body_text": "Our agreement caps increases at 4% a year with 60 days' notice. They gave 30. Push back?" },
{ "id": "m4", "ts": "2023-05-03T09:02:40Z", "from": "p_1187", "to": ["p_0412"], "reply_to": "m2",
"body_text": "Yes. Accept 4%, effective 60 days after their notice, and ask for their cost index." },
{ "id": "m5", "ts": "2023-05-03T09:47:13Z", "from": "p_0412", "to": ["p_9d03"],
"cc": ["p_1187"], "reply_to": "m1", "mentions": ["p_1187"],
"body_text": "Thanks for the notice. Under our agreement we can accept 4% from [DATE_2]. Could you share the index behind the increase?" }
],
"attachments": [
{ "message": "m1", "type": "application/pdf", "status": "withheld", "reason": "third_party_price_list" }
],
"exclusions": [ { "message": "m3", "reason": "privilege_screen" } ],
"redaction": { "scheme": "consistent_pseudonyms",
"entity_types": ["PERSON", "ORG", "EMAIL", "PHONE", "DATE"] },
"language": "en"
}Synthetic record for illustration. Field names, structure and format are agreed per order.
What AI teams use it for
Draft replies that fit the thread
Real threads show how people reply given everything above them: who is copied, what was already agreed, and when a reply belongs with a colleague instead.
Learn handoffs and escalation
Forwards, new recipients and mentions mark the moment a request moves to someone with authority, a routing call an enterprise assistant has to make on its own.
Summarize threads and extract commitments
Multi-party threads that end in a decision or a date are hard summarization tests, with the deciding message as ground truth.
Evaluate enterprise search and RAG
Questions asked and answered inside a channel become retrieval test items whose correct source is a known message.
Triage and prioritize inboxes
Response times, who replied and which threads were escalated give weak labels for urgency and ownership without hand annotation.
Use-case guides: Enterprise and computer-use agents, Sales agents
What makes this data valuable
Intact reply structure
Reply chains rebuilt from headers and thread references, so each message keeps its parent.
Roles without identities
Role, team and internal or external status survive after names are removed.
Consistent pseudonyms
One person maps to one placeholder across threads, so working relationships stay visible.
Decisions in the text
Threads ending in an approval, a refusal or a committed date carry outcomes without annotation.
Deduplicated content
Quoted history and copies of one message held in several mailboxes are counted once.
Linked artifacts
Attachments and references to tickets or documents tie a conversation to the work it was about.
Why the scope is narrow by design
An approved export is a deliberately small slice of a company's communication systems, because those systems hold everything at once: customer escalations, supplier disputes, HR cases, legal advice and personal conversations. What can be released is defined by teams, mailboxes or channels, a date window and a list of excluded categories, and each of those choices shapes the data you receive.
The resulting distribution leans toward functions whose daily work is routine enough to release, such as operations, procurement, implementation and support. The conversations around the most sensitive decisions are the least likely to be present, so a model trained only on approved exports sees fewer high-stakes judgment calls than real work contains. Threads also arrive with gaps where messages were withheld. Keep the exclusion markers rather than stitching the remaining messages together: a model trained on stitched threads learns replies to messages that, in its training data, were never sent.
What survives de-identification
Email and chat are harder to de-identify than structured records, because the text was written for people who already know each other. Nicknames, initials and in-jokes can identify someone without containing a name, and a project codename, a customer's industry and a date can be enough to identify a deal. Typed, consistent placeholders keep more signal than blanket redaction: "[PERSON_3] approved it" still tells a model that the same person approved last week's request, while "[REDACTED] approved it" does not.
Dates need an explicit rule. Shifting every date in a thread by a fixed offset reduces linkage risk but breaks references such as "end of quarter"; keeping real dates preserves temporal reasoning but raises re-identification risk. Agree the rule per order and apply it to attachments as well as message text.
Chat adds problems of its own. Messages are short and lean on shared context, threads branch, and much of the meaning sits in reactions and in links to other systems. A channel export without the linked tickets or documents suits conversation modeling but is weak for task modeling, so scope a chat request around what the conversations are about, not only around channel names.
What to check before licensing
- Ask for the export approval itself: who signed it, which teams, mailboxes, channels and dates it covers, and what it excludes. Compare it with the sample.
- Ask how employees were informed. In some countries, works councils or other employee representatives have information, consultation or consent rights over some uses of employee data.
- Find out how privilege was screened: which counsel identities and legal terms were filtered, and who reviewed borderline threads.
- Test the de-identification yourself: search the sample for signature blocks, quoted replies, display names in email headers, out-of-office notices and text inside attachments.
- Ask whether threads covered by customer or supplier confidentiality terms were excluded or masked, since those terms bind the partner, not you.
- Check how channels shared with other companies were handled, since another organization's staff wrote part of them.
- Confirm that direct messages, HR cases, health details and credentials pasted into chat are excluded, and ask how each was detected.
How licensing works through SourceX
- 1
Define
Send the domain, modality, volume, format, timeline and permitted use you need.
- 2
Source
SourceX identifies businesses that hold matching data and are open to licensing it.
- 3
Qualify
Fit, rights and quality are checked, and you review samples before committing.
- 4
License
Scope, permitted use, exclusivity, price and obligations are agreed in writing.
- 5
Deliver
Approved data is prepared, de-identified where required and transferred securely.
Questions buyers ask
Can company email and Slack data be licensed for AI training?
Yes, as an approved export with a narrow scope. The company that owns the systems decides which teams, mailboxes or channels and which dates can be released, reviews employee notice, privilege and confidentiality, and approves delivery. Communications are among the most sensitive data a company holds, so expect them to be harder to source than operational records such as tickets.
What is usually left out of an approved export?
Anything the owner cannot clear. Typical exclusions are direct messages, executive and HR mailboxes, conversations with lawyers, health and personal matters, threads under confidentiality terms the owner cannot waive, and credentials pasted into chat. Contact lists and address books are out of scope. Exclusions can be marked with reason codes, so you can see where a thread lost a message.
How is attorney-client privilege handled in email exports?
Privileged material is screened out before delivery rather than redacted in place. Screens usually combine filters on counsel identities and legal terms with manual review of borderline threads. Disclosing privileged communications to a third party can waive privilege, so partners screen conservatively and legal topics end up underrepresented.
What happens to customer and supplier personal data in the threads?
It is de-identified along with employee data. External correspondents become pseudonymous participants flagged as external, names and contact details in the text are replaced, signatures are stripped, and attachments are converted to de-identified text or withheld. Organization names are masked too, because they can reveal people and confidential relationships.
Do employees know their messages are being licensed?
That depends on the partner, and it is part of the rights review. Ask how employees were informed and, where employee representatives have rights over the processing of employee data, how those rights were respected. If a partner cannot show this, treat it as a blocking issue, however clean the export looks.
How does a licensed export compare with the Enron email corpus?
A licensed export can be recent and unpublished; the Enron corpus is neither. Enron's email was released by the Federal Energy Regulatory Commission during its investigation, was written more than two decades ago, comes mostly from senior managers at one company, and sits in open pretraining collections such as The Pile, so models may have seen it. A licensed export can also include chat.
Can I get email and chat from the same team, linked to its other work?
Sometimes. When a partner approves both, a team's mailboxes and channels can share one pseudonym scheme, and references to tickets or documents can link to other data from the same partner, turning conversations into context for workflow data. Each added system brings its own rights and de-identification review.
Related datasets
- Enterprise document archives
A company's working files with folders, versions and sharing metadata
- Enterprise workflow and task execution histories
Linked task trajectories from request to outcome, across every tool the work touched
- Sales CRM pipeline histories
Opportunity histories with stage changes, logged activities and won or lost outcomes
- SOPs, playbooks and internal knowledge bases
Written procedures with page history, ownership and links to execution records
Evaluating this data for procurement?
Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.
Request dataset diligenceTell us what your models need
Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.
Updated 3 October 2026.