Systems and records
How to archive Intercom conversations before switching tools
By SourceX Editorial · Updated
Short answer
To archive Intercom conversations before switching tools, plan on the API for full transcripts and treat report exports as metrics, not history. Capture every message, internal note, tag, attribute, rating and attachment with its conversation ID, reconcile counts against Intercom's reporting, and store the result in formats your company controls.
Key takeaways
- Report exports answer how many and how fast; Intercom's S3 Historical export covers two years, so older transcripts usually need the API.
- Internal notes must be exported and labeled separately from customer-visible replies.
- Download attachments while the workspace is active, because links inside a transcript may stop working later.
- Keep raw JSON per conversation, a flattened message table and a readable transcript for people.
- Reconcile conversation counts by month before cancelling, because gaps cannot be fixed afterward.
What a complete Intercom archive contains#
A complete Intercom archive contains each conversation in full, the people and companies involved, and the context agents added along the way. Missing any of these turns a support history into a pile of fragments that nobody can read with confidence.
Inventory these elements before choosing a route. Workspaces that have run for years often carry custom attributes, retired tags and inboxes from teams that no longer exist, and each one needs a place in the archive.
- Conversations: every message from customers, agents and bots, with authors and timestamps.
- Internal notes and mentions, marked as internal.
- Assignment, state and priority changes, which show how a conversation moved.
- Tags and conversation attributes, including custom attributes.
- Customer satisfaction ratings and remarks.
- Attachments and inline images.
- Contacts and companies, with their IDs and attributes.
- Help center articles and saved replies, which explain the answers agents gave.
Report export or API: what to use for each element#
Intercom offers three routes, and each one serves a different purpose: report exports measure support work, the Data Export feature copies conversation data to your own Amazon S3 bucket, and the API retrieves individual conversations in full. Check Intercom's current documentation for your plan and API version, then assign each element to a route and a completeness check.
Intercom's Data Export to Amazon S3 offers a Historical export covering two years of all Conversations data, Periodic exports (hourly or daily) of new and updated data, and Attachments, in JSON or JSON Lines only. It needs an S3 bucket and a CloudFormation stack that your team sets up. Because the historical export covers two years, a workspace with a longer history usually still needs an API pull for the older conversations.
Writing the check down before the export starts matters more than the route. It turns a vague sense that the export looks right into a test someone can pass or fail.
| Element | Where to look first | How to confirm it is complete |
|---|---|---|
| Conversation metadata | Report or CSV export | Monthly counts match Intercom's reporting |
| Full message text, last two years | Data Export to Amazon S3 (Historical export), or the API | Long threads read end to end against the inbox |
| Full message text, older than two years | API conversation retrieval | Monthly counts for older years match Intercom's reporting |
| Internal notes | API, labeled separately from replies | Notes never appear as customer-facing text |
| Tags and attributes | API or CSV export | Tag totals match the tag report |
| Ratings | Report export or API | Counts of rated conversations match |
| Attachments | Data Export Attachments option, or files referenced in conversations | Sample files open from the archive folder |
| Contacts and companies | Contact export or API | Every conversation resolves to a contact ID |
Step-by-step: building the archive#
Building the archive is a small project with an owner, a developer and a verifier. Start with a scoped sample to test the pipeline, then run the full history period by period.
- List workspaces, inboxes and teams, and note which hold customer conversations.
- Confirm API access, the API version and the date your Intercom access ends.
- If you use Data Export, set up the S3 bucket and run the Historical export with Attachments.
- Pull a sample month, check it against the inbox and fix gaps in the script.
- Pull older history month by month through the API, keeping raw JSON per conversation.
- Download attachments into folders keyed by conversation ID.
- Export contacts, companies, tags, attributes, help articles and saved replies.
- Reconcile counts, sign the export log and lock the archive to read-only.
Gaps that show up during verification#
Verification tends to turn up the same few gaps, and each is easy to fix while the workspace is open and impossible to fix after. Look for them deliberately rather than waiting to stumble on them.
| Gap | Symptom | Fix |
|---|---|---|
| Truncated threads | Long conversations stop mid-exchange in the archive | Page through every message, then re-pull the affected conversations |
| Notes mixed with replies | Internal remarks read as if they were sent to the customer | Re-map message types and re-render the transcripts |
| Missing attachments | Transcripts reference files that are not in the folder | Download from the referenced locations while access lasts |
| Orphaned conversations | Conversations with no matching contact record | Export contacts again, including archived or merged ones |
| Time zone drift | Timestamps disagree with the inbox | Store UTC and record the workspace time zone in the readme |
| Missing inboxes | Monthly counts fall short for some teams | Check the export account's permissions across every inbox |
Formats that stay readable after Intercom is gone#
Readable formats are open, documented and independent of the vendor. Keep three layers: raw JSON for fidelity, a flattened message table for analysis and a rendered transcript for people.
The message table should hold one row per message with the conversation ID, message ID, author type, timestamp, an internal flag and the text. The rendered transcript, HTML or PDF per conversation, lets support or legal staff read a thread without tools. A readme explains fields, IDs, time zones and the export date.
Add a checksum for each file and a manifest listing every conversation ID, so anyone can later show the archive has not changed since export. Keep two copies in separate company-controlled locations.
Privacy and customer terms before any reuse#
Support conversations hold personal information by design: names, emails, phone numbers, order details and sometimes payment or account details that customers paste in. The archive inherits your obligations under your privacy notice, customer contracts and any laws that may apply, so limit access by role from the start.
If the archive is ever considered for reuse, including AI data licensing, de-identification comes first and needs human review. Automated tools help but are not complete: the documentation for Presidio, an open-source de-identification SDK, warns that automated detection cannot guarantee it finds all sensitive information and that additional protections should be used.
Illustrative: a scheduling software company moves helpdesks#
Illustrative: a fictional scheduling software company serving fitness studios is moving from Intercom to a helpdesk bundled with its new CRM. The new vendor offers to import conversations from recent years.
The COO accepts the import for recent history but asks IT for a full archive first. A developer pulls every conversation through the API as raw JSON, downloads attachments by conversation ID and builds a message table with internal notes flagged. The support lead reconciles monthly counts and reads long threads by hand against the inbox.
When a studio owner later disputes what an agent promised, support reads the rendered transcript from the archive. When the company asks whether its support history has outside value, it can describe the archive precisely: years covered, which conversations link to engineering issues and how notes are labeled.
How SourceX assesses archived support conversations#
SourceX assesses support conversations with the SourceX Enterprise Data Value Framework, a SourceX methodology whose drivers include human-generated signal, domain expertise, scale, recency, data cleanliness, rights and AI utility, weighed against preparation cost and privacy burden. Complete transcripts linked to outcomes such as bug fixes or refunds score better on several of those drivers. The first fit check works from a short written description of the archive; no transcripts change hands.
Any package that proceeds follows the SourceX five-step transaction, with Preparation removing personal and confidential details and the company approving every release. The archive stays in company storage until delivery is authorized.
Frequently asked questions
Does Intercom delete conversations when we cancel?
Check your agreement and Intercom's data retention terms rather than assuming either way. Whatever they say, export before cancelling, because access to the workspace and its attachments may end with the subscription and any retrieval after that depends on the vendor.
Should we import old conversations into the new helpdesk?
Import what agents need day to day, often open and recent conversations for active customers. Keep the full history in the archive, keyed by Intercom conversation ID, and store that ID on any imported record so the two can be matched later.
What happens to conversations from deleted contacts?
They may already be partly anonymized or removed, depending on how deletion requests were processed. Do not try to restore removed personal data. Note in the readme that some conversations reflect deletions so later users do not mistake gaps for export errors.
How do we export bot and automated messages?
Treat them like any other message, with the author type preserved. They show which questions automation handled and where it handed off to an agent, which matters for support analysis and for any later review of the archive.
Do we need to archive closed or inactive inboxes?
Yes. Inactive inboxes often hold the oldest history, including conversations about products or teams that no longer exist. Include them in the inventory, confirm the export account can read them and reconcile their counts like any other inbox.
Who should have access to the archive?
Only roles that need it: support leadership, legal and the archive owner in IT. Log access, keep the archive read-only and record the retention rule agreed with counsel, including when personal details should be removed.
Sources
- Intercom's Data Export sends Conversations data to the customer's Amazon S3 bucket, with a Historical export (two years of all Conversations data), Periodic exports (hourly or daily) and Attachments, in JSON or JSON Lines format only. Source
- Presidio's documentation warns that because it uses automated detection mechanisms, there is no guarantee it will find all sensitive information, and additional systems and protections should be employed. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.