Skip to content

Leadership and readiness

How to archive a retired system so its records stay usable

By SourceX Editorial · Updated

Short answer

To archive a retired system so its records stay usable, export the records in open formats such as JSON or CSV with their native IDs, timestamps, thread links, attachments and user mapping intact, then reconcile the archive against the live system before access ends. A PDF-only export keeps a record's appearance but drops the structure that makes it reusable.

Key takeaways

  • An archive is usable when records can be found, read in context and joined to each other without the original application.
  • Native record IDs, timestamps, parent-child links, attachments and a user mapping table are the fields most often lost in a careless export.
  • Export and reconcile while the old system still runs, because access often narrows or ends once a subscription is cancelled.
  • Keep PDFs only as a reading copy next to structured data, never as the archive itself.
  • Decommissioning is the last moment when full access, system knowledge and a funded project line up, so it is the cheapest time to package history for reuse.

What makes an archived system usable later?#

An archived system is usable when someone years from now can find a specific record, read it in context, connect it to related records and trust that nothing is missing, all without logging in to the retired application. Those tests, often summarized as discoverability, trustworthiness and interoperability, apply to an archive as much as to any live dataset.

Most archives pass the first test and fail the other two. A folder of exported files is easy to find, but if ticket comments have lost their parent ticket, or orders have lost their line items, the archive answers only the narrowest questions. Teams usually discover this when a dispute, an audit or a buyer conversation needs the full story of a customer or a job.

Treat the archive as a small product with users you have not met yet: finance tracing an old invoice, product teams studying past defects, counsel answering a claim, or a buyer evaluating a license. Each of them needs structure, not screenshots.

Which fields must survive the export?#

The fields that must survive the export are the ones that hold records together: identifiers, time, relationships, files and people. Losing any of them turns a connected history into loose documents.

Check each row below against a sample export before committing to the full run. Vendors name these fields differently, so map them by meaning rather than by label.

Which fields must survive the export?
FieldWhy it mattersWhat breaks without it
Native record IDs such as ticket, order, job or issue numbersEmails, invoices and other systems reference these numbersCross-references in old emails and reports point nowhere
Created, updated and closed timestamps with time zoneSequence and response times depend on exact timeYou cannot rebuild what happened first or how long it took
Parent-child and thread linksComments, line items and linked issues belong to a parent recordConversations and orders fall apart into fragments
Status and field historyShows decisions and changes, not just the final stateOnly the last status survives and the path to it is lost
Attachments with stable keysPhotos, logs and documents often hold the real evidenceFiles exist but nobody can tell which record they belong to
User and team mappingTurns internal user IDs into names, roles, teams and active datesEvery action is attributed to an unreadable code
Custom field and picklist definitionsExplains internal codes, reason codes and categoriesCoded values become guesswork

Why PDF-only exports are a trap#

PDF-only exports are a trap because a PDF preserves how a record looked on screen while discarding the fields, links and codes underneath. Many helpdesks, CRMs and project tools offer a print or PDF view, and it is tempting to bulk-print everything in the final days before shutoff.

A PDF archive cannot be filtered by date or status, cannot be joined to another system and is slow to search at scale. It also makes later privacy work harder: removing personal details from structured fields is routine, while redacting names across thousands of page images is slow and error-prone.

Check what a vendor's built-in archive tool actually produces before relying on it. Procore's Extracts documentation, for example, says most Procore items download as PDFs, while documents and photos keep their original file type. That is useful as a reading copy, but RFIs, submittals and daily logs exported that way need separate structured exports, such as API or report exports, if you want them filterable later.

Keep PDFs if counsel or auditors want a human-readable copy of key records. Store them beside the structured export and name them by record ID, so the two stay connected.

Step-by-step: archiving before the shutoff date#

Archiving before the shutoff date works best as a short project with a named owner, a fixed scope and a reconciliation step at the end. Start when the retirement decision is made, not when the contract is about to lapse.

The reconciliation step is the one most often skipped. Once the account is closed, a missing attachment folder or a truncated date range usually cannot be recovered.

  • Confirm the contract end date and read the vendor's terms on exports, read-only access and data deletion after cancellation.
  • List every object the system holds, such as tickets, comments, users, macros, custom fields and attachments, with a record count for each.
  • Choose the export route for each object: bulk export, API, database backup or a vendor-assisted export.
  • Export structured records to open formats such as JSON, CSV or Parquet, keeping native IDs and timestamps unchanged.
  • Download attachments as files and store them under keys that match their parent records, rather than relying on links back to the vendor.
  • Export configuration tables: users, teams, custom field definitions, picklists, tags, automations and templates.
  • Reconcile record counts against the live system and spot-check complete threads from several different years.
  • Write a README and data dictionary, set access permissions, and only then approve the shutdown.

How should the archive be stored and documented?#

The archive should sit in company-controlled storage with restricted access, encryption and checksums, and it should be documented well enough that a stranger could use it. A data dictionary written by the people who ran the system is worth more than any storage choice.

A useful reference point is the Data & Trust Alliance's Data Provenance Standards, which group dataset metadata into Source, Provenance and Use. Recording where the records came from, how they were exported and what they may be used for gives the archive the same three layers.

How should the archive be stored and documented?
DocumentWhat it records
READMESystem name and version, export date, who ran the export, date range covered and export method
Data dictionaryEach file and field, internal codes, custom field meanings and known quirks
Reconciliation logRecord counts in the live system versus the archive, with an explanation for any difference
Known gapsRecords purged before export, history lost in earlier migrations and attachments that failed to download
Access and retention noteWho may open the archive, any legal holds and the retention period set by policy

Illustrative: retiring a helpdesk after a platform consolidation#

Illustrative: a fictional vertical software company moves its support team from an older helpdesk to the ticketing module of its new CRM. The team decides not to migrate closed tickets, because the migration tool would assign new IDs and flatten comment threads into a single text field.

Instead, the IT lead exports closed tickets and comments as JSON through the old helpdesk's API, keeps the original ticket numbers, and pulls the agent list, macros, tags and custom field definitions as separate tables. A sample check shows that the export references attachments through links that would stop working once the account closed, so the team downloads every attachment and stores it under its ticket number.

The archive later settles a warranty dispute quickly. When leadership considers licensing support history, the archive already shows ticket-to-engineering links through stored issue keys, and a metadata-only fit check can describe the records without opening a single file.

Why decommissioning is the moment to package records for reuse#

Decommissioning is the moment to package records for reuse because it is the last time full access, people who know the system and a funded project exist together. After shutdown, each of those becomes harder to recover.

If licensing the records is a possibility, add three small things to the archive plan: keep fields that show which customer or contract a record belongs to, keep links to related systems such as an issue tracker or ERP, and note any customer contracts or vendor terms that may restrict reuse. None of this commits the company to anything.

SourceX works from that kind of archive through the SourceX five-step transaction: Supply, Rights, Preparation, Approval and Delivery. The fit check uses metadata only, the supplier approves each step, and large archives stay in the company's own storage or ship on encrypted drives, because SourceX never hosts multi-terabyte datasets. The README and reconciliation log become early inputs to the SourceX Evidence Packet.

Frequently asked questions

How long should we keep an archived system's records?

Retention is a policy decision, not a format decision. Your records owner and counsel set it based on contracts, legal holds, privacy obligations and business need. Write the decided retention period into the archive's access note and schedule a review, so the archive is neither kept forever by default nor deleted by accident.

Is a database backup of the old system enough?

A native backup is worth keeping, but restoring it usually needs the original software, often in a specific version that may not be available later. Pair any backup with an open-format export of the records and a data dictionary, so the history can be read without rebuilding the old application.

Should we migrate closed records into the new system instead?

Migration suits open records that teams still work on. For closed history, migration tools often assign new IDs, drop custom fields and merge comment threads. Many teams migrate active records, archive the rest and keep a mapping table between old and new IDs.

What about personal information in the archive?

The archive carries the same privacy obligations as the live system. Restrict access, keep deletion requests and legal holds in scope, and avoid copying personal details into extra places. If records are later prepared for reuse, personal details are removed from the prepared copy while the original archive stays under its own controls.

Who should own the archive once the system is gone?

Name a business owner, usually the leader of the team that used the system, and an IT custodian who manages storage and access. Without a named owner, archives drift into shared drives where nobody knows what they hold or who may open them.

Sources

  • Procore's Extracts documentation says most Procore items download as PDFs, while documents and photos keep their original file type. Source
  • The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify