Software companies
Decommissioning a self-hosted GitLab or Jira server: archive checklist
By SourceX Editorial · Updated
Short answer
Before you decommission a self-hosted GitLab or Jira server, take an archive that can be restored without the old machine: the native backup, a separate database dump, attachments and uploads, configuration and secrets, and a user mapping. Prove it by restoring to a test instance. A cloud migration is not an archive, so keep both.
Key takeaways
- Moving to cloud tools carries over mapped projects; only an archive preserves everything the old server recorded.
- Attachments, uploads and configuration often sit outside the database, so a database dump alone is incomplete.
- Native backups are often tied to the application version, so keep the installer or image with the archive.
- Check for legal holds and scan for live credentials before any disk is wiped.
Why is a cloud migration not an archive?#
A cloud migration is not an archive because migration tools carry over the projects, issues and repositories that map to the new platform and leave the rest. CI pipeline history and job logs, merge request discussions on deleted branches, custom workflows, plugin data and audit logs are common gaps.
Identities change too. Users who left are mapped to placeholder accounts or merged, which breaks the record of who reviewed what. Inactive projects are often left behind on purpose to keep the migration small, and then the server they live on is switched off.
An old self-hosted server frequently holds the earliest and most detailed engineering history a company has, including decisions made before current processes existed. Once the disks are wiped, that history is gone.
The ten-step archive checklist#
Work through the checklist in order. The first three steps prevent mistakes that cannot be undone, the middle six build the archive, and the last one proves it.
- Freeze: set the instance to read-only or announce a change freeze, and record the exact application and database versions.
- Inventory: list groups, projects and repositories in GitLab, projects and installed apps in Jira, integrations, storage locations and disk mounts.
- Legal hold: ask counsel whether any project, issue or repository is subject to a legal hold or retention duty before anything is deleted.
- Native backup: run the vendor's own backup process and keep its output with the installer or container image needed to restore that version.
- Database dump: take a separate dump of the database so records remain readable without the application.
- Files: copy attachments, uploads, large file storage objects, CI artifacts, packages and wiki files, and checksum every copy.
- Configuration and secrets: copy configuration files and encryption keys separately and store them securely, apart from the archive.
- User mapping: export usernames, emails, display names, groups and status, with a table mapping old accounts to current identities.
- Plain exports: mirror every repository with full history and export issues with comments to JSON or CSV, so the archive is usable without a restore.
- Verify and sign off: restore into an isolated test instance of the same version, compare counts and samples, record the sign-off, then wipe disks under your disposal policy.
Where do GitLab and Jira keep the data teams miss?#
The data teams miss is usually the data outside the main database. Both products store files, configuration and some plugin data in places a database dump never touches.
Check the vendor's documentation for your exact version, because what the native backup includes and excludes has changed across releases. Treat anything the documentation lists as excluded as a separate copy job with its own checksum.
For per-project plain copies, GitLab's file export covers project and wiki repositories, uploads, issues with comments, merge requests, labels, milestones, releases, LFS objects and issue boards, though contents vary by version. Default rate limits allow 6 project exports and 1 export download per minute per user, so script the run and allow time when hundreds of projects are involved.
| Data | GitLab self-managed | Jira Server or Data Center | Archive note |
|---|---|---|---|
| Core records | Database: issues, merge requests, comments | Database: issues, comments, workflows, fields | Dump separately from the native backup |
| Code | Git repositories in repository storage | Linked through development tool integrations | Mirror each repository with full history |
| Files | Uploads, large file objects, CI artifacts and packages, on disk or object storage | Attachments in the Jira home directory | Copy and checksum; confirm object storage buckets |
| Configuration | Configuration files and a separate secrets file | Configuration in the home and install directories | Without keys, some encrypted fields cannot be restored |
| Apps and integrations | Integration settings and webhooks | Marketplace apps, some with their own tables or storage | Check each app's export options |
| Users | User records and directory or single sign-on links | Internal directory or connected directory | Export and map before the directory changes |
How do you keep links between issues, code and reviews?#
Links between issues, code and reviews survive when the identifiers stay stable. Jira issue keys appear in commit messages, branch names and merge request titles, so preserving the old keys keeps years of cross-references readable.
If keys or project names change in the new tools, export a mapping table from old IDs to new ones and store it with the archive. Do not rewrite git history to update references. Keep an index or redirect for the old server's URLs, because links in documentation, tickets and chat history will keep pointing at them.
These linked chains, from issue through code change and review to release, are what make an engineering archive useful for later incident investigations, audits and any AI coding work that depends on real tasks with outcomes.
Secrets, personal data and access to the archive#
An engineering archive concentrates credentials and personal data, so scan it before storing it. Repositories, CI logs and issue comments often contain API keys, passwords and connection strings, and any live secret found should be rotated, not just noted.
Open-source scanners can help. TruffleHog says it can verify detected credentials by attempting to log in with them, so use that feature with care on archived data. Gitleaks is a widely used MIT-licensed scanner, but its maintainer stated in May 2026 that it is feature complete and that future releases will be security patches only, which is worth weighing when choosing a tool.
The user table, commit metadata and comments also contain names and email addresses. Store the archive encrypted, restrict access to a named group, and apply your retention policy to it like any other record system.
Illustrative: a hospitality software company retires its server room#
Illustrative: a fictional hospitality software company is moving to cloud tools and retiring the server room that ran its self-hosted GitLab and Jira Server for most of its history. The migration covers active projects only.
Working through the checklist, the CTO's team finds that Jira attachments sit on a separate network share, that a test management app keeps its data outside the main schema, and that the first restore test fails because the secrets file was not copied. Each gap is fixed and the restore is repeated until counts match.
The company keeps the verified archive in encrypted company storage with a user mapping and an index of projects. Engineers can still trace an old outage back through the issue, the fix and the review that approved it.
How SourceX treats an archived engineering server#
SourceX treats an archived GitLab or Jira server as a common source at the Supply stage of the SourceX five-step transaction. An initial fit check needs only descriptive answers: which products and versions, how many years of history, and whether issues link to code and reviews. Nothing is shared at that point.
If a company proceeds, the archive stays in its own storage, and very large datasets ship on encrypted drives rather than being hosted by SourceX. The restore record, user mapping and scan results feed directly into the provenance and privacy record of the SourceX Evidence Packet.
Frequently asked questions
Can we keep the old server running in read-only mode instead?
You can, but an unmaintained server still needs patches, licenses, backups and hardware, and it becomes riskier as versions age. A restore-tested archive plus plain exports usually gives the same access to history with much less exposure. Keep the server only as long as the verification step needs it.
Do we need the same version to restore a backup?
Often, yes. Native backups for tools like GitLab and Jira are commonly tied to the version that created them. Keep the matching installer or container image with the archive, and record the operating system and database versions as well.
Should we archive CI job logs?
Yes, where you can. CI logs record which tests failed, why and what fixed them, which is valuable engineering history. They also commonly contain credentials and internal hostnames, so scan them before storing.
Does this checklist apply to Bitbucket or Confluence?
The same principles apply: native backup, separate database dump, files outside the database, configuration and keys, user mapping and a restore test. Check each product's documentation for where its files and plugin data live.
How long should we keep the archive?
As long as your retention policy, contracts and legal obligations require. Engineering history often has long-term reference value, but it also holds personal data, so document the decision and review it on a schedule.
Sources
- TruffleHog, an AGPL-3.0 open-source secret scanner, says that for each secret it can classify, it can log in to confirm whether the secret is live, and it scans sources including Git, logs and filesystems. Source
- GitLab's file-export migration includes project and wiki repositories, uploads, issues with comments, merge requests, labels, milestones, releases, LFS objects and issue boards; default rate limits are 6 project exports per minute per user and 1 export download per minute per user. Source
- Gitleaks is an MIT-licensed tool for detecting secrets in git repositories, files and stdin; on May 21, 2026 its README was updated to state that Gitleaks is feature complete and future releases will be security patches only. Source
Related resources
See if your company qualifies
A short company assessment. No data uploads are needed.