Skip to content

Software companies

GitHub organization at shutdown: archive repos, issues and PR reviews

By SourceX Editorial · Updated

Short answer

To archive a GitHub organization at shutdown, mirror-clone every repository and separately export what git does not hold: issues, pull requests, review comments, releases, wikis and project boards. The rule that preserves the most value: pull request review threads live outside git, so export them with their commit, file and line references before the organization is deleted.

Key takeaways

  • A git clone captures commits, branches and tags, but not issues, pull requests, review comments or discussions.
  • Review comments are anchored to commits, files and lines, so export those references along with the text.
  • Marking a repository as archived on GitHub makes it read-only; it is not a backup.
  • Scan the archive for secrets, customer data and third-party code while engineers who know the code are still available.
  • Delete or downgrade the organization only after the archive has been verified against the live site.

What does a GitHub organization archive need to capture?#

A GitHub organization archive needs to capture two different things: the git data in each repository and the collaboration data GitHub stores around it. A mirror clone handles the first. The second, which includes issues, pull requests, review comments and discussions, lives in GitHub's platform rather than in git, so a clone leaves it behind.

At a closing software company, the second category is often the more valuable one. It records why code changed, who questioned it and how problems were resolved, which matters to a future acquirer of the code, to counsel answering ownership questions and to any later licensing review.

Repo vs metadata: where each record lives#

Each record in a GitHub organization lives either in git, in GitHub's platform, or in both, and the capture method follows from that. The table maps the common record types to the export that preserves them.

Repo vs metadata: where each record lives
RecordIn git?How to captureNotes
Commits, branches and tagsYesMirror clone of every repositoryInclude forks and archived repositories
Files tracked with LFSPointers onlyFetch all LFS objects alongside the cloneMissing objects leave broken files
WikiSeparate git repositoryClone each wiki on its ownEasy to miss because it has its own address
Issues and commentsNoAPI export or an organization migration exportKeep labels, milestones, assignees and timestamps
Pull requests and review threadsOnly the resulting commitsAPI export of pull requests, reviews and review commentsKeep commit SHA, file path and line for each comment
Releases and release assetsTags onlyDownload release notes and attached filesAssets are not stored in the repository
ActionsWorkflow files yes; run logs noDownload the logs and artifacts you needLogs and artifacts can expire on their own schedule
Projects and discussionsNoAPI exportBoards link issues across repositories
Teams, members and permissionsNoExport membership and permission listsNeeded later to interpret who reviewed what

Why PR review threads need a separate export#

Pull request review threads need a separate export because GitHub stores them as platform records attached to a diff, not as files in the repository. A merged pull request leaves a merge or squash commit in git, but the reviewer's questions, the requested changes, the author's replies and the approval exist only in GitHub.

Each review comment is anchored to a commit, a file path and a line or range. Export only the text and the conversation loses its meaning: a note warning that a query will lock a table under load is useless without the code it refers to. Keep the commit SHA, path, line, diff hunk where available, reviewer, timestamp and review state for every comment.

Rebases and force pushes complicate the picture. Comments on superseded commits are marked outdated, and those commits may survive only in GitHub's pull request references, so make sure the clone includes those references and the commits stay resolvable in the archive.

Step-by-step: archiving the organization#

Archiving the organization is easier to run as a checklist with one owner than as a shared task in a wind-down channel. Revoke tokens, deploy keys and GitHub Apps last, because the exports usually depend on them.

Plan for API limits. GitHub's REST API allows 5,000 requests per hour with a personal access token, and a GitHub App installation on GitHub Enterprise Cloud gets 15,000, with secondary limits on top. An organization with years of pull requests and review comments can take many hours to export, so start early and make the script resumable. GitHub also documents an organization migration archive generated through the API, which needs owner permissions and a token with the repo and admin:org scopes; it is built for import into GitHub Enterprise Server, so keep a readable export as well.

  • Name an archive owner and confirm who holds organization owner rights through the wind-down.
  • Inventory every repository, including private, archived, forked and template repositories, with size and last activity.
  • Mirror-clone each repository with its pull request references, and fetch all LFS objects.
  • Clone each wiki repository separately.
  • Export issues, pull requests, reviews, review comments, discussions, projects and releases through the API or a migration export.
  • Download release assets and any Actions logs or artifacts worth keeping.
  • Record teams, members, outside collaborators and permissions.
  • Write a readme covering structure, export dates and how issue and pull request numbers map to files.
  • Store the archive encrypted in company-controlled storage with checksums and an access list.

Secrets, customer data and third-party code in the archive#

Secrets, customer data and third-party code are the three things most likely to cause trouble when the archive is later reviewed, shown to an acquirer or assessed for licensing. Find them while the engineers who understand the code are still around.

For secrets, scanners such as gitleaks and TruffleHog search full git history for API keys, passwords and tokens. The gitleaks maintainer has said the tool is feature complete, with future releases limited to security patches. TruffleHog says it classifies over 800 secret types and can check whether a found secret is live by attempting to log in, which calls for care because those checks send real authentication requests.

Customer data hides in test fixtures, seed files, sample exports and debug logs committed by mistake. Third-party code arrives through vendored libraries, copied snippets and contractor work. Record what you find, rotate live credentials and flag the affected files rather than rewriting history inside the archive.

Verify the archive before deleting the organization#

Verification is the step that makes deletion safe. Compare the archive against the live organization while it still exists, and keep the comparison with the archive as part of the company's records.

Verify the archive before deleting the organization
CheckHow to confirm it
Every repository presentRepository list matches the inventory, including archived and forked repositories
Full historyCommit counts and the latest commit per branch match the live repositories
Issues and pull requests completeCounts per repository match, with no gaps in numbering
Review threads intactSampled comments carry commit, path and line references
Large files restoredLFS-tracked files open as real content, not pointer files
Archive usableA second engineer restores one repository and reads its pull requests from the export

Illustrative: a retail analytics software company closes its organization#

Illustrative: a fictional retail analytics software company is closing after a planned sale of its customer contracts falls through. Its GitHub organization holds a monorepo, several service repositories, a set of archived experiments and a documentation wiki.

The CTO mirror-clones everything with LFS objects, then runs a script against the API to export issues, pull requests, reviews and review comments with their commit and line references. A full-history scan finds an old cloud key in a deleted config file and two customer spreadsheets in a test folder; the key is confirmed revoked and the spreadsheets are flagged for exclusion.

A second engineer restores the monorepo from the archive and reads sampled review threads from the export before the organization is downgraded and later deleted. The board now holds a complete engineering record it can preserve, show to a future buyer of the code or assess for licensing.

How SourceX treats code and review archives#

SourceX treats code review threads linked to issues and commits as some of the most informative records a software company holds. The fit check works from metadata, such as the number of repositories, years of history and whether review comments were exported, so a CTO can learn what the archive might support before the organization is deleted.

The archive stays in the company's storage. Packages that proceed go through the SourceX five-step transaction, with secrets, personal details and customer data removed during Preparation and documented in the SourceX Evidence Packet.

Frequently asked questions

Is marking repositories as archived on GitHub the same as backing them up?

No. Archiving a repository on GitHub makes it read-only but leaves it on GitHub, so it disappears if the organization is deleted or access is lost. A backup is a copy the company controls outside GitHub, including the issues, pull requests and review threads the repository page displays.

Should we move repositories to a founder's personal account instead?

Usually not without board approval. Code and its history are company assets, and moving them to a personal account can cloud ownership, a later asset sale or a license. If the organization must stay live, keep it under company control with named owners, or archive it to company storage.

What about copies held by former employees?

Depending on settings, forks of private repositories may be removed when a person loses access, but local clones on laptops are not. Offboarding should include written confirmation that company code was deleted from personal devices, and contractor agreements should be checked for return-of-materials terms.

Do we need GitHub Actions logs in the archive?

Only selectively. Build and deployment logs can support an audit trail or show how releases were produced, but most are noise. Keep logs and artifacts for releases that matter, such as the last version shipped to each customer, and note the retention settings so any gaps are explained.

Can we delete the organization once the archive is verified?

Check obligations first. Customer contracts, source code escrow agreements, open-source commitments and any pending sale may require keeping the code available or depositing it elsewhere. Once counsel confirms nothing requires the live organization, delete it and keep the verification record with the archive. Organization owners can restore individually deleted repositories within 90 days, with limits such as fork networks and lost team permissions, but treat that as a recovery window, not a backup.

Sources

  • On May 21, 2026, the gitleaks README was updated to state that Gitleaks is feature complete, that future releases will be security patches only, and that the maintainer is shifting focus to Betterleaks. Source
  • GitHub's REST API allows 5,000 requests per hour for requests made with a personal access token, while a GitHub App installation on a GitHub Enterprise Cloud organization gets 15,000 requests per hour; secondary rate limits also apply. Source
  • GitHub documents exporting migration data from an organization on GitHub.com through the API into a migration archive for import into GitHub Enterprise Server, requiring owner permissions and an access token with the repo and admin:org scopes. Source
  • GitHub organization owners can restore deleted repositories within 90 days of deletion; repositories in an active fork network cannot be restored directly, and restoring does not restore team permissions. Source
  • TruffleHog, an AGPL-3.0 open-source secret scanner from Truffle Security, says it classifies over 800 secret types and, for each secret it can classify, can log in to confirm whether the secret is live; it scans Git, chats, wikis, logs, object stores and filesystems. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify