Skip to content

Software companies

Data inventory for software companies: Jira, GitHub, Zendesk and Slack

By SourceX Editorial · Updated

Short answer

A data inventory for a software company lists each system that holds records, the record families inside it, the years still accessible, the export route and the join keys that connect systems. The join keys matter most: a ticket linked to a Jira issue, a commit and a reviewed pull request says far more than four separate archives.

Key takeaways

  • Build the first inventory at the metadata level; no exports are needed to map systems and history.
  • Record the join key for every system: ticket ID, issue key, commit SHA, pull request number and thread timestamp.
  • Linkage between support, issues and code is the strongest signal that engineering records are worth a closer look.
  • Flag sensitive areas such as security projects, HR channels and customer-owned code while you map, not afterward.

What should a software company data inventory capture?#

A software company data inventory captures, for each system, what records it holds, how far back they go, how to get them out, how they connect to other systems and what restrictions apply. It is a map, not a copy: the first version is built from admin consoles and vendor documentation without exporting anything.

Use the same columns for every system so gaps stand out. A blank cell for join key or export route is often the most useful finding in an early inventory, because it shows where history is isolated or hard to retrieve.

  • System and owner: the tool, the plan, and the admin who can answer questions.
  • Record families: issues, comments, commits, reviews, tickets, threads, pages.
  • Accessible history: the earliest record that can still be exported, not the date the tool was adopted.
  • Volume: rough counts or storage size from the admin console.
  • Export route: API, native export, admin backup or a vendor request.
  • Join keys: the identifiers that link records to other systems.
  • Sensitivity and rights notes: customer content, personal data, security material, third-party code.

Pre-filled inventory rows for the core systems#

The core systems in most software companies are an issue tracker, a code host, a helpdesk, a chat tool, a wiki and a CRM. The rows below are starting points; replace them with what your own admins actually see.

Fill the export route column from the plan you actually hold. Slack offers exports of private channels and direct messages only on Business+ and Enterprise, and owners must apply for them; a Free workspace permanently deletes messages and files older than one year. Zendesk account exports must be enabled through Zendesk Customer Support and are not available on Team plans, though the REST API works on every plan. GitLab's project file export includes repositories, issues with comments, merge requests, releases and wikis, which is useful when history lives there rather than on GitHub.

Pre-filled inventory rows for the core systems
SystemRecord familiesJoin keyExport routeSensitivity notes
Jira or LinearIssues, comments, changelogs, sprints, linksIssue keyREST API or native exportSecurity and HR projects; customer names in descriptions
GitHub or GitLabCommits, branches, pull requests, review comments, releasesCommit SHA, pull request number, issue key in messagesAPI plus git cloneSecrets in history; vendored third-party code; customer-specific forks
Zendesk or IntercomTickets or conversations, internal notes, ratings, macrosTicket ID; linked issue key from an integration fieldAPI exportCustomer personal data in most threads
SlackPublic and private channels, threads, filesChannel ID and thread timestamp; pasted issue keys and pull request linksAdmin export, scope set by planPrivate channels and direct messages; employee personal data
Confluence or NotionSpecs, runbooks, postmortems, decision recordsPage ID; embedded issue linksSpace or workspace exportCustomer names in postmortems; HR pages
Salesforce or HubSpotAccounts, opportunities, cases, activitiesAccount ID mapped to the helpdesk organizationAPI or data exportContact personal data; deal terms

The join keys that connect tickets, issues, code and threads#

Join keys are the identifiers that let a reviewer follow one problem from the customer's first message to the shipped fix. In a well-linked company, a Zendesk ticket carries a Jira issue key, branch names and commit messages contain that key, the pull request references it, and the release notes list it.

Slack fills the gaps between those systems. Engineers paste issue keys and pull request links into threads while debugging, so thread timestamps plus the keys they mention can attach informal discussion to the formal record.

  • Ticket to issue: an integration field, or a link in an internal note.
  • Issue to code: the issue key in branch names, commit messages or pull request titles.
  • Code to review: the pull request number joined to review comments and approvals.
  • Code to release: tags or release notes that list merged pull requests.
  • Discussion to issue: Slack threads or wiki pages that mention the key.

How to measure linkage without exporting data#

Linkage can be measured with searches inside each tool rather than exports. Search Jira for issues with linked development activity, search the code host for pull requests whose titles contain an issue key, and use helpdesk reporting to count tickets with a linked issue.

Record the result as a qualitative rating per period, because linkage usually changes when an integration was installed or a team adopted a naming convention. Many companies find strong linkage in recent years and much weaker linkage earlier.

Check identity linkage as well. If the same engineer appears under a work email in Jira, a personal handle in GitHub and a display name in Slack, record the mapping now; reviewers and preparation teams need it later to remove or pseudonymize names consistently.

How to measure linkage without exporting data
Linkage ratingWhat you seeWhat it means for the inventory
StrongMost fixes trace from ticket to issue to pull requestRecords can be packaged as connected histories
PartialIssue keys in code, few ticket linksEngineering history connects; support stands alone
WeakKeys appear only occasionallyTreat systems as separate record families

Sensitive areas to flag while you map#

Sensitive areas are easiest to flag while you are already inside each admin console, and deciding exclusions later is much faster when the flags exist. Mark them in the sensitivity column as you go.

In code, the main risks are credentials committed to history and third-party code. Gitleaks, an MIT-licensed tool for detecting passwords, API keys and tokens in git repositories, and TruffleHog, which scans Git, chats, wikis and other sources, are common first checks. Gitleaks' maintainer stated in May 2026 that it is feature complete and will receive security patches only. TruffleHog can also log in to confirm whether a found secret is live, so run that step with care.

In Jira, flag security, legal and HR projects. In Slack, flag private channels, direct messages and channels shared with customers. In the helpdesk, note which customers have contracts that restrict reuse of their data.

Illustrative: a construction scheduling software company maps its records#

Illustrative: the CTO of a fictional construction scheduling software company builds an inventory from the admin consoles of Jira, GitHub, Zendesk, Slack and Confluence. No data is exported during the mapping.

The inventory shows strong linkage between Jira and GitHub across the company's history, but the Zendesk to Jira integration was installed only partway through, so older tickets stand alone. Slack history is limited by an earlier plan, and one Jira project holds security reports.

The CTO rates engineering history as the strongest connected record family, marks recent support tickets as a possible second package, and flags the security project and private channels for exclusion. The finished inventory becomes the input to a metadata fit check.

How SourceX uses a software company inventory#

SourceX uses the inventory as the input to a metadata-only fit check, and nothing is shared during that initial assessment. Linkage, accessible history and sensitivity flags map to drivers in the SourceX Enterprise Data Value Framework, including human-generated signal, domain expertise, data cleanliness, rights, preparation cost and privacy burden.

Buyers increasingly expect provenance metadata too. The Data & Trust Alliance's Data Provenance Standards group dataset metadata into Source, Provenance and Use, and the Use group includes elements such as consent documentation location and license to use. For each package, the SourceX Evidence Packet records provenance, licensing rights, permitted use, the privacy record and release authorization.

Frequently asked questions

Do we need to export data to build the inventory?

No. The first version comes from admin consoles, plan settings and vendor documentation: which systems exist, how far back history goes, rough volumes and how exports work. Exports come later, and only for record families that move forward after a fit check and rights review.

Who should build the inventory?

The CTO or head of engineering usually owns it, with the head of support filling in helpdesk rows and IT confirming plans and admin access. Legal or the privacy lead should review the sensitivity column. A single owner keeps the format consistent across systems.

Should archived or retired systems be included?

Yes. Retired helpdesks, old version control servers and exports from tools you no longer pay for often hold the longest history. List them with their storage location and format, and note whether the original tool still exists to read them.

How do we handle repositories that belong to customers?

Record them, but mark them as customer-owned. Code written for a customer under a services agreement, or stored in a customer's own organization, is usually excluded from any license. The inventory should make that boundary visible before anyone discusses scope.

What about repositories under personal GitHub accounts?

List any repositories that live under personal accounts rather than the company organization, and move them into company control if the code belongs to the company. Check that the relevant employee or contractor agreements assign IP to the company. Code that stays under someone else's control is hard to export and harder to license.

How often should the inventory be updated?

Update it when a system is added, retired or changes plan, and before any data discussion. A short dated changelog at the top of the spreadsheet is enough. Many companies refresh it during their regular security or vendor reviews.

Sources

  • Gitleaks is an MIT-licensed tool for detecting secrets such as passwords, API keys and tokens in git repositories, files and stdin; on May 21, 2026 its README stated it is feature complete with security patches only. Source
  • Slack exports of private channels and direct messages are not offered on Free or Pro plans; they are available on Business+ and Enterprise, and owners must apply to use them. Source
  • Slack Free workspaces permanently delete messages and files more than one year old. Source
  • Zendesk data exports must be enabled through Customer Support and are not available on Team plans, while the REST API is available on every plan. Source
  • GitLab's file-export migration includes repositories, issues with comments, merge requests, releases, wikis and other project data. Source
  • TruffleHog, an AGPL-3.0 open-source secret scanner, can log in to confirm whether a secret it classifies is live, and scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
  • The Data & Trust Alliance Data Provenance Standards define dataset metadata in three groups, Source, Provenance and Use; the Use group includes consent documentation location and license to use. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify