Software companies
Illustrative workflow data package: a SaaS company's support-to-fix histories
By SourceX Editorial · Updated
Short answer
A workflow data package for a SaaS company bundles linked records that follow one problem from the customer's first ticket through the engineering issue, the pull request, the release and the final reply. This illustrative example shows the record types, join keys, manifest and exclusions. The core rule: a case belongs only if system IDs, not guesswork, prove its links.
Key takeaways
- A support-to-fix case is a chain of five record types: ticket, issue, pull request, release and customer reply.
- Join keys stored by the systems themselves make a chain trustworthy; links guessed from text belong in a separate, labeled subset.
- Pseudonyms must stay consistent across every system, or the chain breaks during preparation.
- The manifest describes the package at two levels: each case, and the dataset as a whole.
What is a support-to-fix workflow package?#
A support-to-fix workflow package is a set of linked cases, each tracing one customer problem from report to resolution across the systems a SaaS company uses. It differs from a ticket export because every case carries the engineering work and the customer communication that closed it.
AI developers look for this shape because it records the full path of real work: how a symptom was described, how it was diagnosed, what code changed, how reviewers judged the change, when it shipped and what the customer was told. That sequence is useful for training and evaluating both coding agents and support agents.
Everything below describes a fictional company. The structure is the point: record types, join keys, a manifest and written exclusions that anyone could apply again.
Illustrative: the company and its systems#
Illustrative: a fictional accounts payable automation vendor has sold to mid-market finance teams for many years. Support runs in Zendesk, engineering tracks work in Jira, code lives in GitHub, release notes are published from Confluence, and escalations are discussed in a shared Slack channel.
When the CEO and CTO decide to explore a license, they agree on one scoping decision up front: the package will contain only cases where a closed ticket links to a Jira issue. Tickets that never touched engineering stay out, which keeps the package focused on support-to-fix work rather than general support.
The record types and their join keys#
The record types in this package are joined by IDs that the systems stored at the time, not by matching words after the fact. The support team used a Jira integration that wrote the issue key onto the ticket, and engineers put the issue key in branch names and pull request titles by convention.
Where a key was missing, the team did not guess. Cases linked only by a Slack thread or a phrase in a comment went into a separate subset marked as inferred links, so a buyer can include or ignore them.
| Record type | Source system | Join key | Kept | Removed or replaced |
|---|---|---|---|---|
| Support ticket | Zendesk | Ticket ID and Jira key field | Subject, conversation, tags, resolution code, timestamps | Customer names, emails, phone numbers, invoice attachments |
| Engineering issue | Jira | Issue key | Summary, description, comments, status history, priority, fix version | Reporter and assignee names, replaced with pseudonyms |
| Pull request | GitHub | Issue key in branch or title | Diff, review comments, approvals, merge event, CI status | Secrets and customer data in test fixtures |
| Release | GitHub tags and Confluence notes | Fix version and tag | Release note text and release date | Internal remarks about named customers |
| Customer reply | Zendesk | Ticket ID | Final public reply and closing status change | Signatures and contact blocks |
The manifest: what each case file contains#
The case manifest gives every chain a new pseudonymous case ID and lists the records in that chain, the timestamp of each step and how each link was established. Record IDs are re-keyed for delivery, and the crosswalk back to the original Zendesk, Jira and GitHub IDs stays with the company. A buyer can then filter by product area, resolution code or link method without opening a single record.
At the dataset level, the manifest follows the shape of published provenance standards. The Data & Trust Alliance Data Provenance Standards, for example, group dataset metadata into Source, Provenance and Use, and the specification says that metadata is needed to enable proper dataset selection for AI model training.
The table shows one fictional case entry. Field names are illustrative; what matters is that every hop records its source and its link method.
- Case ID, issued for the package and unrelated to any customer account.
- Re-keyed ticket, issue, pull request and release references for every record in the chain.
- Timestamps for ticket creation, escalation, merge, release and final reply.
- Resolution code and reopen flag from the help desk.
- Link method for each hop: system field, branch naming convention or inferred.
- Inclusion notes, such as a rejected first pull request before the one that shipped.
| Manifest field | Illustrative value | How it was set |
|---|---|---|
| case_id | case-0417 | Issued for the package |
| ticket_ref | tkt-a91c | Re-keyed; original Zendesk ID kept in the company's crosswalk |
| issue_ref | iss-5b20 | Linked by the Jira key field written by the help desk integration |
| pr_refs | pr-77e1 closed unmerged; pr-77f4 merged | Linked by the issue key in branch names |
| release_ref | rel-2c08 | Linked by the Jira fix version and the matching Git tag |
| step_times | Created, escalated, merged, released, replied | Taken from system timestamps, not edited |
| outcome | Solved, not reopened, resolution code Known defect | Copied from help desk fields as recorded |
| link_method | System field for every hop | Cases with any inferred hop go to the inferred subset |
How cases were selected and excluded#
Case selection in this example follows written rules, so the package can be rebuilt from the source systems and checked by someone who was not involved. Each rule names what is in, what is out and why.
The customer restriction rule came out of the rights review, not the engineering team. Counsel reviewed order forms and master agreements and listed the accounts whose tickets had to be removed, and the export script applied that list before preparation began.
| Rule | Included | Excluded or flagged |
|---|---|---|
| Chain completeness | Ticket, issue, merged pull request and release all present | Chains missing a release, flagged as unshipped |
| Link method | Keys stored by Zendesk, Jira or GitHub | Text-only links, moved to the inferred subset |
| Customer restrictions | Tickets from customers with no contractual limit on this use | Tickets from customers whose contracts restrict secondary use |
| Code ownership | Repositories the company owns outright | Customer-specific code and vendored third-party libraries |
| Security | Ordinary defects and their fixes | Vulnerability reports and incident threads |
| Rejected work | Closed pull requests with review reasons that preceded the merged fix | Closed pull requests with no comment |
Preparing the package without breaking the chain#
Preparation for a linked package has one extra requirement compared with a flat export: every replacement must be consistent across systems. The engineer who appears as an assignee in Jira, a reviewer in GitHub and a voice in the Slack escalation must receive the same pseudonym in all three places, or the chain stops making sense.
Code needs its own pass. Pull requests and abandoned branches can hold API keys, connection strings and sample invoices used as test fixtures. Secret scanners help; TruffleHog, for instance, says it scans Git, chats, wikis and logs and can log in to check whether a found secret is still live, which calls for care when the credentials belong to a third party.
Timestamps stay, because the order and spacing of steps is part of what the package shows. Customer names inside code comments, commit messages and release notes are replaced with role labels such as customer A.
Mistakes that break a support-to-fix package#
The mistakes that most often break a package like this come from treating it as several separate exports instead of one set of chains. Each mistake below removes a link, a step or a person from the story.
- Linking tickets to issues by fuzzy text matching and presenting the result as fact.
- Exporting current issue fields without the status and field change history.
- Dropping closed, unmerged pull requests, which removes the record of what reviewers rejected.
- Pseudonymizing each system separately, so one engineer becomes three people.
- Leaving customer-specific code or vendored libraries in the repositories that feed the diffs.
- Skipping the release step, so a buyer cannot tell whether a fix ever reached customers.
How SourceX would handle a package like this#
SourceX would run this package through the SourceX five-step transaction. In Supply, a metadata-only fit check asks which systems hold the records, how many years are accessible and how the systems link; in Rights, contracts and repository ownership are reviewed; in Preparation, personal and confidential details are removed with consistent pseudonyms.
In Approval, the company signs off on the manifest and a prepared sample before anything is released, and in Delivery the package stays in the company's own storage or ships on an encrypted drive. The SourceX Evidence Packet then records provenance, licensing rights, permitted use, the privacy record and release authorization.
Frequently asked questions
How many linked cases does a package need?
There is no fixed minimum. A buyer weighs depth and consistency alongside volume: a smaller set of complete chains with stable join keys can be more useful than a large set of partial ones. The fit check looks at how many years of linked history exist before anyone discusses scope.
Do we have to include source code?
Not always. Some packages include pull request metadata, review comments and limited diffs rather than full repositories, and some leave code out in favor of tickets and issues. Each choice changes what a buyer can do with the package, so the decision is made repository by repository with the CTO.
What if our help desk never stored Jira keys?
Links can sometimes be rebuilt from escalation threads, issue comments that quote ticket numbers, or integration logs. Rebuilt links should be labeled with the method used and kept apart from links the systems stored, so a buyer can judge how reliable they are.
Can Slack escalation threads be part of the chain?
Threads from shared escalation channels can add diagnosis that never reached Jira, provided employee notices and workspace policies allow it. Direct messages and private channels need a separate review and are often excluded. Names in threads take the same pseudonyms used everywhere else in the package.
Does licensing a package like this transfer ownership of the records?
No. The records are licensed, not sold outright. The company keeps ownership and grants a buyer a limited right to use a defined copy for a stated purpose and term, with deletion and other obligations set out in the contract.
Sources
- The Data & Trust Alliance's Data Provenance Standards (version 1.0.0 specification) define dataset metadata in three groups: Source, Provenance and Use, and say this metadata is needed to enable proper dataset selection for AI Model Training. Source
- TruffleHog says it can log in to confirm whether a classified secret is live, and it scans sources including Git, chats, wikis, logs, object stores and filesystems. Source
Related resources
- IndustryFintech software data
- InsightCan property management companies sell their data to AI companies?
- InsightMaintenance-mode software products: what their engineering histories hold
- InsightAI features in acquired products vs licensing records out: a holdco rule
- IndustryBPO & contact centers data
- GlossaryEnterprise data
See if your company qualifies
A short company assessment. No data uploads are needed.