Skip to content

IT service management and incident datasets for AI training

An ITSM ticket dataset is the record of how an IT organization handled its incidents, service requests, problems and changes, exported from its service management tool with work notes, handoffs, SLA timers and links to the affected configuration items. SourceX sources and licenses these histories, with alerts and postmortems where they exist, from established companies and managed service providers running ServiceNow, Jira Service Management, BMC Helix or Freshservice. Secrets are removed, and hosts, IP addresses and people are pseudonymized consistently before delivery.

Dataset manifest

Sourced to your spec
What it is
Incidents, problems, changes and requests with work notes, CI links and outcomes
Typical systems
ServiceNow, Jira Service Management, BMC Helix, Freshservice, PagerDuty
Typical history
Often spans years of operations records; varies by partner
Modality
Structured ticket fields and free-text work notes; alerts and postmortems if kept
Delivery formats
JSONL per record, or relational tables joined on record ID
Preparation
Secrets removed; hosts, IPs and people replaced with consistent pseudonymous tokens
Licensing
Permitted use and derivative rights set per license; exclusivity possible for an agreed snapshot
Availability
Depends on partner IT teams that hold matching records; not guaranteed

What a delivery contains

Fields vary by source system and are fixed per order. A typical delivery includes:

FieldTypeWhat it holds
record_idstringPseudonymous ID with a type prefix such as inc_, prb_ or chg_, stable across every table in the delivery.
record_typeenumIncident, service request, problem or change as the source tool classified it, with a flag for major incidents.
timestampsobjectOpened, acknowledged, resolved and closed times in UTC, plus on-hold periods; MTTA and MTTR derive from these.
priorityobjectImpact, urgency and resulting priority, captured at open and at close so re-prioritization is visible.
categorystringCategory and subcategory from the organization's taxonomy, which often changes when ITSM tools are replaced.
affected_ciobjectPseudonymous configuration item with its class, business service and CMDB relationships.
assignment_historyarrayEvery assignment-group handoff in order with timestamps, the basis for routing labels and reassignment counts.
journalarrayInternal work notes and user-visible comments, kept as separate entry types with author roles.
audit_trailarrayField-level changes with timestamps, used to rebuild the record as it stood at any point.
alertsarrayMonitoring alerts and on-call pages correlated with the incident, with source, signal and acknowledgment time.
slaarraySLA targets attached to the record, elapsed business time, pauses and breach flags.
related_recordsobjectParent and child incidents, the problem record, and the change that caused or fixed the incident.
resolutionobjectClose code, close notes, fix type, reopen count and the runbook or knowledge article used.
postmortemobjectTimeline, contributing factors and action items from the post-incident review, where one was written.

Example record

{
  "record_id": "inc_8d41c2",
  "record_type": "incident",
  "timestamps": { "opened": "2024-02-06T07:58:12Z", "acknowledged": "2024-02-06T08:03:40Z",
                  "resolved": "2024-02-06T09:41:05Z", "closed": "2024-02-09T10:00:00Z" },
  "priority": { "at_open": { "impact": 2, "urgency": 2, "priority": "P3" },
                "at_close": { "impact": 2, "urgency": 1, "priority": "P2" } },
  "category": "application / authentication",
  "affected_ci": { "id": "ci_svc_3f9a", "class": "application_service",
                   "depends_on": ["ci_db_77c0", "ci_lb_12e8"] },
  "alerts": [
    { "source": "apm", "t": "2024-02-06T07:55:31Z", "signal": "login p95 latency above 4s",
      "dedup_key": "al_5b2e" }
  ],
  "assignment_history": [
    { "t": "2024-02-06T07:58:12Z", "group": "service_desk" },
    { "t": "2024-02-06T08:10:47Z", "group": "identity_platform" },
    { "t": "2024-02-06T08:52:19Z", "group": "database_ops" }
  ],
  "journal": [
    { "t": "2024-02-06T08:21:33Z", "kind": "work_note", "author_role": "engineer_l2",
      "text": "Session store timing out. Connection pool to ci_db_77c0 saturated since 07:50." },
    { "t": "2024-02-06T09:02:10Z", "kind": "work_note", "author_role": "dba",
      "text": "chg_2e71ab cut max connections on ci_db_77c0 from 200 to 50 last night. Rolling back." },
    { "t": "2024-02-06T09:44:58Z", "kind": "comment", "author_role": "service_desk",
      "text": "Sign-in is working again. We are reviewing the cause." }
  ],
  "sla": [{ "name": "P2 resolution", "elapsed_min": 103, "breached": false }],
  "related_records": { "caused_by_change": "chg_2e71ab", "problem": "prb_0c55d3" },
  "resolution": { "close_code": "change_rolled_back", "reopened": 0,
                  "kb_used": "kb_session_store_triage" },
  "postmortem": { "id": "pm_1a90",
                  "contributing_factors": ["limit change not load tested",
                                           "change rated low risk with no peak-time review"],
                  "action_items": ["load test connection limit changes before approval",
                                   "alert on connection pool saturation"] }
}

Synthetic record for illustration. Field names, structure and format are agreed per order.

What AI teams use it for

Route and prioritize incoming tickets

Assignment-group handoffs label where each ticket should have gone first, and reassignment counts show how long misrouted tickets bounced between teams.

Train IT and SRE agents on real diagnosis

Work notes record the checks engineers ran, in order, against real dependencies, including the dead ends that ruled a cause out.

Predict change risk

Incidents linked to the change that caused them label which change types, configuration items and timing windows tend to end in outages.

Evaluate agents on held-out incidents

Give an agent the alerts, CI context and first notes of a resolved incident, then grade its diagnosis against the cause the team actually found.

Draft postmortems and incident summaries

Incidents with written reviews pair the raw material, meaning alerts, notes and the timeline, with the summary the team actually produced.

Use-case guides: Enterprise and computer-use agents, Coding agents, Private evaluation sets

What makes this data valuable

Change linkage

Caused-by and fixed-by links tie outages to the deployments and maintenance behind them.

CMDB context

Affected CIs with their dependencies let a model reason about blast radius instead of guessing from text.

Separated note types

Internal work notes and user-facing comments, kept apart, show the diagnosis and how it was communicated.

Audit history

Field-level changes make it possible to rebuild what was known at triage, not just at closure.

Specific close codes

Codes such as change rolled back, configuration fix or vendor defect make outcomes learnable; a bare resolved code does not.

Postmortem coverage

Written reviews add timelines, contributing factors and action items that tickets rarely contain.

Why linked records matter more than ticket text

An incident ticket on its own is a thin record, a short description, a category, a few work notes and a close code, and most of its value sits in the records it links to. The configuration item shows what failed and what depended on it. The change record shows what someone did to that system shortly before. The problem record says whether the team found a root cause or settled for a workaround, and the postmortem explains the sequence in a way work notes rarely do. Unlinked ticket text teaches a model to write plausible notes; the linked graph lets it learn to move from symptom to cause.

Public sources do not provide that graph. Status pages and published postmortems cover a small number of large outages, written after the fact for an outside audience. They leave out the dead ends, the handoffs between teams and the routine incidents that fill most of an operations queue, which is the work an IT or SRE agent will actually be given.

Reconstructing what was known at triage

ITSM records are edited as an incident progresses, and a plain table export holds only the final state. Teams often correct the category, affected CI and priority at closure, so those fields record what was learned, not what was known when the ticket arrived. A triage model trained on final values learns from the answer without anyone noticing. CMDB joins have the same problem: attaching a three-year-old incident to today's dependency map adds infrastructure that did not exist at the time.

The remedy is to request the audit trail with the records, meaning field-level changes with timestamps, so each ticket can be rebuilt as of any moment, plus historical CMDB relationships where the partner keeps them. For evaluation, cut each item at a fixed point such as first acknowledgment and withhold everything after it, including close notes, linked problems and the postmortem. Whether a candidate dataset has usable audit history often decides if it can serve as an eval source at all, so put it in the request.

What to check before licensing

  • If the partner is a managed service provider, confirm its client contracts allow licensing tickets about client systems; clients often have to authorize it.
  • Scan a sample for secrets, such as passwords, API keys, tokens and connection strings pasted into work notes and attachments, and ask how the partner found and removed them.
  • Check that hostnames, IP addresses, internal domains and CI names are pseudonymized consistently, so one server keeps one token across incidents, changes and the CMDB.
  • Measure the machine-generated share of the sample. Monitoring auto-tickets, duplicates and auto-closed records can crowd out the human-worked incidents you want.
  • Ask whether categories, assignment groups or priority schemes changed during the period, for example after a tool migration, and whether old values map to new ones.
  • Confirm how security incidents are treated. They can describe unpatched vulnerabilities and attacker activity, and are often excluded or reviewed separately.
  • Check link integrity — the share of incidents with an affected CI, a problem record or a caused-by change, and whether the linked records are inside the delivery.

How licensing works through SourceX

  1. 1

    Define

    Send the domain, modality, volume, format, timeline and permitted use you need.

  2. 2

    Source

    SourceX identifies businesses that hold matching data and are open to licensing it.

  3. 3

    Qualify

    Fit, rights and quality are checked, and you review samples before committing.

  4. 4

    License

    Scope, permitted use, exclusivity, price and obligations are agreed in writing.

  5. 5

    Deliver

    Approved data is prepared, de-identified where required and transferred securely.

Questions buyers ask

Can I license real ServiceNow or Jira Service Management incident data?

Yes, when the organization that owns the records agrees to license them. The two platforms store the same work differently: ServiceNow keeps incidents, problems and changes in separate tables tied to its CMDB, while Jira Service Management records them as linked work items of different types and keeps configuration items in Assets, where a team uses it. Both return records and their field-change history through APIs, subject to audit settings, so the harder questions are ownership, secrets in free text and whether links survive the export.

Which record types can an ITSM dataset include?

Incidents are the core, usually alongside the record types they link to: problems with root causes and known errors, changes with approvals and implementation results, service requests, and configuration items with their CMDB relationships. Alerts, on-call pages, runbooks and postmortems can be added where the partner keeps them and the license covers them. Listing the record types and links you need lets candidates be screened on link coverage early.

How are hostnames, IP addresses and secrets handled?

Secrets are removed and infrastructure identifiers are replaced with consistent pseudonyms. Passwords, API keys, tokens and connection strings are targeted wherever they appear, including free-text work notes and attachments. Hostnames, IP addresses, internal domains and CI names map to stable tokens, so one server can be followed across incidents, changes and the CMDB without exposing the partner's network. Detailed rules are agreed during scoping and checked on a sample.

Can tickets from a managed service provider be licensed?

Sometimes, depending on its client contracts. A managed service provider's tickets describe its clients' systems and users, and those contracts may give clients ownership of that data or confidentiality over it. SourceX reviews the contract position first. Where client consent is needed, a dataset can be limited to the clients who agree, or to the provider's own internal IT records.

Are security incidents included?

Not by default. Security incidents can describe unpatched vulnerabilities, attacker activity and affected users, so partners often exclude them or release them only after a separate review. If you need security operations data, raise it at the start so it can be scoped with the partner instead of filtered out late.

How far back do ITSM histories go?

It varies by partner, and tool migrations are a common limit. Organizations that moved between ITSM platforms may have migrated older records with partial fields, archived them elsewhere or left them behind. Records created after the last migration usually carry the cleanest categories, links and audit trails. Each candidate dataset's manifest states the years covered and any gaps.

Evaluating this data for procurement?

Diligence packets are prepared per dataset. Rights, privacy processing and quality differ between datasets.

Request dataset diligence

Tell us what your models need

Send your spec — domain, volume, format, timeline and permitted use — and SourceX will match it against partner data and come back with what can be licensed.

Updated 3 October 2026. Own data like this? See how companies license it to AI developers.

See if you qualify