Skip to content

Data quality, coverage and contamination

Checking Completeness of Case and Ticket Records: Missing Steps, Truncated Threads and Absent Outcomes

Quick answer

Incomplete records in training data are cases or tickets that lack part of their path from request to outcome: the opening message, the intermediate actions, the closing resolution, or the timestamps that order them. To check a delivery, define what a complete case means for your use, reconstruct each case from its event history, and compute record-level metrics such as share complete, share with an outcome and share truncated at the export window. Agree acceptance thresholds before the data arrives, not after.

By SourceX Editorial · Updated

Why record-level completeness matters more than field-level completeness

Field-level null rates miss the failure that hurts agent training most: a case can have every column populated and still be missing half its steps. Standard data quality models treat completeness as a measurable characteristic [1], but for workflow data the unit of measurement has to be the case, not the cell. An SFT example built from a ticket whose resolution step was dropped teaches the model to stop early; a trajectory missing the middle teaches it to jump from complaint to refund with no lookup in between.

The effect is asymmetric. Duplicates and noisy labels dilute a dataset; truncated cases actively teach wrong behavior, because the model sees an apparently finished episode. If you plan to turn tickets into trajectories, read reconstructing agent trajectories from ticket and case histories alongside this page, and use the broader training data quality metrics page for the field-level checks this one does not cover.

Defining a complete case before you measure anything

A complete case is one you can replay from opening request to recorded outcome without guessing what happened in between. Write the definition down per source system, because "complete" in a Zendesk support queue differs from "complete" in a ServiceNow incident or a CRM opportunity. A practical working definition has four parts:

  • Opening event: the original request, inbound email, form submission or call note, with a creation timestamp.
  • At least one action: an agent reply, internal note, status transition, assignment or tool call recorded as an event.
  • Closure with outcome: a terminal status (solved, closed, won, lost, denied) plus an outcome field or a final message that states what was done.
  • Ordered timestamps throughout: every event carries a timestamp, and the sequence is monotonic once timezone and clock issues are normalized.

ISO/IEC 5259-4 frames data quality as a process that applies across supervised, unsupervised and reinforcement learning [2], which is a useful reminder that the definition depends on the training objective. An evaluation set for resolution-quality scoring needs outcomes; a dataset for routing classifiers may only need the opening event and first assignment.

Where completeness breaks in ticket and case exports

Most gaps come from how the data left the source system, not from how the work was done. Knowing the export path tells you where to look.

Event history stored outside the main record. In Zendesk, a ticket's history lives in ticket audits, which record creation, field changes with their previous values, and comments [3]. An export of the tickets table alone gives you final state with no steps. ServiceNow keeps comments and work notes in the sys_journal_field table rather than on the incident row [4], so a task export without a journal join drops the conversation entirely.

Internal notes filtered out. Exports often keep only public comments. The customer sees "We've fixed this," but the internal work notes that show the diagnosis are gone, which is exactly the reasoning an agent model needs.

Export window truncation. A delivery bounded by date range cuts cases that opened before the start date (missing their beginning) or closed after the end date (missing their outcome). This looks like a large number of short, unresolved cases clustered at both edges of the window.

Email threads split or merged. When cases arrive by email, the reply chain is reconstructed from Message-ID, In-Reply-To and References headers [5]. Forwarding, mail clients that strip headers, or subject-line edits break the chain, so one case becomes several orphan fragments.

System migrations. A helpdesk or CRM migration often resets internal IDs. If the supplier matched records with an external ID field during the move, as Salesforce upsert does [6], you can rejoin old and new halves; if not, pre-migration history may sit under different IDs or be gone. Look for a cluster of cases whose first event is a bulk "imported" action on the migration date.

Attachment and transcript loss. Screenshots, log files, call recordings and chat transcripts are often stored by reference. If the export keeps the reference but not the object, the record looks complete while the evidence is missing.

Metrics that quantify incomplete records

Measure completeness with a small set of record-level metrics, computed per source system and per time slice. A single overall percentage hides the edge-of-window and migration effects described above.

Illustrative example: invented to show structure; it does not describe an available dataset.

MetricDefinitionWhy it mattersExample threshold to agree in advance
Share completeCases meeting all four parts of your definition / total casesHeadline usability for SFT and trajectoriesAt least 90%
Share with outcomeCases with terminal status and non-empty outcome field or final messageLabels for evaluation and reward signalsAt least 95% of closed cases
Median and p10 steps per caseCount of events between open and closeVery short cases often indicate truncationp10 at least 3 events
Edge truncation rateCases opened before or still open at the export boundary / totalDetects window cutsUnder 5%, flagged not dropped
Orphan fragment rateRecords with a reply or follow-up but no reachable opening eventDetects thread splits and migration breaksUnder 2%
Timestamp integrityCases with all events timestamped and monotonic after normalizationOrdering for replay and step labelsAt least 99%
Attachment resolution rateReferenced attachments or transcripts actually present / referencedHidden evidence lossAt least 95% where attachments are in scope
Internal-note presenceCases with at least one internal note where the source system supports themReasoning content for agent trainingReport only; compare to supplier's stated policy

The thresholds above are placeholders. Set yours from the training objective and the sample size you can review; the sample size for estimating a dataset's error rate page covers how many cases to inspect to trust a measured rate.

Running the check on a delivery

Run the check as a short, repeatable pipeline: rebuild each case from events, classify it against your definition, then sample the failures by hand. The steps below assume the delivery includes an event or audit table; if it does not, that absence is itself the first finding.

Illustrative example: invented to show structure; it does not describe an available dataset.

Completeness check: ticket delivery, one source system
1. Join case header to event history on case_id
   (Zendesk: tickets + ticket_audits; ServiceNow: incident + sys_journal_field + sys_audit).
2. Normalize timestamps to UTC; flag non-monotonic sequences.
3. For each case, compute: has_open_event, n_actions, has_terminal_status,
   has_outcome, first_event_ts, last_event_ts.
4. Flag edge truncation: first_event_ts < window_start + 1 day
   OR (no terminal status AND last_event_ts > window_end - 1 day).
5. Flag orphans: records whose parent or In-Reply-To target is absent.
6. Flag migration seams: first event is a bulk import action on a known migration date.
7. Resolve attachment references against the delivered file manifest.
8. Report each metric per system and per month; compare to agreed thresholds.
9. Draw a random sample from each failure class and review by hand
   to confirm the flag is real, not a parsing error.

Step 9 matters because automated flags overcount. A case with two events may be a genuine one-touch resolution ("password reset link sent, solved"), not a truncated one. Hand review tells you the true defect rate for each class, which is what you compare to the threshold.

Deciding what to do with incomplete cases

Not every incomplete case should be dropped; route each failure class to the use it can still serve. Cases missing outcomes cannot be evaluation items or reward examples, but their opening events may still train intent classification or routing. Edge-truncated cases can often be completed by a follow-up delivery that extends the window, which is a cleaner fix than deletion.

Dropping cases also changes the distribution. If complex, long-running cases are the ones most often cut at the window edge, filtering them leaves a dataset skewed toward easy tickets. Check whatever you remove against the coverage gap analysis you run for deployment fit, and against long-tail and edge-case coverage, since rare escalations are disproportionately long.

For high-risk systems in scope of the EU AI Act, As amended by Regulation (EU) 2026/1744, Article 10 requires training, validation and testing data to meet quality criteria under data governance practices [8]. As of October 2026, the high-risk application dates have reportedly moved to 2 December 2027 for Annex III systems under Regulation (EU) 2026/1744. Even outside that scope, recording your completeness definition, metrics and filtering decisions in a datasheet or Data Card [7] makes later debugging and audits far easier.

What to ask a supplier before the delivery is cut

Most completeness problems are cheaper to prevent in the export specification than to repair afterwards. Ask the supplier, or the intermediary managing the deal, to state how each case was assembled:

  • Which source tables or APIs were exported, and whether event history, audit logs and journal entries are included.
  • Whether internal notes are included, excluded or redacted, and why.
  • How the date window was applied (by creation date, close date or last update) and whether straddling cases are complete.
  • Whether any migrations occurred in the window, and how old and new IDs were linked.
  • How email threads were reconstructed and what happens to orphan replies.
  • Whether attachments and transcripts are delivered as files or references.
  • What the supplier's own completeness checks measured, with counts.

These answers belong in the dataset documentation, and they make your acceptance criteria testable. For how buyers think about usable business records more broadly, see what makes business records usable for AI and the supplier-side view in what if my data contains errors. The quality assessment hub links the remaining checks, including acceptance sampling for dataset deliveries.

How SourceX handles case and ticket history requests

SourceX sources operational datasets, including support and sales histories and other workflow records, from US companies on request; categories are not inventory, and a request does not guarantee a match. Each dataset is rights-reviewed and delivered under a license that defines the records, uses, term and delivery, with diligence materials on source and preparation prepared per dataset. If your completeness definition is already written, include it when you describe the data you need, so suppliers can judge whether their exports can meet it. Examples of the kinds of workflow data involved are on the workflow task histories page.

Sourcing complete case histories for agent training

If you need end-to-end case or ticket histories for agent training, SFT or evaluation, describe the records, the completeness definition and the intended uses, not the businesses you hope hold them. SourceX looks for US companies that hold the described data, and nothing is contracted until a supplier agrees and approves the release. Start a buyer request.

Sources

  1. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-2:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 2: Data quality measures" (2024). https://www.iso.org/standard/81860.html
  2. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Artificial intelligence - Data quality for analytics and machine learning (ML) - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  3. Zendesk Developer Docs, "Ticket Audit events reference". https://developer.zendesk.com/documentation/ticketing/reference-guides/ticket-audit-events-reference/
  4. CData Software, "ADO.NET Provider for ServiceNow: sys_journal_field table". https://cdn.cdata.com/help/BNK/ado/pg_table-systemjournalfield.htm
  5. IETF, "RFC 4021: Registration of Mail and MIME Header Fields" (2005). https://datatracker.ietf.org/doc/rfc4021
  6. Salesforce Developers, "upsert() - SOAP API Developer Guide". https://developer.salesforce.com/docs/atlas.en-us.api.meta/object_ref/sforce_api_calls_upsert.htm
  7. Google Research, "Data Cards: Purposeful and Transparent Dataset Documentation for Responsible AI" (2022). https://arxiv.org/pdf/2204.01075
  8. European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
  9. IETF, "RFC 5322: Internet Message Format" (2008). https://datatracker.ietf.org/doc/html/rfc5322

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data