Data quality, coverage and contamination
Templates, Boilerplate and Canned Replies in Business Records: Detecting and Handling Them
Quick answer
Boilerplate in tickets, email and chat (signatures, legal disclaimers, quoted reply history, auto-acknowledgments and agent macros) should be handled at the span level, not by dropping whole records. Detect it by counting how often each n-gram span recurs across records, then strip machine noise, mask reversible spans, and keep macro replies with a label because they are real agent behavior. Measure the boilerplate token share before and after cleaning so the change is auditable.
By SourceX Editorial · Updated
Why templated text is a different problem from duplicate records
Templated text is repetition inside records that are otherwise unique, so document-level deduplication misses most of it. Lee et al. found that common language-modeling datasets contain long repeated substrings, with one sentence repeated more than 60,000 times in C4, and that over 1% of unprompted output from models trained on them was copied verbatim from training data [1]. Their near-duplicate method also targets documents that are identical except for templated fields [1].
Business records amplify this. A help desk with 400,000 tickets may carry one confidentiality footer on every outbound email, a "We received your request" auto-reply on every ticket and a few dozen agent macros covering a large share of first responses. MinHash-based near-duplicate detection collapses whole tickets that differ only by order number; it does not remove the footer from the other 399,999 unique tickets.
The training failure modes are concrete:
- SFT: the model learns to close every answer with "Best regards, Support Team" and a privacy notice.
- Agent training: trajectories overweight the "acknowledge, apologize, link the KB article" pattern because macros dominate first responses.
- RAG: chunk embeddings cluster on shared disclaimers, so retrieval returns the footer-heavy chunk instead of the answer.
- Evaluation: templated replies inflate overlap metrics such as ROUGE because predicted and reference text share the same boilerplate.
Where boilerplate hides in ticket, email and chat exports
Each source system injects templated text in predictable places, so start with the export schema rather than the text. In ServiceNow, comments and work notes are stored as entries in sys_journal_field rather than on the incident row [7], which means a join that ignores the parent table and record keys can attach journal entries, including system-generated ones, to the wrong incidents or repeat them. Zendesk, Freshdesk and Salesforce Service Cloud exports carry similar artifacts in comment bodies, trigger notifications and email-to-case threads.
Common span types, roughly in order of how much token volume they consume:
- Quoted reply history (
On Tue, ... wrote:,>prefixed lines, OutlookFrom: / Sent: / To:header blocks). Each email reply re-embeds the full thread, so a 12-message thread can contain the first message 11 times. - Signatures and footers: name, title, phone, address, social links, "Sent from my iPhone".
- Legal and confidentiality disclaimers appended by mail gateways.
- System notifications: auto-acknowledgments, SLA reminders, satisfaction-survey prompts, status-change messages ("Ticket status changed from Open to Pending").
- Agent macros and canned responses: prewritten replies, often with merge fields like
{{ticket.requester.first_name}}already filled. - Chat widget scaffolding: pre-chat form echoes, bot greetings, "An agent will be with you shortly", transfer notices.
- Form templates in case notes: repeated field labels such as "Issue: / Steps taken: / Resolution:" that carry structure but no content.
Signatures and footers also concentrate personal data (names, direct phone numbers, office addresses), so boilerplate handling and de-identification overlap. Treat the two as separate passes with separate logs; see indirect identifiers in business text and scanning for credentials and secrets.
Detecting templated spans with frequency, position and structure
The most reliable detector is cross-record span frequency combined with position, backed by a few structural rules. Free text written by a customer rarely repeats verbatim across hundreds of unrelated records; templated text does.
Span frequency. Tokenize each message, then count n-grams (8 to 13 tokens is a reasonable starting range for English business text; tune it on a labeled sample) across the corpus, keyed by distinct record rather than raw occurrences. Spans that appear in more than a threshold of distinct records, for example 0.5% of tickets or 50 records, whichever is larger, are template candidates. At large scale, suffix-array indexes such as the one behind infini-gram make arbitrary-length n-gram counting tractable over very large corpora [1][2]. Merge overlapping high-frequency n-grams into maximal spans before labeling.
Position. Signatures and disclaimers cluster at the end of a message; acknowledgments and greetings at the start. Score each candidate span by its relative offset within the message. A span with high frequency and a consistent tail position is almost always a footer.
Structure. Regex rules catch quoted history (^>, On .* wrote:$, -----Original Message-----, Outlook header blocks) and signature delimiters (-- on its own line). Use the ticketing system's own metadata where it exists: author type (agent, end user, system), channel, via or source fields, and whether a comment was created by a trigger or automation. System-authored comments can usually be classified without reading the text.
Macros. If the supplier can export the macro library (titles and bodies) or the macro-applied events from the audit log, match replies against it directly. Without that, cluster high-frequency agent-authored first responses after normalizing merge fields (names, order numbers, dates replaced with placeholders), then review the clusters by hand.
Two traps are worth checking. Short legitimate phrases ("Thanks for your help") also recur, so require a minimum span length. Machine-written text from chatbots or generative drafting tools can look human but repeat at template-like rates; the companion guide on detecting model-generated content covers that case.
Strip, mask, label or down-weight: a decision table
The right action depends on whether the span carries behavior you want the model to learn. Signatures and quoted history are noise for almost every use; macros and form labels are signal in some uses and noise in others. Apply the decision per span type and per downstream use, not once for the whole corpus.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Span type | Detection rule | SFT / chat models | Agent training | RAG index | Default action |
|---|---|---|---|---|---|
| Quoted reply history | Regex on > and wrote: headers; offset match to earlier message in thread | Strip | Strip; keep thread order via message IDs | Strip | Strip, store offsets |
| Signature and footer | Tail position + frequency > threshold | Strip | Strip | Strip | Mask with [SIGNATURE] token |
| Legal disclaimer | Exact-match library of gateway footers | Strip | Strip | Strip | Strip, store hash |
| System notification | Author type = system or trigger | Drop message | Keep as event, not utterance | Drop | Convert to structured event |
| Agent macro | Match to macro library or normalized cluster | Keep a capped sample | Keep, labeled macro_id | Keep one canonical copy | Label, then down-weight |
| Edited macro | Macro match with edit distance above floor | Keep | Keep, flag macro_edited=true | Keep | Keep with label |
| Form field labels | Fixed label set per form | Keep as structure | Keep | Keep as metadata | Normalize to fields |
A useful rule for macros is to cap each macro_id at a fixed number of SFT examples (for instance 20 to 50) and keep every edited instance. Edited macros show the agent adapting a policy answer to a specific case, which is exactly the judgment the ABCD dataset was built to capture: agents following company guidelines and taking actions, not just producing fluent text [3]. Deleting all macros removes that policy signal and leaves a skewed sample of only the unusual, freehand replies.
For summarization work such as the customer-service dialogs in TWEETSUMM [4], strip greetings and closings from the source dialog before generating or scoring summaries, or the summaries will learn to restate them.
Keeping removal reversible and auditable
Every removal should be recorded as a span annotation, not applied destructively, so the supplier and buyer can both audit what changed. Store the raw text once in access-controlled storage, then publish the cleaned view plus a sidecar file of span records.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"record_id": "tkt-000184233",
"message_id": "msg-03",
"span_start": 812,
"span_end": 1047,
"span_type": "signature",
"detector": "tail_freq_v2",
"evidence": {"distinct_records": 18240, "rel_offset": 0.91},
"action": "mask",
"replacement": "[SIGNATURE]",
"macro_id": null,
"span_sha256": "9f3c...e1"
}
The hash lets a reviewer confirm which template was removed without re-exposing its contents, and the detector version lets you re-run a single rule when it proves too aggressive. Keep the boilerplate pass separate from PII redaction: redaction tools, whether curation pipelines like NVIDIA's NeMo PII identification and removal module [6] or LLM-based redactors, have measurable miss rates [5], and a single combined log makes it impossible to tell which pass caused a defect.
Measuring boilerplate share before and after cleaning
Report boilerplate as a share of tokens by span type, by channel and by year, both before and after cleaning. A single corpus-wide number hides the problem; email-to-case tickets can be dominated by quoted history while chat transcripts are dominated by bot scaffolding.
A minimal acceptance report for a ticket, email or chat delivery:
- Token share by span type (quoted history, signature, disclaimer, system, macro) before and after.
- Records emptied by cleaning: messages whose remaining content is under, say, five tokens. These are often pure acknowledgments and should become events or be dropped.
- Over-stripping rate: from a random sample, the share of removed spans that contained case-specific content (an order number in a signature line, a customer's answer typed below the quoted history). Size the sample using error-rate sampling guidance or acceptance sampling plans.
- Under-stripping rate: remaining n-grams that still exceed the frequency threshold.
- Macro distribution: top macro IDs by count before and after capping.
- Thread integrity: after quoted-history removal, every message still maps to its parent, so completeness checks on case records still pass.
Inline replies are the most common over-stripping failure: a customer types answers between quoted lines, and a "remove everything after the first >" rule deletes them. Diff each message against the prior message in the thread rather than truncating at the first quote marker.
What to ask a supplier before you license ticket, email or chat data
Ask for the structure that makes boilerplate detectable before you ask for volume. Records exported as flattened text blobs, with author type, channel and automation flags removed, are far harder to clean than the native export.
Useful requests:
- Per-message rows with
author_type,channel,created_at,viaor source, and parent message ID. - The macro library export and any macro-applied audit events, so replies can be matched to
macro_id. - The list of mail-gateway disclaimers and auto-reply templates in use over the date range.
- Whether quoted history was already removed, and by what rule, with before and after samples.
- Whether chatbot or AI-drafted replies are present and flagged.
For broader context on judging these datasets, see the data quality hub, the overview of customer support ticket datasets, the ticket data glossary entry and the page on licensing chat logs. If you are scoping a request for operational records like these, you can describe the data you need through SourceX for buyers.
Sourcing support, email and chat records for training
SourceX sources operational datasets, including support and sales histories, from US companies on request and manages the licensing process; a request does not guarantee a match, and every release is approved by the supplying company. Personal details such as names, emails and phone numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. Describe the records, channels and fields you need at https://sourcex.si/buyers.
Sources
- Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
- arXiv, "Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens" (2024). https://arxiv.org/html/2401.17377v2
- Chen et al. (arXiv; NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
- Feigenblat et al. (arXiv; Findings of EMNLP 2021), "TWEETSUMM - A Dialog Summarization Dataset for Customer Service" (2021). https://arxiv.org/abs/2111.11894v1
- arXiv, "PRvL: Quantifying the Capabilities and Risks of Large Language Models for PII Redaction" (2025). https://arxiv.org/pdf/2508.05545
- NVIDIA, "PII Identification and Removal (NeMo Framework user guide 25.07)". https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
- CData Software, "CData Python Connector for ServiceNow: sys_journal_field table". https://cdn.cdata.com/help/BNM/py/pg_table-systemjournalfield.htm
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.