Fine-tuning and post-training data
Turning business records into instruction-response pairs
Quick answer
To create an instruction dataset from company data, start with records in which a person answered a request and the business accepted the answer: resolved tickets, approved final drafts, posted forms and endorsed expert replies. Put the request and the context that person saw at the time into the prompt, and the accepted output into the target. Keep only outcome-confirmed records, de-identify before pairing, strip macros and signatures, deduplicate, and hold out the most recent period for evaluation before you train.
By SourceX Editorial · Updated
This page covers the record-to-pair transformation. For buying finished data, see how to source supervised fine-tuning data; for the wider map, the fine-tuning and post-training data guide.
What a business record needs before it can become a training pair
A record converts into a usable supervised fine-tuning (SFT) pair only if it holds three things: a request, the output a person produced in response, and evidence that the business accepted that output. Without the third, you train on whatever was typed, including the replies that failed.
SFT trains a model to produce a target given a prompt; InstructGPT's first stage fine-tuned GPT-3 on demonstrations written by labelers [1]. Operational records are demonstrations produced as a by-product of work by people who knew the domain, often under time pressure and with templates.
Target quality matters more than volume. LIMA fine-tuned a 65B-parameter model on 1,000 curated prompt-response pairs, and its authors concluded that most knowledge comes from pretraining while instruction data mainly teaches the format and style of responses [2]. The voice and habits in your targets are what the model copies most reliably; for sizing, see how much data you need to fine-tune an LLM, and for the term itself the instruction tuning glossary entry.
Open instruction sets such as Alpaca store each example as an instruction, an optional input and a response [3]. Records also need the context made explicit:
- Instruction: what was asked: the customer's message, the reviewer's request, the form to complete.
- Context: what the person could see when answering: earlier turns, account state, attachments, the policy version in force.
- Target: the accepted output, cleaned of templates and identifiers.
- Metadata: record ID, timestamps, outcome signal, de-identification method and split, kept outside the text the model sees.
Four record patterns and where their fields go
Most operational systems hold one of four patterns, and each maps its fields to prompt and target differently; filter on the acceptance signal.
| Pattern | Typical systems | Prompt: instruction and context | Target | Acceptance signal | Where it breaks |
|---|---|---|---|---|---|
| Request → resolution | Help desks and IT service management tools such as Zendesk or ServiceNow; chat platforms | Customer message, earlier turns, account state at that time, linked help article | The agent's public reply | Solved, not reopened within your window, no later escalation | The fix was an action in another tool (refund, configuration change) the prompt never shows |
| Draft → approved final | Contract lifecycle management, document version history, marketing approval workflows | The draft plus an instruction rebuilt from reviewer comments or the style guide | The approved version | Approval or signature event on that version | Changes came from an off-record negotiation, so no instruction explains them |
| Input document → completed output | Accounts payable, claims intake, CRM lead capture, order entry | Document text (OCR for scans) plus the output schema | Field values as finally posted, after corrections | Posted, paid or reconciled, not reversed | The target depends on lookups (vendor master IDs, internal codes) the document cannot supply |
| Question → expert answer | Internal Q&A forums, engineering escalation queues, legal and compliance help desks | The question plus the documents the answer cites | The accepted or endorsed answer | Accepted flag, no later correction in the thread | Relies on tacit knowledge, a private follow-up or a since-changed policy |
Adjacent patterns have their own pages: case notes paired with summaries (summarization fine-tuning data), document versions (draft-to-final pairs for rewriting models), extraction into JSON (structured-output fine-tuning data) and decisions with stated reasons (decision records with rationale). For record-level fields, see SourceX's overviews of customer support ticket datasets and chat logs for AI training.
Rebuild the context as of the moment the answer was written
The prompt should contain what the person could see when they wrote the target, no more and no less. Too little teaches the model to invent facts it was never given; too much leaks information from after the answer that production will not have.
Point-in-time state. A common error is joining today's account snapshot to a ticket from last year. If the customer has since upgraded, the prompt says "Enterprise" while the agent was answering a "Starter" customer. Rebuild state from field-history or audit tables as of the reply timestamp; packaging linked records from multiple systems covers the keys and history this needs.
Fields that leak the outcome. Keep resolution codes, final status, close-time tags, satisfaction scores, later customer messages and anything written after the reply out of the prompt. They belong in metadata, where they drive filtering.
Internal notes. If your production assistant will see agent notes, include the notes written before the reply. If it will not, exclude them and keep a diagnostic note as a separate rationale field.
Multi-turn threads. A thread with three agent replies yields three candidate pairs, each with all earlier turns as context; drop bare acknowledgments. Multi-turn conversation data for chat fine-tuning and chat format, roles and loss masking cover formatting.
Context budget. Check rebuilt prompt lengths against the target model's context window. Keeping the first message, a state block and the latest turns usually beats cutting from the top.
Keep only outputs the business accepted
Filter to outcome-confirmed records (resolved and not reopened, approved, posted and not reversed), because the target is exactly what the model learns to reproduce. Business records offer a filter that public instruction sets lack: the organization's own verdict on each answer.
The AlpaGasus authors found that the 52,000-example Alpaca set contained many low-quality instances with incorrect or irrelevant responses; after ChatGPT scored each example from 0 to 5, a model trained on about 9,000 high-scoring examples outperformed the original in GPT-4-judged and human evaluations [3].
Filters that apply to most operational sources:
- Drop records closed by automation (inactivity auto-close, merges, spam); that "resolution" is not an answer.
- Drop cases reopened within a window you define; the reply did not work.
- Tag or drop replies marked as bot-generated or AI-suggested, or you end up distilling another model's output.
- Drop targets that are entirely macro text or too short to carry information ("Closing this out.").
- Cap the share from any single agent, account or template, so a few prolific writers do not define the model's voice.
- Set rejected drafts and reopened replies aside: they are negative examples for binary-feedback KTO data or rejected responses in DPO preference datasets.
An AlpaGasus-style LLM grader, checked against a human-reviewed sample, makes a sensible second pass; filtering instruction-tuning data for quality covers the trade-offs.
Strip templates, signatures and quoted threads, then deduplicate
Remove repeated template text before pairing and deduplicate near-identical targets, because boilerplate repeated thousands of times dominates the training signal and is the text a model most readily memorizes. Lee et al. found many near-duplicates in common language-model datasets, including one sentence repeated more than 60,000 times in C4; deduplicated training produced models that emitted memorized text about ten times less often and reached the same or better accuracy in fewer steps [4].
Business text concentrates the problem:
- Strip quoted replies (lines prefixed with ">" or below "On … wrote:"), disclaimers and signature blocks before pairing.
- Ask for macro identifiers or the macro library, so inserted template text can be removed or tagged in targets.
- Run near-duplicate detection on targets and on whole pairs. Lee et al. paired exact-substring matching with a MinHash-based method for documents that are identical except for templated fields [4], which is the shape of a macro reply with a different name filled in.
- Cap how many members of each near-duplicate cluster survive.
De-identify the whole record before splitting it into pairs
Run de-identification on the raw record, with one consistent surrogate per identifier, before you split it into prompts and targets. Pairing copies the same names and account numbers into several prompts, and fine-tuned models can reproduce training text.
Carlini et al. extracted hundreds of verbatim training sequences, including personal contact information, from GPT-2 by querying it [5], and later research reports that fine-tuning can amplify such privacy risks [6]. Repetition compounds it: a thread split into three pairs repeats a first-message identifier in all three prompts, and repeated text is memorized more often [4].
Surrogates, not tags. Replacing every name with [NAME] turns "Dana asked Priya to approve" into "[NAME] asked [NAME] to approve", which loses who did what and teaches the model to emit bracket tags. Consistent, realistic surrogates (the same invented name for the same person across the record, account numbers in the original format) keep the text natural. The surrogate map stays with the data owner, never in the training set.
Detection has limits. The open-source Presidio SDK detects and anonymizes personal data, but the project states there is no guarantee it will find all sensitive information and recommends additional protections [7]. NVIDIA documents PII identification and removal in NeMo Curator and reports that its LLM-based approach outperformed Presidio in NVIDIA's own evaluation [8]. Measure recall per entity type on a labeled sample of your own records; tickets carry identifiers default recognizers miss, such as order numbers and internal hostnames.
Scan prompts as well as targets. Training stacks differ on whether loss is computed on prompt tokens; torchtune's chat dataset documentation, for example, exposes a setting for training on input messages [9]. With it on, the model learns to generate prompt text, identifiers included. A security vendor's vulnerability database also describes leakage of personal data that appeared only in fine-tuning inputs [10].
Regulated records. HIPAA de-identification uses either Safe Harbor, which removes 18 listed identifiers, or Expert Determination [11]; for health records, SourceX requires one of these methods before anything is considered for a license. As of October 2026, California's CCPA definition of deidentified information requires, among other conditions, that the business contractually obligate recipients to comply (Cal. Civ. Code §1798.140) [12]. See what CCPA-deidentified data obliges a buyer to do and the privacy-safe data guide.
Worked example: one support ticket becomes two training records
The invented ticket below shows the full transformation: point-in-time context, consistent surrogates, macro removal, an excluded internal note and split metadata.
Illustrative example: invented to show structure; it does not describe an available dataset.
Raw record as exported:
ticket 48213 | channel: email | created 2025-03-04T14:02Z
status history: new 14:02 > open 14:09 > pending 15:31 > solved 2025-03-05T09:40Z | reopened: no
requester: Dana Whitfield <dana.w@acme-fab.example> | org: Acme Fabrication | account: 00-731-552
plan at reply time (history table): Business annual | product: Invoicing 4.2
linked article: KB-1172 "Tax profiles after migration"
[1] customer 14:02 Since we switched to the new tax profile, March invoices for
customers in Texas show 0% sales tax. Dana Whitfield, Finance
[2] internal 14:20 TX rate dropped from profile in migration, see KB-1172.
[3] agent 14:25 Hi Dana, thanks for the details. The migration left the Texas
rate empty in your tax profile. Open Settings > Tax > Profiles,
edit "US default" and add the state rate; new invoices will
use it. {macro: signature_v3}
[4] customer 15:10 Done, new invoices look right. The ones already sent still
show 0%.
[5] agent 15:31 Sent invoices keep the rate in force when they were issued.
Use Credit & Reissue on each affected invoice; your customer
receives the credit note and the corrected invoice together.
{macro: signature_v3}
First training record, pretty-printed (in the .jsonl file each record is one line):
{
"id": "src-7d41a0-r1",
"source_record": "src-7d41a0",
"pattern": "request_resolution",
"created_at": "2025-03-04T14:25Z",
"split": "train",
"outcome": {"status": "solved", "reopened": false},
"deid": {"method": "consistent_surrogates", "entities": ["PERSON", "EMAIL", "ORG", "ACCOUNT_ID"]},
"rationale": "TX rate dropped from profile in migration, see KB-1172.",
"messages": [
{"role": "system", "content": "Billing support for an invoicing product. Account: plan=Business annual; product=Invoicing 4.2. Linked article: KB-1172 Tax profiles after migration."},
{"role": "user", "content": "Since we switched to the new tax profile, March invoices for customers in Texas show 0% sales tax. Maria Lopez, Finance"},
{"role": "assistant", "content": "Hi Maria, thanks for the details. The migration left the Texas rate empty in your tax profile. Open Settings > Tax > Profiles, edit \"US default\" and add the state rate; new invoices will use it."}
]
}
The second record, src-7d41a0-r2, carries messages 1, 3 and 4 forward as context and takes reply 5 as its target. Plan and product version come from history tables as of each reply, and the internal note stays out of the prompt (the production assistant will not see notes) but survives as the rationale field. "Dana Whitfield" becomes "Maria Lopez" in both the customer message and the agent's greeting, and the signature macro is gone from both targets. Both records inherit the ticket's split, so no thread straddles training and evaluation.
Store the result as JSON Lines: UTF-8, one JSON object per line, conventionally with the .jsonl extension [13].
Split by time and by thread before any training
Hold out the most recent period of records, and keep every pair from the same ticket, contract or customer account on the same side of the split, so evaluation measures performance on new cases instead of recall of near-copies. A random split puts turn two of a thread in training and turn three, with nearly identical context, in evaluation.
A time split also exposes policy drift: answers citing last year's prices or procedures pass a random split and fail on recent cases. After splitting, check for overlap. N-gram overlap is the most widely used contamination check [14], but paraphrased or translated items can slip past it [15], so add an embedding-similarity pass for templated replies. Lee et al. found train-test overlap affecting over 4% of the validation sets of standard datasets [4].
Once domain experts verify its references, the held-out slice can seed a golden evaluation dataset built from business records.
What to require when you license records for conversion
When the records come from another company, write the fields and history that conversion needs into the delivery specification and the license, because most of them cannot be rebuilt later:
- Full threads with a timestamp and author role on every message: customer, agent, bot, internal note.
- Status and event history (reopens, escalations, approvals, reversals), not only the final status.
- Point-in-time state: history tables or snapshots for plan, product version and policy version.
- Macro, template and signature identifiers, or the macro library itself.
- A flag for AI-suggested or bot-generated replies, where the source system recorded one.
- Attachments or extracted text, marked by whether the person answering could see them.
- Stable join keys across systems.
- The de-identification method: surrogates or tags, consistency within a record, and how a sample was checked.
- License scope covering training on derived pairs, keeping the transformed dataset, any synthetic expansion from seed pairs, and evaluation use of the holdout slice.
- The notices under which the records were collected. In January 2024, FTC staff warned that companies may be liable if they break promises not to use customer data for purposes such as training models [16].
- A datasheet covering composition, collection process and recommended uses [17], extended with your mapping rules.
Measure conversion yield on a sample before committing: the share of delivered records that survive your filters as usable pairs. Evaluating a fine-tuning dataset before you buy it covers sample checks. First-party records also avoid a known weakness of open collections: an audit of more than 1,800 text datasets reported license omission above 70% and license errors above 50% on popular hosting sites [18] (see open instruction datasets that allow commercial fine-tuning).
If your own systems lack enough resolved or approved records, SourceX sources operational datasets from US companies, such as support histories, documents, and finance and legal workflows, on request. Each dataset goes through rights review and is delivered under a license that defines which records are included and what they can be used for. Personal details are removed or replaced before delivery and the method is recorded per dataset; no method is perfect, so run your own scan too.
You can describe the records and fields your conversion needs. Background is in SourceX's pages on training data for domain-specific fine-tuning and why AI buyers value operational context.
Source operational records for your instruction dataset
Describe the record types, fields, history and outcome signals your conversion needs, and the uses you need licensed. SourceX looks for US businesses that hold those records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Start a data request with SourceX.
Sources
- Ouyang et al., OpenAI, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
- Zhou et al., Meta AI and collaborators, "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
- arXiv, "AlpaGasus: Training A Better Alpaca with Fewer Data" (2023). https://arxiv.org/pdf/2307.08701v1
- Lee et al., "Deduplicating Training Data Makes Language Models Better" (2021; ACL 2022). https://arxiv.org/abs/2107.06499v1
- Carlini et al., USENIX Security, "Extracting Training Data from Large Language Models" (2021). https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting
- arXiv, "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
- Microsoft Presidio project, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
- NVIDIA, "PII Identification and Removal" (NeMo Framework User Guide 25.07). https://docs.nvidia.com/nemo-framework/user-guide/25.07/datacuration/personalidentifiableinformationidentificationandremoval.html
- PyTorch torchtune, "Chat Datasets" (torchtune 0.3 documentation). https://docs.pytorch.org/torchtune/0.3/basics/chat_datasets.html
- Promptfoo (vendor page), "LLM input PII leakage" (LM Security Database). https://promptfoo.dev/lm-security-db/vuln/llm-input-pii-leakage-75a0bd54/
- U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification
- California Legislature, "California Civil Code section 1798.140 (California Consumer Privacy Act definitions)". https://leginfo.legislature.ca.gov/faces/codes_displaySection.xhtml?lawCode=CIV§ionNum=1798.140
- jsonlines.org, "JSON Lines". https://jsonlines.org/
- arXiv, "Benchmark Data Contamination of Large Language Models: A Survey" (2024). https://arxiv.org/pdf/2406.04244
- LMSYS Org, "LLM Decontaminator" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
- Federal Trade Commission, Office of Technology, "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
- Gebru et al., "Datasheets for Datasets" (2018; CACM 2021). https://arxiv.org/pdf/1803.09010
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023; Nature Machine Intelligence 2024). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.