Skip to content

Fine-tuning and post-training data

Prompt datasets for RLHF and post-training: sourcing real user requests

Quick answer

A prompt dataset for RLHF is a pool of inputs, with no responses attached, that your model answers during training while a reward model, verifier or rater scores the answers. Its distribution decides which behaviors get practiced. Real requests from operating businesses, such as customer tickets, IT service requests and employee questions, give an authentic distribution. Buying them well means fixing the taxonomy and difficulty mix, collapsing templated duplicates, de-identifying without erasing intent, holding out an evaluation slice and confirming the supplier may license what its users wrote.

By SourceX Editorial · Updated

This guide, part of the fine-tuning and post-training data hub, covers the prompt side only; for a team's wider data plan, see data sourcing for post-training teams, and for the term, RLHF.

Why post-training buyers license prompts without responses

In on-policy post-training the model writes its own responses, so the prompts are the input you choose directly: they decide which tasks get sampled, scored and reinforced. That makes a prompt pool a separate, longer-lived purchase than demonstrations or preference pairs, because one pool serves several training rounds.

InstructGPT made the split explicit: labeler demonstrations for supervised fine-tuning, rankings of model outputs for the reward model, and a set of prompts with no human labels as inputs to the reinforcement learning stage. Its prompts came from labelers and from requests submitted through the OpenAI API [1].

Business requests usually arrive with a human reply attached, such as an agent's answer to a ticket. You can license the request side only, or keep the resolution as reference metadata for rubric writers; the license should say which.

Post-training stepWhat the prompts feedWhat must travel with each promptBuying implication
Comparison collection for a reward modelSeveral sampled responses per prompt, ranked by raters [1]Any context the request depends onThe mix sets what the reward model learns to judge; see buying RLHF comparison data
Online RL against a reward modelPolicy rollouts scored during trainingPolicy text or account statePrompts unlike the reward model's training data get less trustworthy scores
RL with verifiable rewards (RLVR)Rollouts checked by a program or reference answerA reference answer, test or checkerFew raw requests have checkable answers; see RLVR datasets
DPO and other offline methodsCandidate pairs generated before training; DPO itself does not sample from the model during fine-tuning [2]Context at pair generation; chosen and rejected responses for trainingSee DPO preference datasets and on-policy vs off-policy preference data
Held-out evaluationWin rates or rubric scores between checkpointsA frozen list of prompt IDsSplit off before any training use

Four places real prompts come from, and the catch with each

Prompts come from shared chatbot conversations, open instruction and preference sets, commissioned writing, and the request queues of operating businesses, and each trades authenticity, domain coverage and licensability differently.

SourceWhat the prompts look likeMain limitationFirst check
Shared chatbot conversationsReal users addressing a general assistantA third-party catalog lists ShareGPT 52K under CC-BY-NC 4.0, limited to research and non-commercial use [3]Original license and the chat provider's terms
Open instruction and preference setsHuman-written and model-generated promptsOther models have likely seen them; license omission above 70% and error rates above 50% on popular hosting sites [4]Per-source licenses; see open datasets that allow commercial fine-tuning
Commissioned prompt writingAnnotators writing to a scenario listCleaner phrasing than real users; the writer sets the distributionA sample compared with real traffic
Business request queuesCustomers and employees asking for something they needPersonal data, templated text, missing context, rights tied to noticesSource systems, notices in force, de-identification method

For retrieval, see synthetic vs real user queries for retriever training; assistant chat logs are covered in user-assistant conversation logs.

Business systems that hold request text

  • Customer support tickets and chats. The customer's first message, with product, contact reason, channel, priority, escalation and reopen fields. Noise: auto-acknowledgments, macros, signatures, quoted threads. See customer support ticket datasets.
  • IT service requests and incidents. Short description and description, with category, configuration item and assignment group. Noise: monitoring alerts, form-filled catalog requests, pasted credentials; see IT service tickets and secrets in ticket and chat data.
  • HR and employee service cases. Pay, leave and policy questions by case type; leave and accommodation cases often carry health details.
  • Sales inquiries and internal help channels. Quote, compatibility and contract-change requests from web forms and shared inboxes; questions in IT or finance help channels, which often need the earlier thread as context.

SourceX sources operational datasets from US companies, including support and sales histories, engineering records, documents, and finance and legal workflows, where these requests are recorded. Buyers describe the data, not the businesses; SourceX looks for US businesses that hold it, and every release is approved by the supplying company. Datasets are sourced on request, not held in stock, so a request does not guarantee a match. To scope a pool of real requests, describe the request types and systems you need.

Choosing the distribution: mirror production traffic or cover the space

Decide before sourcing whether the pool should reproduce real traffic frequencies or cover the space of request types, because the two train different models. A paper on data representativity contrasts a miniature of the target population with coverage of the input space, arguing coverage helps robustness to distribution shift and narrows accuracy gaps between groups, while mirroring serves average-case error [5].

Business traffic is head-heavy: password resets, order status and invoice copies can fill most of a queue, and an unweighted sample spends most rollouts on tasks the model already handles. Stratify by taxonomy cell, cap the head, set a floor for the tail, and keep each prompt's natural frequency as a field so you can reweight later. InstructGPT capped prompts at 200 per user ID [1]; the business equivalent is a cap per customer account or requesting employee.

The mix will not resemble a general assistant's. InstructGPT's API prompts were led by open-ended generation, then open QA, brainstorming and chat [1]. Business requests skew toward closed tasks bound by policy and account state, such as refunds, access changes and eligibility questions. Assessing whether a dataset is representative covers measurement.

Tags that make the pool steerable

  • Task type and intent: question, troubleshooting, transactional change, complaint or policy exception, plus your taxonomy, the source system's original category and the mapping between them.
  • Context dependency: general knowledge, policy text, or account and system state.
  • Workflow difficulty proxies: escalation tier, reassignment count, reopen flag, agent turns, time to resolution, exceptions granted. They measure difficulty for the people who handled the request, so recalibrate against your checkpoint's reward or pass rate after a first rollout.
  • Locale, channel, length and risk flags such as legal threats, self-harm mentions or regulated advice.

Illustrative record: one billing request prepared as a prompt

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "prompt_id": "rq-2025-04-118302",
  "source": {"system_type": "customer_support_ticketing", "channel": "email", "created_month": "2025-04"},
  "messages": [
    {"role": "user", "content": "Hi, we were charged twice for the March renewal on account [ACCOUNT_ID], both $1,188.00 on [DATE_1]. Can you refund the duplicate and confirm we are still on the annual plan? Thanks, [PERSON_1]"}
  ],
  "context_required": "account_state",
  "context": {"plan": "annual", "invoices_in_period": 2, "refund_policy_version": "BILL-REF v7"},
  "taxonomy": {"task_type": "transactional_change", "domain": "billing", "intent": "duplicate_charge_refund", "source_category": "Billing > Refund request"},
  "difficulty": {"escalated": false, "agent_turns": 3, "reopened": false, "policy_exception": false},
  "sampling": {"stratum": "billing/duplicate_charge", "natural_frequency": 0.0031},
  "dedup": {"exact_hash": "sha256:9f2c41d0", "near_dup_cluster": "nd-55120", "template_family": null},
  "privacy": {"method": "typed_placeholders", "replaced": ["account_id", "date", "person_name"], "residual_review": "sampled"},
  "original_response_included": false,
  "split": {"key": "account", "assignment": "train"},
  "license_scope": ["rl_prompt_pool", "comparison_sampling"]
}

In a JSON Lines file the record sits on one line. The fields buyers most often forget are context_required, the near-duplicate cluster and original_response_included: without them you cannot exclude unanswerable prompts, keep template families on one side of a split, or show which part of the source record the license covers.

Cleaning request text without erasing what users asked

Raw queues need four passes before they work as prompts: collapse duplicates and templates, strip machine-generated and quoted text, handle hidden context, and de-identify while keeping intent.

  • Duplicates and template families. Lee et al. found language-modeling datasets full of near-duplicates, including one sentence repeated over 60,000 times in C4, and models trained on deduplicated data emitted memorized text about ten times less often [6]. Their near-duplicate pass targets documents that match except for templated fields [6], the shape of form-generated requests such as "Access request for [NAME], start date [DATE]". Keep one representative per template family with its family size, and add an embedding pass for paraphrases; see semantic deduplication.
  • Machine and quoted text. Remove auto-acknowledgments, alert payloads, footers, signatures, quoted messages and bot turns; see separating bot, macro and human turns.
  • Hidden context. "Why was I charged twice?" has no good answer without the account's invoices. Sent to RL bare, such prompts can reward a confident invented answer over a clarifying question. Attach a de-identified context snapshot, keep only requests answerable from supplied policy text, or tag them and train clarification deliberately.
  • Personal data. Replace identifiers with typed, consistent placeholders so the request keeps its structure; see masking vs surrogate replacement. The Presidio project itself states there is no guarantee it finds all sensitive information [7], and free text carries indirect identifiers such as a job title plus a rare event; see indirect identifiers in business text.

Rare requests can identify a person by content alone. The ORCAS authors noted that click logs are usually withheld as too revealing, and released theirs after aggregation and filtering that included a k-anonymity requirement [8]; a frequency threshold or human review of the long tail does that job for request text. Prompt tokens are usually not RL training targets, but one pool often feeds SFT data and reward-model comparisons too: InstructGPT filtered its training-split prompts for personally identifiable information [1], and research shows fine-tuning can make a model more likely to reveal personal data from its training data [9].

Holding back a prompt slice for evaluation

Freeze an evaluation slice before the pool touches any training run, and split by requester and by time rather than by random row. Random splits put near-identical requests from one customer on both sides and inflate win rates.

InstructGPT split train, validation and test sets by user ID [1]. In business data, split on customer account or requesting employee, keep each near-duplicate cluster on one side, and test on the most recent months so drift shows. String overlap alone cannot verify the split: LMSYS showed that paraphrased or translated test items pass n-gram decontamination [10], so compare splits with embedding similarity too. See contamination-resistant evaluation design and evaluation datasets built from real business work.

Rights in text written by customers and employees

A supplier can license request text only as far as the privacy notices, terms and customer contracts in force when each request was written allow, so ask for those documents by collection period rather than accepting a general warranty.

  • Notice changes. Two 2024 FTC staff blog posts warn that moving to AI training or third-party sharing through a surreptitious retroactive change to terms or a privacy policy may be unfair or deceptive [11], and that model-as-a-service companies that promised not to use customer data for training can be liable if they break that promise [12]. See matching records to the notice in force at collection.
  • Business customers and employees. Requests from a supplier's business customers can fall under agreements that restrict secondary use; see customer contracts and DPAs. Internal help requests are employee communications.
  • Downstream disclosure. As of October 2026, providers of general-purpose AI models on the EU market must publish a sufficiently detailed training-content summary under Article 53(1)(d) of the AI Act [13]. California's AB 2013 required developers of generative AI systems offered to Californians to post training-data documentation by 1 January 2026, including whether datasets contain personal information or licensed material [14]. Per-prompt source, collection period and license ID make those answers possible.

For datasets SourceX sources, rights review checks that the business owns or may share the records and that required consents are in place, and the license defines which records are included, their permitted uses, the term and delivery. Personal details such as names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked; no de-identification method is perfect.

Red flags in a prompt dataset offer

Weak offers fail on authenticity, counts or rights, and a sample review exposes them:

  • "Real" prompts rewritten by an LLM for privacy or style, with no per-record rewrite flag.
  • Counts for the whole set only, not per taxonomy cell after deduplication.
  • Non-commercial terms inherited from public chat logs [3].
  • Chat-template tokens, system prompts or agent macros baked into the text, or uniform [REDACTED] masks.
  • No split key or per-prompt source field, so leakage and public benchmark items cannot be ruled out.

Specification template for a real-world prompt pool

A usable request covers use, sources, distribution, metadata, privacy and licensed uses; copy these rows into an RFP or data request.

FieldWhat to state
Post-training useComparison sampling, online RL, RLVR, DPO pair generation, evaluation
Request types and systemsFor example inbound billing tickets and IT access requests; types to exclude
Prompt unitFirst user message only, or thread context up to the request; responses excluded or metadata only
VolumeUnique prompts after deduplication, per taxonomy cell, for a pilot and for full delivery
Distribution ruleMirror traffic, cover the space, or capped mix; natural frequency recorded per prompt
Taxonomy and contextYour taxonomy, source-category mapping, difficulty proxies, context snapshots
De-identificationMethod, entity list, placeholder style, residual review, rare-request handling
DeduplicationExact and near-duplicate thresholds, template-family handling
Held-out sliceSize, split key (account or requester), time window
FormatJSON Lines, UTF-8, one record per line [15]; a messages array without chat-template tokens
DocumentationA datasheet covering motivation, composition, collection process and recommended uses [16]; machine-readable metadata such as Croissant-RAI [17]
Licensed uses and refreshRL prompt pool, comparison sampling, evaluation, synthetic expansion, vendor sharing, retention after the term; whether later periods can be added

Source real request distributions for your post-training runs

Tell SourceX the request types, source systems, volume, distribution rule and licensed uses you need. SourceX looks for US businesses that hold matching records, checks the data and the supplier's licensing permissions, and manages the license and delivery; nothing is contracted until a supplier agrees. Describe the request distribution you need.

Sources

  1. Ouyang et al., OpenAI (arXiv:2203.02155), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  2. Rafailov, Sharma, Mitchell, Ermon, Manning, Finn, Stanford (arXiv:2305.18290; NeurIPS 2023), "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  3. LLM Configurator (third-party dataset catalog), "ShareGPT 52K: LLM Instruction / SFT Dataset". https://llmconfigurator.com/en/datasets/sharegpt
  4. Longpre et al. (arXiv:2310.16787; journal version in Nature Machine Intelligence 6, 2024), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  5. arXiv (2203.04706), "Data Representativity for Machine Learning and AI Systems" (2022). https://arxiv.org/pdf/2203.04706
  6. Lee, Ippolito, Nystrom, Zhang, Eck, Callison-Burch, Carlini (arXiv:2107.06499; ACL 2022), "Deduplicating Training Data Makes Language Models Better" (2021). https://arxiv.org/abs/2107.06499v1
  7. Microsoft (microsoft/presidio project), indexed on pkg.go.dev, "Presidio - Data Protection API". https://pkg.go.dev/github.com/microsoft/presidio
  8. Craswell et al., Microsoft (arXiv:2006.05324), "ORCAS: 18 Million Clicked Query-Document Pairs for Analyzing Search" (2020). https://arxiv.org/abs/2006.05324v2
  9. arXiv (2310.15469), "The Janus Interface: How Fine-Tuning in Large Language Models Amplifies the Privacy Risks" (2023). https://arxiv.org/pdf/2310.15469
  10. LMSYS Org (blog), "LLM Decontaminator: rethinking benchmark and contamination with rephrased samples" (2023). https://www.lmsys.org/blog/2023-11-14-llm-decontaminator
  11. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI (and other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing-your-terms-service-could-be-unfair-or-deceptive
  12. Federal Trade Commission, Office of Technology (Tech@FTC staff blog), "AI Companies: Uphold Your Privacy and Confidentiality Commitments" (2024). https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/01/ai-companies-uphold-your-privacy-confidentiality-commitments
  13. European Commission, AI Act Service Desk, "AI Act Article 53: Obligations for providers of general-purpose AI models". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-53
  14. California Legislature (California Legislative Information), "AB-2013 Generative artificial intelligence: training data transparency (Chapter 817, Statutes of 2024)" (2024). https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202320240AB2013
  15. jsonlines.org (community format specification), "JSON Lines". https://jsonlines.org/
  16. Gebru, Morgenstern, Vecchione, Wortman Vaughan, Wallach, Daumé III, Crawford (arXiv:1803.09010; Communications of the ACM 2021), "Datasheets for Datasets" (2018). https://arxiv.org/pdf/1803.09010
  17. Jain et al., MLCommons Croissant RAI task force (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data