Skip to content

Data quality, coverage and contamination

Detecting Model-Generated Content in Purchased 'Human' Data

Quick answer

You cannot reliably detect model-generated text in training data with a detector score alone. Published detectors miss much AI text and falsely flag human writers, especially non-native ones [3][7]. The workable approach combines process evidence (keystroke and timing logs, paste events, tool-use policies), statistical signals (style uniformity, overlap with outputs from known models), record creation dates, and a stratified manual review sample. Contract for that evidence before delivery, then test a holdout yourself before acceptance.

By SourceX Editorial · Updated

Why undisclosed model content breaks human-data purchases

Model-written records defeat the purpose of buying human data, because the value of SFT demonstrations and preference rankings is that they encode human judgment rather than another model's distribution. InstructGPT-style pipelines fine-tune on labeler-written demonstrations and train a reward model on human rankings of model outputs [4]. If a share of those demonstrations came from a chat model, you are partly distilling that model, including its style tics, refusals and errors. You may also import output terms from the model's provider that your license never contemplated.

The risk is not hypothetical. EPFL researchers reran an abstract-summarization task on Amazon Mechanical Turk and, using keystroke detection plus a synthetic-text classifier, estimated that 33-46% of workers used LLMs [1]. The authors caution that the figure comes from one LLM-friendly task, so treat it as evidence that the incentive is real, not as a base rate. A follow-up study by the same group examined prevalence and prevention measures in crowd work [2].

The same exposure applies to different record types in different ways:

  • Free-text responses and demonstrations (SFT, reasoning traces): highest risk, since a model can produce the whole record.
  • Preference judgments and rubric scores (RLHF, DPO, evaluation): the label is short, but the rationale field, or the judgment itself, can be delegated to a model acting as judge.
  • Annotations (spans, categories, extracted fields): lower text risk, but bulk pre-labeling by a model followed by rubber-stamp acceptance produces model labels with human timestamps.

For the noise side of preference labels, see measuring noise and agreement in preference data. This page covers authenticity: who or what produced the record.

Why AI-text detectors fail as a sole acceptance test

A detector score is a weak, biased signal that should trigger review, never rejection on its own. Perplexity-style and classifier detectors are trained on particular model families and writing populations, and their error rates move when either changes; research on author roles shows detector behavior varies with who wrote the text [3]. Low perplexity, the core signal for many tools, also describes careful, plain professional writing and much non-native English prose. Even the EPFL crowd-work study paired its synthetic-text classifier with keystroke evidence rather than trusting the classifier alone [1].

Three failure modes matter for purchased data:

  1. False positives cluster on legitimate writers. Rater pools often include non-native speakers, and detector behavior shifts with author characteristics [3]. Rejecting batches on detector scores can quietly remove a population you deliberately paid for.
  2. False negatives are cheap to produce. A worker who prompts for a different style, paraphrases model output, or edits it lightly removes much of what a detector keys on. Detectors calibrated on one model's outputs also lag newer models.
  3. Templated human text looks synthetic. Support macros, rubric-driven rationales and style-guide prose are uniform by design. See handling templates and canned replies in business records before you treat uniformity as a model signal.

Use detectors as one ranked feature to choose which records humans review first.

Process evidence: the strongest authenticity signal

The most reliable evidence is how the record was created, captured at collection time, because it is hard to fabricate after the fact. The EPFL study relied on keystroke-level capture, not text analysis alone [1]. Ask the supplier which of these were logged per record, and in what format:

  • Session timing: task open, first keystroke, submit. A 400-word answer submitted 40 seconds after opening is a flag.
  • Paste and focus events: clipboard paste count and character length, window blur and focus changes. A single paste equal to most of the final text is the clearest signal available.
  • Edit trace: keystroke count versus final character count, deletions and revision bursts. Human composition produces a non-trivial edit ratio; transcription of a model answer produces a near-linear trace.
  • Tool environment: whether the annotation tool blocked paste, whether browser extensions were restricted, and whether any built-in AI assist (suggestions, auto-rationales, pre-labels) was enabled for that task.
  • Worker and batch identifiers: pseudonymous rater ID, batch ID, guideline version and tool version, so you can trace patterns to a person or a configuration change.

Also request the annotator tool-use policy that applied when the data was produced: whether LLM assistance was banned, allowed for drafting, or allowed with disclosure, plus how violations were detected and handled. Disclosure practice belongs in provenance documentation; annotator agreements and AI-assistance disclosure covers that side. Note that process logs may contain personal data about raters, so agree what is pseudonymized before it is shared.

Statistical signals inside the delivered records

Text-level signals help rank records for review when process logs are missing or incomplete. Treat each one as a hypothesis to confirm on a sample, not a verdict.

  • Style uniformity across workers. Distinct humans vary in length, punctuation, list usage and register. If rater A and rater B share near-identical sentence-length distributions, openers ("Certainly", "Great question") and closing summaries, look closer.
  • Overlap with known model outputs. Regenerate answers to a sample of the same prompts with several current and older chat models, then measure n-gram and embedding similarity between delivered responses and those generations. The tooling from MinHash and LSH near-duplicate detection applies directly.
  • Within-batch near-duplicates. Multiple workers submitting paraphrases of one answer suggests a shared generator.
  • Rationale-label mismatch. In preference data, a fluent, generic rationale that does not reference specifics of either response, or that contradicts the chosen side, suggests delegated judgment.
  • Agreement that is too high. Pairwise agreement far above your gold-set baseline on ambiguous items can mean raters are all asking the same model. Compare against rater calibration results.
  • Shifts by date. Plot each signal by submission week; a step change after a pay-rate or deadline change is informative.

Creation dates as a hard boundary in operational data

For operational records, the creation date is often the most defensible authenticity evidence available. A support ticket thread, engineering postmortem or legal memo authored before public chat models were widely available in late 2022 is very unlikely to contain chat-model output (earlier GPT-3-based writing tools existed but saw limited business use), provided the timestamp is the original system-of-record value and not an export or migration date. Ask which field carries the date (for example created_at from the ticketing system versus exported_at) and whether edits after that date are tracked.

Post-2022 operational text needs its own check. Many business tools now embed drafting assistants, and support platforms may insert AI-suggested replies that agents send unchanged. Ask whether the source system flags AI-drafted or AI-suggested messages and whether that flag survives export. For conversation data, separating bot, macro and human turns describes the turn-level fields to request.

A practical authenticity audit before acceptance

Run authenticity as a defined step in delivery acceptance, with a written threshold agreed in advance. ISO/IEC 5259-4 frames data quality as a managed process covering training and evaluation data, including labelling [5], which is a useful anchor when you ask suppliers to document their controls. A workable sequence:

  1. Contract for evidence. Specify which process fields ship with each record, the tool-use policy, and the remedy for records later shown to be model-generated.
  2. Profile the full delivery. Compute timing, paste, uniformity and model-overlap features for every record.
  3. Draw a stratified sample. Over-sample the highest-risk decile by feature score, plus a random slice; acceptance sampling for dataset deliveries covers plan sizing.
  4. Review by humans with evidence. Reviewers see the record, its timing trace and the closest model generation, and classify it as human, model-assisted or model-generated.
  5. Decide at the batch and worker level. One flagged record is noise; a worker or batch with a concentrated pattern is a finding.
  6. Document the result. Record method, sample, rates and decisions in machine-readable form; Croissant-RAI provides a vocabulary for life cycle and labeling documentation [6].

Illustrative example: invented to show structure; it does not describe an available dataset.

Field (per record)Example valueWhat it tells you
record_idpref-000482Join key to text and labels
rater_pseudo_idr-0193Worker-level pattern analysis
guideline_versionv3.2Ties behavior to instructions
ai_assist_policyprohibitedWhat the worker was told
task_open_ts / submit_ts2026-03-04T14:02:10Z / 14:09:55ZTime on task (7m45s)
keystrokes / final_chars2,140 / 1,610Edit ratio, about 1.33
paste_events / max_paste_chars1 / 38Small paste, likely a quote
focus_changes3Tab switching during task
max_model_overlap0.21Low similarity to reference generations
review_outcomehumanReviewer classification

Set the rejection rule in advance, for example: reject a batch if the reviewed model-generated rate in the random slice exceeds an agreed threshold, and exclude all records from any worker with confirmed undisclosed model use.

How this differs from synthetic-data and quality checks

Authenticity checking asks whether a record is what it was sold as, which is separate from whether it is accurate, clean or uncontaminated. Synthetic data sold as synthetic needs due diligence on generation and model terms. Reasoning traces raise their own human-versus-model questions, covered in human-written versus model-generated reasoning trace datasets. For pre-purchase authorship questions on long-form text, see verifying authorship before you buy, and for the wider framework start at the data quality assessment hub. Definitions of the category live on the human data glossary entry and the human-generated data overview.

Operational data is one structural answer to the problem: records created in the normal course of business, with system-of-record timestamps, carry process provenance that task-based collection has to build deliberately. SourceX sources operational datasets from US companies on request, such as support and sales histories, engineering records, documents and finance and legal workflows, plus new recordings of hands-on work. Each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery; buyers can describe the data they need.

Sourcing human-originated data for post-training

SourceX finds US businesses holding the operational data you describe and manages licensing, with every release approved by the supplying company. Personal details are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. A request does not guarantee a match; start by describing your data needs to SourceX.

Frequently asked questions

Should we reject records that a detector flags as AI-generated?

No. Detector errors vary with author characteristics [3], and false positives tend to land on plain-style and non-native writers you may have recruited on purpose. Use the score to prioritize human review alongside timing and paste evidence.

Is light LLM editing by annotators acceptable?

It depends on what you contracted for. Decide in advance whether grammar fixes, drafting help or no assistance is allowed, require disclosure per record, and price model-assisted records differently if you accept them.

What if a supplier has no keystroke or paste logs?

Fall back to creation dates, worker-level statistical signals, model-overlap testing and a larger review sample. For future collection, make process logging a delivery requirement.

Sources

  1. Veselovsky, Horta Ribeiro, West (EPFL), arXiv, "Artificial Artificial Artificial Intelligence: Crowd Workers Widely Use Large Language Models for Text Production Tasks" (2023). https://arxiv.org/pdf/2306.07899
  2. Veselovsky et al., arXiv, "Prevalence and prevention of large language model use in crowd work" (2023). https://ar5iv.labs.arxiv.org/html/2310.15683
  3. arXiv, "Who Writes What: Unveiling the Impact of Author Roles on AI-generated Text Detection" (2025). https://arxiv.org/pdf/2502.12611
  4. Ouyang et al. (OpenAI), arXiv, "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  5. ISO/IEC JTC 1/SC 42, "ISO/IEC 5259-4:2024 Data quality for analytics and machine learning - Part 4: Data quality process framework" (2024). https://www.iso.org/standard/81093.html
  6. Jain et al. (MLCommons), arXiv, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  7. Liang et al., arXiv (2304.02819), "GPT detectors are biased against non-native English writers" (2023). https://arxiv.org/abs/2304.02819

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data