Skip to content

Industry-specific operational data

Newsroom fact-checking and corrections records as AI evaluation data

Quick answer

Newsroom fact-check logs and published corrections give factuality teams something public benchmarks rarely do: real claims that professional writers got wrong, the checker query that caught the problem, the evidence type consulted and the final verdict. Scoped well, a record links the drafted claim, the query, the outcome (verified, changed or cut) and the published text, plus correction notices with before and after wording. Buyers should exclude source identities and legally sensitive corrections, and confirm the publisher's rights over checker notes before training.

By SourceX Editorial · Updated

Why human-caught newsroom errors differ from benchmark claims

Newsroom errors are naturally occurring mistakes, which makes them harder and more realistic negatives than the constructed claims in most open fact-verification sets. Widely used verification benchmarks such as FEVER build claims by having annotators rewrite encyclopedia sentences and then label them against evidence [6]. MiniCheck reaches strong grounding-check results with small models trained on GPT-4-generated synthetic errors and evaluated on LLM-AggreFact [1]. Both are useful, but neither shows the error distribution of a reporter on deadline: transposed figures, wrong titles, stale job roles, misattributed quotes and dates shifted by a time zone.

That distribution is the hypothesis buyers are paying for. A hallucination detector tuned on synthetic perturbations may learn the artifacts of the perturbation generator rather than the shape of plausible human error. Checker logs also record claims that turned out to be correct, which gives you calibrated positives drawn from the same drafts, not a separate easy pool.

The record unit: from drafted claim to published correction

The most useful unit is one claim with its full verification trail, not one article. Newsroom fact-checking typically happens in a shared draft (Google Docs comments, a CMS such as WordPress VIP or Arc XP with editorial notes, or a spreadsheet the checker keeps per story), so the raw material is usually comments anchored to text spans. Published corrections live separately: an appended correction notice, a corrections page or a CMS revision history.

Buyers should ask for two linked record types:

  • Pre-publication check records: claim as drafted, claim span offsets, checker query text, evidence consulted described by source type (public record, document provided by subject, interview recheck, prior reporting), outcome code and final published text.
  • Post-publication correction records: original published sentence, corrected sentence, correction notice text, publication and correction timestamps, error type and how the error was surfaced (reader email, subject complaint, internal review).

Some publishers already mark up public fact-check articles with schema.org ClaimReview (claimReviewed, reviewRating, itemReviewed) [7], and rating scales vary by publisher. That markup is a useful label scaffold for public fact-check verdicts, but it does not capture internal checker notes, so treat it as one input rather than the dataset.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "record_id": "chk-000417",
  "record_type": "pre_publication_check",
  "desk": "business",
  "draft_claim": "The plant employed 1,400 people before it closed in 2019.",
  "claim_span": [212, 271],
  "checker_query": "Company filing says 1,140 at closure. Where does 1,400 come from?",
  "evidence_types": ["regulatory_filing", "prior_reporting"],
  "outcome": "changed",
  "error_type": "numeric_transposition",
  "final_text": "The plant employed about 1,140 people when it closed in 2019.",
  "source_identity_removed": true,
  "draft_timestamp": "2024-03-11T14:02:00-05:00",
  "publish_timestamp": "2024-03-12T06:00:00-05:00"
}

A matching correction record would add original_published_text, corrected_text, correction_notice, correction_timestamp and surfaced_by.

Rights and sensitivity limits specific to newsrooms

Newsroom records carry three limits that generic editing data does not: source protection, legal exposure and contributor copyright. Checker notes routinely name confidential sources or quote unpublished allegations, so buyers should require that source identities, contact details and unpublished claims about named people are removed, keeping only the evidence type. Corrections tied to defamation demands, retraction requests or settlements may be confidential under the settlement terms; ask for them to be excluded unless the publisher's counsel clears them.

Copyright in the notes themselves also matters. Freelance writers and contract checkers may own their contributions unless their agreements assign them, and the Third Circuit held on 29 September 2026 that Westlaw editorial headnotes were sufficiently original to be copyrightable [5]. That ruling concerned a non-generative tool, but it is a reminder that editorial annotations can be protected work product, so confirm the publisher holds rights in checker comments, not only in published articles.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Open alternatives and their license limits

Open factuality and expert-text datasets exist, but many cannot be used commercially, and license metadata is often wrong. Meta's LCFO, an expert-written summarization set, is licensed CC-BY-NC 4.0 [2]. The Data Provenance Initiative audited more than 1,800 text datasets and found license omission above 70% and license error rates above 50% on popular hosting sites [3]. For scholarly retractions rather than news, Crossref announced in 2023 that it had acquired the Retraction Watch database and is making it openly available [4], which suits retraction-reason classification but not newsroom claim checking.

The practical split: use open sets for baselines and ablations, and license newsroom records when you need commercial training rights or a held-out eval that has not leaked into pretraining corpora. Pre-publication checker notes, unlike published corrections, were never on the open web, which makes them strong contamination-resistant test material. For time-based holdouts, see post-cutoff evaluation data.

Buyer scoping checklist for fact-check and corrections data

A tight request names the record unit, the label set and the exclusions before any sample is reviewed.

Illustrative example: invented to show structure; it does not describe an available dataset.

Scoping itemWhat to specifyCommon failure if skipped
Record unitClaim-level, with span offsets into the draftArticle-level dumps where errors cannot be aligned
Outcome labelsverified / changed / cut, plus a mapping documentFree-text outcomes that need relabeling
Error taxonomyNumeric, name or title, date, attribution, causal, quoteLabels inferred later by an LLM, adding noise
Evidence fieldSource type only, never source identityConfidential sources exposed in training data
Correction linkageOriginal, corrected, notice, both timestampsCorrections that cannot be matched to the claim
Desk and beatBusiness, politics, health, sports, localEval set dominated by one desk's error profile
ExclusionsLegal holds, settlements, unpublished allegationsSensitive material reaching a vendor or model
BoilerplateStrip CMS templates and standard notice wordingModels learning notice templates, not errors
RightsPublisher rights in checker notes and freelance workUsable articles but unusable annotations

Template and notice wording repeats heavily in corrections feeds; detecting boilerplate in business records covers deduplication. For holdout design, see building a golden evaluation dataset from business records, and for broader acceptance criteria see training data quality assessment.

How SourceX approaches newsroom fact-check requests

SourceX sources operational datasets from US companies on request, including documents, and manages the licensing and ongoing purchases; it does not hold stock, and a request does not guarantee a match. Buyers describe the data they need, not specific publishers, and every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery.

Names, emails, phone numbers and account numbers are removed or replaced before delivery, the method is recorded and a sample is checked, though no method is perfect. SourceX does not source scraped web content, so public fact-check pages crawled from the web are out of scope. Media-specific context is on the media and publishing buyers page, and the supply side is discussed in can media and publishing companies sell data to AI companies. To start, describe your fact-check data needs to SourceX.

Related: evaluation datasets built from real business work, enterprise document datasets and the industry-specific operational data hub, part of the wider AI data buyer guides.

Find newsroom fact-check and corrections data

SourceX looks for US businesses that hold the records you describe, assesses data and licensing permissions, and agrees pricing and allowed uses in a license before anything is delivered; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement. Describe your claim-verification data request.

Frequently asked questions

Are published corrections enough without checker notes?

Corrections alone show what escaped checking, which skews toward subtle errors and post-publication complaints. Checker notes add the much larger set of errors caught before publication and the claims confirmed correct, which you need for precision and recall estimates.

Can ClaimReview pages substitute for internal fact-check logs?

ClaimReview pages describe verdicts on claims made by others, such as politicians, with publisher-specific rating scales. They suit claim-verification on public statements, not detection of errors in a writer's own drafts.

How should evidence be represented if sources are removed?

Keep a controlled source-type field (filing, dataset, document from subject, interview recheck, prior reporting) and, where permitted, links to public records only. That preserves the retrieval signal without revealing who supplied the information.

Sources

  1. Tang, Laban and Durrett (arXiv), "MiniCheck: Efficient Fact-Checking of LLMs on Grounding Documents" (2024). https://arxiv.org/html/2404.10774v2
  2. Meta (Hugging Face), "facebook/LCFO dataset card (README)". https://huggingface.co/datasets/facebook/LCFO/blob/9109c9f66e74a0b78b82aba2b54ffeecf3b57c44/README.md?code=true
  3. Longpre et al. (arXiv), "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
  4. Crossref, "News: Crossref and Retraction Watch" (2023). https://crossref.org/blog/news-crossref-and-retraction-watch
  5. U.S. Court of Appeals for the Third Circuit, "Thomson Reuters Enterprise Centre GmbH v. ROSS Intelligence Inc., No. 25-2153 (3d Cir.)" (2026). https://www2.ca3.uscourts.gov/opinarch/252153p.pdf
  6. Thorne et al. (arXiv), "FEVER: a large-scale dataset for Fact Extraction and VERification" (2018). https://arxiv.org/abs/1803.05355
  7. Schema.org, "ClaimReview Schema Definition". https://schema.org/ClaimReview

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data