Skip to content

Text and language data

Q&A Threads With Accepted Answers and Votes as Training Signal

Quick answer

Q&A threads with accepted-answer flags and vote scores give post-training teams three things at once: question-answer pairs for supervised fine-tuning, ranked answers to the same question for preference and reward modeling, and question-to-answer relevance labels for retrieval evaluation. The signal is only as good as its metadata. Require the full thread (every answer, not only the winner), per-answer scores, the accepted flag, timestamps, edit history, tags and a deletion feed, and treat votes as noisy, biased human judgments rather than ground truth.

By SourceX Editorial · Updated

What a usable Q&A thread record contains

A usable record is the whole thread with its social metadata attached, because the training value sits in the comparison between answers, not in any single answer. A file of "best answer" pairs alone supports SFT but throws away the preference signal and hides how the winner was chosen. Ask for the question, every non-deleted answer, the comments if they are licensed, and per-post metadata at the time of export.

Public community dumps have long used a relational layout that is a reasonable reference even for private or enterprise Q&A systems: a posts table where a type field separates questions from answers, a parent ID linking answers to their question, an accepted-answer ID on the question, a net score, view counts, tags, creation and last-edit timestamps, and separate tables for votes, post history and links between posts. Internal knowledge platforms, customer support communities and product forums usually hold the same concepts under different names, so the first diligence task is a field mapping, not a format conversion.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "thread_id": "q-48213",
  "source_system": "internal_eng_qa",
  "snapshot_date": "2026-08-31",
  "question": {
    "title": "Retry storm after upgrading the payments client to v4",
    "body_markdown": "...",
    "tags": ["payments-client", "retries", "kubernetes"],
    "created_at": "2025-11-03T14:22:10Z",
    "last_edited_at": "2025-11-04T09:01:44Z",
    "author_reputation_band": "1k-10k",
    "accepted_answer_id": "a-90551",
    "accepted_at": "2025-11-05T16:40:02Z",
    "closed_reason": null,
    "duplicate_of": null
  },
  "answers": [
    {"answer_id": "a-90551", "score": 14, "upvotes": 15, "downvotes": 1, "is_accepted": true,
     "created_at": "2025-11-03T15:10:00Z", "edit_count": 2, "author_reputation_band": "10k+", "is_self_answer": false},
    {"answer_id": "a-90560", "score": 3, "upvotes": 4, "downvotes": 1, "is_accepted": false,
     "created_at": "2025-11-03T18:47:31Z", "edit_count": 0, "author_reputation_band": "100-1k", "is_self_answer": false}
  ],
  "rights": {"contributor_license": "employee work product", "removal_feed": true},
  "deidentification": {"method": "names and emails replaced with tokens", "sample_checked": true}
}

Two design choices in that record matter more than they look. Reputation is shipped as a band rather than a raw number or user ID, which keeps a useful quality feature while reducing re-identification risk. Upvotes and downvotes are split, because a score of 3 from 4 up and 1 down is a different signal from 3 from 40 up and 37 down. For a broader field list across text corpora, see metadata fields to require with licensed text corpora.

How accepted answers feed supervised fine-tuning

Accepted answers make good SFT targets when you filter them, because acceptance records that one asker was satisfied, not that the answer is correct, current or well written. The LIMA work showed that a small set of carefully curated prompt-response pairs can carry much of the alignment effect, and its authors noted the curation was labor-intensive [5]. Community Q&A is attractive for exactly that reason: the curation signal is already partly there.

Practical SFT filters for Q&A threads:

  • Score floor plus acceptance. Keep accepted answers that also clear a minimum net score, so a single asker's click is corroborated by other readers.
  • Prefer the top-scored answer when it differs from the accepted one. Askers often accept the first answer that unblocks them; a later answer can overtake it in score.
  • Drop link-only and "see docs" answers. These train the model to deflect. A minimum body length and a ratio of prose to URLs catch most of them.
  • Strip or rewrite conversational residue. Signatures, "thanks, this worked", and edit notes such as "EDIT: fixed typo" are noise in a response target.
  • Respect closure and duplicate state. Closed and duplicate-marked questions are weak SFT prompts but useful for retrieval evaluation (see below).
  • Date-gate by technology version. An accepted 2019 answer about a library API may be wrong for the current release; carry the timestamp so you can filter or weight by recency.

Building preference pairs from votes and acceptance

The standard construction is a pair of answers to the same question, labeled chosen and rejected, where the chosen answer has a meaningfully higher score or the accepted flag. Pairs drawn within one thread control for the prompt, which is what preference objectives need. The InstructGPT recipe trained a reward model on human rankings of several outputs for the same prompt, and a multi-answer thread with scores is structurally the same object: one prompt, several ranked responses [4]. Direct Preference Optimization trains a policy directly on chosen and rejected pairs without a separately trained reward model [6], so the same Q&A pairs can serve either a reward model or a DPO run. Meta's Llama 2 reward models combined in-house binary comparisons with open-source preference datasets, a published example of vote-derived pairs serving as one input among several rather than the whole preference budget [7].

Illustrative example: invented to show structure; it does not describe an available dataset.

Pair ruleChosenRejectedKeep whenMain risk
Accepted vs. low-scoredAccepted answerAnswer with score at least N belowAccepted also has positive scoreAsker accepted an outdated or partial fix
Score marginHigher-scored answerLower-scored answerMargin and total votes both above thresholdsEarly answers accumulate votes by exposure
Score ratioHigher scoreLower scoreRatio test plus minimum vote countSmall-count threads give unstable ratios
Accepted vs. top-scored disagreementTop-scoredAcceptedLarge score gap, later timestampEncodes "popular" over "worked for asker"
Negative-score answerAny positive answerAnswer with net negative scoreDownvotes present, not only lack of upvotesDownvotes may reflect tone, not correctness

Thresholds should be tuned on a held-out set where reviewers judge pairs blind; report the agreement rate between vote-derived labels and reviewer labels before training. Teams building a reward model from this data often also cap the number of pairs per thread so a few very popular questions do not dominate.

Known biases in vote and acceptance signals

Votes are a crowd signal with predictable distortions, and a reward model trained on them will learn those distortions unless you correct for them. The most common failure modes:

  • Position and exposure bias. Answers posted early get more views and therefore more votes. Normalize by time since the question was posted, or compare only answers posted within a similar window.
  • Length and formatting preference. Longer answers with code blocks and headings attract votes. Without length controls, the reward model learns verbosity. Match pairs on length bands or add length as a covariate when evaluating.
  • Popularity of the question. A score of 5 on a question with 200 views means more than 5 on one with 200,000. Ship view counts so you can compute per-view rates.
  • Asker-only acceptance. The accepted flag is one person's judgment, often made before better answers arrive, and in many systems the asker can accept their own answer. Carry an is_self_answer flag.
  • Author reputation halo. High-reputation authors attract votes independent of content. Using a reputation band as a feature is fine; letting it leak into labels is not.
  • Staleness. Votes accumulate over years while the underlying technology changes. A highly voted answer can be confidently wrong today.

Using Q&A threads for retrieval evaluation

Q&A threads give retrieval evaluation a natural query and a natural relevance label: the question is the query, the accepted or top-scored answer is a positive passage, and other answers to the same question are graded positives or hard negatives depending on score. Duplicate links between questions add a second label type, since two questions marked as duplicates should retrieve each other's answers. That makes duplicate and related-post links worth requiring even if you never use them for training. For the corpus-side fields a RAG system needs, see metadata a licensed RAG corpus should ship with.

Keep evaluation threads out of training data by thread ID and by near-duplicate question text, not by random row split. Q&A content is heavily duplicated across mirrors, so contamination checks should run against your pre-training corpus as well.

Rights, removals and freshness

Rights in community Q&A are layered, and the contributor license is not the whole story. Stack Exchange contributions are licensed under CC BY-SA (the version depends on when a post was made), yet in July 2024 access to the data dump was restricted with terms aimed at AI training, which critics argued conflicts with the license's reuse permissions [1]. Platform operators have also licensed structured Q&A to AI providers directly, which shows that buyers pay for structured, API-delivered Q&A rather than only scraping it [2]. For the full rights analysis, see licensing forum and community content for AI training.

Plan for removals from the start. When a platform announced an AI licensing deal, some contributors edited or deleted their posts in protest and were suspended [3]. A buyer should require a removal or tombstone feed keyed by post ID, a contractual process for applying it to retained copies, and a snapshot date on every record. Private enterprise Q&A, such as internal engineering forums or support communities, avoids public-contributor objections but brings its own questions: whether employees or customers wrote the content, what the platform terms say, and whether personal details in bodies and comments have been removed.

Freshness is a separate risk. If a community's new-question volume declines, the date range of the corpus matters more; record the earliest and latest question dates and the distribution by year so you can see whether the data reflects current tools.

Supplier checklist for Q&A thread datasets

Use this checklist when a supplier offers Q&A data, and attach it to the request so answers come back in a comparable form.

Illustrative example: invented to show structure; it does not describe an available dataset.

  • Full threads: question, all surviving answers, optional comments, with stable IDs.
  • Per-answer score, and split upvotes and downvotes where the system records them.
  • Accepted flag, acceptance timestamp and self-answer flag.
  • Creation, last-edit and snapshot timestamps; edit history or at least edit count.
  • Tags or categories with the tag vocabulary.
  • View counts per question.
  • Duplicate and related-post links.
  • Author reputation as bands, not user IDs; no profile data.
  • Close and deletion state, plus a removal feed for later deletions.
  • Contributor license and platform terms in force when content was posted.
  • De-identification method for names, emails and account numbers in bodies, with a sample check.
  • A dataset card covering source system, date range, filters applied and known biases [8].

Where SourceX fits in sourcing Q&A data

SourceX sources operational datasets from US companies on request, and those kinds of data include support and sales histories, engineering records and documents, which is where private Q&A threads such as internal engineering forums and support communities tend to live. It does not source scraped web content, and data is not held in stock, so a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents and delivered under a license that defines records, uses, term and delivery, with personal details removed or replaced before delivery and the method recorded. Buyers describe the data they need through the SourceX buyer intake, not the businesses that might hold it.

Related reading in this cluster: the text datasets hub, proprietary text data beyond web crawls, the glossary entry on preference data, and licensing knowledge base articles for AI training.

Request Q&A thread data with accepted answers and votes

If your post-training or retrieval work needs licensed Q&A threads with acceptance and vote metadata, describe the fields, date range and allowed uses you need. SourceX looks for US businesses that hold matching data, assesses data and licensing permissions, and nothing is contracted until a supplier agrees. Describe your Q&A data requirement to SourceX.

Sources

  1. DevClass, "Stack Exchange restricts access to dump of user-contributed data as critics complain license permits reuse for any purpose" (2024). https://devclass.com/2024/07/30/stack-exchange-restricts-access-to-dump-of-user-contributed-data-as-critics-complain-license-permits-reuse-for-any-purpose
  2. The National CIO Review, "Google's Strategic Partnership with Stack Overflow: A New Era of AI Data Licensing" (2024). https://nationalcioreview.com/articles-insights/extra-bytes/googles-strategic-partnership-with-stack-overflow-a-new-era-of-ai-data-licensing/
  3. The Register, "Stack Overflow banning users who protest AI deal by editing or deleting posts" (2024). https://www.theregister.com/2024/05/09/stack_overflow_banning_users_who/
  4. Ouyang et al. (OpenAI), "Training language models to follow instructions with human feedback" (2022). https://arxiv.org/pdf/2203.02155
  5. Zhou et al., "LIMA: Less Is More for Alignment" (2023). https://arxiv.org/pdf/2305.11206
  6. Rafailov et al., "Direct Preference Optimization: Your Language Model is Secretly a Reward Model" (2023). https://arxiv.org/abs/2305.18290v1
  7. Touvron et al. (Meta), "Llama 2: Open Foundation and Fine-Tuned Chat Models" (2023). https://arxiv.org/pdf/2307.09288
  8. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data