Industry-specific operational data
Coded open-ended survey responses and codeframes for AI
Quick answer
A useful coded open-ended survey dataset pairs each verbatim with the exact question wording, the project codeframe version (nets and sub-codes), every human code assigned, and evidence of coder agreement. Public alternatives are mostly product reviews with sentiment labels, not codeframes. To train or benchmark an auto-coder, license real project exports from research agencies or insights teams, confirm respondent consent covers secondary AI use, and mask end-client brands before delivery.
By SourceX Editorial · Updated
Why review and sentiment datasets do not substitute for coded verbatims
Review corpora teach polarity, while auto-coding requires assigning a response to codes defined by a researcher for one question in one study. Most public opinion-text sets are product reviews with binary or star-based sentiment, and some of the large ones carry terms that restrict commercial use [1]. Small permissive sets exist, such as a 5,000-train and 1,000-test review set under MIT, but they include no codeframe or net structure [2]. Commercial aspect-based sentiment data is the nearest analogue; aspects are fixed taxonomy labels rather than project codeframes.
The gap matters because the hard parts of verbatim coding sit outside sentiment. They include multi-code assignment ("too expensive and the app crashes" takes two codes), "Other – specify" residuals, net roll-ups, and codes that only make sense against the question stem. A model trained on reviews will not learn when a coder creates a new code mid-field or merges two thin codes into a net.
What a coded verbatim record should contain
Each record should let you reconstruct the coding decision without the original project files. Request these fields, ideally as JSONL with the codeframe shipped as a separate versioned file:
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"project_token": "P-0412",
"project_category": "consumer_packaged_goods/concept_test",
"wave": 3,
"question_id": "Q7",
"question_text": "Why did you rate [BRAND_A] that way?",
"question_type": "open_end_after_rating",
"linked_closed_answer": 4,
"verbatim": "Tastes fine but [BRAND_A] costs more than the store brand",
"language": "en-US",
"codeframe_id": "CF-0412-Q7",
"codeframe_version": "v2.1",
"codes_assigned": ["NET_TASTE.101", "NET_PRICE.204"],
"coder_token": "C-17",
"double_coded": true,
"second_coder_codes": ["NET_PRICE.204"],
"adjudicated_codes": ["NET_TASTE.101", "NET_PRICE.204"],
"coding_method": "manual",
"coded_at": "2025-03-14"
}
The codeframe file should list each code ID, label, parent net, coder instructions and example verbatims, plus the version history: which codes were added, split, merged or retired between waves. If a supplier ran assisted coding (a rules engine or machine suggestions reviewed by a human), the coding_method field must say so, because machine-suggested labels that coders accepted will look more consistent than fully manual work.
How to judge label quality and coder agreement
Ask for double-coded subsets and compute agreement yourself, because a single supplier-reported percent agreement hides multi-label disagreement. Verbatim coding is multi-label with a varying number of codes per response, so per-code Cohen's kappa or Krippendorff's alpha with a set-based distance is more informative than exact-match agreement; the right metric depends on the annotation design [3]. Our guide to choosing inter-annotator agreement metrics walks through the options.
Expect agreement to differ sharply by code. Broad nets such as price or taste agree well, while thin codes, sarcasm and "Other" residuals do not. Automated coders show the same unevenness: a model can score well on aggregate F1 while failing the thin codes that analysts care about. That pattern is the reason to buy human-coded data with per-code agreement rather than a single headline number, and to report per-code precision and recall in every eval.
For an evaluation set, adjudicated labels matter more than volume. Test-set label errors can reorder model rankings [4], so reserve double-coded, adjudicated responses for your held-out split and keep codeframes from the same project out of training.
Codeframe drift across waves and projects
Codeframes are project-specific and evolve, so the codeframe version must be tied to every label. Trackers re-field the same question across waves, and coders add codes as new themes appear (a competitor launch, a product recall). If you train on wave 1 labels and evaluate on wave 4 without mapping, the model is penalized for codes that did not exist yet.
Value cross-project variety over raw verbatim count. For example, a dataset of 300 codeframes from many categories and question types teaches a model to read coder instructions and generalize, which is what an LLM-as-coder or codeframe-conditioned SFT setup needs. A dataset of one tracker with millions of responses mostly teaches that tracker. Our note on code-set revisions in multi-year datasets covers mapping tables between versions.
Consent, research-purpose rules and confidential client content
Respondent data collected for market research carries purpose limits that a buyer must review before AI training. Market research codes of conduct, such as the ICC/ESOMAR International Code, restrict how personal data collected for research may be reused, so have counsel check the exact current wording and how the supplier's privacy notice and panel terms described secondary use. Where the agency fielded the study for an end client, the client contract may also govern who owns verbatims and codeframes.
Three content risks recur in open ends:
- Self-identification: respondents write names, employers, towns, health conditions and phone numbers into free text. Field-level masking misses these; you need text-level detection plus a sampled manual check.
- Confidential client material: brand names, unreleased concepts, ad copy and pricing tested in the study may be the end client's confidential information. Tokenize brands and products consistently (
[BRAND_A]) so the model still learns brand-relative reasoning. - Sensitive categories: healthcare, political and employee surveys can carry health data or opinions that trigger stricter rules; health-related verbatims from covered entities need HIPAA de-identification.
Our guide to de-identifying survey and research data covers masking methods. Treat any "free" coded corpus with the same skepticism: dataset hosting sites often omit or misstate licenses [5], and our open dataset license audit explains how to check them.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Request checklist for coded verbatim data
Use this checklist to scope a request so suppliers can tell whether they hold a match.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Requirement | What to specify | Why it matters |
|---|---|---|
| Question context | Full question text, question type, linked closed-end answer | Codes depend on the stem |
| Codeframe | Code IDs, labels, nets, instructions, version history | Enables codeframe-conditioned training |
| Labels | All codes per response, not just primary | Multi-label task |
| Agreement | Double-coded share, second coder labels, adjudication | Per-code reliability |
| Coding method | Manual, assisted or machine-reviewed | Avoids learning another model's output |
| Coverage | Categories, languages, question types, waves | Generalization |
| Privacy | Masking method, sample check results | Free-text identifiers |
| Client content | Brand and concept tokenization | Confidentiality |
| Rights | Consent basis and agency or client ownership | Secondary AI use |
For broader survey licensing questions, see licensing survey and research data for AI and what makes survey and research data valuable for AI. If you are validating synthetic respondents rather than coders, see real survey data for validating LLM personas; for adjacent free-text labels, see rubric-scored written responses and after-call work notes with disposition codes.
How SourceX handles requests for coded survey data
SourceX sources operational datasets from US companies on request; it holds no stock, and a request does not guarantee a match. You describe the data, such as coded verbatims with codeframes and double-coding, and SourceX looks for US businesses that hold it; every release is approved by the supplying company. Each dataset is rights-reviewed for ownership and consents, personal details such as names, emails and phone numbers are removed or replaced with the method recorded and a sample checked, and no method is perfect. Delivery follows an executed license that defines records, uses, term and delivery. Teams in the market research buyer section can describe a request on the buyers page, and the industry data hub and AI data guides cover related categories.
Source coded open-ended survey responses with SourceX
Describe the verbatims, codeframes and agreement evidence your auto-coding model needs, and SourceX looks for US companies that hold them and manages the licensing. Nothing is contracted until a supplier agrees, and terms are set per deal. Start a coded survey data request.
Sources
- HyperAI, "Amazon product reviews sentiment dataset listing". https://hyper.ai/en/datasets/5481
- Hugging Face (Cleanlab), "Cleanlab/amazon-reviews dataset card". https://huggingface.co/datasets/Cleanlab/amazon-reviews/blob/main/README.md
- arXiv, "Counting on Consensus: Selecting the Right Inter-annotator Agreement Metric for NLP Annotation and Evaluation" (2026). https://arxiv.org/pdf/2603.06865
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- Longpre et al., "The Data Provenance Initiative: A Large Scale Audit of Dataset Licensing & Attribution in AI" (2023). https://arxiv.org/abs/2310.16787
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.