Skip to content

Consulting and recruiting

Synthetic respondents: what they mean for market research agencies

By SourceX Editorial · Updated

Short answer

Synthetic respondents are AI-generated survey answers from models conditioned on personas or trained on past research. For market research agencies they work best for early exploration, such as screening concepts or testing questionnaires, while human samples remain the standard for validation and decisions. Agencies with consented, well-documented human data hold the input synthetic models depend on.

Key takeaways

  • Synthetic respondents come from three build methods: prompted general models, models tuned on research data and statistical synthesis of a real dataset.
  • Use synthetic output to narrow options early; use human samples to validate, size and track.
  • Synthetic results should be labeled in reports, with the validation method described.
  • A synthetic panel trained on past studies inherits those studies' consent and client-contract limits.
  • Consented, documented human data gains importance as synthetic tools spread, because every credible model needs it for training and testing.

What are synthetic respondents?#

Synthetic respondents are simulated survey participants whose answers come from an AI model rather than a person. A vendor or agency describes a target audience, the model answers the questionnaire as members of that audience might, and the output arrives as a dataset that looks like fieldwork.

The term covers very different products. Some are general language models prompted with a persona description. Others are tuned on years of real survey data in one category. A third kind is synthetic data in the statistical sense: a generated copy of a real dataset designed to keep its patterns while protecting the people in it.

Inputs shape outputs. A persona described only by age, gender and region produces generic answers, while one conditioned on category purchase history and past attitudes behaves more like a real segment, provided the data behind it is current and was collected with consent that covers the use.

How synthetic respondents are built#

Synthetic respondents are built in three main ways, and the build method predicts where they will be reliable. Ask any vendor which method it uses and what data it was trained or conditioned on.

The third method has the longest history in privacy practice. NIST guidance for government agencies lists publishing synthetic data as one of several data-sharing models, alongside de-identified data, query interfaces and protected enclaves. That use protects a real dataset, which is a different job from asking a model to predict opinions nobody has expressed yet.

How synthetic respondents are built
MethodHow it worksStrengthWeakness
Prompted general modelA large language model answers as a described personaFast and cheap for brainstormingReflects the model's general training, not your category or market
Model tuned on research dataA model is adjusted using past survey responses in a categoryCloser to real answer patterns in that categoryOnly as good, current and consented as the studies behind it
Statistical synthesisA generator reproduces the structure of a specific real datasetLets teams share or test data with less privacy exposureCannot answer new questions the original data never asked

Where synthetic respondents work and where human samples stay the standard#

Synthetic respondents work where being somewhat wrong is cheap and speed matters. Human samples stay the standard where a client will commit budget, launch a product or publish a claim based on the numbers.

The dividing line is the cost of error, not the topic. The same concept can go through a synthetic screen early and a human validation later, and the report should say which result came from which source.

Where synthetic respondents work and where human samples stay the standard
Use caseSynthetic fitHuman sample role
Drafting and stress-testing a questionnaireGood: finds confusing wording and missing answer optionsPilot with real respondents before launch
Early screening of many conceptsReasonable for narrowing a long listValidate the shortlist with real sample
Hypotheses for qualitative workUseful for discussion guides and probesInterviews and groups supply the evidence
Hard-to-reach or niche audiencesWeak: least training data where it is most neededRecruit real respondents, even in small numbers
Market sizing, pricing and forecastingPoor: errors compound into business casesRequired
Claims substantiation and published statisticsNot appropriateRequired, with a documented method
Tracking change over timePoor: models lag real shifts in opinionRequired

Questions to ask a synthetic data vendor#

Questions to ask a synthetic data vendor should test the method, the training data and the evidence of accuracy. A vendor that cannot answer them in writing is asking your clients to take results on trust.

Keep the written answers on file with every project that used the tool, next to your own validation notes. When a client later asks how a concept screen was produced, that file is the record, and it is also what a future buyer or licensee will ask to see.

  • Which build method does the product use, and which base model sits underneath it?
  • What data was the model tuned or conditioned on, from which years and markets, and under what consent?
  • Did any of that data come from client-commissioned studies, and with whose permission?
  • How was accuracy tested: against which human studies, on which question types, and where did it fail?
  • Can the model reproduce real respondents' verbatims, and how is that tested?
  • How are outputs labeled, and can we show the method to our clients?
  • Who owns outputs and any tuning done on our data, and may the vendor reuse our inputs?

What synthetic respondents mean for an agency's business#

Synthetic respondents put pressure on fast, low-stakes work first: quick-turn concept screens, simple screeners and early exploration that once justified a small study. Clients may expect that work faster and cheaper whether or not the agency adopts synthetic tools itself.

They also raise the importance of what synthetic tools lack. Real, recent, consented human data in a category, with clean records of how it was collected, is the input any credible synthetic model needs and the benchmark used to test one. An agency with a well-kept archive is closer to being a supplier of that input than a casualty of the output.

That only holds if the rights are clean. A synthetic panel or a license built on studies whose consent never covered model building, or on client-owned sample, inherits those problems.

Illustrative: a packaged-goods agency adds a synthetic stage#

Illustrative: a fictional agency serving packaged-goods brands, with about 120 employees, starts losing quick concept screens to a client's in-house synthetic tool. It decides to add its own synthetic stage at the front of its concept testing process rather than compete on price for screens alone.

The agency tunes a model only on studies from its own panel fielded under notices that covered model development, and excludes client-list sample. It tests the model against recent human studies, shares the comparison with clients and labels every synthetic result. Shortlisted concepts always go to human sample for validation.

Clients get faster early answers, and the agency keeps the validation work and gains a documented method it can explain in pitches.

How SourceX views research archives in a synthetic market#

SourceX views a research archive the way the SourceX Enterprise Data Value Framework does: human-generated signal, domain expertise, recency and rights increase value, while reproducibility reduces it. Synthetic output is reproducible by design; consented human responses and an agency's own process records are not.

Buyers carry their own paperwork too. Under California's AB 2013, developers of generative AI systems must post documentation about their training data, including whether datasets were bought or licensed and whether personal information is included. Any license SourceX arranges therefore begins with rights review and privacy preparation, the second and third stages of the SourceX five-step transaction, and the agency approves every step.

Frequently asked questions

Will synthetic data replace surveys?

Not for decisions that need evidence about real people. Synthetic data can speed up early stages and replace some small studies, but models learn from past human data and drift as opinions change. Validation, tracking and claims work still need human samples, and synthetic tools need fresh human data to stay credible.

Do we have to tell clients when results are synthetic?

Yes, as a matter of honest reporting and under research codes that require transparency about methods. Label synthetic results in the deliverable, describe how the model was built and validated, and never blend synthetic cases into human sample counts without saying so.

Can a synthetic panel trained on client studies be sold to other clients?

Only if the client contracts and respondent consent for those studies allow it. A model tuned on one client's commissioned research can carry that client's confidential findings into answers for a competitor. Most agencies should exclude client-funded studies unless the client has agreed in writing.

Is synthetic data a privacy fix for sharing our archive?

It can help, but it is not automatic. Statistical synthesis can reduce exposure when sharing a dataset, yet a generator fitted tightly to a small dataset may echo unusual answer patterns that point back to real people. Test outputs for memorized records and keep the source data's de-identification record.

Should we build our own synthetic tool or work with a vendor?

It depends on your archive and team. Agencies with deep, consented category data and analytics staff can tune their own models and keep the method in-house. Others may use a vendor, but should read its terms on training with agency inputs and ownership of tuned models before sending any data.

Sources

  • NIST SP 800-188 tells agencies to choose a data-sharing model such as publishing de-identified data, publishing synthetic data, offering a query interface that applies de-identification, or using a protected enclave. Source
  • California AB 2013 requires developers of generative AI systems to post training-data documentation stating whether datasets were purchased or licensed and whether they include copyrighted material or personal information. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify