Skip to content

Consulting and recruiting

Can an agency use past survey responses to train AI or build synthetic panels?

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

An agency can use past survey responses to train AI or build a synthetic panel only when four gates are cleared: the consent and notice respondents saw covered the use, client contracts allow it, the data is de-identified to a defensible standard, and professional code duties are met. A study that fails any gate stays out of scope.

Key takeaways

  • Consent is judged by the notice each respondent saw at the time, not by today's privacy policy.
  • Client-supplied sample and client-funded studies usually need the client's written permission before any AI reuse.
  • Under California law, deidentified data must come with a public no-reidentification commitment and matching contract terms for recipients.
  • An internal model, a synthetic panel product and an outside license are three different uses with three different risk levels.
  • Studies that fail a gate are excluded, not stretched; the remaining archive is often still useful.

Why past survey responses carry more strings than other agency records#

Past survey responses carry promises that most agency records do not. Each study came with a screener, a privacy notice and often panel terms, and many were paid for by a client under a contract that says who controls the data and what happens to it after the project.

Open-ended answers add another layer. Respondents write names, employers, towns, medical conditions and complaints into free-text boxes, and those details survive in exports long after the closed-ended data has been coded. A training set built from verbatims inherits both the consent question and the identification question.

Agency process records are different. Proposals, questionnaire drafts with review comments, fieldwork logs and quality-control notes describe how the agency works. They still need a confidentiality review, but they rarely depend on respondent consent.

The four-gate decision tree#

The four-gate decision tree runs each study, or each group of studies fielded under the same terms, through the same questions in order. A study moves forward only when it clears a gate; when it fails, it is excluded from that use and logged with the reason.

Run the gates once per use, not once per study. A study can clear them for training an internal coding model and still fail them for an outside license.

  • Gate 1, consent scope: did the notice and consent respondents saw cover AI training, model building or product development, or only the named research project?
  • Gate 2, client contract: does the client own or control the data, restrict reuse, or require return or destruction after delivery?
  • Gate 3, de-identification: can direct identifiers and identifying free text be removed to a standard that holds up under the laws that may apply?
  • Gate 4, code and disclosure: does the use fit the professional codes the agency follows, and will outputs be labeled honestly to clients?
  • Result: studies that clear all four gates form the candidate set; everything else stays out until terms change or new consent is collected.

Gate 1: what respondents actually agreed to#

Respondent consent is judged by the notice each person saw when they answered, which means pulling archived screeners and panel terms by date. Today's privacy policy does not reach back to studies fielded under older wording.

FTC staff have warned that adopting more permissive data practices, such as using consumer data for AI training, and announcing them only through a quiet, retroactive change to terms or a privacy policy may be unfair or deceptive. New consent language can cover future studies; it does not repair past ones.

Gate 1: what respondents actually agreed to
Where the respondents came fromWho set the consent termsTypical position on AI reuse
Agency-owned panelThe agency, through panel terms and study noticesPossible if the terms covered model building or product development; check the wording by date
Client customer listThe client, under its own privacy noticeGenerally needs client permission and may need fresh notice to respondents
Third-party panel providerThe provider, under its member termsCheck the sample contract for reuse limits beyond the commissioned study
Intercept or river sampleThe agency, often with thin noticeWeakest footing; usually excluded

Gate 2: what the client contract allows#

The client contract decides whether a client-funded study can be reused at all. Look for clauses on ownership of data and deliverables, confidentiality, use restrictions, return or destruction at project end, and any carve-out letting the agency keep aggregated data for norms or benchmarks.

A norms carve-out is not automatically an AI-training right. Language allowing aggregated, anonymized results in benchmark databases was usually written with averages and indices in mind, not respondent-level records feeding a model. Where a clause is ambiguous, written client permission removes the doubt.

Some newer master service agreements also add AI-use clauses that bar the agency from putting client data into AI tools or training sets. Those clauses can reach studies already delivered, so check amendments as well as the original agreement.

Gate 3: is de-identified data enough?#

De-identified data clears the third gate only when the method holds up under the laws that may apply. California's privacy law, for example, treats information as deidentified only if the business takes reasonable measures against reidentification, publicly commits not to reidentify it, and contractually obliges recipients to the same terms.

For respondents in Europe, GDPR treats pseudonymized data that could be re-linked to a person with additional information as personal data. Replacing respondent IDs with new codes while keeping a lookup table does not take a study out of scope.

Synthetic panels raise a related problem. A model trained closely on a small study can reproduce rare combinations of answers, so the training data still needs de-identification and the outputs need testing for memorized records.

Three uses, three levels of risk#

Training an internal model, launching a synthetic panel product and licensing data to an outside AI developer are three different uses, and the gates apply more strictly as data moves further from the original research context.

Developers carry their own documentation duties. California's training-data transparency law requires developers of generative AI systems to publish documentation stating, among other things, whether training datasets were purchased or licensed and whether they include personal information, so expect a licensee to ask for exactly that record.

Three uses, three levels of risk
UseWho sees the data or outputsMain risksMinimum before proceeding
Internal model, such as open-end codingAgency staff and the tool vendorVendor training rights, scope creep into client workConsent scope check, vendor no-training terms, documented purpose
Synthetic panel sold to clientsClients receive model outputsMemorized records, undisclosed synthetic data, client data reusedAll four gates, validation against human data, labeling in reports
License to an AI developerAn outside company receives recordsSale definitions under privacy laws, client confidentiality, reidentificationAll four gates, counsel review, recipient no-reidentification terms

Illustrative: a consumer insights agency sorts its archive#

Illustrative: a fictional consumer insights agency with about 180 employees holds years of concept tests and usage-and-attitude studies. It wants to launch a synthetic panel for early concept screening and is also approached about licensing its archive.

The team groups studies by consent wording and sample source. Studies fielded on its own panel after it added model-building language to its notice pass Gate 1. Studies built on client customer lists fail Gate 2 unless the client signs off, and two clients do. Open-ends go through redaction, and the synthetic model is tested for reproduced verbatims before launch.

The outside license is narrowed to the agency's process records: questionnaire templates, review comments and fieldwork quality logs. Respondent-level data stays inside the agency.

How SourceX handles survey archives#

SourceX runs survey archives through the Rights and Preparation steps of the SourceX five-step transaction before anything is offered to a buyer. The gate results become part of the SourceX Evidence Packet, so the permitted use and privacy record for each included study are written down, and the agency approves every step from scoping to delivery.

Frequently asked questions

Can we ask past respondents for new consent?

Yes, where you still have a lawful way to contact them, usually through an active panel. Re-consent should explain the new use plainly, be optional and be recorded per person. Respondents who do not answer stay excluded. Re-consent rarely works for one-off studies or client-supplied sample, where the agency has no ongoing relationship with respondents.

Does aggregating results into norms remove the consent question?

Aggregated norms, such as average scores by category, carry far less risk than respondent-level data because individuals cannot be seen in them. They do not help train a model that needs individual answers, though, and the client contract still decides whether client-funded results can enter a norms database at all.

Are synthetic respondents personal data?

Synthetic outputs are not automatically free of personal data. If a model reproduces a real respondent's answers or a rare combination of traits, those outputs can relate to an identifiable person. Test outputs against the training set and keep the de-identification record for the source data alongside the model.

What about B2B studies of IT buyers or executives?

B2B respondents are still people with privacy rights, and in some states their data is covered even when they answer in a work capacity. Professional verbatims often name employers, vendors and projects, which makes them easier to identify than consumer answers. Put them through the same four gates.

Who should sign off inside the agency?

Usually the CEO or managing director, with the compliance lead and outside counsel for anything leaving the agency. Client relationship owners should confirm Gate 2 for their accounts, since they know which clients have added AI-use restrictions since the original contract was signed.

Sources

  • FTC staff warned on February 13, 2024 that adopting more permissive data practices, such as using consumer data for AI training, through a surreptitious, retroactive change to terms or a privacy policy may be unfair or deceptive. Source
  • Under Cal. Civ. Code 1798.140(m), deidentified information requires reasonable measures against reidentification, a public commitment not to reidentify, and contractual obligations on recipients. Source
  • GDPR Recital 26 states that personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information, should be considered information on an identifiable natural person. Source
  • California AB 2013 requires developers of generative AI systems to post training-data documentation stating whether datasets were purchased or licensed and whether they include copyrighted material or personal information. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify