Skip to content

Consulting and recruiting

Market research data retention: how long to keep respondent data

By SourceX Editorial · Reviewed by Noah Loul ·

Short answer

A market research data retention policy should set a separate clock for each type of respondent data rather than one period for everything. Identifying records such as screeners, contact details and recordings go first; de-identified study data can be kept longer where contracts and notices allow; a legal hold pauses every clock until counsel releases it.

Key takeaways

  • Give each data type its own retention clock and its own trigger event, such as fieldwork close or report delivery.
  • Identifying records, including contact details, recordings and identified transcripts, should have the shortest clocks.
  • A legal hold overrides every deletion rule, including automatic deletion settings inside tools.
  • Pseudonymized data with a stored key is still personal data under the GDPR, so it does not earn a longer clock.
  • Records deleted on schedule cannot be licensed later, and records kept beyond policy should not be.

How long should an agency keep respondent data?#

An agency should keep respondent data only as long as each data type has a documented purpose, which means a market research data retention policy needs several clocks, not one. Contact details exist to recruit and pay incentives; raw response files exist to check quality and answer client questions; de-identified study data can support trend analysis and norms for much longer.

The period for each clock comes from three places: the client contract, the privacy notice and consent text respondents saw, and the laws that may apply to the respondents in the study. Where those sources differ, the shortest applicable period usually wins. Write the source next to every period so the next reviewer can see why it was chosen.

Retention schedule template by data type#

Use the template below as a starting schedule. Each row names the record, where it usually lives, the event that starts its clock and the action when the period ends. Add a period column filled from your own contracts and notices; the template leaves periods out on purpose because they differ by client and jurisdiction.

Retention schedule template by data type
Data typeTypical locationClock starts atEnd action
Screener answers and contact detailsRecruitment database, survey platformFieldwork closeDelete; keep a suppression list only if needed
Incentive and payment recordsIncentive platform, finance systemPayment reconciledKeep for accounting rules, then delete
Consent recordsSurvey platform logs, signed formsEnd of the covered data's lifeKeep as long as the data they cover
Raw response files with respondent IDsSurvey platform, processing driveReport deliveryDelete or strip identifiers
Cleaned de-identified datasetsAnalysis drive, warehouseReport deliveryKeep per contract; review on a set cycle
Open-ended verbatimsCoding tool, tabulation filesCoding completeScrub names and identifiers, or delete
Audio and video recordingsZoom, facility portal, file storageSession dateDelete after analysis or client handover
Identified transcriptsTranscription vendor, shared driveAnalysis completeReplace with a de-identified version
De-identified transcriptsProject folderDe-identification sign-offKeep per contract and consent
Questionnaires, stimuli and reportsProject folder, client portalProject closeKeep per contract; stimuli may belong to the client

How do you set the period for each row?#

Set each period by reading the documents that bind the agency, then choosing the shortest period they allow for that data type. Work through them in the same order for every client so the schedule stays consistent and the reasoning is easy to audit.

  • Client MSA and SOW: look for retention, return-or-destroy and audit clauses, and note whether destruction covers derived datasets.
  • Privacy notice and consent text: check what respondents were told about how long data is kept and for which purposes.
  • Sample and panel supplier terms: some suppliers limit how long respondent-level files may be kept.
  • Laws that may apply: the EU GDPR or UK GDPR for respondents in those regions, and US state privacy laws depending on where respondents live, assessed with counsel.
  • Your professional code: check what the research code your agency follows says about keeping personal data.
  • Accounting rules: incentive payment records may need to outlast the respondent data they relate to.

A legal hold suspends deletion for any record that may be relevant to a dispute, investigation or regulatory inquiry, and it overrides every row in the schedule until counsel lifts it. The policy should name who can place a hold, how it is recorded and which systems it reaches.

The weak point is automatic deletion inside tools, which runs regardless of the written policy. Zoom, for example, lets admins delete cloud recordings after a specified number of days counted from each recording's creation, with an option to exempt individual recordings. When a hold starts, check those settings in every affected tool and exempt the held files.

Keep a simple hold log: matter, scope, systems, date placed, approver and date released. When a hold ends, apply the normal clocks again: delete what has passed its period and leave the rest on schedule.

Retention and value: what de-identification changes#

De-identification changes which clock applies, but only if it is real. A cleaned dataset with names, emails and panel IDs removed can often move from the short identifying clock to the longer study-data clock, provided the contract and notice allow it.

Pseudonymization is not the same thing. GDPR Recital 26 says personal data that has been pseudonymised and could be attributed to a person using additional information should be considered information on an identifiable person. If the agency keeps a key linking respondent IDs back to its panel, personal-data rules still apply to that dataset.

California sets its own bar. Under the CPRA amendments, information counts as deidentified only if the business takes reasonable measures against re-identification, publicly commits to keep it in deidentified form and contractually binds recipients to the same rules. An agency that wants long-term value from its archive, including possible licensing, needs that discipline in place while it keeps the data.

Retention and value: what de-identification changes
SituationWhat it usually means for retention
Contract requires return or destruction at project endDerived datasets usually go too, unless the contract carves them out
Contract silent, notice allows further research useDe-identified study data may be kept; ask counsel about identifiable files
Respondent asks for deletionDelete identifiable records; data that cannot be linked back is usually unaffected
A key linking IDs to people is keptThe data remains personal data; identifying clocks apply
Legal hold in placeKeep everything in scope until counsel releases it

Illustrative: a consumer insights agency rebuilds its schedule#

Illustrative: a fictional consumer insights agency runs quantitative studies in a survey platform, recruits qualitative participants from its own database and stores Zoom recordings and transcripts on a shared drive. Its old policy kept every project folder for one fixed period, with no distinction between data types.

The operations lead maps each data type to a row in the new schedule. Screener and contact data now delete at fieldwork close, recordings delete after analysis, and identified transcripts are replaced by de-identified versions. Several enterprise contracts require return or destruction, so those studies are flagged and excluded from long-term keeping.

The review also shows that Zoom auto-deletion had already removed older recordings, which settles the question for those projects. The result is a smaller, documented archive of de-identified study data the agency can use for trend work and, where rights allow, consider for licensing.

How SourceX treats retention in a licensing review#

SourceX only considers records an agency has kept lawfully and in line with its own policy, and it never asks an agency to keep data longer than its schedule allows. In the Supply and Rights steps of the SourceX five-step transaction, the retention schedule is one of the first documents reviewed, because it shows which studies still exist and on what terms.

For any package that proceeds, the privacy record in the SourceX Evidence Packet notes the de-identification method, the retention basis and any hold or deletion obligations, so the agency and the licensee work from the same facts.

Frequently asked questions

Should we delete raw data as soon as the report is delivered?

Not always. Clients often ask follow-up questions or request extra cuts after delivery, so raw files usually need a short window after the report. What matters is that the window is written down, starts on a defined event and ends with a deletion or de-identification step someone actually performs.

Does a client's destruction request cover our de-identified copy?

It may. Some clauses cover all copies and derivatives; others cover client materials and identifiable data only. If the wording is unclear, ask the client to confirm in writing what the agency may keep. Do not assume de-identification takes a dataset outside a return-or-destroy clause.

What do we do when a respondent asks us to delete their data?

Find every system that holds the person's identifiable records, including the recruitment database, survey platform, recordings and transcripts, and delete them. Data de-identified so it cannot be linked back is usually outside the request, but log what was deleted and where.

Do backups have to follow the schedule?

Backups should be covered by the policy even if they cannot be edited file by file. A common approach is to let backups expire on their own rotation and make sure deleted records are never restored into live systems. Write down how long backups persist so the policy matches reality.

Can we keep data for unspecified future research?

Keeping data with no defined purpose is hard to justify under most privacy frameworks and many client contracts. If the agency wants de-identified data for trend analysis, norms or possible licensing, name that purpose in the policy, check that notices and contracts support it and apply real de-identification.

Sources

  • Zoom lets account owners, admins and licensed users enable deletion of cloud recordings after a specified number of days counted from each recording's creation, with an option to exempt individual recordings. Source
  • GDPR Recital 26 states that pseudonymised personal data which could be attributed to a natural person by the use of additional information should be considered information on an identifiable natural person. Source
  • Under Cal. Civ. Code 1798.140(m), as amended by the CPRA, deidentified information requires reasonable measures against re-identification, a public commitment not to re-identify and contractual obligations on recipients. Source

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify