Speech and audio data
Speech Emotion Recognition Data: Acted vs Naturalistic Corpora and Label Schemes
Quick answer
A speech emotion recognition dataset for production should be mostly naturalistic, in-domain speech with multi-annotator labels and a documented label scheme, with acted corpora used only for coverage of rare emotions and for pre-training. Acted sets are clean and balanced but small and exaggerated; naturalistic sets match real calls but are skewed toward neutral and need agreement statistics. Before buying, map every label set to one target scheme, confirm commercial rights, and check whether your deployment context restricts emotion inference.
By SourceX Editorial · Updated
This guide sits in our speech and audio data buyer's guide and focuses on emotion and paralinguistic labels. Style control for synthesis belongs to consented TTS voice recordings, and conversational dynamics belong to spontaneous conversational speech data.
Acted vs naturalistic emotional speech: what each is good for
Acted corpora give you controlled, balanced emotion classes, while naturalistic corpora give you the distribution your model will actually see. CREMA-D, a common acted benchmark, has 7,442 utterances from 91 actors across six emotions: anger, disgust, fear, happiness, sadness and neutral [1]. That is enough to benchmark, but it is a few hours of scripted sentences spoken with deliberate, often exaggerated affect.
Naturalistic corpora are built from speech that was not produced to display emotion. MSP-Podcast, a widely used naturalistic set, contains about 104,000 podcast segments labeled by crowdsourced annotators [2]. Its class distribution is dominated by neutral and mild states, and full-blown anger is rare, which is also true of contact-center audio.
The failure mode buyers hit most is a model that scores well on acted benchmarks and then over-fires on raised voices in noisy calls. Acted anger is loud, slow and pitch-extreme; real customer frustration is often quiet, clipped and lexical ("I have called three times"). Paralinguistic features alone carry part of the signal, and research comparing speech representations across emotion corpora finds that results vary by corpus [1], so test on your own audio rather than relying on benchmark rankings.
| Property | Acted (studio) | Elicited (induced tasks) | Naturalistic (in-the-wild) |
|---|---|---|---|
| Class balance | Designed balance | Partial | Heavy neutral skew |
| Intensity | High, prototypical | Moderate | Mostly low to moderate |
| Acoustic conditions | Clean, single mic | Lab or remote | Telephony codecs, noise, overlap |
| Label source | Intended emotion plus ratings | Task condition plus ratings | Perceived emotion from raters |
| Lexical leakage | Same sentences across emotions | Varies | Words correlate with emotion |
| Best use | Pre-training, rare-class coverage, probing | Controlled ablations | Fine-tuning and evaluation for deployment |
Why synthetic emotional speech is not a substitute
Synthetic emotional speech from TTS is useful for augmentation experiments but generalizes poorly to natural speech. Work on speech language models for emotional conversation reports that approaches trained on TTS-synthesized emotional speech struggle to generalize to natural speech [4]. The cause is the same as with acted data, only stronger: a TTS engine renders a stylized prototype of each emotion, so the classifier learns the engine's prosody template.
If you use synthetic data, hold out a naturalistic, in-domain test set that contains no synthetic or acted audio and report results on it separately. For the broader argument on where TTS augmentation stops helping, see real vs synthetic speech data.
Label schemes: categorical, dimensional and how to map them
Emotion label schemes differ across corpora, so you must map them to one target scheme before combining sets. Corpora use different categorical sets (six basic emotions, eight with contempt and surprise, or "other" buckets), and many naturalistic corpora add dimensional ratings for arousal, valence and dominance on Likert-style scales. A speech-LLM study combined several emotion corpora into roughly 70,000 utterances only after mapping their labels to common categories [3].
Decide your target scheme from the product, not from the corpora. A contact-center analytics model may need only frustration, satisfaction and neutral plus an arousal score; an empathetic voice agent may need finer categories that drive response style. Then write the mapping down and keep the original labels in every record so the mapping can be revised.
Three mapping pitfalls recur:
- Many-to-one collapses that hide disagreement. Mapping "frustrated" and "angry" to one class is fine only if raters in both corpora used them similarly.
- Unmappable classes. "Contempt," "surprise" and "other" often have no target; drop them or keep a catch-all, but never force them into neutral.
- Scale mismatch. Dimensional scores on different ranges need rescaling, and some corpora report the mean of raters while others report a single adjudicated value.
Illustrative example: invented to show structure; it does not describe an available dataset.
{
"segment_id": "call_88213_seg_014",
"audio_uri": "audio/call_88213.flac",
"start_s": 212.40,
"end_s": 218.95,
"speaker_role": "customer",
"channel": "stereo_left",
"sample_rate_hz": 8000,
"source_type": "naturalistic",
"elicitation": "none",
"transcript": "I have called three times about this.",
"labels_original": {"scheme": "vendor_v2", "raters": ["r07", "r12", "r31", "r44", "r52"], "votes": ["frustrated", "angry", "frustrated", "neutral", "frustrated"]},
"label_target": {"scheme": "cc_emotion_v1", "category": "frustration", "vote_share": 0.6, "arousal_mean": 3.4, "valence_mean": 2.1, "scale": "1-5"},
"agreement": {"segment_entropy": 0.95, "batch_krippendorff_alpha": 0.41},
"pii_handling": "account numbers replaced with tags",
"consent_basis": "documented in license schedule"
}
Keeping vote distributions, not just the majority label, lets you train with soft labels and filter low-agreement segments later.
Annotation quality: agreement, rater pools and label noise
Annotator agreement is the main quality signal for naturalistic emotion data, so require it per class and per batch. Naturalistic corpora depend on crowdsourced perceived-emotion labels [2], and agreement on subtle categories is typically much lower than on ASR transcripts. Ask for Krippendorff's alpha or Fleiss' kappa by class, the number of raters per segment, and how raters were screened and calibrated.
Treat low agreement as information, not only as a defect. Segments where five raters split three ways are genuinely ambiguous and belong in a separate pool or in soft-label training. Label errors also reach benchmark test sets: an audit of ten widely used vision, NLP and audio datasets estimated an average test-set label error rate of at least 3.3% [6], so budget for an independent gold-label audit of any evaluation split.
Practical checks before acceptance:
- Re-label a random sample with your own raters and compare class confusion against the vendor's labels.
- Check whether raters heard audio only, read transcripts only, or both; transcript-aware raters inflate lexical cues.
- Confirm rater pools match the language and dialect of the speakers.
- Require speaker-disjoint train, validation and test splits; emotion models memorize speaker identity easily.
For acceptance thresholds and sampling plans, use acceptable label noise tolerances and gold-label audits for delivered eval sets. Croissant-RAI offers a machine-readable way to document labeling process and rater details alongside the dataset [11].
In-domain sources for contact-center and voice-agent models
For contact-center and voice-agent models, the most representative emotion data is real, consented call audio from your target domain with emotion labels added after collection. Public naturalistic corpora come from podcasts, TV or interviews, which differ from narrowband telephony in codec, turn length, overlap and the reasons people are upset.
Real call recordings bring their own requirements. Agents and customers are usually on separate channels in stereo recordings, which makes per-speaker labeling easier; mono mixes need diarization first. Redaction of account numbers and names can insert tones or silence that a model may learn as a cue, so check how audio redaction affects training. The recording itself also needs a documented basis for reuse in model training, not just for quality monitoring.
SourceX sources operational datasets, including support histories, from US companies on request, and every dataset is rights-reviewed for ownership and consents before delivery under a license. See call center audio datasets for that category, or describe your target data to SourceX; categories are not inventory, and a request does not guarantee a match.
Commercial rights: what to check on public emotion corpora
Many well-known emotion corpora restrict commercial use, so verify the license of each corpus and of each upstream source before training a commercial model. Academic emotion sets are often released under non-commercial or no-derivatives terms, or under custom academic agreements, and commercial vendors publish their own dataset-card terms [5]. A no-derivatives license raises a separate question about whether relabeled or re-segmented versions may be shared internally or with vendors.
Check four layers: the corpus license, the terms of the underlying media (podcasts, broadcast or film clips), the speaker or actor consent scope, and any label-set license if annotations were added by a third party. Our open speech corpora commercial license audit walks through the checks, and fine-tuning-only data licenses covers narrower grants.
Illustrative example: invented to show structure; it does not describe an available dataset.
| Request field | What to specify |
|---|---|
| Source type | Naturalistic in-domain calls; acted share capped at a stated percentage |
| Speakers | Customer and agent roles; languages, accents, age bands where consented |
| Audio | Sample rate, codec, channel layout per role |
| Target scheme | Categories plus arousal and valence scale, with mapping from source labels |
| Annotation | Raters per segment, rater screening, audio-only or audio-plus-text |
| Quality | Per-class alpha, vote distributions retained, gold sample size |
| Splits | Speaker-disjoint, call-disjoint test set |
| Rights | Commercial training use, derivatives, consent scope, deployment geography |
Legal limits on emotion inference
Some jurisdictions restrict emotion recognition itself, so check your deployment context before you buy training data. The EU AI Act prohibits placing on the market or using AI systems to infer the emotions of natural persons in workplaces and education institutions, except for medical or safety reasons [7]. That matters for agent-coaching and employee-monitoring products built on call audio, even when the customer side is the main target.
Outside those prohibited contexts, deployers of emotion recognition systems must inform people exposed to them [8]. Emotion recognition systems are also listed among high-risk uses, which brings Article 10 data-governance and quality duties for training, validation and test sets [9]; as of October 2026, Regulation (EU) 2026/1744 reportedly moved the Annex III high-risk start date to 2 December 2027.
In the US, voice data can implicate biometric privacy law. Illinois BIPA section 15 sets retention, written-release and disclosure rules for biometric identifiers, which the Act defines to include voiceprints [10], and the 2024 amendment treats repeated collection from the same person by the same method as one violation. Whether an emotion model extracts a voiceprint is a fact question for counsel, so record which features your pipeline derives.
This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.
Sourcing emotion-labeled speech for your model
SourceX looks for US businesses that hold the operational data you describe, assesses the data and licensing permissions, and agrees allowed uses in a license before anything is delivered. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Describe the emotion-labeled speech data you need.
Frequently asked questions
How much acted data should be in a production training mix?
There is no fixed ratio, but treat acted data as a minority used for rare classes and pre-training, and select the ratio by measuring on an in-domain naturalistic test set that contains no acted audio.
Should I use categorical or dimensional emotion labels?
Use both if you can. Categories drive product logic, while arousal and valence ratings transfer better across corpora with different category sets and help with segments raters cannot agree on.
Can transcripts replace audio for emotion labeling?
No. Text-only labels miss prosodic cues and over-weight words, so a model trained on them inherits lexical bias; require raters to hear the audio, and record whether they also saw a transcript.
Sources
- arXiv, "Are Paralinguistic Representations all that is needed for Speech Emotion Recognition?" (2024). https://arxiv.org/pdf/2402.01579
- arXiv, "ParaLBench: A Large-Scale Benchmark for Computational Paralinguistics over Acoustic Foundation Models" (2024). https://arxiv.org/pdf/2411.09349
- arXiv, "BLSP-Emo: Towards Empathetic Large Speech-Language Models" (2024). https://arxiv.org/pdf/2406.03872
- arXiv, "Dual Information Speech Language Models for Emotional Conversations" (2025). https://arxiv.org/pdf/2508.08095
- Hugging Face / TrainingDataPro, "speech-emotion-recognition-dataset (dataset card)". https://www.huggingface.co/datasets/TrainingDataPro/speech-emotion-recognition-dataset
- Northcutt, Athalye, Mueller, "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
- European Commission, AI Act Service Desk, "AI Act Article 5: Prohibited AI practices". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-5
- European Commission, AI Act Service Desk, "AI Act Article 50: Transparency obligations for providers and deployers of certain AI systems". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50
- European Commission, AI Act Service Desk, "AI Act Article 10: Data and data governance". https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-10
- Illinois General Assembly, "Biometric Information Privacy Act, 740 ILCS 14/15". http://www.ilga.gov/legislation/ilcs/fulltext.asp?DocName=074000140K15
- Jain et al., MLCommons Croissant RAI task force, "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
Tell us what your models need
Share scope, volume, language, format, timing and licensing requirements.