Skip to content

Speech and audio data

Speech Dataset Metadata: Speaker, Accent, Device and Environment Fields Buyers Should Require

Quick answer

A speech dataset is only as usable as its metadata. At minimum, require a stable pseudonymous speaker ID, L1 and self-described accent, age band, and speaker role at the speaker level, plus capture device, channel map, sample rate, codec, acoustic environment, domain or call type, and recording date at the file level. Without these fields you cannot build speaker-disjoint splits, slice word error rate by accent or device, or prove coverage to a fairness reviewer.

By SourceX Editorial · Updated

Why speech metadata decides whether the audio is usable

Speech metadata is what turns hours of audio into a training and evaluation asset you can slice, split and defend. Two corpora with identical hour counts can differ completely in value: one lets you report WER for Indian English on a headset in a car, the other is an undifferentiated pile of WAV files. Documentation frameworks such as Data Statements were designed for exactly this, asking curators to describe speaker demographics, language variety, speech situation and recording conditions [1].

Generic dataset cards help but stop short of audio needs. The Hugging Face Hub card format records license, language and size in a YAML block on the README [2], and Croissant-RAI makes responsible-AI documentation machine-readable [3]. Neither tells you which microphone captured a file or which channel holds the agent. That audio-specific layer is what this page covers; for general card structure see dataset cards for licensed enterprise data and the dataset card glossary entry.

Per-speaker fields: identity, language background and role

The speaker table should be a separate file keyed by a pseudonymous ID that never changes across sessions, files or delivery tranches. If one person appears under two IDs, a "speaker-disjoint" test set leaks voices from training and your eval numbers inflate quietly. Ask suppliers how IDs were assigned (account ID hash, enrollment record, manual clustering) and whether re-recordings by the same person were merged.

Language background is where most corpora are weakest. Researchers building L2 English data note that existing learner corpora often lack detailed linguistic speaker profiles [5], which makes accent coverage claims impossible to verify. Require L1 (ISO 639-3 code), other languages spoken, self-described accent or region, and years of exposure to the target language where accent robustness is the goal; the accented English speech data guide covers coverage targets.

Demographic fields should be banded and consented. Use age bands (for example 18-29, 30-44, 45-59, 60+) rather than birth dates, and treat gender as optional, self-reported and consented, with "not provided" as a valid value. Crowdsourced corpora such as Common Voice collect this kind of contributor information voluntarily [4], which is a reasonable pattern to ask commercial suppliers to match.

Role matters for operational audio. In call center or clinic recordings, tag each speaker as agent, customer, clinician, patient or other, because the agent population is small and repeated while the customer population is large and diverse. A dataset that is 40% speech from 30 agents behaves very differently from its headline speaker count.

Per-file fields: device, channel layout and capture path

Every file should state how it was captured, because device and capture path shift the acoustic distribution more than most buyers expect. Vendors already sell on this axis: one physician dictation catalog reports its audio split by capture device, with telephone dictation at 54.3% and digital recorders at 24.9% (self-reported) [6]. If your deployment is smartphone dictation, a telephone-heavy corpus is a domain mismatch even when the specialty mix is perfect.

Channel layout must be explicit, never inferred. Telephony platforms make this a configuration choice: as of October 2026, Twilio's RecordingChannels parameter accepts mono or dual and defaults to mono, and dual keeps each call leg in its own channel [8]. Amazon Connect lets administrators record the customer, the agent or both during agent interactions [9], so confirm which party lands on which channel rather than assuming it. Require a channel_map field per file (for example ch0=agent, ch1=customer) and spot-check it by listening.

Technical audio parameters belong in the same row: sample rate, bit depth, codec, and whether the file was transcoded from an earlier format (G.711 at 8 kHz upsampled to 16 kHz is still narrowband audio). The audio file specs guide covers format choices, and what file formats AI buyers accept covers container conventions.

Environment and acoustic fields for far-field and noisy use

Environment metadata lets you build robustness slices instead of guessing them. At minimum, tag each file with a controlled vocabulary such as quiet_office, open_office, vehicle, street, home, clinic_room or call_center_floor, plus a near-field or far-field flag. Estimated SNR per file, computed with a stated method, is more useful than a supplier's subjective "clean" or "noisy" label.

Far-field corpora need geometry. A room-acoustics database from Brno University of Technology ships microphone and loudspeaker positions alongside its recordings [7], and the AMI Meeting Corpus synchronizes close-talking and far-field microphones on a shared timeline across 100 hours of meetings [10]. If you are training for smart speakers or meeting rooms, require mic array type, mic-to-speaker distance band and room type, and ask whether a close-talk reference channel exists. The noisy speech data guide covers SNR coverage specifications.

Domain and processing fields record what the speech is about and what has been done to it. Tag each file with domain (retail support, insurance claims, radiology dictation), call or session type (inbound, outbound, dictation, meeting), speaking style (spontaneous, read, scripted), and recording date or quarter so you can detect drift in product names and phrasing.

Processing history must be per file. Record redaction method (silence, tone, noise fill), redacted span timestamps, any resampling or loudness normalization, and transcript standard (verbatim or clean). Redaction artifacts change what the model learns; see how audio redaction affects speech model training and speaker anonymization for speech datasets.

Consent and rights scope should travel with the record, not only with the contract. Voice can be treated as biometric data: Texas law, for example, lists a voiceprint as a biometric identifier and requires notice and consent before capture for a commercial purpose [12]. A per-speaker consent_scope and per-file permitted_use field, as described in the rights metadata schema, lets you filter out records whose scope excludes voice cloning or TTS. Contract-side terms are covered in license terms for speech and voice recordings.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

A speech metadata schema you can paste into an acceptance spec

The schema below splits metadata into a speaker table and a file table joined by speaker_id. Mark each field required, required-if-applicable or optional in your statement of work, and state the allowed values so the supplier cannot invent new labels mid-delivery. If you are sourcing licensed operational recordings, the same table works as the data description you share with SourceX as a buyer.

Illustrative example: invented to show structure; it does not describe an available dataset.

TableFieldType / allowed valuesRequirementWhy it matters
speakersspeaker_idStable pseudonymous stringRequiredSpeaker-disjoint splits
speakersl1ISO 639-3 codeRequiredAccent and L2 slicing
speakersaccent_regionControlled list plus free textRequired for accent workCoverage proof
speakersage_band18-29 / 30-44 / 45-59 / 60+ / not_providedRequiredFairness eval
speakersgenderSelf-described / not_providedOptional, consentedFairness eval
speakersroleagent / customer / clinician / patient / otherRequired for multi-partyPopulation balance
speakersconsent_scopeASR / TTS / cloning / eval flagsRequiredUse filtering
filesfile_id, session_idStringsRequiredJoins and dedup
filesdevice_typeheadset / handset / smartphone / recorder / array / laptopRequiredDevice slicing
filescapture_pathPSTN / VoIP / app / localRequiredCodec and bandwidth
fileschannel_mape.g. ch0=agent;ch1=customerRequired if >1 channelDiarization and role labels
filessample_rate_hz, bit_depth, codec, transcoded_fromNumeric / enumRequiredTrue bandwidth
filesenvironment, field_type, snr_db_estEnum / near or far / float plus methodRequiredRobustness slices
filesdomain, session_type, speaking_styleEnumsRequiredDomain match
filesrecorded_quarterYYYY-QnRequiredDrift checks
filesredaction_method, redacted_spansEnum / JSON list of [start,end]Required if redactedArtifact handling

Acceptance checks that catch bad metadata

Metadata errors are as common as transcript errors, so test them before you sign off on a delivery. Even widely used benchmark test sets, including audio sets, carry an estimated average label error rate of at least 3.3% [11], and supplier metadata usually gets less scrutiny than labels.

  • ID integrity: run speaker embedding clustering on a sample and flag clusters that span two speaker_id values, or one ID that splits into two clear voices.
  • Channel map: listen to 50 or more dual-channel files and confirm the agent is where channel_map says.
  • Bandwidth truth: check spectral energy above 4 kHz on files labeled 16 kHz; flat spectra above 4 kHz indicate upsampled telephony audio.
  • Distribution match: compare delivered shares of device, environment, accent and age band against the shares in the spec, by hours and by speaker.
  • Null rates: reject or renegotiate when required fields exceed an agreed not_provided rate.
  • Enum drift: fail any value outside the allowed list rather than silently mapping it.

Pricing often depends on these same dimensions; see how speech data is priced. Broader sourcing options, including licensed operational recordings, are compared in off-the-shelf vs custom vs licensed speech data and in the speech and audio data hub.

Specify the speech metadata you need with SourceX

SourceX sources operational datasets, including support and sales call histories and new recordings of hands-on work, from US companies on request; it does not hold speech data in stock, and a request does not guarantee a match. Describe the audio and metadata fields you require, and each dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery once a supplier agrees. Describe the speech data and metadata you need.

Frequently asked questions

Should accent labels be self-reported or annotator-assigned?

Collect both where possible and keep them in separate fields. Self-reported L1 and region are ground truth about background; annotator-assigned accent labels describe perception and drift between annotators, so record annotator ID and the label guideline version.

Is gender a required field for ASR fairness evaluation?

It is useful but should stay optional, self-described and consented. Age band, L1 and device often explain more WER variance in operational audio, and forcing a gender value invites guessed labels.

How should metadata be delivered alongside audio?

Deliver the speaker and file tables as CSV or Parquet with a data dictionary, plus a dataset card that states collection method, value definitions and known gaps [2][3]. Keep per-file metadata out of filenames, which break on renames and cannot carry nulls.

Sources

  1. University of Washington Tech Policy Lab, "Data Statements" (2021). https://techpolicylab.uw.edu/data-statements/
  2. Hugging Face, "Dataset Cards (Hub documentation)". https://huggingface.co/docs/hub/en/datasets-cards
  3. Jain et al., MLCommons (arXiv:2407.16883), "A Standardized Machine-readable Dataset Documentation Format for Responsible AI" (2024). https://arxiv.org/pdf/2407.16883
  4. Ardila et al., Mozilla (arXiv:1912.06670), "Common Voice: A Massively-Multilingual Speech Corpus" (2019). https://arxiv.org/abs/1912.06670v1
  5. Heinrich Heine University Dusseldorf, SLaM lab, "AnglistikVoices: an L2 English speech dataset for educational and technological advancement in speech" (2024). https://slam.phil.uni-duesseldorf.de/publication/akhilesh-2024-l2english/
  6. Shaip, "License High-quality Healthcare/Medical Data for AI & ML Models (physician dictation audio)". https://shaip.com/offerings/physician-dictation-audio-data-medical-data-catalog
  7. Brno University of Technology, Faculty of Information Technology, "Building and Evaluation of a Real Room Impulse Response Dataset" (2019). https://www.fit.vut.cz/research/publication-file/c159973/280107/IEEE_Final_Published_08717722-1.pdf
  8. Twilio, "Recordings resource". https://www.twilio.com/docs/voice/api/recording-resource
  9. Amazon Web Services, "When, what, and where for contact recordings in Amazon Connect". https://docs.aws.amazon.com/connect/latest/adminguide/about-recording-behavior.html
  10. TensorFlow Datasets, "AMI (community catalog)". https://tensorflow.org/datasets/community_catalog/huggingface/ami
  11. Northcutt, Athalye, Mueller (arXiv:2103.14749), "Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks" (2021). https://arxiv.org/abs/2103.14749
  12. Texas Legislature, "Texas Business and Commerce Code Section 503.001 - Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data