For AI teams
License voice and audio data for AI training
Short answer
SourceX sources voice and audio data from U.S. businesses that own them. Each dataset is rights-reviewed, prepared with names spoken aloud, account numbers read out and voices where consent is unclear removed, and approved by the supplying company before delivery under a use-limited license. Teams use it for speech recognition, voice agents and call analysis.
What's included#
Typically calls, voice notes and dictation.
Why it matters for models#
These records are valuable because real speech covers accents, noise and phrasing that scripts miss.
Common uses#
- Speech recognition
- Voice agents
- Call analysis
Removed before delivery#
- Names spoken aloud
- Account numbers read out
- Voices where consent is unclear
How a request works#
- Describe the records, volume and uses you need.
- SourceX matches suitable U.S. suppliers and runs a rights review.
- A prepared sample is shared under NDA after supplier approval.
- The license sets allowed uses, term and deletion terms; resale and re-identification are barred.
Frequently asked questions
Who supplies the data?
U.S. businesses that own the records. Every release is approved by the supplying company.
Can we see a sample?
Yes, once an NDA is in place and the supplier approves a prepared sample.
Is pricing published?
No. Terms depend on volume, history, exclusivity and allowed uses, and are agreed per deal.