Skip to content

Speech and audio data

Drive-Thru and Voice Ordering Audio for Restaurant Voice AI

Quick answer

A useful drive-thru speech dataset is real lane and phone-order audio, captured on the actual speaker-box and crew-headset channels, time-aligned to transcripts and joined to the point-of-sale ticket that the order produced. That join turns audio into labeled slot data: items, sizes, modifiers, removals and quantities. Buyers should also require dated menu versions, recording-consent evidence and a written assurance that no voiceprints were used to recognize repeat customers.

By SourceX Editorial · Updated

Why drive-thru audio is its own acoustic domain

Drive-thru audio differs from contact-center audio because the customer is outdoors, often in an idling vehicle, speaking toward a weatherproof post microphone from a variable distance. Engine idle, wind buffeting the mic, rain, music from the car, passengers talking over the driver and the next car's horn all sit under the speech. Vendor guidance on robust ASR data treats car and outdoor noise as conditions separate from office or call-center noise, and treats far-field capture as its own case [1].

The crew side is different again. Order-takers wear headsets with close-talk mics, but they often hear kitchen noise, fryer timers and a second lane bleeding into the channel. A model trained on clean contact-center calls tends to fail exactly where drive-thru orders are hardest: low-SNR customer turns, clipped first words after the greeting, and overlapping "and also, can I get" corrections.

Phone orders add a third profile: narrowband telephony, often 8 kHz, with in-store background noise on the restaurant end. Treat phone and lane audio as separate strata in any spec. Our guides on telephony versus wideband ASR training and on audio file specs, sample rates and channels cover the format side in detail.

Key acoustic fields to request per recording:

  • Channel layout: separate customer (post mic) and crew (headset) tracks, or a single mixed track, and which one the vendor's transcripts were made from.
  • Capture point: is the audio tapped from the drive-thru timer or headset base station, or from a separate recorder, and what codec was used along the way?
  • Lane type: single lane, tandem or side-by-side lanes, which can produce crosstalk between order points.
  • Conditions metadata: time of day, weather where known, and whether a vehicle was present at the order point.

The final order is the label that matters

The most valuable label in ordering audio is the POS ticket, because it records what the restaurant actually rang up after every correction. Transcripts alone tell you what was said; the ticket tells you what the customer meant to order. Spoken language understanding work frames this as inferring meaning directly from audio, and public benchmarks typically pair audio with generic intent and entity labels rather than restaurant tickets.

For restaurant ordering, the entity schema is richer than generic assistant commands. A single utterance can carry an item, size, a combo upgrade, two removals, an addition and a quantity, and the customer may revise any of them three turns later. The ticket-to-audio join lets you score whether an agent captured the final state of the order, not just individual words.

Ask suppliers how the join was built. Common approaches are matching by store ID, lane, and the timestamp window between the greeting and ticket close, or by an order ID written to both the voice system log and the POS. Timestamp joins fail at busy stores where two lanes close orders seconds apart, so request the join method and a match-confidence field.

Treat these failure modes as known label noise:

  • Crew corrections: the order-taker rings up something different from what was said, and the customer accepts it.
  • Window changes: the customer changes the order at the pay or pickup window, and the change never appears in the lane audio.
  • Voids and comps: manager voids, coupon redemptions and app-order pickups that create tickets without matching speech.
  • Upsell acceptance: "Make it a large?" followed by a nod at the window rather than a spoken yes.

Menu vocabulary changes constantly through limited-time offers, renamed items, regional specials and price-driven combo changes, so every recording should carry a date and a menu version. Without that, an evaluation set can quietly reward a model for knowing a discontinued sandwich or penalize it for an item that did not exist yet.

We recommend requesting a menu snapshot table keyed by store and effective date: item IDs, display names, spoken aliases ("number three," brand nicknames), modifier groups and allowed combinations. That table becomes your grammar for constrained decoding, your slot vocabulary and your test for whether held-out audio covers recent items. It also lets you build time-split evaluation: train on earlier menu periods, test on later ones, and measure how gracefully the agent handles unseen items.

Spontaneity is the other coverage gap. Most public speech corpora are read or scripted speech, which underrepresents the hesitations, restarts and filler of real conversation [2]. Ordering speech is extreme on this axis: "uh, lemme get, no wait, the spicy one, no pickles." Synthetic or TTS-generated orders rarely reproduce this, a limit discussed in our page on real versus synthetic speech data for ASR.

Requirements spec for an ordering-audio request

A written spec lets suppliers say quickly whether they hold matching data and lets your counsel see what personal data is in scope. Use the table below as a starting point and adapt the counts to your own training plan; we do not suggest hour targets because they depend on your base model and task.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldWhat to specifyWhy it matters
ChannelsCustomer post mic and crew headset as separate tracks; phone orders as a separate stratumLets you train and score each side; mixed tracks hide customer errors
FormatOriginal codec and sample rate, plus any resampling appliedRe-encoding artifacts can be mistaken for noise robustness
SegmentationTurn boundaries with speaker role (customer, crew, passenger)Needed for turn-taking and barge-in evaluation
TranscriptsVerbatim, including fillers, restarts and crosstalk tagsClean transcripts erase the hard cases
Order labelPOS line items with modifiers, voids and final total structure, joined per orderGround truth for slot extraction and end-to-end eval
Join metadataJoin method, match confidence, unmatched-order rateMeasures label noise before you train on it
Menu versionStore-level menu snapshot keyed by effective dateSupports time-split eval and vocabulary drift tests
ConditionsLane type, daypart, weather flag where availableLets you stratify error rates
PrivacyRedaction method for names, phone numbers, card digits and loyalty IDsPhone orders often contain callback numbers and payment details
Consent evidenceSignage, IVR disclosure and recording-policy documentation by locationRecording law differs by state

For a broader template that covers contact-center fields, see how to specify contact-center audio data.

A good delivery ships each order as one record that ties audio segments to transcript turns and to the final ticket lines. The record below shows the shape; packaging conventions for manifests and segment timing are covered in speech dataset manifest packaging.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "order_id": "ord_000184",
  "channel": "drive_thru_lane",
  "lane_config": "side_by_side",
  "menu_version": "store_017_2026-08-04",
  "audio": {
    "customer_track": "audio/ord_000184_post.flac",
    "crew_track": "audio/ord_000184_headset.flac",
    "original_codec": "g722",
    "sample_rate_hz": 16000
  },
  "turns": [
    {"role": "crew", "start": 0.00, "end": 1.85, "text": "welcome, order when you're ready"},
    {"role": "customer", "start": 2.40, "end": 7.10, "text": "uh can I get the number three, medium, no onions [crosstalk]"},
    {"role": "customer", "start": 9.30, "end": 11.20, "text": "actually make that a large"}
  ],
  "ticket_lines": [
    {"item_id": "combo_03", "size": "L", "modifiers": ["no_onion"], "qty": 1}
  ],
  "join": {"method": "order_id_log", "confidence": 0.98},
  "redaction": {"names": "replaced", "phone_numbers": "tone_masked", "method_doc": "redaction_v2.pdf"}
}

Note how the later "actually make that a large" changes the size slot. An agent evaluated only on word error rate could transcribe every word correctly and still ring up a medium. For outcome-based scoring across many scenarios, see voice agent evaluation sets with real call outcomes.

The two legal questions buyers should resolve first are whether the audio was lawfully recorded and whether anyone derived voiceprints from it. On recording, California's Penal Code section 632 prohibits recording a confidential communication, including a phone call, without the consent of all parties [8]. Phone-order lines therefore usually need an upfront disclosure, and buyers should ask for the IVR script or greeting and the dates it was in place. Whether lane audio counts as confidential is a fact question for counsel, so ask for drive-thru signage and recording policies by location too.

On biometrics, Illinois BIPA defines voiceprints as biometric identifiers and requires a retention schedule and a written release before collection [3]. A 2024 amendment treats repeated collection of the same identifier from the same person by the same method as a single violation [4], which narrows damages but does not remove liability. Texas likewise lists voiceprints and requires notice and consent before capture for a commercial purpose [7].

Voice AI deployed with consumers has already drawn BIPA claims [5], and as of October 2026 several class actions filed in May 2026 are testing whether voiceprints used to train AI fall under the statute; none has been decided on the merits [6]. The practical risk in drive-thru deployments is customer recognition: some ordering systems may have been configured to recognize returning voices for personalization. Ask suppliers to confirm in writing that no speaker identification or voiceprint enrollment ran on the recordings, and treat any loyalty-linked audio as higher risk. Our page on voiceprints and BIPA in licensed voice datasets goes deeper.

Speaker anonymization can reduce linkability, but it trades off against utility. The VoicePrivacy challenge measures both privacy and downstream ASR performance on anonymized speech [9], and its results are worth reading before you request voice-converted audio for training. See speaker anonymization metrics for speech datasets.

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Diligence questions before you license ordering audio

Most problems with ordering audio surface in the first sample, so ask these questions before negotiating terms:

  1. Who owns the recordings: the restaurant operator, the franchisee, or the vendor running the order-taking system, and who can license them?
  2. What notice did customers receive on the phone line and at the lane, and over which dates?
  3. Were any voiceprints, speaker embeddings for identification or loyalty voice-matching created?
  4. How were callback numbers, card digits, names and license plates (if spoken) removed or replaced, and was a sample checked?
  5. What fraction of orders failed to join to a ticket, and why?
  6. Which menu versions does the audio span, and are store identifiers available or generalized?
  7. Is any audio from an existing automated ordering agent rather than human crew, and is it flagged?

The last question matters for training: audio where a prior AI agent took the order carries that agent's prompts and failure patterns, which can bias a new model. Mark those sessions separately or exclude them from evaluation.

Restaurants and hospitality groups hold far more than ordering audio; our overview of AI use cases for restaurant and hospitality data covers adjacent datasets, and the voice agent training data page explains how SourceX approaches conversational audio. When you are ready to describe a specific request, start on the SourceX buyer page.

How SourceX approaches ordering-audio requests

SourceX sources operational datasets, including new recordings of hands-on work and support and sales histories, from US companies on request; nothing is held in stock and a request does not guarantee a match. Buyers describe the data they need, each release is approved by the supplying company, and personal details such as names, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked. Every dataset is rights-reviewed and delivered under a license that defines records, uses, term and delivery.

For more speech categories and licensing context, see the speech and audio datasets buyer's guide and the AI data hub.

Request drive-thru ordering audio

SourceX looks for US businesses that hold the data you describe and manages the process from Find and Assess through Agree, Transact and Manage; nothing is contracted until a supplier agrees. Delivery runs through private, access-controlled workflows after an executed agreement and supplier approval. Describe the ordering audio you need on the SourceX buyer page.

Frequently asked questions

Can contact-center audio substitute for drive-thru audio?

Only partly. Contact-center audio helps with turn-taking and telephony phone orders, but it lacks outdoor and vehicle noise, post-mic distance effects and dense menu-modifier speech. Use it as pretraining or for the phone-order stratum, and evaluate on real lane audio.

Should I train on transcripts or on POS tickets?

Use both. Verbatim transcripts train and evaluate ASR, while the joined ticket is the ground truth for slot extraction and order accuracy. Report both word error rate and order-level accuracy, because they diverge on corrections.

Do drive-thru recordings include children's voices?

They can, since passengers often speak. Ask how passenger speech is tagged and whether the supplier's notices and redaction covered it, and have counsel review whether any child-privacy rules apply to your use.

Sources

  1. AIxBlock, "Noisy and Far-Field Speech Data for Robust ASR (2026)" (2026). https://aixblock.io/blogs/noisy-speech-data-asr
  2. arXiv, "CASPER: A Large Scale Spontaneous Speech Dataset" (2025). https://arxiv.org/pdf/2506.00267
  3. Illinois General Assembly, "Biometric Information Privacy Act (740 ILCS 14/)". https://www.ilga.gov/legislation/ilcs/ilcs3.asp?ActID=3004
  4. Illinois General Assembly, "SB 2979 (103rd General Assembly), BIPA amendment, engrossed text" (2024). https://www.ilga.gov/documents/legislation/103/SB/PDF/10300SB2979eng.pdf
  5. Lewis Rice, "AI Transcription Tools Give Rise to BIPA Claims". https://www.lewisrice.com/publications/ai-transcription-tools-give-rise-to-bipa-claims
  6. Biometric Update, "Tech giants sued under BIPA over voiceprints used to train AI" (2026). https://www.biometricupdate.com/202605/tech-giants-sued-under-bipa-over-voiceprints-used-to-train-ai
  7. Texas Legislature, "Texas Business and Commerce Code Section 503.001, Capture or Use of Biometric Identifier" (2026). https://statutes.capitol.texas.gov/Docs/BC/htm/BC.503.htm
  8. Justia (reproduction of California Penal Code), "California Penal Code Section 632" (2024). https://law.justia.com/codes/california/code-pen/part-1/title-15/chapter-1-5/section-632
  9. arXiv, "The VoicePrivacy 2024 Challenge Evaluation Plan" (2024). https://arxiv.org/pdf/2404.02677

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data