Skip to content

Industry-specific operational data

Travel reservation change and disruption records for AI booking agents

Quick answer

Booking-change agents need real servicing records that link a traveler's request to the fare rule applied, any waiver used, the actions taken in the reservation system, and the money outcome: fee, fare difference, refund or credit. Synthetic benchmarks such as the τ-bench airline domain test policy-following with simulated users [1], but they cannot show how human agents actually resolved voluntary changes, schedule changes and irregular operations. Buy records that keep the request, the rule, the action and the outcome together, with identities tokenized.

By SourceX Editorial · Updated

Why synthetic airline benchmarks are not enough for booking-change agents

Public benchmarks give a useful baseline but too few, too clean cases for production servicing agents. The τ-bench airline domain gives an agent a policy document, a reservation database and tool APIs, then grades whether the final database state matches the expected outcome after a conversation with a simulated user [1]. That design is the right shape, and it exposes the core weakness: reported pass^k on the airline domain drops from about 0.46 at k=1 to about 0.23 at k=4, so an agent that succeeds once often fails the same task on repeat attempts [2].

Real servicing data adds what simulations lack. It carries ambiguous requests ("move me to the earlier one"), partial fare-rule compliance, supervisor overrides, waivers issued during weather events, and the tail of exceptions that synthetic task writers do not imagine. Earlier customer-service corpora such as ABCD showed the value of pairing dialogue with agent guidelines and actions rather than slots alone [3]; travel servicing records extend that idea with money and inventory consequences. For the general pattern, see exception handling records for agents.

What a reservation servicing record should contain

A usable record joins the request, the rule basis, the system actions and the final financial outcome under one case key. Most of this lives across separate systems: the PNR history in a GDS or airline passenger service system, ticketing and EMD records, the agency mid-office or CRM case, and the contact-center transcript. Records that keep only the conversation, or only the PNR history, cannot teach an agent why an action was allowed.

Ask suppliers which of these fields exist and how they are linked:

  • Original itinerary: segments, booking classes, fare basis codes, ticket and coupon status, tokenized record locator.
  • Change request: channel (voice, chat, email, self-service fallback), timestamp relative to departure, and the request text or transcript.
  • Change type: voluntary change, voluntary cancellation, carrier schedule change, or irregular operations (cancellation, misconnect, long delay).
  • Rule basis: fare-rule categories consulted (for example ATPCO Category 16 penalties, Category 31 voluntary changes and Category 33 voluntary refunds where the supplier uses them), and any waiver or policy code.
  • Actions: rebook, reissue, revalidate, exchange, refund, EMD issued, travel credit created, queue moves, and the order they happened in.
  • Money: change fee, fare difference and taxes as bands or normalized values, with refund method or credit amount.
  • Outcome and quality: final itinerary, whether the traveler accepted, recontacts within a set window, and any QA or audit flag.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "case_id": "c_7f31",
  "pnr_token": "pnr_9a02",
  "traveler_token": "trv_44be",
  "channel": "chat",
  "hours_to_departure": 52,
  "change_type": "carrier_schedule_change",
  "schedule_delta_minutes": 205,
  "original_segments": [{"o": "ORD", "d": "DEN", "class": "K", "fare_basis": "KA21NR"}],
  "rule_basis": {"fare_rule_cats": ["16", "31"], "waiver_code": "SKD_CHG_WAIVER"},
  "request_summary": "Traveler asks for earlier flight same day",
  "actions": ["offer_alternatives", "rebook_same_cabin", "reissue_ticket"],
  "fee_band": "0",
  "fare_difference_band": "0",
  "refund_outcome": "none_rebooked",
  "recontact_7d": false,
  "agent_override": false
}

How voluntary changes, schedule changes and irregular operations differ

Separate the three change types in labels and splits, because the traveler's rights and the agent's permitted actions differ for each. A voluntary change runs under the fare rules the traveler bought; a carrier schedule change or cancellation can trigger refund rights regardless of fare rules. Mixing them teaches an agent to charge fees where a refund or free rebooking was owed.

US rules matter here. As of October 2026, the DOT refund rule codified at 14 CFR Part 260 requires automatic refunds [5] when a carrier cancels or significantly changes a flight and the passenger does not accept alternative transportation or compensation, and it defines what counts as a significant change, including large departure or arrival time shifts, airport changes, added connections and cabin downgrades (verify the current text in the Federal Register). Records created before the rule took effect in late 2024 may reflect older practices, so tag each case with its servicing date and let your eval set weigh post-rule behavior. Check the Federal Register for later amendments or enforcement notices before encoding the rule in policy prompts.

Irregular operations records are the hardest to get and the most valuable. During a mass disruption, agents apply event-specific waivers, protect onto partner carriers and override normal rules, and the waiver text often lives only in an internal bulletin. Ask whether the supplier can provide the waiver bulletin alongside the cases that used it.

Privacy, payment and distribution-terms scoping

Exclude Secure Flight passenger data, passport and payment card data entirely, and tokenize names and record locators consistently across every table. Consistent tokens let you join a PNR history to its transcript and ticket records without exposing the traveler; inconsistent tokens break multi-turn evaluation. Frequent-flyer numbers, phone numbers, emails and street addresses also need removal or replacement, and free-text transcripts need a separate scrubbing pass because travelers read out confirmation codes and card digits.

Distribution and content terms are a separate rights question from privacy. Fare, schedule and availability content that a travel agency received through a GDS or an airline direct connection may be governed by agreements that restrict reuse, so a rights review should check those agreements before fare basis codes or fare amounts leave the supplier. Banding fees and fare differences reduces both commercial sensitivity and re-identification risk without losing the signal an agent needs. NIST AI 600-1 lists data privacy and confabulation among the generative AI risks to manage, which maps directly to an agent quoting a refund amount it invented [4].

Building an evaluation set from servicing records

Turn records into tasks with an expected end state, then run each task several times. Reconstruct the starting reservation, give the agent the applicable fare rules and waivers, replay the traveler's request through a simulator seeded from the real transcript, and grade the final itinerary, ticket actions and money outcome against what policy allowed. Because reliability across repeated trials is where agents fail [2], score pass^k rather than single-run accuracy.

Illustrative example: invented to show structure; it does not describe an available dataset.

Eval sliceWhat it testsGrading signal
Voluntary change, non-refundable fareApplying change fee and fare difference correctlyFee band and reissue match policy
Carrier schedule change beyond thresholdOffering refund before creditRefund offered; no fee charged
Irregular operations with waiverUsing event waiver, protecting onto alternativesWaiver code applied; final itinerary valid
Ambiguous requestClarifying before actingNo write action before confirmation
Supervisor override casesRecognizing when to escalateEscalation, not unilateral override

Keep a held-out slice from a different period or supplier to catch policy drift. Related setups appear in e-commerce order-support conversations with order state and approval and rejection records for approval-routing agents. For broader context on tool-calling data, see training data for tool use and function calling.

Request checklist for travel servicing data

Describe the data precisely so suppliers can tell whether they hold it. Use this checklist when writing a request:

  • Change types needed and target mix (voluntary, schedule change, irregular operations).
  • Channels and whether transcripts, PNR histories and ticket records must be linked.
  • Fare-rule and waiver fields required, and acceptable banding for money fields.
  • Date range, including whether pre- and post-2024 refund rule cases must be labeled.
  • Tokenization requirements and fields to drop entirely.
  • Intended use: SFT for servicing assistants, agent eval, or both.

SourceX sources operational datasets from US companies on request, including support histories and workflow records; it holds no stock, and a request does not guarantee a match. Buyers describe the data rather than naming businesses, and you can start a buyer request. More industry examples are in the industry-specific operational data guide, and supplier-side context is in what AI companies build with travel agencies data. See also customer support AI training data and eval sets, buyers by industry: customer support, and the AI data hub.

Sourcing travel booking-change records for your agents

SourceX looks for US businesses that hold the servicing records you describe, reviews ownership and consents, and removes or replaces personal details before delivery under a license that defines records, uses, term and delivery. Nothing is contracted until the supplying company agrees and approves the release. Describe the travel servicing records you need.

Sources

  1. arXiv (Sierra Research), "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains" (2024). https://export.arxiv.org/pdf/2406.12045
  2. Steel.dev leaderboard, "tau-bench benchmark registry entry". https://leaderboard.steel.dev/registry/benchmarks/tau-bench
  3. arXiv (NAACL 2021), "Action-Based Conversations Dataset: A Corpus for Building More In-Depth Task-Oriented Dialogue Systems" (2021). https://arxiv.org/abs/2104.00783v1
  4. National Institute of Standards and Technology, "Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1)" (2024). https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf
  5. Department of Transportation, Federal Register, "Refunds and Other Consumer Protections (89 FR 32760)" (2024). https://www.federalregister.gov/documents/2024/04/26/2024-07177/refunds-and-other-consumer-protections

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data