Skip to content

Industry-specific operational data

Carrier shipment status (EDI 214) data for exception and ETA models

Quick answer

EDI 214 data is usable for machine learning only after you rebuild it into labeled, time-correct event sequences. Each 214 carries a shipment status code paired with a status reason code in the AT7 segment, a local date and time with a time-zone code, and location and equipment segments. The training value lives in what surrounds those codes: the time each message was received, every appointment revision, the source of each event, and the resolved exception outcome joined back to the shipment.

By SourceX Editorial · Updated

What a 214 message actually gives you

A 214 is a status notification from carrier to shipper, broker or 3PL, not a ground-truth trace, so treat each message as one party's report at one moment. The ANSI X12 214 transaction set exists to convey shipment status between trading partners [1]. In carrier implementation guides, the core is the AT7 segment: AT7-01 is the shipment status code, AT7-02 the status reason code, AT7-03 and AT7-04 an appointment status and its reason, AT7-05 and AT7-06 the date (CCYYMMDD) and time, and AT7-07 a time code such as ET, PT, UT or LT.

Location rides in MS1 (city, state, country, and in some guides latitude and longitude), and equipment in MS2 (SCAC and equipment number). Typical status codes include X3 (arrived at pickup), AF (departed pickup), X1 (arrived at delivery) and CD (departed delivery), while reason codes include NS (normal status) and delay causes such as accident, driver-related, mechanical breakdown and weather. Verify every code against the specific partner guide: some carriers publish only a base status list and supply further codes on request. That variability is the first thing a training pipeline must normalize.

Buyers should also expect the reference loop: shipment identification (often the carrier PRO and the shipper's bill of lading or load number in L11-style references), which is the only reliable way to join 214 events to tenders (204), invoices (210) and the shipper's own order records.

Where the label noise comes from

Most 214 label noise is systematic, not random, and it concentrates in reason codes and timestamps. The patterns below are common in multi-carrier feeds and should be measured before any model is trained.

  • Default reason codes. Many carrier systems send NS (normal) on every event, including late ones, because the dispatcher never selected a delay reason. A delivered-late shipment with only NS codes is not a "no exception" example; it is an unlabeled one.
  • Back-filled events. Drivers or dispatchers enter arrival and departure after the fact, so an X1 can arrive hours after the reported event time, sometimes in the same batch as the departure event. Without a receipt timestamp you cannot tell a fast carrier from a back-filling one.
  • Manual check calls. Phone or portal check calls typed by broker staff carry rounded times, free-text locations and inconsistent reason codes.
  • Time-zone errors. AT7-07 set to LT, or a missing code, turns a cross-country load's timeline into a mix of local and zoned times; a three-hour error is enough to flip an on-time label.
  • Code-set drift. Carriers add or retire codes and change mappings when they switch TMS vendors, which breaks multi-year label consistency (see code-set revisions in multi-year operational data).
  • Duplicates and corrections. Resends after acknowledgment failures, and corrected events that replace rather than supersede the original, distort event counts and dwell times.

The practical fix is to define labels from outcomes, not from the carrier's own reason code. "Late" should be computed from actual delivery time against the appointment in force at tender or at a defined cutoff, and the AT7-02 reason should be a feature or a weak label, with a measured agreement rate against resolved cases.

Preventing leakage in ETA and delay models

ETA leakage is prevented by reconstructing what was knowable at each prediction time, which requires receipt timestamps and full appointment history. Event-log prediction research has found that random splits leak information across cases and that temporal, case-level splits are the standard remedy [3]. ETA prediction from status events is a close cousin of remaining-time prediction in process monitoring, where a cross-benchmark comparison found that accuracy varies widely across methods, prefix encodings and event logs [4].

In freight data the specific traps are concrete:

  1. Final appointments only. If the dataset stores just the last delivery appointment, a model learns from appointments that were rescheduled after the delay was already known. Require original and every revised appointment with the time each revision was made.
  2. Event time versus receipt time. Training on reported event times makes back-filled events look like real-time signal. Store both, and build prediction-time snapshots using receipt time.
  3. Post-hoc corrections. Corrected delivery times and OS&D notes entered days later must be excluded from features at earlier snapshots.
  4. Lane leakage. Repeating the same shipper-consignee lane across train and test inflates accuracy; split by time first, then check lane overlap.

Machine learning over EDI streams to predict events is an established idea in practice, as patent filings on EDI event prediction show [2]. The difference between a working model and a misleading one is almost entirely in this snapshot discipline.

Mixing EDI, API, telematics and check-call sources

A multi-source status stream needs a per-event source field, because cadence, latency and reliability differ sharply by channel. A shipment may have 214 events from the carrier, API pings from a visibility platform, ELD or telematics positions, and broker check calls, all describing the same movement at different resolutions. Models trained on a mix without a source column can learn the channel rather than the shipment, for example that API-tracked loads are rarely late if larger carriers are the ones integrating by API.

Telematics trails need separate handling. The ELD functional specifications in Appendix A to 49 CFR Part 395, Subpart B, treat vehicle position as a recorded data element, and 49 U.S.C. 31137(a)(2) requires that ELDs not be used to harass vehicle operators. For owner-operators, a location trail is effectively a record of one person's movements, so ask for positions generalized to city or a coarse grid, and drop off-duty and personal-conveyance segments. For the broader handling of timestamped sequences, see timestamped business event sequences.

Labels for exception classifiers

Exception classifiers learn best from resolved cases, so require outcomes joined to the status stream rather than alert logs alone. An alert says a load looked late at 14:00; a resolved case says it delivered 26 hours late, was partially refused, and the OS&D claim was settled as carrier liability. The second is a label; the first is a prediction someone else made.

Useful outcome fields include: actual delivery against each appointment, refused or partial delivery, OS&D (over, short and damaged) flags with claim disposition, detention and layover accruals, re-consignment, and the final responsible party. These usually live outside the 214 feed, in the TMS, claims systems and broker notes, which overlaps with freight broker and carrier communications data and with the workflow covered in logistics exception resolution. Rare exceptions such as cargo theft or hazmat holds need deliberate oversampling; see long-tail and edge-case coverage.

Requirements table for a 214 training dataset

The table below is a checklist of what to require from a supplier and what each field prevents. It is meant to be pasted into a data request or diligence questionnaire.

Illustrative example: invented to show structure; it does not describe an available dataset.

RequirementField or artifactFailure it prevents
Raw AT7 codes plus normalized codeat7_status, at7_reason, norm_status, code_map_versionSilent mapping drift across carriers and years
Reported event time with zoneevent_ts_local, at7_time_code, event_ts_utcTime-zone label flips
Receipt timereceived_ts_utc, isa_control_numberBack-fill leakage, duplicate resends
Appointment historyappt_type, appt_ts, appt_set_ts, revision_seqTraining on post-delay appointments
Event sourcesource (edi_214, carrier_api, eld, check_call, visibility_platform)Channel learned instead of shipment
Join keyspro_number, bol_number, load_id (pseudonymized consistently)Unjoinable outcomes
Resolved outcomedelivered_late_min, refused, osd_flag, claim_dispositionLabels from alerts instead of outcomes
Location precision policylocation_precision (city, grid, exact), personal_conveyance_removedOwner-operator movement trails
Coverage profilecarriers, modes (TL, LTL, intermodal), lanes, date rangeHidden sampling bias
Provenance of each feedcontract or platform terms per sourceUnlicensable platform-sourced events

Illustrative example: invented to show structure; it does not describe an available dataset.

A minimal record after normalization:

{
  "shipment_key": "S-7f3a91",
  "source": "edi_214",
  "at7_status": "X1",
  "at7_reason": "NS",
  "event_ts_local": "2026-03-04T15:42",
  "at7_time_code": "CT",
  "event_ts_utc": "2026-03-04T21:42:00Z",
  "received_ts_utc": "2026-03-05T02:10:00Z",
  "location": {"city": "Joliet", "state": "IL", "precision": "city"},
  "scac_hash": "c19e0b",
  "appointment_in_force": {"type": "delivery", "appt_ts_utc": "2026-03-04T19:00:00Z", "revision_seq": 2},
  "outcome": {"delivered_late_min": 162, "osd_flag": false, "reason_label_source": "computed"}
}

Note what this record shows: the carrier reported NS, the event was received more than four hours after it happened, and the late label is computed from the appointment rather than taken from the reason code.

Rights and licensing checks specific to status data

Status data often reaches a shipper or broker through a visibility platform or carrier portal, so the holder's own rights may be narrower than its possession of the data suggests. Before licensing, establish for each feed whether events arrived directly over EDI under a trading-partner agreement, through a platform whose terms govern reuse, or from telematics shared under a carrier agreement. Each path can carry different restrictions on training use and on disclosure to third parties.

Also confirm how carrier identities are treated. SCACs and carrier names are commercially sensitive to the data holder, and consistent pseudonymization keeps carrier-level effects learnable without exposing relationships. Personal data in this stream sits mostly in driver names and phone numbers in check-call notes, and in location trails; ask how each was removed or generalized and whether a sample was checked. Related record types and their buyers are summarized on the supply chain and logistics datasets page and in logistics operations AI use cases.

How SourceX approaches carrier status data

SourceX sources operational datasets from US companies on request, so a buyer describes the status stream and outcomes needed rather than naming a supplier, and a request does not guarantee a match. Each dataset is rights-reviewed for ownership and consents, and personal details such as names, phone numbers and account numbers are removed or replaced before delivery, with the method recorded and a sample checked; no method is perfect. Delivery happens only through private, access-controlled workflows after an executed agreement and supplier approval. Logistics teams can see how buyers work with SourceX on the logistics buyers page, or start a request on the SourceX buyer page. For neighboring record types, browse the industry-specific operational data hub and the AI data guide index.

Request carrier status and exception data for your models

Describe the carriers, modes, lanes, date range and outcome fields your ETA or exception models need, and SourceX will look for US businesses that hold that data. Every release is approved by the supplying company and delivered under a license that defines the records, uses, term and delivery. Describe the shipment status data you need.

Sources

  1. Data Interchange, "EDI 214 T-Set: Optimising Shipment Status Communication". https://datainterchange.com/edi-214-t-set-optimising-shipment-status-communication/
  2. USPTO, "Rules/model-based data processing system for intelligent event prediction in an electronic data interchange system (US 11200370)". https://image-ppubs.uspto.gov/dirsearch-public/print/downloadPdf/11200370
  3. arXiv (Weytjens and De Weerdt), "Creating Unbiased Public Benchmark Datasets with Data Leakage Prevention for Predictive Process Monitoring" (2021). https://export.arxiv.org/abs/2107.01905
  4. arXiv, "Survey and cross-benchmark comparison of remaining time prediction methods in business process monitoring" (2018). https://arxiv.org/pdf/1805.02896

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data