Skip to content

Industry-specific operational data

Freight rate quote and award history for pricing models

Quick answer

Freight rate data for machine learning is most useful when it is a broker's or shipper's own transactional history: every quote sent, every tender offered, the carrier cost paid, and whether the load was won, lost, accepted or rejected, by lane and date. Public index series show market direction but not win probability or margin. Buy lagged, de-identified, lane-level records that include losses and rejections, spanning both tight and loose capacity markets, and have counsel review any rate data that originates with a competitor.

By SourceX Editorial · Updated

This page is general information, not legal advice. Confirm requirements with counsel for your jurisdiction and use case.

Why index series cannot train a quote or win-probability model

Index data answers "where is the market going", while a quoting model needs "what price wins this load at this margin". A systematic review of machine learning freight rate forecasting found 28 studies from 2012 to 2024, most of them on dry bulk shipping using weekly or daily index series [1]. Comparative work on classical versus ML models follows the same index-forecasting pattern [2]. None of that gives you the conditional outcome a broker cares about: given a quoted rate on a lane, was it accepted.

Practitioner projects show the shape of truckload demand. Open trucking rate challenges build features from lane, equipment type, market and time attributes of posted rates [3], but posted board rates are asking prices, not awarded prices. For pricing and margin work you need the internal records that sit in a TMS (for example MercuryGate, Turvo, Revenova or McLeod), a quoting tool, or a shipper's procurement platform.

For a broader view of how this category compares with other operational data, see the industry-specific operational data hub and the cross-industry guide to price history and price-change data.

The records that matter: spot quotes, contract bids, tenders and carrier cost

A usable freight pricing dataset joins four record types at the load or quote level. Each answers a different modeling question, and missing any one of them limits what you can train.

  • Spot quotes: quote requests from shippers (portal, API, email) and the broker's response, with timestamp, quoted all-in or linehaul rate, and outcome. This is the core of a win-probability model.
  • Contract bids: annual or quarterly RFP lane submissions with bid rate, awarded volume and award rank. Contract rates often sit under shipper-carrier or shipper-broker agreements with confidentiality clauses, so the supplier must confirm it may disclose them.
  • Tenders: load tenders against contracted lanes, typically EDI 204 load tender and EDI 990 response messages, with accept or reject and reason codes. Rejection rates by lane are a direct capacity signal.
  • Carrier cost: what the broker paid the carrier (the buy rate), plus accessorials such as detention, layover and TONU. Without buy rate you can model price but not margin.

Linked context helps. Rate confirmations and bills of lading carry the agreed rate and accessorials (see freight document extraction data), negotiation threads explain why a quote moved (see broker and carrier communications data), and status events from EDI 214 shipment status data tell you whether a won load actually delivered on time.

A field list for lane-level quote and award history

The minimum schema keys each row to a quote or load, a lane at a stated geographic grain, a date pair, a price, and an outcome. The table below is a starting specification you can adapt in a request.

Illustrative example: invented to show structure; it does not describe an available dataset.

FieldExample valueNotes for modelers
quote_idQ-00418823Hashed, stable key
quote_ts2025-03-14T09:22:05-05:00Request time with time zone
pickup_date2025-03-17Lead time = pickup minus quote
origin_zip3 / dest_zip3606 / 303Or KMA-style market code
equipment53' dry vanVan, reefer, flatbed, power-only
distance_mi716State the mileage engine used
mode / serviceFTL spotSpot, contract, mini-bid, dedicated
quoted_rate_usd2,140All-in or linehaul; flag which
fuel_treatmentincludedIncluded, separate FSC table, or per-mile
carrier_cost_usd1,860Buy rate; null if lost
accessorials_usd75Detention, layover, TONU, lumper
outcomelostwon, lost, expired, no-response
competing_rate_banded1,900-2,099Only if the shipper shared it; banded
tender_statusn/aaccepted, rejected, cancelled (contract)
reject_reasonn/aCarrier code if present
customer_segmentretail, mid-sizeReplaces shipper identity

Two field decisions break models quietly. If fuel surcharge treatment is not explicit, a 2022 quote and a 2025 quote are not comparable on linehaul. If the lane grain changes partway through history (for example from 5-digit to 3-digit ZIP), lane-level features shift for reasons unrelated to price.

Survivorship bias and market-cycle coverage

The biggest failure mode is a dataset that only contains won loads. Many TMS exports are built from the load table, which records booked freight, so lost quotes and expired requests never appear. A win-probability model trained on that data has no negative class; a margin model trained on it learns only from prices the market accepted.

Ask explicitly for:

  1. Lost and expired quotes, with the quoted rate that lost.
  2. Rejected tenders and the carrier that rejected them (tokenized).
  3. Quote requests the broker declined to price.
  4. The share of quotes with recorded outcomes, by month, so you can see logging gaps.

Coverage across the capacity cycle matters as much as row count. Truckload rates swing between tight markets (high tender rejections, rising spot premiums) and loose markets (low rejections, spot below contract). A history covering only one regime will overfit to it; seek at least one tight and one loose period, and hold out a regime change as a test split. The same logic applies to building eval sets for a quoting agent: score it on weeks where the market turned.

Competition-law limits on sharing rate data

Current, firm-specific pricing exchanged between competitors can raise antitrust risk, and the older US safe harbors are gone. In December 2024 the FTC and DOJ withdrew the 2000 Guidelines for Collaborations Among Competitors, which had covered information sharing among other collaborations [4]. The withdrawal removed the safe harbors and safety zones those guidelines offered [5]. Earlier agency statements on information exchange had also been withdrawn, so ask counsel which historical conditions (third-party management, data age, aggregation) still serve as useful reference points.

As of October 2026, check with counsel whether the agencies have issued replacement guidance; until then, those historical conditions are reference points, not a shield. The risk profile depends on who is buying. A brokerage or carrier buying another broker's rates is a competitor receiving competitor pricing. A TMS vendor or a model developer outside the market carries a different profile, but if its model's outputs are then distributed to competing brokers, the model can become a conduit for pricing information.

Mitigations each cost model utility in a specific way:

MitigationWhat it doesEffect on the model
Time lag (e.g., records older than a set number of months)Removes current pricingFine for training; weak for live nowcasting
Aggregation thresholds (minimum contributors per cell)Hides any one firm's priceDrops thin lanes; biases toward dense corridors
Banding ratesReplaces exact price with a rangeLoses fine-grained price elasticity
Removing counterparty identitiesDrops shipper, carrier, broker namesLoses customer-level effects; keep segments
Coarser lane grain3-digit ZIP or market instead of addressReduces lane specificity, especially for short haul

Document which mitigations were applied and why in the license file. If you later need fresher data for production inference, treat that as a separate review rather than an extension of the training license. The license-terms guide on data use, exclusivity and deletion covers the contract side of these choices.

Rights and confidentiality checks before you license

Rate data has more contractual strings than most operational data. Before signing, confirm with the supplier:

  • Customer contracts: shipper agreements and RFP terms often mark rates confidential; the supplier must be able to disclose them in the form delivered.
  • Carrier agreements: broker-carrier contracts may restrict disclosure of buy rates.
  • Third-party rate feeds: if quotes were benchmarked against licensed market data, those benchmark values usually cannot be resold.
  • Personal data: owner-operator carriers are often individuals, so MC numbers, driver names, phones and emails in notes need removal or tokenization.
  • Allowed uses: training, evaluation, and whether derived model outputs can be offered to third parties.

Using the data: pricing models, quoting agents and eval sets

Different applications need different slices of the same history. A win-probability model needs balanced outcomes and quote-time features only (no leakage from post-award fields such as carrier_cost). A margin model needs won loads with buy rate and accessorials. A quoting agent that reads a shipper email and returns a price benefits from pairing the request text with the structured quote and outcome, similar to how RFQ and quote histories are used in industrial distribution.

For evaluation, freeze a time-based holdout after your training cutoff and score calibration, not just accuracy: a model that says 60% win probability should win roughly 60% of those quotes. When comparing supplier offers, normalize cost to usable records with outcome labels, as described in comparing data vendor quotes.

How SourceX approaches freight rate history requests

SourceX sources operational datasets from US companies on request, including sales and finance workflow records of the kind brokers and shippers hold, and manages the licensing process; it does not hold data in stock, and a request does not guarantee a match. Every dataset is rights-reviewed for ownership and consents, released only with the supplying company's approval, and delivered under a license that defines records, uses, term and delivery. Personal details are removed or replaced before delivery, with the method recorded and a sample checked. Buyers can describe the lanes, fields and history they need on the SourceX buyer page; see also freight brokerage buyers, logistics buyers and supply chain and logistics datasets.

Request freight quote and award history data

Describe the lane grain, record types, outcome fields and date range you need, and SourceX will look for US businesses that hold matching data. The process runs Find, Assess, Agree, Transact and Manage, and nothing is contracted until a supplier agrees. Start your request at sourcex.si/buyers.

Sources

  1. Copenhagen Business School (CBS Research Portal), "Machine learning in freight rate forecasting: a systematic literature review". https://research.cbs.dk/en/publications/machine-learning-in-freight-rate-forecasting-a-systematic-literat/
  2. Copenhagen Business School (CBS Research Portal), "Freight rate dynamics: a comparative study of classical and machine learning models". https://research.cbs.dk/en/publications/freight-rate-dynamics-a-comparative-study-of-classical-and-machin/
  3. GitHub (jennabeachcodes), "FreightRateChallenge". https://github.com/jennabeachcodes/FreightRateChallenge
  4. Morgan Lewis, "FTC and DOJ Withdraw Guidelines for Collaboration Among Competitors" (2024). https://www.morganlewis.com/pubs/2024/12/ftc-and-doj-withdraw-guidelines-for-collaboration-among-competitors
  5. Foley & Lardner, "FTC and DOJ Withdraw Guidelines for Collaborations Among Competitors" (2025). https://www.foley.com/insights/publications/2025/01/ftc-doj-withdraw-guidlines-for-collaborations-among-collaborators/

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data