Skip to content

Logistics and distribution

How to test an AI agent on your own logistics exceptions before you buy

By SourceX Editorial · Updated

Short answer

To evaluate an AI agent for logistics before you buy, test it on your own resolved exceptions instead of a vendor demo. Pull 100 to 300 closed cases, hide the outcomes, ask the agent what it would do at each decision point, score its answers against what your team did, review every miss, then decide.

Key takeaways

  • Vendor demos use tidy cases; your exceptions carry your carriers, customer rules and missing information.
  • A blind test hides resolution notes, final charges and later emails, so the agent sees only what your team saw at the time.
  • A written rubric scores the action, the party contacted, policy compliance, cost treatment and escalation judgment.
  • Reviewing misses by cause shows whether the problem is the agent, your data or your undocumented policies.
  • De-identify the test set and bar the vendor from training on it before anything leaves your systems.

Why test on your own exception history?#

Testing on your own exception history is the only way to see how an AI agent handles the situations your team actually faces. A demo shows a late pickup with complete data and an obvious answer; your real cases include a consignee who changed the appointment by phone, a carrier with a history of missed check calls and a customer whose routing guide overrides the usual rule.

Your history also gives a fair comparison. Every vendor can be scored on the same cases, with the same rubric, against the same ground truth: what your team did and how it turned out.

Step 1: pick the exceptions#

The test set should be a deliberate sample of closed, resolved exceptions that reflects your mix of customers, modes and problem types. We suggest 100 to 300 cases: enough to cover the common exception types, small enough for a supervisor to grade by hand.

Include hard cases on purpose. Exceptions that needed escalation, ran into a customer-specific rule or ended in a claim tell you far more than routine late-delivery notices.

  • Choose cases closed under current processes, so the right answer reflects today's policies.
  • Cover the main types: late pickup, missed appointment, refused delivery, shortage or damage, detention, address correction and temperature excursion if you haul reefer.
  • Spread cases across several customers and carriers rather than one large account.
  • Include escalated cases, cases that became claims and cases where the first response was wrong.
  • Exclude cases still in dispute or under a legal hold.

Step 2: hide the outcomes and package the inputs#

A blind test gives the agent only what your team knew at the decision point. For each case, package the original alert, the shipment record, the relevant customer instructions and the emails or notes up to that moment, then remove everything that came after.

Hidden material includes resolution notes, final accessorial charges, claim outcomes and later messages. Leaving any of it in lets the agent read the answer rather than reason toward it, and the scores become meaningless.

Prepare the package before it leaves your systems. Replace customer, carrier and driver names with consistent labels, remove phone numbers and personal remarks, and keep the mapping in-house. The pilot agreement should state that the vendor may use the test set only for the evaluation, may not train on it and must delete it afterward.

  • Show: the original alert or customer complaint, with its timestamp.
  • Show: the load or order record as it stood then, including lane, equipment, appointment and carrier.
  • Show: customer instructions that applied, such as routing guides, SOPs and accessorial rules.
  • Show: emails, check calls and notes up to the decision point.
  • Hide: resolution notes, later messages, final charges, claim outcomes and closure codes.

Step 3: score the answers with a rubric#

A scoring rubric turns judgment into comparable numbers across vendors. Write it before the test and have the graders agree on how to apply it using a handful of sample cases.

Grade against what a good dispatcher or account rep would do, not only against what your team happened to do. Sometimes the agent's answer is better than the historical one, and the rubric should allow credit for that.

Step 3: score the answers with a rubric
CriterionPassPartialFail
ActionCorrect next step for the situationReasonable step, but not the best oneWrong or harmful step
Party contactedRight carrier, consignee or customer contactRight party, wrong channel or timingWrong party or none
Customer rulesFollows the routing guide or SOPPartly follows itIgnores or contradicts it
Cost treatmentCorrect accessorial or chargeback handlingFlags cost but misjudges itMisses or invents a charge
CommunicationClear, accurate message ready to sendNeeds light editingUnclear or inaccurate
EscalationEscalates exactly when a person should decideEscalates late or too oftenActs alone on a decision it should not make

Steps 4 and 5: review misses and decide#

Every failed or partial answer gets a cause, because the cause decides what you do next. Sort misses into four groups: information the agent never had, a policy that exists only in people's heads, reasoning the agent got wrong, and actions that would have been risky if taken automatically.

The first two groups point at your data and documentation, not the vendor. The last two are the vendor's problem, and risky actions matter most, since one wrong automatic message to a customer can cost more than many correct ones save.

Steps 4 and 5: review misses and decide
Result patternDecision
Strong across types, few risky actionsPilot in production with human approval on every action
Strong on some exception types, weak on othersPilot only on the strong types
Misses traced mostly to missing data or unwritten policiesFix records and SOPs, then retest
Frequent risky actions or invented chargesDo not proceed with this vendor

Illustrative: a 3PL tests two exception agents#

Illustrative: a fictional 3PL with a freight brokerage arm keeps exceptions in its TMS, with notes and customer emails in shared inboxes. Two vendors offer exception agents, and the COO asks for a blind test before either contract is signed.

The operations team builds a test set from the past year of resolved cases, de-identifies it and grades both agents with the same rubric. One vendor scores well on appointment changes and detention but repeatedly ignores a large customer's routing guide on refused deliveries. The other writes fluent messages but escalates almost nothing, including two cases that became claims.

The COO pilots the first vendor on appointment and detention exceptions only, with a supervisor approving each message, and writes the missing routing-guide rules into the SOP library before any retest.

Your exception history is an asset in its own right#

A well-built test set shows how valuable resolved exception history is: situations, decisions and outcomes, linked and explained. That is why the pilot agreement should keep vendors from training on it, and why the same records interest developers who build logistics models.

SourceX works on that second use. It does not build or sell agents. Through the SourceX five-step transaction of Supply, Rights, Preparation, Approval and Delivery, a logistics company can license de-identified exception records for training or evaluation, with the SourceX Evidence Packet documenting provenance, permitted use and release authorization.

Frequently asked questions

How many vendors should we test at once?

Two or three is manageable. Each extra vendor adds grading time and coordination, and the test set must be prepared and governed for each. If more vendors are interested, screen them first on a short written questionnaire about integrations, data handling and escalation controls.

Should vendors see the scoring rubric?

Share the criteria, not the answers. Vendors should know you will grade action, party contacted, customer rules, cost treatment, communication and escalation, so they can configure sensibly. They should never see the hidden outcomes or the grades on individual cases until the test is done.

Can we reuse the test set for later evaluations?

Yes, and you should. A fixed test set lets you compare vendor updates, new products and your own process changes over time. Add new cases periodically as policies change, and retire cases that no longer reflect how you work.

What if our exception notes are too thin to test with?

Thin notes are a finding in themselves. An agent cannot learn your judgment if nobody wrote it down. Start recording a cause, an action and an outcome on every exception now, and build the test set from the newer, better-documented cases.

Should we tell customers we are testing an AI agent?

A blind test on de-identified history does not involve customers directly, so notice is usually not needed at that stage. Check customer contracts for limits on sharing their shipment information with vendors. Once an agent will contact customers in production, some 3PLs choose to tell key accounts how AI-drafted messages are reviewed.

Who should grade the agent's answers?

Use two graders with operational experience, such as a senior dispatcher and an account manager, and have them grade a shared sample first to align. Where they disagree, a third person decides. Keep graders blind to which vendor produced which answer.

Related resources

See if your company qualifies

A short company assessment. No data uploads are needed.

See if you qualify