FreighAI
FreighAI · Playground article
Playground · Checklist article

Why should you ask an AI vendor to show you a failure?

Question A3 of the twelve. Anyone can show you a product working. What tells you whether it has run at real volume is how the vendor talks about the day it did not work.

This article goes deeper than the sheet: what actually breaks when software reads a freight enquiry, the four traps in the answer, and how to turn what you heard into a written decision.

The short answer

Ask any AI vendor to show you one enquiry it got wrong and what happened next. A vendor running at real volume has failures and can describe one in detail. A vendor who says it has never happened is telling you they have not run enough work to find out.

Why does a failure story tell you more than a demo?

A demo is a win reel. That is not a criticism, it is what a demo is for.

A demo shows the product working when everything it needs is present. The enquiry has a port pair, a weight, dimensions and an Incoterm. The rate is in the system. Nothing in your week looks like that.

Failure is where the edges of a product sit. It is the only place you find out what the software does when the weight is in pounds, the port is one of four with the same name, and the enquiry arrived as a photograph of a screen.

There is a second reason, and it is not about the software. Volume produces failure. A vendor who has processed a few hundred thousand enquiries has a ranked list of things that went wrong, with fixes against most of them. A vendor who has processed four hundred has anecdotes. When somebody tells you nothing has ever gone wrong, the useful reading is not that the product is unusually good. It is that they have not run enough work to find out.

A3 is really two questions joined together. “Show me one it got wrong” asks for a case. “Tell me what happened next” asks whether they have a process. The second half does the work, and only if you ask it in the same breath.

Asking A3 inside a real call
    1. 01Ask both halvesThe case and what happened next, in one breath.
    2. 02Let the silence runThe pause before the answer is worth as much as the answer.
    3. 03They name a caseOne enquiry, one customer, one lane.
    4. 04Ask what changedThe fix, and how they know it holds.
    1. 05Ask what reached a customerThe follow-up that separates a save from a failure.
    2. 06Write it downOne line, their words, the same day.
    3. GATEYou decideAgainst a rule you set before the meeting.
The question inside a vendor meeting: ask both halves together, let the silence run, hear the case, ask what changed, ask what reached a customer, write it down the same day. Watch it fail, then decide for yourself.

A vendor who has a case ready reaches for it in about two seconds. A vendor who is constructing one takes longer, and you will hear it.

What does a good answer to A3 actually sound like?

It is specific, slightly uncomfortable, and it ends with a change rather than an apology.

A strong answer has six parts. Few vendors give all six unprompted. Ask for the missing ones by name.

The six parts of a failure story worth accepting.
  1. The case
    Why it matters
    One enquiry, one customer, one lane. If it stays general, it did not happen to anybody in particular.
  2. The detection
    Why it matters
    Who noticed, and how long it took. A failure found by the customer is a different class of problem from one found by a check.
  3. The cost
    Why it matters
    What it actually cost. A margin, a re-quote, a relationship, a credit note, or nothing at all.
  4. The fix
    Why it matters
    What changed in the product afterwards, described so that you could recognise it in a release note.
  5. The proof
    Why it matters
    How they know the fix holds. Usually a test case that now exists because of that enquiry.
  6. The threshold
    Why it matters
    The size of thing that triggers a written follow-up at all.

That last row is the one most buyers never ask about, and it is the most revealing. Teams that take failure seriously write the trigger down in advance. The Google SRE book, in its chapter on postmortem culture, lists triggers such as data loss of any kind, a resolution time above some threshold and a monitoring failure. A team which has thought about failure can tell you where their line sits without inventing one on the spot.

Two things a good answer will not do. It will not name the customer, because a vendor who discusses one customer’s bad week with you will discuss yours with somebody else. And it will not indict a person. The same chapter makes the standard explicit: a review is blameless when it focuses on the contributing causes “without indicting any individual or team”.

What actually goes wrong when software reads a freight enquiry?

Ask the question knowing what an honest answer looks like, or you cannot tell a real one from a plausible one.

These are the shapes failure takes in this particular job. None of them is pinned on any product. They are what a forwarder’s own enquiry inbox does to software.

  • The wrong place with the right name. Valencia in Spain and Valencia in Venezuela. Santiago in Chile, Cuba, Panama and Spain. Portsmouth twice. The enquiry says one word and the software picks a side.
  • Units read as the wrong units. Dimensions in inches read as centimetres, or the other way. Gross weight taken as chargeable weight. Volume weight computed on a divisor that belongs to a different mode.
  • The number that was never a number. A customer purchase order treated as a container number. A date written 03/09 in a market that reads it the other way round.
  • The rate that had already gone. A rate applied after its validity ended, or a spot rate quoted as if it were contracted, or a surcharge from the previous month carried into this month’s quote.
  • The thing nobody read. The rate that was inside a photograph of a spreadsheet in the third attachment. The condition in the last line of the agent’s email, under the signature.
  • Two enquiries wearing one coat. A forwarded WhatsApp message containing two shipments, quoted as one. A reply-all thread where the shipper changed the destination in the eleventh message.
  • The right answer to the wrong sender. A quote drafted for the customer’s competitor because both are on the thread.

Handling these afterwards is not a courtesy. NIST’s AI Risk Management Framework, which is intended for voluntary use, names manage alongside govern, map and measure as one of its four functions. A vendor who can describe what happens after something goes wrong is describing a function of a public framework, not doing you a favour.

Notice how many of these are not really software failures at all. They are the mistakes a tired coordinator makes at four on a Friday. You are not looking for a product that never errs. You are looking for one whose errors are the ordinary ones, caught in the ordinary places.

What are the four traps in the answer you get?

The answer will sound fine in the room. These four are why it can still mean nothing afterwards.

Trap one. A bug is not a failure

“We had an outage in March and fixed it in two hours” is an availability story. It tells you nothing about what the software does to an enquiry. Push it back: “I mean a quote. One that went out wrong, or should not have gone out at all.”

Trap two. Caught at the gate is a save, not a failure

Most vendors will reach for an example where the software got something wrong and the approving person spotted it. That is the gate working, and it is good news. It is also not what you asked. The checklist follow-up exists for exactly this: what is the worst thing that has ever reached one of your customers? Ask it every time.

Trap three. The failure is from the pilot

A story from an early trial two years ago, before the product was really the product, costs the vendor nothing to tell. Anchor it: “What is the most recent one? Say, in the last three months.”

Trap four. The data gets the blame

“The enquiry was badly formed” is true and irrelevant. Your enquiries are badly formed. That is the job. The version worth accepting names the badly formed input and then says what the product now does about it, which is usually to ask a question rather than guess.

Reading the answer you got
Start here

What kind of answer did the vendor just give you?

PATH 01A named case with a change behind itSpecific, recent, something changed

The strongest signal in the whole checklist. Ask for the release note or the test case that came out of it.

PATH 02A case with no changeSpecific but nothing followed

They have volume and no process. Ask what their trigger threshold is before you decide anything.

PATH 03No failure at all“That has not really happened”

Read it as an answer about volume. Fold it into D3 and ask how many forwarders are live on the product.

Three kinds of answer to A3 and what each one is evidence of. None of the three is automatically a rejection.

The third answer is only fatal when the vendor also claims heavy usage, because both cannot be true.

How do you write the answer down as a decision?

Write it the same day, in the vendor’s own words, in one line. Write the failure down while it is in front of you. By next week the tone of the room is all that will be left.

One line per vendor with three columns is enough:

The whole record of question A3, per vendor, on one line.
  1. The case
    What goes in it
    The failure they named, in their words, with the date they gave you.
  2. What changed
    What goes in it
    The fix, and whether they offered evidence for it without being asked.
  3. Checkable
    What goes in it
    Yes or no. Can somebody outside their company confirm any of it.
Recommendation

Decide the rule before the meeting, not after it. A workable one: a vendor passes A3 if the case is from this year, the change is describable, and the follow-up about what reached a customer produced a second answer rather than a repeat of the first.

Carry the same standard into the trial. If you run a pilot, agree at the start what you will both do when something goes wrong in it, because something will. A written note, who writes it, and how quickly. The SRE chapter puts the reason plainly: unless there is some formalised process of learning from incidents, they may recur indefinitely. A vendor who already has that process will agree in a sentence. A vendor who does not will want to talk about it later.

One last use for this question. Ask it of the system you already run. What went wrong on your own quotes last quarter, who noticed, what changed. If nobody can answer that either, the problem you are shopping for is not entirely a software problem.

What if the vendor is genuinely new?

Then the question changes shape, and it is fair to let it.

A company with six live forwarders has not had the chance to fail in the ways a company with six hundred has. Holding that against them turns a good question into a way of choosing the largest supplier.

What to accept from a young vendor: a failure from a trial rather than a paying customer, or one found in their own testing and described with the same specificity. What matters is whether they treat it as information or as an embarrassment.

What not to accept from anybody, at any size, is the claim that it has never happened. A vendor with six customers who says “we have had three things go wrong and here they are” is more credible than a vendor with sixty who says nothing has. A3 and D3 hold each other honest. If the failure list is empty and the live-customer count is large, one of the two is wrong, and you should ask which.

Questions

Common questions

01

Is it unreasonable to ask a vendor to describe a failure in a sales meeting?

No, and the reaction is part of the answer. Vendors who run real volume expect the question and usually have a case ready. If asking it changes the temperature of the room, that is information too.

02

What if the failure they describe sounds serious?

A serious, specific, well-handled failure is a better signal than a small vague one. Look at detection and response rather than severity. A vendor who found it themselves, told the customer first and changed something afterwards has shown you their process working, which is more than a clean record shows you.

03

Should I ask this in the first meeting or the second?

First. It belongs to part A of the checklist with the other evidence questions, and it costs five minutes. If the answer is thin there may not need to be a second meeting.

04

Does this question work on a vendor selling something other than quoting?

Yes. Change the noun. Ask a tracking vendor about the update that was wrong, an invoice vendor about the bill that matched the wrong job. The shape of the question does not change with the job.

READY WHEN YOU ARE

Take question A3 into your next vendor call

It is the cheapest question on the sheet and the one vendors prepare for least. Ask both halves in one breath, ask what reached a customer, and write the answer down before the day ends. Then ask us the same thing.