FreighAI
FreighAI · Playground article
Playground · Checklist article

Run it on twenty of my real enquiries from last week. Can I watch, and can I pick which twenty?

This is question A1 on the checklist, and it decides whether the other eleven are worth asking. Every vendor selling AI to freight forwarders has a demo that works. The demo is not the product. It is a film of the product, shot on a set the vendor built.

Twenty of your own enquiries, chosen by you, run in front of you, this week. That is the whole test. This article is how to run it and how to read the result.

The short answer

Ask the vendor to run twenty enquiries you choose, from last week, unedited, while you watch. It is the only test that uses your own mess rather than their clean sample. A vendor who can do it this week is showing you the product. A recorded demo is showing you a film.

Why does a demo prove so little?

Because the enquiry in a demo was chosen by the person selling you the software.

Look at what a demo enquiry usually contains. A clean subject line. An origin and a destination, spelled correctly. A commodity. Gross weight and dimensions. An Incoterm. A ready date. Nothing contradicts anything else, and nothing important is missing.

Now look at Monday morning in your own pricing inbox. A WhatsApp forward with a photograph of a packing list, taken at an angle. A one-line message that says “pls quote urgent” with an attachment called final_final_2.xlsx. A reply on a thread from March, where the new enquiry sits three lines under six screens of quoted history. An enquiry with dimensions but no weight. Two of the twenty are not enquiries at all. One is a rate sheet from an agent, and one is somebody asking where their cargo is.

The gap between the demo enquiry and Monday morning is the entire product. Everything a vendor is charging you for lives in that gap. A demo steps over it.

The NIST AI Risk Management Framework, which is intended for voluntary use, puts the same idea in more careful language. Under its measure function, “AI system performance or assurance criteria are measured qualitatively or quantitatively and demonstrated for conditions similar to deployment setting(s)”. Conditions similar to the deployment setting. Your pricing inbox is the deployment setting. Their sample data is not.

Public buyers work the same way. The UK guidance on responsible AI in recruitment, written for a very different kind of purchase, still says that “prior to the deployment of this system, it is recommended that your organisation pilots the technology ...”, and that “understanding how the model performs in your organisation’s real-world environment ... is essential before deploying any model at scale”. A forty-person forwarder cannot run a six-month pilot. Twenty enquiries and an afternoon is the version that fits.

How do you run the twenty-enquiry test?

The test has seven steps. It costs you about two hours in total, and most of that is spent watching.

The protocol. The third column is what each step is actually testing.
  1. 1. Pick the twenty
    What you do
    Open last week’s pricing mailbox and take the first twenty enquiries in date order. Do not curate them.
    What it tests
    Whether the sample is yours or theirs.
  2. 2. Keep them exactly as they arrived
    What you do
    No renaming, no retyping, no cleaning, no trimming the quoted history. Forward the original messages with the attachments intact.
    What it tests
    Whether the product reads real input or prepared input.
  3. 3. Send them the same day as the call
    What you do
    Not a week before. A vendor with seven days can hand-build twenty configurations.
    What it tests
    Whether you are seeing the product or a weekend of work.
  4. 4. Ask to watch it run live
    What you do
    Screen shared, one enquiry at a time, in the order you sent them.
    What it tests
    Whether anything is happening off-screen.
  5. 5. Watch what it asks for
    What you do
    A good product stops on the ones missing weight or dimensions and drafts the question rather than guessing.
    What it tests
    Whether it invents numbers when it is short of them.
  6. 6. Record every one
    What you do
    One line each. What it produced, what it got wrong, and whether a person would have had to redo it.
    What it tests
    Whether you remember anything a week later.
  7. 7. Keep the twenty
    What you do
    Use the same twenty on the next vendor, in the same order.
    What it tests
    Whether you are comparing products or comparing salespeople.

The last step is the one people skip and the one that pays. Twenty enquiries used once is an anecdote. The same twenty run by three vendors is comparable evidence.

Recommendation

Book two hours, not one. The conversation after the run is where you learn what the vendor actually understood about the ones that failed.

The test, end to end
    1. 01Pick twentyLast week, in date order, uncurated.
    2. 02Send uneditedOriginal messages, attachments intact.
    3. 03They run it liveScreen shared, in the order you sent them.
    4. 04Watch each oneEspecially the ones missing a weight.
    1. 05Record the outcomeOne line per enquiry, the same day.
    2. 06Repeat, same twentyThe next vendor gets the identical sample.
    3. GATEYou decideNobody signs on the strength of a demo.
The twenty-enquiry test from picking the sample to the decision. Every step before the decision belongs to your business, and the decision itself is the only gate.

Twenty is a practical number, not a statistical one. Four or five of them will decide the result, and they are never the four or five anybody would have chosen.

Which twenty enquiries should you pick?

Take them in date order and you will get a fair sample by accident. If you would rather choose, choose for variety of mess rather than variety of lane.

The kinds worth having in the twenty.
  1. A forwarded WhatsApp with a photo attachment
    Why it belongs in the twenty
    It is how a large part of your work arrives, and it defeats anything that only reads email text.
  2. An enquiry with dimensions but no weight
    Why it belongs in the twenty
    It tests whether the product asks for what is missing or quietly assumes it.
  3. A reply buried under six screens of quoted history
    Why it belongs in the twenty
    It tests whether the product reads the conversation or only the last message.
  4. A revision request on a quote you sent three weeks ago
    Why it belongs in the twenty
    It tests whether the product can find the earlier quote and change one line of it.
  5. A multi-leg or transhipment enquiry
    Why it belongs in the twenty
    It tests whether the routing survives anything that is not port to port.
  6. A lane you hold no rate for
    Why it belongs in the twenty
    It tests the honest answer, which is what the product does when it cannot price something.
  7. An enquiry from a customer with a special commercial arrangement
    Why it belongs in the twenty
    It tests whether your own rules reach the draft.
  8. Something that is not an enquiry at all
    Why it belongs in the twenty
    It tests whether the product knows the difference. Two of the twenty should be this.

That last row is worth insisting on. A product that prices an agent’s rate sheet as a customer enquiry will do it again at eight in the morning, when nobody is watching.

The formats matter more than most buyers expect. FreighAI, for example, states that it “Reads PDF, Excel, Word, forwarded email, WhatsApp messages, photos and scans”. Whatever the vendor in front of you claims on that point, your twenty are where you find out, because a real week contains at least three formats nobody chose to demonstrate.

What does a good run look like, and what does a weak one look like?

You are not scoring accuracy. You are watching behaviour.

What to watch for while the twenty run.
  1. The ones with something missing
    A good run
    Stops, names what is missing, drafts the question to the customer.
    A weak run
    Fills the gap with a plausible number and carries on.
  2. The unreadable attachment
    A good run
    Says it could not read it and routes it to a person.
    A weak run
    Produces a quote anyway, or fails without saying so.
  3. The enquiry that is not an enquiry
    A good run
    Recognises it, classifies it, sends it somewhere sensible.
    A weak run
    Prices the agent’s rate sheet.
  4. Speed
    A good run
    Roughly the same on the twentieth as on the first.
    A weak run
    Fast on the first three, then the vendor starts talking over it.
  5. The vendor’s commentary
    A good run
    Names which ones it will struggle with before running them.
    A weak run
    Explains each failure only after it appears on screen.
  6. The output itself
    A good run
    A draft your pricer would recognise and could approve or amend.
    A weak run
    Something that has to be rewritten, which is not a draft.

The strongest signal is in the fifth row. A vendor who looks at your twenty, says “four, eleven and seventeen will be poor, and here is why”, and is then proved right, has run at real volume. One who explains each failure afterwards is meeting your kind of work for the first time.

The worst signal is a run with no failures at all. Twenty real enquiries from a real forwarding mailbox contain at least two the product should not get right. If nothing failed, you are not watching your enquiries.

What are the traps?

Four, and all four are ordinary rather than dishonest. They happen because the person selling wants a good hour.

  • 01The overnight trap. You send the twenty on Monday for a Thursday call, so what you watch includes three days of configuration and rate loading done for your twenty. Send them the same day.
  • 02The sandbox trap. The screen is not the product a customer runs. It is an internal build with the parts that break switched off. Ask whether this is the environment their existing customers are on, while the screen is still shared.
  • 03The self-selected sample. You send twenty and eleven get run, quietly, because nine were awkward. Number them before you send them, and ask for all twenty back in order.
  • 04The quiet person in the loop. Between your file arriving and the draft appearing, somebody on the vendor’s side is finishing the work by hand. The hardest one to see, and the easiest to expose.

The fourth trap has a counter that costs you nothing. Keep two enquiries back and produce them in the meeting. Nothing on this list survives an enquiry the vendor has never seen.

Note what the test does not tell you. Nothing about behaviour at volume in peak season, nothing about how long it takes to run inside your business, nothing about who owns the data it just read. Those are questions B1, B3 and C1.

How do you turn the run into a decision?

Write four things down the same day, for each vendor. Not later, and not from memory.

  • How many of the twenty produced a draft a pricer could have approved.
  • How many failed, and whether the vendor named them in advance.
  • What the product did when something was missing.
  • What the vendor said they would change, and by when.

Then read which of the three endings you actually got.

Reading the result
Start here

The vendor has run your twenty. What did the hour prove?

PATH 01It ran live, on your twenty, and you watchedNothing was prepared

The failures you saw are real failures and the drafts you saw are real drafts. This is evidence. Take the same twenty to the next vendor.

PATH 02It ran, but on a prepared copyThey held your files for days

You have learned what the product does with several days of setup per twenty enquiries. Worth knowing, and not the same claim. Ask for five fresh ones live.

PATH 03It did not run at allYou were offered a recording

You have learned the most useful thing of the afternoon. Ask the follow-up from the checklist: then let me send you twenty right now.

The three ways a twenty-enquiry test ends, and what each one entitles you to conclude.

The strongest position a buyer can be in is the same twenty enquiries, run by two vendors, in front of the same three people from your business. It takes a week and costs nothing. The UK Guidelines for AI procurement describe themselves as best practice for public-sector buyers, written to help them “evaluate suppliers”. The twenty-enquiry test is what that looks like at the scale a forwarder buys.

One more thing to carry into the next call. The vendor who has just run your twenty will now quote you a percentage. Ninety of a hundred went straight through, or eighty, or seventy-five. That number is question A2, and it means nothing until somebody defines it.

Questions

Common questions

01

Why twenty, and not five or a hundred?

Five is a demo with extra steps. A hundred turns into a project, gets scheduled and quietly becomes a paid pilot. Twenty fits in an afternoon and still contains the mess.

02

What if the vendor wants to load our rate cards first?

Let them, but run the twenty twice. Once before the setup and once after. The gap between the two runs is the honest answer to how much of the result depends on configuration, and that belongs to question B3.

03

Should we anonymise the enquiries first?

Only where a customer name or a rate is genuinely confidential, and then change the name and nothing else. Every other edit removes the mess you are testing for. If a vendor cannot handle real work under a mutual confidentiality agreement, that is itself an answer.

04

The vendor ran our twenty and got sixteen right. Is that good?

It is unreadable on its own. The four failures matter more than the sixteen. Ask what would have happened to those four in a normal week, who would have picked them up and how anyone would have known. A product that fails safely beats one that fails silently.

READY WHEN YOU ARE

Take the test into your next vendor call

This is one question of twelve. The others cover where the rates come from, whether your TMS has to move, what happens to your data and who you are dealing with. Read all twelve before the meeting.