Why we chose where a model belongs in an operational workflow for freight logistics
The reasoning behind where a model belongs in an operational workflow, including the part we expect to age badly.
Written in the form we use internally, because the useful part of a decision record isn't the decision. It's the context that made it reasonable. In two years someone will want to reverse this, and they deserve to know what we knew at the time.
A statistical component is a good fit for suggesting, ranking and drafting, and a poor fit for being the last word on anything with a consequence. Most disappointing deployments are placement errors rather than accuracy problems.
SUGGEST, RANK, DRAFT
Those three placements share a property: a wrong answer is visible and cheap. A bad suggestion gets ignored. A bad ranking costs a scroll. A bad draft gets edited. Nothing irreversible happens because the component was wrong, and the value shows up anyway as saved time.
Compare that with deciding, sending or committing, where a wrong answer is expensive and often invisible until much later. Same accuracy, completely different risk profile, because the difference is placement rather than quality.
UNCERTAINTY IS THE INTERFACE
The most important thing a statistical component can report is how sure it is, and the second most important is that this number is calibrated. When it says seventy percent it's right about seventy percent of the time. Without that you can't route anything, and routing is where the operational value is.
With it, the design gets much simpler. High confidence proceeds, low confidence goes to a person, and you can set that threshold from the cost of being wrong rather than from a hunch. You can also move it as you gather evidence.
We modelled it at roughly 40,000 loads a month and the difference only showed up in the tail. At median load you couldn't tell them apart. At the ninety-ninth percentile, one of them stopped being able to explain itself.
EVALUATE ON YOUR OWN WORK
General benchmarks tell you very little about performance on one organisation's documents, vocabulary and edge cases. What tells you something is a few hundred examples from their actual work, labelled by someone who knows the domain, held back and re-run on every change.
Building that set is the least glamorous and most valuable part of the project. It's also what turns "it seems better" into a number, which is the only way these conversations stay honest over time.
The deciding factor was regulatory, not technical. Dispatchers have to be able to reconstruct why a given load tender was handled the way it was, months later, in front of someone unfriendly. That killed two of the three options on the spot.
WHERE IT GOES WRONG
- Placing a statistical component where a wrong answer is irreversible and invisible.
- Uncertainty scores that aren't calibrated, so no threshold can be set from them.
- Judging suitability on general benchmarks rather than a few hundred of the client's own cases.
- No held-back evaluation set, so every change is assessed on impressions.
Put it where being wrong is cheap and visible. Route on calibrated uncertainty.
CONSEQUENCES WE ACCEPTED
We took a slower first two months in exchange for a system you can still reason about in year three. On an eighteen-month horizon we'd have chosen differently, and we said so at the time.