Where a model belongs in an operational workflow: a cost model for insurance claims
The arithmetic behind where a model belongs in an operational workflow, with the assumptions written out so you can disagree honestly.
The question is never whether something is worth doing in the abstract. It's whether it's worth doing at your volume, with your failure rate, valuing your team's time properly. So let's do the sum.
A statistical component is a good fit for suggesting, ranking and drafting, and a poor fit for being the last word on anything with a consequence. Most disappointing deployments are placement errors rather than accuracy problems.
SUGGEST, RANK, DRAFT
Those three placements share a property: a wrong answer is visible and cheap. A bad suggestion gets ignored. A bad ranking costs a scroll. A bad draft gets edited. Nothing irreversible happens because the component was wrong, and the value shows up anyway as saved time.
Compare that with deciding, sending or committing, where a wrong answer is expensive and often invisible until much later. Same accuracy, completely different risk profile, because the difference is placement rather than quality.
UNCERTAINTY IS THE INTERFACE
The most important thing a statistical component can report is how sure it is, and the second most important is that this number is calibrated. When it says seventy percent it's right about seventy percent of the time. Without that you can't route anything, and routing is where the operational value is.
With it, the design gets much simpler. High confidence proceeds, low confidence goes to a person, and you can set that threshold from the cost of being wrong rather than from a hunch. You can also move it as you gather evidence.
At about 15,000 claims a month with a two percent exception rate, adjusters were absorbing about nine hours of manual reconciliation a week. That's the number the build had to beat, and it's a lower bar than anyone in the room expected.
EVALUATE ON YOUR OWN WORK
General benchmarks tell you very little about performance on one organisation's documents, vocabulary and edge cases. What tells you something is a few hundred examples from their actual work, labelled by someone who knows the domain, held back and re-run on every change.
Building that set is the least glamorous and most valuable part of the project. It's also what turns "it seems better" into a number, which is the only way these conversations stay honest over time.
The interesting term wasn't engineering time. It was the cost of a reserve was released twice against the same loss landing once in the wrong quarter, which the client could size to the pound and we couldn't size at all.
WHERE IT GOES WRONG
- Uncertainty scores that aren't calibrated, so no threshold can be set from them.
- Judging suitability on general benchmarks rather than a few hundred of the client's own cases.
- No held-back evaluation set, so every change is assessed on impressions.
- Placing a statistical component where a wrong answer is irreversible and invisible.
Put it where being wrong is cheap and visible. Route on calibrated uncertainty.
WHERE THE MODEL BREAKS
Do the numbers before the meeting, not during it. A decision that survives arithmetic tends to survive the next reorg too, because the reasoning outlives the people who made it.