FIELD NOTEAUTOMATION

Where a model belongs in an operational workflow, seen up close in discrete manufacturing

What where a model belongs in an operational workflow actually looked like from inside a contract manufacturer running three lines.

FILED
READ
AUTHOR
REF

We spent six weeks inside a contract manufacturer running three lines before we proposed anything. The brief said the problem was reporting. It wasn't reporting. Two days of watching production planners work through about 30,000 units a week made that obvious, and dashboards had nothing to do with it.

A statistical component is a good fit for suggesting, ranking and drafting, and a poor fit for being the last word on anything with a consequence. Most disappointing deployments are placement errors rather than accuracy problems.

SUGGEST, RANK, DRAFT

Those three placements share a property: a wrong answer is visible and cheap. A bad suggestion gets ignored. A bad ranking costs a scroll. A bad draft gets edited. Nothing irreversible happens because the component was wrong, and the value shows up anyway as saved time.

Compare that with deciding, sending or committing, where a wrong answer is expensive and often invisible until much later. Same accuracy, completely different risk profile, because the difference is placement rather than quality.

UNCERTAINTY IS THE INTERFACE

The most important thing a statistical component can report is how sure it is, and the second most important is that this number is calibrated. When it says seventy percent it's right about seventy percent of the time. Without that you can't route anything, and routing is where the operational value is.

With it, the design gets much simpler. High confidence proceeds, low confidence goes to a person, and you can set that threshold from the cost of being wrong rather than from a hunch. You can also move it as you gather evidence.

The clearest thing we saw was how production planners handled a contested work instruction. On paper it's one step. In practice it's five, three of them over the phone, none of them written down. Which is why nobody could ever explain schedule adherence to their director.

EVALUATE ON YOUR OWN WORK

General benchmarks tell you very little about performance on one organisation's documents, vocabulary and edge cases. What tells you something is a few hundred examples from their actual work, labelled by someone who knows the domain, held back and re-run on every change.

Building that set is the least glamorous and most valuable part of the project. It's also what turns "it seems better" into a number, which is the only way these conversations stay honest over time.

Here it showed up as a queue nobody owned. About 30,000 units a week went through it, and production planners had learned to check it twice a day because the alternative was a revised instruction reached the floor after the batch had run. A better queue wasn't the answer. Making ownership a property of the work instruction was.

WHERE IT GOES WRONG

  • Judging suitability on general benchmarks rather than a few hundred of the client's own cases.
  • No held-back evaluation set, so every change is assessed on impressions.
  • Placing a statistical component where a wrong answer is irreversible and invisible.
  • Uncertainty scores that aren't calibrated, so no threshold can be set from them.

Put it where being wrong is cheap and visible. Route on calibrated uncertainty.

WHAT WE TOOK AWAY

The work shipped and schedule adherence moved, but the thing we're proudest of is smaller than the system: production planners stopped keeping a private spreadsheet. That's usually the honest signal that the model finally matches the job.

RELATED
SAME GROUND, DIFFERENT ANGLE
ALL TRANSMISSIONS