Across 47 production inspection contracts, the most expensive misread in a vendor evaluation is the same one every time: treating "99.5% accuracy" as a fact about the model. It isn't. It is a fact about a threshold someone chose on a validation set you never saw. An accuracy number quoted without its operating point is not a claim you can check. It is a number the vendor picked.
Here is the part most buyers never get shown. A vision model does not output "pass" or "fail." It outputs a score — a number between 0 and 1 for how confident it is that a part is defective. The pass/fail verdict only exists because an engineer drew a line across those scores and said: above this, reject; below this, keep. That line is the threshold. Move it, and every headline number moves with it, on the same model, on the same images.
What the score actually is
Think of two piles of parts coming off a line: good ones and bad ones. The model scores every part. Plot those scores and you get two humps — most good parts cluster low, most bad parts cluster high. If the two humps were perfectly separated, you could draw the line anywhere between them and never be wrong. In the real world they overlap. Some good parts score high (cosmetic marks, odd lighting, an unusual but acceptable variant). Some bad parts score low (a hairline defect that barely registers).
The threshold is where you cut through that overlap. And wherever you put it, you are trading two kinds of error against each other:
- Escapes (false negatives): bad parts scored below the line, shipped as good.
- False rejects (false positives): good parts scored above the line, scrapped or sent to manual review.
You cannot minimise both at once with a fixed model. Push the line down to catch more defects, and more good parts get rejected. Push it up to stop scrapping good product, and more defects escape. This is the whole game, and it is invisible in a one-number accuracy claim.
The same model, two different "accuracies"
A worked example makes it concrete. Take one model on one line and set the threshold at 0.62. Say it escapes roughly 3 defects in every 10,000 parts and false-rejects about 1.1% of good parts. Now move the single threshold up to 0.80 — nothing else changes, same weights, same camera. Escapes climb, because borderline defects now fall below the line; false rejects drop, because fewer good parts clear the higher bar. Report the first setting and you can advertise a very low escape rate. Report the second and you can advertise a very low false-reject rate. Both are "the model's accuracy." Both are true. Neither is the whole truth.
This is why an operating point matters more than a percentage. The honest version of an accuracy claim is three numbers, not one: the threshold, the escape rate at that threshold, and the false-reject rate at that threshold — all measured on parts the model was not trained on.
What this means when you are buying
The category error has a cost. A plant manager who signs off on "99.9% accurate" and later finds the line scrapping 4% of good product has not been lied to, exactly. They were shown one end of a trade-off and left to assume it held everywhere. The fix is not to distrust the number. It is to ask the question that pins it down: at what operating point, and measured on which parts?
The operating point is also a business decision, not a purely technical one, and it should sit with the people who own the consequences. A line where an escaped defect reaches a customer and triggers a recall wants the threshold low and accepts more manual review. A line where good product is expensive and an escape is caught cheaply downstream can afford the threshold higher. The model does not know your cost of a miss versus your cost of a false reject. You do. A platform that lets your own team see the score distribution and move the line is giving you that decision. One that hard-codes the threshold in the vendor's environment has made the decision for you — and will revisit it only through a change order.
That control is the same architectural pattern that lets the Auto Parts customer (Client A) hold 99% detection at 270 units per hour across 8,000-plus variants: the operating point is tuned per part family by the people running the line, not fixed once by an integrator and left to drift. The threshold is a dial, and the plant holds it.
The one diagram worth asking for
When a vendor quotes accuracy, ask to see the score distribution (the two humps) with the proposed threshold drawn on it, measured on a held-out set of your kind of parts. If they can show it, you can see exactly how much margin sits between good and bad, and what moving the line would cost you in either direction. If they can only show you a single percentage, you have learned something too: the number was chosen, and you are not being shown the choice.
FAQ
What is a confidence score in machine vision? It is the model's output for a single part: a value, usually between 0 and 1, representing how strongly the model believes the part is defective. The pass/fail verdict is produced by comparing that score to a threshold, not by the model directly.
What is a detection threshold? The cut-off applied to confidence scores to turn them into decisions. Scores above the threshold are flagged as defects; scores below are passed. Changing the threshold changes the reported accuracy without changing the model.
Why does the same model report different accuracy numbers? Because accuracy depends on the threshold. A lower threshold catches more defects but rejects more good parts; a higher threshold does the reverse. A quoted accuracy is only meaningful alongside its operating point and the dataset it was measured on.
Send us a labelled batch from one line, a few hundred confirmed-good and confirmed-defective parts the model has never seen, and within two weeks we return the score distribution for your parts, the escape-versus-false-reject trade-off plotted across thresholds, and the operating point we would recommend for your cost of a miss. No contract until you have seen where your line actually sits.
Send a labelled batch and get your line's score distribution and operating point back in two weeks.
