Skip to main content
Technical Analysis
11 min read

What goes in the sample box: preparing parts for an AI vision trial that does not waste 4 weeks

The quality of parts in a sample box determines whether an AI vision model learns your actual defect taxonomy or a statistical artifact. This post identifies the most common pilot failure—uncurated sample boxes that extend timelines by three to four weeks—and maps what a production-ready training set requires: boundary cases, class-labelled defects, good-part variation across production conditions, and the questions to ask vendors before acceptance. Proper sample preparation is the rate-limiting factor in AI vision deployment, not the model.

What goes in the sample box: preparing parts for an AI vision trial that does not waste 4 weeks

A vision trial that starts with 1,000 loosely selected production parts and no labelled defect examples will produce a model with unverifiable performance and an acceptance test that cannot be written. The four-week extension is then a data-collection exercise the integrator should have completed before the trial began. The parts in the sample box determine whether the model learns your defect taxonomy or a statistical proxy for it.

This is the most common failure mode in AI vision pilots. In trials where sample boxes arrived without boundary cases or class-labelled defects, the initial model required an additional collection and retrain cycle before shadow-mode results were usable for acceptance review. Three to four weeks on the schedule, attributable to the sample box, not the model. The AI is blamed; the sample preparation is not examined. This post covers what a sample box actually needs to contain, how to structure it for a model that performs reliably in production, and the questions to ask any vendor who accepts an uncurated box of parts and promises results.


Why sample quality is the rate-limiting factor

An AI vision model learns the boundaries between acceptable and unacceptable. Those boundaries are defined by the samples it is trained on. If the training set shows only central-distribution good parts, the model will not recognise marginal good parts that fall near tolerance limits. If the training set contains only one defect type, the model will flag similar-looking defects and miss dissimilar ones.

One practitioner responsible for qualifying inspection equipment across multiple electronics lines described the pattern: "Every pilot that struggled came in with a bin of production rejects from the last three months. Half of them were the same defect class, photographed under different shifts' lighting. The model learned to find that one defect class with high confidence. In production, three defect classes it had never seen proceeded straight through."

HyperQ AI Vision initialises on 1,000 images. That number assumes the images cover the range of conditions the deployed model will encounter: multiple defect classes, boundary cases at tolerance edges, and good-part variation that reflects actual production spread. A 1,000-image set weighted entirely toward one defect class is not a representative 1,000 images; it is a highly confident single-class detector and an unreliable everything-else detector.


The four components of a sample box

1. Representative good parts

Good parts should not all come from a single production run, a single shift, or a single operator. The goal is to represent the full visual spread of acceptable output.

Collect at least 400 good-part images from a minimum of 3 different production runs, spanning shifts if the line operates across multiple shifts. If surface appearance is expected to vary by batch, raw material lot, or tooling age, include parts from the beginning and end of a tooling life.

For lines with high SKU count, the sample requirement scales per SKU — 400 good-part images per SKU class, or per the distinct visual profile if multiple SKUs share a surface treatment. On a high-mix line, a pragmatic first deployment targets the 10 highest-volume SKUs and builds a shared-defect model where surface appearance clusters.

2. Boundary cases: worst-tolerable good parts

Boundary samples are the most valuable and most commonly omitted component of a sample box. A boundary case is a part that a quality engineer, looking at it on a light table, would classify as acceptable — but only just. It falls at the outer edge of the tolerance band.

The model must learn that boundary cases are good parts. Without them, the decision threshold calibrates to the centre of the good-part distribution, and marginal-but-acceptable parts generate false positives in production. This is the primary driver of excessive false-reject rates in early production deployments.

Collect a minimum of 50 boundary good-part examples per defect class. The quality engineer responsible for the specification should physically select these parts and mark them as boundary-pass before they enter the sample set.

3. Confirmed defect examples, labelled by class

The model needs confirmed defects for each defect class in the defect catalogue. Each defect example should:

  • Be confirmed as a reject by a quality engineer, not assumed
  • Be labelled with the defect class (scratch, void, pit, flash, discolouration, dimensional out-of-spec, etc.)
  • Cover the visible severity range within the class — a hairline scratch and a deep gouge are both "scratch" but look different to the model

Minimum sample targets by defect class:

Defect class Minimum confirmed samples Notes
Surface scratch / scuff 80–100 Include directional variation
Void / pit / porosity 80–100 Include size range from minimum-reject up
Dimensional / flash 60–80 Cover both under- and over-tolerance
Discolouration / stain 60–80 Cover lighting angle variation
Assembly / positional error 40–60 Include marginal-fail cases
Rare / critical defect class 20–40 Can be supplemented with synthetic augmentation

For defect classes where the line generates fewer than 20 confirmed examples in a reasonable collection window, note this as a risk before the trial begins. A class with fewer than 20 examples will produce a model with low confidence on that class; the acceptance test should reflect this.

4. Worst-tolerable reject examples (near-boundary rejects)

The mirror of boundary good parts is the worst-tolerable reject: a part that fails, but only barely. Including near-boundary rejects teaches the model where the reject threshold starts, not just where it is obvious.

A minimum of 20–30 near-boundary rejects per defect class improves performance on the cases most likely to generate false accepts in production. These are the defects that escaped manual inspection, the ones that field-service logs traced back to inspection misses. If the defect catalogue tracks escapes, those confirmed escapes should form the nucleus of this set.


Capture conditions matter as much as part selection

The model will be deployed under production lighting. It should be trained under production lighting, or a defined proxy with documented deviation.

Three common capture errors that extend pilots:

Overhead office lighting. Shadow angle, colour temperature, and intensity differ from production station lighting. A model trained under fluorescent ceiling lights deployed under LED strip lighting on a dark machine enclosure will encounter images it has not seen.

Single-angle capture. Many defect classes (scratches, die marks, sink marks) are angle-dependent. A scratch invisible at 45° may be highly visible at 10°. If the production camera captures parts at an acute angle, the training captures should replicate that angle, not a convenient perpendicular setup.

Inconsistent background. If the production station uses a black or white background behind the part, the training images should use the same. A model trained on parts photographed against cardboard will see a different feature landscape than the same model sees in a dark machine enclosure.

The capture specification for the sample box should state: camera model and focal length, lighting type and position, working distance, background, and the number of images per part position if multi-angle inspection is planned.


What to do when a defect class is underrepresented

Short-run production, low overall defect rates, or rare critical defect classes may not yield sufficient confirmed samples before the trial deadline. The options, in order of preference:

  1. Extend the collection window. If the trial deadline is flexible, 4 weeks of systematic reject collection at a facility running 11,520 units per day will generate substantial confirmed defect data even at low defect rates. Calculate the expected yield before deciding this is infeasible.

  2. Source from engineering specimens. Quality engineers at most facilities have physical reference standards — parts deliberately manufactured out of spec for inspection gauge calibration. These can supply boundary and near-boundary examples that production rejects do not.

  3. Document the gap and adjust the acceptance test. If a defect class arrives at the trial with fewer than 20 confirmed samples, the acceptance test should explicitly exclude or modify the detection-rate requirement for that class. A pilot that produces high confidence on 4 of 5 defect classes and acknowledged uncertainty on the fifth is a useful result; a pilot that produces claimed confidence on 5 of 5 with an undertrained class is a liability.

  4. Synthetic augmentation. For geometric defects with well-defined visual signatures — voids, flashes, dimensional errors — image augmentation on confirmed examples can supplement a thin training set. This should be disclosed in the acceptance test documentation as augmentation-supplemented, not treated as equivalent to confirmed physical rejects.


The vendor question to ask before the trial starts

Vendors who accept any box of parts and promise a detection rate within two weeks have not told you what they will do when the model encounters a defect class absent from your box. The question to ask before the trial begins:

What is your minimum sample requirement per defect class, and what is your policy when a required defect class arrives with fewer than the minimum?

A vendor who answers "we need X confirmed examples per class, and if you cannot supply them, the trial acceptance criteria should be scoped to exclude that class pending further data" is being honest about what AI vision training requires. A vendor who answers "we can work with whatever you send us" has either a very forgiving product — or has outsourced the risk to your acceptance test.


Sample box checklist

Item Target quantity Format Notes
Good-part images, central distribution 400+ Labelled "pass", tagged with production date + run Minimum 3 production runs
Boundary good-part images 50+ per defect class Labelled "pass-boundary", engineer-approved Selected by QC engineer
Confirmed defect images per class 60–100 per class Labelled with class name, severity Confirmed by QC, not assumed
Near-boundary reject images 20–30 per class Labelled "fail-boundary", engineer-approved Priority: escaped defects
Capture specification document 1 document Camera, lens, lighting, working distance, background Must match production station
Defect catalogue 1 document List of all classes, definitions, engineering references Flag any class with < 20 samples

Frequently asked questions

How many parts do I need for an AI vision trial?

HyperQ AI Vision initialises on 1,000 images. The breakdown that produces a reliable deployment is roughly: 400 good parts spanning multiple production runs, 50+ boundary good-part examples per major defect class, and 60–100 confirmed defect examples per class. A set of 1,000 undifferentiated production parts does not satisfy this requirement.

Can I use historical reject photos from our internal QC database?

Yes, subject to two conditions: they must be labelled by defect class (not just "rejected"), and they should have been captured under conditions consistent with the production inspection station. Smartphone photos under overhead office lighting will introduce a domain shift the model must compensate for. Usable archive images are a useful supplement, not a substitute for a structured sample collection.

What if we only have one or two examples of a rare defect class?

Flag the class as undertrained before the trial begins. Adjust the acceptance criteria to reflect this — require demonstration of detection capability only on classes with sufficient training data, and document the rare-class gap as a known limitation. Collecting additional examples from engineering reference standards or extended production monitoring is preferable to accepting an unsubstantiated detection claim on a thin sample set.

Should the sample box include parts that passed inspection but later failed in the field?

If confirmed field-failure parts are available, yes — these are the highest-value defect examples because they represent the exact condition the production system previously missed. Confirmed escapes should form the nucleus of the near-boundary reject set.

Who should select the boundary good-part samples?

A quality engineer with authority over the specification, not a production operator or the AI vendor's integration team. Boundary-pass classification is a quality decision, not a visual judgement call. The person who signs the FAI is the appropriate selector.

How long does sample collection typically take?

On a line running hundreds of units per day, a systematic 2-week collection yields a workable initial set for most defect classes. For lines with low defect rates or rare defect classes, 4–6 weeks is a realistic window. The collection period is shorter than the delay caused by starting a trial with an insufficient sample box.


For the method comparison between reference matching and learned models, including the decision table by SKU count, defect variety, and audit regime, see golden-sample matching vs learned defect models. For the HyperQ AI Vision product overview and typical deployment scope, see the AI Vision solution page.


Send us your defect catalogue before the box ships. We will review the class list, flag any classes likely to arrive undertrained, and return a structured sample-collection plan within 5 working days — specifying target quantities and capture conditions for your line. No contract until the trial passes your own acceptance criteria.

Request the sample plan

Written by

Hypernology Team

August 25, 2026

Share

Continue Reading

Translate Insight
to Infrastructure.

Interested in deploying these solutions to your facility? Let's discuss the technical requirements.

Initiate Briefing