Skip to main content
Technical Analysis
13 min read

What to put in the sample box for an AI vision trial

This post gives a practical guide to preparing physical samples for an AI vision trial so the evaluation reflects real production risk instead of idealized demo conditions. The key takeaway is that a good sample box should include borderline parts, known escapes, and variation that truly tests defect detection performance.

What to put in the sample box for an AI vision trial

1,000 labeled images is all HyperQ AI Vision needs to initialize a production-ready defect model. Whether those 1,000 images actually cover your worst-case parts depends on what gets shipped in the sample box. Most buyers send ten pristine parts from last week's golden batch — the vendor runs a demo on material that was never going to fail anyway, the system reports 100% detection, and six weeks later the production deployment misses the defect class that escaped last quarter. The root cause is not the AI. It is the sample box.

This post is a practical guide to building a sample submission that gives a vision trial actual diagnostic power. The principles apply whether you are evaluating HyperQ AI Vision or any other system. The difference is that HyperQ's evaluation process is explicitly built around physical samples — send a sample, results in 2 weeks, no contract until the system performs against your data — which means the quality of what you ship determines the quality of the answer you receive.

Buyers control trial quality more than vendors typically admit.


Why most sample boxes prove nothing

A vision trial is a test of a hypothesis: "This system can detect the defect classes that matter on my specific parts, at my production rate, with acceptable false-positive rates." Testing that hypothesis requires samples that contain the actual variation you encounter in production — including, and especially, the variation that has caused escapes.

Vendors have an incentive to produce clean results. Clean results close deals. A buyer who ships ten perfect parts and ten parts with obvious, photographically clear defects will always get a strong demo. The system has been optimized to perform well on clear cases; the demo confirms it. But the industrial inspection problem is not about clear cases. It is about the borderline call — the part your most experienced inspector argues about with the newest hire, the surface mark that is cosmetic at 10-micrometer tolerance but critical at 5-micrometer tolerance, the lot that passed last month's AQL sample and shipped with two escapes.

A Tier-1 automotive parts supplier running HyperQ AI Vision across six production lines, handling more than 8,000 product variants and 11,520 units per day, did not get there by submitting a clean sample box. The initial trial was built around their worst-case parts: edge cases from historical escape logs, units that had triggered inspector disagreements, and parts from a supplier process-change event that had slipped through the previous system. That is the sample set that proved the hypothesis.

If your sample box does not contain that material, you are not running a trial. You are running a marketing demonstration.


The principle: worst-case parts, historical escapes, argued units

Before you build the sample box, you need three source lists.

Source 1: Your escape log. Pull the last 12 months of customer escapes or internal rework events. For each escape, identify which part type was involved and what the defect looked like. If you have physical examples, set them aside. If you do not, note the defect class — you may need to deliberately stage a comparable unit (more on that below).

Source 2: Your inspector disagreement log. Most QA teams have a de facto record of the parts that trigger disagreements among inspectors — parts that one inspector passes, another rejects, and a third sends to supervisor review. These are your borderline cases. They are the hardest class for any system to handle correctly, and they are the class most likely to expose the real performance floor.

Source 3: Your recent supplier non-conformance reports. Supplier process changes, batch variation, and substitute raw materials produce visual patterns that differ from your reference. The AI system needs to see those patterns to classify them correctly.

These three sources give you the real distribution of what your line handles. A trial run on that distribution tells you something. A trial run on golden parts tells you nothing.


The sample-box packing checklist

Build the box in five layers. You do not need to label every part in advance; a simple log sheet in the box is enough to track provenance.

Layer 1: Golden reference parts (10–15 units)

These are your baseline. Parts that definitively conform to spec, from a stable lot, with no cosmetic variation. You need them so the system can calibrate the boundary between acceptable and unacceptable. Do not over-represent this layer. Ten to fifteen units is sufficient.

Include:

  • 10–15 known-good units from the most recently accepted lot
  • At least 2 part variants if your product line has significant geometry differences (e.g., two connector pitches, two surface finishes)
  • Lot number and acceptance date noted on the log sheet

Layer 2: Known-bad parts — confirmed escapes or rework pulls (10–20 units)

These are parts you know are defective. The defect should be verified: either confirmed by a customer complaint, an internal rework decision, or a quality engineer's disposition.

Include:

  • At least 3 distinct defect classes if your escape history includes multiple types (e.g., scratch, burr, contamination, marking error)
  • For each defect class, at least 3 physical examples if available
  • Severity range within each class: at least one "severe" (unambiguous) and one "marginal" (borderline by your current standard)
  • Defect class and severity noted on the log sheet

Layer 3: Borderline units — the argued cases (5–10 units)

These are the parts your inspectors disagree about. Pull them from hold bins, supervisor-review queues, or a deliberate re-sample of recent borderline lots.

Include:

  • 5–10 units with documented inspector disagreement (pass/fail split of at least 2 inspectors)
  • A note on the log sheet indicating the disagreement: "Inspector A: pass. Inspector B: fail. Disposition: shipped."
  • If you cannot source physical borderline parts, ask your QA supervisor to select the three most ambiguous units from the next production run before they are dispositioned

Layer 4: Supplier variation parts — same drawing, different process (5–10 units)

If you have experienced a supplier process change in the last 18 months, include parts from the affected lot. These parts may be dimensionally within tolerance but visually distinct from your reference. They are exactly the case where template-matching systems produce high false-positive rates.

Include:

  • 5–10 units from a known supplier process-change lot, if available
  • Parts from alternative suppliers for the same part number, if dual-sourced
  • Supplier name (or code) and lot date on the log sheet

Layer 5: Accelerated-condition units — if relevant to your failure mode (5 units)

For some inspection applications, failure is not cosmetic but functional — a seal that will leak under pressure, a solder joint that will fail under vibration. For those cases, consider whether you can include parts that have been stress-tested to early-failure. This is optional but valuable for applications where the defect is subsurface or develops under environmental load.

Include (if applicable):

  • 5 units that have been thermal-cycled, humidity-exposed, or mechanically stressed according to your product's typical failure mode
  • Note the conditioning method and duration on the log sheet

Sample quantity and imaging implications

HyperQ AI Vision's training architecture requires approximately 1,000 images per defect class — a patented approach that is roughly 10 times less data than conventional deep-learning inspection systems need. That means a well-composed sample box of 40–60 parts, imaged from the appropriate angles, can seed the training set for a production deployment.

This has a direct implication for how you pack the box: part quantity matters less than class coverage. Ten well-chosen borderline units across three defect classes are worth more than 50 golden parts. The training pass that follows an evaluation is not starting from zero; it is extending what the system learned during the trial. A sample box with good class diversity compresses the time from trial approval to production go-live — and HyperQ's standard timeline from contract to live deployment is approximately 4 weeks, with 2 days of on-site setup.

If your sample box covers the five layers above, the imaging session during evaluation is effectively the first imaging pass for production training. That means the 2-week results window is not just a proof-of-concept: it is the first iteration of your actual production model.


What the vendor does with the box

When you ship to HyperQ for a 2-week evaluation, the process is:

  1. Intake and imaging. Every part in the box is imaged under controlled lighting at the relevant inspection angles. The log sheet you included is used to label each image with ground-truth disposition (known-good, known-bad, borderline, supplier-variation).

  2. Model initialization. The labeled images are used to initialize a defect-detection model for your specific part and defect class combination. Inspection runs at 0.3–1.0 seconds per unit; the initial model pass generates a detection and false-positive rate against your labeled ground truth.

  3. Performance report. The 2-week deliverable is not a demo video. It is a performance report showing detection rate by defect class, false-positive rate on the known-good and supplier-variation parts, and a recommendations section on sample gaps — defect classes where more examples are needed before the model is production-ready.

  4. Gap identification. If a defect class is underrepresented in your sample (fewer than 10–12 examples of a given severity), the report will flag it and request supplementary samples before the production model is finalized. This is common for rare escape classes.

The performance report is what you take to your quality director or plant manager to make the go/no-go call on deployment. It is grounded in your own parts, your own defect history, your own borderline cases. That is why the sample box you build is the primary determinant of trial quality — not the system's underlying architecture.


What makes the trial inconclusive

Even with HyperQ's process, a trial can return inconclusive results. The most common causes:

Insufficient defect examples. If a critical defect class has fewer than 5 physical examples in the box, the detection rate for that class is statistically unreliable. The report will say so, but you may not have the production time to source more samples quickly.

Missing severity range. If every example of a given defect class is severe, the model learns to detect severe instances. The marginal cases — the ones that actually escape — are not covered. Include at least one marginal example per defect class.

No baseline for false positives. If the golden-reference layer is absent or too small, there is no reliable basis for measuring false-positive rate. A system that catches 99% of defects but flags 30% of good parts is not production-deployable. You need the golden reference to measure the false-positive floor.

Single part variant. If your line runs multiple product models under the same part family, a sample box with only one variant does not test the model's generalization capability. HyperQ AI Vision supports auto-switching between over 8,000 product models with no manual configuration change, but that capability needs representation in the trial to be validated.

Parts that are too clean. This is the most common issue. Vendors love clean golden parts. The 2-week evaluation is more useful — and ultimately results in a stronger production model — if the box contains the ugliest unit you have ever shipped. Not because the AI needs a challenge, but because the production system will encounter that unit eventually. Better to find the edge now.


Honest positioning: where a sample trial has limits

A sample-based trial has boundaries that any buyer should understand before committing to a deployment decision.

It does not test throughput. A 2-week lab evaluation runs at controlled imaging rates. Your production line may run at a rate that requires parallel imaging stations, specific conveyor integration, or triggering logic tied to your PLC. Throughput validation requires on-site setup — which is why the 2-day on-site setup phase exists before production go-live.

It does not test all environmental variation. Lab imaging conditions are controlled. Your production floor has variable lighting, vibration, and temperature fluctuation. The on-site calibration pass addresses this, but the sample trial will not reveal every environmental challenge.

It is a model initialization, not a finished model. The production model that goes live after 4 weeks is not identical to the trial model. It has been refined on additional imaging from the on-site setup phase. The trial proves the hypothesis; production calibration optimizes the execution.

These limits are not reasons to avoid a trial. They are reasons to treat the trial as step one of a two-step validation, and to ask the vendor specific questions about what the on-site setup phase adds to the trial results. For a full grounding in the underlying technology and what to expect at each validation stage, see AI visual inspection evaluation basics.


The log sheet: what to include

Include a simple printed log sheet in the box. It does not need to be elaborate. A single A4 page with the following fields is enough:

Part number Qty Layer Defect class (if applicable) Severity Notes
[Your P/N] 12 Golden reference Lot 2026-07, accepted 15/07
[Your P/N] 8 Known-bad Surface scratch Severe Customer escape, July 2026
[Your P/N] 5 Known-bad Surface scratch Marginal Internal rework pull
[Your P/N] 6 Borderline Surface scratch Ambiguous Inspector A: pass, B: fail
[Your P/N] 7 Supplier variation Supplier batch change, May 2026

Add a contact name, phone number, and preferred return format (PDF report, shared drive folder) so the evaluation team can reach you if a part class needs clarification during imaging.


Five things that decide trial quality

  1. Your escape log drives layer 2. Known-bad parts from real production failures are irreplaceable. Source them before you build the box.
  2. Your inspector disagreements drive layer 3. Borderline cases are the hardest and most important class. Do not omit them.
  3. Defect class diversity matters more than total part count. 40 well-classified parts outperforms 100 golden parts.
  4. The log sheet is the ground truth. Without it, the evaluation team cannot train a model; they can only run a generic demo.
  5. The ugliest unit in your facility belongs in the box. If you have shipped it, the system needs to see it. The trial exists to find the edge of the model's capability on your specific parts — not to confirm the system works on easy cases.

The sample box you build is the seed of your production inspection model. A well-composed 40–60 part box covering the five layers above can initialize a deployment-ready model in the 2-week evaluation window.


Ship us your 3 worst parts alongside 15 clean reference units — we'll run the full evaluation and return a defect-class performance report in 2 weeks, no contract until the results meet your spec.

Written by

Hypernology Team

September 2, 2026

Share

Continue Reading

Translate Insight
to Infrastructure.

Interested in deploying these solutions to your facility? Let's discuss the technical requirements.

Initiate Briefing