Skip to main content
Research
13 min read

Image classification, object detection, and segmentation: what each one sees on a production line

This post explains the differences between image classification, object detection, and segmentation in manufacturing inspection. It shows how each computer vision task answers different quality questions and why choosing the wrong one leads to systematic escapes or false rejects even with good hardware and models.

Image classification, object detection, and segmentation: what each one sees on a production line

Image classification, object detection, and segmentation are three distinct computer vision task types. Classification assigns a single pass/fail label to an entire image. Object detection returns bounding boxes that locate and name specific regions within an image. Segmentation traces a pixel-level boundary around each identified region. On a production line, the choice of task type determines which quality questions the system can and cannot answer.

99% defect detection accuracy is what the right architecture delivers — and the task type is the variable that determines whether the architecture matches the question. Vendors who describe their system as "AI vision" without specifying which of the three tasks it runs are omitting the part of the specification that determines whether your line question gets answered. That gap does not show up in a sales demonstration. It shows up at your factory acceptance test, when the system produces escapes or false rejections that do not match your specification.

Most buyers encounter classification, detection, and segmentation as interchangeable labels for the same capability. They are not. Each task type has a defined input, a defined output, and a bounded set of quality questions it can resolve. Running the wrong task type on the right hardware — good cameras, edge inference, trained model — produces a system that operates at full production speed and produces systematically wrong answers.

Classification: pass or fail on the whole image

Classification takes an entire image and returns a single class label. In a production inspection context, that label is typically pass or fail, or a defect category: "marking complete," "marking missing," "surface acceptable."

What classification does not return is location. The model assigns a label to the image as a whole. It cannot tell you where on the part the defect appears, how large it is, or whether it falls inside or outside a tolerance zone. That is not a deficiency — it is the correct behavior for applications where location does not affect the disposition decision.

If your quality standard says any surface mark anywhere on the part rejects it, classification handles the decision correctly. If your standard says a surface mark in the thread zone rejects the part but the same mark on the end face is acceptable, classification cannot carry the decision. The model has no access to location information.

Use classification when:

  • The inspection question is binary: "Is this feature present or absent?"
  • Any occurrence of the defect anywhere on the part produces the same disposition
  • Marking verification: "Is the lot code complete and legible?"
  • Lot-level go/no-go screening where part geometry is irrelevant to the decision

Where classification breaks down: Position-sensitive quality standards, multi-defect routing where different defect types have different dispositions, and any specification expressed as an area or perimeter threshold.

Worked example — OCR marking verification on a laser-engraved fastener: A laser engraver marks each fastener with a lot code before it enters the packaging line. The line question is: "Is the marking present and complete?" Classification is the correct task. The model trains on complete markings (pass class) and missing or partial markings (fail class). At 270 items per hour, the system returns a binary label per part. No localization is needed — the marking occupies a fixed zone on the fastener face and the question is all-or-nothing. Adding detection or segmentation to this application increases inference cost and label preparation time without improving the quality decision by a single defect.

Detection: where is it, and what type?

Object detection returns bounding boxes with class labels and confidence scores. For inspection, the output is a structured record per detected region: defect type, bounding box coordinates, confidence score. The model does not label the whole image; it labels specific regions within it.

Detection is the right task when disposition depends on where the defect is located, when multiple defect types can appear simultaneously on the same part, or when downstream handling depends on defect type — a scratch routed to the rework station, a void routed to scrap.

Use detection when:

  • Location determines the disposition: a defect in zone A fails; the same defect in zone B is acceptable
  • Multiple defect classes appear on the same surface and carry different consequences
  • Part routing depends on defect type or defect position
  • You need a defect catalog for process control: "porosity cluster, lower-left quadrant"

Where detection breaks down: Bounding boxes are rectangles. If your quality standard requires the exact area or perimeter of a defect — whether a chip is within a 0.25 mm² threshold, for instance — a bounding box cannot give you that measurement. You need segmentation.

Worked example — scratch inspection on a reflective automotive fastener body: The quality specification accepts surface marks in the mid-shaft zone but rejects scratches in the thread zone and at the mating face. The line question has a location component: the same surface mark produces two different dispositions depending on where it falls. Detection returns: "surface mark, mid-shaft zone, confidence 0.91" or "scratch, thread zone, confidence 0.87." Disposition logic runs against the bounding box coordinates. Classification is the wrong task here — it can confirm that a scratch exists but cannot answer where. The escape pattern for a classification model on this application is predictable: scratches in the thread zone sometimes pass (the model averaged them against mid-shaft marks in training), and mid-shaft marks sometimes fail (the model is over-sensitive from thread-zone examples).

Segmentation: what is the exact shape and area?

Segmentation assigns a class label to every pixel in the image. The output is a mask: each pixel is either background or a specific defect class. This gives the exact shape, area, and perimeter of each defect in calibrated units.

Segmentation is the most computationally intensive of the three tasks. It is the correct choice when the accept/reject threshold is expressed in area or perimeter, when defect boundaries need to map against process zone coordinates, or when the downstream action requires a precise target region — a repair system that needs to know exactly where to apply material, not a bounding box estimate.

Use segmentation when:

  • Area threshold compliance: "Reject if total defect area in the sealing zone exceeds 0.5 mm²"
  • Leak-path identification on gasket or sealing surfaces where the path geometry determines the risk
  • Repair-zone targeting where downstream automation needs exact boundary coordinates
  • Regulatory documentation requires precise defect area measurements per part

Where segmentation is not needed: Binary pass/fail applications, or applications where defect type and location — not area — determine the disposition. Running segmentation on a classification question adds inference latency without improving the quality decision.

Worked example — edge chip measurement on a display panel: The quality specification accepts edge chips up to 0.3 mm² in the non-active border zone; chips above 0.3 mm², or any chip in the active display zone, reject the panel. Segmentation traces the pixel boundary of each chip, calculates the area in calibrated mm², and maps each chip against the zone boundary. Detection would return "chip present, upper-left quadrant" — the bounding box area is not the chip area, and the bounding box cannot give you a reliable sub-mm² measurement. Classification would return "panel fails" — with no chip size or zone information. Only segmentation answers the actual quality question the specification asks.

Decision table: matching the task to the line question

Task Line question it answers Model output Use when
Classification "Does this part pass or fail?" Single label per image Presence check, binary surface pass/fail, marking verification, lot-level screening
Object detection "Where is it and what type?" Bounding box + class label + confidence per region Position-sensitive disposition, multi-defect cataloging, part routing by defect type
Segmentation "What is its exact shape and area?" Per-pixel class mask with calibrated area/perimeter Area threshold compliance, leak-path identification, repair-zone targeting

What the wrong task type looks like in production data

A classification model running on a position-sensitive application produces systematic patterns in the defect data. Parts with surface marks in tolerable zones fail inspection — false rejection rates climb. Defects in critical zones with low contrast may produce inconsistent results — some escape, some fail, depending on which training examples had similar visual features. The escape rate does not track defect severity, and the false rejection rate does not track defect frequency.

The standard misdiagnosis is a training data problem. More labeled images are added to the training set. The model is retrained. Performance improves slightly and then plateaus below the manual inspection baseline. Iteration cycles repeat. The root cause — task-type mismatch — is not visible in training metrics. Recognizing a task-type mismatch before multiple retraining cycles have been completed saves substantial time and avoids the false conclusion that the defect class is not learnable.

The signal to look for: if your defect escape pattern varies by location on the part (the same defect class escapes in some positions but fails in others), your model is running classification on a detection application. If your false rejection rate is high on parts that pass manual inspection, and those parts have surface features that look similar to real defects in a different zone, your model cannot separate location from appearance — again, a detection application running as classification.

Cross-industry architecture proof

At a Tier-1 automotive parts supplier running 11,520 units per day across 8,000+ product variants, HyperQ AI Vision deploys all three task types by inspection recipe. Marking checks run classification. Mating-face inspection runs detection for scratch position. Sealing-surface checks run segmentation for area measurement against the specification threshold. The same deployed system handles plastic and metal components on the same line at 99% detection accuracy, with changeover time under 2 seconds between product variants. This is cross-industry proof — the multi-task architecture that handles mixed materials and 8,000+ variant recipes in an automotive context applies to any production environment where the quality standard requires more than a binary answer.

The architectural requirement is the same regardless of the part type: inference fast enough to keep pace with line speed, task selection matched to the quality question, and training data requirements low enough to be practical across a large SKU library. The same patented training approach that reaches production-ready performance with 1,000 images — versus the 10,000 images that conventional systems require — applies across all three task types.

For background on how rule-based and learning-based inspection approaches differ in handling mixed defect classes, see the rules-first vs. learning-first inspection guide. Full specifications for deploying HyperQ AI Vision across multiple task types in a production environment are on the HyperQ AI Vision solution page.

Honest positioning

Classification is underused for applications where it is genuinely the correct task — pure presence checks, marking verification, binary go/no-go on stable surfaces — and overused for position-sensitive applications where it systematically cannot carry the decision. Vendors who default to classification for everything are choosing the cheapest training pipeline, not the architecture that matches your specification.

Detection handles the majority of manufacturing inspection cases in practice. Its limitation is that bounding boxes are rectangles and cannot report area. For the subset of applications where the specification is expressed in area units, detection produces an incorrect output format regardless of model quality.

Segmentation is the most capable task type and the most expensive to run per inference cycle. It is the correct answer for area-threshold and leak-path applications and an unnecessary overhead for everything else.

Rule-based vision systems handle geometric measurement tasks with high precision for single-SKU applications with stable geometry and predictable defect positions. They are the right choice for those applications. They cannot handle the position ambiguity or atypical defect variety that trained AI models address in high-mix production.

Frequently asked questions

Can one inspection system run all three task types?

Yes. A production inspection platform assigns the task type per product recipe, not as a global system setting. A marking check on one SKU runs classification. The next SKU's sealing surface check runs segmentation. Changeover loads the correct task type automatically from the inspection recipe. The practical requirement is that the platform supports all three task type architectures within the same inference engine — not all systems do.

Which task type requires the most labeled training data?

Segmentation requires the most label preparation time because each training image needs per-pixel annotations. Bounding box annotation for detection is faster. Classification labels are the fastest to prepare — a single pass/fail tag per image. A patented low-data training approach reaches production-ready performance for all three task types with 1,000 labeled images, compared with 10,000 that conventional approaches require — a 90% reduction in training data. That difference is meaningful when a product library has 8,000+ SKUs.

How do I know which task type my vendor's system runs?

Ask for the inference output format. If the output is a single pass/fail signal per image, the system runs classification. If it returns bounding box coordinates with class labels, it runs detection. If it returns a pixel mask or a calibrated area measurement, it runs segmentation. A vendor who cannot answer this question clearly is running classification and hoping your quality standard does not require location data.

Does segmentation always perform better than detection?

No. Accuracy is task-specific, not task-type-specific. On a binary application where location is irrelevant, classification achieves the same quality outcome as segmentation at a fraction of the inference cost. "Better" means the output format matches the line question — not that the algorithm uses more sophisticated operations.

What are the warning signs that the wrong task type is running?

High false rejection rates on parts that pass manual inspection. Escape rates that do not correlate with defect severity. Inconsistent results on the same defect class appearing in different positions on the part. These patterns point to a task-type mismatch more reliably than a training data gap. Adding more labeled images to a mismatched architecture does not fix the specification problem.

Can the task type change automatically when the product changes on a high-mix line?

Yes — with the right platform. A correctly configured system stores the task type assignment in the product recipe and switches automatically at changeover, triggered by the PLC changeover signal. Before accepting any high-mix vision proposal, confirm that the vendor can specify task type per SKU rather than per system. A global setting is not a high-mix-capable architecture.


Send a sample batch and a copy of your quality specification to Hypernology. Within two weeks we will run a task-type evaluation against your actual defect library — classification, detection, and segmentation — and confirm which architecture your line question requires. No contract until the specification match is confirmed.

Written by

Hypernology Team

September 28, 2026

Share

Continue Reading

Translate Insight
to Infrastructure.

Interested in deploying these solutions to your facility? Let's discuss the technical requirements.

Initiate Briefing