TB-2614 · REV B · Technical newsheet
Machine Vision & InspectionDevice profile
One-Class Anomaly Detection Screens New Lines Before Failure History Exists
One-class anomaly detection trains on confirmed-good time-series traces to screen production lines at launch, before SPC limits or supervised ML classifiers have failure history to learn from.
By Amara Osei5 min read927 words
Features
- One-class models train only on confirmed-good units, requiring zero failure examples to begin screening at new product introduction.
- The method flags units whose time-series trace shapes depart from the learned baseline even when all summary metrics remain within specification.
- Baselines should span several hundred confirmed-good units to capture shift-to-shift and lot-to-lot variation; single-run baselines over-flag normal production.
- Dispositioned flags accumulate into labeled failure history, supplying the training data supervised classifiers need later in the product lifecycle.

At new product introduction, defect screening is weakest exactly where defect rates are highest. Statistical process control needs a stable process and enough production volume to estimate variation before control limits mean anything. A supervised machine learning classifier needs labeled failure examples before it can recognize one. At launch, neither exists. One-class anomaly detection closes that gap: trained exclusively on units that passed test and were confirmed healthy, it learns the shape of normal production behavior and scores each new unit by how far that unit departs from the learned baseline.
The method matters because end-of-line test stands already record more information than their limits interrogate. On a heavy-duty diesel fuel injector line described by the author, units passed every established bench criterion yet later exhibited field issues. The recorded results sat within specification. The emerging failure signature simply had no representation in the existing screening limits — the test performed as designed, but it had never been designed to recognize that failure.
A third question
Conventional screening asks whether a unit violates a set limit. Supervised machine learning asks whether a unit resembles a defect seen before. Both require knowing what to hunt. One-class detection asks a different question: does this unit resemble the good ones? The model learns typical relationships between signals, how they move together, and the natural spread of the process. The author likens the result to an experienced test technician who has listened to ten thousand units run — the technician often cannot name the defect but can tell that something is off, which is enough to pull the unit for review.
Feed it traces, not summary metrics
Test stands typically produce two outputs: summary metrics compared against limits, and the underlying time-series traces those metrics came from. Exporting the tabulated summary metrics into the model is tempting. The article argues against it. Those metrics were designed to answer specific questions, and the entire purpose of the exercise is to catch what those questions miss. A unit can produce a perfectly ordinary summary value from a trace that looks nothing like a healthy one; the flagged unit in the author's example passed every limit, with its departure visible in the shape of the signal rather than its endpoint. Coverage planning should mirror a control plan: which test points exercise which functions, and which failure modes have no signal represented at all. A model can only flag departures in what it can see.
The baseline matters as much as the input. Plan on enough confirmed-good units — in many manufacturing applications, several hundred rather than a single well-behaved run — to capture shift-to-shift and lot-to-lot variation. Train the model on one quiet production day and it will flag Tuesday.
Threshold setting is an economics problem
The model outputs a continuous score, not a verdict. Where to draw the flag line is not a tuning parameter; it is the same trade-off an engineer already makes when setting an inspection level. Tighten the threshold and more potential escapes are caught at the cost of pulling more good units for review; loosen it and the review queue empties while catches drop. In many applications the cost of a field return substantially exceeds the cost of a pre-shipment teardown or review, an imbalance that can justify a tighter threshold and a higher initial false-alarm rate. The article offers a first-pass shortcut: set the threshold so the flag rate lands near the scrap and rework tolerance the process already absorbs, then tighten as review results accumulate.
Reaction plan before switch-on
The biggest deployment failure mode has nothing to do with the algorithm: a flag arrives on the floor and nobody knows what to do with it. A flagged unit is not a reject. It may be a real defect, an untrained legitimate design variant, or a fixture problem. The control plan should specify in advance who reviews a flagged unit, against what criteria, within what time, and where disposition gets recorded. That record-keeping compounds: every dispositioned flag becomes a labeled example, so after a few months of production the plant holds the failure history it lacked on day one — exactly the training data a supervised classifier requires. The anomaly detector is both a screen and a data collection mechanism that pays for the next tool.
Distinguish drift from change
Flag rates drift for two opposite reasons. The process may have genuinely moved and the model is reporting something true. Or the process has settled into a new, acceptable normal and the model is simply out of date — tooling changes, supplier changes, and test stand maintenance all produce the second kind. A sustained shift in flag rate should trigger engineering review, not automatic retraining; retraining on unexamined current production teaches the model that whatever is happening now is normal, including a drift that should have been caught. The quieter failure deserves watching too: if flagged units are reviewed and released without documented reason, the screen has been bypassed in practice even while it keeps running.
The method displaces nothing already in place. SPC still governs the process once enough history exists, and limits still catch what limits catch well. One-class detection covers the launch interval — and hands back a defect record that makes every downstream tool easier to build. The open question for any quality team standing up a new line: can you justify the review labor of a tighter threshold against the cost of a field return you have not yet seen?
via bnpmedia.com (Original)
Filed under
- anomaly-detection
- machine-learning
- quality-control
- end-of-line-testing
- manufacturing
More from Amara Osei
Show full bio
Senior reporter covering industry trends and analytics at Testbench Report.
22 articles
Application notes
- Synthetic Data Rewires Machine Vision for Quality Inspection
- Siemens and P&G Scale AI Quality Inspection Globally
- AI Trends Reshape Industrial Inspection and Robotics Roadmaps
- India fields first indigenous condition-monitoring system for marine diesels
- Machine Vision Moves From Defect Detection to Autonomous Control