Industrial anomaly detection · Machine learning · Predictive maintenance

An anomaly is not a failure

Anomaly detection has become the magic phrase of predictive maintenance. ‘Our artificial intelligence finds anomalies’: everyone says it. It is true. It finds them. It finds a great many.

The problem is that almost everything in a factory is anomalous.

Change the product on the machine: anomaly. Restart the press after the weekend: anomaly. The operator takes a break and the machine stays on without producing: anomaly. Summer arrives: anomaly. A new compound enters: anomaly.

None of these is a failure. A system that reports them all does not find failures. It finds the normal life of a factory and calls it a problem.

Anomalous does not mean broken
Product changeNormal factory life
RestartNormal factory life
Operator breakNormal factory life
SummerNormal factory life
New compoundNormal factory life
Real degradationA seal leaks and keeps getting worse

Five unusual events can be normal. A failure can begin without a spectacular spike.

The anomaly detection trick

Most systems work like this: look at the data, learn what is ‘usual’, and report anything unusual. The more sophisticated the algorithm, the better it becomes at finding the unusual.

But unusual and broken are different things. A machine doing something new is unusual. A machine beginning to fail often is not.

You know the result: an anomaly score rising and falling without explanation, alarms nobody can justify, and a maintenance technician who stops looking after a month. When someone asks, ‘why is this machine red?’, the answer is a number.

A number is not a diagnosis. It is a bet.

Three mistakes almost everyone makes

01

They learn normal from dirty data

If the machine was already sick during training, the system learns the illness as normal. From then on, it will never see it again.

02

They chase the failure

Many systems continuously update their reference to ‘adapt’. A slowly deteriorating machine moves normal with it, day after day. The alarm switches off exactly when it was needed.

03

They judge without context

They compare every reading with one normal for the machine's entire life. Yet the same reading is perfect with one product and wrong with another.

When normal chases degradation
Adapting referenceAlarm disappears
Stable referenceAlarm stays active

If the reference follows the worsening signal, degradation gradually becomes normal.

How Atherya works

We built our engine around one rule: every alarm must be able to say why.

01

Every reading is judged against the right normal

Atherya knows what the machine is producing, which month it is, and whether it has just restarted. Each situation has a signature learned from that machine's history. When the product changes, the reference changes at the same moment. A new sample is never judged against the previous product's band.

02

A product change is not an anomaly

Every change has a settling period learned from that machine's real changes, not chosen on paper. Atherya observes but does not alarm during it. Only physical limits stay active, because physics does not pause.

03

What it does not know, it declares

If a never-seen product arrives, Atherya does not pretend to know it or judge it against other product bands, which would guarantee noise. It abstains and keeps only physical plausibility checks active.

04

A set parameter is not a sensor

A curing time chosen by the operator does not ‘drift’: it is changed. Atherya checks it as a perimeter of legitimate settings, never as a failure signal.

05

A stopped machine is not a broken machine

An activity detector separates a working press, an operator break, a machine left on unattended and a switched-off machine. Breaks do not create alarms.

06

The setpoint is the strongest available reference

When the management system records the set value, the gap between actual and setpoint works from the first reading, without waiting months for learning.

A real failure has a shape

An isolated spike is not a failure. A failure is recognised by how it behaves over time, and that is precisely what our engine reads.

  • Persistence. An anomaly lasting hours weighs more than a momentary spike.
  • Trend. Drift is recognised only when there is statistical evidence that the signal is truly moving in one direction, using robust methods that noise cannot fool.
  • Multiple signals together. When several sensors on the same machine leave normal at once, the engine measures it: the probability of coincidence falls as the number of involved signals rises.
  • Multiple horizons. The engine reads sudden spikes, patterns over the last hour and slow degradation over days, and keeps them separate.
  • Direction matters. Rising vibration is a problem. Vibration falling to zero usually means a disconnected sensor. Treating both alike creates phantom alarms.
The signature of a real failure
01Persistence
02Trend
03Multiple signals
04Multiple horizons
05Direction

Not one isolated red point: a coherent shape that withstands noise and unfolds over time.

Never chase degradation

There is one point on which we do not compromise.

Some machines have a chronic defect: they have always worked slightly outside nominal, steadily, without getting worse. Atherya recentres their bands on real behaviour and records the defect as a known, visible condition. From then on, an alarm means something again.

But if a machine is outside nominal and moving, we adapt nothing. Adapting there would switch off the alarm while the failure grows. The difference between a stable defect and ongoing degradation is exactly the difference between a system that learns and one that gets used to the problem.

The same applies to initial learning: normal is calculated with robust statistics and anomalous historical episodes are excluded, so a past failure does not widen the bands used to judge the next one.

No black boxes

Every Atherya engine produces a verdict with its reason. A final arbiter combines them and preserves the trace of who said what. When a technician asks, ‘why is this machine critical?’, the answer is not a number. It is a sentence: which signal, against which normal, by how much, since when, and which engines confirm it.

Atherya also counts the work nobody sees. Whenever a reading would have triggered a fixed-threshold alarm but sits inside the normal for that product in that month, the system records it. These are useless alarms that never reached anyone's phone.

From signal to explanation
01Signal
02Right normal
03Deviation
04Duration
05Confirmations
Pressure outside the product signature, rising for hours, confirmed by multiple signals.

The technician receives the reason, not merely a score.

Value is measured, not declared

Our industry circulates beautiful numbers. We chose another path: compare Atherya's warnings with real failures and interventions recorded in the customer's maintenance systems. How many real failures had a prior warning. How many warnings were followed by an intervention. How much advance notice the customer had.

The technician can say when an alarm was wrong, and that judgement returns to the system to improve the thresholds. Sensitivity is corrected with the people who know the machine, not against them.

The truth in one line

Finding anomalies is easy. A factory has them on every shift.

The hard part is knowing which ones matter: separating a new product from a leaking seal, a cold restart from failing cooling, a different machine from a machine getting worse.

That is why we do not sell anomaly detection. We say, with a reason, when a machine is leaving its true normal and when it is simply doing its job.

An anomaly is not a failure. A system that does not know the difference is not protecting your factory. It is only watching it.

The Atherya team

Do not search for everything unusual.

Find what is truly leaving your factory's normal.

Discover Atherya

Related guides