Analyzing breakdowns so they don't repeat
What to do with corrective maintenance history: finding the repeat failure, deciding whether to repair or replace, and turning a breakdown into a preventive plan.
Updated on 6 min read
- Corrective maintenance
- Indicators
- Costs
- Continuous improvement
The first part covers how a breakdown is managed: from the alert to the close-out. This one is about what comes next and is almost never done: looking at the whole picture.
A single breakdown is an incident. Fifty breakdowns logged with their cause, time, and cost are a diagnosis, and from there come the decisions that actually reduce next year’s corrective maintenance.
First: having something to analyze
Worth saying up front, because it conditions everything else. If all that’s left of a breakdown is “it got fixed,” there’s no analysis possible.
The minimum that has to be recorded on every work order: the cause of the stoppage, the actual time measured with a timer, the material consumed, and, where relevant, the anomalies detected with their values.
That’s achieved by recording in the field from the app — which works without signal — not by reconstructing it from memory in the afternoon. A rounded-off figure isn’t useful for comparison.
The repeat failure
It’s the most profitable finding, and the easiest to spot.
The reports for anomalies and order history let you group by equipment and by cause. What you’re looking for are three kinds of patterns:
By unit. A specific machine that concentrates the breakdowns. Usually it’s installation, location, or use — not the model.
By model. The same failure across every unit of a model, spread across the site or the network. It stops being bad luck and becomes an argument with the manufacturer or for the next purchase.
By timing. Breakdowns that cluster in a time slot, a season, or after a certain type of intervention. This last one is uncomfortable and very useful: sometimes maintenance itself is what’s causing the corrective work.
Repair again or replace
The most expensive question, and the one that gets answered by gut feeling when there’s no data.
What decides it is the accumulated cost per asset — hours plus material, over years — set next to the replacement cost and the downtime it generates. A piece of equipment that costs three times as much as its twins has an explanation, and it’s usually in how it’s used or how it’s installed.
GMAO CLOUD has an automation for this: an asset can have its replacement cost and a warning percentage recorded, and the system notifies when accumulated repair spend exceeds it. It doesn’t generate an order — replacing something is a business decision, not a task — but it puts the figure in front of you right when it matters, instead of inside a report nobody opened.
Turning the breakdown into a plan
This is the step that closes the loop and actually reduces corrective work.
When a failure repeats and is predictable, it stops being corrective: it becomes preventive. There are three ways to do it, depending on what it depends on.
Add the point to the checklist template. If the failure is in a component that isn’t being checked, more visits won’t change anything: what needs to change is the checklist. Adding two points to an equipment family takes five minutes and rolls out to all of them through the cascade.
Raise the frequency. If the equipment breaks down between checks, the periodicity is set wrong.
Switch to a counter. If wear depends on use, the asset can carry a counter — hours, cycles, kilometers — with a limit and a warning percentage, so the preventive order generates automatically once the threshold is crossed.
The material that’s always missing
Analyzing corrective maintenance also tells you what needs to be in the warehouse.
The list of spare parts worth stocking isn’t guessed: it comes from the history. The reports on materials used show which references are actually consumed, and cross-referencing with orders flagged as pending material shows which ones are causing second visits.
Those are the ones that need minimum stock in warehouses and items. The rest don’t: a system that forces you to inventory everything gets abandoned.
Measuring whether it’s working
Four figures, compared on the same equipment over time.
MTBF per critical piece of equipment: if it rises, it fails less.
MTTR: if it falls, response is faster. These are different things.
Accumulated downtime, which translates everything into lost service.
Ratio of corrective to preventive hours, the most honest indicator because it takes months to move and can’t be dressed up.
A word of caution when interpreting these: look at the equipment family you’ve touched, not the site-wide total. The overall figure has too much noise to attribute anything to it. And give it one or two cycles: measuring after three weeks and concluding it didn’t work is the most common mistake.
The most common analysis mistake
Blaming the breakdown on the last intervention. If a piece of equipment was repaired in March and failed in April, the easy conclusion is that the repair was bad. Sometimes it is; many other times, the March repair was the first symptom of a degradation that had been building for months.
That’s why it matters that checklists record values, not checkboxes: with a series of readings you can see whether the equipment had been giving warning signs, and that changes the conclusion and the action.
What to do with what can’t be avoided
Not everything corrective converts into preventive, and it’s worth accepting that so effort isn’t wasted where it doesn’t pay off.
There are sudden failures — a brittle fracture, an electronic component — that have no prior degradation to measure. There’s no predictive maintenance possible there, by definition, and what’s needed is something else: redundancy if the equipment is critical, or spare-part stock so downtime is short.
And there are assets where repairing is cheaper than preventing: cheap, redundant, or ones whose failure has no consequence. For those, corrective is the right strategy and checking them is wasted hours.
Telling one from the other is what asset criticality is for, which in asset management is a field on the equipment. Without it, the discussion repeats case by case; with it, it’s decided once per family.
How often to do this analysis
An hour a quarter is enough, and it’s the first thing to disappear from the schedule once the day fills up with emergencies. Setting it up in the calendar with its own periodicity, like any other task, is a dumb trick that works.
There’s more on the method in continuous improvement in maintenance. If you want to see what would come out of your own history, you can request a demo.