Skip to content
GMAO CLOUD
en

Continuous improvement in maintenance: the method

How to apply continuous improvement to maintenance: what to measure, how often to review it, how to decide on a change and how to check afterward whether it worked.

Updated on 6 min read

  • Continuous improvement
  • Indicators
  • Maintenance management

Continuous improvement in maintenance usually stays an intention. Not because nobody wants to improve, but because the day fills up with urgent tasks and reviewing what’s working is the first thing to get postponed. And by the time it finally happens, there’s no data to do it with.

This article is about the method: what to measure, how often to stop and look at it, how to decide on a change and how to check whether it worked. The part about what to change specifically is in the second part.

The cycle, applied to maintenance

It’s the usual one — plan, do, check, act — but it helps to translate it into concrete terms, because in the abstract it doesn’t apply to anything.

Plan means defining the checklist and the frequency for each equipment family. Do means work orders get closed with real data. Check means looking at the indicators. Act means changing the frequency, the checklist, the stock or the assignment.

The link that almost always breaks is the third one, and it breaks for a very specific reason: the time was never set aside.

No logging, no improvement

Worth saying up front, because it explains why many initiatives never get off the ground. If hours are jotted down from memory at the end of the day and materials are noted on Friday, the data exists but isn’t accurate, and deciding based on it is worse than not deciding at all.

What makes logging reliable is that it happens where the work happens: a timer inside the order itself in the app, materials consumed against the warehouse, a completed checklist and a signature collected. The app works without coverage precisely so that this can happen in a basement.

What to measure: four figures, not thirty

A dashboard with thirty indicators doesn’t get looked at. These four do, and all of them come from the reports:

Ratio of preventive to corrective hours. The one that best describes whether the operation is proactive or reactive. It takes months to move, and that’s exactly what makes it the most honest one: you can’t fake it in a quarter.

MTBF, mean time between failures per critical asset. Measures reliability: if it goes up, the equipment is failing less.

MTTR, mean time to repair. Measures response capacity. These are two different things and shouldn’t be mixed up: you can repair something very quickly that still fails too often.

Accumulated cost per asset. The one that lets you decide whether a machine gets repaired again or replaced.

And one control figure: the percentage of the annual plan executed, which tells you whether the plan is being followed or just drawn up on paper. Without it, the other four get misread.

Measure with values, not checkboxes

This is where the difference between being able to improve and not being able to lies.

A checkbox-based checklist says someone looked. One with a minimum and maximum value says what they saw, and logs any out-of-range reading as an anomaly on the spot. That turns an inspection round into a time series, and a time series lets you see a trend before it turns into a breakdown.

It’s also what lets you answer, after a failure, the question that teaches the most: was it giving warning signs?

The one-hour quarterly meeting

This is the mechanism, and it’s surprisingly cheap.

Once a quarter, with the five indicators in front of you, answer four questions:

  1. What percentage of the plan has really been executed?
  2. Which equipment concentrates the anomalies?
  3. Which frequencies need to go up, and which need to go down?
  4. Which assets have accumulated a cost that no longer justifies continuing to repair them?

That one hour produces decisions worth more than weeks of operational work. A silly trick that works: schedule it in the calendar with its own frequency, like any other task. What isn’t scheduled doesn’t happen.

Raising and lowering frequencies

The most common change and the one that pays off the most, and it works in both directions.

Equipment that fails between reviews needs more frequency. Equipment that never causes problems is probably over-maintained, and that also costs money: technician hours, materials and downtime that weren’t necessary.

Both situations exist side by side in almost every facility, and without a history you can’t tell them apart. That’s why continuous improvement in maintenance is almost never about “reviewing more”: it’s about reviewing where things fail and stopping reviewing where they don’t.

When wear depends on usage rather than time, there’s a better path than adjusting the calendar: attach a counter to the asset — hours, cycles, kilometers — with a limit and a warning percentage, so that the preventive order is generated automatically once the threshold is crossed.

Closing the loop on anomalies

The point where most improvement systems die.

A technician spots something, logs it, and nothing happens. By the third time they stop logging it, and rightly so. An anomaly has to end up as an incident with its priority and its owner, or as an explicit decision to do nothing. Both outcomes are valid; silence isn’t.

The same goes for improvement suggestions from people in the field: they’re the cheapest source of useful ideas, and the first one to dry up if nobody responds to them.

How to know if the improvement was real

By comparing the same series before and after, on the same equipment. If you touched the frequency of a family, look at the MTBF and cost of that family, not the plant’s overall figures: the overall figures have too much noise to attribute anything to a single change.

And give it time. A frequency change takes one or two cycles to show up. Measuring after three weeks and concluding it didn’t work is the most common mistake.

Where to start if you don’t have a history

The reasonable objection: all of the above assumes data, and if you’ve just gone live you don’t have any.

Start with the only thing that doesn’t depend on history: the percentage of plan executed. From the first month you can know how many of the preventive orders generated were closed on time, and that figure alone already forces decisions. If it’s below 70%, the problem isn’t the frequency: the plan is bigger than the team can execute, and that’s fixed by cutting scope, not by pushing harder.

By three months, the first aggregated anomalies show up. By a year, MTBF starts to make sense. And by two years, you can compare full periods, which is when continuous improvement stops being an exercise and starts saving money.

In the meantime, what matters is that logging is accurate from the start: the data from the first months is what later serves as the baseline, and it can’t be reconstructed.

A warning about how to use the data

It’s meant for sizing equipment, budgeting and deciding on investments. Not for watching anyone. A team that senses the system exists to monitor them stops feeding it accurately, and at that point all continuous improvement runs out of raw material.

If you want to see what indicators would come out of your own data, you can request a demo.

← All articles

Shall we look at it with your way of working?

Leave your details and we will get in touch to see whether we fit. No commitment, no lock-in period.

We reply within one working day.