Reliability Centred Maintenance · Chapter 6

Common Pitfalls

The mistakes seen in practice and the guards against them.

RCM has a failure mode of its own, and it is not getting the answers wrong. It is producing a two-hundred-page analysis that nobody implements, three years after the maintenance programme was needed.

PitfallWhat it looks likeThe guard
Analysis paralysisEighteen months on the first subsystem, the fleet already flyingScope by consequence: analyse what can hurt somebody or stop the fleet, first
Analysing everything at the same depthThe cabin light fittings given the same treatment as the brakesSignificant-item selection, applied honestly and recorded
Functions without performance standards"Provide cooling" with no numbersA function nobody can measure cannot have a functional failure
Functional failures skippedStraight from function to failure modeThe middle step is what makes the mode list complete rather than plausible
Modes at the wrong level"Pump fails" as a mode; or every solder joint as a modeDeep enough to choose a task, no deeper
Effects written as diagnoses"Sensor fault" instead of what the crew and the aircraft seeThe effect is the evidence for the consequence, and the consequence decides
Hidden failures missedProtective devices analysed as though the crew would noticeAsk the evident question of every mode, and expect protection to be hidden
Failure-finding intervals guessedA round number, no arithmeticDerive from the tolerable multiple-failure rate, the demand rate and the hidden rate
P-F intervals assertedA number with no source, and no statement of how it was measuredSay where it came from, and be honest that it is the softest number in the method
Intervals set by the calendarEverything at 500 hours because that is when the aircraft comes inPackage deliberately, and only ever shorten; a packaged interval is still the analysis's
Age-based tasks by defaultScheduled overhaul because it feels prudentAge-related failure has to be demonstrated, not assumed
Run to failure treated as failureEvery mode given a task, so the programme looks thoroughIt is a decision with a reason, and it is the right answer for most modes
Redesign avoidedA safety consequence with no effective task, and a task written anywayDefault actions exist for exactly this case; the answer is the design
The operating context omittedOne analysis used for two fleets flown differentlyThe context is an input, and it is what makes the answers transferable or not
No cost caseTasks justified because they are good practiceAn operational or economic consequence needs the task to cost less than the failure
The analysis frozen at entry into serviceIntervals from the original study, ten years of data ignoredAge exploration and interval revision are part of the method, not an afterthought
Results that never became tasksAn RCM report, and a maintenance programme that predates itThe output is a task requirement with an interval, handed to task analysis
Compliance claimed, criteria unread"RCM" on the cover of something that skipped consequencesThe evaluation criteria exist precisely so that claim can be checked

Four deserve more than a row.

Analysis paralysis is the reason streamlined variants exist. A full analysis of a complex system is measured in person-years, and the value is concentrated in a small fraction of the modes. Every practical programme is therefore some form of triage, and the honest ones say which triage they used rather than presenting the result as a complete analysis. The danger of triage is that it quietly drops the hidden failures, which are exactly the ones nobody misses until they are needed.

The P-F interval is where the method's rigour runs out. Everything downstream of it is arithmetic; the number itself is an engineering claim about how an item degrades and about how early a particular check can see it. It is usually estimated, rarely measured, and almost never revisited. Where a task's whole justification is a P-F interval, the report should say how the number was obtained and what would change it.

Run to failure is the most common correct answer, and a programme that never reaches it is not being thorough, it is being uncritical. Most failure modes in most systems are evident, cost only the repair, and have no task that costs less than the failure. Writing a task anyway consumes hours, generates its own failures through maintenance-induced damage, and buys nothing.

A task that cannot be delivered is not a task. When the arithmetic demands a nine-and-a-half-hour interval, or a check nobody can perform, or a P-F interval shorter than the inspection cycle, the analysis has produced a design requirement. Recording it as a maintenance task with a comfortable interval is not a compromise: it is a false statement about the risk that somebody will rely on.


Want to see this on a live system model? Request a walkthrough.