RCM has a failure mode of its own, and it is not getting the answers wrong. It is producing a two-hundred-page analysis that nobody implements, three years after the maintenance programme was needed.
| Pitfall | What it looks like | The guard |
|---|---|---|
| Analysis paralysis | Eighteen months on the first subsystem, the fleet already flying | Scope by consequence: analyse what can hurt somebody or stop the fleet, first |
| Analysing everything at the same depth | The cabin light fittings given the same treatment as the brakes | Significant-item selection, applied honestly and recorded |
| Functions without performance standards | "Provide cooling" with no numbers | A function nobody can measure cannot have a functional failure |
| Functional failures skipped | Straight from function to failure mode | The middle step is what makes the mode list complete rather than plausible |
| Modes at the wrong level | "Pump fails" as a mode; or every solder joint as a mode | Deep enough to choose a task, no deeper |
| Effects written as diagnoses | "Sensor fault" instead of what the crew and the aircraft see | The effect is the evidence for the consequence, and the consequence decides |
| Hidden failures missed | Protective devices analysed as though the crew would notice | Ask the evident question of every mode, and expect protection to be hidden |
| Failure-finding intervals guessed | A round number, no arithmetic | Derive from the tolerable multiple-failure rate, the demand rate and the hidden rate |
| P-F intervals asserted | A number with no source, and no statement of how it was measured | Say where it came from, and be honest that it is the softest number in the method |
| Intervals set by the calendar | Everything at 500 hours because that is when the aircraft comes in | Package deliberately, and only ever shorten; a packaged interval is still the analysis's |
| Age-based tasks by default | Scheduled overhaul because it feels prudent | Age-related failure has to be demonstrated, not assumed |
| Run to failure treated as failure | Every mode given a task, so the programme looks thorough | It is a decision with a reason, and it is the right answer for most modes |
| Redesign avoided | A safety consequence with no effective task, and a task written anyway | Default actions exist for exactly this case; the answer is the design |
| The operating context omitted | One analysis used for two fleets flown differently | The context is an input, and it is what makes the answers transferable or not |
| No cost case | Tasks justified because they are good practice | An operational or economic consequence needs the task to cost less than the failure |
| The analysis frozen at entry into service | Intervals from the original study, ten years of data ignored | Age exploration and interval revision are part of the method, not an afterthought |
| Results that never became tasks | An RCM report, and a maintenance programme that predates it | The output is a task requirement with an interval, handed to task analysis |
| Compliance claimed, criteria unread | "RCM" on the cover of something that skipped consequences | The evaluation criteria exist precisely so that claim can be checked |
Four deserve more than a row.
Analysis paralysis is the reason streamlined variants exist. A full analysis of a complex system is measured in person-years, and the value is concentrated in a small fraction of the modes. Every practical programme is therefore some form of triage, and the honest ones say which triage they used rather than presenting the result as a complete analysis. The danger of triage is that it quietly drops the hidden failures, which are exactly the ones nobody misses until they are needed.
The P-F interval is where the method's rigour runs out. Everything downstream of it is arithmetic; the number itself is an engineering claim about how an item degrades and about how early a particular check can see it. It is usually estimated, rarely measured, and almost never revisited. Where a task's whole justification is a P-F interval, the report should say how the number was obtained and what would change it.
Run to failure is the most common correct answer, and a programme that never reaches it is not being thorough, it is being uncritical. Most failure modes in most systems are evident, cost only the repair, and have no task that costs less than the failure. Writing a task anyway consumes hours, generates its own failures through maintenance-induced damage, and buys nothing.
A task that cannot be delivered is not a task. When the arithmetic demands a nine-and-a-half-hour interval, or a check nobody can perform, or a P-F interval shorter than the inspection cycle, the analysis has produced a design requirement. Recording it as a maintenance task with a comfortable interval is not a compromise: it is a false statement about the risk that somebody will rely on.