Reliability Centred Maintenance · Chapter 3

The Method

How the analysis actually runs, step by step.

Ten steps. The first decides how much of the system gets the full treatment, and it is the step that decides whether the analysis finishes.

1. Write the operating context down

Roles, duty cycles, environment, redundancy, dispatch allowances, what a stoppage costs, and what the organisation treats as tolerable risk. This is an input to every subsequent answer, and it is the reason the same equipment gets different programmes in different fleets.

2. Select what to analyse, and at what depth

Full analysis of everything is how RCM programmes die. Select the items whose failure can hurt somebody, breach a regulation, stop the operation or cost more than the analysis, and record the rule you used. Everything else gets a lighter treatment, and the fact that it did is part of the record rather than a secret.

Two things must not be filtered out by the triage: protective and other hidden functions, and modes with safety or environmental consequences. Those are exactly the items where the value is concentrated and exactly the ones a cost-driven filter tends to drop.

3. Functions, with performance standards

Primary and secondary. Include containment, protection, support, compliance and comfort. Every one needs a number or a stated condition, because the next step is defined against it.

4. Functional failures

The ways the function stops meeting its standard: total loss, partial loss, and the failure to stay inside limits. Write them against the function rather than against the hardware.

5. Failure modes, from the FMECA

Take them from the FMECA rather than re-deriving them, and take the rates with them. Modes belong at the level where a maintenance decision differs: deep enough to choose a task, no deeper. Include modes caused by human error and by maintenance itself where they are credible, because those get tasks too, and the tasks are usually procedural.

6. Effects, written for the consequence question

What happens, in what order, in what time, and what evidence the crew has. The effect has to answer three things: whether the failure is evident, whether it hurts anybody, and what it costs. If the effect column cannot answer those, the consequence column is guesswork.

7. Consequences, on both axes

Evident or hidden first, then safety and environmental against operational and non-operational. Hidden failures are assessed on the multiple failure, not on themselves, which usually moves them a whole category.

8. Choose the task, and prove both halves

Applicability and effectiveness, task type by task type. Both have to be true, and the effectiveness test depends on the consequence category from step 7.
Applicability and effectiveness, task type by task type. Both have to be true, and the effectiveness test depends on the consequence category from step 7.

For each mode, in this order:

  1. Is a condition-based task applicable? Is there a potential failure, an identifiable P-F interval, and an interval shorter than the shortest likely one, with enough left to act?
  2. Is scheduled restoration or discard applicable? Is there a demonstrable age at which the conditional probability of failure rises, and does the task restore original resistance to failure?
  3. Is the failure hidden? Then a failure-finding task, at an interval derived from the tolerable multiple-failure rate.
  4. Is any of those effective, against the test the consequence category demands, assessed as if nothing were being done today?
  5. If not, take the default action: run to failure where only money is at stake, redesign where safety or the environment is.

9. Derive the intervals, and say where they came from

A failure-finding interval is computed from three numbers. Reporting the three, and not just the interval, is what lets somebody re-derive it when one of them changes.
A failure-finding interval is computed from three numbers. Reporting the three, and not just the interval, is what lets somebody re-derive it when one of them changes.
Task typeInterval comes from
Condition-basedThe P-F interval, less the time needed to act, conventionally halved
Scheduled restoration or discardThe age at which the conditional probability of failure rises
Failure-findingThe tolerable multiple-failure rate, the demand rate and the hidden failure rate
ServicingThe consumption or contamination rate of whatever is being replenished

Then package, and only ever downward. A task can be pulled in to join an existing visit; it cannot be pushed out past the interval its arithmetic allows. Record the computed maximum next to the adopted interval so the margin is visible.

10. Hand it on, and plan to revise it

The output is a set of maintenance requirements: task, interval, reason, and the inputs the interval depended on. They go to task analysis to be turned into work with resources attached.

Then plan the revision from the start. Age exploration is how an interval earns the right to be extended, and it is a sampling programme rather than a piece of analysis:

What is sampledHigh-time units, inspected, tested or stripped at increasing ages, with the condition found recorded against the age it was found at
How it movesIn bounded steps. Extend the interval by one increment, inspect the sample that reaches the new age, and let what those items look like decide whether there is a next step
What counts as evidenceThe condition of the sampled items. No failures at the current interval is what you would see whether or not the interval could be extended, so on its own it proves nothing
Who signs itWhoever approves the programme, because an extension changes an approved document rather than an engineering opinion

In-service data through FRACAS is how an assumed rate gets corrected, and a design change or a change in the operating context invalidates the answers that depended on it. An analysis with no revision plan is a snapshot of what was believed on the day it was signed.


Want to see this on a live system model? Request a walkthrough.