Reliability Centred Maintenance · Chapter 4

Worked Example

The method applied end-to-end on a concrete system, with numbers.

The same fleet as the task analysis and repair level modules: 36 medium transport aircraft, 600 flight hours each per year, two environmental control system packs per aircraft, so 21,600 flight hours and 43,200 pack operating hours a year. This is the analysis that produced the scheduled tasks those two modules were already using. Every value is an illustrative teaching figure.

The operating context, first

Before any function is written: the aircraft operates from three bases, two of them hot and humid; dispatch with one pack inoperative is permitted for ten days; there is no cabin air quality instrumentation on the flight deck; and the fleet flies an average sector of two hours. Change any of those and some of the answers below change, which is why the context is recorded rather than assumed.

Functions and functional failures

FunctionPerformance standardFunctional failures
Deliver conditioned air to the cabin0.35 kg/s per pack, delivered between 5 and 30 °CNo air at all; flow below 0.28 kg/s; temperature outside the band
Limit duct temperatureTrip and isolate the pack above 88 °C at the ductFails to trip when the temperature is exceeded
Remove ozoneCabin ozone below the certified limit at cruise altitudeOzone above the limit
Contain bleed airNo hot leak into the equipment bayHot bleed leak into the bay

The second function is the one worth pausing on. It does nothing in normal operation, it has no cockpit indication, and the crew will discover it has failed only on the day the first function has already failed in a particular way. That makes it hidden, and hidden functions are where this method earns its reputation.

Modes, consequences and answers

Nine failure modes came out of the FMECA with their rates:

Modeλ per 10⁶ hArisings a yearEvident?ConsequenceAnswer
ACM bearing degradation552.4evidentoperationalOn-condition, 1,500 FH
Heat exchanger core fouling301.3evidentoperationalOn-condition, 3,000 FH
Pack control unit fault1205.2evidentoperationalRun to failure
Flow control valve fails closed954.1evidentoperationalRun to failure
Ozone converter catalyst exhausted401.7evidentoperationalRun to failure
Temperature sensor drift1506.5evidentoperationalRun to failure
Ram air door actuator jam1104.8evidentoperationalRun to failure
Overtemperature switch fails to trip251.1hiddensafetyFailure-finding, 500 pack hours
Duct coupling seal ageing180.8evidentsafetyScheduled discard, 6,000 FH

Two things in that table are the whole method. Five modes get no scheduled task, because none is both applicable and effective and because the consequence is operational rather than safety: that is a recorded decision with a reason, not an omission. And the two safety-consequence modes get tasks whose intervals are computed rather than chosen, which is the next two sections.

The on-condition intervals come from the P-F interval

The bearing gives 3,600 flight hours of warning. Take off the time needed to plan and do the removal, halve what is left, and the interval falls out; the number that is actually hard to defend is the 3,600.
The bearing gives 3,600 flight hours of warning. Take off the time needed to plan and do the removal, halve what is left, and the interval falls out; the number that is actually hard to defend is the 3,600.
ACM bearingHeat exchanger core
P-F interval3,600 FH7,000 FH
Time to plan and act250 FH400 FH
P-F interval less the time to act3,350 FH6,600 FH
Interval under the halving convention1,675 FH3,300 FH
Adopted1,500 FH3,000 FH
Opportunities inside the window22

That third row is not the net P-F interval of the foundations, which is what is left of the warning after the check rather than after the action: 3,600 − 1,500 = 2,100 FH on the bearing and 7,000 − 3,000 = 4,000 FH on the core. Both are far longer than the 250 and 400 hours the removal takes, which is the requirement the standard actually states. Two different quantities, and a report that calls them by the same name will eventually have one of them read as the other.

It is also the longest interval the standard's requirements allow: a check at 3,350 flight hours is still shorter than the 3,600 of warning and still leaves the 250 hours the removal needs. The fourth row is half of it, and halving is convention rather than requirement, as the foundations chapter says. What has to be defended if a programme ever wants the longer number is the requirements, not the habit.

The adopted intervals are shorter again, and deliberately so: an existing check package sits at 1,500 and 3,000 flight hours, and a task that rides on a visit somebody is already making costs a fraction of a task that generates its own. Packaging can only ever shorten an interval, never lengthen it, which is the one direction the arithmetic allows.

What the ACM task buys is worth stating precisely, because it is routinely oversold:

arisings, with or without the task: 2.4 a year

caught at the potential-failure stage: 2.0 a year · still fail in service: 0.4 a year

The task does not make the bearing last longer and does not reduce removals. It converts an unplanned in-service failure into a planned removal, which is worth having for the consequence and for the secondary damage avoided, and worth nothing at all if somebody expected the removal count to fall.

The failure-finding interval is arithmetic

The multiple-failure rate against the check interval. The tolerable rate fixes the maximum interval at 952 flight hours; the programme adopts 500 because a visit already exists there, and gains a factor of 1.9 in margin.
The multiple-failure rate against the check interval. The tolerable rate fixes the maximum interval at 952 flight hours; the programme adopts 500 because a visit already exists there, and gains a factor of 1.9 in margin.

The hidden function protects against duct overheat. The multiple failure is the overheat event happening while the protection has already failed, and the FHA classified that as hazardous, which on this aircraft is a tolerable rate of 1 × 10⁻⁷ per hour.

InputValue
Hidden failure rate of the switch, λh2.5 × 10⁻⁵ per hour
Rate of the event it protects against, λd8.4 × 10⁻⁶ per hour
Tolerable multiple-failure rate1.0 × 10⁻⁷ per hour

maximum tolerable unavailability U = 1.0 × 10⁻⁷ ÷ 8.4 × 10⁻⁶ = 1.19 per cent

average unavailability of a hidden item checked every T ≈ λh · T ⁄ 2, so T = 2U ⁄ λh = 952 flight hours

IntervalUnavailabilityMultiple-failure rateVerdict
500 FH0.63%5.3 × 10⁻⁸within target, 1.9× margin
952 FH1.19%1.0 × 10⁻⁷exactly at target
1,500 FH1.88%1.6 × 10⁻⁷over target

Now change one input. Had the multiple failure been classified catastrophic rather than hazardous, the tolerable rate would be 1 × 10⁻⁹ per hour, the maximum interval would be 9.5 flight hours, and no maintenance programme on earth delivers that. The analysis would then have produced a design change rather than a task, which is not a failure of the method: it is the method working, and it is the reason the default action for hidden safety consequences is what it is.

The scheduled discard, and what it can and cannot do

The duct coupling seal is an elastomer whose conditional probability of failure rises after about 7,000 flight hours of seal age. A life limit at 6,000 hours removes the age-related part of the mode:

Without a life limitWith discard at 6,000 FH
Hot leaks a year0.780.26
Planned removals a year07.2

The task cannot touch the random part of the mode, which is what the residual 0.26 is, and it buys the reduction with 7.2 planned removals a year that would not otherwise happen. A scheduled discard is a trade of failures for planned work, and it is worth making here because the consequence is a hot leak in an equipment bay rather than a pack shutdown.

The cost test, worked on one mode

The two safety modes got the two risk tests. The five decisions to do nothing came from the other two tests, the cost ones, and those are the tests most often waved through with "it is good practice". Here is one of them in full.

Temperature sensor drift. Evident, because the cabin temperature leaves the band and the pack fault indication comes up. Operational, because the pack is shut down and the aircraft carries on with the other one. The effectiveness test is therefore the third of the four: the task has to cost less than the operational consequence plus the repair, over comparable periods of time.

What the failure costs in a year
Arisings6.5
Unscheduled sensor change2.4 MMH each
Repair6.5 × 2.4 = 15.6 MMH
Operational consequenceNothing. Dispatch with one pack inoperative is permitted for ten days, so the change waits for the next convenient stop
What the task would cost in a year
Candidatecalibrate the sensor against a reference at each 1,500 FH visit
Occurrences43,200 ⁄ 1,500 = 28.8
Duration0.75 MMH. Unlike the switch check it needs the pack running and a reference source, so it cannot ride on a panel somebody already has open
Task28.8 × 0.75 = 21.6 MMH

The task does not remove the 15.6. A drifting sensor still has to be changed, and the check makes no difference to how many changes there are. What it buys is the difference between an unscheduled change at 2.4 MMH and the same change worked into a planned visit at 1.4 MMH, on the six arisings in ten it would catch before the band is broken:

3.9 catches a year × 1.0 MMH avoided each = 3.9 MMH

21.6 man-hours of task to avoid 3.9 man-hours of unplanned work, with no operational consequence to avoid at all. The task fails the effectiveness test by a factor of five and a half, and the mode is recorded as run to failure with that arithmetic beside it. That is the difference between a decision and an omission: whoever proposes the same calibration check in three years' time has to beat 21.6 against 3.9, not win an argument about good practice.

Then change the operating context, which is why it was written down first. Take away the second pack, or take away the ten-day dispatch allowance, and each of those 6.5 arisings costs a delay or a cancellation instead of nothing. The repair term does not move. The operational term goes from zero to something that dwarfs twenty-one man-hours, and the identical task passes the identical test. The other four run-to-failure modes were tested the same way and failed for the same reason: on this aircraft the redundancy and the dispatch allowance are already doing the work the task would have been bought to do.

The programme it produced

Five scheduled tasks and five deliberate decisions to do nothing. Two of the tasks are new work, and they are what goes to task analysis next.
Five scheduled tasks and five deliberate decisions to do nothing. Two of the tasks are new work, and they are what goes to task analysis next.
TaskTypeIntervalMMH a year
Pack filter servicingservicing500 pack hours103.7existing
ACM vibration and debris checkon-condition1,500 FH28.8existing
Heat exchanger core inspectionon-condition3,000 FH32.4existing
Overtemperature switch functional checkfailure-finding500 pack hours34.6new
Duct coupling seal replacementscheduled discard6,000 FH10.1new
209.6

44.7 MMH a year of new scheduled work, a 27 per cent increase on the 164.9 the programme already had

Not every row in that table is on the same clock, and the interval column has to say which. The two per-pack tasks are counted against pack hours, 43,200 ÷ 500 = 86.4 a year; the aircraft-level checks are counted against aircraft flight hours, 21,600 ÷ 1,500 = 14.4 and 21,600 ÷ 3,000 = 7.2. The interval itself is the same number on either clock, because a pack runs whenever its aircraft flies and so accrues one operating hour per flight hour. What changes is how many times the fleet does the task: read the filter servicing off aircraft hours and you get 43 a year instead of 86, and the workload comes out at half. The task analysis carries the same three clocks and the same counts.

The failure-finding check is the interesting line. At 500 pack hours it runs 86 times a year, which sounds expensive until you notice it is 0.4 man-hours because it rides on the filter servicing visit: the same panel is already open, the same technician is already there. A task written as a separate visit would have cost three times as much and bought exactly the same protection.

The same two modes, through the civil aviation logic

If this were a transport aircraft going through certification, the analysis above would be run to the sector procedure rather than to the generic criteria. The shape does not change; the vocabulary, the numbering and the approval route do.

Level 1 sorts each functional failure onto one of five numbered routes, 5 to 9, using the same evident-or-hidden question generic RCM asks first. Level 2 then works the candidate tasks for that route, and the answer travels through an approval chain that ends at the authority.
Level 1 sorts each functional failure onto one of five numbered routes, 5 to 9, using the same evident-or-hidden question generic RCM asks first. Level 2 then works the candidate tasks for that route, and the answer travels through an approval chain that ends at the authority.

Step 1, select the significant item. Working down the aircraft by system, the pack is a maintenance significant item: its failure is not always obvious to the crew, it can affect operating capability, and it is expensive. The analysis is done at the highest level that can be managed, rather than on every part inside it.

Step 2, list the functions, functional failures and causes. The same three columns as before: deliver conditioned air, limit duct temperature, remove ozone; the ways each fails; the causes from the FMECA.

Step 3, Level 1: put each functional failure on a route. One question first, and it is the same question generic RCM asks: is the failure evident to the operating crew in normal duties?

FailureEvident?ThenRoute
Reduced airflow, ACM bearing degradationYes: pack fault indication and cabin temperatureSafety? No, the other pack carries it. Operating capability? Yes, it restricts dispatch6, evident operational
Loss of duct overheat protectionNo: nothing shows until the day it is neededWould this plus one more failure hurt somebody? Yes8, hidden safety

Step 4, Level 2: pick the task for that route. The candidate task types are worked in order, and what the route changes is how hard the bar is:

Route 6, the bearingRoute 8, the overheat switch
Is a task required?Only if it pays for itselfYes: a hidden safety route must end in a task
Task selectedInspection or functional check for bearing conditionOperational check that the protection still trips
Interval1,500 FH, from the P-F interval500 pack hours, from the tolerable multiple-failure rate
If nothing workedAccept the failures, run to failureRedesign. Not an option among several: the answer

Those are the same two answers the generic analysis reached earlier in this chapter, arrived at through different paperwork. That is the point of running one engine under several rule sets: the engineering is portable and the compliance artefact is not.

Step 5, the approval chain. The working group for that system does the analysis, the industry steering committee agrees it, the authority's review board approves it, and the result becomes the board report. The manufacturer publishes it as the planning document, and each operator then builds its own approved programme from that. Nothing in this chain re-does the engineering; it decides whose signature the answer carries.

Alongside the systems analysis run three others that generic RCM has no equivalent of: structures, with its own damage categories, zonal, which looks at what shares a space rather than at what shares a function, and the wiring and lightning analyses added later. A pack analysis that stopped at the systems logic would have covered one quarter of the aircraft's scheduled programme.

What is reproduced above is the shape of the procedure and its route numbering. The exact wording of the questions, the task codes and the structural terminology are in the published specification, which is a paid document and is not reproduced here.

Where it goes next

OutputConsumer
Five task requirements, with intervals and the reason for eachMaintenance task analysis, which prices them
Two of them new, needing steps, tools and resourcesThe same, as new work
A design change request, had the classification been one class higherThe design authority
Five run-to-failure decisionsThe corrective task set, and the repair level analysis that places it
Every interval, with the assumption behind itThe in-service review, which is where the assumptions get tested

Want to see this on a live system model? Request a walkthrough.