RAMSynapse
Log inSign up

Maintainability Prediction · Chapter 4

Worked Example

The method applied end-to-end on a concrete system, with numbers.

Consider the traction converter cabinet of a rail vehicle: the power electronics that take the line supply and drive the motors, maintained in a depot road with the vehicle out of service. Seven replaceable items, 123 failures per 10⁶ hours between them. Every value is an illustrative teaching figure; the arithmetic is MIL-HDBK-472's.

The maintenance concept, settled first: repair by replacement at item level, by a depot technician with the vehicle stationary and the cabinet accessible from one side, complete when the converter has passed its power-on check.

The task times

Each repair broken into its elemental activities, in minutes:

ItemλlocisodisasinterreasaligncheckTotal
A · Cooling fan unit40436860532
B · Gate driver card224141061041260
C · Control card1849858181567
D · Power module stack1548342632618128
E · Line contactor1247129125857
F · DC-link capacitor bank94622142001076
G · Current sensor742695971171

Three rows are worth reading before any arithmetic. The fan is the most frequent failure and the quickest repair, because it was designed to be swapped. The power module is a two-hour job dominated by getting at it and putting it back. And the current sensor has a 26-minute isolation time, four times anyone else's, which is not a mechanical property at all: it is the ambiguity group the testability analysis would have found.

The mean, weighted properly

Mct = Σ λᵢ tᵢ ÷ Σ λᵢ = 7,591 ÷ 123 = 61.7 minutes = 1.03 hours

The unweighted average of those seven totals is 70.1 minutes, which is the number a spreadsheet gives and no technician experiences. The difference is the fan: the most common repair is also the shortest, and the weighting is what notices.

The distribution, and the number the contract wants

The lognormal fitted to the seven times, weighted by failure rate. The median is 55.7 minutes, the mean 61.7 and the 95th percentile 116.4, so a slot booked at the mean is overrun by more than a third of the work.
The lognormal fitted to the seven times, weighted by failure rate. The median is 55.7 minutes, the mean 61.7 and the 95th percentile 116.4, so a slot booked at the mean is overrun by more than a third of the work.

Fitting on the logarithms, weighted the same way:

μ = 4.0203, σ = 0.4481

QuantityValue
Median, e^μ55.7 min
Mean, e^(μ + σ²/2)61.6 min
Mmax at 90 per cent99.0 min
Mmax at 95 per cent116.4 min

The fitted mean of 61.6 against the weighted arithmetic mean of 61.7 is the internal check that the lognormal is describing this mix honestly. And the 95th percentile is nearly twice the mean, which is the number a depot slot has to be sized on.

Where the time actually goes

The power module's 128 minutes by activity, and the whole cabinet's weighted average underneath. Disassembly and reassembly together are 39.4 per cent of the technician's time.
The power module's 128 minutes by activity, and the whole cabinet's weighted average underneath. Disassembly and reassembly together are 39.4 per cent of the technician's time.

Two sorts of the same table, and they say different things.

By item, weighted by λ·t, the ranking is not the failure-rate ranking:

ItemShare of the technician's yearFailure-rate rank
D · power module stack25.3%4 of 7
B · gate driver card17.4%2
A · cooling fan unit16.9%1
C · control card15.9%3
E · line contactor9.0%5
F · DC-link capacitors9.0%6
G · current sensor6.5%7

The power module fails a third as often as the fan and consumes half again as much of the year.

By activity, summed across the cabinet and weighted:

ActivityWeighted meanShare
Disassembly12.3 min20.0%
Reassembly12.0 min19.4%
Checkout10.3 min16.7%
Interchange9.8 min15.8%
Isolation8.4 min13.6%
Alignment5.0 min8.0%
Localisation4.0 min6.5%

Access is 39.4 per cent of the time and diagnosis is 20.1. That is the finding, and it belongs to whoever designs the enclosure.

What a perfect diagnostic would be worth

Capping every isolation time at five minutes takes Mct from 61.7 to 57.7 minutes and M_max(95) from 116.4 to 106.1. Worth having, and not the largest thing on the table.
Capping every isolation time at five minutes takes Mct from 61.7 to 57.7 minutes and M_max(95) from 116.4 to 106.1. Worth having, and not the largest thing on the table.

Suppose the testability work succeeds completely: every fault isolated to one item in five minutes, including the current sensor's 26.

Mct: 61.7 → 57.7 minutes, a 6.5 per cent cut · Mmax(95): 116.4 → 106.1 minutes

Worth having, and smaller than most people expect. Diagnosis is a fifth of the technician's time on this cabinet, so perfecting it cannot buy more than a fifth, and it buys a third of that because most isolation times were already short. Halving the power module's disassembly and reassembly, 66 minutes of its 128, would be worth considerably more, and that is a decision about fasteners and clearances taken while the cabinet is still a drawing.

What it means for availability

At 123 per 10⁶ hours the cabinet's MTBF is 8,130 hours, so at Mct = 1.03 h:

A = 8,130 ÷ (8,130 + 1.03) = 0.999873, about 66 minutes of downtime a year

and with the isolation fix, 62 minutes. Four minutes a year, which is the honest scale of the availability argument and a reminder of what this analysis is really for: it is a design tool for the maintenance burden, not an availability lever. The availability case for this vehicle is dominated by what happens before the technician arrives, not by what happens afterwards.


Want to see this on a live system model? Request a walkthrough.