Reliability Growth Analysis · Chapter 6

Common Pitfalls

The mistakes seen in practice and the guards against them.

Growth programmes fail in a way that is almost invisible from the inside: the curve goes up, the reviews go well, and the fleet is disappointed three years later.

PitfallWhat it looks likeThe guard
Cumulative MTBF reported as the result278 h quoted at the end of a test worth 427Report instantaneous, with cumulative beside it if you must
Test without fixHours accumulated, corrective actions deferred to the next phaseGrowth comes from design changes; the curve will be flat and it should be
Failures scored awayRelevance argued after the fact, and the intensity fallsClassify against rules written before the test, as in FRACAS
The plan curve shown as dataA rising idealised curve at a design review, with no failures on itPlanning, tracking and projection are three different pictures
Growth potential never computedA target above what the design can reach, chased with more hoursCompute the ceiling from the A-mode share and the fix effectiveness
Fix effectiveness assumed to be oneEvery corrective action credited in fullReal effectiveness is well below one, and it is the ceiling's biggest lever
A-modes not declaredEvery mode implicitly assumed fixableSay which modes will not be fixed, and why, before the test
Growth rate taken as a targetA programme committing to an alpha of 0.5Alpha is an outcome of finding and fixing, not a knob
The likelihood bias left inβ̂ quoted straight from the estimatorIt is high by exactly n⁄(n−1), or n⁄(n−2); the correction is arithmetic, not judgement
Test type not recordedA growth report with no statement of how the test stoppedIt sets the correction and the bounds, and it cannot be recovered from the failure times
Model fitted to too few failuresTwo parameters from six failures, quoted to three digitsReport the interval; early estimates are almost uninformative
A passed fit test read as validationCramér-von Mises passes, so the model is declared soundAround twenty failures it misses a modest step change more often than not; look at the plot as well
Duane r² quoted as evidenceA high correlation on the log-log plot offered as proof of growthThe cumulative points are not independent, so the statistic does not mean what it means elsewhere
Test environment gentler than the worldLaboratory conditions, benign duty cycle, clean powerThe environment is part of the result and belongs in the report
Prototype not productionGrowth demonstrated on hand-built articlesThe manufacturing modes arrive later, and they are not in this data
The curve extrapolatedA projection carried far past the test to meet the requirement on paperExtrapolation past the data is a plan, not a measurement
Late failures reset nothingA significant redesign mid-test, with the model carried straight throughA step change in the design breaks the model's assumption
Field data never comparedThe growth result filed, and the fleet's actual MTBF never checked against itClose the loop; the comparison is the only validation the model gets

Five are worth expanding.

The likelihood estimate is biased and the bias is not a matter of opinion. β̂ comes out high by exactly n⁄(n−1) on a time-terminated test and n⁄(n−2) on a failure-terminated one, which on eighteen failures is 5.9 and 12.5 per cent. Because β is below one during growth, an inflated β means an understated growth rate: 0.349 rather than 0.385 on the worked example. The correction takes one multiplication, and once it is applied the likelihood answer and the Duane plot usually stop disagreeing, which removes an argument that has no content in it.

Reporting the cumulative figure is the oldest trick in this field, and it is often not a trick at all but a misunderstanding. The cumulative MTBF averages the whole test, including the failures that were designed out in the first week. It is lower than the instantaneous figure, so quoting it feels conservative and safe, and it systematically understates what the design is worth on the day. The two are related by a single factor, 1 ÷ β̂, which is 1.54 on the worked example, so there is no excuse for confusing them.

A test that finds failures and defers the fixes is not a growth test. It is a reliability measurement with an optimistic label. The whole model assumes each corrective action removes or reduces a mode; if the actions are deferred to a later build, the intensity does not fall, the fitted β sits near one, and the curve is honest even though the programme is not.

The growth potential is the number that should be computed first and is usually computed last. It answers whether the target is reachable by testing at all. If the answer is no, more test time is money spent proving something the arithmetic already knew, and the real answer is a design change. That conversation is far cheaper before the test than after it.

The test is not the fleet. On the worked example the fielded system runs at nearly four times the end-of-test figure; on other programmes the fleet is worse by a similar margin. Both are normal, because the test environment, the build standard, the duty cycle and the failure-counting rules all differ. A growth result is evidence about a test article in a test environment, and treating it as a prediction of field reliability is the assumption that makes the whole activity look dishonest in hindsight.


Want to see this on a live system model? Request a walkthrough.