Growth programmes fail in a way that is almost invisible from the inside: the curve goes up, the reviews go well, and the fleet is disappointed three years later.
| Pitfall | What it looks like | The guard |
|---|---|---|
| Cumulative MTBF reported as the result | 278 h quoted at the end of a test worth 427 | Report instantaneous, with cumulative beside it if you must |
| Test without fix | Hours accumulated, corrective actions deferred to the next phase | Growth comes from design changes; the curve will be flat and it should be |
| Failures scored away | Relevance argued after the fact, and the intensity falls | Classify against rules written before the test, as in FRACAS |
| The plan curve shown as data | A rising idealised curve at a design review, with no failures on it | Planning, tracking and projection are three different pictures |
| Growth potential never computed | A target above what the design can reach, chased with more hours | Compute the ceiling from the A-mode share and the fix effectiveness |
| Fix effectiveness assumed to be one | Every corrective action credited in full | Real effectiveness is well below one, and it is the ceiling's biggest lever |
| A-modes not declared | Every mode implicitly assumed fixable | Say which modes will not be fixed, and why, before the test |
| Growth rate taken as a target | A programme committing to an alpha of 0.5 | Alpha is an outcome of finding and fixing, not a knob |
| The likelihood bias left in | β̂ quoted straight from the estimator | It is high by exactly n⁄(n−1), or n⁄(n−2); the correction is arithmetic, not judgement |
| Test type not recorded | A growth report with no statement of how the test stopped | It sets the correction and the bounds, and it cannot be recovered from the failure times |
| Model fitted to too few failures | Two parameters from six failures, quoted to three digits | Report the interval; early estimates are almost uninformative |
| A passed fit test read as validation | Cramér-von Mises passes, so the model is declared sound | Around twenty failures it misses a modest step change more often than not; look at the plot as well |
| Duane r² quoted as evidence | A high correlation on the log-log plot offered as proof of growth | The cumulative points are not independent, so the statistic does not mean what it means elsewhere |
| Test environment gentler than the world | Laboratory conditions, benign duty cycle, clean power | The environment is part of the result and belongs in the report |
| Prototype not production | Growth demonstrated on hand-built articles | The manufacturing modes arrive later, and they are not in this data |
| The curve extrapolated | A projection carried far past the test to meet the requirement on paper | Extrapolation past the data is a plan, not a measurement |
| Late failures reset nothing | A significant redesign mid-test, with the model carried straight through | A step change in the design breaks the model's assumption |
| Field data never compared | The growth result filed, and the fleet's actual MTBF never checked against it | Close the loop; the comparison is the only validation the model gets |
Five are worth expanding.
The likelihood estimate is biased and the bias is not a matter of opinion. β̂ comes out high by exactly n⁄(n−1) on a time-terminated test and n⁄(n−2) on a failure-terminated one, which on eighteen failures is 5.9 and 12.5 per cent. Because β is below one during growth, an inflated β means an understated growth rate: 0.349 rather than 0.385 on the worked example. The correction takes one multiplication, and once it is applied the likelihood answer and the Duane plot usually stop disagreeing, which removes an argument that has no content in it.
Reporting the cumulative figure is the oldest trick in this field, and it is often not a trick at all but a misunderstanding. The cumulative MTBF averages the whole test, including the failures that were designed out in the first week. It is lower than the instantaneous figure, so quoting it feels conservative and safe, and it systematically understates what the design is worth on the day. The two are related by a single factor, 1 ÷ β̂, which is 1.54 on the worked example, so there is no excuse for confusing them.
A test that finds failures and defers the fixes is not a growth test. It is a reliability measurement with an optimistic label. The whole model assumes each corrective action removes or reduces a mode; if the actions are deferred to a later build, the intensity does not fall, the fitted β sits near one, and the curve is honest even though the programme is not.
The growth potential is the number that should be computed first and is usually computed last. It answers whether the target is reachable by testing at all. If the answer is no, more test time is money spent proving something the arithmetic already knew, and the real answer is a design change. That conversation is far cheaper before the test than after it.
The test is not the fleet. On the worked example the fielded system runs at nearly four times the end-of-test figure; on other programmes the fleet is worse by a similar margin. Both are normal, because the test environment, the build standard, the duty cycle and the failure-counting rules all differ. A growth result is evidence about a test article in a test environment, and treating it as a prediction of field reliability is the assumption that makes the whole activity look dishonest in hindsight.