Almost every FRACAS that fails does so in the same place. The reporting works, the analysis is competent, the corrective actions are written, and the loop is open at the far end: nothing is measured afterwards and nothing goes back into the analyses. What is left is an expensive register of things that broke.
| Pitfall | What it looks like | The guard |
|---|---|---|
| The open loop | Actions closed on a date, never on evidence | The Poisson probability of the observed count computed at closure, or the exposure shortfall recorded |
| Classification written afterwards | Relevance argued report by report | Rules issued before the first report arrives |
| Relevance conflated with chargeability | Engineering data shaped by a warranty argument | Two independent fields, two owners |
| No operating time | A Pareto chart and no rate | The denominator is a mandatory field, filled from the fleet record |
| Under-reporting | A rate that looks good and cannot be defended | Six fields, prefilled, reportable from the field; cross-check against returns |
| No fault found discarded | Units that test good on the bench leave the data entirely | NFF is a diagnostic result; route it to testability |
| Secondary failures double-counted | One event, three data points | Primary or secondary decided at verification |
| The Pareto mistaken for the analysis | The top bar is a symptom, and it stays one | Each bar taken down to a cause somebody can change |
| Only the first question asked | Every failure fixed, the process that produced it kept | The escape question on every report: which analysis should have caught this |
| Corrective action with no scope | Nobody can tell which units have it | Population and effectivity recorded at the time |
| Trend read from small counts | Three this quarter against one last quarter, called a trend | Same Poisson test; small counts rarely survive it |
| Nothing fed back | Prediction and FMECA unchanged after two years in service | Corrections addressed to named documents, with owners |
| The loop unmeasured | A closure rate quoted alone, with recurrence never counted | Closure, backlog age, count over threshold, recurrence, published |
| Hard reports parked | The backlog ages into the interesting ones | Count beyond the age threshold, reviewed by the board |
| A board without authority | Minutes instead of directed action | Membership that can commit design, manufacturing and support |
| Data thrown away at handover | Development-test findings never reach service | One system across the lifecycle, or a defined transfer |
Four of these deserve a sentence more.
Under-reporting is the largest error nobody can see. Every other pitfall leaves a trace in the data. This one improves all the numbers at once, is invisible from inside the system, and is best bounded from outside it: compare the report count against warranty returns, spares consumption or maintenance labour records for the same period. Where they disagree, the failure reports are the ones that are low.
No recurrence to date is not evidence. It is a statement about how long the programme has waited, and it is the direct cause of the recurrence rate. The arithmetic that converts it into evidence takes two minutes, and the honest alternative is also a result worth writing down: the exposure needed is 0.5 million unit-hours and we have 0.1.
A cause that names no document is a description. "Operator error", "component quality" and "environmental conditions" are report categories, not causes. The test is whether somebody can be handed a change: a drawing, a process, a specification, an inspection, a manual.
The escape question is what makes the loop pay for itself twice. Fixing the seal helps this product. Discovering that α = 0.05 was assumed for a mode that turned out to be 0.33, and that the environment driving it was never in the derating analysis, changes how the next three products are analysed. Programmes that ask only "why did it fail" pay full price for their field data and take delivery of half of it.