FRACAS · Chapter 6

Common Pitfalls

The mistakes seen in practice and the guards against them.

Almost every FRACAS that fails does so in the same place. The reporting works, the analysis is competent, the corrective actions are written, and the loop is open at the far end: nothing is measured afterwards and nothing goes back into the analyses. What is left is an expensive register of things that broke.

PitfallWhat it looks likeThe guard
The open loopActions closed on a date, never on evidenceThe Poisson probability of the observed count computed at closure, or the exposure shortfall recorded
Classification written afterwardsRelevance argued report by reportRules issued before the first report arrives
Relevance conflated with chargeabilityEngineering data shaped by a warranty argumentTwo independent fields, two owners
No operating timeA Pareto chart and no rateThe denominator is a mandatory field, filled from the fleet record
Under-reportingA rate that looks good and cannot be defendedSix fields, prefilled, reportable from the field; cross-check against returns
No fault found discardedUnits that test good on the bench leave the data entirelyNFF is a diagnostic result; route it to testability
Secondary failures double-countedOne event, three data pointsPrimary or secondary decided at verification
The Pareto mistaken for the analysisThe top bar is a symptom, and it stays oneEach bar taken down to a cause somebody can change
Only the first question askedEvery failure fixed, the process that produced it keptThe escape question on every report: which analysis should have caught this
Corrective action with no scopeNobody can tell which units have itPopulation and effectivity recorded at the time
Trend read from small countsThree this quarter against one last quarter, called a trendSame Poisson test; small counts rarely survive it
Nothing fed backPrediction and FMECA unchanged after two years in serviceCorrections addressed to named documents, with owners
The loop unmeasuredA closure rate quoted alone, with recurrence never countedClosure, backlog age, count over threshold, recurrence, published
Hard reports parkedThe backlog ages into the interesting onesCount beyond the age threshold, reviewed by the board
A board without authorityMinutes instead of directed actionMembership that can commit design, manufacturing and support
Data thrown away at handoverDevelopment-test findings never reach serviceOne system across the lifecycle, or a defined transfer

Four of these deserve a sentence more.

Under-reporting is the largest error nobody can see. Every other pitfall leaves a trace in the data. This one improves all the numbers at once, is invisible from inside the system, and is best bounded from outside it: compare the report count against warranty returns, spares consumption or maintenance labour records for the same period. Where they disagree, the failure reports are the ones that are low.

No recurrence to date is not evidence. It is a statement about how long the programme has waited, and it is the direct cause of the recurrence rate. The arithmetic that converts it into evidence takes two minutes, and the honest alternative is also a result worth writing down: the exposure needed is 0.5 million unit-hours and we have 0.1.

A cause that names no document is a description. "Operator error", "component quality" and "environmental conditions" are report categories, not causes. The test is whether somebody can be handed a change: a drawing, a process, a specification, an inspection, a manual.

The escape question is what makes the loop pay for itself twice. Fixing the seal helps this product. Discovering that α = 0.05 was assumed for a mode that turned out to be 0.33, and that the environment driving it was never in the derating analysis, changes how the next three products are analysed. Programmes that ask only "why did it fail" pay full price for their field data and take delivery of half of it.


Want to see this on a live system model? Request a walkthrough.