RAMS Core

PublishedMIL-STD-2155 · MIL-HDBK-2155 · IEC 60300-3-2

FRACAS

The closed loop between the field and the design analyses: report, verify, find the cause, act, and prove the action worked.

Every other analysis in this knowledgebase is a statement about a machine that has not run yet. A prediction estimates a rate, an FMECA divides it into modes, an RBD arranges them, a derating report claims the parts are inside their limits. FRACAS is the one that measures, and the grades it hands back are rarely close.

The letters stand for failure reporting, analysis and corrective action, and the word doing the work is system. Any organisation can collect failure reports; a FRACAS is the machinery that carries each one from a symptom in the field to a change in the design or the process, and then proves the change worked. The last clause is where most implementations stop being a system and become a register.

The loop, and where it opens

Detect, report, verify, analyse, act, and prove. Everything up to and including Act is what most systems already do in some form; the sixth step is a measurement on the fleet months later, and without it nobody can say whether the last twenty corrective actions changed anything.
Detect, report, verify, analyse, act, and prove. Everything up to and including Act is what most systems already do in some form; the sixth step is a measurement on the fleet months later, and without it nobody can say whether the last twenty corrective actions changed anything.
StepWhat it meansWhat goes wrong
DetectSomething happened, and somebody noticedFailures that nobody reports because reporting is slow
ReportOne record, with enough fields to be analysable laterFree text, no operating hours, no configuration
VerifyDid the item actually fail, and howNo fault found, counted as a failure or discarded silently
AnalyseThe root cause, not the symptomThe Pareto chart mistaken for the analysis
ActA change with an owner and a dateAn action that is a note to be careful
ProveThe rate afterwards, measuredSkipped, so the loop never closes

Where the discipline comes from

MIL-STD-2155 established FRACAS as a required system rather than a habit: a closed loop with defined states, a failure review board with the authority to direct action, and the requirement that corrective actions be verified rather than asserted. MIL-HDBK-2155 carries the guidance that replaced it. MIL-STD-785B puts the same thing in a reliability programme as a task in its own right, alongside the failure review board.

On the international side, IEC 60300-3-2 is about the collection of dependability data from the field: what a record has to contain, how operating time is accounted for, and the classification rules without which two organisations' failure rates are not comparable. That last point is the one this module spends most of its arithmetic on.

Programme stageWhat FRACAS is doing
Development testTest-analyse-and-fix: the loop running fast, feeding reliability growth
QualificationEvery test failure classified and dispositioned; the evidence a review needs
Early serviceThe first real grades on the prediction, the mode ratios and the diagnostics
Mature serviceTrend detection, recurrence control, and the data that makes the next design's prediction honest

What it corrects

One year of reports from a 400-unit fleet against the analyses that preceded it: a rate out by 37 per cent, a mode ratio out by a factor of six, and a dominant failure driven by an environment nobody modelled.
One year of reports from a 400-unit fleet against the analyses that preceded it: a rate out by 37 per cent, a mode ratio out by a factor of six, and a dominant failure driven by an environment nobody modelled.

The loop's output is not a report. It is a set of corrections to documents that already exist:

AnalysisWhat the field corrects
PredictionThe item rates, and the environment factor that was guessed
FMECAThe α mode ratios, and the modes nobody listed
TestabilityThe false alarm rate and the no-fault-found share, which only the field can measure
MaintainabilityThe task times, against a technician who was hurried and cold
DeratingOverstress found in service, and the environments the analysis never had
RBD and FTABlock and event rates, and any common cause that actually happened

Want to see this on a live system model? Request a walkthrough.