RAMSynapse
Log inSign up

FMECA · Chapter 2

Theoretical Foundations

Definitions, units, models, and the assumptions that bound them.

An FMECA has four moving parts: a structure to walk, a vocabulary for what can go wrong, a scale for how much it matters, and an arithmetic for how often. The first three are what make the fourth mean anything.

Indenture levels, and the effect that has three of them

The analysis walks a hierarchy, and MIL-STD-1629A calls each step an indenture level: the part, the assembly it sits in, the subsystem that contains it, the system. Every mode is therefore stated at one level and has consequences at others, which is why the worksheet asks for the effect three times:

EffectStated atThe question it answers
LocalThe level of the item itselfWhat does this mode do to this item?
Next higherOne level upWhat does the assembly do about it, or fail to?
EndThe system boundaryWhat does the operator experience?

The end effect is the column that carries the severity, and it is the one most often filled in with a restatement of the local effect. "Capacitor short-circuits" is a mode; "DC link collapses" is a local effect; "blade stuck at the last commanded angle during a gust" is an end effect, and only the third can be classified.

Severity: four classes, no arithmetic

Severity classifies the end effect, never the mode. It is a statement about consequence and needs no failure rate, which is why it is stable while everything numerical is still moving.

ClassMIL-STD-1629AIn practice
ICatastrophic: may cause death or system lossA different kind of event, not a worse one
IICritical: severe injury, major damage, or mission lossExpensive and survivable
IIIMarginal: minor injury or damage, delay, degradationAn availability problem
IVMinor: no injury or damage, but unscheduled maintenanceA cost-of-ownership problem

Two rules keep the column honest. Classify with the compensating provision removed, because a mode's severity is what it does, not what the guard against it does; the guard belongs in its own column and gets its own analysis. And classify per mission phase, since the same mode is routinely Class I in one phase and Class IV in another.

Where a mode's share of the rate comes from

The four terms of the criticality expression and their four different sources. Only λp is measured; α is usually generic, t is declared, and β is an engineer's judgement on a four-value scale. The result is exactly as good as its weakest term.
The four terms of the criticality expression and their four different sources. Only λp is measured; α is usually generic, t is declared, and β is an engineer's judgement on a four-value scale. The result is exactly as good as its weakest term.

An item has one failure rate and several failure modes, so the rate has to be divided. The failure mode ratio α is that division: the fraction of the item's failure rate that appears as this mode, with Σα = 1 over the item's modes. It comes from one of three places, in descending order of comfort: the item's own field returns, a published failure-mode distribution for that part family, or an engineering estimate that should be labelled as one.

The conditional probability β is the second judgement, and the sharper one: given that this mode occurs, how often does the stated end effect actually follow? MIL-STD-1629A offers four values, and the coarseness is deliberate:

βMeaning
1.00Actual loss: the effect follows certainly
0.50Probable loss: more likely than not
0.10Possible loss
0No effect

The criticality arithmetic

For a mode, the mode criticality is the expected number of failures of that mode, producing that end effect, in the exposure considered:

Cm = β · α · λp · t

and for an item, the item criticality in a severity class is the sum over that item's modes in the class:

Cr = Σ Cm, over the modes of one item in one severity class

Two properties of that definition catch people out. Criticality is summed within a severity class and never across classes, because a Class I number and a Class IV number are not the same kind of quantity and adding them produces nothing. And Cm is an expected count over an exposure, so t has to be the exposure that mode actually sees: a mode that can only occur during a start-up sequence does not get the mission's whole duration.

Where failure rates are not available, MIL-STD-1629A defines a qualitative criticality analysis instead, which places each mode in a probability-of-occurrence level, A (frequent, more than 0.20 of the item's failure probability) through E (extremely unlikely, under 0.001), and plots those against severity. The matrix is the same shape and the ranking survives; only the resolution changes.

The criticality matrix

The criticality matrix with the worked example's card on it. The IGBT is highest on the page and is a Class II problem; the four items in the Class I column are an order of magnitude below it and are the whole safety case.
The criticality matrix with the worked example's card on it. The IGBT is highest on the page and is a Class II problem; the four items in the Class I column are an order of magnitude below it and are the whole safety case.

The matrix plots severity across and criticality (or the qualitative probability level) up, and its diagonal is the ranking. It exists to stop two failure modes of the technique itself: ranking by rate, which puts availability problems above safety ones, and ranking by severity alone, which treats a Class I mode occurring once in ten thousand years as urgent.

Reading it is a matter of distance from the origin, with one asymmetry: the severity axis is not linear and cannot be traded against the criticality axis. A Class I mode at 10⁻⁴ and a Class III mode at 10⁻¹ are not comparable quantities; the first is a hazard to be eliminated or shown to be adequately remote, the second is a cost to be managed.

What the automotive line does instead

The other dialect ranks with a risk priority number, RPN = S × O × D, from three ten-point scales for severity, occurrence and detection. It needs no failure rate at all, which is its attraction, and it multiplies three ordinal scales, which is its defect: the product of ranks is not a rank, 8 × 2 × 8 and 4 × 8 × 4 both give 128 and mean entirely different things, and a threshold on RPN ends up hiding a severity 9 behind two low scores. The 2019 AIAG-VDA handbook replaced it with an action priority table that is explicitly ordered by severity first. The MIL-STD-1629A form avoids the problem by not multiplying ordinals: α and β are probabilities, λp is a rate, and severity stays a separate axis rather than a factor.

What an FMECA cannot do

  • It is single-fault. Each row considers one mode in isolation. Combinations are what a fault tree or an RBD is for, and an FMECA that has been asked about combinations has been asked the wrong question.
  • It does not find systemic causes. A mode caused by a specification error, a shared calibration or a common environment appears once per affected item, and nothing in the worksheet joins them up.
  • It is only as complete as its mode list. The analysis is exhaustive over what is in the list and blind to what is not, which is the argument for a piece-part pass and for a library that accumulates.
  • It says nothing about timing or sequence. "A fails then B" is outside the worksheet's grammar.

Want to see this on a live system model? Request a walkthrough.