The analysis is a walk through a structure with a fixed set of questions asked at every node. Most of its failures are procedural: the ground rules were never written, the indenture level drifts, the end effect column never leaves the item, or the worksheet is finished after the design has stopped listening.
1. The ground rules come before the first row
The ground rules are the part of an FMECA that makes two analysts produce the same worksheet. They are short and they are contractual:
| Ground rule | Why it has to be decided once |
|---|---|
| The indenture levels and the coding scheme | Otherwise half the worksheet is written at part level and half at assembly level |
| The mission phases and the exposure time of each | t in the criticality expression comes from here, and so does phase-dependent severity |
| The severity definitions, in this system's language | "Major system damage" means something specific per programme |
| What is in scope: hardware, software, human action, interfaces | An interface nobody owns is where the missed modes live |
| Whether compensating provisions are credited in severity | They are not, in this module's convention: they get their own column |
| The α source and the failure-rate source | A worksheet mixing generic and measured ratios without saying which is which cannot be reviewed |
2. Choose the approach, per assembly
Functional where the design is not yet settled, piece-part where it is. Mixing them across assemblies is normal and should be recorded per assembly rather than argued about globally. The one combination to avoid is a functional analysis whose criticality numbers are quoted as though they were piece-part results: a function has no failure rate of its own until something underneath it does.
3. Build the item list from the structure that already exists
The list comes from the BOM or the functional breakdown, not from a fresh document, because an FMECA on a private list of items is unmaintainable and cannot be checked for coverage. Coverage is checkable: every item in the structure either has rows or has a recorded reason for having none.
4. For each item, enumerate the modes
The library is the memory. Standard mode sets exist per part family, and the analysis starts from them rather than from a blank column: an electrolytic capacitor opens, shorts and drifts; a connector opens and goes intermittent; a microcontroller halts, corrupts and produces a plausible wrong answer. The last of those is the class most often missed, and it has a name worth using in reviews: the mode where the item continues to work and lies.
5. Follow each mode up to an end effect
Three columns, three levels, and a discipline: the end effect must be phrased in the operator's terms, at the system boundary, in a mission phase. If the end effect cannot be written without saying "depending on", the row needs to become two rows.
6. Classify severity, and record the detection and the compensation
Severity on the end effect with the compensating provision removed; then, separately, how the failure is detected (built-in test, an operator's observation, a periodic check, or nothing) and what compensates for it if it happens. The detection column is the single most valuable output of Task 101 for the rest of the programme, and it is the one an analyst under time pressure fills with "BIT" without asking whether the BIT actually covers this mode.
7. Task 102: put numbers on it
For every row: α from the library or the field data, λp from the prediction for that item, t from the phase, β from judgement against the four-value scale. Then Cm = β·α·λp·t, and Cr per item per severity class.
Two checks catch most arithmetic errors. Σα = 1 over each item's modes, which fails immediately when a mode has been added without rebalancing. And Σ(α·λp) over the worksheet returns the system failure rate the prediction produced, which fails when items have been missed or double-counted.
8. Plot the matrix and read it in the right order
Class I first regardless of criticality, then Class II, then by criticality within each class. The output of this step is not a ranked list but a set of findings, each of which is one of four things:
| Finding | The action it implies |
|---|---|
| A single-point failure with a Class I or II end effect | Eliminate it, add redundancy, or justify it explicitly in the safety case |
| A Class I mode with no detection | Add detection, or add an independent compensating provision, and price both |
| A high-criticality mode with a large α | Attack the mode: derating, a different part family, a screen |
| A cluster of Class III and IV modes | An availability and cost problem, and the input to maintenance planning |
9. Task 103: read the same rows for the repair case
The worksheet already knows what fails and how it is detected. Task 103 adds what it takes to put right: how the fault is isolated, what has to come off, how long it takes, what spares it consumes. That is the input maintainability needs and the reason a maintainability prediction should not be built on a separate mode list.
10. Task 104, when the threat is deliberate
Damage mode and effects analysis asks the same three-column question about damage from a stated threat rather than about a random failure: what a fragment, a blast overpressure or a directed-energy exposure does to this item, and what follows from it. It shares the worksheet's grammar and almost nothing else, since its "rate" is a probability of being hit rather than a failure rate, and its purpose is survivability and vulnerability reduction rather than reliability. It applies to combat systems and to little else, which is why it is the task most often marked not applicable, and it should be marked so explicitly rather than forgotten.
11. Close the loop, or the worksheet is an archive
Every finding gets an owner, a decision and a date, and the analysis is re-run when the design moves, when the prediction moves, and when the field says something different from the library. The α ratios are the first thing real returns correct, and a worksheet whose ratios have never been touched by field data is still carrying the assumptions it was born with.