RAMSynapse
Log inSign up

FMECA · Chapter 4

Worked Example

The method applied end-to-end on a concrete system, with numbers.

Consider the pitch-control drive card of a wind turbine: the power and control electronics that hold each blade at the commanded angle and feather it when the turbine has to stop. It is deliberately none of the systems the industry examples carry, and its numbers belong to no source. Every value below is an illustrative teaching figure; the arithmetic and the column definitions are MIL-STD-1629A's.

The ground rules, settled before the first row:

Ground ruleThis analysis
Indenture levelsPart → card → pitch axis → turbine
ApproachPiece-part, from the card's parts list
Exposure t8,760 hours, one turbine-year: the pitch system is energised whenever the turbine is
SeverityClassified on the end effect at turbine level, with compensating provisions removed
α sourcePart-family mode distributions, replaced by field returns where the fleet has them
λp sourceThe parts-count prediction for this card

The card predicts at 38.3 failures per 10⁶ hours, an MTBF of 26,110 hours, distributed over eight items:

Itemλp per 10⁶ hShare of the card
V1 · IGBT module12.031.3%
C1–C4 · DC-link electrolytic bank9.625.1%
U1 · microcontroller5.013.1%
U5 · resolver interface4.511.7%
U2 · gate-drive IC3.28.4%
PCB · board and solder joints2.05.2%
X1 · 24-way connector1.23.1%
R7 · current-sense resistor0.82.1%

Task 101: the modes and what follows from them

The worksheet is 22 rows, one per mode-and-effect pair. Six of them carry the analysis:

Item · modeLocal effectEnd effect at the turbineDetectionSev
U5 · plausible but wrong angleangle reported wrong by up to 8°blade pitched to the wrong angle, rotor overspeed on a gustnone: the value passes every range checkI
C1–C4 · shortconverter shoot-throughdrive destroyed, blade stuck at the last anglenone before the eventI
U1 · erroneous output, undetectedwrong pitch command issuedblade driven off the commanded anglenone: no independent check of the commandI
U2 · output stuck highshoot-through in one leguncommanded torque on the pitch axisnone before the eventI
V1 · short-circuitleg fails, drive tripsblade held at the last angle, turbine stopsdesaturation detectII
C1–C4 · ESR driftride-through shortenspitch rate falls, turbine deratesDC-link ripple trendIII

Two things are already visible without any arithmetic. Every Class I row's detection column says "none", and the reason is the same in each case: these are the modes where the card keeps working and reports something false. And the item with the largest failure rate on the card, the IGBT at 31.3 per cent, does not appear in the Class I list at all: its modes stop the turbine, which is expensive and safe.

Task 102: the criticality numbers

Each row gains four values and produces one. The resolver's dangerous mode, in full:

Cm = β · α · λp · t = 0.5 × 0.20 × 4.5 × 10⁻⁶ × 8,760 = 3.94 × 10⁻³ per turbine-year

The α of 0.20 says a fifth of the resolver interface's failures present as a plausible wrong angle rather than as a dead signal; the β of 0.5 says that when that happens, the wrong angle produces an overspeed event about half the time, the other half being caught by wind conditions that make it harmless. The four Class I rows:

Modeαλm = α·λpβCm
C1–C4 · short0.100.960.504.20 × 10⁻³
U5 · plausible wrong angle0.200.900.503.94 × 10⁻³
U1 · erroneous output0.150.750.402.63 × 10⁻³
U2 · output stuck high0.250.800.302.10 × 10⁻³
Class I total1.29 × 10⁻²

Item criticality is the sum within a class, so those four numbers are also the four items' Cr in Class I. Across all four classes:

ClassTotal Cr per turbine-yearDominated by
I · catastrophic1.29 × 10⁻²the capacitor bank, 32.7% of the class
II · critical4.97 × 10⁻²V1, the IGBT, 86.7% of the class
III · marginal2.92 × 10⁻²the capacitor bank's ESR drift, 79.1% of the class
IV · minor0the three nuisance-trip modes, all at β = 0

One Category I event every 78 turbine-years, which on a 200-turbine farm is 2.6 a year and is the number the operator will eventually experience. Class IV coming out at zero is not an error: those modes produce spurious trips and no damage, so their β against a damaging end effect is zero. They still cost money, and that cost is Task 103's problem rather than criticality's.

What the two rankings disagree about

Share of the card's failure rate against Class I criticality. The IGBT is the biggest single contributor to the rate and contributes nothing catastrophic; the resolver interface fails a third as often and carries 31 per cent of the Class I criticality.
Share of the card's failure rate against Class I criticality. The IGBT is the biggest single contributor to the rate and contributes nothing catastrophic; the resolver interface fails a third as often and carries 31 per cent of the Class I criticality.

Ranked by failure rate the answer is the IGBT, then the capacitors, then the microcontroller. Ranked by Class I criticality it is the capacitors, the resolver interface, the microcontroller and the gate drive, and the IGBT is not on the list. Both rankings are correct and they are answering different questions: the first is where the turbine's downtime comes from, the second is where its hazard comes from. A programme that spends its budget on the top of the rate list buys availability and leaves the safety case untouched.

The finding that pays for the analysis

All four Class I modes share one property: nothing on the card can tell that they have happened. A second resolver channel with a plausibility vote moves one β from 0.5 to 0.05 and removes 28 per cent of the card's catastrophic criticality.
All four Class I modes share one property: nothing on the card can tell that they have happened. A second resolver channel with a plausibility vote moves one β from 0.5 to 0.05 and removes 28 per cent of the card's catastrophic criticality.

One hundred per cent of the Class I criticality sits in modes with no detection. That is not a coincidence and it is the general shape of the result: a mode that announces itself gets a compensating provision, and what remains is what the card cannot see. The four modes have the same character, which the testability page calls the same thing: the item continues to operate and reports something false.

The design action falls out of the number. Adding a second resolver channel and voting the two for plausibility moves that mode's β from 0.5 to about 0.05, because a wrong angle now has to survive a comparison to reach the controller:

Class I criticality per turbine-year
As built1.29 × 10⁻²
With a voted second resolver channel0.93 × 10⁻²

A 28 per cent cut in the card's catastrophic criticality for a connector, a converter and some code, which is more than halving the failure rate of any part on the card would deliver. That comparison is the argument for doing Task 102 rather than stopping at Task 101: the qualitative worksheet identifies the mode, and only the arithmetic says what fixing it is worth against everything else on the list.

Task 103, on the same rows

The maintainability reading uses the worksheet already written and adds three columns:

ModeHow it is isolatedWhat comes offTime
V1 · short-circuitdesaturation flag names the legthe whole card3.5 h, one tower visit
C1–C4 · ESR driftripple trend, weeks of warningthe card, on a planned visit3.5 h, no lost production
X1 · intermittent contactnot isolated: the trip log is ambiguousoften the wrong card first3.5 h, twice

The third row is the one worth reading. An intermittent connector is a Class IV mode with a criticality of zero and a maintenance cost that exceeds several Class III modes together, because it produces repeat visits to a nacelle. Criticality ranks hazard, not cost, and Task 103 is where the second ranking gets built, feeding maintainability and the spares case rather than the safety case.

Task 104, and why this analysis does not have one

Damage mode and effects analysis asks what a stated threat does to each item: a fragment, a blast overpressure, a directed-energy exposure. A wind turbine has no threat specification, so Task 104 is marked not applicable, with that reason recorded, which is the honest treatment. Had this been a naval or airborne installation, the same eight items would be re-analysed against a damage mechanism rather than a failure mode, and the answer would not resemble the criticality ranking above: the resolver interface is a small part in a protected housing and the DC-link capacitor bank is a large one that fragments, so the rankings invert. That inversion is the reason the task exists as a separate analysis rather than as another column.


Want to see this on a live system model? Request a walkthrough.