RAMSynapse
Log inSign up

Safety · Worked example

Electronics and high tech

Storage array controller pair

Industry overview: Electronics and high tech at RAMSynapse

Two active-active controller boards, dual hot-swappable power supplies, drive shelves, and a lithium backup unit that holds the write cache alive when the mains goes. The array is engineered to five nines, which is 5.3 minutes of downtime a year, and every analysis the programme runs is about keeping a service available. The safety analysis is looking at a different object entirely.

Nothing in this rack can hurt anybody by stopping. Losing the array breaches a service level agreement. What can hurt somebody is the one component that stores energy: a lithium unit whose failure mode is thermal runaway, mounted in a chassis, in a rack, in a room full of other racks, in a building. This page is about a hazard that is not a loss of function, and about what happens to the usual engineering instincts when the thing you are analysing is a source rather than a path.

The technique, and why this one

An energy-based hazard identification, followed by a barrier argument against release, with the item failure rate converted to a mode rate before any probability is quoted. The functional methods that serve the rest of the array, RBD success paths and a Markov model of the repairable controller pair, ask what happens when a block stops working. The hazard here is produced by a component doing something, not by a component stopping, and no success-path method has a place to put that.

ObjectModelValueWhat the analysis wants from it
Controller board, 2 fittedFIT parts count, exponential3,000 FIT = 3.0 per 10⁶ h eachfunction only
Power supply, 2 fittedexponential, hot-swappable1,500 FIT eachfunction only
Driveannualised failure rate0.4% AFRfunction only
Single controller unavailability, 4 h responseMarkovq = 1.2 × 10⁻⁵function only
Controller pair, independent repair1.4 × 10⁻¹⁰function only
Lithium backup unitenergy source900 FIT = 0.9 per 10⁶ h, all modesthe hazard, after a mode split
Barrierslayer creditcell protection, containment, state-of-health monitorhow release is prevented and limited

The functional answer, and why it ends the wrong conversation

Work the availability side first, because it is short. A single controller with a four-hour on-site response sits at q = 1.2 × 10⁻⁵. Two of them with independent repair give

q_pair = q² = (1.2 × 10⁻⁵)² = 1.44 × 10⁻¹⁰

and the two-state repairable form used in the reliability treatment gives 2q² = 2.9 × 10⁻¹⁰. Both are small enough to stop being interesting: the five-nines target of 0.99999 corresponds to (1 − 0.99999) × 525,600 = 5.3 minutes a year, and the controller pair contributes microseconds of it. The real constraints on availability are the shared backplane, the shared firmware build and the disruptive update, none of which the pair arithmetic sees.

That is a complete and correct analysis of the wrong object. It says nothing at all about the lithium unit, because the lithium unit's contribution to loss of function is negligible and its contribution to loss of a building is not in the model.

The number the table does not contain

Start from what is given. The lithium backup unit runs at 900 FIT, which is 0.9 per 10⁶ hours, and over a five-year service life of 43,800 hours the expected failures per unit are

0.9 × 10⁻⁶ × 43,800 = 0.0394

about one unit in twenty-five, or 39.4 unit failures per thousand arrays over five years. Now the essential move. That 900 FIT covers every way the unit can fail, and the overwhelming majority of those ways are benign: it stops holding charge, it reports a degraded state, it is swapped at the next service call. The hazardous mode is some fraction f of the 900, and the table does not give f.

Watch what the answer does when f moves:

f = 1 (item rate treated as hazard rate): 0.9 per 10⁶ h → 39.4 thermal events per 1,000 arrays in 5 years

f = 10⁻²: 9 FIT → 0.39 events per 1,000 arrays in 5 years

f = 10⁻⁴: 0.09 FIT → 3.9 × 10⁻³ events per 1,000 arrays in 5 years

Four orders of magnitude of answer, driven entirely by a number that has to be established mode by mode. The mode split is the analysis; the item rate is merely an input to it, and no honest argument may quietly treat one as the other. This is the step an FMEA exists to produce and the step most often skipped, precisely because the component rate is available and the mode split has to be worked for.

Barriers, and the arithmetic of layers

Three layers, each credited an order of magnitude, and two of them coupled by the same propagation mechanism. A stored-energy hazard is a barrier stack rather than a success path, and the layer credit is worth only as much as the independence claim under it.
Three layers, each credited an order of magnitude, and two of them coupled by the same propagation mechanism. A stored-energy hazard is a barrier stack rather than a success path, and the layer credit is worth only as much as the independence claim under it.

With an initiating frequency established, the residual risk is the initiating frequency reduced by the independent layers between the event and the harm. The design response named for this unit is three layers, and they make three different assumptions about what has already gone wrong: cell-level protection tries to prevent the runaway, enclosure containment accepts that prevention can fail and limits what the failure reaches, and the state-of-health monitor detects the approach in time for a human to act.

Credited at the conventional order of magnitude per independent layer, and taking f = 10⁻² as a working figure:

f_release = 9 × 10⁻⁹ × 10⁻¹ × 10⁻¹ × 10⁻¹ = 9 × 10⁻¹² per hour per array

over 5,000 array-years: 9 × 10⁻¹² × 43,800 × 1,000 = 3.9 × 10⁻⁴ events

one event per 12.7 million array-years. The credit is only as good as the independence, and in a battery pack the independence has a specific enemy: cell-to-cell propagation. A single cell going into runaway that ignites its neighbours converts one cell's event into the whole unit's, which means the cell-level protection and the containment are not independent layers at all but two halves of one. Propagation is the beta factor of a battery pack, and it is attacked with spacing, thermal barriers and vent routing rather than with better cells.

Where the two properties point in opposite directions

Here is the finding that makes this row worth its place. Redundancy, the move that improves every functional number in the table, makes the energy hazard strictly worse.

function: one controller q = 1.2 × 10⁻⁵ → pair q = 1.4 × 10⁻¹⁰, better by a factor of 8.6 × 10⁴

energy: one lithium unit 9 × 10⁻⁹ per hour → two units 1.8 × 10⁻⁸ per hour, worse by a factor of 2

A second backup unit doubles the stored energy in the chassis and doubles the frequency of the initiating event, while improving write-cache protection and every availability figure the programme reports. This is the only place in the chapter where an engineer's reflex, add another one, is quantifiably the wrong move, and it is worth noticing that the availability model would recommend it and never mention the cost.

Detection has a similar asymmetry. In-service telemetry and drive self-monitoring reach 97 per cent of failure rate, so the uncovered fraction of the lithium unit is

0.03 × 0.9 = 0.027 per 10⁶ h = 27 FIT

over 43,800 h: 27 × 10⁻⁹ × 43,800 = 1.18 × 10⁻³ per unit, about one unit in 850

which announces nothing at all. And the state-of-health monitor should not be measured by its coverage percentage in the first place. Its value is warning time, and a monitor that flags an approaching runaway ten minutes ahead and one that flags it ten days ahead score identically on coverage and are worth completely different amounts. The testability page works the coverage arithmetic that this hazard needs restated.

What the analysis tells the engineer to do

Get the mode split before quoting any hazard probability. The four-order-of-magnitude spread above is the finding, and closing it is worth more than any other analysis activity on this component.

Do not add a second lithium unit for availability. The availability it buys is already below the measurement floor and the hazard it adds is real and doubles.

Specify containment against propagation explicitly, and test it as a pack rather than as a cell, because the pack is where the independence claim is either true or false.

Measure the state-of-health monitor in warning time and design the response around the time it actually gives.

Stop returning no-fault-found lithium units to stock. At 24 per cent, the no-fault-found rate on returns is the dominant warranty cost and it is also a safety problem: a unit whose fault was not found and whose state of health is therefore asserted rather than measured is going back into a rack on the strength of a diagnostic that failed once already.

What a different technique would have given

Take the lithium unit through the same treatment as everything else in the array. An availability model calls it a block whose failure loses write-cache protection, applies the four-hour response, and reports

unavailability = 0.9 × 10⁻⁶ × 4 = 3.6 × 10⁻⁶

notes that the array survives it comfortably, and closes the item as non-critical. Every step of that is correct and the conclusion is upside down, because the model asks what happens when the battery stops working and the hazard is what happens when it works too hard. A function-loss analysis cannot see an energy-release hazard, and the reason is structural rather than a matter of effort: success-path methods enumerate the ways a system fails to deliver, and this failure delivers something extra.

An FMEA on loss of function fails the same way for the same reason. It would list "backup unit fails to hold charge", assign the effect "write cache unprotected during mains loss", note the mitigation, and never write a line about heat. The five-nines target that drives every other decision on this array is silent on the only object in the rack that can start a fire, and a programme that never runs an energy-based identification will not discover that, because nothing in its method asks the question.


Want to see this on a live system model? Request a walkthrough.