Safety's mathematics has an awkward job: it must combine a quantity that is often measurable (how often something fails) with a quantity that is never measurable (how bad the outcome would be), and produce a decision that a regulator, a court, or a bereaved family might one day examine. Every framework in this chapter is an attempt to do that honestly. They differ in vocabulary and in industry, and they agree on far more than their notation suggests: classify the consequence, bound the likelihood, demand more rigour as the consequence worsens, and never let a single failure produce a catastrophe.
Notation used in this topic
| Symbol | Reads as | Meaning |
|---|---|---|
| RAC | risk assessment code | one severity category paired with one probability level |
| DAL | development assurance level | the aviation ladder, A (most severe) to E |
| FDAL, IDAL | function and item DAL | the level carried by a function, and by each item implementing it |
| SIL | safety integrity level | the functional safety ladder, SIL 1 to SIL 4 |
| ASIL, QM | automotive SIL, quality management | the automotive ladder, determined from severity, exposure and controllability |
| THR | tolerable hazard rate | the railway expression of an integrity target, per hour |
| PFD, PFH | probability of failure on demand, per hour | the low-demand and high-demand target measures |
| λ_DU, T | dangerous undetected rate, proof-test interval | the two terms behind the λT/2 approximation |
| ALARP, SFAIRP | as low as reasonably practicable | the duty to reduce risk short of gross disproportion |
| CCA, ZSA, PRA, CMA | common cause analysis and its triad | zonal, particular risks and common mode studies |
Every abbreviation used anywhere in RAMS Core is defined in the glossary.
Risk: two dimensions, one decision
Risk is the pairing of severity and likelihood, and the discipline's first move is always to put both on a defined scale so that different hazards can be compared by different people and reach the same answer.
The defence standard practice for system safety (MIL-STD-882E) gives the canonical four-level severity scale, defined by the worst credible outcome of the potential mishap:
| Severity | Category | The consequence that defines it |
|---|---|---|
| Catastrophic | 1 | Death, permanent total disability, irreversible significant environmental impact, or loss above the standard's top monetary threshold |
| Critical | 2 | Permanent partial disability, injuries or illness hospitalising at least three people, reversible significant environmental impact, or major monetary loss |
| Marginal | 3 | Injury or illness causing lost work days, reversible moderate environmental impact, or moderate monetary loss |
| Negligible | 4 | Injury or illness not causing a lost work day, minimal environmental impact, or minor monetary loss |
and a six-level probability scale from Frequent (A) through Probable (B), Occasional (C), Remote (D) and Improbable (E) to Eliminated (F), the last being reserved for hazards that have been designed out entirely. The pairing of one severity with one probability is a Risk Assessment Code, and a matrix maps every code onto a risk level of High, Serious, Medium or Low, which in turn determines who is allowed to accept it. That last link is the point of the whole apparatus: the matrix is not a calculator, it is a routing table that sends each hazard to the right decision-maker.
Two properties of the scales repay attention. Severity is assessed on the worst credible outcome, not the average one, because averaging consequence across scenarios is how a catastrophic hazard gets diluted into a marginal one on paper. And the probability scale is explicitly qualitative-if-necessary, quantitative-if-possible: where real rate data exists the standard prefers it, and the Improbable band is conventionally anchored around a one-in-a-million order of magnitude, but a scale that forces invented numbers produces false precision rather than insight.
Reading a risk matrix honestly
The matrix is the most widely used and most widely abused instrument in the discipline, so its limits belong in the foundations rather than in a footnote:
- The scales are ordinal, not arithmetic. Severity 2 is not "twice" severity 4, and probability C is not a number. Multiplying category indices to get a "risk score", then averaging or summing those scores across hazards, is arithmetic on labels and produces conclusions the inputs cannot support.
- Bands compress enormous ranges. A single probability band can span two orders of magnitude; two hazards in the same cell can differ in real risk by a factor of a hundred.
- The colouring encodes a policy, not a physical fact. Where the boundary between Serious and Medium falls is an organisational risk-appetite decision that the matrix silently naturalises. Different sectors draw it differently, and correctly so.
- Matrices tempt severity deflation. Because a Catastrophic assessment triggers the most expensive obligations and the most senior acceptance, there is chronic organisational pressure to argue a hazard down one category. The countermeasure is procedural: severity is assigned from the consequence definitions before any mitigation is considered, and mitigation moves the probability, not the severity, unless the design genuinely changes what happens.
Used properly, the matrix is excellent at what it is for: triaging a large hazard inventory quickly, forcing an explicit consequence judgment, and routing acceptance authority. It is not a quantitative risk assessment, and the frameworks in the next sections exist because for the highest-consequence hazards, "High risk, needs work" is not a sufficient answer.
Quantitative targets: the inverse relationship
Where consequences are severe enough and the industry mature enough, qualitative bands give way to numerical probability targets, built on a principle every sector states in almost identical words: the more severe the consequence, the less frequently it may be permitted to occur.
Civil aviation gives the most fully worked example, set out in the airworthiness authorities' guidance on system design and analysis. Failure conditions are classified by the severity of their effect on aeroplane and occupants, from No Safety Effect through Minor, Major and Hazardous to Catastrophic, and each class carries an average probability per flight hour it must not exceed, stepping by roughly two orders of magnitude per class down to the famous 10⁻⁹ per flight hour for catastrophic failure conditions. The number is not arbitrary and its derivation is instructive: starting from the historical accident rate for large transport aircraft, allocating the portion of that rate attributable to systems, and dividing it among the order of a hundred potential catastrophic failure conditions in a modern aeroplane, one arrives at a per-condition budget in the region of one in a billion flight hours. It is a budget allocation, exactly like the reliability allocation in this knowledgebase, with the top-level number set by what society has come to expect of air travel.
That target comes with a qualitative partner requirement that matters just as much: catastrophic failure conditions must not result from a single failure. The probability arithmetic and the no-single-point rule work together, because a lone number can always be argued down with optimistic data, while a structural rule about cut-set order cannot.
| Aviation failure-condition class | Effect, in the regulatory language | Order-of-magnitude probability target per flight hour |
|---|---|---|
| Catastrophic | Loss of the aeroplane, multiple fatalities | 10⁻⁹ |
| Hazardous (Severe-Major) | Large reduction in safety margins, serious or fatal injury to a small number of occupants | 10⁻⁷ |
| Major | Significant reduction in safety margins, discomfort or injuries to occupants | 10⁻⁵ |
| Minor | Slight reduction in safety margins, routine crew action | 10⁻³ |
| No safety effect | No effect on operational capability or safety | No requirement |
Railway practice expresses the same idea as a Tolerable Hazard Rate: a permitted rate of occurrence per hour for each identified hazard, apportioned down from a system-level safety target, with the deepest integrity band conventionally covering the 10⁻⁹ to 10⁻⁸ per hour region. Process and machinery sectors express it as the required risk reduction between the unprotected process risk and the tolerable risk. Different units, one grammar.
Integrity levels: the ladders
Once the consequence classification exists, every sector attaches to it a level that governs how much rigour the design and its evidence must carry. The four best known:
| Framework | Levels | Anchored on | What the level actually demands |
|---|---|---|---|
| Aviation development assurance (the ARP4754 lineage) | DAL A (most severe) to DAL E | Failure condition classification | Rigour of the development process for the function and its items: requirements, verification, independence of review |
| Functional safety (IEC 61508) | SIL 1 to SIL 4 | Required risk reduction | A probability target band plus systematic capability and architectural constraints |
| Automotive (ISO 26262) | QM, ASIL A to ASIL D | Severity, exposure and controllability, combined through a determination table | Process rigour and technical measures across the automotive safety lifecycle |
| Railway (EN 50126/50129 family) | SIL 1 to SIL 4 with tolerable hazard rates | Apportioned hazard rate | Quantified THR for random failures plus process requirements for systematic ones |
The functional safety ladder is the one with the crispest numbers, because it separates the two failure classes explicitly. For a protective function operating in low demand mode (it sits dormant and is called on rarely), the level is defined by bands of average probability of dangerous failure on demand: each SIL step is one order of magnitude, with SIL 4 the most demanding band and SIL 1 the least. For functions in high demand or continuous mode, the same ladder is expressed as an average frequency of dangerous failure per hour instead. The low-demand quantity is exactly the on-demand unavailability from the availability foundations, which is why proof-test intervals appear inside safety integrity calculations.
The automotive ladder is the one with the most interesting anchoring: rather than starting from a probability target, it derives the required level from three factors assessed per hazardous event: severity of harm, exposure (how much of the time the operational situation occurs), and controllability (whether a typical driver can act to avoid the harm). Combining the three through the standard's determination table lands the event on QM (quality management measures suffice) or on ASIL A through D; the classes are ordinal labels looked up in a table, never multiplied. It is a more honest reflection of how automotive risk actually works, where the same failure is trivial in a car park and lethal at motorway speed.
Two warnings about ladders. First, the levels do not translate between sectors: a SIL 3 argument is not a DAL B argument wearing different clothes, and cross-mapping tables between frameworks (which circulate widely) are conveniences, not equivalences. Second, and more fundamentally, a level is not a measurement of the delivered product. Assigning DAL A or SIL 3 to a function specifies how carefully it must be developed and evidenced; it does not by itself demonstrate that the resulting item is that good.
Random failures and systematic failures
Underneath every framework above lies the distinction that shapes modern safety practice more than any other:
| Random hardware failures | Systematic failures | |
|---|---|---|
| Origin | Physical degradation of components in service | Errors designed or built in: specification, design, software, manufacture, procedure |
| Behaviour | Occur at a rate; statistically predictable in populations | Present from the start; occur whenever the triggering conditions arise |
| Can you compute a probability? | Yes, from prediction and field data | No, in any defensible way |
| How you fight them | Redundancy, margin, monitoring, proof testing | Process rigour, independence, verification, simplicity, review |
This is why software has no failure rate. A program does not wear out; it does exactly what it was written to do, and its "failures" are latent design errors waiting for the input that reveals them. No amount of running time gives a meaningful statistical basis for claiming a software function will fail less than once per billion hours. The discipline's response was to stop trying: instead of predicting the probability of a systematic error, the frameworks regulate the rigour of the process in proportion to the consequence, which is precisely what a DAL, an ASIL or the systematic-capability half of a SIL is.
The defence standard practice takes the same route with a different mechanism, and it is a neat illustration of the logic. Instead of asking how likely a software function is to fail, it asks two answerable questions: how severe would the resulting mishap be, and how much control does the software actually have over the hazardous outcome (from fully autonomous authority, through arrangements where independent mechanisms or an operator can intervene, down to software that merely displays information or has no safety impact). Crossing those two answers produces a software criticality index, which in turn prescribes a level of rigour: which analyses and which depth of safety-specific testing must be performed. The standard is careful to note that this criticality index is not itself a risk assessment; it is a specification of the work needed to make a risk assessment credible.
Tolerability: how safe is safe enough
Every quantitative target eventually rests on a societal judgment, and the UK regulatory tradition articulated it most explicitly in a framework that has since been borrowed worldwide. Its legal root is the duty to reduce risk so far as is reasonably practicable, and the classic judicial interpretation established the shape of the test: the duty holder must weigh the risk against the sacrifice (in money, time or trouble) needed to avert it, and must act unless there is a gross disproportion between them, with the burden of proof on the duty holder. The disproportion is deliberately asymmetric: a mere balance of cost against benefit is not enough to justify inaction.
The regulator's tolerability framework turns that test into three regions:
- Unacceptable. Risk above this line is refused irrespective of benefits, and the published guidance frames the boundary in terms of individual risk of death per year, at an order of magnitude around one in a thousand per year for workers and one in ten thousand per year for members of the public.
- Tolerable, subject to reduction. The broad middle, where risk is accepted only in return for benefit and only if it has been driven down as far as reasonably practicable. This is where nearly all real engineering lives, and it is why safety arguments must document what was considered and rejected, not merely what was done.
- Broadly acceptable. Below an order of magnitude around one in a million per year, further reduction is not normally required, though good practice still applies.
Three consequences for practice. ALARP is an argument, not a calculation: the deliverable is a documented demonstration that credible further measures were identified and shown to be grossly disproportionate, and a bare assertion is worthless. Good practice sets a floor: where established practice exists (codes, standards, sector norms), meeting it is normally the minimum, and a cost argument does not excuse falling below it. And the units must survive the argument: individual risk per year, societal risk over a population, and per-hour or per-demand failure rates are different quantities, and sliding between them is the most common way a tolerability argument quietly breaks.
Exposure, units, and the arithmetic that connects them
The final piece of the mathematics is the least glamorous and the most frequently botched: converting between the rates the engineering produces and the risk the tolerability judgment consumes. A hazardous failure rate is not a risk until it has been through three multiplications:
risk of harm = failure rate × exposure to the hazardous state × probability the exposure produces harm
Each factor is its own analysis. The rate comes from the reliability chain. The exposure is how long the system spends in the state where the failure matters (an aircraft's approach phase, a level crossing's occupancy, the fraction of driving time at motorway speed, which is exactly the automotive exposure parameter). And the conditional probability of harm covers whether someone is present, whether they can escape, and whether protection or an operator intervenes: the controllability factor by another name. Skipping the middle and last terms turns every failure into a fatality and produces safety cases nobody believes; assuming them optimistically produces safety cases nobody should believe.
The dormant-function correction from the availability book belongs here too, because most protective functions are dormant: a protection channel with undetected dangerous failures arriving at rate λ and proof tested every T hours contributes an average unavailability of about λT/2 to the demand, and that unavailability, multiplied by the demand rate, is the frequency of the unprotected event. Test interval, detection coverage and redundancy therefore sit inside the safety arithmetic, not beside it, and the systems chapter takes up what happens when the channels are not as independent as the equation assumes.