RAMSynapse
Log inSign up

RAMS Core

PublishedIEC 60050-192 · MIL-HDBK-338B

Availability

The fraction of time a system is fit for use: where reliability and maintainability meet, how logistics turns inherent availability into operational availability, and why whole sectors are built on the number.

Ask a customer what they bought and the answer is rarely a failure rate or a repair time. They bought a system that is there when needed: a radar that is watching, a call line that answers, a plant that is generating. That property is availability, and it is the point where the two previous concepts in this row (reliability and maintainability) stop being separate disciplines and become one number. Reliability decides how often the system goes down; maintainability and the support system decide how long it stays down; availability is the balance the operator lives with.

Availability is also where engineering meets the enterprise. The inherent form of the number is set by design, but the operational form (the one missions and contracts actually feel) is set jointly by design, spares, crews, transport, and paperwork, which is why this topic spends as much time on logistics delay as on formulas. A programme that specifies availability without deciding who owns the logistics has specified nothing.

What availability is

The international dependability vocabulary (IEC 60050-192) defines availability as an ability: to be in a state to perform as required, under given conditions. The quantified reading is a fraction of time: over a stated operating period,

A = uptime / (uptime + downtime)

and the classic military definition (MIL-STD-721C) adds the operational nuance: a measure of the degree to which an item is in an operable and committable state at the start of a mission, when the mission is called for at a random time. That phrase separates the three RAMS time-properties cleanly:

PropertyThe questionThe moment it cares about
ReliabilityWill it keep working, once started?The mission, after it begins
MaintainabilityHow fast does it come back?The repair, after the failure
AvailabilityIs it up, right now?The random instant the demand arrives

Because it is a ratio of times, availability comes in flavours defined by what the clock counts: whether preventive maintenance counts as downtime, whether waiting for a spare counts, whether the measurement window is a mission, a year, or the steady-state limit. The foundations chapter builds those forms precisely; the one distinction to carry from the start is inherent versus operational: inherent availability (Ai) is the design's promise under ideal support, and operational availability (Ao) is what the fleet actually delivers once logistics and administration wrap around every repair.

Sectors built on the number

The nines ladder. Each added nine divides permitted downtime by ten, and each is bought differently: the first nines come from component reliability, the middle ones from maintainability and spares, the last ones from redundancy, failover, and the discipline to keep planned work invisible. The ladder is climbed with money and architecture, not wishes.
The nines ladder. Each added nine divides permitted downtime by ten, and each is bought differently: the first nines come from component reliability, the middle ones from maintainability and spares, the last ones from redundancy, failover, and the discipline to keep planned work invisible. The ladder is climbed with money and architecture, not wishes.

Every industry cares about availability; a few are constructed on it, and they are worth studying because they show the property in its purest forms.

  • The emergency call line. An emergency answering service has no product except availability: the service is the answered call, and downtime is measured in lives rather than revenue. Everything in such a centre's engineering follows from that: duplicated call-routing paths, geographically separate centres that overflow to one another, independent power, and the operational rule that even planned maintenance must be invisible to a caller. It is the cleanest example of availability as a system property: the phones, the network, the building, and the staffing roster are one availability chain, and the weakest link sets the number.
  • The surveillance radar. A gap in radar coverage is not lost output; it is blindness, and an adversary chooses when to exploit it. This is why defence programmes write operational availability into the contract as a key performance parameter: the number that matters includes the two days a spare transmitter spends in transit, not just the four hours of wrench time. Radar stations are the textbook case of the Ai-versus-Ao gap, and of buying the last nines with dual channels, on-site spares, and built-in test that catches the silent failure in the standby unit before it is needed.
  • The nuclear power plant. A nuclear unit is an economics machine with enormous fixed costs and near-zero marginal fuel cost, so every hour off the grid is almost pure loss; the industry tracks availability factors fleet-wide as a first-class indicator. At the same time, its safety systems embody the other face of availability: protective functions that must be available on demand while sitting dormant, kept honest by periodic proof testing. One site, two availability disciplines: production availability managed through compressed, critical-path planned outages, and on-demand availability of safety functions managed through test intervals and redundancy.

The three cases triangulate the topic: the call centre shows availability as chain-of-everything, the radar shows the logistics-dominated operational form, and the plant shows the split between continuous availability and on-demand availability. All three run through the rest of this book.

The nines and what they cost

Availability classes are quoted in "nines", and the arithmetic deserves to be internalised because it converts an abstract percentage into a budget measured in hours:

AvailabilityDowntime per year (8,766 h)What typically lives here
99%87.7 hoursWell-run single-string industrial equipment
99.9%8.8 hoursRedundant hardware or excellent logistics
99.99%53 minutesRedundancy plus fast, rehearsed failover
99.999%5.3 minutesArchitecture designed around continuity; no single repair on the critical path

Two readings matter. First, the ladder is logarithmic: each nine divides the downtime budget by ten while costing more than the last, because the mechanisms change. The early nines are bought with component quality and margins (reliability's levers); the middle nines with repair speed and spares (maintainability's levers); the final nines only with architecture, because once the budget is minutes per year, no repair of any speed fits inside it, and the system must survive failures without stopping. Second, at the top of the ladder the enemy is no longer the failure but the single long event: one botched changeover or one part stuck in customs can spend a decade of five-nines budget at once, which is why high-availability operations obsess over rehearsals and rollback plans as much as hardware.

Availability among R, A, M and S

Availability is the derived member of the quartet: it has no physics of its own, only the balance of its parents, A = MTBF / (MTBF + MTTR) in its simplest steady-state form. That derivation is its strength as a requirement: an availability target leaves the designer an honest trade space (fail rarely with slow repair, or fail more often with instant recovery), and the right mix is an economic choice, not a mathematical one. The allocation discipline exists partly to manage exactly this trade, flowing an availability commitment down into paired reliability and maintainability budgets.

The safety connection is subtler and runs through the hidden-function idea from the nuclear case: for a protective system, "availability" means probability of working when demanded, and its complement (the probability of failure on demand) is a core currency of functional safety. A dangerous failure that hides in a dormant channel is unavailability that no uptime dashboard shows, which is why testability and proof-test intervals appear throughout this book, and why the fault tree topic treats the unavailability of safety functions as a quantity of its own. The chapters that follow build the mathematics first, then the composition rules and the worked radar station, then the design and logistics machinery, then the toolkit map.


Want to see this on a live system model? Request a walkthrough.