Functional Hazard Analysis · Chapter 2

Theoretical Foundations

Definitions, units, models, and the assumptions that bound them.

The assessment has four moving parts: what counts as a function, what counts as a failure of it, how the effect is described, and how the description becomes a class. The fourth is where the money is, and it is decided entirely by the third.

A function, not a box

The unit of analysis is what the aircraft does, stated so that it survives any implementation choice. Provide attitude information to the flight crew is a function. Attitude and heading reference unit is not: it is one of the things that performs the function, alongside the displays, the standby instrument and the crew's own scan. An assessment built around the boxes misses both the combinations across boxes and the conditions that appear when the allocation changes.

Functions decompose. The aircraft-level tree runs from a handful of top functions down two or three levels, and the assessment attaches to the level where a failure has a describable effect on the aircraft. Too high and the effects are unclassifiable; too low and the analysis has quietly become an FMEA of an architecture nobody has agreed to yet.

The ways a function fails

One function, five failure modes, six phases and two annunciation states. The candidate list is mechanical; the analysis is deciding which candidates are physically distinct and which collapse into one condition.
One function, five failure modes, six phases and two annunciation states. The candidate list is mechanical; the analysis is deciding which candidates are physically distinct and which collapse into one condition.
ModeThe questionWhy it is on the list
LossThe function is not there at allThe obvious one, and usually not the worst
Partial lossSome of it, degradedOften a different class from total loss, in both directions
Misleading or erroneousThe function is performed wrongly, and plausiblyAlmost always the worst row on the page
InadvertentThe function is performed when it was not commandedThe row an assessment built around loss will miss entirely
Out of timeLate, early, or out of sequenceMatters wherever timing is part of the function

Two crossings turn this list into failure conditions. Phase, because the same functional failure is trivial in the cruise and catastrophic on the approach: a condition can appear twice at two classifications, and the worksheet is expected to carry both rows. And annunciation, because what the crew can do about a failure depends on whether they know about it, and a loss the crew is told about is routinely one class less severe than the same loss they discover for themselves.

Misleading and unannunciated are the two words that drive programme cost. Between them they generate the conditions that need independence, monitoring and the highest assurance levels, and an assessment that lists only losses will produce an architecture that is comfortable and wrong.

More than one function at a time

The grid is single-function by construction, and the aircraft level needs conditions the grid cannot generate. Two functions that are each Major on their own can be Hazardous or Catastrophic together, because the second one was the answer to losing the first: attitude and airspeed, thrust and wheel braking, navigation and communication. The question that finds these rows is not what else might fail, it is what the crew was going to use instead.

Two things keep the list finite. A combination is worth carrying only where the members can plausibly fail in the same flight, whether from one cause or from two that happen to coincide, and a combination whose members are already required to be independent is a claim for the common cause analysis to test rather than a fresh row on the worksheet. The system-level assessment then arrives at the same conditions from the other direction, because allocating two functions to one item creates the combination in hardware.

Describing the effect before classifying it

The effect column is written in terms of what happens to the aircraft, the crew and the occupants, in language a pilot would recognise. Not "loss of the function". Not "the system fails". What the aeroplane does, what the crew is left to do about it, and where that ends.

This ordering is the discipline of the whole method: the effect is the evidence, and the classification is the conclusion drawn from it. An assessment that writes the class first and the effect afterwards produces a document that cannot be reviewed, because there is nothing in it to disagree with. Where the effect depends on a crew action, the action is stated as an assumption, and it becomes a requirement on procedures and training that somebody has to own.

The ladder

The five classes with the average probability bands the airworthiness guidance attaches to them. The bands are orders of magnitude rather than hard edges: the guidance treats a factor of two as on the order of for the remote band, and a factor of three for the two below it.
The five classes with the average probability bands the airworthiness guidance attaches to them. The bands are orders of magnitude rather than hard edges: the guidance treats a factor of two as on the order of for the remote band, and a factor of three for the two below it.
ClassEffect, in the guidance's own termsAverage probability per flight hour
CatastrophicLoss of the aeroplane, or multiple fatalitiesExtremely improbable: ≤ 10⁻⁹
HazardousLarge reduction in safety margins, serious or fatal injury to a small numberExtremely remote: 10⁻⁷ to 10⁻⁹
MajorSignificant reduction in margins, higher workload, discomfort or minor injuriesRemote: 10⁻⁵ to 10⁻⁷
MinorSlight reduction in margins, routine crew actionProbable: 10⁻³ to 10⁻⁵
No safety effectNo effect on safety or operational capabilityNo requirement

Three notes on using it honestly. The bands are guidance, not a rule: an advisory circular offers an acceptable means of compliance, and the classification argument still has to be made. Minor carries no obligation to quantify at all. And a calculated probability comfortably below its band is compliant, not suspicious.

Where 10⁻⁹ came from

It is an allocation, and knowing its derivation is the fastest cure for treating it as a law of nature. The airworthiness guidance sets it out: a severe accident rate of about one per million flight hours was taken as the starting point in the 1960s, roughly a tenth of those accidents were attributed to systems design, giving about 10⁻⁷ per flight hour for all systems-related catastrophic conditions together, and that budget was divided among an assumed hundred or so catastrophic failure conditions on an aeroplane. One in a billion flight hours is what falls out of that division.

Two consequences follow, and both are in the guidance. The number was never intended to apply to a single failure condition on its own, and these values are not accident-rate goals; they are apportionment devices for judging one condition at a time.

The rule that is not a probability

Alongside the number sits a qualitative requirement that no arithmetic can buy relief from: a catastrophic failure condition must not result from a single failure. The rule is absolute in the sense that matters, because a single failure has to be assumed regardless of how improbable it is, and a purely quantitative demonstration that a condition is extremely improbable is generally not sufficient on its own.

The definition of "single failure" is the part that does real work: it includes any set of failures that cannot be shown to be independent of each other, which is why common cause analysis is not an optional refinement. Two channels that share a power supply, a cooling path, a maintenance action or a calibration constant are one failure wearing two names.

At-risk time

Twenty seconds of a two-hour flight. The per-flight probability budget does not change; what changes is the failure rate an item may be allowed, because its failure only matters inside that window.
Twenty seconds of a two-hour flight. The per-flight probability budget does not change; what changes is the failure rate an item may be allowed, because its failure only matters inside that window.

The objective is stated as an average probability per flight hour, defined over a flight representing the average at-risk time of the fleet. Where a condition is dangerous only in one phase, the guidance is explicit that the criterion should be applied per flight or per flight cycle rather than diluted across a flight of mean duration.

Working the conversion in the useful direction: a catastrophic objective of 10⁻⁹ per flight hour on a two-hour flight is 2 × 10⁻⁹ per flight, and if the condition is only catastrophic for twenty seconds, an item whose failure causes it during that window may carry a rate of

2 × 10⁻⁹ ÷ (20 ⁄ 3600) h = 3.6 × 10⁻⁷ per hour

That is 360 times looser than demanding 10⁻⁹ per hour of the item itself, and it is legitimate: the item is only at risk for twenty seconds per flight. The misuse runs in the opposite direction, averaging a phase-limited condition over the whole flight, which is exactly what the guidance tells you not to do.

Development assurance, and why it is a separate column

The classification sets an assurance level as well as a probability objective. Where the functional failure set has more than one member and the members' development can be shown to be independent, the level can be shared out between them.
The classification sets an assurance level as well as a probability objective. Where the functional failure set has more than one member and the members' development can be shown to be independent, the level can be shared out between them.

A probability objective governs random hardware failure. It says nothing about a mistake in the requirements, because a design error has no failure rate: it is present from the first unit, and no amount of operating time makes it less likely. The discipline's answer is to regulate the rigour of the development process in proportion to the consequence, which is what a development assurance level is.

The mechanism runs through the functional failure set: the set of function failures that together produce one failure condition. If a single function failing is enough, that function carries the level for the classification directly. If the condition needs a combination and the members' development can be shown to be independent, the level can be shared: a catastrophic condition can be met by one member at A and another at C, or by two independent members both at B. The classes below work the same way one step down, until the ladder runs out at Minor: that set can hold one member at D and let its partner sit as low as E, and there is no pair below that to offer as a second option. Higher levels are always acceptable; the pairs are minima.

Two things a level is not. It is not a probability budget, and no assurance level implies a failure rate. And it is not a property of the delivered item: it describes how carefully the item was developed and evidenced, which is a different claim from how well it works.


Want to see this on a live system model? Request a walkthrough.