A gas turbine driver, a centrifugal compressor, a lube oil system with a duty and a standby pump, three vibration sensors voted two out of three, a trip logic solver and a shutdown valve. The machine runs continuously for four years between turnarounds and the protective function at the end of it moves perhaps never. The safety question on this train is not whether the compressor fails; it is whether the valve that stops it will close when it is finally asked.
That question cannot be answered with a failure rate, because the valve is not operating. It sits, it degrades while it sits, and nothing reveals the degradation until someone tests it or a real demand arrives. The quantity the process sector uses instead is the average probability of failure on demand, and its interesting property is that it is set as much by the test schedule as by the hardware. This train's shutdown function fails its target on a calendar rather than on a fault, and recovers it the same way.
The technique, and why this one
Average probability of failure on demand for a low-demand safety instrumented function, with the proof-test interval treated as the governing design variable. A dormant item tested at interval T is on average halfway through its exposure window when a demand arrives, which gives PFD ≈ λ_DU × T / 2. Everything below follows from that expression being linear in both T and the undetected dangerous rate, and from what the site can and cannot do to either.
| Element | Model | Value | Role |
|---|---|---|---|
| ESD shutdown valve | dormant, PFD = λ_DU·T/2 | λ total 40, λ_DU 12 per 10⁶ h | final element, the constraint |
| Trip logic solver | exponential | 30 per 10⁶ h | not the constraint |
| Vibration sensor, 3 fitted | exponential, 2oo3 vote | 200 per 10⁶ h each | initiator |
| Proof test, full stroke | interval | 12 months, 8,760 h | reveals the valve |
| Partial stroke test | credited as automatic diagnostic | 3 months | reduces the undetected fraction |
| Turbine hot section | Weibull, β = 3.1, η = 52,000 h | 350 per 10⁶ h book value | source of demands |
| Target integrity | SIL 2 | PFD between 10⁻³ and 10⁻² in this sector's scheme | risk reduction 100 to 1,000 |
Two lines deserve comment before the arithmetic. The valve's total rate is 40 per 10⁶ hours and only 12 of it is dangerous and undetected; the rest fails safe, announces itself, or is found by ordinary maintenance. Quoting 40 would overstate the answer by 3.3 and would be wrong in principle, because a revealed failure gets repaired and a safe failure trips the train. And the SIL bands belong to the process sector's scheme; they do not transfer as written to the aircraft, the vehicle or the crossing elsewhere in this chapter.
The valve fails its target on the calendar
With an annual full stroke test:
PFD_valve = λ_DU × T / 2 = 12 × 10⁻⁶ × 8,760 / 2 = 12 × 10⁻⁶ × 4,380 = 5.3 × 10⁻²
The SIL 2 band tops out at 10⁻², so the final element misses its target by a factor of 5.3 before the sensors, the logic solver, the wiring, the actuator air supply or the common-cause contribution have been considered at all. Expressed as risk reduction, 1/5.3 × 10⁻² = 19, against the 100 the band's lower edge requires.
Ask the obvious question. What test interval would fix it?
T = 2 × 10⁻² / (12 × 10⁻⁶) = 1,667 h
about ten weeks, and only to reach the edge of the band; a defensible mid-band 3 × 10⁻³ would need 500 hours, roughly three weeks. A full stroke means closing the valve, which means stopping the train, and a plant that stops its compressor seventeen times a year to prove its trip has solved the safety problem by abolishing the production. The test interval is the right lever and the plant cannot pull it, which is the entire reason partial stroke testing exists.
Partial stroke: the same lever, applied differently
A partial stroke moves the valve a few per cent of its travel while the train keeps running, revealing the modes that stop a valve moving at all: a seized stem, a lost air supply, a failed solenoid, a packing gland gone solid. It does not reveal everything, so it is credited as an automatic diagnostic with a coverage rather than as a shorter interval, converting part of the undetected dangerous rate into a detected one.
The quoted loop figure with quarterly partial stroke is 6.4 × 10⁻³. Back out the coverage it implies, treating the valve as dominating the loop:
(1 − C) = 6.4 × 10⁻³ / 5.3 × 10⁻² = 0.122, so C = 88 per cent
residual λ_DU = 0.122 × 12 = 1.46 per 10⁶ h
PFD_valve = 1.46 × 10⁻⁶ × 4,380 = 6.4 × 10⁻³
Risk reduction 1/6.4 × 10⁻³ = 156, comfortably inside the SIL 2 band, achieved by a factor of eight improvement without changing a single piece of hardware. The integrity level is a property of the loop plus its test regime. It is never a property of the valve, and a certificate that describes the valve alone is answering a different question.
Two operational consequences that are easy to lose
The first is that the SIL 2 claim now rests on two schedules at once, and is linear in both. Skip the quarterly partial stroke for a year and λ_DU reverts from 1.46 to 12, taking the loop back to 5.3 × 10⁻². Stretch the full stroke from twelve months to twenty-four to fit the turnaround cycle and the PFD doubles to 1.3 × 10⁻², out of band again. Neither event produces a fault, an alarm or a finding. An engineering assumption has become an operating requirement, and it must be written down as one, tracked as one, and raised through FRACAS when a test is missed or a stroke comes back slow.
The second is the vote, which encodes a trade rather than an improvement. Three vibration sensors at 200 per 10⁶ hours each, arranged so that any one could trip, would produce
3 × 200 × 10⁻⁶ × 8,760 = 5.3 spurious trips a year
on instrumentation alone, on a machine whose unplanned stop is enormously expensive. The 2oo3 vote removes essentially all of that, and the price is that one dangerous sensor failure consumes its margin. The degraded behaviour is then a configuration choice: vote the failed channel out and the arrangement becomes 1oo2, more likely to trip spuriously and less likely to fail to trip; leave it in and it becomes 2oo2, the reverse. A setting in a logic solver now carries part of the integrity argument.
The demand rate is half the answer, and it is not constant
A PFD on its own is not a hazard rate. The hazardous event rate is the demand rate multiplied by the probability the function fails on demand, and this machine's demand rate rises with age, because the hot section is a strong wear-out item:
h(t) = (β/η)(t/η)^(β−1) = (3.1 / 52,000) × (t / 52,000)^2.1
At the age-replacement optimum of 28,600 hours that is 5.96 × 10⁻⁵ × 0.2849 = 1.70 × 10⁻⁵ per hour, 17.0 per 10⁶ hours. Run the section to its characteristic life of 52,000 hours, where the ratio is 1, and the hazard is 5.96 × 10⁻⁵ per hour, 59.6 per 10⁶ hours, three and a half times higher. Insofar as hot-section events demand the trip, the hazardous event rate moves with them:
at 28,600 h: 17.0 × 10⁻⁶ × 6.4 × 10⁻³ = 1.09 × 10⁻⁷ per hour, 9.5 × 10⁻⁴ per year
at 52,000 h: 59.6 × 10⁻⁶ × 6.4 × 10⁻³ = 3.81 × 10⁻⁷ per hour, 3.3 × 10⁻³ per year
Deferring the hot-section replacement multiplies the hazardous event rate by 3.5 without anyone touching the safety instrumented function, and the deferral is normally argued on production grounds by people who have never seen the PFD. The decision behind the 28,600 hours is worked on the reliability page.
What the analysis tells the engineer to do
Put the twelve-month full stroke and the three-month partial stroke into the operating procedure as safety requirements with named owners, and treat a missed test as a hazard-log entry rather than a maintenance backlog item.
Spend on the final element, not on the sensors. The logic solver's 30 per 10⁶ hours is a larger number than the valve's 12 and contributes far less, because it is continuously operating and diagnosed while the valve is dormant. A loop analysis that follows the biggest λ spends its effort in the wrong place.
Bring the hot-section replacement age into the safety conversation: it is currently a cost optimisation and it is also a demand-rate control.
Verify the 88 per cent partial-stroke coverage against the valve's actual failure modes rather than accepting a supplier figure, because the SIL 2 claim is one subtraction away from it. Where such figures come from is the subject of reliability prediction and of the testability column.
What a different technique would have given
Treat the valve like the rest of the train and the answer changes beyond recognition. An availability model takes the valve's 40 per 10⁶ hours, pairs it with a repair time near the lube system's six hours, and reports an unavailability of
40 × 10⁻⁶ × 6 = 2.4 × 10⁻⁴
219 times better than the untested valve's true 5.3 × 10⁻² and 27 times better than the properly tested loop. It is arithmetically impeccable and it describes a valve that is repaired promptly when it breaks, which is not this valve. The same model quotes an MTBF of 1/40 × 10⁻⁶ = 25,000 hours, about three years, a figure that sounds reassuring and is irrelevant, because the valve is not accumulating operating time. It is sitting.
A layer of protection analysis fails more gently. It would identify the initiating event, credit the trip and the relief and containment layers around it, and arrive at a required risk reduction, which is a target rather than a demonstration. It would tell the site that a SIL 2 function is needed. It would not tell them that the one they have sits at 5.3 × 10⁻² until somebody works the interval, and the interval is the only part of this system that turns out to be adjustable.