RAMSynapse
Log inSign up

Testability · Worked example

Energy and resources

Gas compressor train

Industry overview: Energy and resources at RAMSynapse

A gas turbine driving a centrifugal compressor, a lube oil system with a duty and a standby pump, three vibration sensors voted two out of three, and an emergency shutdown function that closes a valve on demand. The machine runs continuously for four years between turnarounds. Two entirely separate test regimes live inside it, and they share no equipment, no interval, no budget and, in most organisations, no owner.

The first regime is continuous condition monitoring, and it protects production. It watches a machine that is running, using signals the machine generates by running, and its currency is unplanned outage hours avoided. The second is periodic proof testing, and it protects people. It exercises a valve that spends its entire life doing nothing, and its currency is probability of failure on demand. The interesting engineering happens when the second regime runs into a constraint the first never meets: the only honest test of a shutdown valve is to shut the plant down.

The technique, and why this one

Partial stroke testing analysed as a coverage-raising device rather than as an interval change, sitting alongside rate-weighted condition-monitoring coverage for the running machine. The distinction is the point of the page. Faced with a dangerous undetected failure rate that fails its safety integrity target, there are only two mathematical levers: test more often, or make each test catch more. The valve forbids the first, so the industry engineered the second, and the arithmetic that justifies it is a coverage argument, not a scheduling one.

Itemλ per 10⁶ hModel usedTest means credited
Turbine hot section350Weibull, β = 3.1, η = 52,000 h; strong wear-outvibration spectra, exhaust gas temperature spread, performance trending
Lube oil pump (each, 2, one running)500constant λ; repairable duty and standby pairdischarge pressure, filter differential, oil temperature and debris
Vibration sensor (each of 3)200constant λ2oo3 cross-comparison: the channels diagnose each other
ESD shutdown valve40, of which 12 dangerous undetectedconstant λ, dormantpartial stroke quarterly, full stroke at proof test
Trip logic solver30constant λcontinuous internal diagnostics and watchdog

Constant rates are right for the instrumentation and wrong for the hot section, which is why the hot section carries a Weibull. The turbine's β = 3.1 is not a statistical decoration: it is the reason condition monitoring works on it at all, developed below.

The production side, where coverage changes the category of the work

The train's production-affecting failure rate is 1,180 per 10⁶ hours, and vibration and lube monitoring detect 88% of it:

FFD = Σλ(detected) / Σλ(all) = 1,038.4 / 1,180 = 0.880

leaving an undetected rate of 141.6 per 10⁶ hours. Over a continuous year that is

unannounced production-affecting failures per year = 141.6 × 10⁻⁶ × 8,760 = 1.24

against 10.34 in total. What makes the compressor unlike every other row in this chapter is what the 88% buys. Elsewhere detection shortens a repair. Here it changes the category of the work: a hot section whose degradation is trended is repaired in a planned window, while the same hot section discovered by failure imposes the 340-hour unplanned outage that the maintainability page prices. Coverage is the mechanism by which reliability-centred maintenance converts calendar-based intervention into condition-based intervention, and the 12% that monitoring misses is exactly the population that appears as an unwelcome discovery when the casing comes off at turnaround.

The hot section is monitorable because it wears. Its hazard rate

h(t) = (β/η)·(t/η)^(β−1)

evaluated at half the characteristic life and at the characteristic life gives

h(26,000) = (3.1/52,000) × (26,000/52,000)^2.1 = 5.96 × 10⁻⁵ × 0.233 = 1.39 × 10⁻⁵ per hour

h(52,000) = (3.1/52,000) × (52,000/52,000)^2.1 = 5.96 × 10⁻⁵ per hour

13.9 per 10⁶ hours rising to 59.6, a factor of 4.3 across the second half of the machine's life, and rising steeply enough that the physical mechanisms behind it (blade creep, coating loss, tip clearance growth) produce measurable drift for thousands of hours before anything breaks. That drift is what the monitor reads. An item with β close to 1 offers a monitor nothing to trend, which is why the same investment in instrumentation pays very differently on the turbine and on the logic solver.

A third test structure is hiding in the instrument set and costs nothing extra. The 2oo3 vibration arrangement exists for voting, but voting is also diagnosis: three channels that should agree provide continuous cross-comparison, so a sensor drifting away from its peers is detected by its peers with no dedicated monitor at all. Redundancy used as a diagnostic mechanism is among the cheapest coverage available anywhere a design already carries parallel channels for other reasons.

The valve, where the entire safety problem lives

Of the shutdown valve's 40 per 10⁶ hours, 12 are dangerous and undetected: 30% of the item's failure rate sits in the one category that neither announces itself nor fails to the safe side. Proof tested annually, the valve alone gives

PFD = λ_DU × T / 2 = 12 × 10⁻⁶ × 8,760 / 2 = 5.3 × 10⁻²

The SIL 2 band for a low-demand function runs from 10⁻² to 10⁻³. At 5.3 × 10⁻² the valve fails its target by a factor of five before any other element of the loop has contributed anything, and the safety page has nowhere to put the shortfall. The obvious response, testing more often, is unavailable: a full stroke closes the valve, and closing the valve stops the train.

Raising the coverage instead of shortening the interval

Partial stroke testing converts most of the dangerous undetected fraction into a detected one without ever closing the valve. It cannot prove the valve will seat, so it supplements the proof test rather than replacing it.
Partial stroke testing converts most of the dangerous undetected fraction into a detected one without ever closing the valve. It cannot prove the valve will seat, so it supplements the proof test rather than replacing it.

Partial stroke testing moves the valve through a fraction of its travel, far enough to prove that the actuator develops force and the stem is free, not far enough to interrupt flow. It runs every three months without touching production. Its effect is not on T but on λ_DU itself, because the mechanisms it exercises (stem seizure, packing friction, actuator spring set, solenoid sticking) are the ones that dominate the dangerous undetected population. Crediting it with 88% coverage of that population:

λ_DU(residual) = λ_DU × (1 − C_PST) = 12 × (1 − 0.88) = 1.44 per 10⁶ h

and the annual proof test now works on the residual only:

PFD = 1.44 × 10⁻⁶ × 8,760 / 2 = 6.3 × 10⁻³

The sensors and the logic solver, both continuously diagnosed and voted, add the small remainder that brings the loop to the 6.4 × 10⁻³ the safety case quotes: a factor of 8.2 improvement, comfortably inside SIL 2, achieved without a single extra shutdown. The coverage is 88% and not 100% for an honest reason worth stating in the safety case: a partial stroke does not prove final seat closure or tightness, so leakage-class failures survive the partial test and wait for the full one. Claiming 100% for a partial stroke is the standard way this argument is abused.

What the analysis tells you to do

The valve's coverage credit has to be earned and maintained, not assumed. That means the partial stroke must be instrumented well enough to measure something (breakaway torque, travel time, position achieved) rather than merely to command something, because an uninstrumented partial stroke proves only that a solenoid energised. It means the claimed 88% belongs in the safety requirements specification with the mechanisms it covers listed, so that a later change of valve or actuator triggers a review. On the production side, the priority is the 141.6 per 10⁶ hours the monitoring misses, and the useful question is which of it is unmonitorable in principle and which is merely uninstrumented: lube system failures that show first as debris are already covered, while auxiliary and control-air failures typically are not and are cheap to add. Finally, the 2oo3 cross-comparison should be exploited deliberately rather than incidentally, with the disagreement threshold set and alarmed rather than left as a side effect of the voting logic.

What a different technique would have given

The alternative is arithmetically obvious and operationally impossible: keep full-stroke testing as the only test and shorten the interval until the target is met. Inverting the same formula for the required interval,

T = 2 × PFD_target / λ_DU = 2 × 6.4 × 10⁻³ / (12 × 10⁻⁶) = 1,067 h

which is six weeks, or

8,760 / 1,067 = 8.2 full shutdowns per year

on a train the operator plans to stop once every four years. No production organisation will accept it, and a safety case built on an interval that the plant will quietly not honour is worse than one that is honestly short of target, because the paper says the risk is controlled while the machine says otherwise. Where the test conflicts with the function, the answer is a different test rather than a shorter interval, and the move that partial stroke embodies, finding a partial exercise that reaches the dominant dangerous mechanisms, generalises well beyond valves.


Want to see this on a live system model? Request a walkthrough.