RAMSynapse
Log inSign up

Reliability · Worked example

Medical devices

Infusion pump fleet

Industry overview: Medical devices at RAMSynapse

Six hundred volumetric infusion pumps in one hospital, each running about 2,500 hours a year. A pumping mechanism, an occlusion sensor, an air-in-line detector, a dose controller and a rechargeable battery. The fleet generates enough returns to fit distributions properly, and the first fit produced is wrong in an instructive way, which is what this page is about.

The technique, and why this one

Field-data fitting with the population separated by mechanism before any distribution is fitted. The hospital's biomedical department has years of work orders, so prediction is unnecessary. What is necessary is resisting the temptation to fit the fleet as a whole.

ItemModelParametersNote
Pumping mechanismWeibullβ = 1.6, η = 26,000 hplunger-drive wear
BatteryWeibullβ = 3.4, η = 17,500 hstrong wear-out, on a replacement cycle
Occlusion sensorexponentialλ = 150 per 10⁶ hno wear signature
Air-in-line detectorexponentialλ = 120 per 10⁶ hno wear signature
Dose controller boardexponentialλ = 90 per 10⁶ helectronics in useful life

Three items keep a constant rate and two do not, and the split is not a matter of taste. The sensors and the controller board are screened electronics in a temperature-controlled building well inside their useful life, the one condition under which the exponential model is a description rather than a convenience; the storage array controller page makes that choice for a whole system and defends it. The drive and the battery fail by accumulating damage, so their hazard is a function of age and an exponential fit discards the only variable a policy can act on. The Weibull items carry 1,100 of this pump's 1,460 per 10⁶ hours: three quarters of the failure rate sits on items whose hazard is moving, the reverse of the storage array, which is why the two pages end up opposite.

The mixed-population trap

One fleet, two answers. Fitting every return together gives a near-straight line at β ≈ 1.0 and the comfortable conclusion that failures are random. Separated by mechanism, the same data resolves into a mild wear-out at β = 1.6 and a steep one at β = 3.4, and the maintenance policy that follows is completely different.
One fleet, two answers. Fitting every return together gives a near-straight line at β ≈ 1.0 and the comfortable conclusion that failures are random. Separated by mechanism, the same data resolves into a mild wear-out at β = 1.6 and a steep one at β = 3.4, and the maintenance policy that follows is completely different.

Take every pump return from the fleet, rank them, and fit a single Weibull. The result comes out close to β = 1.0, and it looks clean: a straight line on the probability plot, a comfortable interpretation that pump failures arrive at random, and a clear conclusion that preventive maintenance beyond the statutory annual service would be wasted money.

That conclusion is an artefact. A mixture of several distributions with different shapes and different characteristic lives tends to flatten towards an apparently exponential aggregate, because the early failures of the steep population overlap the later failures of the shallow one and the combination washes out both curvatures. The straight line is not evidence of randomness; it is evidence that several mechanisms have been added together.

Separate the returns by the part that actually failed and the structure reappears:

  • The battery fits β = 3.4, η = 17,500 hours. That is aggressive ageing: charge cycling degrades capacity, and the hazard climbs as the cube and a bit of age.
  • The pumping mechanism fits β = 1.6, η = 26,000 hours. A mild wear-out on the plunger drive, real but far less urgent.
  • The sensors and the controller board show no curvature at all and keep their constant rates honestly.

Where the two curves put their failures

A shape is only worth fitting once percentiles are read off it, because a percentile is what a replacement interval is written in:

t_p = η(−ln(1 − p))^(1/β)

For the plunger drive at β = 1.6, η = 26,000 operating hours:

B10 = 26,000 × (0.10536)^(1/1.6) = 26,000 × 0.2450 = 6,370 h

B50 = 26,000 × (0.69315)^(1/1.6) = 26,000 × 0.7953 = 20,680 h

MTTF = η·Γ(1 + 1/β) = 26,000 × Γ(1.625) = 26,000 × 0.8966 = 23,300 h

At 2,500 hours a year: 2.5 years to the first tenth of the population, 8.3 to the median, 9.3 to the mean. The battery at β = 3.4 is far tighter, with B10 = 17,500 × 0.5158 = 9,026 hours and a mean of 17,500 × Γ(1.2941) = 15,720 hours falling within ten hours of its own median, both of them short of the 17,520-hour replacement date. A high β is a narrow distribution, which is what makes a scheduled replacement worth setting at all.

The hazard rate says it in the planner's units:

h(t) = (β/η)(t/η)^(β−1) so h(2t)/h(t) = 2^(β−1)

plunger drive 2^0.6 = 1.52 battery 2^2.4 = 5.28 any exponential item 2^0 = 1

The testability page evaluates the battery ratio directly, 36.9 rising to 195 per 10⁶ hours across the two-year cycle. The third line decides more policy than the other two: for the electronics the ratio is one at every age, so no life limit could reach them.

Why the battery hides so well

The battery's characteristic life is 17,500 hours. The hospital replaces batteries on a two-year cycle, which at 2,500 operating hours a year is 5,000 hours, well inside the steep part of the curve; expressed against calendar time on a 24-hour-a-day ward pump the cycle sits near 17,520 hours, essentially at η itself. Either way, the policy truncates the distribution before most of its mass arrives, so the observed battery failures are the left tail of a steep curve, and a left tail of a steep Weibull looks very much like a flat rate.

Both readings are worth numbers, because between them they bracket the policy:

F(5,000 operating h) = 1 − exp(−(5,000 / 17,500)^3.4) = 1 − exp(−0.0141) = 0.014

F(17,520 calendar h) = 1 − exp(−(17,520 / 17,500)^3.4) = 1 − exp(−1.0039) = 0.634

A pump used a few hours a day reaches its replacement date with fourteen batteries in a thousand gone; a pump running around the clock reaches it with nearly two thirds gone, the figure the safety page works from. Same battery, same schedule, a factor of 45 apart, decided by which clock the interval is written against.

This is the general trap worth carrying away: a wear-out item under an effective replacement policy masquerades as a constant-rate item in the aggregate data. The policy is doing the work, the data has been shaped by the policy, and an analyst who removes the policy on the strength of the data will rediscover the distribution the hard way.

Rolling the fleet up

For the fleet arithmetic the mechanisms are combined at their average rates: 300 per 10⁶ hours for the mechanism, 800 for the battery, and the constant items as listed, giving 1,460 per 10⁶ hours per pump. At 2,500 operating hours that is 3.65 failures per pump-year, and across 600 pumps roughly 2,190 failures a year arriving at one biomedical department.

Itemλ per 10⁶ hShare of pump rateFleet events a year
Battery80054.8%1,200
Pumping mechanism30020.5%450
Occlusion sensor15010.3%225
Air-in-line detector1208.2%180
Dose controller board906.2%135
Pump total1,460100%2,190

The last column is the rate column times 1.5, since 600 pumps at 2,500 hours is 1.5 × 10⁶ pump-hours a year. The three constant-rate rows are the only ones that behave as a textbook series calculation assumes, and over an operating year they multiply out term by term:

R_E(2,500) = e^(−150×10⁻⁶×2,500) × e^(−120×10⁻⁶×2,500) × e^(−90×10⁻⁶×2,500) = 0.687 × 0.741 × 0.799 = 0.407

so 0.9 electronic failures per pump-year, 540 across the fleet, and a 59% chance a given pump sees one. That block will not worsen with age and answers to no interval.

The aggregate is also measured very precisely, which is the last part of the trap. With 2,190 events a year the count is Poisson, so a 90% interval is 2,190 ± 1.645√2,190 = 2,190 ± 77 and the pump rate is pinned to

1,460 × (1 ± 0.035) = 1,409 to 1,511 per 10⁶ h

three and a half per cent either side, and it was that well determined before anyone separated the population. The aggregate fit was precise and wrong at the same time, and no confidence interval warns you, because the interval is computed on the quantity you chose to estimate and never on the choosing.

At six failures a working day the department behaves like a queue rather than a workshop, and the operational consequence of a reliability improvement becomes nonlinear: shaving 20% off the arrival rate shortens the backlog by considerably more than 20%, because the department was operating near its service capacity. That is the point at which a reliability number stops being a property of the device and becomes an input to an operations model, and the availability page picks it up there, with and without the loaner pool. The maintainability page enlarges the same 2,190, because 17% of returns have nothing wrong with them: 2,190 / 0.83 = 2,640 units reach the bench a year, about 450 of them healthy and consuming a full service time regardless.

What the number does not cover

The dominant hazard of this device is over-infusion, and the analysis on the safety page finds that device failures cause a minority of it: most over-infusion events begin with a correctly functioning pump programmed wrongly. A reliability improvement programme aimed at the 1,460 per 10⁶ hours would leave roughly four fifths of the harm untouched.

That is not a criticism of the reliability work; it is a statement about scope. The 1,460 is the right answer to the question "how often does this device fail", and the device's most important risk is not answered by that question. Recognising which of your numbers is authoritative for which decision is most of the skill.

What the analysis tells the engineer to do

Change the battery at B10 rather than at η. The comparison is 0.634 of the population failed at the scheduled 17,520 hours against 0.10 at 9,026, on an item worth 1,200 of the 2,190 fleet events a year. The battery itself does not improve; what changes is which side of the schedule its work lands on, and the availability page turns that into utilisation and ward-visible downtime.

Forecast the plunger drive against fleet age instead of holding it flat. B10 at 6,370 operating hours is two and a half years and the hazard rises 1.52 times per doubling of age. Replacing at B10 is not the answer, since that exchange is a workshop job; the flat arrival forecast is the thing that is wrong, and a fleet bought in one procurement will deliver its mechanism failures as a wave.

Spend nothing on the electronics. Those three items are 360 per 10⁶ hours, 24.7% of the pump, at a hazard ratio of one at every age. Halving all three removes 270 fleet events from 2,190 and changes no policy, because there is no interval to set against a flat hazard. Their leverage is detection, which the testability page works.

Require the work order to name the part. Every decision above rests on one field distinguishing five items; without it the only fit is the fleet fit, the only answer β ≈ 1.0, and the finding that nothing needs a schedule.

What a different technique would have given

Accepting the aggregate β ≈ 1.0 would have produced three wrong decisions in sequence: that preventive battery replacement is unnecessary (it is the single most effective maintenance action on this fleet), that the plunger drive needs no life limit (it has a real, if mild, ageing mechanism), and that the failure arrival rate will stay flat as the fleet ages (it will not, because two of the five contributors are climbing). The fit was not incorrect arithmetic. It was correct arithmetic on the wrong population, which is harder to catch and does more damage.


Want to see this on a live system model? Request a walkthrough.