skip to content

For disk time-to-failure, when is an exponential model wrong and a Weibull right?

level: seniorimportance: should knowfreq 34%

answer

  1. does age change the risk?
  2. hazard rate as a function of age
  3. flat hazard versus rising hazard
  4. shape above one means wear-out

basics

~20 s

An exponential lifetime assumes the failure rate never changes with age, so an old disk is as likely to fail next month as a new one. When wear-out makes failures rise with age, use a Weibull.

solid answer

~50 s

The deciding question is whether the hazard, the instantaneous failure rate given survival so far, depends on age. An exponential commits you to a constant hazard: age carries no information, and no age-based replacement policy can be justified under it. A Weibull adds a shape parameter `k` that makes the hazard a power of age: `k = 1` reduces exactly to the exponential, `k` greater than 1 gives a rising hazard, which is wear-out, and `k` below 1 gives a falling hazard, which is infant mortality. Disk fleets typically show a bathtub over full life, early failures, a flat middle, then wear-out, so a single exponential across the whole life understates late-life failures. Check it empirically by bucketing units by age and computing failures per drive-month among the units alive in each bucket; if that rate trends upward, the exponential is wrong. Handle the units still running as censored rather than discarding them.

go deeper

for a junior

Be ready to state that an exponential lifetime means the failure rate does not change with age, and that real hardware often fails more as it wears out.

for a middle

Explain the hazard rate as risk given survival so far, and map the Weibull shape parameter to flat, rising and falling hazards, noting that shape 1 is exactly the exponential.

for a senior

Demonstrate the empirical check on real fleet data with age buckets and exposure time, handle censored units correctly, and connect a rising hazard to an age-based replacement policy.

for a principal

Own the operational consequence: batch purchasing concentrates wear-out failures in time, so procurement staggering, spares inventory and replacement policy all follow from the hazard assumption you sign off on.

## Hazard is the quantity that decides the family For lifetime data, the most useful way to compare families is not the shape of the density but the **hazard rate**: the instantaneous rate of failure at age `t` *given that the unit has survived to `t`*. It answers the operational question directly, namely what is the risk for the drives I have running right now, given how old they are. Two families differ exactly in what they assume about that curve. **Exponential.** The hazard is a constant. The risk of failing in the next month is the same for a brand-new drive and one that has run for five years. Survival to date carries no information about remaining life. **Weibull.** The hazard is proportional to age raised to the power `k - 1`, where `k` is the shape parameter. Three regimes follow: - `k = 1`: the hazard is flat and the Weibull collapses exactly to the exponential. The exponential is a special case, not a rival. - `k > 1`: the hazard rises with age. This is wear-out, the physical picture of bearings, actuators and media degrading. - `k < 1`: the hazard falls with age. This is infant mortality, where manufacturing defects surface early and survivors are progressively more reliable. ## Why the choice changes what you do on Monday This is not a modelling nicety, it drives policy. Under a constant hazard, age-based proactive replacement is pointless: swapping a healthy four-year-old drive for a new one buys nothing, since the new one has the same forward risk. Under a rising hazard, proactive replacement is exactly the right lever, and the optimal replacement age is a real calculation. Under a falling hazard, the lever is burn-in: stress units before deployment so that defective ones fail in the factory rather than in the rack. A second consequence is correlation of failures in time. A fleet purchased in one batch and deployed together ages together. With a rising hazard, their failures bunch around a common age, which means a cluster of simultaneous replacements and a real availability risk. A constant-hazard model spreads the same total failures evenly through time and will not warn you about that cluster at all. ## Diagnosing the hazard shape from your own fleet You do not need a formal fit to answer the question. Bucket the drives by age, in six-month bands, say. For each band, compute the number of failures observed in that band divided by the total drive-months that units actually spent alive within that band. That ratio is an empirical hazard. Plot it against age: - Flat within noise: an exponential is adequate, and its simplicity is a virtue. - Trending upward: shape above 1, wear-out, and the exponential will understate later-life failures. - Falling then flat: infant mortality followed by a stable period. Bathtub shapes are common over full life. A pragmatic response is not to force one family across the whole span but to model the useful-life plateau with a constant hazard while treating the burn-in and wear-out ends separately, and to say out loud which age range the model is claimed to cover. ## Censoring: the mistake that quietly biases everything Most of the fleet has not failed yet. Those units are **censored**: you know their lifetime exceeds their current age, but not what it will be. The tempting and wrong move is to compute the average lifetime of drives that have already failed. That average is systematically too low, because the long-lived units are precisely the ones excluded from it, and it can be off by a large factor in a young fleet. Correct handling keeps every unit in the calculation, contributing its exposure time whether or not it has failed. The drive-months denominator described above does this naturally. A related trap: survivorship in the reporting pipeline. Drives that are decommissioned for capacity reasons, or removed during a rack migration, exit the population for reasons unrelated to failure. Treat those exits as censoring too, not as failures and not as silent deletions. ## Neighbouring families If the hazard rises and then falls, or the population is a mixture of a fragile and a robust batch, a single Weibull will not capture it either, and a mixture of two lifetime distributions is the honest description. If you have per-drive covariates such as temperature, workload or vendor, the question shifts from picking a single family to modelling how the hazard shifts with those covariates. Recognising when the one-family question has been outgrown is itself a senior signal. ## How to answer State that the hazard is the deciding quantity. Say what the exponential assumes and why it is a special case of the Weibull. Map the shape parameter to wear-out and infant mortality. Describe the age-bucketed empirical hazard as the check. Mention censoring, because a candidate who computes lifetimes only from failed units has made an expensive error that no amount of family selection can repair.

  • How would you check empirically whether the hazard is constant?
    Bucket units into age bands and, for each band, divide the failures observed in that band by the drive-months that surviving units actually spent inside it. Plot that empirical hazard against age. Flat within noise supports a constant hazard; a clear upward trend means wear-out and rules the exponential out for the later bands.
  • What does a Weibull shape parameter below 1 imply operationally?
    The hazard falls with age, which is infant mortality: defective units surface early and survivors get progressively more reliable. That justifies burn-in testing before deployment and a short, aggressive early-life replacement window, and it argues against replacing older healthy units, since they are the most reliable ones you own.
  • Why does a constant-hazard model understate the risk for a fleet bought in one batch?
    Because it spreads the same expected failures evenly across time. If the hazard actually rises with age, a batch deployed together reaches its wear-out region together, so failures arrive in a correlated cluster. The exponential predicts a steady trickle and gives no warning of the simultaneous replacements and capacity crunch.

A constant hazard is a lottery drawn fresh each month regardless of how long you have played; a rising hazard is a machine visibly wearing down.

saying these in an interview costs you the question

  • Assumes a constant failure rate without ever checking against age
  • Computes average lifetime using only the units that already failed
  • Thinks the Weibull and the exponential are unrelated families
  • Reads a rising hazard as a data-quality problem rather than wear-out
  • Ignores that a batch deployed together fails together under wear-out

context