Reliability: failure rate, bathtub curve and MTBF
Reliability and the lifetime functions f, F, R and hazard rate, the bathtub curve and its causes and remedies, the exponential and Weibull models, MTTF, MTBF, MTTR and availability, and estimating hazard rates from test data.
Drafted with Aria, reviewed by the AiCanCode.org team. Spotted an error? Use Give Feedback at the bottom of the page.
Why it matters
Quality asks whether a product is right when it leaves the factory; reliability asks whether it keeps working afterwards. Warranty budgets, spare-parts stocking, maintenance intervals and the design of safety systems all rest on failure-rate data. The bathtub curve, the exponential model and MTBF are the vocabulary every maintenance, quality and design engineer uses.
Key ideas
Reliability. The probability that an item performs its required function, under stated conditions, for a stated period of time. All four parts matter: a pump's reliability for 1000 h at rated head is not its reliability for 5000 h at overload.
The lifetime functions. Let T be the time to failure of an item.
- Failure density f(t): the probability per unit time of failing at t (a probability density).
- Unreliability F(t) = P(T ≤ t), the fraction failed by time t.
- Reliability R(t) = 1 − F(t), the fraction still working at t.
- Hazard rate (instantaneous failure rate) h(t) = f(t)/R(t): the probability per unit time that an item that has survived to t fails in the next instant. It is a conditional rate, which is why it is the right quantity to plot over life.
The bathtub curve. For many populations h(t) plotted against age has three regions:
- Infant mortality (early failure, burn-in): h decreasing. Causes: manufacturing defects, weak components, assembly and installation errors. Remedies: burn-in or run-in testing, better process control and inspection, supplier quality.
- Useful life (random or chance failures): h roughly constant. Causes: random overloads and stress peaks. Preventive replacement does not help here, because an old item is as good as a new one; derating and redundancy do.
- Wear-out: h increasing. Causes: fatigue, wear, corrosion, ageing. Remedy: preventive replacement or overhaul before wear-out begins. Electronic components tend to have long flat regions; mechanical parts often show little flat region and a gradual wear-out.
Constant failure rate: the exponential model. If h(t) = λ is constant, R(t) = e^(−λt). The mean time to failure is 1/λ. The model is memoryless: the probability of surviving the next 100 h is the same for a new item and for one that has already run 5000 h. Only 36.8 % of items survive to the mean life (R = e⁻¹), so MTBF is not "the time by which most items will still be working".
Weibull model. h(t) = (β/η)(t/η)^(β−1). The shape parameter β describes the bathtub region: β < 1 decreasing hazard (infant mortality), β = 1 constant (exponential), β > 1 increasing (wear-out). η is the characteristic life, at which 63.2 % have failed. Parameters are fitted to test or field data.
MTTF, MTBF, MTTR, availability.
- MTTF (mean time to failure) is used for non-repairable items: MTTF = ∫₀^∞ R(t) dt.
- MTBF (mean time between failures) is used for repairable systems: total operating (up) time divided by number of failures. In the constant-λ region, MTBF = 1/λ.
- MTTR (mean time to repair) measures maintainability.
- Inherent availability A = MTBF / (MTBF + MTTR): the long-run fraction of time the system is up.
Estimating the failure rate from data. From a test of N₀ items with failures recorded in time intervals, the hazard in an interval is the number failing in it divided by the number surviving at its start and the interval length. From field data on repairable equipment with constant λ, λ = number of failures / total operating time (all units added together).
Formulas
R(t) = 1 − F(t), h(t) = f(t) / R(t)
- R, F dimensionless; f and h in failures per hour (h⁻¹).
R(t) = exp[−∫₀ᵗ h(t) dt]
- General link between hazard and reliability.
R(t) = e^(−λt), F(t) = 1 − e^(−λt), f(t) = λ·e^(−λt)
- Constant failure rate λ (h⁻¹), time t (h).
MTTF = ∫₀^∞ R(t) dt = 1/λ (exponential)
λ̂ = r / ΣTᵢ and MTBF = ΣTᵢ / r
- r = number of failures; ΣTᵢ = total operating time of all units (h).
h(interval) = n_f / (N_s · Δt)
- n_f = failures in the interval; N_s = survivors at the start of the interval; Δt = interval length (h).
A = MTBF / (MTBF + MTTR)
- Inherent availability (dimensionless).
R(t) = exp[−(t/η)^β]
- Weibull reliability; η = characteristic life (h); β = shape parameter (dimensionless).
Worked examples
Example 1 (standard): exponential reliability Given: a control valve in its useful-life period has λ = 2 × 10⁻⁴ h⁻¹. Find (a) MTBF, (b) reliability for a 1000 h mission, (c) the longest mission with reliability at least 0.95, (d) the probability of surviving to the MTBF.
- (a)
MTBF = 1/λ = 1 / (2 × 10⁻⁴) = 5000 h. - (b)
R(1000) = e^(−2 × 10⁻⁴ × 1000) = e^(−0.2) = 0.8187. - (c)
e^(−λt) = 0.95givest = −ln(0.95)/λ = 0.05129 / (2 × 10⁻⁴) = 256.5 h. - (d)
R(5000) = e^(−1) = 0.368.
MTBF = 5000 h; R(1000 h) = 0.819; maximum mission for 95 % reliability ≈ 256 h; only 36.8 % survive to the MTBF.
Example 2 (GATE level): hazard rate from test data Given: 1000 identical components are put on test. Failures: 60 in 0–100 h, 40 in 100–200 h, 30 in 200–300 h. Find the hazard rate in each interval, the reliability at 300 h, and the region of the bathtub curve.
- Interval 1: survivors at start = 1000.
h₁ = 60 / (1000 × 100) = 6.0 × 10⁻⁴ h⁻¹. - Interval 2: survivors = 1000 − 60 = 940.
h₂ = 40 / (940 × 100) = 4.26 × 10⁻⁴ h⁻¹. - Interval 3: survivors = 940 − 40 = 900.
h₃ = 30 / (900 × 100) = 3.33 × 10⁻⁴ h⁻¹. - Survivors at 300 h = 900 − 30 = 870, so
R(300) = 870 / 1000 = 0.87. - The hazard rate falls with age.
h = 6.0, 4.26 and 3.33 × 10⁻⁴ h⁻¹; R(300 h) = 0.87; decreasing hazard means infant mortality, so a burn-in period is worth considering.
Common mistakes
- Dividing failures in an interval by the original population instead of the survivors at the start of the interval.
- Treating MTBF as a guaranteed life; with constant λ, 63 % of items fail before the MTBF.
- Applying R = e^(−λt) in the infant-mortality or wear-out regions, where λ is not constant.
- Expecting preventive replacement to help in the constant-λ region; it only helps for wear-out.
- Mixing hours and years, or failures per 10⁶ h with failures per hour.
- Using MTBF for non-repairable parts (MTTF is the correct term).
For GATE PI
Expect numericals on R(t) = e^(−λt), MTBF/MTTF, the mission time for a target reliability, λ from total operating time, hazard rates from tabulated failure data, and availability with MTTR. Conceptual MCQs ask which phase of the bathtub curve has which hazard trend and cause, and what the Weibull shape parameter indicates.
Quick check
- λ = 0.001 h⁻¹. What is R(100 h)?
- What fraction of exponential items survive to the MTTF?
- A Weibull fit gives β = 0.6. Which region of the bathtub curve is this?
- MTBF = 400 h and MTTR = 20 h. What is the availability?
Answers: 1. e^(−0.1) = 0.905. 2. 36.8 %. 3. Infant mortality (decreasing hazard). 4. 0.952.
Interview questions
All Metrology, Quality and Reliability interview questionsTry answering each one aloud before you open it.
1.What is the failure rate in reliability engineering?Concept
The (instantaneous) failure rate or hazard rate h(t) = f(t)/R(t) is the probability per unit time that an item which has survived to time t fails in the next instant; it is usually quoted in failures per hour or per 10⁶ h. It is a conditional rate, so it describes how failure risk changes with age, which is what the bathtub curve plots. For a constant failure rate λ, reliability is R(t) = e^(−λt) and MTBF = 1/λ; from field data λ is estimated as failures divided by total operating hours.
2.Explain the bathtub curve in the context of reliability engineering.Concept
The bathtub curve is a graphical representation of the failure rate of a product over time. It consists of three distinct phases: the infant mortality phase with a decreasing failure rate, the normal life phase with a constant failure rate, and the wear-out phase with an increasing failure rate. This curve helps in understanding and predicting the lifecycle of a product.
3.What does MTBF stand for, and how is it calculated?Concept
MTBF stands for Mean Time Between Failures. It is calculated by dividing the total operational time of a system by the number of failures that occurred during that time. MTBF is used to predict the reliability of a system and is expressed in hours.
4.Why is the bathtub curve important in product design and maintenance?Application
The bathtub curve is important in product design and maintenance because it helps engineers understand the different phases of a product's lifecycle. By recognizing these phases, engineers can design products to minimize early failures, plan maintenance during the constant failure rate phase, and anticipate wear-out failures, thereby improving product reliability and customer satisfaction.
5.What happens if a system has a high failure rate during the infant mortality phase?Application
If a system has a high failure rate during the infant mortality phase, it indicates that there are likely design flaws, manufacturing defects, or issues with the initial setup. This can lead to increased warranty claims, customer dissatisfaction, and higher costs for repairs and replacements. Identifying and addressing these issues early can improve the overall reliability of the system.
6.How can MTBF be used to improve system reliability?Application
MTBF can be used to improve system reliability by identifying components or systems with lower MTBF values and targeting them for design improvements or preventive maintenance. By analyzing MTBF data, engineers can prioritize resources to address the most critical reliability issues, thereby enhancing the overall performance and lifespan of the system.
7.Calculate the MTBF for a system that operates for 10,000 hours and experiences 5 failures.Numerical
MTBF = Total operational time / Number of failures = 10,000 hours / 5 failures = 2,000 hours. Therefore, the Mean Time Between Failures for the system is 2,000 hours.
8.If a product has an MTBF of 1,000 hours, what is its failure rate?Numerical
The failure rate (λ) is the reciprocal of MTBF. Therefore, λ = 1 / MTBF = 1 / 1,000 hours = 0.001 failures per hour.
9.Explain how reliability testing can be used to estimate the failure rate of a new product.Application
Reliability testing involves subjecting a product to stress conditions to simulate its lifecycle and observe its performance. By recording the time to failure and the number of failures during testing, engineers can estimate the failure rate. This data helps in predicting the product's reliability and identifying potential improvements before market release.
10.What are some common methods to reduce the failure rate during the wear-out phase?Application
Common methods to reduce the failure rate during the wear-out phase include implementing regular maintenance schedules, using higher quality materials, redesigning components to withstand wear, and employing condition monitoring techniques to predict and prevent failures before they occur. These strategies help extend the product's useful life and improve reliability.
Finished this topic? Mark it so your progress, study plan and readiness keep up.
Stuck on something here?