Written by David Rodgers

Manufacturing Quality Perspective

Written by David Rodgers, Lean Six Sigma Black Belt and ASQ-certified manufacturing quality leader with experience in enterprise storage hardware, quality systems, process improvement, training, and production operations.

Last editorial review: September 24, 2026. Reviewed for statistical accuracy, shop-floor practicality, and educational clarity.

The guides on SixSigmaKaizen.com are written from practical manufacturing experience and are intended to help teams apply Lean, Six Sigma, quality engineering, training, and operations methods more effectively in real production environments.

  • Lean Six Sigma Black Belt
  • ASQ CQE
  • ASQ CMQ/OE
  • Manufacturing leadership
  • Training and operations

MTBF and MTTR are the two numbers behind most equipment reliability discussions: how long a machine runs between failures, and how long it takes to get it running again. Together they determine availability, the fraction of time equipment is ready to work.

This guide gives the definitions and formulas, warns about the most common way MTBF is misread, and works through a packaging machine example that compares two ways to cut downtime. It also covers how to collect data that makes the numbers trustworthy.

Open the Reliability Calculator Open the KPI Calculator

Why These Numbers Matter

They Turn Downtime Into a Measurable Quantity

MTBF and MTTR separate how often equipment fails from how long it takes to get back to work, which point to different fixes.

They Drive Availability

Availability is set by the ratio of MTBF to MTTR. Improving either one raises it.

They Guide Spares and Staffing

MTTR shows where repair time goes, and MTBF shows how often the crew and spares will be needed.

They Are Easy to Misread

An MTBF is an average, not a guaranteed life. Misreading it leads to wrong maintenance intervals.

A maintenance technician repairing a motor on a conveyor system while a colleague holds a flashlight
Every repair has a clock on it: the time to restore is what MTTR measures.

Definitions

TermFormulaMeaning
MTBFTotal operating time / number of failuresAverage operating time between failures of a repairable item.
MTTRTotal repair time / number of repairsAverage time to restore an item after a failure.
MTTFTotal operating time / number of items failedAverage time to failure of a non-repairable item, such as a light bulb.
Failure rate (λ)1 / MTBF (constant-rate case)Failures per operating hour.
Inherent availabilityMTBF / (MTBF + MTTR)Fraction of time up, counting only failures and repairs.
Reliability R(t)e−t / MTBFProbability of no failure over time t, if the failure rate is constant.

The Most Common Misreading

An MTBF of 450 hours does not mean an item will last 450 hours. With a constant failure rate, the probability of surviving a full MTBF is e−1, about 36.8%, so roughly 63% of items fail before reaching it. MTBF is an average of a spread of failure times, not a guaranteed life.

It also assumes failures happen at a roughly constant rate. Equipment that wears out, or that fails mostly early in life, needs a different model, such as Weibull analysis. See the Product Lifecycle Reliability Analysis tool and the Weibull Analysis entry.

Worked Example: A Packaging Machine

A packaging machine ran for 4,050 hours over six months and failed 9 times. Repairs took a total of 72 hours. The numbers are illustrative.

MeasureCalculationResult
MTBF4,050 / 9450 hours
MTTR72 / 98 hours
Availability450 / (450 + 8)98.25%
Reliability over 100 hourse−100/45080.1%
Expected downtime per 8,760-hour year8,760 × (1 − 0.9825)about 153 hours

Two improvement options are compared, each of which improves the MTBF-to-MTTR ratio from 56 to 113:

  • Halve MTTR from 8 to 4 hours, through spares kits, a better procedure, and pre-staged tools.
  • Double MTBF from 450 to 900 hours, through better preventive tasks or a design change.
153 h TodayMTBF 450 h, MTTR 8 h 77 h Halve MTTRMTBF 450 h, MTTR 4 h 77 h Double MTBFMTBF 900 h, MTTR 8 h
Halving repair time and doubling time between failures give the same availability, because availability depends on the ratio of the two.

The lesson is that neither number is automatically the best lever. Compare the cost and effort of each option. Halving MTTR by reorganizing the spares crib may cost far less than redesigning a component to double its life, and in the packaging example it delivers the same 76 hours less downtime a year.

Try your own numbers in the Reliability and Availability Calculator, and track them monthly with the Maintenance KPI Calculator.

A maintenance planner in a control room reviewing equipment status indicators on wall monitors
Availability is the ratio of uptime to total time, and it depends on both how often things fail and how fast they are restored.

Getting Good Data

  • Define a failure. Decide whether a short stop, a quality reject, or a minor jam counts, and apply the same rule every time.
  • Separate repair time from waiting time. MTTR that includes waiting for parts hides the reason for delay. Record active repair, waiting for parts, and waiting for people separately.
  • Use operating time, not calendar time. Equipment that runs one shift a day accumulates hours differently from one that runs continuously.
  • Track by failure mode. An overall MTBF averages very different problems. Break it down with a Pareto of causes.
  • Use enough data. Nine failures give a rough estimate. Report the sample size along with the MTBF.

Improving Availability: Fail Less Often or Repair Faster

Because availability depends on both MTBF and MTTR, there are two levers. Which one to pull depends on where the losses come from, and on cost.

Current (MTBF 50 h, MTTR 2 h) 96.2% Double MTBF (100 h, 2 h) 98.0% Halve MTTR (50 h, 1 h) 98.0% Inherent availability = MTBF ÷ (MTBF + MTTR)
Halving repair time gives the same availability as doubling the time between failures, and is often cheaper to achieve.

The figure uses the simple formula for inherent availability, MTBF divided by MTBF plus MTTR: 50 / 52 = 96.2%, 100 / 102 = 98.0%, and 50 / 51 = 98.0%. The two options are equal on paper, but they cost very different amounts and act on different causes.

LeverTypical actionsWatch for
Raise MTBFRoot cause analysis of repeat failures, better lubrication and cleaning, condition monitoring, design changes, operator careCostly redesign when a simple procedure would do; failure modes that are random, not age-related
Cut MTTRSpares at the point of use, standard repair procedures, diagnostics and alarms that point to the fault, modular parts, cross-trained crews, quicker access to specialistsFast repairs that are not durable, so MTBF falls

Look at the breakdown of repair time. MTTR includes detection, response, diagnosis, waiting for parts, the repair itself, and testing. Often the actual repair is a minority of the time, and waiting for parts or people is the big bar. A simple breakdown of a dozen repairs shows where to start.

Making the Numbers Trustworthy

Reliability figures are only as good as the records behind them. A few habits keep the measures from misleading.

  • Define a failure. Decide whether minor stops and micro-stoppages count, and use the same rule for every machine. Without this, MTBF from two sites is not comparable.
  • Record start and end times of each downtime event, and the reason and the part that failed, in a system that operators actually use. Reason codes should be few and clear.
  • Use enough history. A handful of failures gives a wide range. Report the count of failures along with the average.
  • Separate planned from unplanned downtime. Planned maintenance lowers operational availability but is not a failure; mixing them hides both.
  • Look at the distribution, not only the mean. A few long repairs can distort MTTR; the median and the longest events tell you more about what to fix.

Know which availability you are quoting. Inherent availability counts only corrective repair time. Operational availability includes every source of downtime, such as preventive maintenance, waiting for parts, and administrative delay, and it is what production experiences. If the two differ by much, the gap points to logistics and planning, not to the machine itself.

See the Reliability and Availability Calculator for calculations, the Total Productive Maintenance Guide and the Reliability-Centered Maintenance Guide for improvement methods.

Self-Assessment Questions

  • Do we have a clear, written definition of a failure?
  • Do we record operating hours and repair hours for each event?
  • Do we separate active repair from waiting time?
  • Do we analyze by failure mode, not just overall?
  • Do we test whether a constant failure rate is reasonable before relying on MTBF?

Common Mistakes

Treating MTBF as a Guaranteed Life

About 63% of items fail before reaching the MTBF when the failure rate is constant. Do not set replacement intervals from MTBF alone.

Mixing Failure Modes

Averaging unrelated problems hides the ones that matter. Analyze by mode.

Including Waiting in MTTR Without Saying So

It is fine to measure it, but separate it so you can fix delays in spares and staffing.

Comparing Numbers With Different Definitions

MTBF from two sites means little if they define failure differently.

MTBF, MTTR, and Availability: Frequently Asked Questions

What is the difference between MTBF and MTTR?

MTBF, mean time between failures, is the average operating time between failures of a repairable item, and measures how often it fails. MTTR, mean time to repair, is the average time to restore it after a failure, and measures how long each failure lasts. Availability combines them: MTBF divided by MTBF plus MTTR.

Does an MTBF of 1,000 hours mean the equipment lasts 1,000 hours?

No. With a constant failure rate, the probability of running a full MTBF without failure is about 37%, so most items fail before reaching it. MTBF is an average of a spread of failure times, and equipment with wear-out behavior needs a different model such as Weibull analysis.

How do I improve availability?

Availability depends on the ratio of MTBF to MTTR, so either fewer failures or faster repairs improve it. Compare the cost of each option, since shortening repair time through spares, procedures, and staging is often cheaper than extending time between failures through design changes.

Sources and Further Reading

  • NIST/SEMATECH e-Handbook of Statistical Methods, reliability chapter.
  • Patrick D. T. O'Connor and Andre Kleyner, Practical Reliability Engineering.
  • IEC 60050-192, International Electrotechnical Vocabulary: Dependability.
  • ASQ Certified Reliability Engineer Body of Knowledge.