TechKnowSurge
NIST NICE K0740 NIST NICE K0741
VideoSecurityFree

MTTR and MTBF

Mean Time to Recovery (MTTR) and Mean Time Between Failures (MTBF) are key reliability metrics used to evaluate system health beyond simple uptime and downtime tracking. Understanding how to calculate and consistently apply these measurements helps organizations assess and improve the dependability of their systems and services.

Complete this video to capture a CTF flag worth 1 point.

About this video

System reliability is measured using more than just total downtime. Two of the most widely used metrics in IT operations are Mean Time to Recovery (MTTR) and Mean Time Between Failures (MTBF), each of which captures a different dimension of how a system performs over time. MTTR represents the average duration it takes to restore a system or service following a failure. It is calculated by summing all individual recovery times across a set of incidents and dividing by the total number of incidents, yielding a single average recovery duration that helps teams understand how quickly they can respond to and resolve outages. The R in MTTR can stand for recovery, repair, or restore depending on the source, but the calculation method remains the same. MTBF measures the average time a system operates between failures—essentially the uptime intervals between successive outages. By adding all operational periods together and dividing by the number of failures, organizations can gauge how stable and dependable a system is under normal conditions. Definitions of MTBF can differ across industry sources. Some define it strictly as the time between the end of one failure and the start of the next, while others calculate it as total operating time divided by the number of incidents. Sources using the latter definition often distinguish a separate metric, Mean Time to Failure (MTTF), to describe the time from recovery to the next failure event. Regardless of which specific definition an organization adopts, consistency in measurement is what makes these metrics actionable. Teams tracking MTTR and MTBF must align on a shared standard so that data collected over time remains comparable and accurately reflects changes in system performance. These metrics together provide a more complete picture of system health than downtime alone, supporting better decisions around maintenance, infrastructure investment, and service reliability goals.

What you'll learn

What's covered

MTTR and MTBF

Aligned to

NIST NICE
K0740 Knowledge of system performance indicators
K0741 Knowledge of system availability measures
K0740 Knowledge of system performance indicators
K0741 Knowledge of system availability measures

Key terms

Mean Time to Recover
MTTR
Mean Time to Recover is the average time required to restore a system or service to normal operation after a failure or security incident, used as a key metric for measuring incident response effectiveness and system resilience.
Mean Time Between Failures
MTBF
Mean Time Between Failures is a reliability metric representing the average operational time between component failures, used in availability planning and security resilience design.
Mean Time to Failure
MTTF
Mean Time to Failure is the expected operational lifespan of a non-repairable component before it fails, used in availability and resilience planning to predict when hardware components in security infrastructure should be replaced.
Time to Recovery
TTR
The measured duration it takes to restore a system or service following a single failure or incident.

Topics

Mttr Mtbf System Reliability Reliability Metrics Incident Management Systems Administration

Transcript

There's actually quite a few different measurements that we can use to measure how healthy our systems are. Not just uptime and downtime, but there are a whole bunch of different measurements. And a couple of those measurements are MTTR and MTBF: mean time to recovery and mean time between failures.

Let's say we have a system or services that our customers are relying on, and every time it goes down it creates a lot of frustration for our customers or our users. And how long it's down also creates frustration for our users. And so what we want to do is we probably want to measure downtime, to understand how much downtime we have as a measurement of how healthy our systems are.

But that's not the only measurement. We could have a system where a big downtime is much worse than lots of little downtimes, or vice versa — that lots of little downtimes are worse than a big downtime. So we have other measurements that we can use to measure the health of a system. One of them is mean time to recovery, MTTR, or another one is mean time between failures, MTBF.

Time to recover, and what mean means

We have something called time to recover, or TTR. Time to recovery is just the time it takes us to recover from an incident or a failure. So we have a system that goes down — there's some sort of failure, the system goes down — and the time to recover is this distance right here. And so we have a measurable time to recover.

When we talk about the mathematical term mean, what we're talking about is an average. So we're just averaging numbers together. So when we're finding the mean of two numbers, we're just averaging those two numbers. So let's say it's the difference between two and four. What is the average of those? Well, the average of that is three. Or another example is, let's say, what is the average between 6 and 12? Well, we could add those together. There's two numbers, so we divide by two, and the average of those two numbers is nine.

So when we're talking about mean time to recovery, what we're doing is we're just adding up all of the recovery times that we have, dividing it by the number of recovery times we have, and then that will give us an average time to recover, or MTTR. And so in this case right here, maybe this is 12 minutes, 8 minutes, and then 10 minutes. So we add this up, and 12 minutes plus 8 minutes plus 10 minutes, we've got 30 minutes total, divided by three incidents here, and so the 10 minutes is our average, or mean time to recovery, MTTR.

Now one thing to note is that the R in this could mean a few different things depending on what source you're looking at. So it's mean time to recovery, or repair, or restore.

Time between failures

The time between failures is just the time from the time we go up to the time we go down. So it's the remaining time here that we're actually up and running here. So we've got maybe this right here is, let's say, 18 minutes, and this is 12 minutes. We add those together, so we've got 18 + 12, so 18 + 12 gives us 30. We have two times here, so we divide that by two, so 15 minutes. The average time, or mean time between failures, is 15 minutes.

Now, not everybody has the same definition of mean time between failures. What we said is from the uptime of one failure to the downtime of the next, that is the TBF. And so if we were to average all of those together, then we would have the mean time, or the average time, between failures.

Now what some people will actually do is define it that, well, the failures are the red arrows here. So here's a failure, and then here's another failure, and then here's another failure. And so the mean time between failures is the average of these times right here. So essentially take all of the operating time right here, all the time in general, and divide it by the number of incidents you have, and you come up with the MTBF.

Now, sources that say that this is the definition of MTBF also have a time to failure. And so this MTTF, time to failure, is then between the uptime and the downtime. So this is TTF. So if we were to average those together, then we would have MTTF, mean time to failure.

Which one do you use?

So which one do you use? Now, the important thing is just that you're consistent with your measurement, where you set a standard of what it's going to be. That's what you're aiming for and you're measuring. So the important part is the consistency and how you're measuring things, and making sure that whoever is doing the measurement, if there's multiple people, are on the same page with this. So realize there's a little different definition out there, but essentially these are a couple of different methods of measuring or interpreting the mean time between failures.

About TechKnowSurge

TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.

Explore free tools and programs →