Mean Time to Recovery (MTTR) and Mean Time Between Failures (MTBF) are key reliability metrics used to evaluate system health beyond simple uptime and downtime tracking. Understanding how to calculate and consistently apply these measurements helps organizations assess and improve the dependability of their systems and services.
MTTR and MTBF
There's actually quite a few different measurements that we can use to measure how healthy our systems are. Not just uptime and downtime, but there are a whole bunch of different measurements. And a couple of those measurements are MTTR and MTBF: mean time to recovery and mean time between failures.
Let's say we have a system or services that our customers are relying on, and every time it goes down it creates a lot of frustration for our customers or our users. And how long it's down also creates frustration for our users. And so what we want to do is we probably want to measure downtime, to understand how much downtime we have as a measurement of how healthy our systems are.
But that's not the only measurement. We could have a system where a big downtime is much worse than lots of little downtimes, or vice versa — that lots of little downtimes are worse than a big downtime. So we have other measurements that we can use to measure the health of a system. One of them is mean time to recovery, MTTR, or another one is mean time between failures, MTBF.
We have something called time to recover, or TTR. Time to recovery is just the time it takes us to recover from an incident or a failure. So we have a system that goes down — there's some sort of failure, the system goes down — and the time to recover is this distance right here. And so we have a measurable time to recover.
When we talk about the mathematical term mean, what we're talking about is an average. So we're just averaging numbers together. So when we're finding the mean of two numbers, we're just averaging those two numbers. So let's say it's the difference between two and four. What is the average of those? Well, the average of that is three. Or another example is, let's say, what is the average between 6 and 12? Well, we could add those together. There's two numbers, so we divide by two, and the average of those two numbers is nine.
So when we're talking about mean time to recovery, what we're doing is we're just adding up all of the recovery times that we have, dividing it by the number of recovery times we have, and then that will give us an average time to recover, or MTTR. And so in this case right here, maybe this is 12 minutes, 8 minutes, and then 10 minutes. So we add this up, and 12 minutes plus 8 minutes plus 10 minutes, we've got 30 minutes total, divided by three incidents here, and so the 10 minutes is our average, or mean time to recovery, MTTR.
Now one thing to note is that the R in this could mean a few different things depending on what source you're looking at. So it's mean time to recovery, or repair, or restore.
The time between failures is just the time from the time we go up to the time we go down. So it's the remaining time here that we're actually up and running here. So we've got maybe this right here is, let's say, 18 minutes, and this is 12 minutes. We add those together, so we've got 18 + 12, so 18 + 12 gives us 30. We have two times here, so we divide that by two, so 15 minutes. The average time, or mean time between failures, is 15 minutes.
Now, not everybody has the same definition of mean time between failures. What we said is from the uptime of one failure to the downtime of the next, that is the TBF. And so if we were to average all of those together, then we would have the mean time, or the average time, between failures.
Now what some people will actually do is define it that, well, the failures are the red arrows here. So here's a failure, and then here's another failure, and then here's another failure. And so the mean time between failures is the average of these times right here. So essentially take all of the operating time right here, all the time in general, and divide it by the number of incidents you have, and you come up with the MTBF.
Now, sources that say that this is the definition of MTBF also have a time to failure. And so this MTTF, time to failure, is then between the uptime and the downtime. So this is TTF. So if we were to average those together, then we would have MTTF, mean time to failure.
So which one do you use? Now, the important thing is just that you're consistent with your measurement, where you set a standard of what it's going to be. That's what you're aiming for and you're measuring. So the important part is the consistency and how you're measuring things, and making sure that whoever is doing the measurement, if there's multiple people, are on the same page with this. So realize there's a little different definition out there, but essentially these are a couple of different methods of measuring or interpreting the mean time between failures.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →