Availability measurement covers the key metrics used to quantify and communicate system uptime, including downtime, uptime percentage, the nines, mean time to recovery, and mean time between failures.
Measuring Availability
Anytime that we're trying to achieve something like high availability, we need to have some sort of target. What are we trying to achieve? Well, we need to be able to measure it then.
Let's say we have customers that are connecting into our system. And what we're going to do is we want to inform those customers on what we expect our uptime to be, or our availability to be. We want to set the expectation for our users so they understand what to expect when it comes to accessing our services.
Now, no system is perfect, and to have 100% uptime is just not realistic. So what happens is that there are times when we have some sort of incident. So this is a timeline right here. Let's say we have an incident, we go down, and at some point in time we recover from that incident and we come back up. And that's going to happen throughout the lifetime of whatever service that we're offering. But how do we measure this?
There are several different measurements that we're going to talk about here. There are more measurements than this, but we're going to go over downtime, uptime, the nines and what those actually mean. We're going to talk about mean time to recovery and mean time between failures.
Downtime is a key measurement to really understanding what is the availability of our systems. For instance, let's say we have a system that over time it's supposed to be up during business hours. So business hours is 8 hours a day. Maybe we have some sort of definition of what those hours are supposed to be. We're setting those expectations for our customer. Well, if during that day we're down for two hours, we now have two hours of downtime during that day.
And often we will measure this in days, or we will measure it in weeks or months or years, whatever the case may be. But it is a measurement of how much time the system is down in the period of time that it should be up. So it's not including things like maintenance windows. So we measure this downtime over a period of, once again, days or months or years, or whatever the case may be.
Now 2 hours in a day would be outrageous. That is a huge amount of time for most systems, I would say. So that is kind of an extreme situation. Hopefully we're not having two hours of downtime in a day.
Now downtime is usually measured in how many minutes it went down. And so the use case for this, I generally find that it's for a per incident. So for instance, we have 30 minutes of downtime right here. Maybe we have 10 minutes of downtime right here, and 20 minutes of downtime right here. So what we would refer to this as being is, well, we had a 30 minute downtime here, we had a 10-minute downtime here. So this is where I find generally that we use this downtime measurement.
Uptime I feel like is a more comprehensive measurement, and more standard, to measure how our systems are doing overall. So now we have that 30 minute, that 10 minute and that 20 minute. Well, what is the overall performance? And so uptime is a little more comprehensive in how we measure it.
Now, uptime is measured in the form of a percent. You can think of it as like a grade. You get a grade on a paper while you're in school. So maybe you got nine out of 10 points on a paper. Well, now you've received 90% on that paper. It gives you an idea of what your score is. Uptime gives us that same type of measurement here.
So what do we do? We take the total time that we should be up and that's divided into the total time that we were actually up. So maybe on this system we were supposed to be up for 10 hours, but we were only up for 9 hours. Well, now we were up 90% of the time. So it gives us the total amount of time that we were up divided by the total time that we are supposed to be up.
Here's an example of maybe we needed to be up 24 hours out of the day. Well, here we went down two hours, then we went down another two hours, and then we went down another 2 hours. So, 6 hours total. So, what do we do? Well, 6 hours. So we were actually up 18 hours out of the day, and we are supposed to be up 24 hours out of that day. And so now we received a grade of 75%.
Now 75% is not good when it comes to system uptime. I hope we can achieve something much greater than that. But what should we achieve? Well, that's where the nines come into play.
The nines is the idea that we want to achieve a certain level of uptime that we have. And if we say that level is 90%, that means that we could be down 36 days out of a year. And if you find that number staggering, like being down 36 days, it is. We don't want those numbers. If we wanted a 95%, that means that we're down 18 days out of the year. Still again way too — in most cases, I'll say in most cases, some systems that might be okay, but 95% for most systems would be way too low.
So the 90% would just be considered one nine. The 99%, which says 3.65 days out of the year, that would be considered what we call two nines. And then if we want to achieve three nines, which is 99.9% — we've got one, two, three nines — that means 8 hours. Now that sounds a little bit reasonable, more reasonable, to have our systems down 8 hours out of the year. Sounds much more achievable. And also our customers are probably going to accept that there are going to be some times that we might be down for a few hours. So this is going to be a little more realistic here.
Maybe we achieve four nines. If we do that, then that means 52 minutes out of the year. Now that's going to be a little bit harder to achieve. 52 minutes — that means we're only down for about an hour out of the year. Or five nines, that means only five minutes out of the year. So it becomes increasingly harder and harder to hit these numbers the more nines that we have.
A common gold standard is this five nines. So often you'll hear this five nines. By the way, even a service like Amazon Web Services, for most of their services, don't report five nines. So this is going to be a hard standard to achieve, but it's something to aspire to, to try to achieve. And then a lot of services sit right around this 99.9, this three nines. So if you look at a lot of Amazon Web Services and a lot of other hosting providers out there, this is what they're guaranteeing, is a 99.9% uptime.
We're going to get into something called mean time to recovery. Well, to understand mean time to recovery, let's first of all understand time to recovery. So right here we have a failure. Our systems go down and they come back up. Let's say this is 30 minutes. So our time to recover took us 30 minutes right there. Here's another one, and that time only took us 10 minutes. And this time right here, it took us 20 minutes to recover. So the time to recover is just per incident: if things were to go down, how much time it took for it to come back up.
Well, as I mentioned, we're going to get into mean time to recover. So what does mean actually mean? Well, it's a mathematical term. Mathematically, mean just means an average. So if we have a two and a six, what we're going to do is we're going to find the average between the two, which the middle of this is a four. So that's the mathematical mean.
So if we're trying to find the mean time to recovery, what we're doing is we're just averaging all of our incident recovery times. So in this case right here, we said it was a 30 minute. In this case right here, 10 minute. In this case right here, 20 minute. Well, we add that all together, 30 + 10 + 20, and that comes up with 60 minutes. So then we divide that by how many we included in here, which was three numbers there. So the average here is 20 minutes. So on average it took a 20 minute time in order to recover between these incidents right here. So that's our mean time to recovery, was 20 minutes.
So what is the mean time between failures? Well, that's just the time between these. So this is this amount right here. So maybe this is 8 hours right here and this is 4 hours right here. Well, what's in between those two numbers? 6 hours. So in this case, 6 hours is our mean time between failures.
Now, realize there's an alternate meaning with this mean time between failures. Most sources that I've seen say that it's between the uptime and the next downtime. So this would be the between failures, and the mean time would be just averaging what those are. But there are some sources that say, well, this is the failure right here and then you have another failure right here, so the mean time between failures is between these two points. It's a little different definition. It actually doesn't matter too much, as long as you're consistent in the way you measure it, and then you also define how you're going to measure it and be deliberate about it. I think this is the more common definition, but this is not wrong either.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →