TechKnowSurge
VideoSecurityFree

Measuring Availability

Availability measurement covers the key metrics used to quantify and communicate system uptime, including downtime, uptime percentage, the nines, mean time to recovery, and mean time between failures.

Complete this video to capture a CTF flag worth 1 point.

About this video

Availability is not an abstract goal but a measurable characteristic of any system, and organizations use a set of standard metrics to define, track, and communicate it. Downtime refers to the total duration a system is unavailable within a period it is expected to be operational, typically excluding scheduled maintenance windows. It is most commonly applied on a per-incident basis to capture how long a specific outage lasted. Uptime, by contrast, is a broader and more widely used measurement expressed as a percentage of total expected operational time during which the system was actually available, functioning similarly to a performance score across a defined window such as a day, month, or year. The nines framework puts uptime percentages into practical perspective by converting them into annual downtime figures. A single nine, or 90% uptime, allows for roughly 36 days of downtime per year, which is unacceptable for most production systems. Three nines, or 99.9%, translates to approximately 8 hours of allowable downtime annually and represents a realistic baseline that many customers and service-level agreements find acceptable. Five nines, or 99.999%, permits only about five minutes of downtime per year and is widely regarded as the gold standard, though even major cloud providers do not guarantee it across all services. Many hosting and cloud platforms commonly target 99.9% as their published uptime commitment. Mean time to recovery and mean time between failures provide additional insight into how a system behaves across multiple incidents over time. Mean time to recovery is the mathematical average of all individual recovery durations, giving teams a reliable benchmark for how quickly service is typically restored after an outage. Mean time between failures measures the average interval between outages, offering a view into how frequently failures occur. Some definitions measure this interval from the end of one outage to the start of the next, while others measure from failure point to failure point. Both interpretations are valid as long as the methodology is clearly defined and applied consistently, since the value of these metrics depends entirely on the discipline with which they are tracked.

What you'll learn

What's covered

Measuring Availability

Key terms

Availability
The assurance that systems and data are accessible and operational when needed by authorized users.
Uptime
The amount of continuous time a network device or service has been running without interruption, often expressed as a percentage of total time available. High uptime percentages (such as 99.9%) indicate a reliable network, while planned and unplanned downtime reduces this figure.
Downtime
The period during which a system, service, or component is unavailable or non-operational; it is measured against uptime targets defined in service agreements.
The Nines
A shorthand notation for expressing uptime targets as percentages containing repeated 9s, such as 99.9% (three nines), where each additional nine reduces the acceptable annual downtime threshold.
Mean Time to Recover
MTTR
Mean Time to Recover is the average time required to restore a system or service to normal operation after a failure or security incident, used as a key metric for measuring incident response effectiveness and system resilience.
Mean Time Between Failures
MTBF
Mean Time Between Failures is a reliability metric representing the average operational time between component failures, used in availability planning and security resilience design.

Topics

Availability Metrics Uptime Measurement Mean Time To Recovery Mean Time Between Failures Nines Notation It Service Reliability

Transcript

Anytime that we're trying to achieve something like high availability, we need to have some sort of target. What are we trying to achieve? Well, we need to be able to measure it then.

Let's say we have customers that are connecting into our system. And what we're going to do is we want to inform those customers on what we expect our uptime to be, or our availability to be. We want to set the expectation for our users so they understand what to expect when it comes to accessing our services.

Now, no system is perfect, and to have 100% uptime is just not realistic. So what happens is that there are times when we have some sort of incident. So this is a timeline right here. Let's say we have an incident, we go down, and at some point in time we recover from that incident and we come back up. And that's going to happen throughout the lifetime of whatever service that we're offering. But how do we measure this?

There are several different measurements that we're going to talk about here. There are more measurements than this, but we're going to go over downtime, uptime, the nines and what those actually mean. We're going to talk about mean time to recovery and mean time between failures.

Downtime

Downtime is a key measurement to really understanding what is the availability of our systems. For instance, let's say we have a system that over time it's supposed to be up during business hours. So business hours is 8 hours a day. Maybe we have some sort of definition of what those hours are supposed to be. We're setting those expectations for our customer. Well, if during that day we're down for two hours, we now have two hours of downtime during that day.

And often we will measure this in days, or we will measure it in weeks or months or years, whatever the case may be. But it is a measurement of how much time the system is down in the period of time that it should be up. So it's not including things like maintenance windows. So we measure this downtime over a period of, once again, days or months or years, or whatever the case may be.

Now 2 hours in a day would be outrageous. That is a huge amount of time for most systems, I would say. So that is kind of an extreme situation. Hopefully we're not having two hours of downtime in a day.

Now downtime is usually measured in how many minutes it went down. And so the use case for this, I generally find that it's for a per incident. So for instance, we have 30 minutes of downtime right here. Maybe we have 10 minutes of downtime right here, and 20 minutes of downtime right here. So what we would refer to this as being is, well, we had a 30 minute downtime here, we had a 10-minute downtime here. So this is where I find generally that we use this downtime measurement.

Uptime

Uptime I feel like is a more comprehensive measurement, and more standard, to measure how our systems are doing overall. So now we have that 30 minute, that 10 minute and that 20 minute. Well, what is the overall performance? And so uptime is a little more comprehensive in how we measure it.

Now, uptime is measured in the form of a percent. You can think of it as like a grade. You get a grade on a paper while you're in school. So maybe you got nine out of 10 points on a paper. Well, now you've received 90% on that paper. It gives you an idea of what your score is. Uptime gives us that same type of measurement here.

So what do we do? We take the total time that we should be up and that's divided into the total time that we were actually up. So maybe on this system we were supposed to be up for 10 hours, but we were only up for 9 hours. Well, now we were up 90% of the time. So it gives us the total amount of time that we were up divided by the total time that we are supposed to be up.

Here's an example of maybe we needed to be up 24 hours out of the day. Well, here we went down two hours, then we went down another two hours, and then we went down another 2 hours. So, 6 hours total. So, what do we do? Well, 6 hours. So we were actually up 18 hours out of the day, and we are supposed to be up 24 hours out of that day. And so now we received a grade of 75%.

The Nines

Now 75% is not good when it comes to system uptime. I hope we can achieve something much greater than that. But what should we achieve? Well, that's where the nines come into play.

The nines is the idea that we want to achieve a certain level of uptime that we have. And if we say that level is 90%, that means that we could be down 36 days out of a year. And if you find that number staggering, like being down 36 days, it is. We don't want those numbers. If we wanted a 95%, that means that we're down 18 days out of the year. Still again way too — in most cases, I'll say in most cases, some systems that might be okay, but 95% for most systems would be way too low.

So the 90% would just be considered one nine. The 99%, which says 3.65 days out of the year, that would be considered what we call two nines. And then if we want to achieve three nines, which is 99.9% — we've got one, two, three nines — that means 8 hours. Now that sounds a little bit reasonable, more reasonable, to have our systems down 8 hours out of the year. Sounds much more achievable. And also our customers are probably going to accept that there are going to be some times that we might be down for a few hours. So this is going to be a little more realistic here.

Maybe we achieve four nines. If we do that, then that means 52 minutes out of the year. Now that's going to be a little bit harder to achieve. 52 minutes — that means we're only down for about an hour out of the year. Or five nines, that means only five minutes out of the year. So it becomes increasingly harder and harder to hit these numbers the more nines that we have.

A common gold standard is this five nines. So often you'll hear this five nines. By the way, even a service like Amazon Web Services, for most of their services, don't report five nines. So this is going to be a hard standard to achieve, but it's something to aspire to, to try to achieve. And then a lot of services sit right around this 99.9, this three nines. So if you look at a lot of Amazon Web Services and a lot of other hosting providers out there, this is what they're guaranteeing, is a 99.9% uptime.

Mean Time to Recovery

We're going to get into something called mean time to recovery. Well, to understand mean time to recovery, let's first of all understand time to recovery. So right here we have a failure. Our systems go down and they come back up. Let's say this is 30 minutes. So our time to recover took us 30 minutes right there. Here's another one, and that time only took us 10 minutes. And this time right here, it took us 20 minutes to recover. So the time to recover is just per incident: if things were to go down, how much time it took for it to come back up.

Well, as I mentioned, we're going to get into mean time to recover. So what does mean actually mean? Well, it's a mathematical term. Mathematically, mean just means an average. So if we have a two and a six, what we're going to do is we're going to find the average between the two, which the middle of this is a four. So that's the mathematical mean.

So if we're trying to find the mean time to recovery, what we're doing is we're just averaging all of our incident recovery times. So in this case right here, we said it was a 30 minute. In this case right here, 10 minute. In this case right here, 20 minute. Well, we add that all together, 30 + 10 + 20, and that comes up with 60 minutes. So then we divide that by how many we included in here, which was three numbers there. So the average here is 20 minutes. So on average it took a 20 minute time in order to recover between these incidents right here. So that's our mean time to recovery, was 20 minutes.

Mean Time Between Failures

So what is the mean time between failures? Well, that's just the time between these. So this is this amount right here. So maybe this is 8 hours right here and this is 4 hours right here. Well, what's in between those two numbers? 6 hours. So in this case, 6 hours is our mean time between failures.

Now, realize there's an alternate meaning with this mean time between failures. Most sources that I've seen say that it's between the uptime and the next downtime. So this would be the between failures, and the mean time would be just averaging what those are. But there are some sources that say, well, this is the failure right here and then you have another failure right here, so the mean time between failures is between these two points. It's a little different definition. It actually doesn't matter too much, as long as you're consistent in the way you measure it, and then you also define how you're going to measure it and be deliberate about it. I think this is the more common definition, but this is not wrong either.

About TechKnowSurge

TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.

Explore free tools and programs →