System availability is measured as a percentage of uptime over a defined period, with industry standards ranging from three nines (99.9%) to five nines (99.999%) determining acceptable downtime thresholds. Service level agreements (SLAs) formalize these commitments and define financial penalties when availability targets are not met.
One of the things that we're trying to do when it comes to incident response is to maintain a high uptime and high performance level, and make sure that our customers remain happy. But how do we measure this? One of the ways that we measure this is by measuring availability, and there's also some concepts around nines and five 9s.
We're going to use this graph to illustrate several of our examples. There is time that elapses, so the line here represents time. Let's say we have a system — maybe it's a web service that we have customers getting on — and sometimes this service goes down. If it goes down we see a red arrow; when it comes back up we have a green arrow. So over time here it went down again and it came back up, and then here it went down a third time and then came back up. This period of time, maybe it's over the day, maybe it's over the week or month, or whatever the case may be — that's not really important for us in this discussion right now — but this is going to be representative of a system that goes down and up occasionally.
One of the measurements we use is downtime, just to figure out how much total down a system has been. Here we have an incident that occurs and we've got this down, and then maybe it comes back up, and maybe in total this was down for 10 minutes. So now we've got 10 minutes of downtime. Here again we see another little 7 minute outage, and then another 10 minute outage. So for this period of time we see that it was down 27 minutes. That's one of the ways we can measure downtime: in minutes.
We can also measure the uptime. Maybe between these two incidents right here, that was up for 10 hours, and between these incidents right here that was up for another 4 hours, so we can see that our uptime during this period of time is 14 hours. The problem is that when we measure uptime in this way, we don't know: is that 14 hours in a period of a day, or a month, or a year? Because that could make a huge difference.
So usually we measure uptime as a percentage, not as based off of an actual time. Here's the calculation for measuring uptime. Usually we just take the total amount of uptime and we divide it by the total amount of time, and we get a percentage: during this time period, of any time period, this is the percent out of that time period that it was up. A similar measurement would be the total time minus the downtime, divided by the total time, and here again we would end up with a percentage — the percentage of how much time it was up during this time period, during this total time.
Let's take a look at an example of uptime and calculating it. Here's an example where we went down 2 hours out of the day and we were up for 18 hours. Total time here is 24 hours; this is during a 24-hour period of time. So we're going to divide the total of uptime, so in this case 18. If we get 18 divided by 24, that would be 75%. So here we have an example of a 75% uptime.
Now, there are some common standards out there, and this is a chart of the nines. What it's showing us is if we take a percentage of uptime, 75% would be considered terrible, absolutely terrible. So we start here in this at 90%. If you had 90% uptime over the year, that means you would be down for 36.5 days — another example that is terrible. We wouldn't want 90% uptime. If we went to 95% uptime, now we're looking at 18.25 days. If we went to a 97% uptime, then that's down for 10 days, which is still a lot. And so we start seeing 98% is down for 7 days, 99% is 3.65 days.
At 99.9% now we're down to 8 hours. This is a common standard — three 9s is a very common standard that's out there, you see it on a lot of different services. What it is saying is that over the year we could be down 8 hours out of the year, or 8.76 hours, or 8 hours and 45 minutes out of the year, and so that's what our allowance is. If we want to shoot for something higher we can go to four 9s, and that's 52 minutes of downtime, and then five 9s is 5 minutes of downtime. You can see as we go up, nine 9s is just milliseconds out of the year, so that's kind of crazy for this type of measurement — that would be a little unreasonable.
Now, I mentioned a common standard out there is three 9s. This is a very acceptable standard to many, where out of the year we can have an 8 hour, 8.7 hour downtime there. But the gold standard is considered around this five 9s. If you can get it down to five 9s, that's only 5 minutes out of the year of downtime, then you are doing excellent. So this is a gold standard here, versus the three 9s is a common standard.
Now let's say our customers are demanding a certain level from us and they want that in writing. How would we deliver that? We'd probably create some sort of service level agreement, or SLA. The SLA is going to be the agreement that we have on what our service levels are going to be.
Usually just saying that we're going to do something is not enough for the customer. Perhaps what they want is some sort of reassurance that you are going to follow through with whatever your service level agreement says. So usually on our SLA we also stipulate some sort of penalties or repercussions if we fall below those numbers. Often it comes in the form of a reimbursement.
So for instance, maybe they're paying us $1,000 a month, and if we operate at, let's say, a three 9s level here — if we operate at this level, then there's no reimbursement. But if we fall below that, then maybe there's some sort of reimbursement. So now let's say we go to two 9s instead, so we had a big outage here. Maybe at that point in time there's a 10% reimbursement, so off of that $1,000 we're going to give a credit back for $100. And then it could be even, if we go to the next step right there, maybe this range right here somewhere between 95 and 99%, then perhaps there is a 50% reimbursement — that we give them $500 back. And then if it's below 90%, then maybe it's a full credit back, that they don't pay for that month. So there's some sort of reimbursement that happens as a penalty.
Often we want to strive for something higher, so we choose this gold standard of five 9s, but in our SLAs we're only going to guarantee three 9s. This makes our target, what we want to achieve internally, much higher, creating more customer satisfaction, but we're not legally or financially responsible at this lower rate here, 99.9%.
One aspect to these SLAs though is that we are going to be considering the duration that we're guaranteeing that for. What I mean by that is the amount of time in a year, quarter and month are obviously different. Here I've calculated out 525,000 minutes in a year, 131,000 in a quarter, and 43,000 in a month.
At 99.999%, are we guaranteeing that in the month period of time, in the quarter period of time, or in the year period of time? Because if it's by the year, we're guaranteeing that we're not going to go below 5.26 minutes out of that year. So if we go 5 minutes of downtime within a month, we've not crossed over that line and so we don't need to worry about it. But if we're guaranteeing it out of the whole year, and maybe we have a downtime in one month and it's 5 minutes, and another downtime in the next month which is 5 minutes, that's a total of 10 minutes, and now we're going to have to pay out for that.
Or we could be guaranteeing it for the quarter. For the quarter, 99.999% of 131,000 is only 79 seconds, and out of the month it's only 26 seconds. So the duration in which we're guaranteeing this uptime matters as well.
One thing to mention with this is that there are times when we intentionally bring the system down — for instance, maybe we're doing some sort of maintenance. For that, we are going to exclude maintenance windows out of our SLA. Here is an example of maybe there's an incident, we've gone down right here, and this is another incident, we've gone down right here, but this downtime right here was a maintenance window that is planned and communicated ahead of time.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →