TechKnowSurge
NIST NICE S0850 NIST NICE S0858 ISC2 CISSP 1.9 CompTIA SecurityX 1.3 CompTIA Security+ 5.2
VideoSecurityFree

Cost-Benefit Analysis (CBA) Example

Cost-benefit analysis applied to IT infrastructure weighs the revenue lost from service downtime against the investment required to prevent it. Finding the optimal balance between these two cost curves determines the most financially sound level of uptime investment.

Complete this video to capture a CTF flag worth 1 point.

About this video

For a SaaS company serving millions of users, unplanned downtime directly translates to lost revenue. Customers have varying tolerances for service interruptions — some will leave after a single significant outage, others will stay through minor disruptions — but as downtime increases, so does churn, and at extreme levels, it can erode an entire customer base. Mapping this relationship produces a downtime cost curve that rises as reliability decreases, representing the revenue a business sacrifices when it underinvests in infrastructure stability. Reducing downtime requires spending on solutions such as load balancers, RAID configurations, failover systems, and high-availability architectures. Early investments in these areas tend to deliver the greatest impact, eliminating the most obvious failure points at relatively low cost. As the system matures, each additional increment of reliability becomes more expensive to achieve, following the law of diminishing returns. Reaching 99% uptime may be straightforward, while pushing to 99.999% demands exponentially greater investment for a comparatively small gain. Overlaying the downtime cost curve with the mitigation cost curve creates a composite view of total expenditure at every level of investment. On the left side of the graph, minimal mitigation spending produces high downtime losses; on the right, heavy mitigation spending reduces those losses but introduces excessive operational costs. The point where the two curves intersect — or where their combined value is lowest — represents the optimal investment level. A cost-benefit analysis uses this framework to guide infrastructure decisions, ensuring that spending on reliability is justified by the revenue it protects rather than driven by the pursuit of theoretical perfection.

What you'll learn

What's covered

Cost-Benefit Analysis

Aligned to

NIST NICE
S0850 Skill in performing cost/benefit analysis
S0858 Skill in performing economic analysis
ISC2 CISSP
1.9 Understand and apply risk management concepts
CompTIA SecurityX
1.3 Explain the importance of risk management for an enterprise
CompTIA Security+
5.2 Explain elements of the risk management process

Key terms

Software as a Service
SaaS
A cloud service model that delivers software applications over the internet on a subscription basis.
Availability
The assurance that systems and data are accessible and operational when needed by authorized users.
Redundancy
The duplication of critical components or systems to increase reliability and availability.
Failover
The automatic switching to a redundant system or component when the primary one fails.
Load Balancer
A device or software that distributes incoming network traffic across multiple servers to ensure availability and performance.
Cost-Benefit Analysis
A financial evaluation method that compares the costs of implementing a solution against the benefits it produces to support decision-making.
Mitigation Cost
The financial investment required to reduce or prevent service downtime, including redundancy, failover systems, and high availability measures.
Optimal Investment Point
The point in a cost-benefit analysis where the combined total of mitigation spending and revenue loss from downtime is at its lowest.

Topics

Cost Benefit Analysis Risk Management Service Availability Downtime Cost It Governance Uptime Optimization

Transcript

There's actually a lot of different ways we can do a cost benefit analysis. Here's one way that we could carry out a cost benefit analysis.

Let's develop a little scenario here. Let's say we're part of a SaaS company, a software as a service company, and we have a bunch of users that gain access, millions of users that gain access to this service. So therefore, if these services go down, we've got a lot of angry customers here that don't like our product if there's too much downtime. If it's just a little bit, they're usually pretty forgiving, but if there's a lot, then a lot of customers may leave us because they're unhappy with the services that we're providing.

The Cost of Downtime

We can plot this out on a graph. Let's say we have very little downtime, so this is how much downtime we have. Maybe it's expressed in minutes, maybe hours, maybe days, who knows what this measurement is, but let's just say we have very, very little downtime as compared to other people in the industry. Now what's going to happen is really there's not going to be very much impact to that, because people are just expecting this, and so there's very little impact.

But let's say we go up a little bit and have some downtime here. What's going to happen is there's going to be a cost to that. The cost is that we're going to have some users that really have very little tolerance to this downtime, and they're going to ditch our services. They're going to leave it, they're going to say I don't want your services because I'm aggravated. Maybe they just happen to hit it during the worst times that we had these downtimes, and there are just going to be a fraction of those users that are going to say I'm going to leave. You're going to have some that are going to say I'm a little concerned about this but I'm not going to leave the services yet, and then you're going to have some that are kind of clueless that anything has happened.

Now we get a little more downtime. Now what happens is it's a little more prevalent, and you're going to have a larger number of users that are going to say I don't like this, this is not good, and you're going to have a larger number that leave, driving your costs up. So now you have more cost involved with this.

Let's say then it gets out of control and you have a huge amount of issues here. What happens here is now there's a bunch of people that are not happy with this and a bunch of people that are leaving, so it drives the cost up. So the further we go right on here, the more cost, the more money that we are losing in revenue, because we've got people that are actually leaving our services and now no longer paying for our services. And so this can be quite devastating, and there comes a certain point where you just lose all your customers and now you're not making any revenue because everybody's pretty much fed up with your services.

The Cost of Mitigation

So what we have here is downtime, and in order to drive the downtime down, to get the downtime as low as possible, we need to increase our measures that we put into place. We need to figure out what's going on with it, we need to figure out how can we create things like load balancers and high availability and ways to mitigate IT issues, failover. We put RAID, extra RAID disks in, we create a lot of redundancy. Well, all of that takes time and money, and it all equates to money. So the mitigation costs go up. So as mitigation costs go up, we drive the downtime down, and so there's this balancing act between these two.

So we've got the cost of how much things are going to cost, and we also have how much downtime we have. We want to get this as low as possible. So what we have here is, if we have a high amount of downtime because we haven't done the proper maintenance on our systems and we haven't been able to troubleshoot what's going on and fix these issues that are happening, what's going to happen is we've essentially spent no money on mitigation.

Now we spend a little bit of money on mitigation, so we spend a little bit and that drives some of this downtime down. Now we actually make significant improvements with just a little bit of money, because we're hitting the low hanging fruit and there's some obvious things that are wrong. And then we spend a little more money and we drive it down even further, and so this also has a huge improvement, because now we're hitting that stuff, maybe it's not the low hanging fruit but it's still stuff that is a little bit obvious, and we are fixing these issues.

Then we spend a lot of money and we see some improvement with that as well, not as much improvement. There's something called the law of diminishing returns that I'm not going to get too much into, but we're going to spend now quite a bit more to get to the next level. And at some point it just becomes kind of ridiculous on how much money we're spending, and we're not getting as big of a return from that investment here.

So we're going to have to figure out how much money we're going to spend to drive our downtime. To get to the 99% might be easy, to get to the 99.9% might be a little more difficult, to get to 99.999% uptime, these are all uptime measurements by the way, it's going to be even more difficult there. So we're going to exponentially have to put cost in there to drive that downtime down, which essentially is increasing that uptime that I just mentioned, which is usually measured in percentage.

Overlapping the Two Graphs

So what we need to do is we need to do an evaluation, a cost benefit analysis, to figure out how much time and energy we spend in these servers. You spend not enough and we're going to start losing revenue from those users, but spend too much and it's extra money that we're spending that we might not need to spend.

So what we do is we overlap these two graphs. One is the cost of downtime, and the cost of downtime is measured in how much revenue we lost because customers are leaving our services, and the cost of mitigation. Now let's take any two points here. Let's say we take A, where we're going to not do any work. Well, what happens is that when we don't do any work, we trace this over, and now the cost comes into play where we're losing customers and we're losing revenue. But there's also that cost of overreacting. That is, let's say we really get this downtime to be very minimal when we spend a lot of money doing it, and here what we have now is we spent a ton of money that really was unnecessary at this point. So what we want to do is we want to find a balance between these two.

To understand what the right balance is, we really have to understand what this chart is telling us. So let's say we spend no money here, so we've really spent no or very little money in creating high availability services that are up and running most of the time for our customers. So because of that we've spent very little money on mitigation, but we can see that no one's going to really like our services, so we're going to have a huge amount of lost revenue here.

But now let's invest just a little bit. So let's say we are investing this amount of money, which would be from the end of the chart to right here, so this is how much we're investing into this. So we're investing that into lessening our downtime, but it's made a huge difference in how much lost revenue we have. So the lost revenue has decreased a ton, and so we've really successfully made some huge changes here. So really what happens is that we follow this arch right here initially in driving our downtime down and this cost mitigation.

But what happens on the other side of the chart? So let's take a look at that on this other side of this chart. So let's say we invest, we're going to say we invest let's say this amount of money right here in mitigating our services, so we've got that right there, and we've got this amount of lost revenue right here. So we've really driven this down a lot, we've got very little lost revenue in this scenario.

But then let's say we spend even more to make these servers really have very little downtime at all. The downtime is so small here, but we spent a ton of money on mitigating costs here, and really we're not losing hardly any customers here. So we see some benefits to it, but look at how much benefit we received from it, this little area right in here. Not that much. This is very little benefit for the extra expense that we had here. So what we see is on this side of the chart it follows this trend line right here.

The Magic Point

So the magic point here, where we spend the least amount of money, is right here. In this scenario we spent this amount of money in mitigating the cost, so we have this amount of downtime and we also have this amount of lost revenue. So we have kind of equal amounts in this scenario right here. So this is the magic point in which we balance out between these two.

About TechKnowSurge

TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.

Explore free tools and programs →