Recovery Time Objective (RTO) and Recovery Point Objective (RPO) are the two core metrics that define how quickly an organization must restore operations and how much data loss is acceptable following a disaster. Together, they drive all disaster recovery planning decisions, from infrastructure investment to staff training requirements.
RPO and RTO Objectives
How much time, money and effort we put into our disaster recovery is going to depend on what our objectives are. We have two objectives that we may want to accomplish: a recovery point objective and a recovery time objective, an RPO and an RTO.
Here's a timeline, and on this timeline we have some sort of incident that's occurred. We're going to say that the incident that occurred was a fire in our data center. It destroyed some equipment, it destroyed some data, and now we need to recover from that. We may have some sort of objective — that's what the O is, objective, what we want to accomplish. We may tell our customers that this is what we're going to accomplish or agree to; we may even have it in our service level agreements, and so we're held accountable to these objectives.
So what is our objective? One of them is our recovery time. The recovery time is how long it's going to take us to recover from this incident. We may say, if we had a fire in our data center, we could be up and running in 24 hours. Now, that's a long time for some customers, and other customers might be just fine with it — it's not a big deal, I'll wait 24 hours. But for big companies this would be unacceptable, so maybe they narrow it down to being up and running in an hour, or maybe it's within minutes. So whatever you have here, you're going to establish some sort of time to recover. It could be even for8 hours or even longer, and it might be different for different services that you have.
Then you have the recovery point objective: that is how much are you going to lose. We had a destruction with our data, so now we have to go pull from backup. Are we willing to lose 15 minutes? So maybe 15 minutes is what we guarantee our customers, that if something were to happen. Maybe it's an hour, maybe it's a day — that would be quite a bit for a lot of customers, to lose a whole day's worth of data. In fact, 15 minutes for some may be a long time; you may want it all to be real time and not lose any data. A bank can't lose anything, and so it needs to be all real time.
So this is the recovery point objective. The recovery point objective goes backwards — how far back can we go from restore points — and the recovery time objective is how long can we be down and still be okay.
There are some considerations when we're weighing this out. Ideally what we'd say is our recovery time objective is going to be less than a minute, we pretty much want to be up right away, and the recovery point objective is going to be less than a minute, we can't lose any data. That's the ideal, and there are some systems out there that require this. However, this is very costly. You have to have a hot site that the data is constantly being replicated to and you can pull that up at any point in time.
What we need to do is balance out what our customers are demanding, what our board of directors is demanding, what the execs are demanding, what the demand on us is, and what the cost involved with this is. We have to weigh those out to determine what this is going to be.
Obviously this is going to be expensive, and the further back we go with the recovery point objective — if we're willing to lose a day's worth of work, 24 hours' worth of data — then that's going to be a lot less expensive than trying to back up a lot more often. Now, the cost for data storage is becoming fairly minimal, so we can actually get the recovery point objective a lot closer to the incident than further away.
Versus the recovery time objective: if you're going to operate in 48 hours, you might be able to just have a plan to say we're not going to have any servers up and running, if something happens we have everything in the cloud and we're going to spin up all new equipment, and in 48 hours then we'll be up and running. That can be very cheap. Then if we get closer and closer to here — let's say we want everything to be in 1 hour — then we're going to have to have machines already set up, the data ready to go. We're going to have to have a hot site that we just need to go over to and flip some switches and go through some processes to bring that live, and so that is going to be more costly.
Not only that, but our training has to go up. We have to devote more time towards training, which takes away from other duties that we have and other services that we need to be doing as job functions. So there is definitely a cost, not just from a hardware and resource standpoint, but also from an employee standpoint: that time is going to have to be devoted towards that training.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →