Disaster recovery planning prepares organizations to restore systems and data following major incidents like ransomware attacks, natural disasters, or utility failures. This content covers DR plan components, roles and responsibilities, and the differences between hot, warm, and cold recovery sites.
Disaster Recovery Planning
There's going to be quite a few different events that could happen that would require us to restore from backup or fail over to a disaster recovery site — something more drastic than just a simple "hey, let's just flip the switch or change this configuration and fix whatever issue it is that broke." So for those cases we have disaster recovery, we have disaster recovery plans, we've got disaster recovery sites, but we need a plan for all of that ahead of time.
A disaster recovery plan helps guide us through recovering from major incidents. In this case right here, let's say we have some major system issues on these servers. What is it going to take to recover those servers? A lot of times we have to think about that beforehand, before there's an incident. If we think of it afterwards, these servers might have some custom data on them and things that we have not captured, and now it could be impossible to fully recover from that. So we really do need to think about this ahead of time, before there's an incident, before there's a disaster.
Disasters that we may see: maybe a fire in the data center, or a natural disaster like an earthquake or a volcano, maybe it's some sort of cyber crime, maybe there's some sort of utility disruption like there is no electricity.
Now, a disaster recovery plan is different than an incident response plan, or a contingency plan, or a business continuity plan. A contingency plan and business continuity plan are more from that business operations level. An incident response plan is at the IT level, the same as the disaster recovery plan, but it's really more to respond initially to an incident, and at some point in time, if that incident gets bad enough, then we go into disaster recovery mode.
So here we have an incident. Maybe it is that ransomware has taken over some of our data and encrypted it, and now we need to do a full system recovery from that. Initially a report comes in and we start addressing it, we start looking into it, we start calling people for troubleshooting, and then once we determine, oh yeah, what's happened here is ransomware has gotten hold of our data and is encrypting it — now what we need to do is go into disaster recovery, where we're going to recover from that. We'll have to completely eradicate this ransomware and then do full system restores of this, and so then we go into that disaster recovery plan and start doing those restores.
Disaster recovery plans can be very complex or very simple, depending on what type of infrastructure you have and how detailed you get with it. But some of the items that we're going to want to look at are who's responsible and what processes they're going to carry out, who's going to do the communication, how are we going to do the communication, who are we going to communicate to, what are the roles and responsibilities, what kind of training do we need to do ahead of time, and what kind of assessments to make sure that the DR plan is efficient and sufficient for what we need to do.
As part of the plan, you're going to have processes in place, you're going to have checklists, you're going to have details of how to do this disaster recovery. So if we're failing over to a DR site, then we're going to need to have that all detailed out of what those steps are.
We also need to think about redundancy. In this case right here, we've got redundancy within our database server, we've got redundancy within our application servers here. So we want to think about redundancy, but we also want to think about site redundancy. We have some sort of original copy, original site — maybe it's a local site, maybe it's in the cloud — but then we have something that we could fail over to, and we could have a hot, warm or cold site.
So this obviously has a longer recovery time, versus this is a very short recovery time.
I don't see this all too often, but there's even the mobile site, where you have the equipment installed on a truck that you can drive around, creating a really dynamic way of being able to have this equipment accessible from wherever you're at and be able to drive it around.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →