TechKnowSurge
NIST 800-53 CP-4 ISC2 CISSP 7.12 CompTIA Security+ 3.4 NIST NICE K0709
VideoSecurityFree

Testing Plans

Disaster Recovery plans must be regularly tested to verify their effectiveness, and several testing methodologies exist to balance thoroughness with operational risk. From tabletop exercises to live failovers, each approach offers different tradeoffs between accuracy and potential customer impact.

Complete this video to capture a CTF flag worth 1 point.

About this video

Testing a Disaster Recovery plan is not optional — without it, there is no reliable way to know whether the plan will work when it is actually needed. A common scenario involves web servers and high-availability database servers that must fail over to a secondary site, and verifying that this failover happens correctly and within defined recovery time objectives requires deliberate, structured testing. Different testing approaches exist on a spectrum from low-disruption to high-fidelity, and choosing the right one depends on acceptable risk, available infrastructure, and organizational readiness. A tabletop exercise is the most accessible starting point, bringing the team together to walk through documented procedures step by step, identify gaps, and make corrections without touching any live systems. Simulated testing goes further by executing procedures in a staging or virtual environment that mirrors production as closely as possible, providing hands-on experience with the actual processes and checklists even if the environment is not a perfect replica. Parallel processing testing isolates the backup environment from the live system entirely, allowing a full technical test of the recovery site without redirecting production traffic or exposing customers to risk. At the far end of the spectrum, a live failover moves actual customer traffic to the backup site, which is the most definitive test of plan viability but also the most consequential if something goes wrong. Even tests that do not touch production systems can introduce unexpected dependencies or partial outages, which underscores the importance of communicating planned tests to all stakeholders who could be affected. Each completed test should be treated as a learning opportunity, with findings used to continuously refine procedures and strengthen the overall plan.

What you'll learn

What's covered

DR Plan Testing

Aligned to

NIST 800-53
CP-4 Contingency Plan Testing
ISC2 CISSP
7.12 Test Disaster Recovery Plans (DRP)
CompTIA Security+
3.4 Explain the importance of resilience and recovery in security architecture.
NIST NICE
K0709 Knowledge of business continuity and disaster recovery (BCDR) policies and procedures

Key terms

Disaster Recovery
DR
The process and procedures for recovering IT systems and data following a disruptive event.
Failover
The automatic switching to a redundant system or component when the primary one fails.
Recovery Time Objective
RTO
The maximum acceptable time to restore a system or service after a disruption.
Tabletop Exercise
A discussion-based DR testing method where participants walk through a disaster scenario step by step to identify gaps in the plan without disrupting live systems.
Simulated Environment
A DR testing approach that uses a virtual or staging environment to rehearse failover procedures without impacting production systems.
Parallel Processing
A DR testing method where a backup copy of systems is brought up and tested independently while the live production environment continues serving customers.

Topics

Disaster Recovery Dr Testing Tabletop Exercises Failover Business Continuity Resilience Planning

Transcript

One thing we're going to want to do is test out our disaster recovery plan. Without testing it out, you have no idea if it's flawed or not, and there are several different ways that we could go about testing it out.

Why Test the Plan

Within our DR plan we are going to have processes that we're going to carry out to carry out our disaster recovery plan, but without testing this out we really don't know how effective it is.

So let's develop a little scenario here. Let's say we have a bank of web servers, we have a couple of database servers that are in high availability, and what we want to do is we want to test out our disaster recovery site and make sure that things can fail over correctly and that there's no issues.

Live Failover

The most basic form of this is just do a live failover. That is, all of our customers are on our live site and we bring it down and move them over to our backup site. So this is a true test to make sure that we can comply with everything that we need to comply with. But the problem with this methodology is our customers could be down, and if anything goes wrong and we go outside our recovery time objective, then this is really problematic for our customers. So this might not be the best solution many times. Although it's great from a making sure that everything fails over correctly standpoint, it could really have a negative impact and really be problematic to us too. It could be very painful.

Walkthrough or Tabletop Exercise

So what are some different ways that we could do this? Well, one, we could just sit down and run through our process, and this is like a walkthrough or a tabletop exercise. We call it a tabletop exercise because we're just bringing out the instructions with the team and then we're going through it step by step and just seeing what it looks like and seeing how it is, and making little corrections and correcting things as we go.

Simulated Exercise

We also have some sort of simulated exercise that we can do. Maybe we have a whole virtual environment and we can test it out in our virtual environment. Or I've had a multiple stage, like a stage environment before production, so we roll things out to stage and then production. Well, I would test out failing over stage, and it wouldn't be perfect, it wouldn't be exactly like production was, but it would give us that experience and make sure that our checklist and everything that we would go through, our processes, were accurate and were going to be effective.

Parallel Processing

There's also parallel processing. Maybe we have that live system out there that our customers are getting into, but we have our backup copy. Maybe we sever this connection between these two, bring this copy up, test it out without actually failing over our customers to it. So we've got some parallel processing going on.

Communicate the Test

Now there have been times that I've been failing over the non-production equipment, our staging equipment, and realized that there was a dependency with our live copy and we needed to fix that, but that did cause some outage and so there was some concerns over that. So even if I'm doing one of these other environments, no matter what environment I'm doing, with the exception of the tabletop here, but whatever environment I'm doing, I want to make sure I communicate these tests out to make sure whoever is affected with these tests, or could possibly be affected with these tests, understands what's happening.

As we go through these processes and test out the plan, we're going to learn a lot of things, so we're going to make adjustments to these processes and plans to really make sure that we're continually improving it and making it better.

About TechKnowSurge

TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.

Explore free tools and programs →