Effective incident response starts long before an incident occurs, requiring clearly defined procedures, assigned roles, and tested plans. This content covers the key elements of incident response preparation, including documentation, tooling, training, and automation with SOAR platforms.
Incident Response Preparation
Anytime I've encountered an incident on the systems I've managed, it's caused a lot of stress. There's always a lot going on. What is the problem? How do I fix the problem? Who do I need to communicate to? What steps do I need to follow? There's a lot that's happening, and one of the things I can do to reduce the amount of stress — I probably can't get rid of it altogether, but one of the things I can do to reduce the amount of stress in these situations and make sure that I can efficiently tackle this — is prior preparation. Preparation is a huge part of success when it comes to incident response.
Doing things ahead of time, preparing things ahead of time, can really set us up for success through this whole process.
The first thing that we're going to want to do is prepare by figuring out what is an incident. We're going to want to define what an incident is, because when we call this an incident, there's going to be a lot of things that we're going to have to do, there's a lot of work that has to be done. So do we call it when there's the first signs of malfunction? When the system alerts? When the system goes offline? When customers report the issue? Where do we actually say, this is an incident, I need to go into incident response? Then, of course, we'll have to figure out how do we respond to an incident, what are the steps that we're going to carry out during this incident.
When it comes to preparation, there's several things that we're going to want to look into:
We're going to want some people troubleshooting, so who's that going to be — technicians, database administrators, developers? We're also probably going to want to involve customers, customer support, those who are reaching out to our customers or our users. We're going to probably want some sort of incident response manager, somebody who's orchestrating everything, making sure that communication is happening, that people are in the right spots doing the right things. And then we might want to include maybe some legal, if there's some sort of legal issue that we're encountering, a release manager if we've got a release code, maybe there's a public relations officer. So we could have a lot of different roles and responsibilities that we would hand out.
Every time I've gotten into an incident, I'm like, "Oh shoot, what do I do now?" and I have to think about what is the first thing that I need to do. Well, what it should be is pull up the documentation that outlines exactly what needs to happen. We need to validate that this is truly an issue, we may need to escalate it, we may need to communicate it. Whatever your policies or your processes and procedures are, you set those up ahead of time, so that way, when it's in the heat of the moment, we just go through and we carry out this plan, we carry out these procedures.
There's also a lot of tools that can help us out through this process. Just to name a few: maybe it's some sort of documentation that we have, and making sure that the documentation isn't on some sort of system that could be broken. For instance, let's say the internet connection is part of the things that can go wrong, that can cause an incident within our business, within our organization — and then are we accessing documentation on the cloud? How are we going to access that documentation? We also have to think about the software that we would use for doing troubleshooting, or the software we'd use to manage these incidents. Do we have access to that software? Is the software in the right place? And similar for hardware and services.
Not only are we going to come up with our procedures and what is our step-by-step process, but there could be flaws in this process, and so what we'll want to do is test this out. Our procedures and our processes — do they function well, and can we follow those functions? Which goes hand in hand with training. Once we've established a great procedure and a process, do people know how to flow through this process? Well, one way we discover how we can flow through this process is through training, and making sure that people are trained and tested according to what our processes are.
A lot of our preparation comes in the form of a plan. That is, we have an incident response plan where we outline everything that we just talked about. There are also other preparations. We could prepare for a disaster recovery plan, or a DRP. We have a contingency plan. We have a business continuity plan. So these are some of the plans that we can come up with in order to prepare ourselves for incidents that come up.
I don't want to get too into this, but the difference between an incident response plan and some of these other plans is that an incident response plan has to do with IT issues that come up. And then, if it's extreme enough that maybe we have to move into a disaster recovery plan, doing recovery, that's where we're doing full system restores, or we're having to go into backup and do a full restore of backups. And then we also have a contingency plan and business continuity plan, which have more to do with the business side of things and business operations — how do we keep business operating despite some sort of issue that comes up? These have to do with contingencies that come up.
Another thing that we could do is automate some of these tasks. We've got validation that needs to happen, escalation that needs to happen, communication that needs to happen, troubleshooting, containment, eradication, and perhaps some of these steps we can identify as steps that could be automated somehow, that we could escalate and do this through some sort of automation. So there are components within what we do that could be automated.
Now, there is software that can help us automate some of these tasks — not only automate these tasks, but also set things up so we can really establish a great flow with all of this and be able to communicate well, troubleshoot well, be able to go through the flow of incident response. We call it a SOAR. SOAR stands for security orchestration, automation and response, and we can implement this software to help us go through this flow much quicker, much faster, and get over these incidents much easier.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →