Network monitoring and alerting are complementary practices that allow IT teams to detect and respond to system issues before they escalate into critical failures. Effective alert systems use tiered notification methods—dashboards, emails, texts, and calls—calibrated to severity to minimize downtime and reduce operational risk.
Network Monitoring & Alerting
When we're monitoring our network, sometimes we want to be alerted right away when something happens. There is some sort of alerting or notification that happens.
Monitoring just means observe or check the progress of something, so we're checking in on different aspects of our network. Alerting is warning of a danger or threat. So we really need to be implementing both monitoring and alerting. If we have a machine right here and we want to make sure that we understand what's happening on it, we need to do monitoring. Then, if it were to be compromised, we need some sort of alert process so the proper person can go in and research and figure out what's going on with that machine and why something has been alerted on that machine.
Let's create a little scenario here. Let's say we're IT personnel at a SaaS company and we've got a web application. Let's take a look at the timeline of events of an issue we have on that application.
Let's say initially it starts out and it's functioning fine, but then it goes through this period of time where things are not functioning quite as designed, and we see some issues cropping up. Finally, they hit a tipping point and our system goes down. Now the system is just flat out down. We start getting customers that are responding to us, so customer support is receiving these. They contact the IT team, so they contact us, and this is when we get notified right here. So we get notified, and then what happens is we go through our troubleshooting process and start trying to fix things. Finally we fix things, and now we are back up and running.
So what we want to do is minimize the time it takes to troubleshoot right here, but we also want to minimize the time it takes for us to get notification. We want to get notification sooner. So maybe we're not relying on customer support, but maybe we have systems that are monitoring to notify us when things go down, or even better yet, to notify us when things start looking like the system is not operating quite as designed.
Now, if you look at how we shift all of this, this troubleshooting period that we have right here now shifts all back, and we can have it maybe even fixed, hopefully even fixed, before this system goes down altogether. So by getting notified early, by identifying issues before they become a serious issue, it can really help us to reduce the impact that this has, or maybe even get rid of this incident altogether.
Remember, mitigating risk is about reducing the likelihood, and if we get notified that things are going to arise sooner rather than later, maybe we fix it before it goes down, or reducing the impact — and part of that is the duration.
Any of the systems that we're doing monitoring on, we're going to also probably want to do some sort of notification. Now, depending on what system it is, maybe we want things to really alert us and say, hey, this is going to be a critical issue, we really need to look into it. Other things maybe are not as high of a concern. So depending on what piece of equipment it is, we may want different levels of notification. One of the things that we'll be looking for is those indicators of compromise, so we can start researching and looking into why systems are going awry.
How are we going to get notified, how are we going to receive alerts? Well, one way it might be is dashboards. One of the things that I did with one of my companies is I put monitors all along the IT department so we could monitor exactly what was happening and see if things were ever going off, and be able to see live stats. So we could have dashboards that display a lot of different counters and a lot of different actions, maybe our network map and what's happening within our network, and what's happening within our event logs.
Then, if it ever escalates to a point where we need to start looking into something, then maybe they get sent out an email. That email is something that's a little more active, that maybe if somebody's at home or somebody is out and about, they receive an email and it says, oh yeah, you probably need to look into this a little sooner rather than later. Then maybe it gets escalated, and now we really have an incident and we want a higher level of notification. That's where text messages come into place, or even a phone call. I've had systems that will give text messages and phone calls to the person on call, so that way they can tackle the issue.
Depending on how we want to be notified depends on the severity and the priority, the priority of the equipment, the priority of whatever it is that we're monitoring. So we're going to determine at what point in time we just get an email letting us know, hey, you should look into this — maybe we don't raise it up to the highest priority at that point, but maybe it's just to a level of priority that when we get around to it we need to start looking into it. Then maybe it gets raised a little bit more when we say things really need to be looked at because we're really seeing some problems. And then a point in time where you drop everything and go and look into what's going on. So we're going to have to determine these thresholds and how we're going to get notified along the way.
So in the example here of notifications, I always use the dashboards as just being something to keep a pulse on what's happening, email is the next level up, and then text messages and calls where you need to drop everything and start researching something.
There is some tuning that needs to happen, because there are times we will get false positives. False positives means it alerts us, but then there is really nothing actually wrong. Or we could get false negatives, meaning that it doesn't alert us when it should have alerted us, when we should have looked into it.
So whenever we have an incident, I get the team together and discuss why didn't we get notified and how can we get notified sooner, so we could get rid of those false negatives, so we could be notified when there are actually issues. And also, having too much false positives can desensitize us to these different alerts. If we are desensitized, and it's just one more time that we're getting a false one — I'm sure it's nothing, we don't need to really worry about it — we get desensitized to that stuff, and we don't want that to happen either. So we need to make sure that our systems are notifying us when there are issues, but also not notifying us when there aren't issues.
There are a lot of different monitoring tools that we are using within our system, and each one of these could have a method of alerting us. In fact, most of these, like network monitoring systems, database monitoring, all of those have their own ways of sending out an email or sending out a text message or sending out maybe perhaps even a phone call. But they have different methods of raising the flag.
The problem is that these are a lot of different systems here, and sometimes that can be really cumbersome in managing. So there is alerting software that's out there that we can make these all report to. In fact, I used to use one up in the cloud, and these would all report to the one up in the cloud individually. Then it would be one system that would manage the on call. So we'd have a person that was on call, and then this system would know who is on call and start notifying them through many different methods. It would call them multiple times again and again, in case they were sleeping and needed to be woken up, or it would also text them as well. So within a 5-minute period it would call them like five different times, and if they didn't respond, then it would go to the secondary person on call. So there is alerting software out there that's separate from all the rest of these systems that can help us manage.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →