System monitoring and alerting are essential practices for maintaining infrastructure reliability and protecting organizational reputation. Effective monitoring requires continuous refinement to minimize false positives and false negatives while ensuring critical issues are caught before they escalate.
Monitoring & Alerting Systems
System issues can really attack our reputation. That is, if we are managing business systems for our end users or our customers, and those systems are going offline, or there's some sort of issue that impedes their work or causes problems for our customers and our users, it is really damaging to our reputation, especially if it's reoccurring and happening quite a bit. So it's really important that we set up proper monitoring.
Monitoring is our ability to take a look at our infrastructure and see what's happening on it: see what systems are being used, what resources are being used, what applications are up and running, what infrastructure equipment is up and running or goes offline or is being over-utilized. So it's really important that we have monitoring of these different systems.
Alerting is that capability, when things get too far out of scope of where it needs to be at, that there is some sort of alert that tells us something is happening. This could be a dashboard that we look on on a regular basis, it could be an email, a text message, a phone call, something along those lines that lets us know that we need to look into our systems because something is off.
There is so much that we could be monitoring, and it becomes a question of what should we be monitoring, what are the most critical items to monitor. The really important part to this is to continually improve, continually adjust it. That is, we don't want too many false positives.
What a false positive is, is if we get alerted, let's say I'm getting text messages or phone calls that there is some sort of issue, and then I look into it and there's not an issue. At some point in time I'm going to start ignoring those, so false positives can be really dangerous. A false negative means that I don't get alerted when I should be alerted, and so if there's an incident that happens I can look at it and say, okay, what do I need to monitor to help alert me before this becomes an actual issue?
Therefore we need to create this process of identifying what we need to monitor. We go into configuring those monitors to monitor it, we do the monitoring and make sure that it's accurate, and assess: is this something that is effective for us or not effective for us?
I've even gotten to the point where I scheduled a weekly meeting with my team and team members from other teams, when we got together and discussed, are we monitoring the right aspects of our environment? We took a look at all the alerts that we got and said, are these false positives, and do we need to dial things back? And we took a look at different aspects of our environment and said, what should we have been alerted on?
The result was there was a lot less noise with our alerting, as well as we became much more responsive when there was issues. We actually reduced the amount of issues that we had within our environment, because we fixed a lot of issues that was making a lot of noise. So there was a huge amount of improvement with our reaction time.
TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.
Explore free tools and programs →