TechKnowSurge
CompTIA Security+ 4.4 NIST CSF DE.CM-09 NIST 800-53 SI-4 ISC2 CISSP 7.2 CompTIA CySA+ 1.5 NIST CSF ID.IM-03 NIST 800-53 CA-7
VideoSecurityFree

Monitoring and Alerting

System monitoring and alerting are essential practices for maintaining infrastructure reliability and protecting organizational reputation. Effective monitoring requires continuous refinement to minimize false positives and false negatives while ensuring critical issues are caught before they escalate.

Complete this video to capture a CTF flag worth 1 point.

About this video

Unreliable systems are a direct threat to organizational reputation, particularly when the affected infrastructure supports end users or external customers. Repeated outages or performance degradations erode trust quickly, making proactive infrastructure management a business-critical priority. Monitoring provides the visibility needed to understand what is running, what resources are being consumed, and where equipment or applications are approaching failure or over-utilization. Alerting extends that visibility by triggering notifications through dashboards, email, text, or phone when conditions drift outside acceptable parameters, giving teams the opportunity to act before problems fully materialize. One of the core challenges in building an effective monitoring program is determining what actually warrants attention. Organizations must prioritize the most critical components of their environment and configure alert thresholds deliberately. False positives—alerts that fire when no real problem exists—are particularly dangerous because repeated unnecessary notifications train teams to ignore warnings, potentially causing them to miss genuine incidents. False negatives carry the opposite risk, allowing real problems to go undetected until they cause significant disruption. Balancing these two failure modes requires ongoing evaluation rather than a one-time configuration. A continuous improvement cycle is the foundation of a mature monitoring practice. That cycle involves identifying what to monitor, configuring and deploying monitors, assessing their accuracy and relevance, and adjusting based on findings. Regular structured reviews—whether through scheduled team meetings or cross-functional discussions—allow organizations to examine alert patterns, eliminate noise, and surface gaps in coverage. Teams that commit to this process consistently see measurable outcomes: fewer unnecessary alerts, faster response times to real incidents, and a reduction in overall environmental issues as underlying problems are identified and resolved rather than repeatedly triggering alarms.

What you'll learn

What's covered

Monitoring & Alerting Systems

Aligned to

CompTIA Security+
4.4 Explain security alerting and monitoring concepts and tools.
NIST CSF
DE.CM-09 Computing hardware and software, runtime environments, and their data are monitored to find potentially adverse events.
ID.IM-03 Improvements are identified from execution of operational processes, procedures, and activities.
NIST 800-53
SI-4 System Monitoring
CA-7 Continuous Monitoring
ISC2 CISSP
7.2 Conduct logging and monitoring activities
CompTIA CySA+
1.5 Explain the importance of efficiency and process improvement in security operations.

Key terms

Availability
The assurance that systems and data are accessible and operational when needed by authorized users.
Infrastructure Monitoring
The continuous observation of IT infrastructure components—such as servers, applications, and network equipment—to track performance, resource usage, and operational status.
Alerting
The automated notification process that triggers when monitored systems deviate from expected thresholds, informing administrators of potential issues.
False Positive
An alert that fires when no actual issue exists, which over time can cause administrators to ignore notifications and reduce monitoring effectiveness.
False Negative
A failure to generate an alert when a real issue exists, leaving problems undetected and unaddressed.
Baseline
A documented set of minimum security standards or performance metrics used as a reference point.

Topics

Infrastructure Monitoring Alerting Systems False Positives False Negatives System Availability Site Reliability

Transcript

System issues can really attack our reputation. That is, if we are managing business systems for our end users or our customers, and those systems are going offline, or there's some sort of issue that impedes their work or causes problems for our customers and our users, it is really damaging to our reputation, especially if it's reoccurring and happening quite a bit. So it's really important that we set up proper monitoring.

Monitoring

Monitoring is our ability to take a look at our infrastructure and see what's happening on it: see what systems are being used, what resources are being used, what applications are up and running, what infrastructure equipment is up and running or goes offline or is being over-utilized. So it's really important that we have monitoring of these different systems.

Alerting

Alerting is that capability, when things get too far out of scope of where it needs to be at, that there is some sort of alert that tells us something is happening. This could be a dashboard that we look on on a regular basis, it could be an email, a text message, a phone call, something along those lines that lets us know that we need to look into our systems because something is off.

What should we monitor

There is so much that we could be monitoring, and it becomes a question of what should we be monitoring, what are the most critical items to monitor. The really important part to this is to continually improve, continually adjust it. That is, we don't want too many false positives.

What a false positive is, is if we get alerted, let's say I'm getting text messages or phone calls that there is some sort of issue, and then I look into it and there's not an issue. At some point in time I'm going to start ignoring those, so false positives can be really dangerous. A false negative means that I don't get alerted when I should be alerted, and so if there's an incident that happens I can look at it and say, okay, what do I need to monitor to help alert me before this becomes an actual issue?

Therefore we need to create this process of identifying what we need to monitor. We go into configuring those monitors to monitor it, we do the monitoring and make sure that it's accurate, and assess: is this something that is effective for us or not effective for us?

Reviewing the alerts with the team

I've even gotten to the point where I scheduled a weekly meeting with my team and team members from other teams, when we got together and discussed, are we monitoring the right aspects of our environment? We took a look at all the alerts that we got and said, are these false positives, and do we need to dial things back? And we took a look at different aspects of our environment and said, what should we have been alerted on?

The result was there was a lot less noise with our alerting, as well as we became much more responsive when there was issues. We actually reduced the amount of issues that we had within our environment, because we fixed a lot of issues that was making a lot of noise. So there was a huge amount of improvement with our reaction time.

About TechKnowSurge

TechKnowSurge builds IT and cybersecurity professionals through hands-on, concept-first training built around real understanding — not memorization. Free interactive tools, structured programs, and 25+ years of real-world experience, all in one place.

Explore free tools and programs →