Network management software watches the infrastructure that carries the traffic: the switches, routers, firewalls, access points and the links between them. Almost all of it is sold on the first of three questions, and the first question is the cheap one.
Is it up comes from polling every device on an interval. Why is it slow needs different data entirely, because a device can report itself perfectly healthy while the path through it is congested.
What changed needs a third source, a stored copy of every configuration and something that compares them. A product answering only the first will tell you a switch is up on the morning somebody quietly changed a VLAN on it.
- Is it up comes from polling on an interval, usually over SNMP
- Why is it slow needs performance and flow data, not device counters
- What changed needs configuration backup, and it is usually off
- Passive monitoring reports a quiet network as a healthy one
- It is not RMM: these devices will never run an agent
On this page
- Three questions, and three different kinds of data
- What network performance actually measures
- What is actually in the box
- The five functions of network management software
- The alert nobody reads
- On premises, in the cloud, or as a managed service
- Buying it, or building it
- Where people go wrong
- Comparison
- FAQ
Three questionsThree questions, and three different kinds of data
Almost every buying mistake with network management software comes from treating these three as one product because they arrive in one console. Network monitoring tools answer the first, and management of a whole network needs all three.
Is it up. A poller asks every device it is set to monitor, on a schedule, for a set of counters: interface state, throughput, errors, CPU, memory, temperature, and whatever else the device exposes.
This is SNMP in most cases, and it produces the graphs. It is the cheapest question to answer and it is what the category is sold on.
Why is it slow. Device counters cannot answer this, because a switch reporting an interface at 40 percent utilization and no errors is telling you the truth and telling you nothing.
Answering it needs network performance data and a record of what the traffic actually was: which conversations, between which hosts, how much of the link each one took. That is flow data, exported by the devices themselves, and it is a separate feature with a separate cost in storage.
What changed. The most useful question and the one least often bought. It needs the software to log in to every device it monitors on a schedule, retrieve the running configuration, store it, and show a difference against yesterday. Without it, network issues that started at four in the afternoon are investigated by asking people what they remember doing.
| The question | The data it needs | What it costs |
|---|---|---|
| Is it up | Counters polled on an interval | Cheap, and every product does it |
| Why is it slow | Performance measurements, and flow records from the devices | Storage, and it grows with the traffic |
| What changed | A stored copy of every device configuration | Almost nothing, and it is rarely enabled |
The third row is worth arguing about with a supplier. It is the least expensive of the three to run and the one that shortens the most outages, and it is routinely left switched off because nothing breaks while it is off.
PerformanceWhat network performance actually measures
Slow is not a measurement, and network performance management is the part of network management that turns it into one. Five numbers do most of the work, and they are not interchangeable.
| Measure | What it is | Why it matters |
|---|---|---|
| Latency | The time for a packet to cross the path and back | Everything interactive feels this one directly |
| Jitter | The variation in that time between packets | Voice and video break on jitter long before latency |
| Packet loss | The proportion that never arrives | A little is invisible, and then it is catastrophic |
| Throughput | How much is actually moving | The number people mean when they say bandwidth |
| Utilization | How much of the link that fills | Averages hide the bursts that caused the complaint |
The bottom row is where five minute polling misleads, and it is a limit of the analysis rather than of the network. A link averaging 30 percent utilization across five minutes can be completely full for eight seconds of it, and the eight seconds is what the user noticed.
Packet loss and latency are the two that get confused with each other most, and they have different causes and different fixes.
There is also a split in how any of this gets seen, and it decides what a monitoring system can tell you at three in the morning.
Passive monitoring watches traffic that is already happening: interface counters, flow records, real user sessions. It gives accurate visibility into what people actually experienced, and it is silent about anything nobody is currently doing.
Active monitoring, also called synthetic, generates its own test traffic on a schedule and measures the result. It costs a little bandwidth and it answers the question passive monitoring cannot: whether the site to site link nobody is using right now is still working, and whether it was working an hour ago.
A network monitoring tool with only the passive half will report a quiet network as a healthy one. They are the same picture, and no amount of analysis separates them after the fact.
The partsWhat is actually in the box
Once the feature lists are stripped down, network management software is built from the same six parts, whichever product is being demonstrated.
Discovery. It finds the devices to monitor rather than being told about them, usually by sweeping address ranges and asking anything that answers what it is.
A topology map. Drawn from what the devices report about their neighbors. It is the visibility people picture when they buy this, it is useful when it is accurate, and it stops being accurate the first time somebody patches a cable without telling anybody.
Polling and storage. The monitoring interval decides what you can see. Five minute polling cannot show a thirty second outage, and no amount of dashboard makes it appear.
Alerting. Thresholds turned into notifications. This is where the product succeeds or fails in practice, and it has almost nothing to do with the software.
Configuration backup. Retrieval, storage and comparison, as above.
Reporting. Trend analysis of network capacity and performance over time, to help the business decide what to buy next year, and availability figures for whoever is owed a service level.
FCAPS modelThe five functions of network management software
The three questions are a buyer's view. The standard view is older and has five parts.
The ISO network management model, usually shortened to FCAPS, divides the work into fault, configuration, accounting, performance and security management. Most network management software is still organized along those lines, because the model describes what network teams need in order to manage networks of any size.
- Fault management. Detects, logs and helps resolve issues on a device or a link, from a failed power supply to a flapping interface. This is where alerts come from.
- Configuration management. Backs up device configurations, tracks changes, and pushes approved changes or firmware to many devices at once. Provisioning a new switch from a template belongs here too.
- Accounting management. Records who and what is using the network: bandwidth by application, by user, by site. In a business network this is usage reporting and capacity planning rather than billing.
- Performance management. Collects latency, loss, throughput and utilization over time, so network teams can monitor trends and fix a bottleneck before users notice it.
- Security management. Controls who can access the network devices themselves, keeps an audit trail of changes, and flags firmware with known vulnerabilities. It supports network security without replacing a firewall or endpoint detection.
Set against the three questions, fault management answers is it up, performance and accounting management answer why is it slow, and configuration management answers what changed. Security management runs across all three, because each function depends on control over who can log in to the devices.
AlertingThe alert nobody reads
This is the failure mode, and it is a management issue rather than a product one.
A network monitoring system installed with default thresholds generates alerts about conditions that are normal for your network, and buries the real issues among them.
Nobody has time to tune it in the first week, so the notifications get filtered to a folder. Six months later the system is technically working, every device is being monitored, and no human being has read an alert since spring.
The discipline that prevents it is unpopular and short. Every alert has an owner and an action, or it is not an alert.
If nobody would do anything at three in the morning about a router at 70 percent CPU, that is a graph rather than a notification. A monitoring tool with twelve alerts people act on beats one with four hundred that they mute.
The test is one question, asked of whoever is on call: what was the last alert you acted on. If the answer takes a while to arrive, the software is not the thing to change.
DeploymentOn premises, in the cloud, or as a managed service
Network management software is deployed in three ways, and for a small organization the choice matters more than the feature list does. Business networks now span offices, cloud services and remote workers, so where the console runs decides what it can reach.
On premises. The software runs on a server inside the network it manages. Data stays in the building and nothing depends on the internet link, which is why larger and regulated organizations still run it this way. Somebody has to patch, back up and size that server.
Cloud based. The vendor hosts the console, and a small collector at each site polls the local devices and sends the results up. There is no server to maintain, remote sites are easy to add, and network teams can manage every site from one console, from anywhere.
The trade is that the subscription never ends, and the monitoring data lives with the vendor. Equipment that is itself cloud managed ships with a console of this kind.
As a managed service. A managed service provider runs the platform for many clients and sells the outcome: monitoring, alert handling, configuration backup and a monthly report. For a business without a network engineer this is often the realistic route, since network management tools only help when somebody responds to them.
Whichever model is chosen, automation is where newer network management solutions differ. Automated discovery, scheduled configuration backup, bulk firmware updates and scripted responses to common issues take routine work off a small IT team.
Automation also adds risk. Software that can change a hundred devices at once can break them at once, so change control still applies.
Buy or buildBuying it, or building it
There are real open source network monitoring tools in this category, which is not true everywhere, and the choice is genuinely about labor rather than about capability.
Zabbix is published under the AGPL 3.0. LibreNMS describes itself as a community based, GPL licensed network monitoring system. Prometheus is Apache 2.0 and comes at the problem from the application and cloud metrics direction rather than the network device one.
None of them are cheap in the sense that matters. The license is free, and the cost moves to somebody installing it, upgrading it, tuning the thresholds, and being the person who knows how it works over time.
For an organization with that person already, this is the better trade. For one without, a commercial product is buying the availability of somebody else's engineering and support rather than the features.
The question to answer before either is which of the three questions you need answered, because it decides the shortlist more than any feature comparison will. That is the test behind the ranked comparison of ten tools, which also prices one 200 device network in each vendor's own unit.
PitfallsWhere people go wrong
Buying on the dashboard. Every product demonstrates well, and every dashboard promises visibility. The differences show up at two in the morning six months later, which no demonstration covers.
Monitoring everything. Every device polled, every counter stored, every threshold on. This is what produces the muted folder. Start with what somebody would act on and add the rest when it is asked for.
Assuming up means working. A device answering a poll is a device with a running management process. It says nothing about the performance of the traffic passing through it, or about whether that traffic is arriving at all.
Skipping configuration backup. It is nearly free, it is usually included, and it turns what changed from an argument into a diff.
Confusing it with RMM. They overlap in the console and not in what they can reach. Neither replaces the other.
Treating the topology map as documentation. It is a picture of what the devices currently report about themselves. Documentation is what somebody wrote down about why the network is that way.
ComparisonNetwork management, RMM and endpoint detection, side by side
| Criterion | Network management | RMM | Endpoint detection |
|---|---|---|---|
| What it watches | Switches, routers, firewalls, network links | Servers and workstations | What runs on a machine |
| Needs an agent installed | No | Yes | Yes |
| Works on a switch or a UPS | Yes | No | No |
| Answers why is it slow | With performance and flow data | No | No |
| Can change configuration | With config management | Yes, it is the point | Limited |
| Who buys it | Whoever owns the network | Whoever owns the endpoints | Whoever owns the security |
The second row is the whole distinction. RMM puts an agent on a machine and can therefore do almost anything to it. Network management software monitors network devices that will never run an agent, which is why the two categories exist separately and why most providers run both.
FAQFrequently asked questions
What is network management software?
Software that watches network infrastructure: switches, routers, firewalls, access points and links. It polls each device for its state, records the results over time, and raises alerts when something crosses a threshold.
What is the difference between network management and network monitoring?
Network monitoring is the watching part. Management adds acting on what is found, which in practice means configuration backup and change, so most network monitoring tools are one part of a management product.
What protocol does network management software use?
Mostly SNMP, on UDP 161 for polling and UDP 162 for traps. Flow export for traffic analysis, SSH for configuration retrieval, and API access to newer and cloud managed equipment fill in the rest.
Do I need an agent on each device?
No, and that is the point of the category. Switches, firewalls and power supplies will never run an agent, so the software monitors them with the protocols they already speak.
Is network management software the same as RMM?
No. RMM manages servers and workstations through an installed agent. Network management software monitors infrastructure that cannot run one. Most providers run both.
What is the best free network management software?
Zabbix, LibreNMS and Prometheus are the well known open source network monitoring tools, published under AGPL 3.0, a GPL license and Apache 2.0 respectively. Free of licensing cost, not free of the labor to run them.
How often should devices be polled?
Often enough to see what you need to see. A five minute monitoring interval cannot show a thirty second outage, and a one minute interval across thousands of devices is a storage and load decision rather than a setting.
Why does my monitoring say everything is fine when it is not?
Usually because it is answering only the first question. Devices report themselves healthy while a path through them is congested, which needs network performance data rather than device counters. A purely passive tool will also report a quiet network as a healthy one.
What is configuration backup and why does it matter?
The software logs in to each device it monitors on a schedule, stores the configuration and shows the difference from last time. It turns what changed into something you can read instead of something you ask people to recall.
Should I buy or use open source?
It depends on whether somebody will own it. Open source moves the cost from licensing to labor, which is a good trade when that person already exists and a poor one when the plan is that somebody will find time.
What should I ask a supplier before buying?
Which of the three questions each part of the quote answers, whether configuration backup is included and switched on, how flow data is stored and for how long, and what happens to alerting when nobody tunes it.
Does network monitoring improve security?
Indirectly. It is built to answer availability and performance questions, and the visibility its change history gives is genuinely useful during an incident. It is not a security product, and the difference between network security and cyber security is a separate question worth reading.
How do network management tools and network management solutions differ?
The terms overlap. Network management tools usually means single purpose programs such as a ping monitor or a configuration backup script. Network management solutions means a platform that combines monitoring, configuration, alerting and reporting. A network management application of either kind still depends on SNMP, flow data and device APIs underneath.
What do providers use for MSP network monitoring?
For MSP network monitoring, providers need a multi-tenant product that keeps each client separate, uses a probe at every site, and opens tickets in their PSA. Some use the network module of their RMM, and others add a dedicated network monitoring product alongside it.
Keep readingRelated concepts
Read next · Managed IT What RMM Is, and What the Agent Can Actually Do The other monitoring category, for the machines that can run an agent, and the reason most providers end up buying both. Open this next9 min- Protocols · 10 min What LLDP Is, and Why It Answers the Port Question Where the topology map comes from, since it is drawn from what each device reports about its own neighbors.
- Ports · 9 min SNMP Ports 161 and 162, and the Firewall Rule That Is Half Right The protocol almost all of the polling runs on, and the two ports it needs open in opposite directions.
- Managed IT · 9 min What an MSP PSA Actually Is, and Why Replacing One Is Harder Where its alerts should end up as tickets.
- Network operations · 11 min Network Automation, and Why the Interface Matters More Than the Tool Changing the devices a management platform watches, through structured interfaces.
- Managed IT · 9 min What a Managed ISP Sells, and How to Tell If It Is Worth It The service that buys the monitoring rather than running it.
- Operations · 9 min MTTR, and Why Two Teams Quoting the Same Number Disagree The metric its alerts end up inside.
- Infrastructure · 10 min Cloud Managed Networking, and What Happens When You Stop Paying The vendor dashboards that network management software sits alongside.
- Network operations · 10 min Bandwidth Management, the Four Levers and When to Pull Each Where the monitoring that has to come first actually runs.
- Network operations · 9 min Syslog, the Standard Way Devices Send Their Logs Where syslog messages become dashboards and alerts.
- Security operations · 8 min SOC vs NOC, the Two Operations Centers Compared Where the network operations half of the picture actually happens.
- Network operations · 9 min NetFlow, the Record of Every Conversation on the Network Where NetFlow data is turned into dashboards and alerts.