Establish who is affected, then work upward from the cable to the name. Link, address, gateway, internet by address, name resolution, application. The first step that fails names the problem, and everything above it fails too and tells you nothing.
- Step zero is how many people are affected, and it names the layer
- Works by address and fails by name is always DNS
- A 169.254 address means DHCP never answered
- Opens then hangs on large transfers is MTU, not bandwidth
- Rebooting first destroys the evidence you will need
On this page
- Step 0 of the checklist, which is not a test
- The six network troubleshooting steps, in order
- The four splits that do the most troubleshooting work
- Common network issues and the step that catches them
- When the checklist is finished: verify, document, escalate
- What each tool proves, and what it does not
- What to have in place before the next network issue
- Where people go wrong
- Comparison
- FAQ
Step zeroStep 0 of the checklist, which is not a test
A network troubleshooting checklist that starts with a command starts one step too late. Before touching one, get two facts. They cost thirty seconds and they routinely cut the search space for most network issues by ninety percent.
How many people are affected. One person, one department, one site, or everybody on the network. This single answer usually identifies the layer, before any troubleshooting tools come out.
| Who is affected | Where the problem almost certainly is |
|---|---|
| One person, one device | That device, its cable, its port, its configuration |
| Everyone on one switch | That switch, or its uplink |
| Everyone on one VLAN | The VLAN, its gateway, its DHCP scope or relay |
| Everyone at one site | The internet connection, the firewall, or the router |
| Everyone everywhere | A shared service: DNS, DHCP, the directory, the internet link |
| Everyone using one application | The application or its server, not the network |
What changed. A network that worked yesterday and does not today changed. A firewall rule, a switch configuration, a certificate that expired, a lease that ran out, a provider maintenance window. Asking is faster than deducing, and the answer is usually available for the asking.
That last row of the table is worth stating separately: a great many issues reported as network problems are one application being down, and establishing that everything else on the network works takes one minute and saves an afternoon.
The six stepsThe six network troubleshooting steps, in order
Each step in the checklist assumes the previous one passed. Stop at the first failure, because every step above it will fail too and testing further tells you nothing about the problem.
The order follows the OSI model from the bottom up. Step 1 is layer 1, the physical layer: cables, ports and hardware. Steps 2 to 4 cover the data link and network layers, where addressing, configuration and routing issues live. Steps 5 and 6 check the services and software on top.
1. Is the link up? Look at the port lights, on the device and on the switch. No light is a cable, a port, or a dead network card, and no amount of software troubleshooting will help.
A link at 10 or 100 megabits where it should be gigabit is usually a damaged cable, and it is worth noticing because it produces a performance problem rather than an outage.
On a wireless client there are no lights to look at, and the equivalent first question is whether it associated at all, which is the first of four steps on a WLAN and the only one of them that is radio.
2. Does the device have a real address? ipconfig on Windows, ip addr on Linux. A 169.254 address means DHCP never answered, which points at the VLAN, the scope or the relay rather than at the device. A correct address with a wrong subnet mask or gateway is a configuration problem that looks like a network routing problem.
3. Can it reach its gateway? Ping the default gateway. Success means the local network works and the problem is beyond it. Failure with a valid address means the VLAN, the switch, or the gateway itself.
4. Can it reach the internet by address? Ping something like 1.1.1.1. This is the step that splits any network issue in half. If an address works and a name does not, you have a DNS problem and can stop testing connectivity entirely.
5. Does a name resolve? nslookup a public name, then an internal one. Which of the two fails tells you whether the DNS resolver is unreachable, the internal zone is wrong, or the forwarder is broken. The DNS layers are worth walking in their own order when this is the step that failed.
6. Does the application work? Only now. If the first five steps of the checklist pass, the network is delivering packets and the problem is the service, its certificate, its authentication, or a firewall rule for that specific port.
link up? -> real address? -> gateway? -> 1.1.1.1? -> a name? -> the app?
| | | | | |
cable DHCP or VLAN switch or firewall or DNS the service
gateway the ISP
The four splitsThe four splits that do the most troubleshooting work
Underneath the troubleshooting checklist are four either/or questions. Each one helps identify the kind of problems you are dealing with, and halves what is left.
By address or by name. Works by IP and fails by name is DNS, always. This is the single most valuable test in network troubleshooting and it takes one command with tools every device already has.
One device or many. One device is a device problem. Many devices means a shared thing, and which devices are affected names which shared thing.
Small packets or large. A connection that opens and then hangs when data flows is MTU, not bandwidth, and no amount of capacity fixes it.
Slow or broken. Slow is a performance problem: capacity, latency or packet loss. Broken is usually a path or a rule. They are different investigations and treating a network performance issue as an outage wastes the afternoon.
Common issuesCommon network issues and the step that catches them
Most network issues that reach a help desk are one of a short list of problems. The table maps each one to its usual symptoms and to the step of the network troubleshooting checklist where it shows up.
| Common issue | What users report | Where the checklist catches it |
|---|---|---|
| Damaged cable or failed hardware | One device offline, or a link running slow | Step 1, port lights and link speed |
| DHCP scope exhausted | New devices cannot connect, old ones are fine | Step 2, a 169.254 address |
| Duplicate IP address | One device drops on and off the network | Step 2, an address conflict warning in the system log |
| Wrong VLAN or switch configuration | A group of devices lost access after a change | Step 3, no reply from the gateway |
| Firewall rule or ISP outage | The whole site has no internet | Step 4, no reply from a public address |
| DNS failure | "The internet is down" while ping by address works | Step 5, names do not resolve |
| Weak wireless signal or interference | Wi-Fi connectivity drops in one area | Step 1, association and signal strength |
| Bandwidth saturation | Everything is slow at certain times of day | The slow or broken split, then monitoring data |
| Security software on the device | One application blocked on one machine | Step 6, the host firewall or endpoint agent |
Two checks sit beside the six steps and often save time. Read the logs on the devices in the path, meaning the switch, the firewall and the system event log on the affected machine, because a port flapping or a blocked connection is usually recorded. And compare with a device that works: same VLAN, same software, different result.
After the fixWhen the checklist is finished: verify, document, escalate
Finding the failed step is not the end of network troubleshooting. Three short tasks close it out.
Verify the fix from the affected devices. Run the failed test again from the machine that reported the problem, then ask the user to check the application they actually needed. A fix verified only from the network team's own laptop is not verified.
Document what happened. Record the time, the scope, the step that failed, the cause and the change that resolved it. That record is the data the next person will search when similar issues come back.
Escalate with evidence. If the checklist points outside your network, call the ISP or the vendor with facts: when it started, which sites are affected, the traceroute output and what the monitoring tools show. A first call that already rules out your own hardware and configuration skips the provider's basic questions.
The toolsWhat each tool proves, and what it does not
Five network troubleshooting tools cover the checklist, and the useful thing to know about each is the limit of what a passing result actually establishes.
| Tool | A pass proves | A pass does not prove |
|---|---|---|
| Port lights | The physical link exists | That it is running at the right speed |
| ipconfig or ip addr | The device has an address | That the address is correct for this VLAN |
| ping to the gateway | The local network delivers packets | Anything at all beyond the gateway |
| ping to a public address | Routing and the internet path work | That names resolve |
| nslookup | The resolver answered | That the answer is the right one |
| traceroute | Where the path stops | Which hop is at fault, on its own |
Two of those rows deserve emphasis, because they are where confident wrong conclusions come from.
A successful nslookup does not mean DNS is correct. It means a resolver replied. An internal name resolving to a public address is a successful lookup and a broken result, and it is the specific failure that split DNS produces when a device is on the wrong resolver.
A traceroute shows where packets stop and not why. Routers generate replies at low priority, so the hop that looks worst is frequently the busiest rather than the broken one. Only behavior that persists through every subsequent hop is a real finding.
Beyond the command line, three kinds of network troubleshooting tools settle the problems the checklist cannot.
Network monitoring software collects SNMP and flow data, which is how you identify performance issues that come and go over time. A packet analyzer such as Wireshark shows what is actually on the wire. Configuration management tools keep a backup of every device configuration and show what changed.
BeforehandWhat to have in place before the next network issue
How long this checklist takes depends almost entirely on things decided months earlier. Four of them do most of the work.
A current network diagram. Which switch feeds which floor, which VLAN is which, where the gateway for each one lives. Without it, the scope question cannot be answered quickly, and the scope question is step zero.
Monitoring that watches the right things. Not a dashboard nobody opens: network monitoring tools with alerts on the internet link, the DHCP scope filling, the switch uplinks, and the DNS resolvers. Monitoring that tells you before the first ticket turns an outage into a maintenance task.
A record of what changed. A change log, however informal, answers the second question of step zero without anybody having to remember. Most network issues are downstream of a change somebody made, and most of those changes are findable if they were written down.
Somewhere the fixes are written. The same faults recur. A page of the last twenty resolutions is worth more than any monitoring tool, and it costs a sentence each time.
PitfallsWhere people go wrong
Rebooting first. It sometimes works and it destroys the evidence, so if the network issue returns you are starting from nothing. Get one fact before restarting anything.
Skipping the scope question. Ten minutes of troubleshooting one laptop before discovering the whole floor is down is the commonest way to waste a morning.
Testing by name when the question is connectivity. A failed name lookup proves nothing about whether packets can reach anywhere. Test by address first.
Trusting a short ping test. Ten packets says nothing about intermittent loss, which is what most network problems reported as slowness actually are. Run it for minutes.
Blaming the slow hop in a traceroute. Routers answer at low priority, so a middle hop showing high latency with normal hops after it is behaving correctly.
Treating an application outage as a network issue. If everything else works from the same device, the network delivered the packets and the service did not answer.
Changing two things at once. Then the network works and nobody knows why, which means the issue will happen again and the notes will be useless.
Not writing down what fixed it. The same network problems recur, and the second occurrence costs as much as the first because the answer left with whoever found it.
ComparisonThree different jobs, and why the scope question comes first
| Criterion | One machine | One VLAN or switch | The whole site |
|---|---|---|---|
| Check the cable and port | Yes, first | No | No |
| Check DHCP scope and relay | Rarely | Yes | Sometimes |
| Check the switch uplink | No | Yes | Sometimes |
| Check the firewall or router | No | Sometimes | Yes |
| Check the internet connection | No | No | Yes |
| Ask what changed | Always | Always | Always |
| Likely time to resolve | Minutes | Under an hour | Depends on the provider |
The bottom row is the reason the scope question comes first. The three columns are different troubleshooting jobs, and starting the wrong one is the expensive mistake rather than any individual test.
FAQFrequently asked questions
What is the first step in network troubleshooting?
Establishing how many people are affected and what changed. Both take under a minute and usually name the layer before any troubleshooting tools come out.
What order should the troubleshooting steps go in?
Upward from the physical layer: link, address, gateway, internet by address, name resolution, then the application. Stop at the first step that fails.
How do I tell a DNS problem from a connectivity problem?
Ping something by address, then by name. If the address works and the name does not, it is DNS, and no further connectivity testing is needed.
What does a 169.254 address tell me?
That DHCP did not answer. The fault is somewhere between the machine and the DHCP server, which usually means the VLAN, the scope or the relay.
Should I reboot first?
No. It sometimes works and it removes the evidence, so a returning fault has to be diagnosed from scratch. Get one fact first.
How long should I ping for?
Minutes, not seconds. A ten packet test cannot see the intermittent loss that most complaints are actually about.
One hop in my traceroute is slow. Is that the problem?
Almost certainly not. Routers generate replies at low priority, so a slow hop with normal hops after it is doing its job. Only delay that persists to the end is real.
The connection opens and then hangs. What is that?
Usually MTU. Small packets fit and large ones do not, so the handshake succeeds and the first real transfer stops.
The network is slow rather than broken. Where do I start?
With whether it is latency, packet loss or capacity, because those are three different performance investigations and the tests distinguish them quickly.
What network troubleshooting tools do I actually need?
ping, traceroute, nslookup, ipconfig or ip addr, and access to the switch. A packet capture when those six steps have not settled it.
When is it not a network problem?
When every one of the six steps passes from the affected device. At that point the network is delivering packets and the service is not answering.
Why write down the fix?
Because the same fault recurs, and the second time costs as much as the first if the answer left with whoever found it.
What is the single most useful test?
Pinging by address and then by name. One command splits every problem into connectivity or resolution, and those go in opposite directions.
How do I run a quick DNS check?
Run nslookup or dig against a name you know, first with your configured server and then with a public resolver such as 1.1.1.1. If the public resolver answers and yours does not, the fault is your DNS server. If an IP address pings but the name does not resolve, the DNS check has found the problem.
Keep readingRelated concepts
Read next · Diagnostics Ping and Traceroute Steps three and four of the checklist are these two commands, run properly rather than briefly. Open this next11 min- DNS · 14 min DNS Not Resolving When step five is the one that failed, DNS has an ordered checklist of its own.
- Diagnostics · 11 min Packet Loss vs Latency, and Why They Feel the Same to Users Slow and broken are different investigations, and this is the one for slow.
- Diagnostics · 11 min CRC Errors, and the Four Moves That Find the Cause Where this check sits in the wider order.
- Diagnostics · 13 min ipconfig and ip, and What the Output Is Telling You Where to go after the address reading.
- Wireless · 9 min A WLAN Is a LAN With a Radio in Front of It The wireless branch in detail, including the three steps that happen after the radio.
- Network services · 9 min NTP Stratum, How Far a Server Is From the Clock Time sync as one of the things to verify when things break.
- Network fundamentals · 10 min Layer 1 Networking, the Physical Layer, and Why It Is Where Most Faults Start Why confirming the physical path first saves the most time.