Debugging Network Connectivity Issues Without Touching the Server
Walk through DNS, process listening checks, and firewall verification step-by-step when a service stops responding from outside.
The message comes in at the worst possible time. Someone reports they cannot reach the service. You check monitoring. Everything looks green. The process is running, logs are clean, and nothing is obviously on fire. But something between that user and your server is failing quietly. The instinct is to restart the service, but restarting without understanding is a guess, not a fix. Here is the ordered sequence that actually narrows it down, starting from the outside and working inward.
Incident ChecklistWhen a service stops responding from outside, the failure is almost never where you first look. DNS can resolve but traffic never reach the port. The port can be open but nothing is listening. The process can be listening but the firewall drops packets before they arrive. Walk the layers in order: name resolution, then listening state, then firewall rules, then external reachability. Each step rules out a whole category of causes.
Why "Unreachable" Covers More Ground Than You Think
Network failures are frustrating because the same symptom, a browser timing out or a curl command hanging, can come from completely different root causes. A DNS failure looks identical to a firewall drop from the client's side. That is why ad-hoc troubleshooting wastes time. Without a fixed order of investigation, you bounce between guesses and extend the outage.
The core insight is that network connectivity is layered. A request has to clear each layer before the next one even becomes relevant. If DNS resolution fails, the TCP handshake never starts. If the TCP handshake never completes, the application layer is not involved at all. Work from the outermost layer inward, and you eliminate whole categories of problems with each confirmed step.
DNS Resolution: The Layer Nobody Checks First
The first question is always whether the hostname resolves to the correct address. This sounds basic, but stale DNS records, misconfigured TTL values, and split-horizon DNS setups catch experienced engineers off guard regularly. A deployment to a new IP with a short TTL can leave some users hitting the old server for hours.
From any machine with a terminal, run this:
dig yourdomain.com +short
If dig is not available, nslookup produces the same answer:
nslookup yourdomain.com
You are checking two things: whether a response comes back at all, and whether the IP it returns matches the server you think is hosting the service. A mismatch points directly to a stale or incorrect record. The way resolvers handle TTL and caching is defined by the domain name system protocol specification, which is why a bad record can survive for hours even after you fix it upstream. If DNS resolves correctly and the IP is right, move on.
Confirming the Process Is Actually Listening on the Right Port
Once DNS is ruled out, the next question is whether your service is actually bound to the port and network interface it should be. Processes crash and restart with changed configurations. A service that was listening on all interfaces might restart and accidentally bind only to localhost. This is one of the most common causes of "it works from the server but not from outside."
On the server, use ss or netstat to inspect the listening state:
ss -tlnp | grep 8080
The output tells you everything you need at a glance:
- Which address the process is bound to.
0.0.0.0means all interfaces;127.0.0.1means localhost only and is unreachable from outside. - Which port it is actually listening on, in case a config variable was silently overridden.
- The process ID that owns the socket, so you can confirm it is the right binary.
- Whether the socket is in LISTEN state, not in some half-open or TIME_WAIT condition.
If the bind address shows 127.0.0.1 when it should be 0.0.0.0, that is your bug. The service is running, the port is occupied, but it is rejecting all external traffic at the socket level. Fix the bind address in your application config and restart.
Ruling Out the Firewall Before Blaming the Application
This is where most teams lose significant time. A firewall that drops packets gives no feedback. The connection just times out. From the client side it looks exactly like the server is down. From the server side the process is healthy. Neither side logs an obvious error. Knowing which tool manages the host firewall is half the battle, because iptables, nftables, ufw, and firewalld can coexist on the same machine in ways that are not obvious.
Check the firewall state in this order:
- List the INPUT chain rules with
iptables -L INPUT -n -vand look for any rule that mentions your port. - Confirm whether an ACCEPT rule exists and that no DROP or REJECT rule appears above it. Rules evaluate top to bottom.
- If the host uses ufw, run
ufw status verbosefor a cleaner summary of allowed ports and source ranges. - If firewalld is active, run
firewall-cmd --list-allto see which ports are open in the active zone. - Check the cloud perimeter separately. AWS security groups, GCP firewall rules, and Azure network security groups all sit outside the OS and block traffic before it reaches iptables. Open the cloud console and verify the rule explicitly.
The combination of an OS-level allow and a cloud-level deny is common after infrastructure migrations. Both need to permit the traffic for the connection to succeed.
Testing External Reachability Without SSH Access to the Target
Sometimes you cannot get onto the target machine. It might belong to a client. The jump host might be unreachable. Or you want to confirm what the outside world actually sees before escalating to someone with access. Running a check from your own workstation is not always enough, because your IP might be allowlisted in ways a normal user is not.
This is where a browser-based port scanner earns its place in the workflow. Unlike running Nmap locally, it checks from an external IP, which reflects what your users actually experience. You enter the host and port and get back a clear result: open, closed, or filtered. That single data point often resolves the argument about whether the issue is inside or outside the network perimeter.
What each result means in practice:
- Open: The TCP handshake completed from outside. Something is accepting connections on that port. If users still report problems, the issue is at the application layer, not the network.
- Closed: The host responded with a RST packet. It is reachable, but nothing is listening on that port, or the application actively rejected the connection.
- Filtered: The connection timed out with no reply. A firewall or network device is silently dropping packets. This result, combined with the process showing as listening inside the server, confirms the block is at the perimeter.
What the Error Message Is Telling You at Each Layer
Error messages are diagnostic signals, not just noise to scroll past. Each type maps to a specific layer of the stack, and reading it correctly points you directly to the next command to run rather than making you repeat the entire sequence from scratch.
Connection Error Reference by Network Layer
| Error Message | Likely Layer | First Place to Check |
|---|---|---|
Name or service not known |
DNS | DNS records and resolver configuration |
Connection refused |
Transport (TCP) | Process listening state and bind address |
Connection timed out |
Network / Firewall | OS firewall rules and cloud security groups |
No route to host |
Routing | Route tables and network interface state |
SSL handshake failed |
Application (TLS) | Certificate validity and cipher negotiation |
A connection refused and a connection timed out look similar from a user's perspective but require completely different fixes. The first means something is there but rejecting you. The second means packets are disappearing before they arrive. Reading the error first saves you from checking the firewall when the real problem is that your process is bound to the wrong address.
What Stays With You After the Incident Closes
The process described here, DNS first, then listening state, then firewall, then external confirmation, takes under ten minutes once it becomes habit. The difficult part is that most of this happens reactively, under pressure, while someone is asking for updates in a shared channel.
The way to make the next incident faster is to document your baseline now, before anything breaks. Know what ports each service is supposed to listen on and which interface it binds to. Know which IP ranges are allowlisted in your cloud security groups. Keep a note of which firewall tool manages each host, because iptables and ufw coexisting on the same box is more common than it should be. When an incident hits, you spend zero time rediscovering that baseline and all your time comparing against it.
None of the steps in this article require elevated access, specialized software, or SSH to the target. Most of them run from any machine on the internet. The right sequence of questions replaces the need for physical access, and that is exactly the point. Debugging should follow the failure, not the infrastructure you happen to have keys to.