A request to https://api.example.com/v1/health fails from one Linux host. Walk through how you would isolate whether name resolution, the path, the TCP port, TLS, or the HTTP application is at fault, and say what each step rules out.
answer
- each layer depends on the one below
- start where a failure invalidates everything after
- the TCP handshake is the pivot
- let curl narrate TLS and HTTP
- compare from a second host
basics
~20 sWork outward one layer at a time: dig for the address, a path probe toward it, nc -zv for the TCP port, then curl -v to watch TLS and the HTTP response. Each step that succeeds removes a layer from suspicion, so the first failure localises the fault.
solid answer
~50 sThe value is in the order, because each step eliminates everything below it. Start with `dig +short api.example.com` — do you get an address at all, and is it the one you expect? Wrong address and everything after it is meaningless. Then probe the path toward that address with `mtr -rwc 20`, keeping in mind that ICMP results are advisory. Next test the port itself: `nc -zv <addr> 443` proves whether a TCP handshake completes, which separates a network or firewall problem from anything application-level. Then run `curl -v --max-time 5 https://api.example.com/v1/health`, whose verbose output narrates the remaining layers in order — the address chosen, the connect, the TLS handshake and certificate, then the status line and headers. Finally, repeat the whole drill from a second host and from the service's own machine against localhost, which tells you whether the fault is specific to this host or global.
code
bash · 3 linesdig +short api.example.com
nc -zv -w 3 api.example.com 443
curl -o /dev/null -s -w 'dns=%{time_namelookup} tcp=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total} code=%{http_code}\n' https://api.example.com/v1/healthgo deeper
Know the sequence and the tool for each rung: dig for the address, nc for the port, curl for TLS and HTTP. Say what a passing step rules out rather than just listing commands.
Explain why a completed TCP handshake is the pivotal result — it clears routing, firewalling and the listener in one shot — and read curl's verbose output as a stage-by-stage narration.
Demonstrate incident discipline: cap timeouts, carry the resolved address forward so later steps test the same target, compare against a second host and localhost, and use the curl timing fields to attribute latency instead of guessing.
Turn the drill into infrastructure: the checks worth baking into a runbook or a synthetic probe, what each layer's failure should page on, and how the same reasoning telescopes hop by hop through a chain of proxies and services.
## Why the order is the answer An interviewer asking this is not testing whether you know `curl`. They are testing whether you narrow a search space instead of guessing. Every layer depends on the one beneath it, so testing bottom-up means each success permanently removes suspects, and the first failure is the fault. Jumping straight to "the app is broken" when DNS is returning a decommissioned address wastes the outage. ## Step 1 — resolution: is there an address, and is it the right one? ```bash dig +short api.example.com ``` Three distinct outcomes. No output at all means no usable answer, and you are debugging naming, not the service. An address you do not recognise means you are about to test the wrong machine — chase the record, not the app. The expected address means resolution is clean and you can stop thinking about it. Carry that address forward explicitly for the rest of the drill so later steps cannot be confused by a different answer. ## Step 2 — the path: can packets get there at all? ```bash mtr -rwc 20 <addr> ``` This is advisory rather than decisive, because ICMP results are shaped by policy. Read the destination row, and treat loss at a middle hop as meaningful only if it persists through the later rows. A completely dead path is a strong signal; a clean one does not prove the port is reachable. Do not linger here — the port test is far more decisive. ## Step 3 — the port: does a TCP handshake complete? ```bash nc -zv <addr> 443 ``` This is the pivotal step. A completed handshake proves routing works in both directions, a firewall is permitting the port, and something is listening and accepting. Everything below the transport is now exonerated, and the fault must be in TLS or above. A failure means the opposite: stop looking at the application. The *way* it fails carries information too — an immediate rejection means you reached the host and nothing accepted, while a silent hang until timeout is the classic signature of a filtering rule that drops rather than rejects. Always pass a short timeout (`nc -w 3`) so a filtered port fails in seconds. On a stripped-down image without `nc`, `curl -v --max-time 3 telnet://<addr>:443` gives the same connect-only test. ## Step 4 — TLS and HTTP: let curl narrate the rest ```bash curl -v --max-time 5 https://api.example.com/v1/health ``` Verbose output walks the remaining layers in sequence: the address it chose, the TCP connect, the TLS handshake with the negotiated protocol and the certificate's subject and validity, then the request headers, the status line and the response headers. Wherever it stops is your layer. A certificate subject-name mismatch or an expired certificate is a TLS finding; a `502` or `503` is an application or upstream finding; a `404` means you reached the right server and asked for the wrong path. For latency rather than failure, replace the guesswork with measurement: ```bash curl -o /dev/null -s -w \ 'dns=%{time_namelookup} tcp=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total} code=%{http_code}\n' \ https://api.example.com/v1/health ``` Those fields split the elapsed time into resolution, handshake, TLS negotiation, server think-time and transfer. "It's slow" becomes "TLS takes 900 ms" or "time to first byte is 4 s" — an attributable number instead of an impression. ## Step 5 — is it me, or is it everyone? Re-run the same commands from a second host on a different network, and from the service's own machine against `127.0.0.1`. Three informative shapes emerge: works everywhere except here (local resolution, local firewall or egress policy), fails everywhere but works on localhost (the service is healthy and the path or the ingress filter is not), fails on localhost too (the application itself). This step converts a single data point into a boundary. ## Working the layers as a habit The same reasoning telescopes. If the 502 came from a proxy, the proxy's own upstream call is the next thing to run the drill against, one hop further in. What makes the answer strong is not the tool list but the discipline: name the layer you are testing, name what a pass rules out, and never skip a layer because you are confident about it.
- DNS, TCP and TLS all pass, and the response is a 502. Is the investigation over?No — it has moved one hop inward. A 502 is generated by a proxy telling you that *its* upstream call failed, so the same drill now runs from the proxy toward the backend: does the backend name resolve there, does the port accept, does the backend answer. The client-side layers are exonerated, which is exactly what the earlier steps bought you.
- How do you decide quickly whether the problem is local to this host or affects everyone?Run the identical commands from a second host on a different network and from the service's own machine against localhost. Working elsewhere but not here points at local resolution, a local firewall or egress policy; failing everywhere while localhost works points at the path or the ingress filter; failing on localhost too means the application. One comparison turns a data point into a boundary.
- What does the curl timing breakdown give you that a stopwatch does not?Attribution. time_namelookup, time_connect, time_appconnect and time_starttransfer are cumulative marks, so the gaps between them isolate resolution, the TCP handshake, the TLS negotiation and server think-time. A five-second request caused by a slow handshake and one caused by a slow backend look identical from outside and demand completely different fixes.
- Why insist on a short --max-time or nc -w during an incident?Because default connect timeouts can be tens of seconds or more, and a filtered port produces exactly that hang. Under pressure people interpret a long silence as evidence about the application when it is really the absence of any evidence. Capping the wait turns the test into a fast, repeatable signal you can run against several addresses in a minute.
saying these in an interview costs you the question
- Start by restarting the application
- A clean ping means the port is reachable
- curl failing proves the server is down
- Slow means the network is slow, no measurement needed
- Testing from one host is enough to declare an outage