Using the OSI layers bottom-up, how do you localise why a client can no longer reach a service on another subnet, and what evidence clears each layer?
answer
- each working layer vouches for those below
- link, neighbours, route, port, response
- a reset is an answer; silence is not
- a handshake can succeed while bulk data hangs
basics
~20 sBottom-up OSI triage checks the link (L1), then reachability of local neighbours (L2), then a route to the remote host (L3), then an open port (L4), then a correct response (L5-L7). The first layer whose evidence fails is where to dig.
solid answer
~50 sWalk up the stack and stop at the first layer that fails. **L1**: is there a link at all, a carrier or a link light? **L2**: with the link up, can the client exchange frames with neighbours on its own subnet, above all its default gateway? A port in the wrong VLAN fails here. **L3**: can it reach the remote host? No route may produce an ICMP Destination Unreachable, which RFC 1122 says to treat only as a hint, or plain silence. **L4**: does the service's port answer? A TCP connection attempt to a closed port gets a reset (RFC 9293); a UDP probe usually gets ICMP Port Unreachable; silence suggests a filter. **L5-L7**: the connection opens but the handshake or the response is wrong. Each success clears the layers beneath it for that path, so an experienced engineer often starts in the middle and branches.
go deeper
Know the order to check: link, local network, route, port, then the application's response, and what a working result at each step looks like.
Explain what each protocol response means: a TCP reset for a closed port, ICMP Port Unreachable for UDP, and why silence is different from a refusal.
Show judgment: start where the evidence points rather than always at L1, treat ICMP net and host unreachable as hints, and recognise size-dependent failures such as a handshake that works while bulk data hangs.
Turn the ladder into team practice: runbooks and alerts that record evidence per layer, so an outage report says which layer failed rather than that the service is down.
## Why bottom-up works Each OSI layer depends on the one below, so **evidence that a layer works vouches for every layer beneath it on that path**. If a TCP connection to the service opens, the link, the local neighbours and the route are all fine for that path; if the link light is off, nothing above it can work. Bottom-up triage turns an unhelpful report ("the service is down") into a question per layer, each with a concrete piece of evidence. ## The ladder | Layer | Question | Typical failure | Evidence that clears it | |---|---|---|---| | **L1 Physical** | is there a working link? | unplugged cable, dead transceiver, no radio association | link is up, carrier present | | **L2 Data Link** | can I exchange frames with neighbours? | access port in the **wrong VLAN**, wrong wireless network | the gateway's hardware address is resolved; frames from neighbours arrive | | **L3 Network** | can I reach the remote host? | **no route**, wrong gateway or mask | an ICMP echo or other reply from the remote host | | **L4 Transport** | is the service's port open? | **port closed**, service not listening, filtered | the TCP handshake completes | | **L5-L7** | is the exchange itself correct? | failed secure handshake, **bad response**, application error | a correct response to a real request | ## Reading each layer's evidence 1. **L1**: an interface with no link means nothing else is worth testing. Fix the medium first. 2. **L2**: the link is up but the host hears nobody on its own subnet. The classic cause is a switch port assigned to the **wrong VLAN**: the host is physically connected to a broadcast domain its gateway is not in, so resolving the gateway's hardware address never succeeds (ARP over IPv4, Neighbor Discovery over IPv6). 3. **L3**: the gateway answers but the remote host does not. A router may return **ICMP Destination Unreachable**; RFC 1122 says codes 0 (net) and 1 (host) "may result from a routing transient" and must be treated as "only a hint, not proof". Often there is only silence. 4. **L4**: the remote host is reachable but the service is not. The protocol tells you a lot here: - **TCP**: RFC 9293 says a segment arriving for a connection that does not exist (CLOSED), other than a reset, "causes a RST to be sent in response". An immediate reset means the host (or a device answering for it) is there and **nothing is listening**. - **UDP**: RFC 1122 says a host **SHOULD** send an ICMP Port Unreachable (code 3) when no process is listening on the port. - **Silence** (no reset, no ICMP, a timeout) points instead to something **dropping** the traffic on the way, such as a filter, or a reply lost on the return path. 5. **L5-L7**: the connection opens, then the secure-channel handshake fails, the server returns an error, or the content is wrong. That is the application and its immediate supporting layers, and the logs of the service are now the right place to look. ## Where layers mislead - **Success at one layer is evidence for one path and one packet size.** RFC 8201 describes connections "that complete the TCP three-way handshake correctly but then hang when data is transferred": when the ICMPv6 Packet Too Big messages that Path MTU Discovery relies on are blocked (for IPv4, RFC 1191's Datagram Too Big messages play the same role), small packets pass and large ones vanish. An L3 size problem then looks like an L7 hang. - **A reset is not always from the server.** A device on the path can answer for it, so "port closed" means "something at L4 refused", not necessarily the host you meant. - **Silence is ambiguous.** A reset or an ICMP error proves a device answered at that layer; a timeout can mean loss at any layer, a filter, or a host that simply sends no ICMP (it is a SHOULD, not a MUST). - **Asymmetric paths** can make a request arrive while its reply takes a broken route back. ## Bottom-up versus starting in the middle Strict bottom-up is exhaustive and ideal when **nothing** on the client works. When some things work, experienced engineers **divide and conquer**: - If the client reaches other services, L1-L2 on the client side are already cleared; start at L3 or L4. - A single TCP connection attempt to the service is the cheapest high-value test: a completed handshake clears L1-L4 for that path, a reset lands you at L4, silence sends you down to L3. - If only one service fails for every client, start at the top. ## What the interviewer is listening for Not the order of seven words, but whether each step produces **evidence** and whether the candidate knows which answers are **proof** (a reset, a completed handshake) and which are **hints** (a net or host unreachable, a timeout).
- In OSI-layer triage, when would you not start at Layer 1?When the evidence already clears the lower layers. If the client reaches other services, its link and local network work, so start at L3 or L4. One TCP connection attempt to the service is a good opening move: a completed handshake clears L1-L4 for that path, a reset puts you at L4, and silence sends you down to routing and filtering.
- In network triage, why is a timeout harder to interpret than a TCP reset or an ICMP error?A reset or an ICMP error is an answer: some device processed the traffic at that layer and said no. A timeout is the absence of an answer, which can mean loss on any link, a filter dropping silently, a broken return path, or a host that sends no ICMP, since RFC 1122 makes Port Unreachable a SHOULD rather than a MUST.
saying these in an interview costs you the question
- Jumping to application logs when nothing on the subnet can connect
- A completed TCP handshake proves the lower layers are fine for all traffic
- A timeout and a TCP reset both mean the server is down
- An ICMP net unreachable is proof the route no longer exists
- A wrong VLAN is a routing problem to fix at Layer 3
- A reset always comes from the server host itself