skip to content

In TCP/IP encapsulation, how does a receiving host know which protocol's header comes next after it strips each layer?

level: middleimportance: should knowfreq 32%

answer

  1. each header names its payload
  2. a type field per layer
  3. IPv4 Protocol, IPv6 Next Header
  4. ports select the endpoint

basics

~20 s

Each header carries a selector for its payload: the Ethernet type field names IPv4 or IPv6, the IP Protocol or Next Header value names TCP (6) or UDP (17), and the transport ports pick the receiving application socket.

solid answer

~40 s

Demultiplexing works because every header names what it wraps. The Ethernet frame's two-octet type field says the payload is IPv4, IPv6 or ARP. The IPv4 **Protocol** field, or the IPv6 **Next Header** field, which uses the same number space, says whether TCP (6), UDP (17) or something else follows; in IPv6 it can also name an extension header, which chains to the next one. At the transport layer, the ports choose the endpoint: for TCP the pair of sockets, meaning both addresses and both ports, identifies the connection. Above that there is no universal selector; the port is only a convention, so the application must validate what it receives.

go deeper

for a junior

Remember that each header names what is inside it: the frame says IPv4 or IPv6, the IP header says TCP or UDP, and the port picks the application.

for a middle

Name the fields and values: the Ethernet type field, IPv4 Protocol and IPv6 Next Header with 6 for TCP and 17 for UDP, and the socket pair for TCP connections.

for a senior

Connect the chain to operations: why extension headers or fragments can blind port-based filters, and why a port is evidence, not proof, of the protocol.

for a principal

Be ready to discuss the design tradeoff of selector fields: cheap, layered dispatch versus ossification, where middleboxes that parse the chain block new transports and headers.

## The problem demultiplexing solves Encapsulation on send is easy: each layer knows which protocol handed it the data. On receive, a layer only has bytes. After stripping its own header it must decide **which protocol above gets the payload**. This is **demultiplexing**, and it works because each header carries a **selector** naming the kind of payload it wraps. RFC 3439 lists multiplexing as one of the generic functions a layer performs; the selector field is how the receiving side undoes it. ## The selector at each layer | Header | Selector field | What it selects | Examples | |---|---|---|---| | Ethernet frame | type field (two octets) | the network-layer protocol | IPv4, IPv6, ARP | | IPv4 | `Protocol` | the transport or other upper protocol | 6 = TCP, 17 = UDP | | IPv6 | `Next Header` | an extension header or the upper protocol | 6, 17, 58 = ICMPv6 | | TCP / UDP | destination port (plus the rest of the socket pair for TCP) | the receiving endpoint | a listening service or an established connection | | Application | none by rule | — | conventions, or negotiation inside the protocol | ## Walking up a received frame 1. **Link layer.** The interface reads the Ethernet **type field**. RFC 1122 §2.3.3 notes that this two-octet field sits where IEEE 802.3 puts a **Length** field; a value of 1500 or less is an 802.3 length, and every valid type value is greater than 1500, so a receiver can tell the two framings apart. 2. **Internet layer.** IPv4 reads its **Protocol** field. IPv6 reads **Next Header**, which RFC 8200 §3 says uses the same values as the IPv4 Protocol field. In IPv6, the value may name an **extension header** instead of a transport, and each extension header has its own Next Header, forming a chain that ends at the upper-layer protocol. 3. **Transport layer.** UDP delivers to whoever is bound to the destination port (and address). TCP finds the connection from the **pair of sockets** — local address and port, remote address and port — because RFC 9293 defines a connection by that pair. Many connections can share destination port 443; the remote side of the pair tells them apart. 4. **Application layer.** Nothing in the transport header says which application protocol is inside. Port numbers are assigned by IANA, but RFC 7605 §7.1 is blunt: an assigned port number is not a guarantee of exclusive use, and traffic for any service can appear on any port. Where a protocol must be chosen on one port, the application negotiates it inside its own handshake, for example TLS's ALPN extension. ## Why every layer needs its own selector - **Independence.** A router reads the IP Protocol field only if it must; the link layer needs its own type field so that IP, ARP and other network protocols can share one wire. - **Parallel stacks.** IPv4 and IPv6 run side by side on one interface, and the frame's type field hands each packet to the right stack before any IP header is parsed. - **Checksum coupling.** The Protocol or Next Header value is fed into the TCP and UDP checksums through the pseudo-header, so a segment handed to the wrong transport will almost certainly fail that transport's checksum. In IPv6 the pseudo-header carries the upper-layer protocol, which differs from the IPv6 header's Next Header value when extension headers sit in between. ## Middleboxes and the chain A device that wants to act on ports has to walk the same chain: parse the frame type, the IP header, any IPv6 extension headers, and only then the transport header. That is why an unfamiliar extension header or an IP fragment without the transport header can leave a port-based filter blind. The detailed inspection capabilities of layer 4 and layer 7 devices belong to the devices topic; the point here is that ports are reachable only by walking the selectors in order. ## Common mistakes - Believing the receiver **guesses** the protocol by inspecting payload bytes. The protocol stack reads the selector field; content inspection is a middlebox or application technique, not how the stack demultiplexes. - Saying the destination port **alone** identifies a TCP connection. It identifies the listening service; established connections are told apart by the full socket pair. - Assuming IPv6 Next Header always names TCP or UDP. It can name an extension header. - Placing ports in the IP header. They live in the transport header, inside the IP payload. - Trusting that traffic on a well-known port must be that port's protocol.

  • Why can many TCP connections share a server's port 443 without confusion?
    RFC 9293 defines a TCP connection by a pair of sockets, each an IP address plus a port. The server side of every connection is the same address and port 443, but the client addresses and ports differ, so each incoming segment matches exactly one connection. Only a listening socket waits on the local half alone.
  • In IPv6, what happens to demultiplexing when extension headers are present?
    The IPv6 Next Header names the first extension header, and each extension header carries its own Next Header naming what follows, until one names the upper-layer protocol. A receiver, or a middlebox that wants the ports, must walk the whole chain to find the transport header.
  • If port numbers do not identify the application protocol, how does a server running several protocols on one port choose?
    It negotiates inside the application exchange. With TLS, the client offers protocol names in the ALPN extension and the server picks one. Without such a mechanism the server must recognise the protocol from what the client sends, which RFC 7605 recommends anyway: validate traffic by content, not by port.

saying these in an interview costs you the question

  • The receiver inspects the payload bytes to guess which protocol it is.
  • The destination port alone identifies an established TCP connection.
  • IPv6 Next Header always names TCP or UDP.
  • Port numbers are carried in the IP header.
  • Traffic on an assigned port is guaranteed to be that port's protocol.