skip to content

In the OSI model, if the data link layer already detects corrupted frames, why does the transport layer check data for errors again?

level: middleimportance: should knowfreq 38%

answer

  1. per link versus end to end
  2. what happens inside a forwarding node
  3. the end-to-end argument
  4. a local check as a performance aid

basics

~20 s

A data link layer check protects one link and is rebuilt on every hop, so corruption inside a router or host escapes it. Only the endpoints can verify data end to end, so the transport layer checks again.

solid answer

~50 s

The layers have different **scopes**. L2 works **node to node on one link**: the sender computes a check over the frame, the next node verifies it and strips the frame, and the following link gets a brand-new frame with a fresh check. Anything corrupted *between* those checks, inside a router's memory or a host's stack, is covered by a perfectly valid new check. L4 is the first layer that, in the model, runs only in the **two end hosts**, so its checksum (TCP and UDP both define one) is the lowest check covering the data from sender to receiver. This is the **end-to-end argument** RFC 1958 cites: integrity can be assured only by the endpoints, and a check inside the network is at most a performance aid. RFC 6936 notes corruption inside routers that the strong link-layer frame checksums did not detect.

go deeper

for a junior

Remember that the data link layer works on one link at a time while the transport layer runs only on the two end hosts, and that both check for errors.

for a middle

Explain why a per-link check is rebuilt at every hop and so cannot catch corruption inside a router, and why that makes the transport checksum the lowest end-to-end one.

for a senior

Know the limits: a 16-bit ones' complement checksum misses some errors, UDP only discards, and data that must be intact needs an application-level or cryptographic check.

for a principal

Apply the end-to-end argument to design choices: decide which guarantees belong at the endpoints and which in-network checks are only optimisations worth their cost.

## Three layers, three scopes The question is really about **what each layer is responsible for**, and the answer is that the lower layers do *not* have the same reach: | Layer | Delivers between | Error handling it owns | Who runs it | |---|---|---|---| | **L2 Data Link** | two nodes on **one link** | detects corrupted frames on that link | every node on the path, per link | | **L3 Network** | two **hosts**, across many links | IPv4 carries a header checksum; IPv6 has none | end hosts and every router | | **L4 Transport** | two **processes** on the end hosts | checksum over header and data; TCP also recovers | the two end hosts only | RFC 1122 states the L3 part plainly: IP is "a connectionless or datagram internetwork service, providing no end-to-end delivery guarantees", and datagrams "may arrive at the destination host damaged, duplicated, out of order, or not at all. The layers above IP are responsible for reliable delivery service when it is required." ## Why the link check is not enough A data link check is computed by the sender of a frame and verified by the receiver **on that same link**. Then: 1. A router receives the frame, verifies its check, and **discards the frame wrapper**. 2. It processes the packet inside: looks up a route, updates fields, copies it between buffers. 3. It builds a **new frame** for the outgoing link, with a **new check computed over whatever is in memory now**. If a bit flipped during step 2, the outgoing frame's check is computed over already-corrupt data, so it verifies perfectly on the next link. Every link check on the path passes; the data is still wrong. RFC 6936 (section 3.1) cites exactly this: evidence of corruption from bad internal processing in routers or hosts that was "not detected by the strong frame checksums employed at the link layer". So link checks protect against **noise on the medium**; they cannot protect against faults **between** links. ## The end-to-end argument RFC 1958 (section 2.3) summarises the principle from Saltzer, Reed and Clark: certain functions "can completely and correctly be implemented only with the knowledge and help of the application standing at the endpoints". Data integrity is the classic example. Consequences: - The **transport layer**, running only in the two end hosts, is the lowest layer that can check the data over its **whole journey**. - TCP's checksum is "the 16-bit ones' complement of the ones' complement sum" of the header and text, and RFC 9293 says it is never optional; UDP uses the same construction (RFC 768). Both also cover a pseudo-header with the IP addresses, so data delivered to the wrong host is caught too. - Because IPv6 has no header checksum, RFC 8200 makes the UDP checksum **not optional** by default over IPv6, whereas over IPv4 a sender may omit it. ## Detecting is not recovering Each layer's responsibility also differs in what happens *after* detection: - **L2**: a frame that fails its check is typically dropped at that hop; the layers above notice only its absence. - **UDP**: RFC 1122 says a datagram with a non-zero, invalid checksum "MUST" be silently discarded. UDP never retransmits; recovery, if any, is the application's job. - **TCP**: discards the damaged segment and, because that data is never acknowledged, the sender **retransmits** it. Reliability is a transport responsibility. ## Why keep the link check at all The same RFC 1958 quote adds that "an incomplete version of the function provided by the communication system may be useful as a performance enhancement". On a noisy link, dropping a bad frame immediately is far cheaper than carrying it across the network, delivering it, and waiting for an end-to-end retransmission. The link check is a **local optimisation**; the transport check is the **end-to-end safeguard**. ## Limits of the transport check - A 16-bit ones' complement sum is weak: it misses some multi-bit error patterns that a link CRC would catch. - RFC 6936 reports measured paths where UDP datagrams arrived with detectably corrupt checksums, which suggests some corruption also arrives undetected. - Applications that cannot tolerate that add **their own integrity check** above transport, such as a strong hash or a cryptographic integrity check. That is the end-to-end argument applied once more, one layer higher. ## The takeaway Two checks at two layers are not redundancy by accident. They answer different questions: *"did this link deliver this frame intact?"* and *"did the other endpoint receive what I sent?"* Only the second is the transport layer's job, and only the endpoints can answer it.

  • If the transport checksum covers the data end to end, why do data link layers bother checking frames at all?
    As a performance aid. RFC 1958, quoting the end-to-end argument, notes that a partial version of a function inside the network can be a useful performance enhancement. Dropping a corrupt frame at the noisy link saves carrying it across the rest of the path and waiting for an end-to-end retransmission, but it is not the integrity safeguard.
  • Does a TCP or UDP checksum make end-to-end data integrity certain?
    No. Both use a 16-bit ones' complement sum (RFC 9293, RFC 768), which misses some error patterns. RFC 6936 reports paths where corrupt UDP datagrams were detected in measurable numbers, so some undetected corruption is plausible too. Applications that need certainty add their own strong or cryptographic integrity check above transport.

saying these in an interview costs you the question

  • The link-layer check protects the data all the way to the receiver
  • The transport checksum is redundant once every link checks its frames
  • A checksum corrects the errors it detects
  • UDP retransmits a datagram whose checksum fails
  • IP guarantees delivery, so transport needs no error checking
  • The transport checksum is stronger than a link-layer frame check