skip to content

Two OSPF routers list each other as neighbours but never get past ExStart or Exchange; what usually causes this, and how do you confirm it?

level: seniorimportance: should knowfreq 30%

answer

  1. Hellos fine, DD packets not
  2. a field Hellos do not carry
  3. larger than I can receive
  4. who is master decides who stalls where

basics

~20 s

Usually an interface MTU mismatch: each OSPF Database Description packet carries its sender's interface MTU, and a router rejects one advertising more than it can receive, so the exchange never completes. Duplicate Router IDs and dropped unicast OSPF packets are the other suspects.

solid answer

~50 s

Hellos carry no MTU, so the pair passes the Hello checks and reaches ExStart normally. Database Description (DD) packets do carry an `Interface MTU` field, and RFC 2328 says a router rejects a DD packet advertising an IP datagram larger than it can accept on that interface without fragmentation. The router with the smaller MTU therefore never leaves ExStart; its peer waits in ExStart if it would be master, or in Exchange if it became slave, and the stuck side keeps retransmitting every `RxmtInterval`. To confirm it, compare the interface MTUs on both ends and capture the DD packets to read their MTU fields. Fix it by aligning the MTUs; an implementation option to skip the check only hides the fault until a large update is dropped. If the MTUs match, check for duplicate Router IDs and for filters that pass multicast Hellos but drop unicast DD packets.

go deeper

for a junior

Remember that OSPF routers stuck in ExStart or Exchange usually have different interface MTUs, even though their Hellos look fine.

for a middle

Explain that only Database Description packets carry the MTU, and that a router rejects one advertising more than it can receive.

for a senior

Predict which side stalls where from the MTUs and Router IDs, confirm it with a capture, and rule out duplicate IDs and unicast filtering before touching configuration.

for a principal

Treat MTU as a fabric-wide standard with drift checks, since a link-by-link mismatch surfaces as routing instability long before anyone looks at frame sizes.

## Why Hellos succeed and the exchange fails Becoming OSPF neighbours and becoming **adjacent** are separate steps. The Hello checks in RFC 2328 cover the area, timers, mask, authentication and stub flag; none of them look at the link's **MTU** (maximum transmission unit, the largest IP datagram the interface sends without fragmentation). So two routers whose MTUs differ see each other normally, reach 2-Way, decide to form an adjacency and enter **ExStart**. Trouble starts with the first packet that carries the MTU: the **Database Description (DD)** packet. ## The MTU check in RFC 2328 Every DD packet carries an `Interface MTU` field, set to the largest IP datagram the sender can send out that interface without fragmentation (0 on virtual links). RFC 2328 section 10.6 then says, in effect: ```pseudocode on receive DD packet from neighbour N on interface I: if DD.interfaceMtu > largest datagram I can receive unfragmented: reject packet # nothing is learned, nothing is acknowledged return process by N.state # ExStart negotiation, Exchange sequencing, ... ``` The purpose is protective. OSPF sizes its own DD and Link State Update packets to fit the sender's MTU. If the receiver's MTU is smaller, those packets would be dropped or fragmented on the way, and an adjacency built on top of that would lose updates later. Refusing the exchange up front makes the problem visible. ## Tracing the stall Take router A with an MTU of 1500 bytes and router B with jumbo frames at 9000. A rejects every DD packet from B, because 9000 is more than A can receive; B accepts A's packets, because 1500 fits. What each router shows depends on which one is **master**, the router with the higher Router ID: | Case | Router A (MTU 1500) | Router B (MTU 9000) | Why | |---|---|---|---| | B has the higher Router ID | ExStart | ExStart | B, as master, waits for A to acknowledge its first DD; A rejects it, so B ignores A's own claim to be master | | A has the higher Router ID | ExStart | Exchange | B accepts A's first DD and becomes slave; A rejects B's reply, so it never confirms the negotiation | Either way the router with the **smaller** MTU never leaves ExStart, and the stuck side retransmits its DD every `RxmtInterval` (RFC 2328 gives 5 seconds as a sample for a LAN). How long an implementation tolerates this before resetting the neighbour is an implementation choice; many cycle the neighbour back down and start again, which looks like a flapping adjacency. ## Other causes of an ExStart or Exchange stall - **Duplicate Router IDs.** The master and slave roles are settled by comparing Router IDs: the slave sees a *larger* ID, the master a *smaller* one. Two equal IDs satisfy neither test, so the negotiation packets are ignored. Router IDs must be unique in the autonomous system for many other reasons too. - **Unicast OSPF filtered, multicast allowed.** On broadcast and NBMA networks, Hellos go to the multicast group `224.0.0.5` but DD packets are sent as **unicasts** to the neighbour's address. A filter that permits the multicast group and drops unicast OSPF (IP protocol 89) produces exactly this picture. On physical point-to-point links everything goes to `224.0.0.5`, so this cause does not apply there. - **Repeated `SeqNumberMismatch`.** If DD sequence numbers or the Options field keep changing mid-exchange, each mismatch sends the neighbour back to ExStart to start again. ## Confirming it 1. Note which side is in which state; the router stuck in ExStart while its peer sits in Exchange points at the MTU of the ExStart side. 2. Compare the configured MTU on both interfaces, remembering that some platforms count the layer-2 header in their MTU figure and some do not. 3. Capture the DD packets and read the `Interface MTU` field each side actually advertises. 4. Check the neighbour's Router ID against the router's own. 5. Check that unicast OSPF packets between the two interface addresses are not filtered. ## Fixing it, and the trap The fix is to **make the MTUs equal** on both ends, which is also what the rest of the traffic on that link needs. Some implementations offer an option to skip the DD MTU check. That gets the adjacency to Full, but the protection disappears: a Link State Update sized for the larger MTU can later be dropped on the way to the smaller-MTU router, leaving a neighbour stuck in Loading or a database quietly missing LSAs. Use it only as a deliberate, documented exception.

  • Why not just tell the OSPF routers to ignore the MTU check?
    The check exists because OSPF sizes its Database Description and Link State Update packets to the sender's MTU. With the check skipped, the adjacency reaches Full, but an update larger than the receiver's MTU can be dropped later, leaving a neighbour stuck in Loading or a database missing LSAs. Aligning the MTUs fixes the cause; ignoring the check only moves the symptom.
  • How does a duplicate Router ID keep two OSPF routers in ExStart?
    ExStart settles master and slave by comparing Router IDs: a router becomes slave when the neighbour's ID is larger and master when it is smaller. Equal IDs match neither rule in RFC 2328, so each router ignores the other's negotiation packets and the pair never reaches Exchange. The duplicate would corrupt the area's database in other ways too, so the IDs must be made unique.

saying these in an interview costs you the question

  • An MTU mismatch stops OSPF routers from seeing each other's Hellos.
  • Two OSPF routers stuck in ExStart are still waiting for the DR election.
  • Telling OSPF to ignore the MTU check is a complete fix.
  • Raising the OSPF dead interval will let the stuck database exchange finish.
  • With an MTU mismatch, both OSPF routers always show exactly the same state.