skip to content

Before maintenance on an OSPF router, how do you move transit traffic off it without dropping packets, and why not simply shut OSPF down?

level: seniorimportance: should knowfreq 20%

answer

  1. converge while the old path works
  2. change only the router's own LSA
  3. non-stub links at maximum cost
  4. cost belongs to the output side

basics

~20 s

Cost the OSPF router out first: reoriginate its router-LSA with every transit link at the maximum cost, 0xffff (RFC 6987 stub router advertisement), let traffic reconverge around it, then work. A hard shutdown drops traffic until the area reconverges.

solid answer

~40 s

A hard shutdown loses traffic through the router until the area learns of it — from flushed LSAs, lost carrier or, at worst, the Dead interval — and every router has flooded, recomputed and installed. Draining first moves that reconvergence to a moment when the old path still forwards. RFC 6987 (Informational, obsoleting RFC 3137) has the router reoriginate its router-LSA with every non-stub link at `MaxLinkMetric` (0xffff). Any transit path through it must leave over one of those links, so SPF prefers alternatives, while its own stub prefixes stay reachable. OSPFv3 can clear the R-bit instead, which forbids transit outright. The same trick protects bring-up: stay costed out until forwarding is ready. To drain a single link, raise its cost on both ends, because cost applies only to the output side.

go deeper

for a junior

Recall that you should move traffic away from a router before working on it, rather than switching routing off while it carries traffic.

for a middle

Explain how advertising a maximum cost on the router's transit links makes other routers choose different paths while its own addresses stay reachable.

for a senior

Show the operational judgment: drain before work, hold the router out on bring-up until forwarding is ready, raise per-link cost on both ends, and know when the R-bit is needed.

for a principal

Frame draining as part of a maintenance discipline: which tools the network must support, how redundancy makes them useful, and how to standardise them.

## Why a hard shutdown loses traffic Turning off OSPF (or the interfaces) on a router in service creates exactly the failure the network is built to survive — but survival is not instant. Packets already headed through the router are lost until: - the rest of the area **learns** of the change — from LSAs the router flushes as it shuts down, from lost carrier, or at worst only after the Dead interval; - each affected neighbour **originates** a router-LSA without the router, and that LSA **floods** across the area; - every router reruns **SPF** and **installs** the new routes. During that window, upstream routers still forward into a router that no longer forwards. The fix is to reverse the order: make the network move traffic away *while the router still forwards*, then do the work. ## Stub router advertisement (RFC 6987) RFC 6987, an Informational RFC that obsoletes RFC 3137, defines how a router announces itself as a **stub** — reachable, but not to be used for transit: 1. The router reoriginates its **router-LSA** with the cost of every **non-stub link** (every link type other than type 3, stub network) set to **`MaxLinkMetric`**, the 16-bit all-ones value **0xffff**. 2. Its **stub network links keep their normal cost**, so its own addresses stay reachable at normal cost — for example the address you manage it through. 3. Other routers run SPF, find any alternative path cheaper, and move transit traffic off it. Only then does the operator start disruptive work. The RFC explicitly lists "graceful introduction and removal of the router" among its uses, along with a router short of CPU or memory. ## Why one router's LSA is enough OSPF costs are **directed**: each router advertises the cost of its own *outgoing* interfaces, and a path's cost is the sum of the output costs along it. A packet that transits router X must **leave** X over one of X's non-stub links, so every transit path through X includes at least one 0xffff hop. X does not need its neighbours to change anything. The flip side matters for a single link. Raising the cost on router A's interface toward B affects only traffic **leaving A** over that link. Traffic from B toward A uses B's output cost and keeps flowing. To drain a link in both directions, raise the cost on **both ends**. ## Options compared | Technique | Transit traffic | Router's own prefixes | When it is the only path | |---|---|---|---| | Hard shutdown | lost until reconvergence | unreachable | unreachable | | `MaxLinkMetric` on non-stub links | moved off before work | reachable | still used by routers following RFC 1583 and later | | OSPFv3 R-bit cleared | never computed through it | reachable | not used for transit | | Higher cost on one interface | moved off in one direction only | reachable | still used | | Flushing its router-LSA | moved off | unreachable (no back link) | unreachable | ## Bring-up and caveats - **Bring-up is the mirror image.** A freshly reloaded router can attract traffic before its forwarding table is complete. Keeping it costed out until it is ready, then restoring normal costs, avoids that black hole. - **0xffff is very expensive, not infinite.** If the stub router is the only path, routers computing SPF as in RFC 1583 and later still use it; only routers following the older RFC 1247 procedure discard such links. Use OSPFv3's R-bit when transit must be refused even then. - **It needs redundancy.** Draining helps only where traffic has somewhere else to go. - **It costs two reconvergences**, out and back, but both happen while every path still works. ## A maintenance sequence 1. Confirm an alternative path exists for the traffic the router carries; draining a single point of failure moves nothing. 2. Advertise the router as a stub (or raise the chosen links' costs on both ends). 3. Wait for the area to reconverge, and verify on the traffic counters that transit traffic has actually left the router. 4. Do the disruptive work: reload, replace hardware, recable. 5. Bring the router back still costed out, wait until its adjacencies are Full and its forwarding table is complete. 6. Restore normal costs and watch traffic return. Step 3 is the one people skip. A drain that was announced but never checked is an easy way for this procedure to still drop traffic.

  • If a costed-out OSPF router is the only path to some destination, is that traffic dropped?
    Not with MaxLinkMetric. 0xffff is a very high cost, not unreachable, so routers following RFC 1583 and later still use the path when no alternative exists; only routers using the older RFC 1247 procedure discard such links. OSPFv3's R-bit is the stricter tool: with it clear, no route transiting the router can be computed even if it is the only one.
  • How does costing a router out compare with OSPF graceful restart for a software upgrade?
    They answer different questions. Graceful restart keeps traffic flowing through the router while only its control plane restarts and its forwarding plane survives; neighbours pretend nothing changed. Costing out moves traffic away for work that does interrupt forwarding — a reload, a hardware swap, recabling — at the price of two reconvergences, both while the old path still works.

saying these in an interview costs you the question

  • Shutting down the OSPF process is a graceful way to remove a router.
  • Raising the cost on one end of a link moves traffic in both directions.
  • Stub router advertisement makes the router's own addresses unreachable.
  • Every router treats a 0xffff link cost as unreachable.
  • Costing out a router requires configuration changes on all its neighbours.