How does OSPF graceful restart (RFC 3623) keep traffic flowing while a router's OSPF software restarts, and what ends it early?
answer
- control plane restarts, forwarding plane stays
- link-local Opaque-LSA with a timer
- neighbours pretend the adjacency is Full
- any real topology change aborts
basics
~20 sIn OSPF graceful restart, the restarting router announces a grace period in link-local grace-LSAs and keeps forwarding on its preserved table; helper neighbours keep advertising it as fully adjacent until it resynchronises, the period expires or the topology changes.
solid answer
~40 sBefore a planned restart the router ensures its forwarding table will survive, then sends a **grace-LSA** on each interface: a link-local Opaque-LSA (LS type 9, Opaque Type 3) carrying a grace period, which should not exceed `LSRefreshTime` (1800 s), and a restart reason. Each neighbour that is Full with it, sees no changed LSAs and allows it by policy enters **helper mode** and keeps listing the router as fully adjacent, so nobody reroutes. Meanwhile the restarting router originates no LSAs of types 1-5 or 7, rebuilds adjacencies and runs SPF without installing routes. It finishes when every adjacency is back. A changed LSA of types 1-5 or 7 that would reach it, an expired grace period, or an LSA inconsistent with its pre-restart router-LSA turns it into a normal restart.
go deeper
Recall that graceful restart lets a router restart its routing software while still forwarding traffic, with neighbours' help.
Explain the grace-LSA, its link-local scope and grace period, and what the restarting router and its helpers each do until the restart finishes.
Judge when it is safe: a forwarding plane that truly survives, planned rather than unplanned restarts, and how detection timers or a topology change can end it early.
Weigh graceful restart against simply routing around the router: hiding a restart saves two reconvergences but risks forwarding on stale state if assumptions break.
## The problem graceful restart solves Many routers separate the **control plane** (the software running OSPF) from the **forwarding plane** (the hardware or kernel table that actually moves packets). Restarting the OSPF software — an upgrade, a process restart, a switch to a redundant control processor — need not stop forwarding. But under plain OSPF, the neighbours see the adjacency drop, reoriginate their LSAs without the router, and the whole area reroutes around it, only to reroute back moments later. **Graceful restart**, specified in RFC 3623 (Standards Track) and also called *non-stop forwarding*, lets the network act as if the router never left, provided nothing else changes. ## The restarting router's side 1. **Prepare.** The router ensures its forwarding table is up to date and will survive the restart. It does *not* flush its own LSAs — the point is that they stay in everyone's database. 2. **Announce.** It originates a **grace-LSA** on each OSPF interface and floods it reliably. A grace-LSA is a **link-local Opaque-LSA** (LS type 9, Opaque Type 3, Opaque ID 0) whose TLVs carry the **grace period** in seconds, a **restart reason** (0 unknown, 1 software restart, 2 software reload or upgrade, 3 switch to redundant control processor) and, on broadcast, NBMA and point-to-multipoint segments, its interface address. The grace period should not exceed `LSRefreshTime` (1800 s), so its LSAs do not age out. 3. **Restart and resynchronise.** After the restart it does not originate LSAs of types 1-5 or 7, accepts the copies of its own pre-restart LSAs that neighbours send back, and runs the normal Hello and database-exchange procedures to rebuild adjacencies. 4. **Compute without installing.** It runs SPF (needed to bring virtual links back) but keeps forwarding on the entries installed before the restart. 5. **Exit.** It reoriginates its router-LSAs and any network-LSAs, reruns SPF and installs the result, removes stale forwarding entries, and flushes its grace-LSAs. ## The helper's side A neighbour enters **helper mode** for the restarting router on a segment only if: - it is currently **Full** with that router on the segment; - the database has seen **no changes** to LSAs of types 1-5 or 7 since the restart, apart from periodic refreshes; - the grace period has not expired; - **local policy** allows helping (for example, only for planned upgrades); - it is not itself restarting. While helping, it keeps advertising the restarting router as fully adjacent in its own LSAs, regardless of database synchronisation. ## What ends graceful restart | Event | Who acts | Result | |---|---|---| | All adjacencies re-established | restarting router | successful exit; grace-LSAs flushed | | Grace-LSA flushed | helper | stops helping (success) | | Grace period expires | both | normal OSPF operation resumes | | Changed LSA of types 1-5 or 7 that would reach the restarting router | helper | stops helping, reoriginates LSAs from real adjacency state | | LSA inconsistent with its pre-restart router-LSA | restarting router | aborts, reoriginates and recomputes | The topology-change rule exists because a restarting router cannot adapt its frozen forwarding table; continuing to pretend would risk loops and black holes. A neighbour that does not support graceful restart simply reroutes, which the restarting router sees as an inconsistency — so the mechanism falls back safely. ## Prerequisites and caveats - **A surviving forwarding plane.** Without it, graceful restart only hides a real outage. - **Unplanned restarts** are allowed but discouraged: the router could not prepare, so its table may not be trustworthy. If used, grace-LSAs must go out before any Hello, to 224.0.0.5 on broadcast networks, with reason 0 or 3. - **Cryptographic authentication** may need sequence numbers preserved across the restart; otherwise re-forming adjacencies can take up to `RouterDeadInterval`, lengthening the grace period needed. - **OSPFv3** uses the same procedure with its own grace-LSA, LS type 0x000b, defined in RFC 5187. - Grace-LSAs use the Opaque-LSA mechanism now specified in RFC 5250, which obsoletes RFC 2370. ## Graceful restart versus routing around the router Graceful restart is a bet that hiding the restart is safer than reacting to it. - **What it saves**: two area-wide reconvergences (away and back), each with flooding, SPF and install on every router, and the transient loss they cause. - **What it risks**: for up to the grace period the network forwards through a router whose table cannot adapt. If a failure elsewhere goes unnoticed by the helpers, or the forwarding plane is not as healthy as assumed, traffic can loop or vanish. - **How the RFC limits the risk**: helpers abort on any relevant topology change, and the restarting router aborts on any inconsistency, so the bet is called off as soon as the world stops matching the frozen state. It is the right tool for a planned software restart on a router whose forwarding plane really survives; for a reload or hardware work, moving traffic away first is safer.
- Should OSPF graceful restart be used to recover from an unplanned crash?RFC 3623 permits it but warns it may not be wise: the router could not prepare, so it is unlikely to guarantee a sane forwarding table. If used, it must send grace-LSAs before any Hello, to 224.0.0.5 on broadcast networks, with reason 0 (unknown) or 3 (switch to redundant control processor), so neighbours can decide whether to help, and operators must be able to turn the option off.
- Why can a fast failure-detection session interfere with OSPF graceful restart?Graceful restart assumes forwarding continues while the control plane restarts. A detection session that shares fate with the control plane fails as a side effect of the restart and can tear down the adjacency, ending the restart; one running in the forwarding plane can instead show the data path really died, which should abort it. How such protocols advertise which case applies belongs to the detection protocol itself.
A shop manager is out for an hour: staff keep following yesterday's delivery routes, and neighbouring branches agree not to redraw the map around the shop. If a road actually closes meanwhile, the pretence ends and everyone replans properly.
saying these in an interview costs you the question
- Graceful restart lets a router fully reboot, forwarding plane included, without loss.
- Grace-LSAs are flooded across the whole area so every router knows.
- Helpers keep covering for the restarting router even after a topology change.
- The restarting router installs fresh routes as soon as its first adjacency returns.
- Graceful restart only works if every router in the area supports it.