In an MPLS core, how does RSVP-TE fast reroute keep traffic flowing within tens of milliseconds when a protected link fails?
answer
- build the detour before the failure
- repair at the router next to it
- point of local repair, merge point
- one bypass, many tunnels, by stacking
basics
~20 sRSVP-TE fast reroute pre-signals a backup path around each protected link or node; when the failure is detected, the adjacent router redirects traffic onto that backup immediately, with no path computation or signalling, while the head-end later re-optimises.
solid answer
~40 sRSVP-TE (RFC 3209) signals **explicitly routed** LSPs: the head-end sends a Path message with an `EXPLICIT_ROUTE` object and a `LABEL_REQUEST`, and labels return in Resv messages, so the path and its bandwidth can differ from the IGP's choice. Fast reroute (RFC 4090) adds backup paths built **before** any failure. The router upstream of a failure, the **point of local repair (PLR)**, switches traffic onto a backup that rejoins the LSP at a **merge point**. In **one-to-one backup** each LSP gets its own detour; in **facility backup** one bypass tunnel protects many LSPs by pushing the bypass label on top of the label the merge point expects. Redirection needs no computation or signalling, which is why RFC 4090 speaks of tens of milliseconds; the PLR then notifies the head-end to re-optimise.
go deeper
Recall that fast reroute protects an MPLS path by building a backup in advance, so the router next to a failure can switch traffic without waiting.
Explain how RSVP-TE signals an explicit path with Path and Resv messages, and name the point of local repair and the merge point.
Contrast one-to-one detours with facility bypass tunnels, explain the label push that lets one bypass carry many LSPs, and cover detection and head-end re-optimisation.
Weigh fast reroute's state and reserved spare capacity against IGP-speed recovery, and decide which traffic classes justify RSVP-TE in the core.
## Why LDP alone is not enough An LDP-signalled LSP follows the IGP's best path, so it recovers only as fast as the IGP reconverges: detect the failure, flood the change, rerun the shortest-path computation and reprogram forwarding. For voice or other real-time traffic that can be too slow, and during that window packets toward the failed link are dropped or loop briefly. RSVP-TE gives the operator two things LDP lacks: **explicit paths** that need not be the shortest, and the ability to set up **protection in advance**. ## RSVP-TE in brief RFC 3209 extends RSVP to set up LSP tunnels: - The **head-end** (ingress) sends a **Path** message carrying an `EXPLICIT_ROUTE` object, a list of strict or loose hops, and a `LABEL_REQUEST` object. - Each downstream router answers with a **Resv** message carrying a `LABEL` object, so labels are still downstream-assigned. - `SENDER_TSPEC` describes the bandwidth to reserve, and a `RECORD_ROUTE` object records the path actually taken. - The head-end usually computes the path from traffic-engineering data the IGP floods; OSPF carries it in Type 10 opaque LSAs with sub-TLVs such as maximum reservable and unreserved bandwidth (RFC 3630). RSVP is soft state: Path and Resv messages are refreshed periodically, and make-before-break lets the head-end move an LSP to a new path without dropping traffic. The price is state: every transit router holds per-LSP reservations, so a full mesh of RSVP-TE tunnels between many PEs grows quickly. ## How fast reroute works RFC 4090 defines backup LSPs for **local repair** of explicitly routed LSPs. Its speed comes from three choices: backups are computed and signalled before any failure, the repair happens at the router nearest the failure, and no failure notice has to travel anywhere before traffic moves. | Term | Meaning | |---|---| | **PLR** (point of local repair) | the router just upstream of the protected link or node; it performs the switch | | **MP** (merge point) | the router where the backup rejoins the protected LSP | | **NHOP bypass** | a backup tunnel around a single link | | **NNHOP bypass** | a backup tunnel around a single node, to the next-next hop | The two methods differ in how much state they create: 1. **One-to-one backup.** Each protected LSP gets its own **detour** LSP from each PLR. On failure the PLR swaps the packet onto the detour using the detour's label; the label stack does not grow. Protecting an LSP across N nodes can need up to N-1 detours per LSP, so state grows with the number of LSPs. 2. **Facility backup.** One **bypass tunnel** protects every LSP that passes through the PLR and the merge point. On failure the PLR swaps each packet's label to the one the merge point expects for that LSP, then **pushes** the bypass tunnel's label on top. The merge point sees its own label again once the bypass label is popped. Label stacking is what lets one bypass carry many LSPs. The head-end asks for protection through the `SESSION_ATTRIBUTE` flags (local protection desired, label recording desired, node protection desired) or a `FAST_REROUTE` object. Label recording is what tells the PLR which label the merge point expects. ## After the switch Local repair is a temporary path, often longer and possibly congested. The PLR SHOULD send a PathErr with error code "Notify" and the sub-code "Tunnel locally repaired", and the head-end then computes a better path and moves the LSP with make-before-break. RFC 4090 calls this global revertive mode; local revertive mode has the PLR re-signal the LSPs once the failed resource returns. ## The timing claim, stated carefully - RFC 4090 says traffic can be redirected "in 10s of milliseconds". It does not state 50 ms; that figure is an operator convention for the target. - The switch is fast only after the failure is **detected**. Loss of signal on a direct link is quick; a failure behind a switch needs a faster detector such as BFD. - RFC 4090 applies **only to explicitly routed LSPs**. LSPs that follow IGP routing, such as LDP's, are outside its scope and are protected by other mechanisms. - Backups cost state and spare capacity. A bypass that reserves no bandwidth keeps traffic flowing but may congest; one that reserves bandwidth holds capacity idle until a failure.
- In RSVP-TE fast reroute, when would you choose an NNHOP bypass over an NHOP bypass?An NHOP bypass goes around one link to the next hop, so it survives a fibre cut but not the loss of the next router. An NNHOP bypass skips the next node and merges at the hop after it, protecting against a router failure too. It needs a path avoiding that node, and the PLR must know the label the next-next hop expects, which label recording supplies.
- Why does the PLR notify the head-end after a successful local repair?The repaired path is a stopgap that may be longer, ignore the LSP's constraints or overload the bypass. RFC 4090 has the PLR send a PathErr with a Notify code saying the tunnel was locally repaired, so the head-end can compute a proper path and move the LSP with make-before-break without dropping traffic.
saying these in an interview costs you the question
- Fast reroute computes a new path after the failure, just very quickly.
- RFC 4090 guarantees restoration within exactly 50 milliseconds.
- Facility backup builds a separate detour for every protected LSP.
- RSVP-TE fast reroute protects LDP LSPs that follow the IGP.
- Once traffic is on the bypass, the job is done until the link returns.
- Fast reroute makes failure detection itself instantaneous.