Strict uRPF goes live tonight on both transit edges of a merged estate to stop forged sources: what breaks, and what is the back-out?
answer
- inbound paths are chosen by other people
- the drop is silent and instant
- strict belongs where there is one path
- smaller provider first, one interface at a time
- the back-out is typed before the change is armed
basics
~20 sOn a dual-transit edge, inbound traffic from any source whose best return path is the other provider is dropped the moment strict mode is armed - potentially a large slice of the internet. Enable it only where routing is symmetric, and pre-stage the removal.
solid answer
~60 sStrict mode assumes the best path back to a source leaves by the interface the packet arrived on. A merged estate with two edges and two providers violates that constantly: inbound paths are chosen by other people's routing policy, so traffic from a source your table reaches via provider B routinely arrives via provider A. Arm strict mode on the provider-facing interfaces and you drop it, silently, at line rate. The first calls are partners and customers whose sessions simply stop, and they will not know why. The correct plan is not a safer window, it is a different placement: strict only where the estate is single-pathed by construction - access and user segments, single-homed branch links - feasible-path on interfaces to multi-homed neighbours whose alternate announcements you actually learn, and on the transit edges an outbound ACL permitting only your own source prefixes plus loose mode inbound. Back out one interface at a time, smaller provider first, with the removal pre-staged and a hard decision time in the change record.
go deeper
Understand that this check compares an arriving packet's source against the routing table, and that dropping happens with no error returned to the sender - the symptom is a session that simply stops.
Be able to explain why two providers make inbound paths asymmetric, and therefore which traffic strict mode drops that a single-homed site would never see.
Show the placement plan rather than the window plan: strict where the path is single by construction, feasible-path where alternates are learned, and an outbound prefix ACL plus loose mode at the transit edges. Bring the baseline, the ordering and the pre-staged back-out.
Be ready to defend keeping the weak mode at the border as the deliberate answer, and to say who owns the hand-maintained prefix list that ends up carrying the real protection.
## What actually happens at 02:00 The change is one line per interface and it takes effect instantly in the forwarding plane. There is no ramp, no learning mode, no log-only stage on most platforms - the check either runs or it does not. Within seconds, every arriving packet whose source prefix your table reaches by the other edge is dropped. On a merged estate the blast radius is larger than people expect, and it is worth being able to enumerate it in an interview: 1. **Ordinary inbound internet traffic with asymmetric paths.** You do not choose how the internet reaches you; the sending networks do, based on their own policy and the announcements they receive. Any source whose traffic prefers provider A while your best return path is provider B breaks. On a two-transit edge this is not an edge case, it is a large fraction of the table. 2. **Multi-homed partners and customers.** A neighbour announcing to you over two links sends over whichever it prefers. Strict at the interface that is not your best return path drops them. 3. **Anything behind a path you deliberately de-preferred.** Traffic engineering that makes one link the backup for return traffic does not stop that link receiving inbound traffic. 4. **Return traffic for flows YOU started.** If your outbound path and the reply path differ - normal on a dual-transit estate - the reply arrives on an interface strict mode does not accept. And the failure is silent. There is no rejection, no ICMP by default worth relying on, no application error that says *filtered*. The user-visible symptom is a session that stops, and the only local evidence is a per-interface drop counter that increments. ## The judgment being tested The interviewer is not asking for a safer maintenance window. They are checking whether you know that strict mode is the wrong control for this position, and that the honest end state is uneven: | Position | What goes there | What it costs | | --- | --- | --- | | Access and user segments, single-homed branch links | Strict | Nothing while the segment stays single-pathed; a future second path breaks it | | Interfaces to multi-homed neighbours whose alternate paths you learn | Feasible-path | Depends on the neighbour continuing to announce on both links | | Provider-facing transit interfaces | Loose inbound, plus an outbound ACL of your own source prefixes | The prefix list needs an owner; loose mode does not stop a spoofer | The last row is the point of the leaf. Across the edge as a whole, the mode you can actually keep is the weak one. The forged-source packets leaving your estate are stopped by the explicit outbound prefix ACL - a hand-maintained list - not by reverse-path checking at the border. Anyone who claims the border reverse-path check solved spoofing has not deployed it. ## Running the night properly Things that separate an operator who has done this from one who has read about it: - **Baseline before you arm anything.** Capture the per-interface drop counters and a window of flow records (a flow record carries the five-tuple, byte and packet counts and timestamps - it carries no payload, so it will tell you WHICH sources you dropped and how much, never what the traffic was). Without the before, the after means nothing. - **One interface at a time, smaller provider first.** If it goes wrong you have lost the smaller share of inbound traffic and you still have a working comparison. - **A hard decision time, written down.** Not *we will see how it looks*. A timestamp at which, absent clean counters, you remove it. - **The removal pre-staged.** The back-out is typed, reviewed and ready to paste before the change is armed, because at 02:20 with partners paging you nobody should be composing configuration. - **A contact list for neighbours.** Their traffic dies inbound. They see an outage in your direction with no explanation and no way to diagnose it from their side. If your change is what broke them, they need to hear it from you, not open a ticket into a queue. - **A second pair of eyes on what the counters mean.** A drop counter climbing is the change WORKING as configured and possibly failing as intended - it does not distinguish a spoofer from a partner. That distinction comes from the flow records, by looking at which source prefixes stopped. ## The uncomfortable summary You set out to stop forged source addresses leaving your estate and you have discovered that the strong mechanism cannot live where the traffic crosses. The control that actually does the job in that position is the one nobody finds interesting: a written list of your own prefixes, applied outbound, owned by a human, updated when the business changes shape. Strict reverse-path checking is excellent - deeper inside, where you built the topology and know it has one path.
- How do you tell, during the window, whether the drops are spoofed traffic or a partner's real flows?The interface counter cannot tell you - it counts packets, not intent. You need the flow records either side of the change: which source prefixes were arriving before and stopped after. A source prefix that was carrying steady, long-lived bidirectional volume is a neighbour with asymmetric routing; forged sources tend to appear as scattered prefixes with no matching return traffic in your own outbound records.
- Why not leave strict mode running in a log-only mode instead of backing out?Because on most forwarding paths there is no such mode - the check runs in hardware and its output is a drop. Where a counting or logging variant exists it is worth using, but you should not plan the change assuming one: the safe equivalent is to test placement on a single low-risk interface first and to derive the expected breakage from flow records beforehand, not from production drops.
- The estate is single-pathed today and strict mode works. What makes it a future outage?Any change that adds a second path: a new circuit for resilience, an acquired site interconnect, a cloud transit attachment, a partner who starts announcing over a backup link. Nothing in the strict configuration references the topology assumption it depends on, so the outage arrives months later attached to a change nobody connected to it. The assumption belongs in the change record for the link, not only in the router.
saying these in an interview costs you the question
- Plans a bigger maintenance window instead of changing the placement
- Believes reverse-path checking at the border stops forged sources leaving
- Expects an error or ICMP message to reveal the dropped traffic
- Arms both edges at once because the change is one line
- Reads a rising drop counter as proof the change is working correctly
- Has no back-out prepared before the change is armed