A Linux server has two NICs on two different networks, each with its own gateway. Clients on the first network reach it fine; clients on the second reach it only sporadically or not at all. Explain what the kernel is doing with the reply packets, and how policy routing fixes it.
answer
- one table means one default route
- replies leave by the wrong door
- destination routing ignores where it came from
- rules select the table, source selects the rule
- test the lookup with a source address
basics
~20 sWith one routing table there is one default route, so replies to off-subnet clients on the second network leave through the first NIC's gateway. That asymmetry is dropped by reverse-path filtering or by stateful middleboxes. The fix is a second routing table selected by a source-based rule.
solid answer
~50 sThe kernel makes its routing decision purely on the destination, and the `main` table holds only one default route. So a request that arrives on eth1 from an off-subnet client is answered out of eth0 via eth0's gateway — the reply carries eth1's source address but leaves through the wrong path. Clients on eth1's own subnet still work, because the connected route matches; anything beyond it breaks. The traffic is then dropped either by reverse-path filtering on this host in strict mode, or by a stateful firewall upstream that never saw the outbound half of the flow. The fix is policy routing: create a second table containing eth1's connected route and its own default, then add `ip rule add from <eth1-address> table <n>` so that packets sourced from that address are routed by it. Verify with `ip route get <client> from <eth1-address>`.
code
bash · 12 lines# name the extra table (optional, numbers work too)
echo '100 uplink2' >> /etc/iproute2/rt_tables
# the table needs the connected route and its own default
ip route add 10.20.0.0/24 dev eth1 src 10.20.0.50 table uplink2
ip route add default via 10.20.0.1 dev eth1 table uplink2
# route anything sourced from eth1's address by that table
ip rule add from 10.20.0.50 table uplink2 priority 1000
# verify the decision for a reply to an off-subnet client
ip route get 10.20.7.9 from 10.20.0.50go deeper
Know that Linux routes on destination alone and that one table holds one preferred default route, so a second NIC does not automatically answer through its own gateway.
Explain asymmetric routing concretely: which packet leaves which interface and why. Describe the roles of ip rule and a second routing table, and how a rule is keyed on source address.
Diagnose from the symptom — same-subnet clients fine, remote clients dead — and name both killers, strict rp_filter locally and stateful devices upstream. Build the second table with connected route plus default, and verify with ip route get ... from <addr>.
Decide whether a multi-homed host is the right design at all versus a single uplink with proper upstream routing, and if it is, how policy rules are declared, persisted and audited so the configuration is not a hand-typed artefact on one machine.
## The default model, and where it breaks Linux routing is destination-based. Given a packet to send, the kernel consults one table and picks the longest-prefix match. On a machine with one uplink that is entirely sufficient. Add a second NIC on a second network and the model quietly stops matching reality, because the answer to "how do I reach the internet" now has two valid answers and the table can hold only one *preferred* default route. Concretely: eth0 is 192.168.1.50/24 with gateway 192.168.1.1, eth1 is 10.20.0.50/24 with gateway 10.20.0.1, and `main` has one default route via 192.168.1.1. A client at 10.20.7.9 opens a connection to 10.20.0.50. The request arrives on eth1 correctly. The reply is generated with source 10.20.0.50 and destination 10.20.7.9 — and the routing decision looks only at 10.20.7.9, which is not inside 10.20.0.0/24, so it falls through to the default route and leaves via eth0. That is asymmetric routing, and it is precisely why the symptom is *partial*: clients on 10.20.0.0/24 itself work, because the connected route matches their destination. Everything beyond that subnet fails, which is why "the second network works from the same rack but not from the office" is the classic report. ## Why asymmetry actually kills the traffic Asymmetric routing is legal IP, so something has to break it: - **Reverse-path filtering on this host.** `net.ipv4.conf.<iface>.rp_filter` in strict mode (1) makes the kernel check, for each incoming packet, whether a reply to its *source* would leave the interface it arrived on. Here it would not, so the inbound packet is silently dropped — no log line, no counter most people look at. Two details matter: the kernel uses the **maximum** of the `all` and per-interface values, so relaxing only one has no effect; and mode 2 is loose validation, which accepts the packet if any route to the source exists. - **Stateful devices in the path.** A firewall or NAT device in front of eth1 sees the inbound SYN but never the outbound SYN-ACK, which departs by a different path entirely. It drops the return traffic or the follow-up packets as out of state. Either way, the flow is dead, the host's own logs show nothing, and the routing table looks correct because it is correct — for a single-homed machine. ## The fix: rules plus a second table The kernel supports many routing tables and a rule set that decides which one to consult. `ip rule show` on an untouched system prints three built-in rules: ``` 0: from all lookup local 32766: from all lookup main 32767: from all lookup default ``` Rules are evaluated by ascending priority; the first match whose table returns a route wins. Inserting a rule below 32766 lets it be consulted before `main`. Give the second uplink its own table. Table numbers work directly; names come from `/etc/iproute2/rt_tables` (or a drop-in in `/etc/iproute2/rt_tables.d/`) and only make the commands readable: ```bash echo '100 uplink2' >> /etc/iproute2/rt_tables # the second table needs both the connected route and its own default ip route add 10.20.0.0/24 dev eth1 src 10.20.0.50 table uplink2 ip route add default via 10.20.0.1 dev eth1 table uplink2 # consult that table for anything sourced from eth1's address ip rule add from 10.20.0.50 table uplink2 priority 1000 ``` The rule is keyed on **source address**, which is the whole trick: a reply generated for a request that arrived on eth1 already carries 10.20.0.50 as its source, so the rule catches it and it leaves via 10.20.0.1. Traffic the server itself initiates has no such source until routing assigns one, so it continues to use `main` and eth0 — which is normally what you want. Omitting the connected route in the new table is the usual mistake: with only a default route there, packets to 10.20.0.0/24 hosts get sent to the gateway rather than delivered on-link. ## Verifying it Do not infer, ask: ```bash ip rule show ip route show table uplink2 ip route get 10.20.7.9 from 10.20.0.50 ``` The last command is the decisive one — it runs the real lookup including the rule set for a packet with that source, and should report `dev eth1 via 10.20.0.1`. Without the `from` clause you are testing locally originated traffic, which will still choose eth0 and mislead you. ## Operational notes Everything above is runtime state and disappears on reboot, so the rules and the extra table must be expressed in whatever configures your interfaces. Relaxing `rp_filter` is a legitimate mitigation when asymmetry is genuinely intended, but on a dual-homed server it usually masks the problem rather than solving it: the reply still leaves the wrong interface, and any stateful device on that path will still object. Fix the path; do not disable the check that noticed it.
- Why do clients on the second NIC's own subnet still work while everyone else fails?Because their destination address matches the connected route for that subnet, which is more specific than the default route and points out of the correct interface. Only off-subnet destinations fall through to the single default route and leave via the wrong NIC, which is why the symptom looks intermittent or location-dependent.
- What does `ip rule` actually decide, and how are rules ordered?It decides which routing table a packet's lookup consults, based on selectors such as source address, incoming interface or firewall mark. Rules are evaluated in ascending priority, and the first one whose table yields a route wins. The built-in rules are local at 0, main at 32766 and default at 32767, so a custom rule needs a lower number to be consulted before main.
- Is turning rp_filter off an acceptable fix here?It can restore connectivity, but it only stops this host from noticing the asymmetry — replies still leave the wrong interface, and any stateful firewall on that path will drop them anyway. Treat it as a mitigation where asymmetry is genuinely intended, and remember the kernel takes the maximum of the `all` and per-interface values.
- Why must the second table contain a connected route as well as a default?Because a table is consulted on its own once its rule matches. With only a default route in it, packets destined for hosts on that NIC's own subnet would be handed to the gateway instead of delivered directly on-link. The table needs every route required to serve the traffic the rule steers into it.
saying these in an interview costs you the question
- Adding a second default route to the main table and expecting per-interface behaviour
- Believing the kernel remembers which interface a request arrived on
- Disabling rp_filter and calling the asymmetry fixed
- Creating the second table with only a default route in it
- Testing with `ip route get` but omitting the source address