skip to content

What does stuck-in-active mean for an EIGRP route, how do SIA-QUERY and SIA-REPLY change the outcome, and what usually causes it?

level: seniorimportance: must knowfreq 18%

answer

  1. a reply that never comes
  2. a timer on every active route
  3. are you still working on it?
  4. the culprit is often hops away

basics

~20 s

An EIGRP route is stuck in active when a queried neighbour fails to reply in time; an SIA-QUERY asks whether it is still computing, an SIA-REPLY says yes, and a silent neighbour loses the route or its whole adjacency.

solid answer

~50 s

When an EIGRP route goes active, the router queries its neighbours and starts an **ACTIVE timer**. A route is **stuck in active** (SIA) when a neighbour's reply does not arrive in time. RFC 7868 splits the wait: at half the interval the router sends an `SIA-QUERY`, which the neighbour must answer with a `REPLY` or an `SIA-REPLY`; an SIA-REPLY with the active flag set says "still computing", so the wait is extended, and up to three SIA-QUERYs may be sent. A neighbour that answers neither in time is declared stuck, and DUAL either treats it as having replied unreachable for that route or deletes all its routes and resets the adjacency; the RFC notes one implementation does the latter. The timer values are implementation choices. The usual causes are a query domain so large that replies wait on long chains of routers, and lost packets, congested links or an overloaded router somewhere downstream.

go deeper

for a junior

Recall that stuck-in-active means a queried EIGRP neighbour has not replied in time, so the router cannot finish recomputing the route.

for a middle

Explain the active timer, what SIA-QUERY asks and what an SIA-REPLY confirms, and why SIA-REPLY is not a real answer.

for a senior

Show how you would trace the outstanding reply to the router actually at fault, and why shrinking the query domain prevents a repeat.

for a principal

Weigh the RFC's two responses, dropping one route or resetting the adjacency, against the blast radius each one has in a large domain.

## What "stuck in active" means An EIGRP route goes **active** when the router loses its successor and has no feasible successor. It sends a `QUERY` to its neighbours and must wait for a reply from **every** one before it can choose a new path. RFC 7868 bounds that wait: when the router goes active for a prefix it starts an **ACTIVE timer**. A route is **stuck in active** (SIA) when a neighbour has not answered within that limit. The RFC defines SIA as a destination that has stayed active longer than a predefined time and gives one implementation's value of 3 minutes. That figure, and the 90 seconds below, are implementation defaults, not protocol constants. ## The SIA-QUERY and SIA-REPLY exchange A missing reply has several possible explanations: a lost packet, a congested link, or a neighbour that is itself waiting on routers further away. EIGRP has two extra messages, which RFC 7868 describes, so the router can tell these apart instead of giving up blindly. | Message | Sent by | Meaning | |---|---|---| | `QUERY` | the active router | "What is your distance to this prefix?" | | `SIA-QUERY` | the active router, at half the active time | "Are you still working on my query?" | | `SIA-REPLY` | the queried neighbour | "Yes, I am still active for this prefix" (active flag set) | | `REPLY` | the queried neighbour | the actual answer, which ends that neighbour's part | The sequence the RFC describes: 1. The router sends a `QUERY` and starts the ACTIVE timer. 2. At half the interval (90 seconds in the implementation the RFC cites), it sends an `SIA-QUERY` for each destination still waiting. 3. The neighbour must answer with a `REPLY` or an `SIA-REPLY` within the next half interval. 4. The SIA-QUERY round resets the ACTIVE timer, and an `SIA-REPLY` with the active flag set shows the computation is still alive, so it may continue. The RFC allows up to three SIA-QUERYs for a destination. 5. An `SIA-REPLY` does **not** answer the original query. Only a `REPLY` lets the route go passive. If the neighbour answers neither in time, or is still active after the allowed rounds, it is declared stuck in active. ## What happens to a stuck neighbour RFC 7868 gives DUAL two options once SIA is declared: - **a)** delete the route from that neighbour, acting as if it had replied unreachable for that prefix; - **b)** delete **all** routes from that neighbour and reset the adjacency, acting as if it had replied unreachable for everything. The RFC notes that one implementation uses option b. That is the expensive one: every route learned through that neighbour disappears, other prefixes may go active, and more queries follow. One slow prefix can become a wide disruption. ## What usually causes it - **A large query domain.** In a big flat network a query can travel through long chains of routers, each waiting for the routers behind it. The originator waits for the slowest reply in the whole region. - **Lossy or congested links.** Queries and replies are delivered reliably, but repeated retransmission on a bad link delays the reply. - **An overloaded router.** A router short of processing time or memory may answer late, or many prefixes may go active at once after one failure. ## How to diagnose and prevent it 1. **Do not stop at the router reporting the problem.** It is often only waiting on a neighbour that is waiting on another. Follow the outstanding replies hop by hop until you reach the router that never answered or the link that lost the packets. 2. **Fix the local cause** there: the bad link, the overloaded device. 3. **Shrink the query domain** so this cannot recur: summarise prefixes at distribution points, filter what is advertised, and use stub routing, an implementation feature outside RFC 7868, on routers that never offer transit. 4. **Treat a longer active timer with caution.** It can let a slow but healthy computation finish, but it also delays the response to a neighbour that has really stopped answering.

  • Where do you look first when an EIGRP router reports a route stuck in active?
    Not at that router alone. Find which neighbour's reply it is still waiting for, go to that neighbour and see what it is waiting for in turn, and repeat until you reach the router that never replied or the link losing packets. The SIA-QUERY exchange helps here: a neighbour that keeps sending SIA-REPLYs is alive and waiting downstream, so the problem lies further along.
  • Why can resetting the adjacency do more harm than the stuck route itself?
    Resetting the neighbour removes every route learned through it, not just the stuck prefix. Those prefixes may lose their successor at once, go active together and send a burst of new queries, which raises the chance of further stuck-in-active events. RFC 7868 allows a narrower option, removing only the stuck route, but notes that one implementation chooses the adjacency reset.
  • Why were SIA-QUERY and SIA-REPLY added at all?
    Without them a router waiting on a reply cannot tell a dead or unreachable neighbour from one that is healthy but waiting on routers further away. The SIA-QUERY asks directly, and an SIA-REPLY proves the neighbour is alive and still working; RFC 7868 lets the router reset its active timer in that case, so a slow but healthy computation can finish.

saying these in an interview costs you the question

  • Stuck-in-active means the neighbour adjacency failed to come up in the first place
  • The router that reports stuck-in-active is always the one with the fault
  • An SIA-REPLY answers the original query and lets the route go passive
  • The 3-minute active timer and 90-second SIA-QUERY are fixed by the EIGRP specification
  • Raising the active timer as high as possible is the proper cure for stuck-in-active