Why does a new BGP session sit in the Active state and never reach Established, and what does Active actually mean?
answer
- the name misleads
- no TCP connection yet
- below BGP: reachability, addresses, port 179
- Connect and Active alternate on a timer
basics
~20 sIn BGP, Active means the speaker has no TCP connection to its peer yet: it listens, and redials when ConnectRetryTimer expires. Stuck there means TCP never completes: no route, wrong peer or source address, port 179 filtered, or mismatched authentication.
solid answer
~50 sDespite the name, the BGP `Active` state is not a working session: RFC 4271 defines it as the speaker trying to acquire the peer by listening for and accepting a TCP connection, and each time `ConnectRetryTimer` expires (suggested 120 s) it goes to `Connect` and dials again. A session that cycles between `Connect` and `Active` therefore never got a TCP connection to port 179 in either direction, so the causes sit below BGP: no route to the peer's address, the neighbour address mistyped, the peer not configured for this speaker's address so it rejects the connection, the speaker sourcing from an address the peer does not expect, a filter dropping TCP 179, or a TCP MD5 key set on one side only. If TCP does come up but the session falls from `OpenSent` back to `Idle`, the problem has moved into the OPEN exchange instead.
go deeper
Recall that BGP Active means there is no working session yet, only a speaker waiting for a TCP connection. The name suggests the opposite.
Explain the Connect and Active pair, the ConnectRetryTimer between them, and why a session stuck there points to TCP rather than to BGP's own messages.
Show a diagnostic order: reachability from the source address, segments to port 179, matching peer and source addresses on both sides, authentication, then multihop for eBGP. Tell this apart from an OPEN rejection.
Talk about removing the whole class of fault: peering templates that pair addresses on both ends, filters that permit port 179 between known peers, and monitoring that alerts on sessions not Established.
## What the Active state means A **BGP speaker** keeps a finite state machine for each configured peer, defined in **RFC 4271**: Idle, Connect, Active, OpenSent, OpenConfirm and Established. The name **Active** misleads people into thinking the session is working. RFC 4271 defines Active as the state in which the speaker is "trying to acquire a peer by listening for, and accepting, a TCP connection." No BGP message has been exchanged yet. Its partner is **Connect**, the state in which an outbound TCP attempt to the peer's port 179 is in flight. The **ConnectRetryTimer** paces the pair: when it expires in Active, the speaker dials the peer again and moves to Connect. RFC 4271 suggests 120 seconds for it. Implementations differ on exactly which state a failed outbound attempt lands in (RFC 1771 moved it to Active; RFC 4271 moves it to Idle unless the optional DelayOpen timer is running), but every one of them shows a session that has never completed TCP as Connect or Active. That is why "stuck in Active" is the classic BGP troubleshooting question: the answer is almost never in BGP itself. ## The scenario A new eBGP session is configured between a speaker in AS 64500 at 192.0.2.1 and a speaker in AS 64501 at 192.0.2.2. Hours later it still shows Active, with an occasional flicker to Connect. Both speakers listen on port 179 and both try to dial, so for the session to stay down, **no TCP connection to port 179 has completed in either direction.** ## The usual causes Each cause below stops the TCP connection before any OPEN can be sent: - **No route to the peer's address.** The speaker cannot reach 192.0.2.2 at all, or reaches it only by a path that does not return. Sessions between loopback addresses are the common case: the loopback must be reachable through some other routing source before BGP can use it. - **The wrong peer address.** The neighbour address is mistyped, so the speaker dials a host that does not answer, or answers with a reset. - **The peer is not configured for this speaker.** In Idle, RFC 4271 refuses all incoming connections for a peer; and a connection request from an address the receiver does not recognise is an invalid connection that it rejects. A one-sided configuration therefore never completes. - **The wrong source address.** The speaker dials from its interface address while the peer has been configured with the speaker's loopback, or the other way round. The peer sees a connection from an address it has no peering for and rejects it. - **TCP port 179 filtered.** A packet filter on either speaker or on the path drops TCP segments to or from port 179. Because both sides dial, a filter that blocks only one direction can still leave one direction working; a filter that blocks both stops the session completely. - **A TCP authentication mismatch.** RFC 4271 requires implementations to support the TCP MD5 signature option. If a shared key is configured on one side only, or the keys differ, the receiving TCP does not accept the segments and the handshake never completes. - **An eBGP peer that is more than one hop away.** Implementations commonly send eBGP packets with an IP TTL of 1 unless the session is configured as multihop (an implementation default, not an RFC 4271 rule); when the peer is several hops off, the TCP segments expire on the way. Multihop sessions belong to the eBGP and iBGP sessions subject, and are named here only as one more cause. ## How to narrow it down 1. **Check reachability from the right address.** Test the path to the peer's address from the address the session is sourced from, not from whatever the default would be. 2. **Look at what TCP sees.** Watch for segments to and from port 179. SYNs leaving with nothing coming back point at routing, a filter or authentication; a reset coming back points at the peer refusing the connection. 3. **Compare both configurations side by side.** The peer address on each side must equal the source address of the other, and authentication must match. 4. **Read the state the session reaches.** Use the table below. | What the session does | Where the problem is | |---|---| | sits in Connect or Active | TCP to port 179 never completes | | reaches OpenSent, then drops to Idle with a NOTIFICATION | the OPEN is rejected: AS number, hold time, BGP Identifier or a capability | | reaches Established, then drops on a timer | hold timer expiry or a lost path to the peer | ## Do not borrow the word The BGP Active state is not an EIGRP route that is active (that router is querying its neighbours for a new path) and not a VRRP router in the Active state (the one forwarding for a virtual address). In BGP, Active means waiting for TCP.
- If both BGP speakers dial each other, why can a filter that blocks only one direction leave the session working?RFC 4271 has each speaker listen on port 179 and also connect to its peer. If a filter drops the segments of the connection one speaker opens, the connection the other speaker opens can still complete, and once TCP is up it does not matter which side initiated. A session that is passive on one side loses that redundancy: only one direction is ever tried.
- How do you tell a BGP session stuck in Active from one rejected during the OPEN exchange?A session stuck in Active or Connect never completes TCP, so no OPEN or NOTIFICATION is ever exchanged. A rejected OPEN means TCP did complete: the session reaches OpenSent, a NOTIFICATION with the OPEN Message Error code and a subcode such as Bad Peer AS is sent, and the FSM falls back to Idle before the next retry.
A BGP session in Active is like a phone that is switched on and waiting for a call, with the caller redialling every two minutes. The phone works; the line between the two parties never connects. Fix the line, not the conversation.
saying these in an interview costs you the question
- A BGP session in Active is up and actively exchanging routes.
- Stuck in Active means the peer rejected the AS number in our OPEN.
- Only one BGP speaker ever opens the TCP connection; the other just listens.
- Active in BGP is the same as an EIGRP route in the active state.
- If ping to the peer works, port 179 and the session addresses must be fine.