skip to content

How would you size a stateful firewall's session table, and why is the idle timeout the cheapest lever?

level: middleimportance: should knowfreq 50%

answer

  1. concurrency is rate times lifetime
  2. most entries are idle, not busy
  3. you cannot cap who arrives
  4. shorten lifetime, not arrivals
  5. idle flows pay in silent reconnects

basics

~20 s

Concurrent entries are roughly the new-flow arrival rate multiplied by how long an entry lives. You cannot cap arrivals from an unbounded population, so shortening the idle timeout is the only lever that shrinks the table without buying hardware.

solid answer

~50 s

Size from measurement, not from the datasheet: peak new-session rate times mean session lifetime gives steady-state concurrency. Lifetime is the half you can influence. Most entries are not busy — they sit idle after the application has finished with them and are reclaimed only when the idle timer expires — so halving the established idle timeout roughly halves that tail at zero capital cost. The price is paid somewhere else: a flow that legitimately idles longer than the new timeout is torn down silently and the client finds out on its next write, so you are pushing keepalives onto application teams whose roadmap you do not own. The same timer also bounds a badly-behaved client that opens connections and never closes them, so it is a security lever as well as a capacity one. Size for peak, then add headroom for the failover case where one member carries everything.

go deeper

for a junior

Know that a firewall's capacity limit is a number of simultaneous connections, and that it is a different limit from bandwidth. Be able to say that entries are released by a timer when nothing closes the flow.

for a middle

Explain concurrency as arrival rate times lifetime, name the separate embryonic, established-idle and close timers, and say why UDP entries always live out their timeout.

for a senior

Demonstrate that you would measure before changing anything, and that you know what a shortened idle timeout breaks in production and who has to fix it. Include failover headroom in the number you quote.

for a principal

Frame the timeout as a shared decision rather than a firewall setting: it trades your capacity against another team's keepalive work, and it also bounds how long a badly-behaved client can hold your resources.

## The arithmetic A session table is a queue in the accounting sense, so the standard relationship applies: **mean concurrent entries is approximately the arrival rate of new flows multiplied by the mean lifetime of an entry**. Ten thousand new flows per second at a mean lifetime of sixty seconds is six hundred thousand concurrent entries. Both terms have to be measured, and both are routinely guessed wrong. **Arrival rate.** Measure the connections-per-second counter, not bandwidth. Bandwidth and concurrency are separate ceilings and they are reached by different traffic: a video stream is enormous bandwidth in one entry, while a chatty API client is trivial bandwidth in thousands. Take the peak, over a period long enough to include a busy day, not the average. **Lifetime.** This is where the surprise lives. An entry's lifetime is not how long the application was busy; it is how long until something releases the entry. TCP usually releases early because a FIN exchange or a reset tells the firewall the conversation is over. When nothing tells it, the entry lives until an idle timer fires. ## Why timeouts are the lever You have exactly three ways to shrink a table: reduce arrivals, reduce lifetime, or stop tracking some flows at all. You cannot reduce arrivals. The population that allocates entries is every source your policy permits, and for a published service that is unbounded — customers, crawlers, scanners, retrying mobile clients, and anyone else who sends a packet. Nothing in the rule base caps the count without also refusing the business. So lifetime is the lever, and it is free in capital terms. Firewalls keep several timers, and they behave differently: | timer | governs | typical scale | who it reclaims | |---|---|---|---| | embryonic / handshake | TCP flows that have not completed | seconds to tens of seconds | connection attempts nobody finished | | established idle | open flows with no traffic | often an hour or more by default | pooled and long-lived connections | | close / TIME-WAIT-ish | flows after FIN or reset | seconds | flows that closed politely | | UDP idle | every UDP conversation | tens of seconds | everything, because UDP never closes | Halving the established idle timeout roughly halves the long tail of entries, and the tail is usually most of the table. The embryonic timer is the one to check first when a table is filling with flows that never became anything. ## What shortening a timeout costs This is the half candidates skip, and it is the half the interview is about. A flow that idles longer than the new timeout has its entry removed. Neither endpoint is told. The firewall now has no entry for a conversation both endpoints still believe is open, so the next packet from either side is an unsolicited mid-stream packet and is dropped or reset. The application discovers this only when it next tries to write, which is why the failure shows up as a delayed error at the worst possible moment rather than at the moment you changed the timer. The classic victims are pooled connections: an application server keeping database connections open and unused between bursts, an administrative session left open on a laptop, a long-poll or message-bus connection that is quiet by design. The fix is keepalives at the application or socket layer, which cost you packets instead of entries — but the change lives in somebody else's codebase and somebody else's release train. That is why a timeout change is a negotiation, not a knob. ## The adversarial half The same timer is what bounds how long a client that behaves badly can hold your capacity. A client that opens a connection and abandons it without closing — a mobile release that leaks sockets, a partner integration that retries without cleanup — parks an entry for the full idle timeout, and nothing in your rule base objects, because the traffic is permitted. Shortening the timer shortens the hold. It is one of the few settings that is simultaneously a capacity decision and a security decision. ## The third lever, and its price You can exempt a flow class from state tracking entirely — a bulk replication or backup path between two known addresses is the usual candidate. You get the entries back, and you give up exactly the property you were paying for: an untracked flow is permitted by a static match on header fields the sender chooses, in both directions, with no check that the traffic was invited. It is a defensible trade for a fixed pair of endpoints; it should be written down with the name of whoever accepted it. ## Headroom Finally, size for the failure case. In a high-availability pair, one member may end up carrying the whole estate, so a table sized at the pair's combined peak leaves you exactly nothing on the day one member goes away. Add headroom for that, plus growth to the end of the hardware's refresh cycle, because the capacity you buy has to last the life of the box.

  • Which timeout would you shorten first, and which one would you leave alone?
    Shorten the embryonic timer for half-formed TCP connections first: it is short already, nothing legitimate needs a handshake to hang around, and abandoned attempts are exactly what it reclaims. The established idle timer shrinks the table most but is also the one that breaks pooled and long-idle flows, so it moves only once you know which applications hold idle connections and can add keepalives.
  • Why is UDP traffic harder to size for than TCP?
    TCP tells the firewall when a flow is finished — a FIN exchange or a reset releases the entry early, so most entries never reach the idle timer. UDP has no close, so every entry lives out its full idle timeout regardless of whether the exchange took a millisecond. The same rate of short UDP conversations therefore produces far more concurrent entries than TCP does.
  • What does it cost to exempt a flow class from state tracking to save entries?
    You recover the entries and lose the property you bought the box for. The exempted class is permitted by a static match on header fields the sender picks, in both directions, with no check that the traffic was invited, so anything that fits the match passes. It is a reasonable trade between two fixed replication endpoints and a bad one for anything facing an open population.

saying these in an interview costs you the question

  • Sizes from average load rather than measured peak
  • Measures bandwidth instead of new sessions per second
  • Thinks buying a bigger box is the only lever
  • Forgets headroom for one member carrying the pair
  • Shortens idle timeouts without warning application owners

context