skip to content

Finance wants the firewall SKU with half the session capacity. How do you defend the larger one?

level: principalimportance: nice to knowfreq 30%

answer

  1. translate entries into a customer outcome
  2. the ceiling refuses new, keeps old
  3. arrivals are not yours to cap
  4. headroom for one member alone
  5. name who accepts the residual risk

basics

~10 s

Translate the number into a business failure. The ceiling counts concurrent flows, the population claiming them is unbounded, and at the ceiling the box refuses new connections while existing ones look perfectly healthy.

solid answer

~50 s

Do not argue in vendor units. Say what the number bounds — concurrent flows — and who fills it: every source the policy permits to reach a published service, an unbounded population no rule of yours can cap. Then state the failure in their language: the box does not slow down, it refuses new sessions while established ones keep working, so the outage reads as "no new customer can sign in" and gets misdiagnosed as an application fault for the first hour. Show the arithmetic — measured peak session-creation rate times mean lifetime, headroom for the failover case where one member carries everything, and growth across the refresh horizon, because you are buying for four years and not for Tuesday. Then offer the cheaper paths honestly: shorter timeouts, untracked bulk flows, or the smaller box with a documented ceiling, an alert and a named person who accepted the risk.

go deeper

for a junior

Know that a firewall's session capacity is a purchased ceiling on simultaneous connections and that it is separate from bandwidth. Be able to name the symptom of hitting it.

for a middle

Be able to derive the number from a measured session-creation rate and mean lifetime, and to explain why the box has no graceful degradation as it approaches the ceiling.

for a senior

Show that your sizing includes failover headroom and growth to the refresh horizon, and that you would bring utilisation history rather than a datasheet quote.

for a principal

Own the translation from a capacity ceiling to a revenue failure mode, put the cheaper alternatives on the table yourself, and make sure a decision to under-buy is recorded with a named owner rather than absorbed by the network team.

## The mistake that loses this argument The losing move is to quote the datasheet. "We need the four-million-session model" is a number with no meaning to anyone outside the network team, and it invites the obvious counter: the other one is cheaper and also has a big number on it. You have to translate before you ask. ## What the number actually bounds, and who spends it The capacity is a count of **concurrent conversations**, and it is usually the model or licence tier that sets it. The part worth saying out loud is that this budget is not spent by your architecture; it is spent by whoever sends packets to services you deliberately published. A policy that permits any source to reach the front door — which is what publishing a service means — names an unbounded population, and each member of it claims a slot with one packet. You do not get to cap that with a rule, because the rule is the front door. So the honest sentence to a budget holder is: *this number is the only thing standing between an ordinary increase in arrivals and a refusal to accept new business, and we do not control arrivals.* ## State the failure mode in their language This is where the argument is won, because the failure mode is genuinely nasty and genuinely explicable without jargon. A firewall at its ceiling does not get slow. It keeps every conversation that already has a slot running perfectly and refuses to start new ones. What the business sees is: - existing sessions fine, dashboards green, bandwidth graphs unremarkable; - new sign-ins, new checkouts and new API clients failing; - an hour spent looking at the application, because that is what the symptom looks like. A budget holder can price that. "Revenue that depends on new connections stops, and it takes us longer than it should to find out why" is a sentence they can act on; "we would exceed our concurrent session limit" is not. ## Bring the arithmetic and bring the evidence Come with four numbers, not one: 1. **Measured peak session-creation rate** — connections per second at the busiest hour, not bandwidth. The two are unrelated ceilings and people conflate them constantly. 2. **Mean entry lifetime**, so the concurrency figure is derived rather than asserted. 3. **Failover headroom.** In a high-availability pair, one member may carry the whole estate. A table sized exactly at the observed peak has nothing left on the day you need it most. 4. **Growth to the end of the refresh horizon.** You are buying capacity for the life of the box, typically several years, and it is far cheaper to buy it once than to replace a platform mid-cycle. Bring a year of utilisation history with the peaks marked, and be specific about what caused past jumps. The most instructive ones are usually not attacks at all: a mobile client release that opened connections and abandoned them without closing, a partner integration that retried aggressively, a marketing event. Those are the events that make the case, precisely because nobody malicious was involved and nobody could have been blocked. ## Offer the cheaper options rather than refusing Credibility is the currency here, and it is spent by asking for the top model "to be safe". Put the alternatives on the table yourself: - **Shorten idle timeouts.** Free in capital terms; costs application teams keepalive work you do not control, and silently breaks long-idle flows if you do it unilaterally. - **Stop tracking a bulk flow class** between fixed endpoints. Recovers entries; gives up the invited-traffic guarantee for that class, which somebody must accept in writing. - **Move the boundary.** A stateful control facing an unbounded internet population is the expensive case. The same control between two internal zones faces a population you own and can count, so it can be a much smaller box. Sometimes the right answer is a cheaper, coarser control at the edge and the stateful spend placed inside where sizing is a known quantity. - **Buy the smaller model deliberately**, with a documented ceiling, an alert well below it, a named owner and a pre-agreed purchase trigger. ## Then let them decide, and make the decision visible If the smaller box is defensible on the evidence, say so. If it is not, say what you cannot promise: you cannot bound the arrival rate, so the residual risk has a trigger outside your control, and it needs an owner with a name attached — ideally the person accountable for the revenue that stops when new connections are refused. Write the acceptance down. That is the whole principal move. The technical answer is a sizing calculation any competent engineer can do. The judgment is translating a capacity ceiling into a business failure mode, offering the honest cheaper paths, and ensuring that a decision to under-buy is recorded as a decision somebody made rather than a surprise you absorb alone at three in the morning.

  • What evidence would you actually bring to that meeting?
    A year of session-table utilisation with the peaks marked and explained, the measured session-creation rate rather than bandwidth, the failover calculation showing what one member has to carry alone, and a growth line to the end of the refresh cycle. If that evidence supports the smaller box, say so — the credibility you keep is what makes the next request land.
  • Finance buys the smaller box anyway. What do you do on Monday?
    Record the ceiling as an operational limit, alert on utilisation well below it with a named owner, shorten whichever timeouts you can shorten and tell the application owners why, and write down who accepted the residual risk. The aim is that the exposure becomes visible and owned rather than silently yours.
  • Is a second smaller pair a real alternative to one larger pair?
    Sometimes. Splitting the estate splits the table and the blast radius, which is genuinely useful when the two halves are independent zones. But you pay twice for operations, policy consistency and licences, and each pair still needs its own failover headroom, so two halves rarely cost less than one whole. It is a poor answer when it is only an accounting trick.

saying these in an interview costs you the question

  • Argues in datasheet numbers with no business translation
  • Asks for the largest model just to be safe
  • Sizes for today's peak with no failover headroom
  • Confuses bandwidth capacity with session capacity
  • Leaves the residual risk unowned and unrecorded

context