Monthly floods make SYN cookies engage and silently cap distant customers' throughput. How do you decide, and who owns it?
answer
- one cost is loud, the other silent
- measure before you argue
- availability traded against quality
- a segment, a ceiling, hours per month
- someone named accepts the residual
basics
~20 sFrame it as a trade the business owns: refused connections for everyone during floods versus invisible throughput loss for a minority of distant customers. Make the invisible half measurable first, then get a named owner to accept the residual.
solid answer
~50 sThe engineering half is easy — leave the degraded path armed, because the alternative is refusing connections for everyone every time a flood lands. The judgment is what to do about the cost nobody can see. I would make it observable before arguing about it: record when the stateless path engaged, tag the connections admitted in those windows, and report throughput bucketed by client round-trip time, so the harm stops being anecdotal. Then take a bounded number to the account and contract owners — this many customers on long paths, this ceiling, this many hours a month. If a contract commits to throughput rather than availability, that clause decides it, not me. What I refuse is silence: the residual gets written down and a named owner accepts it, and if nobody will, the answer becomes capacity or an upstream absorber that somebody funds.
go deeper
Understand that a defence can succeed and still cost customers something, and that the cost may appear as slowness for some users rather than as an error anyone sees.
Be able to describe what you would measure to make an invisible degradation visible: when the degraded path engaged, which connections it admitted, and throughput split by client distance.
Show that you would quantify the affected segment and the hours involved before recommending anything, and that you can state the alternative cost — refused connections for everyone — in the same units.
Own the decision structure: name the trade, produce a bounded number for both sides, get a named owner to accept the residual explicitly, and turn a refusal into a funded case for capacity or upstream absorption.
## Why this is not an engineering decision Both options are defensible and both hurt someone. Leaving stateless admission armed means that whenever the queue of half-completed handshakes overflows, connections are admitted with degraded parameters — no window scaling, so a hard ceiling of 65,535 bytes in flight, which on a long path is a throughput cap the customer feels and nobody logs. Disabling it means that during the same minutes the queue simply fills and new connections are refused, for every customer, near and far, until the flood stops. One of those costs is loud, symmetrical and short. The other is quiet, concentrated on a minority, and long — it lasts as long as the affected connections do, which for pooled API clients can be weeks. Engineering can describe both. It cannot decide which one the business would rather pay, because that depends on who the distant customers are, what was promised to them, and what an outage costs. ## Make the invisible half measurable first An argument about an unmeasured harm is an argument about temperament, and the person with the strongest opinion wins. So the first move is instrumentation, and it is cheap: - A signal for when the degraded path engaged, and for how long. Without it you cannot even say how many hours a month this happens. - A tag on connections admitted during those windows, so you can count them and see how long they persist. - Throughput reported by client round-trip time band rather than in aggregate. The harm is invisible in an average and obvious once you split by distance, because the ceiling is arithmetic: about 1.7 Mbps at 300 ms, about 3.5 Mbps at 150 ms, unnoticeable under 20 ms. Note what is deliberately absent from that list: error rates. This failure produces no errors, so every alert you already own is blind to it. That fact is itself worth putting in front of an owner, because it explains why nobody reported the problem for a month. ## The number you take to the owner Multiply affected customers by the hours their connections stayed capped by the gap between their measured throughput and their normal throughput. It is an estimate, and you say so, but a bounded estimate ends the argument about whether the problem is real and starts the argument that actually matters. Put the alternative in the same units: floods of the size we actually see, times the minutes of refused connections, times all customers rather than the distant ones. Then hand the decision over, with the constraints named: - **Contracts.** If any agreement commits to throughput or to a performance level rather than to availability, that clause decides the question and the decision is not yours. - **Segment value.** Distant customers are frequently the ones who bought reach specifically. Losing throughput for the segment that pays for distance is not a rounding error. - **Measured flood sizes, not feared ones.** The case for disabling the mechanism only exists if there is enough headroom that realistic floods never overflow the queue. That is a measurement, and it changes over time, so it is a commitment to keep measuring rather than a one-off finding. ## What you do not do Do not present this as protection with no downside — the moment a defence is recorded as free, the residual loses its owner. Do not argue that the degradation does not count because nothing failed; that reasoning is exactly why it went unreported. And do not let the decision be made by default. The default in every estate is that the degraded path is armed, which is probably right, but a default that nobody chose is a risk nobody accepted. The outcome you are aiming for is unglamorous: the trade is written down, one named person accepts the residual, the invisible cost now has a graph, and if the owner will not accept it, that refusal becomes the business case for capacity or for absorbing the flood before it reaches your racks — which somebody then has to fund.
- What would change your recommendation to disabling it?A customer segment whose agreement commits to throughput on long paths and whose value exceeds the availability risk, combined with measured headroom showing that floods of the size you actually see never overflow the queue. That is a narrow case, and it has to rest on measured flood sizes rather than on how frightening the last incident felt.
- How do you stop this being invisible next time?Emit a signal when the degraded path engages, tag the connections admitted in those windows, and alert on segment-level throughput by client distance instead of on errors. This failure produces no errors at all, so any alerting built on failure signals will never see it, no matter how good it is.
- The account owner wants a number for the harm. What do you give them?Affected customers, multiplied by the hours their connections stayed capped, multiplied by the gap between the ceiling their path implies and their normal throughput. Say plainly that it is an estimate. A bounded estimate ends the argument about whether the harm is real and moves the conversation to what the business is willing to pay.
saying these in an interview costs you the question
- Treats it as a pure engineering call with no business owner
- Argues the degradation does not matter because nothing failed
- Promises protection instead of naming the residual
- Decides on fear of floods rather than measured flood sizes
- Assumes existing error-based alerting would have caught it