Your DDoS appliance logs zero mitigations while an attacker's crafted export requests keep an API down - what do you tell the service team that filed the ticket, and where must the control move?
answer
- answer with evidence, not a refusal
- same window, both sides of the boundary
- no threshold setting fixes a unit mismatch
- name the next owner explicitly
- admission in front of the expensive call
basics
~20 sAnswer with evidence, not a refusal: show border and service counters for the same window, explain why a traffic-rate meter cannot see per-request cost, and hand the control to an admission check in front of the expensive call.
solid answer
~50 sI would answer with evidence rather than a jurisdictional no. Put the two counter sets side by side for the same window: bits, packets and connections per second orders of magnitude below any threshold, against a shared database pool fully busy with a deep queue and p99 latency in the tens of seconds. That establishes the attack is real and that the boundary is not where it can be seen or stopped - the device that can drop cannot price the request, and price is the only thing separating these requests from a customer's quarterly report. Then name the next owner instead of closing the ticket: the control becomes an admission check in front of the expensive call - bounded parameters, a cost class, a cap on concurrent expensive work - owned by the service team, with the network side's real contribution stated honestly. A refusal with no destination bounces straight back.
go deeper
Know that a mitigation appliance reporting zero events is a statement about its own counters, and that the right first move is to compare border and service numbers for the same window.
Be able to explain to another engineer why no threshold change catches this, using the difference between a traffic rate and per-request work.
Show the whole handoff: correlated evidence, a clear statement of what the control cannot do, a specific replacement control, its cost, and a named owner.
Own the standing gap between what the mitigation control covers and what availability commitments assume, and decide where the second control is funded and who staffs its exceptions.
## The situation, stated plainly Every network graph is flat. The mitigation appliance reports zero events and is behaving correctly. The service is unusable. A ticket lands on the appliance owner's queue that says, in effect, 'we are being attacked, please block it', and the honest answer is that it cannot be blocked where the requester wants it blocked. How that answer is delivered decides whether the outage gets fixed this week or the ticket ping-pongs for a month. ## Step one: prove the attack is real The worst version of this answer is 'nothing on our side', because it reads as a denial that anything is happening. Lead with correlated evidence from both sides of the boundary, in the same time window: - Border: ingress bits per second, packets per second and new connections per second, each shown against its configured threshold, plus the empty mitigation log. - Service: request count barely changed, p99 latency enormous, shared pool fully busy with a deep queue, timeouts at the client. That pairing does the entire argument in one screen. It says the attack is real, the damage is downstream, and the border numbers are not evidence of health - they are evidence that the border was never given a measurement that could carry this. ## Step two: say what the appliance measures, and why this escapes it One sentence, no jargon: the device meters traffic rate, the damage is work per request, and a few hundred well-formed requests a minute move no traffic-rate counter. Add the consequence explicitly, because it is the part people resist: there is no threshold setting that fixes this. Any threshold low enough to fire would also fire on an ordinary busy hour and on any legitimate launch, and it would still be measuring the wrong quantity. ## Step three: name the next owner, and be specific This is the step that gets skipped and the reason tickets bounce. Refusing without a destination is read as territoriality. The destination follows from the vantage argument: the control has to sit where the request's cost is knowable *and* before the scarce resource is committed. In practice that means an admission decision in front of the expensive downstream call, owned by the service team: - Bound what can be asked - a maximum range, a mandatory tenant filter, a bounded result size, a fixed set of report shapes. - Classify requests into cost classes from their declared shape and cap how many expensive ones run at once, so the shared pool cannot be fully consumed by one class. - Move exports off the inline path onto a bounded asynchronous one. Say what this costs, because the receiving team will discover it anyway: added latency and a policy lookup on every ordinary request, a new component in the hot path, and a standing exception list once real customers hit the bounds. ## Step four: state what the network side genuinely gives Refusing prevention does not mean refusing to help, and the credible answer distinguishes the two: | The network side can | The network side cannot | | --- | --- | | Supply per-source request and connection records showing who is asking | Decide from the wire that this query will cost thirty seconds | | Terminate or block specific sources during a live incident, as a stopgap | Keep that working once the attacker moves | | Provide a clean place to insert an application-aware admission point | Enforce a cost policy it cannot evaluate | | Absorb the volumetric attack that may follow this one | Substitute for an admission control at the service | The stopgap deserves care. Blocking source addresses works exactly as long as the attacker stays put, and it damages you when addresses are shared behind one egress - you will drop legitimate customers to buy hours. Do it during the incident if it buys time, and say out loud that it is not the control. ## Step five: write the handoff so it survives without you The artefact is short and it is the deliverable: the two counter sets, one sentence on what the appliance measures and why it cannot see this, the specific control being requested, the named owner, and an offer to join the design review. If an on-call engineer reads that ticket at three in the morning next month, they should reach the same conclusion without re-running the analysis. ## The failure modes to avoid Promising to tune the appliance to buy goodwill is the worst of them: it accepts an obligation you cannot meet and it delays the real fix by however long the tuning takes to fail. Closing the ticket as not-a-network-problem is nearly as bad. Blocking sources and declaring the incident resolved leaves the service exactly as vulnerable to the next set of addresses, and it converts a temporary measure into an unowned permanent one that quietly blocks real customers.
- The service team says just block the source addresses. What do you tell them?That it works exactly until the attacker moves, and that it hurts you when addresses are shared behind one egress. Source blocking has no relationship to a request's cost, so it catches this attacker only while he stays still and takes out legitimate callers alongside him. I will do it as an incident stopgap and say plainly that it is not the control.
- How do you keep the ticket from bouncing straight back to you?Write the handoff so it stands without you: the two counter sets that prove the attack, one sentence on what the appliance measures and why it cannot see this, the specific control being asked for, and a named owner. Offer to join the design review. A refusal with no evidence and no destination is read as territoriality and returns within a day.
- Leadership asks whether the mitigation service you pay for was worth it. How do you answer?It did the job it is scoped for and it has a class of outage it cannot see. Say both: it absorbs volumetric attacks, which are still the common case and would have taken the site off the internet, and it cannot price a request. The correct conclusion is a second control at the service, not cancelling the first.
saying these in an interview costs you the question
- Closes the ticket as not-a-network-problem with no evidence
- Promises the appliance can be tuned to catch it
- Blocks source addresses and calls the incident resolved
- Refuses the ticket without naming who owns the control next
- Presents flat border graphs as proof the service is healthy