Keyed hashing costs throughput on a hot path — how do you set fleet-wide policy against hash flooding?
answer
- is the cost assumed or measured?
- where should the default live
- one layer that does not depend on the hash
- exemptions decided by key provenance
- persisted hashes versus in-memory hashes
basics
~20 sMeasure the real cost before assuming it, make unpredictable hashing the default in shared parsing and container layers, and add a cheap cap on key counts so no team's oversight is fatal. Exempt only maps untrusted keys cannot reach.
solid answer
~50 sRefuse the premise until it is measured: keyed hashing of short keys is usually a small share of request cost next to parsing and I/O, and "too slow" is an assumption a benchmark on real key sizes often overturns. Then set policy by threat surface, not per service: any map built from data the service did not create gets the unpredictable hash, and that default lives in the shared parsing and container layer so no team has to remember. Add a key-count cap at the boundary, which bounds the quadratic term regardless of hash choice. Grant exemptions narrowly, with a written argument for why untrusted text cannot arrive at that map. Finally, sequence the migration: per-process keys make hash values unstable across processes, so anything persisted or used for sharding moves to a separate fixed function first.
go deeper
Take away the habit rather than the policy: assume any map built from data your service did not create needs the safe default, and do not disable a protection you cannot measure the cost of.
Be ready to produce the measurement that settles the argument — hash cost as a share of end-to-end request time on realistic key sizes — and to name the cheap orthogonal control, a cap on how many keys a request can create.
Show you would put the default in the shared parsing and container layer, keep one hash-independent limit, and sequence the rollout so persisted or shard-selecting hash values move to a separate fixed function first.
Own the tradeoff end to end: a measured cost spread across every team versus an asymmetric availability cliff, a secure-by-default platform decision instead of a recurring audit, and exemptions that carry evidence and an expiry date.
## The decision, not the mechanism Every reasonable engineer already knows the menu: unpredictable keyed hashing, caps on key count, cost-bounded buckets. The principal-level question is which of them you make mandatory for a hundred-odd services maintained by a few dozen teams, given a latency budget, a migration cost, and a finite amount of attention. ## Step 1: measure, because the premise is usually softer than it feels "Keyed hashing is expensive" is an intuition formed from cryptographic hashing of large payloads. Table keys are short — tens of bytes — and modern keyed functions designed for short inputs run at a few cycles per byte. Benchmark on your real key distribution and on the real request path, and express the answer as a share of end-to-end request cost rather than as a microbenchmark ratio. Frequently the honest number is a fraction of a percent of a request that also parses text, touches a socket, and waits on a database. That number, not a debate, should decide the policy. Where the cost is genuinely material is a narrow class: extremely hot in-memory maps in tight loops, keyed by short internal identifiers, called millions of times per second. Those exist. They are also usually the maps that never see untrusted text — which is why the exemption criterion should be *provenance of the keys*, not *hotness*. ## Step 2: place the default where forgetting is safe A policy that requires every team to remember something will be violated by the newest service on its busiest week. So the default belongs in the layers everyone already uses without thinking: the request parsing layer, the document decoder, the shared container defaults. A team should have to take deliberate action to become vulnerable, not deliberate action to become safe. This is also the answer to the audit-cost problem. Auditing "which maps see untrusted keys" is not a one-time exercise; it is a question re-opened by every refactor that widens what flows into a cache key. A secure default converts a recurring audit into a one-time platform decision. ## Step 3: keep an orthogonal layer that does not depend on the hash Cap the number of fields a request may produce, cap body size, and give requests a CPU budget that a runaway handler cannot exceed. These bound the quadratic term directly and, crucially, they fail *independently* of the hash decision. If a hash key leaks, is accidentally shared fleet-wide, or is derived from something predictable, the caps still hold. Different failure modes are the whole justification for depth here, and they are why the answer to "which one do we do" should usually be "the cheap one and the strong one". ## Step 4: exemptions with evidence, and an expiry A workable exemption process: a service may opt a specific map out of unpredictable hashing if it documents (a) a measurement showing material end-to-end cost, and (b) an argument that only self-minted identifiers reach that map. Exemptions carry a review date, because the second half rots — the map that only held internal identifiers acquires a client-supplied cache dimension two quarters later. ## Step 5: sequence the migration around what depends on stable hashes The change that surprises people is not throughput; it is that per-process keys make hash values non-reproducible. Anything that assumed stability breaks: iteration order varies run to run, and, more seriously, hash values written to storage, used to select a shard, or exchanged between processes are no longer meaningful. The correct structure is two functions with two jobs — a fixed, documented function for persistence and partitioning, and an unpredictable per-process one for in-memory tables. Doing the separation first, and the key randomization second, is the difference between a quiet rollout and a data-placement incident. ## Step 6: prefer boring mechanisms your team can keep A bespoke cost-bounded container written in-house is a legitimate engineering answer and a poor organizational one if a single engineer understands it. Weigh the ongoing cost of the clever structure against the cost of the attack it prevents and against the simpler layers available. The judgment interviewers are listening for is not "which is fastest" but "which set of defenses will still be correct after two years of staff turnover". ## How to say it in an interview Name the tradeoff explicitly: an availability risk with an asymmetric cost curve, mitigated by measures whose price is small but non-zero and spread across every team. Decide by measurement, default to safe in shared code, keep one hash-independent layer, allow narrow evidence-backed exemptions, and sequence the rollout behind the separation of persisted from in-memory hashing.
- Which internal maps can safely keep a predictable hash?Ones whose key space you fully control — configuration keys, identifiers your own system mints — and only when no untrusted text can reach them through a parser, a cache dimension, or a log-derived key. Because that property erodes with each refactor, the safe institutional answer is to default secure and treat exemptions as reviewable.
- Why isn't a request-size cap a complete answer on its own?It bounds n and therefore the quadratic term, which is real, but only where it is enforced. Nested documents, repeated headers, batch endpoints and downstream caches each build maps of their own, and one uncapped path restores the full exposure. Caps are an excellent independent layer, a fragile sole layer.
- How would you justify the spend to a team that has never been attacked?Frame it as an asymmetry, not a probability: a small request buys the attacker quadratic CPU, so the exposure is a capacity cliff rather than a slow degradation. Then show the measured cost of the default — usually a fraction of a percent of request time — so the argument becomes cheap insurance rather than a security lecture.
saying these in an interview costs you the question
- Assumes keyed hashing is too slow without measuring
- Leaves the choice to each service with no default
- Treats a size cap as the complete answer
- Ignores that per-process keys break persisted hash values
- Adopts a bespoke structure only one engineer understands
- Grants exemptions based on hotness rather than key provenance