skip to content

How much should an architecture hedge against the hardness assumption beneath its security turning out to be false?

level: principalimportance: should knowfreq 33%

answer

  1. sort the tail risks by plausibility
  2. one failing beats all failing
  3. how long must the data stay secret
  4. buy replaceability, not prophecy
  5. negotiation invites a downgrade

basics

~20 s

Hedge against one assumption falling, not against the whole edifice collapsing. Buy replaceability, assumption diversity and an explicit secrecy lifetime for the data; do not buy insurance against a universal collapse, which would have no product-level answer anyway.

solid answer

~40 s

Separate the tail risks by plausibility. A constructive, practical collapse of the two classes would break essentially every computationally secure system at once; there is no design you could ship that survives it, so it belongs in a risk register as a shared, unmitigable event rather than in a backlog. The plausible risk is narrower and real: one specific problem yields to an algorithmic advance or a different machine model while the rest hold. Hedge that. Name the assumption explicitly in the design, keep the primitive behind a replaceable boundary, carry an algorithm identifier on the wire so a successor can be negotiated, and write down how long the data must stay secret, because captured traffic can be attacked years later. Then stop — over-hedging buys new failure modes, including downgrade paths.

go deeper

for a junior

The takeaway is that these guarantees rest on unproven conjectures, not theorems. Knowing that the security of a system is an assumption someone chose is the first step.

for a middle

Be able to explain why a replaceable boundary and an identifier stored alongside the data are what make a future migration possible at all.

for a senior

Drive the design to name its assumption, bound its blast radius, and state the secrecy lifetime of the data it protects. Watch the negotiation mechanism for downgrade paths.

for a principal

Own the judgment about where the budget goes: convert an unmanageable tail risk into managed properties — replacement speed, assumption diversity, and lifetime analysis — and defend that allocation explicitly.

## Name the risk you are actually managing The question sounds like one risk and is really three, with wildly different probabilities and completely different responses. 1. **A constructive, practical collapse of the classes.** Every computational security assumption dies simultaneously. So does everyone else's. There is no product-level mitigation, and a design document claiming one is not credible. 2. **A non-constructive or galactic collapse.** The assumption is known false while no usable attack exists. The response is not technical on day one; it is to start treating the primitive as end-of-life and to make sure you *can* move. 3. **One assumption falls on its own.** An algorithmic advance against a single problem, or a physically different machine model that attacks one structure and not others. This is the historically ordinary outcome, and the only one worth engineering against directly. A lead's job in the review is to say that out loud, because teams routinely spend the anxiety of the first risk on the third's budget, or the reverse. ## What hedging looks like when it is worth paying for - **Write the assumption down.** The design should contain a sentence naming exactly what must stay hard. An unnamed assumption cannot be reviewed, monitored or retired. - **Keep the primitive behind a boundary.** One module that everything calls, so replacement is a contained change rather than an archaeology project across the codebase. - **Carry an identifier on the wire and in storage.** Stored data and messages should say which construction produced them, so two can coexist during a migration. Without it, a change is a flag day. - **Write down the secrecy lifetime.** Data captured today can be attacked with tomorrow's method. If a record must stay confidential for a decade, the assumption must hold for a decade, and that is a far stronger requirement than 'secure today'. - **Prefer assumptions with a long public attack history.** Age is the only empirical evidence anyone has for these conjectures, since none is proven. - **Limit the blast radius.** If a single assumption secures authentication, storage and backup at once, its failure is total. Spreading load across structurally different assumptions turns one catastrophe into a degraded mode. ## What over-hedging costs | Hedge | What it buys | What it costs | |---|---|---| | Replaceable primitive behind one boundary | Migration measured in days | A small indirection, and the discipline to keep it | | Algorithm identifier and negotiation | Coexistence during migration | Downgrade attacks if the weakest option stays acceptable | | Two independent assumptions composed | Survives one falling | Double the implementation and review surface | | Information-theoretic security | Immune to any complexity result | Key material as large as the data, delivered in advance | Negotiation is the clearest example of a hedge that bites back: the mechanism that lets you introduce a successor is the same mechanism an attacker uses to force a retreat to the predecessor. If you build it, you must also build the policy that refuses the old option once the migration is done, and the monitoring that tells you who is still using it. ## How to decide, concretely Three questions settle most of these arguments: 1. **How long must this data stay secret?** Minutes for a session token, years for a health record. Long lifetimes justify conservative assumptions and migration planning; short ones do not. 2. **What is the blast radius if this one assumption falls?** One feature degraded, or every record ever stored readable? 3. **What would replacement actually cost today?** If the honest answer is 'we do not know where all the call sites are', the cheapest hedge available is not a new assumption, it is the boundary. ## Saying it well in a review The answer that lands is the one that refuses both extremes. Dismissing the question — 'that will never happen' — misses that single assumptions do fall, repeatedly. Accepting it whole — 'we should design for the collapse' — commits the organisation to a cost with no payoff, because the scenario that motivates it also removes the world in which the product has customers. The principal-level move is to convert an unmanageable tail risk into a managed engineering property: not 'is this secure forever', but 'how fast can we change it, how much is behind it, and how long does the data need it to hold'.

  • Why does the required secrecy lifetime of the data change the assumption you should pick?
    Because an attacker can store today's traffic and break it later. Security is needed for as long as the data must stay confidential, not merely at the moment of transmission. A session token needs minutes of hardness; a long-lived record needs the assumption to survive decades of algorithmic progress, which is a far stronger bet.
  • What is the downside of making the primitive negotiable on the wire?
    The same mechanism that lets a successor be introduced lets an attacker push both ends back to the weaker predecessor. Negotiation is worth having, but only alongside an explicit policy that stops accepting the retired option, and visibility into who still offers it.
  • Is composing two independent assumptions worth the cost?
    Sometimes. It is worth it where the data's secrecy lifetime is long, the two assumptions rest on structurally different problems, and the composition is done so that breaking either alone gains the attacker nothing. It is not worth it for short-lived data, and it roughly doubles the code and review surface either way.

A building is not designed to survive the continent sinking, but it is designed so a single failed column does not take the floor with it, and so a column can be replaced without demolishing the block.

saying these in an interview costs you the question

  • Proposes designing the product to survive a universal collapse
  • Dismisses the question because the open problem is unresolved
  • Adds algorithm negotiation without a policy for retiring the old option
  • Ignores how long the protected data must remain confidential
  • Treats an unproven conjecture as a guarantee because it is widely believed
  • Lets one assumption secure every subsystem at once