skip to content

A settlement signing key cannot leave its device — how do you plan recovery for the day that device fails?

level: seniorimportance: should knowfreq 40%

answer

  1. you cannot copy what cannot leave
  2. replicate, wrap, or re-key
  3. a wrapped copy is still a copy
  4. back up the capability instead
  5. compare recovery target to verifier lead time

basics

~20 s

You cannot back up material that cannot leave, so you choose between replicating the key into sibling devices at generation, exporting it wrapped under a key that itself never leaves a device, or holding a second accepted key. Each weakens or costs something specific.

solid answer

~50 s

Three shapes, and the choice is arithmetic, not taste. **Replicate at generation**: provision the key into a defined set of devices that share a wrapping relationship, so it exists in more than one device and never outside one. **Wrapped export**: take the material out encrypted under a backup key that itself only lives in devices, with its components split among custodians — the material is never in the clear outside a device, but a copy now exists and its protection is a key you manage. **No backup**: accept that device loss means generating a fresh key, which is fast, and getting the clearing partner to accept the new public half, which is not. The honest fourth answer is to stop backing up the key and back up the capability: a second key in a second device that the partner already accepts, so loss is a failover.

go deeper

for a junior

Recall that a key which cannot leave its device also cannot be copied into a normal backup, so recovery has to be designed rather than assumed.

for a middle

Explain the three shapes — replicate at generation, wrapped export under a backup key, or no backup with re-keying — and what each one does to the claim that the key exists in one place.

for a senior

Compare the recovery target against the lead time of whoever verifies your signatures, state the numbers, and show that a second accepted key turns recovery into a failover.

for a principal

Own the estate view: decide which keys get a second accepted signer as standard, who holds custodian components, and how the untested step in each plan gets rehearsed on a schedule.

## "Back up the key" is the wrong sentence The property you paid for is that the private half cannot leave the device. Every classical backup — copy it to storage, put it in an escrow file, snapshot the host — either cannot be done or undoes the property. So the question has to be restated: **what do you need to be able to do again after the device is gone, and how fast?** For a settlement signing key, that is "sign instructions the clearing partner will accept." Nothing in that sentence requires the same key. ## The three shapes, and what each costs | Shape | Where the material ends up | What it costs you | |---|---|---| | Replicate into sibling devices at generation | Inside two or more devices, never outside one | Custody is multiplied — more devices to control, and the set is usually fixed at generation | | Export wrapped under a backup key held only in devices | A ciphertext copy outside, plus the backup key inside devices | A copy now exists; its protection is a key and a custodian process you must run correctly forever | | No backup; re-key on loss | Only ever in the one device | Recovery time is the verifier's lead time, which you do not control | Two things to be precise about. A wrapped copy is **still a copy**: the material is never in the clear outside a device, but the claim "this key exists in exactly one place" is no longer true, and whoever can assemble the custodian components plus the ciphertext has the key. And replication is usually a decision you can only take **at generation**: a device that will not export the key will not export it to its sibling later either, so if you did not provision the replica on day one, that door is closed. ## Back up the capability, not the material The strongest answer is to make recovery a failover. Generate a second key in a second device, in a different failure domain, and have the clearing partner accept **both** public halves from day one as equally valid signers. Now losing a device costs you a routing change, not a recovery. The material was never duplicated, no wrapped copy exists, and you never have to ask an external party for anything in the middle of an incident. The cost is real but it is the kind you can pay in advance: a second device, a second provisioning, and a conversation with the partner while nothing is on fire. It also gives you something the other shapes do not — a way to retire one key without an outage, because the other is already accepted. ## Do the arithmetic before you choose State your numbers and compare two of them: 1. **Recovery target.** Suppose settlement signing may be unavailable for 30 minutes. 2. **Re-key lead time.** Suppose generating in a spare device takes minutes, but having the partner accept a new public half takes, say, five working days, because it goes through their change process. Five days against 30 minutes rules out the no-backup shape on its own for this key — not because it is insecure, but because it cannot meet the target. Flip the numbers (an internal verifier you control, re-keying in an hour) and the no-backup shape becomes the *right* answer, because it is the only one that keeps the key in exactly one device. This is why the choice is estate-specific and why a blanket rule is wrong in one direction or the other. ## What else has to survive with it - The **credential** the service uses to authenticate to the device, and the process that provisions it. - The device's **access rule** — which identities may invoke which key for which operation. Restoring signing with a rule you reconstructed from memory under pressure is how an over-broad rule gets written. - The **public half** as the partner holds it, and the agreed route for telling them it has changed. A recovery plan that restores the key and not these three restores the ability to sign and loses the ability to say who may sign. ## Rehearse it, or you have a document Every shape above has a step nobody has done: assembling custodian components, provisioning a replacement device, or asking a partner to accept a new public half. The value of the plan is entirely in whether that step has been performed once, on a schedule, by someone who still works there.

  • Why can you usually not add a replica device a year later?
    Because replication depends on the key being provisioned into the device set while it could still be placed there. A device that will not export the key refuses to export it to a sibling too, so the set is fixed at generation. Adding custody later normally means a new key and a new verifier conversation.
  • What does a wrapped backup actually change about your threat model?
    It converts "the key exists in one device" into "the key exists in one device plus a ciphertext whose protection is a backup key and a custodian process." The material is never in the clear outside a device, but a quorum of custodians plus that file reconstitutes it, so those people and that file are now in scope.
  • The partner accepts two of your public halves at once. Does that weaken anything?
    It widens the set of keys that can produce an accepted instruction from one to two, so both devices must be held to the same standard and both must be watched. In exchange, device loss is a failover and key retirement is routine. For most estates that trade is clearly worth it, but state it rather than hiding it.

saying these in an interview costs you the question

  • Snapshot the host's disk; that backs up the key
  • Export an encrypted copy to shared storage — it is encrypted, so nothing is lost
  • If the device dies we just generate a new key; nobody else is involved
  • Any device that will not export a key can still be cloned on demand
  • Recovery is solved once the key is restored; nothing else needs restoring