skip to content

You are designing the on-disk format for encrypting many records with a symmetric key. Walk through the decisions — nonce strategy, what belongs in the associated data, how many records one key may protect — and say what an authenticated mode still does not protect against.

level: principalimportance: nice to knowfreq 28%

answer

  1. layout: version || key_id || nonce || ct || tag
  2. bind record identity or the ciphertext is relocatable
  3. nonce strategy = writer topology; per-record key removes it
  4. per-key message budget drives rotation
  5. gaps: freshness, deletion, length, key commitment

basics

~20 s

Store a versioned header — format version, key id, nonce — and authenticate all of it plus the record's identity as associated data. Bound messages per key so nonces cannot collide. An authenticated mode still misses replay, rollback, record deletion, length leakage and key ambiguity.

solid answer

~60 s

Design the record as `version || key_id || nonce || ciphertext+tag`, with the header **and the record's identity** (row id, path, tenant) passed as associated data so a valid record cannot be relocated to another row or tenant, and a version field cannot be downgraded to select a weaker parser. Choose the nonce strategy against your writer topology: random nonces if writers are uncoordinated (respect the birthday budget), a partitioned counter if a single writer truly owns the state — or sidestep it by deriving a per-record key, which shrinks each key's message count to one and makes rotation a re-wrap rather than a re-encryption. What remains unprotected: **freshness** — an old valid record restored from backup verifies; **deletion or truncation** of whole records, since per-record tags say nothing about the set; **length**, which leaks through ciphertext size and through compressing secret and attacker-influenced data together; and **key commitment** — most common AEADs allow one ciphertext to decrypt validly under two different keys, which matters in multi-key or password-derived-key settings.

code

text · 12 lines
text
stored bytes:
  [ version | key_id | nonce | ciphertext | tag ]
     ^^^^^^^^^^^^^^^^^^^^^^ in the clear

encrypt(key, nonce,
        plaintext = value,
        aad = version || key_id || nonce ||
              tenant_id || table || row_id || column || row_version)

reader:
  expected_aad = rebuilt from WHERE THE RECORD WAS FOUND, not from the record
  decrypt(...) fails if the record was moved, downgraded, or is the wrong version

go deeper

for a junior

Focus on the layout: an authenticated mode, a fresh nonce per record stored alongside, and a key identifier so keys can be rotated.

for a middle

Add associated data — the header plus the record's identity — and explain why a valid ciphertext is otherwise movable between rows or tenants.

for a senior

Drive the nonce strategy from writer topology, connect per-key message budgets to rotation, and name the residual gaps (replay, deletion, length) with concrete mitigations.

for a principal

Treat the format as a protocol with a version field, an explicit downgrade policy, set-level integrity where it is required, a key hierarchy that makes rotation cheap, and a written statement of which properties are deliberately not provided.

## Frame the problem "Encrypt the records" is not a single decision. A record format is a small protocol, and the interesting questions are all about *binding* — what each ciphertext is tied to, how many of them one key may protect, and which properties simply are not on offer no matter what mode you pick. ## The header, and why the associated data is the real design work A reasonable layout is: `version (1 byte) || key_id || nonce || ciphertext || tag` Everything before the ciphertext travels in the clear, which is fine — none of it is secret. What is not fine is leaving it unauthenticated. - **version** tells the reader how to parse. If it is not covered by the tag, an attacker can flip it and steer your reader into an older code path — a downgrade. Authenticate it, and reject unknown versions rather than falling back. - **key_id** lets you rotate without re-encrypting everything at once and lets you find the right key. Unauthenticated, it invites confusion between keys. - **nonce** must be covered because in some constructions a modified nonce silently changes the keystream. Then the decision most designs miss: **bind the record's identity**. An authenticated ciphertext is valid *anywhere* unless something ties it to where it belongs. If a row's ciphertext carries no notion of which row it is, an attacker with write access to the store — or a bug in your own code — can copy tenant A's encrypted balance into tenant B's row, or move an encrypted file to a different path, and every integrity check passes. Passing the row id, the tenant id, the column name, the file path as associated data makes relocation fail the tag check. This is *ciphertext relocation* or confusion, and it is the practical reason associated data exists. ## Nonce strategy follows from writer topology The contract is that a (key, nonce) pair is used once, and how you achieve that depends on who writes. - **Many uncoordinated writers** (horizontally scaled services, clients, serverless): random nonces. Simple, stateless, and bounded by the birthday problem — with a 96-bit nonce, keep the number of messages per key far below the point where collisions become likely, which in practice means a per-key message budget in the low billions at most, and preferably far less. - **A single owner of durable state**: a counter, ideally with a per-writer prefix so that adding a replica or restoring a snapshot cannot collide. The failure modes here are operational rather than mathematical — clones, snapshots, restarts, scale-out. - **Best of both**: derive a per-record data key from a long-lived key and the record identity, then encrypt that record with a fixed or trivially chosen nonce. Each derived key protects one message, so nonce collision stops being a live concern. This is the structural-separation answer: you removed the possibility rather than managing it. The per-key message budget also drives **rotation**. Rotation is not only a hygiene ritual; it is what keeps you inside the nonce space and limits the blast radius of a single key. A wrapped-data-key design (each record's key encrypted under a key-encryption key) makes rotating the outer key a re-wrap of small blobs instead of a re-encryption of the corpus. ## What an authenticated mode does not give you This is usually where the interview is really going. **Freshness / rollback.** Every record is individually valid forever. Restore last month's file, replay an old row from a backup, revert a value to a previous authenticated version — all verify perfectly. If ordering or recency matters, you must supply it: a monotonic version number bound in the associated data and *checked against expected*, or a signed/authenticated summary over the whole set (a root hash, a sequence number, a signed manifest). **Deletion and truncation of whole records.** Per-record tags say nothing about the collection. An attacker who deletes rows, drops log lines, or truncates a file at a record boundary produces a store where every remaining record is authentic. Set-level integrity needs a structure over the set — a chained MAC, a Merkle root, a counted manifest. **Length and compression.** Ciphertext length tracks plaintext length. If the length is sensitive, pad to buckets. And never compress attacker-influenced data together with secret data before encrypting: compression makes the ciphertext length a function of how well the attacker's guess matched the secret, which is a plaintext-recovery oracle. **Key commitment.** Most widely used AEADs are not *committing*: it is possible to construct a ciphertext that decrypts successfully — tag and all — under two different keys, to two different plaintexts. Where every record is decrypted with one server-held key this is a curiosity. Where the key is chosen by an untrusted party, derived from a user-supplied password, or selected from a set the attacker can influence, it becomes a real attack (a partitioning oracle that accelerates password guessing, or a message that means different things to different recipients). If your design has multiple candidate keys, add commitment — for example by binding a key-derived commitment value into the record and checking it. **Everything outside the cryptography.** A tag proves nothing about who was authorised to write the record; access control is a separate mechanism. It does nothing once the process is compromised and holds the key in memory (a separate concern of its own). And it does not protect the key management chain that decides who can call decrypt at all. ## How to present this A strong answer states the layout, then spends its time on the three binding decisions (identity in the associated data, version authentication and no silent fallback, nonce ownership), then names the residual gaps explicitly — freshness, set integrity, length, key commitment — and says which of them this product actually cares about. The weak answer is a list of primitives with no discussion of what each ciphertext is bound to.

  • Why must the reader rebuild the associated data from context rather than reading it out of the stored record?
    Because associated data only binds if the receiver supplies the value it *expects*. If you parse the tenant id out of the record and then pass that same value as associated data, the tag will verify for whatever the record claims, and relocation succeeds. The row id, tenant and path must come from where the record was actually found — the query, the file path, the request context — so that a mismatch fails the tag check.
  • How would you add rollback protection to per-record encryption?
    Bind a monotonically increasing version number into the associated data and store the expected current version somewhere the attacker cannot roll back independently — a separate authenticated metadata store, a counter in a trusted component, or an authenticated manifest over the whole set. Verification then means both "the tag is valid" and "the version is the one I expect," which is what turns integrity into freshness.
  • When does the lack of key commitment actually matter?
    When more than one key is a plausible candidate for a given ciphertext and an untrusted party influences which. Password-derived keys are the sharp case: an attacker can craft a ciphertext that decrypts validly under many candidate passwords, or use the ability to find a key that decrypts successfully as an accelerated guessing oracle. Multi-recipient or attacker-supplied-key designs have the same shape. A single server-held key with no attacker influence is not exposed to it.

An authenticated record is a sealed, signed envelope. The signature proves the contents were not altered — it does not prove the envelope is in the right pigeonhole, that it is this month's, or that three envelopes were not quietly removed from the stack.

saying these in an interview costs you the question

  • Believing an authentication tag prevents an old record from being restored from backup.
  • Leaving the version or key identifier outside the authenticated data, then supporting a silent fallback path for old versions.
  • Passing the record's own claimed identity as associated data instead of the identity of the location where it was found.
  • Assuming per-record tags detect deletion of whole records.
  • Treating key commitment as an academic footnote in a design where the key is derived from a user password.

context