How does an append-only, tamper-evident store answer repudiation, and why is it not tamper-proof?
answer
- the past must not be editable in place
- keep versions, not just the current row
- each entry commits to the previous one
- detective only, unless witnessed outside
basics
~20 sAppend-only storage keeps every version instead of overwriting, so history cannot be quietly rewritten; chaining each entry's hash over the previous one makes any later alteration detectable. Together they preserve and prove history - but detection is not prevention.
solid answer
~50 sTake a clinical laboratory system where a released result can be amended in place and the store keeps only the current row: nothing can show what was reported at the time, so a technician can deny the original value and the store's own history is repudiable. The append-only answer models the record as a sequence of entries - the original result, then the amendment with its author, time and reason - and treats the current value as a projection over that sequence, never as the only copy. Tamper-evidence adds a hash chain: each entry commits to the hash of the previous one, so editing or removing an earlier entry breaks verification from that point on. That is a **detective** control, not a preventive one: whoever can rewrite entry two can recompute the rest of the chain, and can also simply stop writing. It becomes real evidence only when the chain head is witnessed outside their administrative control.
go deeper
Know the two ideas by name: append-only means new entries instead of overwriting a row, and tamper-evident means an alteration can be detected afterwards. Be able to say why an update-in-place record cannot settle an argument.
Explain the chaining mechanism concretely - each entry stores a hash over the previous entry - and what verification finds. Be precise that this detects alteration rather than preventing it.
Show you design against the privileged insider: witness the head outside the boundary, copy across a trust boundary, handle truncation and stopped writes, and separate who administers the evidence from who it holds accountable.
Own the cost and retention tradeoffs: years of immutable records collide with data minimisation and storage budgets. Decide which records deserve external witnessing at all, and who is accountable when verification fails.
## Why repudiation lands on data stores A store's job is remembering. If the only representation of a fact is a row that gets updated in place, then the past is whatever the row says now, and anyone with write access can make yesterday different. That is a repudiation threat against a store: not that data was stolen, but that **what the system said at the time is no longer knowable**, so any claim about it becomes one person's word against another's. In a laboratory system, a released result that can be amended in place with only the current value retained means nobody can show what the clinician actually received. The technician who changed it can deny the original value; equally, an honest technician cannot clear themselves. Both halves of the repudiation threat are live. ## Append-only: keep history, project the present The first control is structural. Model the record as an **immutable sequence of entries**: - the original release, with author, time and content; - the amendment, as a **new** entry carrying author, time, the new content and a reason; - deletion expressed as a tombstone entry, never as a removed row. The "current result" becomes a view computed over the sequence. This alone removes the silent overwrite, and it also gives you something the dispute actually needs: the *reason* an amendment was made, captured at the moment it was made rather than reconstructed afterwards. Note that append-only is a property you have to design for. A history table filled by whatever also has write access to the live row is only as append-only as the permissions around it. ## Tamper-evidence: chaining Append-only stops accidental and casual overwrites. It does not stop someone with write access from editing an old entry. Tamper-evidence addresses that by making each entry commit to its predecessor - the entry stores a hash computed over the previous entry's hash plus its own payload: ``` seq 1 payload A prev = 0000 hash = H1 seq 2 payload B prev = H1 hash = H2 seq 3 payload C prev = H2 hash = H3 <- head ``` Verification walks the chain and recomputes. Change payload B and H2 no longer matches what entry 3 stored as its `prev`, so the break points precisely at the edited entry. This is genuinely useful: it converts an invisible edit into a visible one and localises it. ## Why it is evident, not proof The honest limitation, and the part interviewers probe: whoever can rewrite entry 2 can normally also recompute H2, H3 and every later hash. A self-contained chain in a store one administrator fully controls proves very little against **that** administrator. Two further problems: - **Truncation.** Deleting the newest entries leaves a perfectly valid chain. Nothing internal reveals that entries 198-200 ever existed. - **Silence.** An attacker who simply stops the writes produces no broken hash at all; the absence of entries is the attack, and absence is invisible unless the design expects presence. ## What turns evidence into proof 1. **Witness the head outside the trust boundary.** Periodically copy or publish the current head hash and sequence number to a place the operators of the store cannot administer - a separate account or organisation, a counterparty, a signed periodic attestation. Now a rewritten chain contradicts an external record, and truncation is detectable because the external sequence number is higher. 2. **Ship entries continuously across a boundary.** A near-real-time copy in a store under different administrative control means an edit has to be made in two places by two different administrator populations. 3. **Sign with a key held elsewhere.** If entries are signed by a key the store's operators do not hold, they cannot forge replacements even if they can delete. 4. **Use write-once retention where it exists.** Retention-locked, write-once storage is a genuinely **preventive** control, whereas a hash chain is detective - the distinction is worth stating explicitly, because conflating the two is the classic error here. 5. **Make silence visible.** Monotonic sequence numbers and periodic heartbeat entries mean a gap or a stopped stream is itself an observable fact of the record rather than nothing at all. 6. **Separate duties.** The evidence store must not be administered by the population it holds accountable. That separation is what makes the record credible to a third party, and it is a modelling decision, not an implementation detail. ## Verify, and keep only what you need A chain nobody ever verifies is decoration. Verification should run on a schedule and on demand during a dispute, and someone must own what happens when it fails. And because these records live for years, apply data minimisation deliberately: keep the fields that settle disputes, and where the payload itself is sensitive, store a hash of it plus a pointer rather than the content, so the evidence store does not quietly become the most attractive copy of the data you own. ## What goes in the threat model For each store whose history could be disputed, record three things: the answering control (append-only, chained, witnessed, retained for N years), **who could defeat it**, and how a break would ever be noticed. The second and third are what distinguish a modelled control from a checkbox.
- Someone with full write access rewrites every entry and recomputes the whole chain. What defeats that?Only something outside their control. Witness the head hash and sequence number periodically to a separate account, organisation or counterparty; ship entries continuously into a store administered by different people; sign entries with a key they do not hold; or use write-once, retention-locked storage. A self-contained chain proves nothing against the person who administers it.
- How would the design make it obvious that an attacker simply stopped writing entries?By making presence expected. Monotonic sequence numbers turn a stop into a visible gap against an externally witnessed counter, and periodic heartbeat entries mean silence itself becomes an anomaly in the record. Without either, absence looks exactly like a quiet period, which is why "stop the writes" is the cheapest way to defeat an audit trail.
- Is append-only enough on its own to answer a repudiation threat against a store?It removes silent overwrites, which is the largest single win, but a privileged user can still delete entries or the whole store, and nothing reveals an edit. Append-only needs tamper-evidence to detect alteration, an externally witnessed head or an off-boundary copy to constrain the administrator, and retention that outlasts the dispute window.
A hash chain is a bound ledger with numbered pages: you can see a page was torn out, but the binding alone does not stop the person holding the book from rewriting the whole thing. Someone outside noting the page count is what makes it hold up.
saying these in an interview costs you the question
- Calls a hash-chained log tamper-proof rather than tamper-evident
- Assumes append-only prevents deletion of entries or the whole store
- Never verifies the chain, so a break would never be noticed
- Keeps the only copy under the administrators it holds accountable
- Overwrites the record and calls the resulting diff an audit trail
- Ignores truncation of the newest entries and stopped writes entirely