How do you set retention and access for test-run evidence that holds personal data yet is what makes a failure investigable?
answer
- most of it is never read again
- one short default, few named exceptions
- hold longer only against an open investigation
- redacted evidence stops being the same problem
basics
~20 sTier it. Keep ordinary run evidence for a short default window, hold longer only what is attached to an open investigation, limit reads to the people investigating and record those reads, then shrink the whole problem by redacting as the evidence is written.
solid answer
~50 sStart from how the evidence is actually used: nearly all of it is read within days of the run, or never. So set a short default window - suppose the team agrees on fourteen days - and make the exceptions explicit: evidence linked to an open defect is held until it closes, and anything held for another reason needs a named owner and a review date. Access follows the same shape - readable by the people who investigate rather than by everyone with a login, with reads recorded so that *who has seen this* has an answer. Then make the trade cheaper instead of arguing it: evidence with the personal values removed as it was written keeps its investigative value while dropping out of the personal-data conversation entirely, which is what lets you keep it for as long as debugging genuinely benefits.
code
pseudocode · 14 lines# how long one piece of run evidence is kept
function retentionFor(evidence):
if evidence.linkedDefect exists and evidence.linkedDefect.isOpen:
return UNTIL_DEFECT_CLOSED plus 30 days
if evidence.holdsUnredactedPersonalValues:
return 7 days # shortest tier: the exposure window
return 14 days # team default
# who may read it, and the read is recorded either way
function mayRead(person, evidence):
allowed = person in evidence.run.owningTeam
or person in evidence.linkedDefect.assignees
evidence.accessLog.append(person, now(), allowed)
return allowedgo deeper
Be ready to say that the files a test run produces are not kept forever and are not open to everyone, and to check before sharing one outside the team.
Explain the tiering: a short default window, a longer hold only for evidence tied to an open investigation, and reads limited to the people who actually investigate.
An interviewer expects the tension handled honestly - name the window you would pick and why, say how you verify that deletion really happened across mirrors and backups, and show that redacting as evidence is written is what buys a longer hold.
Own the policy and its accountability: who may grant an exception, how holds expire by default, and how you would measure whether the evidence being kept is read often enough to justify what holding it costs.
## The tension, stated plainly Evidence a test run captures is at once the thing that makes a failure investigable and, when it holds real personal data, a standing exposure that grows with every day it is kept and every person who can open it. Both sides are real. Delete quickly and last Tuesday's intermittent failure becomes unexplainable; keep everything for a year with open access and the team accumulates a large, unreviewed pile of personal data nobody remembers creating. The way out is not to pick a side. It is to notice that the two curves have different shapes. ## The decay curve nobody measures Investigative value falls off sharply: almost every piece of run evidence is read within days of the run that produced it, or is never read at all. Exposure does the opposite - it is flat for as long as the file exists, and it steps up each time the file is copied. That asymmetry is the whole argument, and it is measurable. Count reads of evidence by age; past a week the number is usually indistinguishable from zero, and that measurement settles arguments no amount of policy can. ## A tiering rule Suppose a team settles on the following. The numbers are that team's decision, not a fact about the world: | Tier | What it holds | Window | Who may read | | --- | --- | --- | --- | | Default | Evidence from any run, redacted as it was written | 14 days | The owning team | | Unredacted | Evidence known to carry real personal values | 7 days | The owning team, reads recorded | | Linked | Evidence attached to an open investigation | Until it closes, plus 30 days | The team plus the assignees | | Summary | Durations, pass and fail counts, error kinds, timings | Indefinite | Anyone | Four properties make a rule like this survive contact with a real team: - **A short default and few exceptions.** One window plus a small set of named exceptions is enforceable. Per-project negotiation is not. - **Every exception carries an owner and a review date.** A hold that does not expire by default is permanent retention with extra steps. - **The heavy channels expire fastest.** Images, recordings and full captured traffic carry the most personal data and are read the least. - **The trend data is separated out.** Teams ask for long retention because they want trends. Durations, counts and error kinds give them that and contain no personal data at all - nobody has ever built a trend out of a screen image. ## Access, and being able to answer *who has seen this* Retention limits how long; access limits how many. Read access for everyone with a login is the common failure, and it is usually accidental: the store was set up for convenience before anyone noticed what lands in it. A workable rule is that the people who investigate can read, other requests are granted individually, and every read is recorded. The recording matters less as a deterrent than as an answer - when a copy turns up somewhere it should not be, the access record is the only way to reconstruct how it got there. ## Verifying that deletion happened A retention rule that is not verified is a document. Two checks are worth running: 1. Take runs older than the window and try to open their evidence. If it opens, the deletion path is broken or was never scheduled at all. 2. Do the same for mirrors, replicas and backups. Deleting from the primary store while a nightly snapshot keeps a copy for months is the standard way a rule looks satisfied and is not. ## The move that dissolves the argument Retention is contested only because the evidence carries personal data. Take the personal values out at the moment of capture and most of the tension goes with them: evidence with no real values in it can be kept as long as debugging genuinely benefits, shared with whoever needs it, and argued about on storage cost alone. That is why the retention conversation and the redaction conversation are the same conversation, and why the cheapest way to win the first is to have already won the second. What an interviewer is listening for here is not a number. It is whether you can name a default, justify it from how the evidence is actually used, say what makes an exception legitimate, and admit that the rule is worth nothing until somebody has checked that the files are really gone.
- The team wants a year of run evidence for trend analysis. How do you answer?Separate the trend from the evidence. Durations, pass and fail counts, error kinds and timings are what a trend needs, and none of them is personal data, so keep those indefinitely as small structured records. The heavy captures that carry personal data - images, recordings, full recorded traffic - expire on the short window. Nobody has ever built a trend out of a screen image.
- How do you know the retention rule is actually being applied?Sample it. Take runs older than the window and try to open their evidence; if it opens, the deletion path is broken or was never scheduled. Then check the mirrors, replicas and backups, because deleting from the primary store while a nightly snapshot keeps a copy for months is the standard way a rule looks satisfied and is not.
- What accountability belongs with a longer hold?A name and a date. Every exception should record who asked for it, what it is for and when it is reviewed, so a hold ends by default rather than because somebody remembered. An exception with no owner is permanent retention with extra steps, and those are the holds that are still there when nobody can explain why.
Run evidence ages like a photograph of a whiteboard: genuinely useful for a few days, and awkward to still be holding a year later.
saying these in an interview costs you the question
- Keeps every run's evidence forever in case it helps
- Deletes everything nightly and can investigate nothing
- Gives the whole organisation read access to run evidence
- Quotes a legal retention period as a universal number
- Assumes deletion from the main store clears the mirrors
- Never checks whether old evidence is read at all