Which properties of a rendered notification message can a case assert, and which need human review?
answer
- Determinism draws the line, not difficulty
- Could two people disagree given the inputs?
- Facts to the machine, taste to a person
- Normalise the per-run values, then compare
- Review triggered by change, not by release
basics
~20 sMachines assert facts: every slot filled, seeded values present, link destinations right, both bodies agreeing, formats matching the recipient's profile. People judge appearance and wording — and review a captured sample only when the rendered content changes, not every release.
solid answer
~50 sDraw the line at determinism, not difficulty: if two competent people could not disagree about the right answer given the case's inputs, a machine decides it. That covers every slot rendering a value, the seeded values appearing, link destinations, recipient-appropriate dates and amounts, the two bodies of an email carrying the same facts, and the correlation value in the subject line. What is left is perceptual — does the layout hold together in an unfamiliar reading application, does the wording read as intended, does a long translated string truncate awkwardly. Those need a person, so make the person's pass rare rather than routine: the run captures the rendered content, a normalisation step replaces the per-run values, and the result is compared against an approved reference. Unchanged content costs nothing; changed content fails the case and buys one focused reading of one difference.
code
pseudocode · 14 linescaptured = captureMessageFor(recipient)
// deterministic half - always asserted
assertNoTemplateResidue(captured)
assertSeededValuesPresent(captured, seeded)
assertLinkDestinations(captured, expectedTargets)
assertFormatsMatchProfile(captured, recipient.profile)
assertBodiesCarrySameFacts(captured)
// perceptual half - reviewed only when content moves
normalised = replaceAll(captured, seeded + generatedRefs + timestamps, MARKER)
if normalised != approvedReference(messageType):
fail("rendered content changed - review required",
attach = [normalised, approvedReference(messageType), captured])go deeper
Be ready to sort properties of a message into ones with a single right answer — a slot filled, a value present, a link destination — and ones that need a person's eye, such as whether the wording reads well.
Explain the determinism test and give examples from both halves. Say why the deterministic checks belong in the shared capture step rather than being repeated per case.
Show how to make the human pass rare: capture rendered content, normalise the per-run values, compare against an approved reference, and review only a difference. Name both failure modes — comparing images every run, and eyeballing with no trigger.
Own the policy: what the normalisation list contains, who approves a new reference, and how the manual half is written down so a lapse is visible. Be able to argue why continuous fact checks plus change-triggered review beats either extreme.
Every rendered notification splits into two kinds of property: things with a single right answer computable from the case's own inputs, and things that require taste or an unpredictable rendering surface. The first kind belongs in an assertion. The second kind belongs to a person — and the engineering work is not deciding *that* a person is needed, it is making the person's pass rare, cheap and triggered by change rather than by the calendar. ## The line is determinism, not difficulty A useful test to apply property by property: **could two competent people disagree about the right answer, given only the case's inputs?** If not, a machine decides. If so, a person does. | Property of a rendered message | Decided by | Why | |---|---|---| | Every slot rendered a value | the case | one right answer, no judgement involved | | The values the case seeded appear | the case | the case created them, so it knows them | | Link destinations point where the flow intends | the case | comparable against a stated expectation | | Dates and amounts match the recipient's declared preference | the case | derivable from seeded profile data | | The formatted and plain-text bodies carry the same facts | the case | set comparison over extracted values | | Subject line carries the correlation value the run injected | the case | exact match | | Layout holds together in an unfamiliar reading application | a person | the surface is not controlled by the sender | | Wording reads as intended for a first message | a person | taste, and no stable expected value | | A long translated string overflows or truncates awkwardly | a person | perceptual, and varies per rendering surface | | Emphasis and hierarchy land on the important line | a person | design judgement | Everything in the top half is written once into the shared step that handles a captured message. It costs nothing per case afterwards, and it never gets skipped under release pressure — which is the real argument for automating it, more than the effort saved. ## Make the human pass rare: review on change "Somebody eyeballs the welcome email each release" is a commitment that quietly stops happening around the third release. Replace it with a **review triggered by change**: 1. The run captures the rendered bodies and subject line as artefacts. 2. A normalisation step removes everything that legitimately varies per run — the seeded values, generated references, single-use codes, timestamps — replacing each with a stable marker. 3. The normalised content is compared against an approved reference stored with the code. 4. Identical? No review, no failure, nobody's time spent. 5. Different? The case fails with both versions attached, a person reads the difference once, and either fixes a defect or approves the new content as the reference. That inverts the economics. Unchanged templates cost nothing. A changed template costs one focused reading of one diff, by someone who can see exactly what moved — which is a far better review than a tired scan of a message that looks the same as last time. The normalisation step is the part that decides whether this survives. Under-normalise and every run differs, the case is noisy and gets disabled within a month. Over-normalise and you erase the content you were reviewing. Normalise **exactly** the per-run values and nothing else, and keep the list of them in one visible place. ## Both failure modes are real - **Over-automating appearance.** Comparing images of a rendered message across several reading applications every run produces differences that harm nobody, on a surface the sender does not control. The failures are ignored, then the case is disabled, and the genuinely deterministic checks that were bundled with it go quiet at the same time. - **Under-automating facts.** "We look at the emails" covers a raw placeholder, a stale link and a wrongly formatted amount only if somebody happens to notice. All three are exactly computable, and the machine notices every time. ## Write the human half down so it can fail The part that needs a person still needs to be an obligation rather than a hope. Name who reviews, what triggers it — a change in normalised rendered content — and what a review can conclude: approve, or raise a defect. A review that has no way to fail is decoration. In an interview, saying that out loud is what separates "we automate what we can" from someone who has actually kept the manual half alive past release three.
- How do you keep the change-triggered review from firing on every run?Normalise exactly the values that legitimately vary per run — seeded data, generated references, single-use codes, timestamps — and replace each with a stable marker before comparing. Under-normalising makes the case noisy and it gets disabled; over-normalising erases the content being reviewed. Keep the list of normalised values in one visible place.
- What is wrong with comparing images of a rendered message across several reading applications each run?It fails on rendering differences that harm nobody, on a surface the sender does not control. The team learns to ignore the failures, then disables the case — and any deterministic checks bundled with it go quiet at the same time. Compare facts continuously; look at appearance when content changes.
- Your team says it eyeballs the welcome email every release. Why is that not a control?Because it has no trigger, no owner and no way to fail. It stops happening quietly around the third release and nobody notices. Naming who reviews, what triggers a review — a change in normalised rendered content — and what a review may conclude turns a habit into something that can be seen to have lapsed.
saying these in an interview costs you the question
- Calls appearance untestable and asserts nothing at all
- Compares images of rendered messages on every run
- Relies on someone eyeballing each release with no trigger
- Normalises so much that the reviewed content is erased
- Treats human review as impossible to make repeatable