When an email carries both a formatted body and a plain-text alternative, what should a case compare between them?
answer
- The reader chooses which body to display
- Compare facts, never characters
- Values, link destinations, issued codes
- Both bodies get the residue guard
- Blank or raw alternative is the classic defect
basics
~20 sCompare extracted facts, not text: the same seeded values, the same destination behind every link, the same single-use code. Reading applications choose which body to display, so a defect in either one reaches a real person.
solid answer
~50 sStrip each body to plain content, pull out the values the case seeded, every link destination and any issued code, and assert the two extractions agree. Assert neither extraction is empty, and run the placeholder-residue guard against both bodies independently. Do **not** compare the two texts character for character. They legitimately differ: the plain-text body has no styling, wraps differently, and spells link destinations out instead of hiding them behind labels. A case demanding equality gets weakened until it asserts nothing. The defects this actually finds are consistent: a slot added to the formatted body and never to the plain-text one; two bodies whose links are built by different code paths, so one carries a stale destination; and a plain-text body that is blank or contains the other body's markup as literal characters. All three survive for months because nobody's own reading application shows the neglected body.
code
pseudocode · 15 linesfunction factsOf(body):
text = collapseWhitespace(stripMarkup(body))
return {
values: seededValuesFoundIn(text),
links: normalise(allLinkDestinations(body)),
codes: matchAll(text, ISSUED_CODE_SHAPE)
}
rich = factsOf(captured.richBody)
plain = factsOf(captured.plainBody)
assert plain.values is not empty // catches a blank alternative
assert rich.values == plain.values
assert rich.links == plain.links
assert rich.codes == plain.codesgo deeper
Know that one email can carry two alternative bodies for the same content, that the reading application picks which to show, and that a defect in the one you never look at still reaches somebody.
Explain what must agree — seeded values, link destinations, issued codes — and what may differ, such as styling, wrapping and how a link is presented. Say why a character-level comparison is the wrong instrument.
Show the extraction-and-compare shape, put it in the shared capture step, and name the recurring defects: the forgotten body, the stale link from a second code path, and the blank or raw alternative. Make the failure name which part broke.
Judge how far this is worth taking. Fact parity is cheap and permanent; a full structural comparison is a maintenance sink. Be ready to say where the content check stops and where review of rendered output takes over.
Many products send an email with two alternative bodies for the same content: a formatted one carrying styling, structure and images, and a plain-text one carrying the same information as unadorned text. The reading application on the other end chooses which to display, and the sender does not get a say. Some readers are configured to prefer plain text; some strip formatting for security; some display the plain-text body in a preview line before anyone opens anything. **A defect in either body reaches a real person**, which is why one is not a throwaway copy of the other. ## What "agreeing" means — and what it does not The two bodies are supposed to say the same thing, not to *be* the same text. A character-by-character comparison fails on every legitimate difference and teaches the team to ignore the case. | Property | Must agree | May legitimately differ | |---|---|---| | Dynamic values (names, references, amounts, dates) | yes — the same values, identically formatted | no | | Link destinations | yes — the same target, including any correlation value carried in it | how the link is presented: text-behind-a-label versus spelled out | | Single-use codes | yes — the same code, once | surrounding wording | | Structure | the same facts, in the same order | headings, tables and images versus lines and separators | | Line breaks and spacing | no requirement | freely | | Calls to action | the same destination and intent | button versus a written-out address | So the comparison is not textual. It is a comparison of **extracted facts**. ## How to compare them 1. Parse or strip each body down to plain content: remove markup, collapse runs of whitespace, and pull out the set of link destinations separately from the visible text. 2. Extract the things that carry meaning — the values the case seeded, every link destination, any code of the shape the product issues. 3. Assert set equality on those extractions, and assert each set is non-empty. 4. Run the placeholder-residue guard against **both** bodies independently, because a template that renders one body correctly can leave machinery in the other. Step three is where the real defects surface, and they are consistently the same three: - **The forgotten body.** A slot is added to the formatted body and never to the plain-text one, so the plain-text reader is missing the delivery window or the amount everyone else can see. - **The stale link.** The two bodies build their links through different code paths, one gets a correlation value or an updated destination and the other does not. This is the worst of the three because the message looks fine to whoever tested it. - **The empty or raw alternative.** The plain-text body is blank, or it contains the other body's markup dumped as characters. It survives for months because nobody's own reader shows it. ## What not to do Do not assert the bodies are identical after stripping markup. Wrapping, separators and the spelled-out link destinations mean the stripped texts differ for good reasons, and a case that demands equality gets weakened until it asserts nothing. Do not assert only on the formatted body because it is the one people design; the plain-text body is the one that appears in previews and in readers that refuse formatting. And do not build a second, parallel set of assertions per body: extract once, compare once, and the case stays short. ## How far to take it This check is worth writing once, in the shared step that handles a captured message, and then it costs nothing per case. It is not worth extending into a full structural comparison of the two bodies — that is a maintenance sink for a decreasing return. The rule of thumb is: **compare what a person would be harmed by getting wrong** — values, destinations, codes — and leave presentation to the review pass that looks at rendered content when a template changes. One boundary worth stating out loud in an interview: this is a check on content the product has already produced. If a body is missing entirely, that is a rendering or configuration defect in the sending path, and the useful failure message says *which* body was absent rather than reporting a missing value. Half of the value of this case is that it names the part that failed.
- Which of the two bodies would you assert on if you could only afford one?The plain-text one, slightly counter-intuitively. It is the body nobody designs and nobody looks at, so it decays silently, and it is what appears in preview lines and in readers that refuse formatting. The formatted body gets human eyes on it every time someone edits the template.
- The two bodies carry different link destinations. How do you make that failure diagnosable?Fail with both destinations printed side by side and the part each came from named, plus both bodies attached as artefacts. The usual cause is two code paths building the link separately, so naming which body carries the stale one points straight at the path that was missed.
Two translations of the same notice: you check that both state the same date and the same address, not that they use the same number of words.
saying these in an interview costs you the question
- Compares the two bodies character for character after stripping markup
- Asserts only on the formatted body because it is the designed one
- Treats an empty plain-text alternative as acceptable
- Runs the placeholder guard on one body only
- Assumes both bodies are built by the same code path