skip to content

When a product strips address tags before keying an account, how does that break per-case recipient addresses?

level: middleimportance: nice to knowfreq 21%

answer

  1. The product does not store what you typed
  2. Two variants, one account
  3. Anti-abuse plumbing meets test identity
  4. Move the varying part into the domain

basics

~20 s

Stripping the tag collapses every decorated variant into one canonical address, so identities the suite believes are separate become one account. Registration fails as a duplicate, or worse succeeds and two cases silently share state.

solid answer

~50 s

A scheme that mints identities by decorating one local part depends on the product preserving the decoration. Many products canonicalise first -- lowercase, trim, drop the separator and everything after it -- because it is anti-abuse plumbing, and then every variant is one account. The visible symptoms all lie: a duplicate-registration error in setup looks like dirty data, two cases operating one account look like an ordering bug, and suppression or repeat-send guards keyed on the canonical address look like a slow delivery path. The fix is to move the varying part out of the local part, since canonicalisation almost never touches the domain: give each case a wholly distinct local part on a catch-all domain, or give each run its own subdomain. Verify it with a precondition probe rather than assuming a policy you do not own.

code

pseudocode · 13 lines
pseudocode
# precondition probe: does this environment keep decorated variants apart?
a = "probe-" + runId + "+case-a@" + SUITE_MAIL_DOMAIN
b = "probe-" + runId + "+case-b@" + SUITE_MAIL_DOMAIN

register(a)
outcome = register(b)

if outcome == ALREADY_REGISTERED:
    failPrecondition(
      "account keying normalises the local part; " +
      "switch to distinct local parts or a per-run subdomain")

assert accountsMatching(runId).count == 2

go deeper

for a junior

Recall that a product may not store an address exactly as typed: it can lowercase it and drop a decoration before deciding whether the account already exists.

for a middle

Explain what canonicalisation does to a decorated local part, name the symptoms it produces -- duplicate registration, cross-talk, silent suppression, dropped repeat sends -- and describe the schemes that survive it.

for a senior

Show how you would find out rather than assume: a precondition probe, run when the deployed target is prepared, that registers two variants and fails loudly with the cause so nobody debugs it inside a functional case.

for a principal

Own the assumption itself. Decide which identity key a suite may depend on across environments, and treat any scheme that relies on another team's normalisation policy as a dependency to verify rather than a fact.

## What canonicalisation is, and why products do it Many products do not store an address exactly as it was typed. Before storing it, or before checking whether an account already exists, they reduce it to a **canonical form**: lowercase the whole thing, trim surrounding whitespace, drop a separator and everything after it in the local part, and sometimes remove other punctuation from the local part as well. The motive is reasonable -- it is anti-abuse plumbing meant to stop one person opening unlimited accounts on trivial variants of one address, and it also spares support from *the same customer twice with different capitalisation*. That policy collides head-on with a suite that mints identities by decorating one local part. If the product reduces `base+case-a@d` and `base+case-b@d` to the same canonical string, the two addresses your suite believes are two identities are one account. ## The four symptoms, and how each one lies to you | Symptom | What the run reports | What is actually happening | |---|---|---| | Registration rejected | *address already registered*, in setup | Both variants canonicalise to one existing account | | Cross-talk between cases | Assertions pass against data another case created | Two cases are operating one account | | Sudden silence | Later cases time out waiting for email | Suppression or unsubscribe state is keyed on the canonical address | | Only the first send arrives | A timeout that looks like a slow delivery path | The delivery path treats the repeat send to one recipient as a duplicate | Every one of these presents as an infrastructure problem. The first looks like dirty data, the second like an ordering bug, the third and fourth like a slow or broken delivery path. None of them looks like *your identity scheme has no identities in it*, which is why this costs teams a week rather than an hour. ## Three places the reduction can happen - **The product's account layer**, when it normalises before its uniqueness check. This is the one that produces duplicate-registration errors and cross-talk. - **The delivery path**, when suppression lists, bounce records or repeat-send guards are keyed on a normalised recipient. The product happily keeps two accounts; the email for one of them is silently dropped. - **The receiving mail system**, when it routes several written forms into one store. Delivery still works, but a case that expects *only my message here* sees a neighbour's. Because the three are independent, the answer is not one policy but a property of the specific environment -- and it can change when any of the three is reconfigured. ## Detect it, do not assume it Make it a **precondition probe** rather than a discovery inside a functional case. Register two identities that differ only in the decorated part, then assert that the product reports two distinct accounts. Run the probe when the environment is prepared, so a violation fails loudly, once, with a message that names the cause -- instead of surfacing as an unexplained duplicate-registration error in whichever case happens to run second. ## Schemes ranked by how much they depend on the separator surviving | Scheme | Shape | Survives a stripped tag | What it needs | Triage readability | |---|---|---|---|---| | Decorated local part | one base name, a separator, a per-case tag | No | Receiving system must honour the separator | High | | Distinct local part on a catch-all domain | a wholly different local part per case | Yes | Catch-all routing on one domain | High | | Per-run subdomain | any local part, a distinct subdomain per run | Yes | Wildcard routing plus catch-all | Highest: the run is in the domain | The move that solves it is to **push the varying part out of the local part**. Canonicalisation operates on the local part almost exclusively; the domain part is the address's routing identity and products do not rewrite it. A distinct local part on a catch-all domain already survives tag stripping, and a per-run subdomain survives it while making the run visible in the address itself. Both cost a little routing setup and no per-case provisioning. ## The residual limit worth naming Distinct addresses buy distinct accounts **only when the address is the account key**. A product that de-duplicates on a phone number, a device identifier, a payment instrument or a hashed identity will merge your identities no matter how the addresses are spelled, and no addressing scheme repairs that. At that point the identity has to come from a key the product itself considers distinct, and the recipient address goes back to being a channel rather than an identity -- which is a different design with its own tradeoffs. The general lesson is narrower than *tags are bad*: never treat address distinctness as a guarantee the product has agreed to. It is an assumption about someone else's normalisation policy, it is cheap to verify, and it is expensive to discover late.

  • Why does moving the varying part into the domain survive canonicalisation?
    Because normalisation is aimed at the local part -- that is where the abuse it was written to stop lives. The domain is the address's routing identity, and rewriting it would send mail somewhere else, so products leave it alone. A distinct local part on a catch-all domain, or a per-run subdomain, therefore stays distinct through the same reduction that erases a tag.
  • The product keeps the variants apart but only the first email ever arrives. Where would you look?
    At the delivery side rather than the account side. Suppression lists, bounce records and repeat-send guards are often keyed on a normalised recipient, so a second send to what the delivery path considers the same person is dropped. The account layer and the delivery layer normalise independently, and either one alone is enough to break the scheme.
  • When does no addressing scheme rescue you?
    When the address is not the account key. If the product de-duplicates on a phone number, a device identifier or a hashed identity, distinct addresses merge into one identity regardless of how they are spelled. At that point the distinctness has to come from whatever key the product does treat as unique, and the address goes back to being a channel.

saying these in an interview costs you the question

  • Assumes every product preserves whatever follows the separator
  • Treats a duplicate-registration error in setup as deployment flakiness
  • Believes distinct addresses always mean distinct accounts
  • Ignores that suppression state is keyed on the canonical address
  • Discovers the policy from a failing case instead of a precondition probe
  • Thinks lengthening the tag makes the variants harder to collapse