skip to content

A contractor firm's identity provider cannot meet one of your validation requirements — how far do you bend, and who decides?

level: principalimportance: should knowfreq 25%

answer

  1. the pressure is permanent, plan for it
  2. per connection, as data
  3. owner, reason, review date
  4. bounded and reversible versus removes a property
  5. every exception enlarges the test matrix

basics

~20 s

Rank the requests by blast radius: clock skew is bounded and reversible, endpoint-binding checks cost you once you have more than one endpoint, and audience and single use are not negotiable. Grant the rest as per-connection data with an owner, a reason and a review date — never as a global default or a code branch.

solid answer

~50 s

Every onboarding produces at least one of these, and the pressure is real: the firm will not change its identity provider, and the permit work has a start date. My rule is that a relaxation is **data on one connection**, with a named owner, the reason, and a review date — never a global default and never a branch in code, because both of those quietly apply to counterparties nobody agreed them for. Then I rank by what each costs. Widening clock skew is bounded and reversible, though it lengthens how long a message stays acceptable and therefore how long replay-cache entries have to live. Relaxing the endpoint-binding checks matters more as soon as I run more than one endpoint. Dropping the audience check or single-use enforcement I refuse outright, because those are what tie a message to my service and to one use. And I accept that every exception granted enlarges the matrix I have to test against forever.

code

json · 24 lines
json
{
  "connectionId": "firm-halbrook-civils",
  "relaxations": [
    {
      "check": "clockSkewAllowance",
      "value": "PT4M",
      "default": "PT2M",
      "reason": "Counterparty identity provider clock runs ~3m fast; their change window is Q1.",
      "approvedBy": "platform-security",
      "grantedOn": "2026-09-19",
      "reviewBy": "2026-12-19"
    },
    {
      "check": "destinationMustMatchReceivingUrl",
      "value": false,
      "default": true,
      "reason": "Provider emits a stale value; recipient check remains enforced.",
      "approvedBy": "platform-security",
      "grantedOn": "2026-09-19",
      "reviewBy": "2026-12-19"
    }
  ],
  "nonNegotiable": ["expectedAudience", "singleUseAssertionId"]
}

go deeper

for a junior

Recognise that counterparties differ and that the differences have to be recorded somewhere per counterparty, rather than by changing what the service does for everyone.

for a middle

Explain why an exception belongs in the connection record rather than in code or a global setting, and name at least one request you would refuse outright and why.

for a senior

Rank the common requests by blast radius, describe the ownership and expiry machinery, and account for the second-order cost that a per-connection exception makes the validation test matrix per-connection too.

for a principal

Set the estate-wide position: publish what you require and what you tolerate before a start date exists, keep the refusals short and constant, and make correct configuration cheaper to obtain than an exception.

## The pressure is structural, so build for it The water authority onboards contractor firms continuously, and each brings an identity provider it did not choose and cannot change quickly. Some of them will not set the destination value. One will have a clock four minutes out. One will sign the response but not the assertion inside it. One will publish metadata that expired last year. In every case the request arrives the same way: *this is how our provider works, and the crew starts on Monday.* A team without a position on this makes the decision at the ticket level, under deadline, in a configuration file. Six months later nobody can say which counterparties are running on which exceptions, or why. The principal-level answer is not a list of yes and no; it is a mechanism that makes each answer local, visible and expiring. ## The mechanism: relaxations are data, scoped to one connection Three properties, and all three matter: 1. **Per connection.** A global switch flipped for one firm silently applies to every other. This is how an estate ends up with checks disabled that no customer ever asked for. 2. **Data, not code.** A branch in code cannot be audited by listing rows, cannot carry an owner, and cannot expire. A row can. 3. **With an owner, a reason and a review date.** Exceptions rot. Without a date, the temporary accommodation for one firm's 2026 provider is still there in 2029, long after that provider was replaced. Add the report that nobody asks for until they need it: every connection running on a non-default, how long it has, and who owns it. An alarm when a relaxation passes its review date turns a permanent hole into a recurring conversation. ## Ranking the requests by what they actually cost | Request | What it costs | Position | |---|---|---| | Widen clock skew | lengthens the period a message stays acceptable, and therefore how long replay-cache entries must live | grantable, bounded, with a stated maximum | | Relax the endpoint-binding checks | a message minted for one of your endpoints is accepted at another; harmless with one endpoint, not with several | grantable while you have one endpoint; re-examined the day you add a second | | Accept a response whose inner assertion is not individually signed | forces you to consume strictly the element verification covered, and narrows your library choice to one that returns it | grantable only with that guarantee demonstrated in code | | Accept expired metadata | you are trusting material the publisher has declared stale | refuse; take a fresh copy instead, even manually | | Drop the audience check | nothing ties the message to your service | refuse | | Drop single-use enforcement | a captured message stays usable for its whole window | refuse | The two refusals at the bottom are the ones worth being able to defend in a room with a commercial deadline in it, and the defence is short: one of them is what makes the message about you, and the other is what makes it usable once. ## The costs nobody counts at the time - **Your test matrix becomes per-connection.** A change to the validation path now has to be exercised against every relaxation combination live in production, not against the default path. Keep the set of permitted relaxations small precisely so this stays finite. - **The exception outlives the reason.** The provider is replaced, the firm's team changes, and the row remains. Only a review date catches that. - **Support loses the ability to reason.** *It works for everyone else* stops being informative once every connection is subtly different, and diagnosis becomes per-tenant archaeology. - **Precedent.** The second firm asks for what the first was given, and the argument that it is unsafe is now harder to make. ## What a good outcome looks like Publish the requirements as part of onboarding — what your service provider requires, what it will tolerate, and what it will not — so the conversation happens before a start date exists. Offer the tolerated set explicitly, with limits attached to each. Keep the refusals short and constant. Then make the default path the easy one: if configuring a connection correctly takes ten minutes and an exception takes a signature from a named owner with a review date, most firms will configure it correctly. That asymmetry, rather than any individual decision, is what keeps an estate of federated connections honest over years. ## The judgment being tested An interviewer asking this is not looking for a policy recital. They want to hear that you know the pressure is permanent, that you have separated the reversible accommodations from the ones that remove a security property, that you can name which is which and why, and that you have somewhere for the answer to live other than a config file and somebody's memory.

  • Why is widening clock skew not free, even though it feels like the harmless request?
    It lengthens the period during which a message still passes your checks, which is the same period your replay-cache entries have to survive — so retention and capacity grow with it. It also widens the window in which a captured message is useful. Bounded and reversible, but priced, so state a maximum and hold it.
  • A firm asks you to skip the audience check because its provider cannot set the value. What is the answer?
    No. That check is what says the message was minted for your service, and nothing else in the set substitutes for it. Without it a message intended for another consumer verifies and is accepted here. Offer to help the firm configure the value; do not offer to stop looking at it.
  • How do you stop granted exceptions from accumulating across years of onboarding?
    Give every one an owner and a review date, report on connections running non-defaults, and alarm when a date passes. Keep the permitted set small so the combinations stay testable, and make the correct configuration cheaper to obtain than the exception — most counterparties take the cheaper path when it is genuinely cheaper.

saying these in an interview costs you the question

  • Turn the check off globally; it is only one firm that needs it.
  • Add a branch in the validation code for that counterparty.
  • An exception granted at onboarding does not need a review date.
  • Clock skew is free — widen it as far as the firm asks.
  • If the deadline is tight, drop the audience check and fix it later.
  • Accepting expired metadata is fine as long as the signature still verifies.