skip to content

What does a per-tenant identity-provider connection record hold so onboarding a brand needs no code change?

level: seniorimportance: should knowfreq 40%

answer

  1. if two customers can disagree, it is data
  2. certificates plural, not singular
  3. validate on save, not on first sign-in
  4. one broken connection, one tenant affected
  5. editing it is a privileged, audited action

basics

~20 s

Everything that differs per customer: the issuer identifier, the sign-in URL, the current signing certificates, the audience and endpoint you publish for them, their attribute names, their admission and entry-point settings, plus state and audit columns. The code reads it; it never hard-codes it.

solid answer

~50 s

A connection record is the answer to "what is different about this customer?", held as a row rather than a branch in code. It carries the tenant it belongs to, the protocol, the issuer identifier such as a SAML `entityID`, the sign-in URL to redirect to, the list of currently accepted signing certificates, the audience value and the `AssertionConsumerServiceURL` you publish for that tenant, the attribute names that customer sends, its admission policy, whether unsolicited responses from its own launch tile are accepted, and a lifecycle state - draft, testing, live, disabled. Validate it when it is saved rather than at the first sign-in, so a typo surfaces to the operator who made it. Everything about it should be tenant-scoped at runtime: one broken connection produces an error for one brand, with a correlation id, not a failure the whole estate feels.

code

json · 17 lines
json
{
  "connectionId": "conn-8842",
  "tenantId": "brand-coastal",
  "protocol": "saml2",
  "state": "live",
  "issuer": "https://idp.coastal-stays.example/entity",
  "signInUrl": "https://idp.coastal-stays.example/sso",
  "signingCertificates": ["MIIC...current", "MIIC...incoming"],
  "metadata": { "url": "https://idp.coastal-stays.example/metadata", "validUntil": "2026-12-01T00:00:00Z" },
  "audience": "https://housekeeping.example/sp/brand-coastal",
  "assertionConsumerServiceUrl": "https://housekeeping.example/sso/brand-coastal/acs",
  "attributeNames": { "subject": "NameID", "displayName": "cn", "groups": "memberOf" },
  "acceptUnsolicitedResponses": true,
  "admissionPolicy": "invite-only",
  "updatedBy": "admin:4412",
  "updatedAt": "2026-09-11T08:22:00Z"
}

go deeper

for a junior

Know what the record is for: the parts of a sign-in that differ per customer live in a row, so adding a customer is data entry. The code that reads it is the same for every brand.

for a middle

List the fields and say why each varies: issuer, sign-in URL, accepted certificates, the audience and endpoint you publish, attribute names, and the state that says whether it is live.

for a senior

Demonstrate the operational design: validation at save, a testing state before live, certificates as a list for overlap, per-tenant failure scoping and per-connection caches and timeouts, and last-known-good on a failed refresh.

for a principal

Decide how far the record may vary before it becomes a per-customer code path. Each new field is a permanent support surface; the judgment is which divergences you will absorb as configuration and which you will decline outright.

## What goes in the row The test for a connection field is simple: **if two customers can disagree about it, it is data**. Everything below differs between brands in a hotel group running one housekeeping app. - **Identity of the other side** - the tenant it belongs to, the protocol, the issuer identifier (a SAML `entityID` or an OpenID Connect issuer), and the sign-in URL you redirect the person to. - **Trust material** - the signing certificates you will accept, as a *list*, because a customer rotating one publishes both for a period and a single-valued column forces a flag day. Where the customer publishes an `EntityDescriptor` you can fetch, store its location and honour its `validUntil` and `cacheDuration` rather than copying values once and letting them rot. - **What you publish back** - the audience value the customer must address and the `AssertionConsumerServiceURL` for that tenant. Per-tenant endpoints cost nothing and make "which customer sent this?" answerable from the path alone. - **Per-customer shape** - the attribute names that carry the person's identifier, name and groups, because every directory names them differently. - **Policy** - the admission policy for first sign-ins, the group-to-role mapping this connection uses, and whether unsolicited responses from the brand's own launch tile are accepted. - **Lifecycle and audit** - state (draft, testing, live, disabled), who last changed it, when, and the previous value. ## Validate at save, not at first sign-in A connection saved but never exercised is a support ticket waiting to be filed by an end user at 06:00 on a shift change. On save: 1. Check structural facts you can check alone - the URLs parse and are absolute, the certificates parse and have not expired, the issuer identifier is non-empty and unique among live connections. 2. Fetch the customer's metadata document if there is one, and confirm it matches what was typed. 3. Keep the record in a **testing** state until at least one end-to-end sign-in has succeeded against it, so that going live is a positive event rather than an assumption. The operator who made the typo is on the screen at save time. Nobody is on the screen when the first housekeeper tries to sign in. ## Isolating one broken connection A multi-tenant registry fails as a shared-fate design unless you deliberately stop it: - **Scope every failure to its tenant.** A connection whose certificates no longer match the signature produces a tenant-scoped error with a correlation id, and the response tells the person to contact their own administrator. - **Do not share a mutable parse cache across tenants.** Cache parsed material per connection, keyed by connection id and version, so a poisoned or oversized document from one customer cannot evict or corrupt another's. - **Bound every outbound call.** Metadata refreshes run with timeouts and a per-connection failure counter; a customer whose endpoint hangs must not consume the request capacity other brands need. - **Keep the last-known-good.** A refresh that fails leaves the previous accepted certificates in place and raises an alert, instead of emptying the list and taking a working brand down. ## The per-tenant entry point One brand will insist its own launch tile works, which means responses arriving at your endpoint with no sign-in of yours preceding them. That is a per-tenant setting, and it is a real trade: for that tenant you lose the ability to correlate the response with a request you issued, so you accept a weaker correlation and compensate with tighter validity windows, single-use enforcement and a landing page that cannot be turned into a deep link by the sender. Make it opt-in per connection rather than a global default, and record who asked for it. ## Changing a connection is a privileged action The row decides who may sign in as whom. Edits belong behind the same scrutiny as a role grant: authenticated, authorised to that tenant, audited with before-and-after values, and ideally reversible. "Someone changed a connection" should be answerable from the audit trail, because when a brand reports that sign-in broke at 15:40, that is the first query you will run.

  • Why publish a separate assertion endpoint per tenant instead of one shared endpoint for all brands?
    Because the path then identifies the connection before anything is parsed. You load exactly one tenant's trust material, reject material that does not match it, and never fall into "try every customer's certificate until one verifies", which is both slow and a way to accept a message from the wrong brand. It also makes logs, rate limits and incident scoping naturally per-customer.
  • A customer is rotating its signing certificate next week. What does the record need?
    Room for both. Accept a list and put the incoming certificate alongside the current one before the switch, so either verifies during the overlap, then drop the old one after the customer confirms. A single-valued field forces a flag day coordinated with someone you cannot phone, and the failure mode is every housekeeper in that brand locked out at the same moment.
  • Should a brand's own administrators be able to edit their connection record themselves?
    Usually yes, scoped strictly to their tenant, and it is the difference between onboarding taking an hour and taking a sprint. The conditions are the same as for any privileged write: authenticated administrators of that tenant only, every change audited with before-and-after values, validation at save, and a live connection that cannot be pointed at a new issuer without the domain claim still holding.

saying these in an interview costs you the question

  • Branches on the customer's name in application code
  • Stores a single signing certificate, so rotation is a flag day
  • Shares one parse cache across every tenant's connection
  • Discovers a typo in a connection at the first user's sign-in
  • Lets one customer's slow metadata endpoint block unrelated sign-ins
  • Treats a connection edit as an ordinary settings change, unaudited