skip to content

Multi-Tenant IdP Patterns

One product, many customer identity providers: domain-based routing, per-tenant connection config and just-in-time account creation. Interviewers ask how a new tenant is onboarded without a deploy.

on this pageshow

questions

6

With only an email address typed at sign-in, how do you route the person to the right brand's identity provider?

level: middleimportance: must knowfreq 48%

answer

  1. two lookups, not one
  2. only the domain part routes
  3. domain to tenant, tenant to connection
  4. verified domains, unique per tenant
  5. a hint, not a credential

basics

~20 s

Take the domain from the typed address, look it up in a table of domains a tenant has verified, then follow that tenant to its identity-provider connection record and redirect. An unmatched domain falls back to local sign-in.

solid answer

~40 s

Home-realm discovery is two lookups chained, not one. Normalise the typed address, keep only the domain, resolve `domain -> tenant` against domains that tenant has verified, then `tenant -> connection` against the per-tenant identity-provider registry. Keeping the hops separate is what makes onboarding a brand a data change: a tenant can hold several domains and swap its connection without either table knowing about the other. Match both and you can redirect straight to that brand's provider; miss either and the sign-in continues on your own password form or asks for a tenant identifier explicitly. Whatever the chain returns, it has authenticated nobody - the typed domain is unverified input, and the brand's identity provider is still the party that says who arrived.

code

pseudocode · 14 lines
pseudocode
resolve_sign_in(typed_address):
    domain = punycode(lowercase(domain_part(typed_address)))
    if domain in shared_consumer_domains:
        return LOCAL_SIGN_IN

    tenant = verified_domains.find(domain)          # unique: one tenant at most
    if tenant is null:
        return LOCAL_SIGN_IN                        # or prompt for a tenant identifier

    connection = connections.current_for(tenant)    # may be null, draft, or disabled
    if connection is null or not connection.live:
        return LOCAL_SIGN_IN

    return REDIRECT(connection.sign_in_url, tenant, connection.id)

go deeper

for a junior

Remember that the sign-in form's address is used only to decide which identity provider is asked. It does not log anyone in, and the part before the @ plays no role in that decision.

for a middle

Explain the two hops and why they are separate: a domain points at a tenant, a tenant points at a connection, and a tenant may hold several domains or swap its connection without the other table changing.

for a senior

Show the operational side: cache invalidation on a connection edit, outcome counters per tenant, a rate limit on an unauthenticated endpoint, and the deliberate decision about whether a redirect may disclose that a domain is federated.

for a principal

Argue the entry-point mix as product policy. Silent routing is one field and one screen but discloses your customer list and strands anyone off-domain; the extra prompt buys privacy and a fallback. Say which cost you are choosing and for whom.

## The lookup chain **Home-realm discovery** answers one question before any authentication happens: *whose identity provider should see this person?* In one deployment of a housekeeping-task app serving every brand in a hotel group, that is three steps, and keeping them separate is what keeps a new brand out of the release process. 1. **Address to domain.** Normalise the typed address, discard the local part, keep the domain lower-cased and in its punycode form. Only the domain routes. The local part buys you nothing here and carrying it forward invites logging a full address on a path that has authenticated nobody. 2. **Domain to tenant.** Look the domain up in a table of domains a tenant has *verified*. A domain resolves to at most one tenant, and that is worth a unique constraint rather than an application-level check, because two brands inside one group will eventually both claim a shared domain and the database is the only place that settles it. 3. **Tenant to connection.** Load the tenant's identity-provider connection record - the issuer identifier, the sign-in URL you send the person to, and which connection is current when the tenant holds more than one. Three outcomes are ordinary rather than exceptional: no verified domain matches; the matched tenant has no live connection; the tenant has one but has not switched federation on yet. All three end the same way - the sign-in continues locally instead of failing. ## Silent redirect, interstitial page, or explicit identifier | Entry style | What the person does | What it costs you | |---|---|---| | Silent redirect on the typed address | Types an address and lands at the provider | Tells any stranger which domains are federated; nothing to offer someone whose address is not under a verified domain | | Interstitial "where do you work" page | Types the address, then confirms the brand | An extra screen on every sign-in, and the picker must not enumerate your customers | | Explicit tenant identifier in the link a brand publishes | Arrives already scoped to one tenant | Depends on the brand distributing that link; no help to someone who reached the product's front door | Most deployments end up running all three at once: the brand's own launch link carries the tenant, the typed address routes people who arrive at the front door, and a prompt is the fallback when the address resolves to nothing. ## The typed domain is a hint, not a credential At the moment of routing, the address is unverified form input. Routing on it grants nothing and proves nothing - it selects which party is asked. That has two consequences worth stating out loud in a design review: - The account you eventually create or find is keyed on what the provider asserts, not on what was typed. - A person who types a brand's domain and is redirected has learned only that the domain is federated. Decide deliberately whether you are willing to disclose that; a uniform "continue" that always advances one step hides it, at the cost of a second screen. ## The people this model does not fit - **Contractors and agency staff** whose addresses sit under their own employer's domain, not the brand's. - **Shared consumer mail domains**, which need a denylist so that no tenant can ever claim one and capture everyone on it. - **People who work for two brands** in the group and hold one address. - **Integration and machine accounts** that no identity provider vouches for. Each of these needs a route that does not depend on a verified domain: an invitation link that carries the tenant, or a local credential the tenant's admission policy explicitly allows. ## Operating it - **Cache the resolution, key it per domain, and invalidate it on any connection edit.** If a corrected sign-in URL needs a restart to take effect, onboarding is still a deploy in disguise. - **Log the outcome, not the input:** tenant, connection, outcome and a correlation id - not the full address. - **Count outcomes by reason.** A brand whose share of "no verified domain" jumps has almost always added a domain in its own directory without telling you. - **Rate-limit the endpoint.** It is unauthenticated and it answers questions about your customer list. - **Re-resolve on every sign-in** rather than storing a provider on the account row: people move brand, and a stored provider silently keeps sending them to the old one.

  • A contractor works across three brands in the group on one address under her own agency's domain. How does she sign in?
    Her domain resolves to no tenant, so domain routing cannot help her. She arrives through an invitation link that carries the tenant explicitly, or through a local credential that tenant's admission policy permits, and her account is bound to the tenant that invited her. If she needs all three brands, that is three bindings, each invited separately - not one address claiming three tenants.
  • Two brands in the group both claim the same shared domain. What happens?
    The unique constraint on the verified-domain table rejects the second claim, and the second brand's operator sees a conflict rather than silently overwriting the first. Resolve it as an onboarding exception: either the domain belongs to one tenant and the other brand's people arrive by invitation, or the domain is split by subdomain, each verified separately.
  • Why not store the resolved identity provider on the user's account row after the first sign-in?
    Because it freezes a routing decision that legitimately changes. People move between brands, a brand replaces its provider, and a domain is transferred. Re-resolving from the domain-and-connection tables on every sign-in means those changes take effect the next time someone signs in; a stored provider keeps sending them somewhere that may no longer be right.

A switchboard that routes a call by the number dialled: it decides which office picks up, and it has not the faintest idea who is calling. The person on the other end is the one who establishes that.

saying these in an interview costs you the question

  • Treats the typed domain as proof the person works for that brand
  • Shows a dropdown listing every customer brand on the sign-in page
  • Keeps the domain-to-connection mapping in application code
  • Assumes every user's address sits under a verified domain
  • Caches the resolution with no invalidation, so a config fix needs a restart
  • Stores the chosen provider on the account and never re-resolves it
open as a page

What must a brand's email domain prove before it may route sign-ins to that brand's identity provider?

level: seniorimportance: must knowfreq 42%

basics

~20 s

Control of the domain, proved out of band - typically a record the domain's own operator publishes in DNS, which you check and re-check. An unverified claim lets the claiming tenant's identity provider vouch for every address under that domain.

open as a page

On a first federated sign-in, what decides whether an account is created and which tenant it binds to?

level: middleimportance: should knowfreq 38%

basics

~20 s

The tenant's admission policy decides: create anyone the provider vouches for, create only pre-invited people, or create nobody. The tenant binding comes from the connection the sign-in was routed through, never from an attribute the provider sent.

open as a page

What does a per-tenant identity-provider connection record hold so onboarding a brand needs no code change?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Everything that differs per customer: the issuer identifier, the sign-in URL, the current signing certificates, the audience and endpoint you publish for them, their attribute names, their admission and entry-point settings, plus state and audit columns. The code reads it; it never hard-codes it.

open as a page

Every customer names its directory groups differently, so how do you map those groups to your roles without a deploy?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Hold the mapping as tenant-scoped rows - the group value the customer sends on the left, your role on the right - read on every sign-in and editable at runtime. Grant nothing for an unmapped value, record it, and alert when a rule stops matching.

open as a page

A brand asks you to switch its domain to federated-only sign-in, so how do you sequence that and who must keep a way in?

level: principalimportance: should knowfreq 30%

basics

~20 s

Run both paths in parallel until the connection has proved itself against real traffic, then flip a per-tenant flag that gates new sign-ins for that tenant's verified domains. Keep an explicit route for people no provider vouches for, including that tenant's own administrators.

open as a page