In OpenID Connect, why must a provider settle its public-or-pairwise subject_type policy before the first relying party integrates?
answer
- a name, not a fact about the person
- changing it strands what is stored
- no mapping the protocol will give you
- decided once, before anyone stores it
- public, pairwise, or one sector
basics
~20 sEvery sub a relying party has stored is a product of that policy. Change subject_type, or move a client between sectors, and the provider starts issuing unrelated values, orphaning every account record keyed on the old ones.
solid answer
~50 s`sub` is not a fact about the user; it is a name the provider chose, and `subject_type` and the client's sector are inputs to choosing it. Change either and the provider begins issuing values that match nothing a relying party has stored, so returning users arrive as strangers and every account keyed on `(iss, sub)` has to be re-identified by something else. The protocol offers no migration for this: the old and new values are unrelated strings, and only the provider is in a position to map them. So the decision is made once — `public` where a single shared identity across your own clients is the product, `pairwise` where clients must not be able to pool records, and a sector where you want correlation inside one estate and none outside it. The cost of getting it wrong is paid by every relying party, not by the provider.
go deeper
Understand that the identifier an application stores for a signed-in user comes from the provider, and is only meaningful next to the issuer that produced it.
Explain which inputs change a subject value - the client's sector and the subject type in force - and why that makes the value an outcome of configuration rather than a property of the person.
Show that you establish a provider's subject policy before any relying party stores an identifier, and that you can describe the blast radius of changing it once accounts exist.
Own the trade-off across the estate: which products must recognise one person, which must be unable to, and how little of that remains changeable after the first integration goes live.
## The decision that looks like configuration Setting `subject_type` reads like a switch. It is not: it is the moment a provider decides what its identifiers **mean**, and every relying party that subsequently stores one inherits that decision permanently. The reason is simple and worth stating flatly — a subject value is derived, and the subject type and the client's sector are inputs to the derivation. Change an input and you change every output. That makes this one of the few genuinely one-way doors in an identity integration. Most protocol choices can be revised: an algorithm can be rotated, an endpoint moved, a scope added. A subject identifier policy cannot, because its consequences are already sitting in other people's databases. ## What actually breaks Suppose one relying party has a year of accounts keyed on public subject values and the provider is switched to pairwise. 1. At the next sign-in, the provider derives a value for that client under the new policy. 2. The value is unrelated to the one the client stored. It is not a transformation of it, and nothing in the token signals that it replaced anything. 3. The client's lookup misses, so a returning user is treated as new. 4. Either a duplicate account is created, or the user is stopped at a screen that cannot be satisfied — both of which are visible to every returning user at once, not gradually. 5. Recovery means re-identifying each user through some other verified attribute and rewriting the stored key, which is work on the relying party's side that the provider cannot do for it. Moving a client into or out of a sector does the same damage by the same mechanism. It is not a rollback path either: leaving a sector produces a third set of values unrelated to both earlier sets. ## The trade-off to reason about up front | Policy | You get | You give up | |---|---|---| | `public` | every client recognises one person with no linking work | any client, including third parties, can pool records with any other | | `pairwise`, no grouping | clients cannot correlate at all | your own products cannot recognise a shared user either | | `pairwise` with a sector | correlation inside the grouped set, none outside it | a grouping that is itself fixed once those clients store identifiers | The useful framing for a lead is not "which is more secure" — it is **which parties should be able to prove they are looking at the same person**. That is a product and privacy question with a protocol answer, and it has to be asked before the first integration rather than after. Some concrete inputs to it: - **Who are the clients?** A provider serving only applications one organisation runs has a very different correlation risk from one onboarding third parties. - **Does a product require recognition across clients?** Shared entitlements, a single support view, or cross-product history all assume correlation and quietly get it from `public`. - **What is promised to users?** If the privacy statement implies products cannot pool records, `public` contradicts it whatever the intent was. - **How many relying parties already exist?** The cost of the decision scales with the number of parties holding stored identifiers, which is why "before the first one" is the cheapest moment there will ever be. - **Can the estate be grouped instead of flattened?** A sector is usually the honest answer when the tension is internal convenience versus external correlation. ## What does not rescue a late change - **Rotating gradually does not help.** There is no dual-value period in the protocol; the token carries one `sub`. - **The provider's own mapping is not part of the protocol.** A provider could, as an operational favour, export an old-to-new mapping for one client during a migration, but nothing in the specification does it and it partly undoes the separation for that client. - **Falling back to the email address is a different bug.** An email is reassignable and changeable, which is exactly why it is not the account key; reaching for it under migration pressure trades a one-off outage for a permanent identity defect. - **Adding a sector later does not restore old values.** It produces new ones. ## The answer in one line A subject identifier policy is not a setting that can be revisited, because it is already embodied in every identifier every relying party has stored. Decide what may be correlated with what, write it down, and integrate against that — the first relying party to store a value has already made the decision permanent whether or not anyone chose it deliberately.
- Your own three clients need one shared identity but partners must not correlate - what do you choose?Pairwise as the provider-wide policy, with the three in-house clients registered into one sector so they share a value, and every partner's client left outside it. That buys correlation inside the estate and none across it, without making the permissive option the default for parties you do not run.
- A relying party has stored public sub values and you now want pairwise - what does that cost?Every stored value is orphaned at once. That party must re-identify each returning user by some verified attribute of its own and rewrite its records, and until it does, returning users create duplicates. A provider could export a one-time mapping for that client as an operational favour, but nothing in the protocol does it.
- Does moving a client out of a sector have the same effect as switching it to public?Not the same values, but the same damage. The sector is an input to the derivation, so leaving one yields a fresh set of pairwise values unrelated to the stored ones. The client is stranded exactly as a subject type change would strand it.
saying these in an interview costs you the question
- You can switch a client to pairwise later and migrate the stored values across
- Pairwise is strictly safer, so there is no decision to make
- Keying accounts on the email claim avoids the whole problem
- Only the provider is affected, relying parties just read whatever sub arrives
- A sub value is stable across providers, so the user is recognisable either way