skip to content

questions

3

In OpenID Connect, what does a subject_type of pairwise change about the sub claim compared with public?

level: middleimportance: should knowfreq 48%

answer

  1. one person, two names
  2. who can compare user tables
  3. correlation is the decision
  4. same value everywhere, or per client
  5. subject_type: public or pairwise

basics

~20 s

A pairwise subject_type makes the provider issue a different sub value to each client for the same end user, so no two clients can match their user records. Public gives every client the identical sub.

solid answer

~40 s

`sub` is the identifier a provider uses to name the end user who just authenticated. With `subject_type` set to `public`, every client of that provider receives the same string for that person — convenient, and exactly what makes cross-client correlation possible. With `pairwise`, the provider derives a different `sub` per client, a **Pairwise Pseudonymous Identifier (PPID)**, so two clients comparing their user tables find no overlap. Nothing else about the claim changes: the value is still opaque, still stable for that client, and still only unique within its issuer, so an application still keys accounts on the `(iss, sub)` pair. A provider advertises what it offers in `subject_types_supported`, and where more than one type is listed a client MAY state a preferred `subject_type` at registration. Where no `subject_types_supported` is advertised, `public` applies.

code

json · 16 lines
json
[
  {
    "iss": "https://login.example-theatre.org",
    "aud": "volunteer-scheduler",
    "sub": "a1b6f4c0d2e8934755f0c1aa3b7d2e19",
    "iat": 1758294000,
    "exp": 1758297600
  },
  {
    "iss": "https://login.example-theatre.org",
    "aud": "donor-portal",
    "sub": "7c2e05ab913d4f60b8a1e7d34c05f8b2",
    "iat": 1758294120,
    "exp": 1758297720
  }
]

go deeper

for a junior

Recall that sub names the authenticated user and that the value is chosen by the provider, not by the application. Know that one person can arrive at two applications under two different sub values without anything being wrong.

for a middle

Explain that subject_type selects between one shared value and one value per client, that public is what applies when a provider advertises no supported types, and that the identifier stays opaque and stable either way.

for a senior

Show that you establish which subject type a provider issues before storing anything keyed on it, and that you can state what correlation the choice permits and what it costs support tooling and analytics later.

for a principal

Weigh privacy against the product's need to recognise one person across an estate, and own the fact that the answer is baked into every identifier already stored.

## What `sub` actually names Every ID token an **OpenID Provider** issues carries a `sub` claim: the identifier that provider uses for the end user who has just authenticated. It is an opaque string — the specification gives it no internal structure a client may parse — and it is meaningful only inside that issuer's own namespace, which is why an application stores accounts against the `(iss, sub)` pair rather than against `sub` on its own. The part that catches people out is that `sub` is not a fact the provider looks up and reports. It is a **name the provider chooses to hand to a particular client**, and `subject_type` is the setting that decides how that name is chosen. The same person, authenticating to two clients of one provider in the same minute, can legitimately receive two entirely unrelated `sub` values — not because something is broken, but because the provider was configured to name that person separately for each client. ## The two values | | `public` | `pairwise` | |---|---|---| | Value seen per client | identical for every client | different for each client | | Two clients joining user tables | they match on `sub` | they find no overlap | | Who can still correlate | anyone holding both records | only the provider | | Stability for one client | stable | stable | | Typical fit | clients one organisation runs and wants to share an identity | third-party clients, or anywhere client collusion is a concern | Both values are advertised the same way: the provider lists the types it is willing to issue in `subject_types_supported`, and a client that has a preference and finds more than one type listed MAY state a `subject_type` when it registers. A provider that advertises no `subject_types_supported` is taken to issue `public` values — worth knowing precisely because it is the quiet default that permits correlation. ## What pairwise hides, and from whom - It hides the link **from the clients**. Two relying parties cannot discover that they have the same user by comparing stored identifiers. - It does **not** hide the link from the provider. The provider derives both values and must be able to reproduce each one on every sign-in, so the mapping is retained on its side for the life of the account. - It does not hide anything the user volunteers elsewhere. If both clients also ask for and store an email address, the two records join on that instead, and the pairwise subject has bought nothing. - It is not encryption. A pairwise `sub` is a different value, not a ciphertext addressed to one client; there is nothing to decrypt and nothing a client is expected to unwrap. - It is not a per-session value. The identifier is stable for that client, which is what makes it usable as an account key at all; a value that changed on every login would be useless for lookup. - It says nothing about the other claims. The profile and email claim sets a provider returns are governed by the scopes requested, not by the subject type. ## How a client ends up with one or the other 1. The provider decides which subject types it supports at all and advertises them in `subject_types_supported`. 2. Where more than one is advertised, a client MAY register a preferred `subject_type`. Where only one is advertised, the client takes what it is given. 3. Where the provider advertises no `subject_types_supported`, `public` is what applies. 4. On each authentication, the provider derives the value for that client and puts it in the ID token's `sub` claim. None of those steps is something the relying party computes. The client's job is to store the value it is handed, alongside the `iss` that produced it, and to look the user up by that pair on the next sign-in. ## Why the choice is a design question Three applications run by one organisation against one provider make the trade-off concrete. If the provider issues `public` subjects, all three recognise the same person immediately: a volunteer who also buys a ticket and makes a donation is one identity across the estate, with no linking work. That is a feature if the organisation wants a single view of that person, and a privacy leak if it does not — and it is the same mechanism either way. If the provider issues `pairwise` subjects, each application sees a stranger, and the estate has to either group those clients deliberately or accept three separate identities. The consequence worth internalising is that a subject value is a **product of configuration**, not a property of the human being. Anything a relying party has stored under one configuration stops matching if the configuration changes, and the protocol offers no migration for that. So the correlation question is answered once, before the first relying party stores anything, and the answer is then effectively fixed.

  • Does a pairwise sub stop the provider from knowing that two clients saw the same person?
    No. The provider derives both values and retains the mapping, because it has to reproduce each one on every sign-in. Pairwise withholds correlation from the clients, not from the issuer, so it is a control against client collusion rather than against the provider.
  • Is a pairwise sub stable for a given user and client over time?
    Yes, and it has to be — an account lookup depends on it. It is a stable per-client name, not a per-session one. It changes only when an input to the derivation changes: the client's grouping, or the subject type in force for that client.
  • If a provider supports both types, who decides which one a given client receives?
    The provider publishes what it supports in `subject_types_supported`, and where more than one type is listed a client MAY state a preferred `subject_type` when it registers. A client cannot obtain a type the provider does not advertise.

A writer who publishes under a different pen-name with each house: no two houses can tell they bought from the same person, while the agent who arranged every deal always can. Pairwise withholds correlation from the clients, never from the provider.

saying these in an interview costs you the question

  • Pairwise encrypts the sub so only the right client can read it
  • With pairwise, even the provider cannot link the two accounts
  • Just key accounts on the email claim and subject_type stops mattering
  • subject_type is something the relying party computes on its side
  • A pairwise sub is regenerated on every login, so it cannot be stored
open as a page

In OpenID Connect, why must a provider settle its public-or-pairwise subject_type policy before the first relying party integrates?

level: principalimportance: should knowfreq 34%

basics

~20 s

Every sub a relying party has stored is a product of that policy. Change subject_type, or move a client between sectors, and the provider starts issuing unrelated values, orphaning every account record keyed on the old ones.

open as a page

In OpenID Connect, how can three clients run by one organisation receive the same pairwise sub for one user?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Register all three clients with the same sector_identifier_uri. A pairwise value is derived per sector rather than per client, and the Sector Identifier is that URI's host component, so the three resolve to one sector and one shared sub.

open as a page