skip to content

SAML Service-Provider Integration

Being the service provider: publish metadata, consume a signed assertion at an ACS endpoint, and keep a replay cache. Certificate rotation and signature wrapping are what interviewers push on.

on this pageshow

questions

6

Your permit system's assertion consumer service is an unauthenticated POST endpoint on the public internet — what do you harden on it?

level: middleimportance: must knowfreq 55%

answer

  1. a stranger's browser posts here
  2. hostile input, not a trusted caller
  3. cap the bytes, then the parser
  4. generic failure, correlation id in the log
  5. no session before the checks pass

basics

~20 s

Treat every request as hostile input from a stranger: cap the body size before reading it, parse the XML with document type definitions and external entity resolution disabled, rate-limit per connection, return a generic failure, and issue no session until every check has passed.

solid answer

~50 s

The assertion consumer service is a URL I publish, so anyone can post to it — there is no session and no credential on that request, and the browser sending it was pointed at me by an identity provider I do not run. I cap the request size before I read the body, because a base64 document expands on decode and a large one is a cheap way to burn parser CPU. I parse with a hardened configuration: no document type definitions, no external entity resolution, no network fetches for schemas, an explicit expansion limit. I rate-limit per source and per counterparty connection, and I keep error counters per connection so one misconfigured firm shows up as one tenant failing rather than as noise. Failures return a single generic message plus a correlation id; the detail goes to logs, and the raw message never does, because it carries personal data. Nothing writes a session record until verification and validation have both passed.

code

http · 6 lines
http
POST /sso/saml/acs HTTP/1.1
Host: permits.example
Content-Type: application/x-www-form-urlencoded
Content-Length: 9214

SAMLResponse=PHNhbWxwOlJlc3BvbnNlIC4uLjwvc2FtbHA6UmVzcG9uc2U%2B&RelayState=permit-4812

go deeper

for a junior

Remember that this endpoint is reachable by anybody, and that the browser posting to it was sent by an identity provider you do not operate. Nothing about the request proves anything until your code checks the message it carries.

for a middle

Explain the order out loud: size limit, hardened parse, signature verification, structural checks, single-use claim, then session issuance. Name the parser switches you change and say why an early size cap matters when the payload expands twice.

for a senior

Show the operational half: per-connection rate limits and error counters, generic failure responses carrying a correlation id, and a logging rule that lets support diagnose without turning the log into a store of credentials and personal data.

for a principal

Argue what this single public endpoint exposes across every counterparty you will ever onboard, and what you standardise so one firm's misconfiguration cannot become an availability problem for the rest of the estate.

## The endpoint you publish is a hole you opened on purpose A service provider's **assertion consumer service** is a URL you advertise in your own metadata and hand to every counterparty you onboard. From a browser's point of view it is an ordinary form target: an identity provider you do not run renders a page that posts to it, and what arrives at your server is a form submission carrying a base64 `SAMLResponse` field and an opaque `RelayState` field. There is no session on that request, no bearer credential, no authenticated caller, and no way to restrict who reaches it — a counterparty's people sign in from wherever they are. Anyone on the internet can post anything to it, as often as they like. Everything that protects it is code you wrote and operate. For a regional water authority's contractor permit-to-work system, that is the whole shape of the problem: staff sign in locally, but each contractor firm authenticates its own people at its own identity provider, and the only thing that arrives from those firms is a form post at this one URL. ## Cap the bytes before you have XML The first rule is that the request must be bounded before it becomes a document. A base64 payload expands when decoded, and an XML document expands again when parsed into a tree. A generous or absent body limit turns a single request into a memory and CPU event. - Set an explicit maximum request size on this route, sized against the largest legitimate message a counterparty sends plus headroom — not against your API's general default. - Reject over the limit **before** decoding, with a status and nothing else. - Put a wall-clock budget on parsing and verification so a pathological document cannot hold a worker thread. ## The parser is the part that bites XML parsers are configurable, and several of the switches that matter are permissive unless you turn them off. | Parser behaviour | Common default | What you set on this route | |---|---|---| | External entity resolution | enabled | disabled | | Document type definition processing | enabled | disabled | | Entity expansion | generous or unbounded | small, explicit limit | | Fetching a schema over the network | permitted | forbidden | The reason is blunt: this document came from a stranger, and a parser that will resolve references inside it will make outbound requests, or allocate memory, on a stranger's instruction. This is configuration you assert deliberately, per route, and cover with a test that feeds the endpoint a document with a declaration in it and asserts a rejection. ## Why this endpoint sits outside your cross-site request protection Every other state-changing POST in the permit system is expected to carry a same-site token issued into an existing session. This one cannot. The post is cross-site by construction, and the party that sends it never had a session with you, so there was no moment at which a token could have been issued. Excluding this route from that protection is correct, and saying so out loud in a design review is part of the job. What replaces it is not nothing. Two properties carry the weight instead: 1. The message is **signed by a key you already hold for that connection**, so a forged body does not verify. 2. The message is **single-use**, so a captured one cannot be posted again. A candidate who answers *add a token like everywhere else* has not understood which party is posting. ## The order the endpoint runs in 1. Reject on size. 2. Decode and parse with the hardened configuration. 3. Verify the signature against the certificate configured for this connection. 4. Run the structural checks. 5. Claim the message id as used, atomically. 6. **Only then** look up an account and issue a local session. Nothing before step 6 may write a session record or a cookie, and nothing before step 3 may be treated as content. ## Failing without teaching the sender A rejection should return one generic message and a correlation id the firm can quote to your support desk. Echoing the parser's complaint, the certificate subject that failed to match, or the clock difference that fell outside tolerance hands a prober a checklist. The detail belongs in your logs, and the logs need a rule of their own: record the message id, the issuing counterparty and the failure class, and do **not** record the raw document. A valid one is a credential and carries personal data about a named worker. ## Rate limiting and the metrics that name the firm Limit by source address and by connection. Keep the error counters **per connection**, because the failure you actually get paged for is one firm's identity provider changing something at three in the morning, which looks like a flat line on a global error rate and like a cliff on that firm's.

  • Why can you not simply put the assertion consumer service behind the same authentication filter as the rest of the application?
    Because this endpoint is how a person becomes authenticated. The request that arrives is the credential; requiring an existing session to reach it makes federated sign-in impossible. The route is deliberately public, and the checks that protect it are the signature and the single-use claim rather than an ambient session.
  • What do you log on a rejected message, and what must never reach the log?
    Log the message id, the counterparty connection, the failure class and a correlation id you also return. Do not log the decoded document: a valid one is a credential, and it carries a named worker's attributes. If support needs more, capture it behind an explicit, time-boxed diagnostic switch on that one connection.
  • The endpoint starts returning failures for one contractor firm only. Which metric tells you that fastest?
    A per-connection failure counter split by failure class. A global error rate hides one tenant entirely, because the rest of the estate is fine. Per-connection counters turn the page into a sentence: this firm, signature failures, starting at this minute — which is usually a certificate that moved.

It is a letterbox cut into the front of the building, at street level, with your address printed in a public directory. You cannot stop anyone posting through it. What you can do is make the slot too small for anything dangerous, open the envelopes in a room with nothing flammable in it, and check the seal before you act on the contents.

saying these in an interview costs you the question

  • It sits behind the identity provider, so requests arriving there are already authenticated.
  • Add a same-site request token to this endpoint like every other POST.
  • Return the parser or validation error so the firm's team can debug it.
  • The XML library's defaults are safe enough for a document from a partner.
  • Log the whole decoded message so support can read what the firm sent.
  • Issue the session first and finish validating afterwards, to keep sign-in fast.
open as a page

Which checks on an inbound SAML response must you configure yourself, and which does an integration library leave off by default?

level: middleimportance: must knowfreq 50%

basics

~20 s

A library verifies the signature; the bindings to you are yours to configure. Set the expected audience to your own entityID, require the recipient and destination to be your endpoint, bound the validity window with an explicit skew, correlate InResponseTo against a stored request id, and process only the element verification returned.

open as a page

Where does your service provider's assertion replay cache live when several replicas sit behind a load balancer, and how long do entries last?

level: seniorimportance: must knowfreq 45%

basics

~20 s

In a store every replica shares, claimed with one atomic set-if-absent keyed on issuer plus message id, with a time-to-live running to the end of the message's validity window plus the clock skew you allow. A per-process cache leaves a replay hole at the load balancer.

open as a page

A contractor firm rotates its assertion-signing certificate without telling you — how does your service provider survive that?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Hold a set of accepted signing certificates per connection rather than one, refresh each counterparty's metadata on a schedule and honour its cacheDuration and validUntil, keep the last good copy when a fetch fails, and alarm per connection so one firm's rotation is visible before its people are.

open as a page

What does your service provider store at login so that a later logout message naming a SessionIndex can find the right session?

level: seniorimportance: should knowfreq 32%

basics

~20 s

Record, against the local session you issue, the issuing counterparty, the subject identifier as sent, and the SessionIndex from the assertion — indexed so a lookup by those three is a single read. Without that row the inbound message names a session at the other side that you cannot resolve to one of yours.

open as a page

A contractor firm's identity provider cannot meet one of your validation requirements — how far do you bend, and who decides?

level: principalimportance: should knowfreq 25%

basics

~20 s

Rank the requests by blast radius: clock skew is bounded and reversible, endpoint-binding checks cost you once you have more than one endpoint, and audience and single use are not negotiable. Grant the rest as per-connection data with an owner, a reason and a review date — never as a global default or a code branch.

open as a page