skip to content

Which SCIM 2.0 filter and pagination subset do real provisioning clients send, and how do you serve it safely?

level: seniorimportance: should knowfreq 26%

answer

  1. clients send one operator, not a grammar
  2. parse it, never paste it
  3. refuse before you scan
  4. one-based, and short is not last
  5. invalidFilter and tooMany

basics

~20 s

In practice clients send little more than an equality filter on userName or externalId, occasionally co or sw. Serve it by parsing to a bound, parameterised query on indexed columns, refuse what you cannot evaluate with scimType invalidFilter, and cap results rather than scanning.

solid answer

~50 s

The `filter` grammar in SCIM 2.0 is large — comparison operators including `eq`, `co`, `sw` and `pr`, boolean composition, grouping and value paths — and almost none of it arrives. What real provisioning clients send is an equality match on the attribute they key on: `filter=userName eq "r.okonkwo"`, sometimes `externalId eq` or `displayName eq` for a group. Implement that narrow subset properly rather than the grammar badly: parse to an abstract expression, map attributes onto **indexed** columns, and bind values as parameters — never interpolate the filter text into a query. Refuse what you will not serve with a 400 and `scimType` `invalidFilter`, and cap an over-broad result with `tooMany`. On pagination, `startIndex` is **1-based**, and the response carries `totalResults`, `itemsPerPage` and `startIndex`. `itemsPerPage` may be smaller than the `count` asked for, so a short page does not mean the end of the collection.

code

http · 15 lines
http
GET /scim/v2/Users?filter=userName%20eq%20%22r.okonkwo%22&startIndex=1&count=50 HTTP/1.1
Authorization: Bearer <credential scoped to one practice group>

HTTP/1.1 200 OK
Content-Type: application/scim+json

{
  "schemas": ["urn:ietf:params:scim:api:messages:2.0:ListResponse"],
  "totalResults": 1,
  "startIndex": 1,
  "itemsPerPage": 1,
  "Resources": [
    { "id": "2819c223-7f76-453a-919d-413861904646", "userName": "r.okonkwo", "active": true }
  ]
}

go deeper

for a junior

Know that a provisioning client finds people by sending a filter such as an equality match on userName, and that paging is done with a one-based startIndex and a requested count.

for a middle

Explain why the filter must be parsed into a bound query rather than pasted into one, and why a page shorter than the requested count does not mean the collection has ended.

for a senior

Show the subset you support mapped onto indexed attributes, an explicit refusal path using invalidFilter and tooMany, and an understanding that totalResults moves under a client that is paging while writes continue.

for a principal

Treat the supported subset as a published contract with a cost: every operator you accept is load a customer's administrator can turn on without telling you, so the limits belong in the configuration document rather than in an incident review.

## Two halves of one read contract Everything a provisioning client does that is not a blind write depends on being able to find a person in your radiography viewer. After a conflict on create it looks the user up. Before an update it confirms what it holds. When an administrator repairs a broken sync it walks the whole collection. All three go through `/Users` or `/Groups` with a `filter` and pagination parameters, and those two mechanisms are where a SCIM implementation most often becomes either an injection hole or a source of load nobody sized. ## The filter subset that actually arrives SCIM 2.0 defines a full expression grammar: attribute operators such as `eq` (equal), `co` (contains), `sw` (starts with) and `pr` (present), ordering comparisons, logical `and`, `or` and `not`, parenthesised grouping, and value paths in brackets for multi-valued attributes. The honest observation from running one of these endpoints is that a provisioning client sends a vanishingly small corner of it, because it is not exploring your data — it is resolving one identity it already knows about: - `userName eq "..."` — by far the most common, and the lookup after a create conflict. - `externalId eq "..."` — used by clients that key on their own identifier. - `displayName eq "..."` — the group equivalent. - occasionally `co` or `sw` behind a search box an administrator typed into. So the design decision is not *how do I implement the grammar* but *which corner do I implement well, and what do I do with the rest*. Three rules hold that together: 1. **Parse, do not substitute.** Turn the filter into an abstract expression and then into a bound, parameterised query. A filter is attacker-influenced text arriving from another organisation's software; concatenating it into a query language is the same defect as concatenating any other untrusted string, and the fact that it looks like a small DSL does not change that. 2. **Map only onto attributes you can serve from an index.** `userName` and `externalId` should be indexed per tenant anyway because uniqueness depends on it. An operator that cannot reach an index is a collection scan you have handed to a caller you do not control. 3. **Refuse loudly.** A filter you cannot evaluate gets a 400 with `scimType` `invalidFilter`. A filter whose result set exceeds the maximum you are willing to materialise gets `tooMany`. Both are far better than the two alternatives — quietly interpreting it as something adjacent, or serving it with a full scan that degrades the endpoint for every other customer. That pair of refusals is also what stops the contradiction between rules 2 and 3 from arising in production: a `co` on an unindexed attribute is not "served slowly", it is served up to an explicit cap and refused beyond it. ## Pagination, and the one rule clients get wrong A list response is a `ListResponse` document carrying `totalResults`, `itemsPerPage`, `startIndex` and a `Resources` array. The request parameters are `startIndex` and `count`. | element | where it lives | what it means | the trap | |---|---|---|---| | `startIndex` | request and response | **1-based** index of the first result | assuming 0, which silently skips or repeats the first person | | `count` | request | how many the client would like | it is a request, not a guarantee | | `itemsPerPage` | response | how many you actually returned | **may be fewer than `count`** | | `totalResults` | response | the size of the matching set | it moves while the client pages, because writes continue | The rule worth stating plainly is the third row: a page smaller than the requested `count` **does not mean the collection has ended**. A client that stops there stops early and concludes that half a practice group does not exist. And the same discipline binds any reader you write yourself: code in your own product that walks this endpoint must page until it has covered `totalResults` or received an empty page, not until it sees a short one. `totalResults` moving under a paging client is real and unavoidable — people are being created and deactivated while the walk runs — so offset-based paging can repeat or skip a person at a page boundary. You cannot fix that inside the specification's parameters. What you can do is make it harmless: never let a partial walk be interpreted as authoritative absence, and hold a stable sort so that at least the ordering does not shift under the offsets. ## The failure at 3 a.m. The pathology worth recognising is a customer's sync that has "worked for months" suddenly loading your database. Nothing changed on your side; the customer's administrator turned on a wider scope, and a filter that used to match one person now matches four hundred, paged fifty at a time, every fifteen minutes. If your filter path can only touch indexed attributes and refuses over-broad result sets with `tooMany`, this is a support conversation. If it cannot, it is an incident with your name on it and a customer who believes your product is slow.

  • Why is interpolating a SCIM filter string into your query language a defect even though the grammar looks constrained?
    Because the values inside it are free text chosen by whoever can reach the endpoint, and the grammar's shape gives you no guarantee about them. Parse the filter into an abstract expression, then build a parameterised query whose attribute names come from a fixed mapping you control and whose literals are bound values. The filter grammar is a parsing problem; treating it as a string-building problem is how it becomes an injection one.
  • A client asks for count=100 and you return 20 items, with more to come. What must the response say, and what must the client not conclude?
    The response carries `itemsPerPage` 20, the `startIndex` it began at, and `totalResults` for the whole matching set. The client must not conclude that the collection has ended: a page shorter than the requested `count` is explicitly permitted. It continues from the next index until it has covered `totalResults` or receives an empty page.
  • What do you do when a filter is valid and supported but matches far more resources than you want to materialise?
    Return 400 with `scimType` `tooMany`, and publish the maximum you will serve in your configuration document so the client can see the limit before it asks. That is preferable to streaming a very large result or silently truncating one, because truncation without a counter is indistinguishable, to the caller, from the people simply not existing.

saying these in an interview costs you the question

  • Builds the query by string-concatenating the filter expression
  • Assumes startIndex is zero-based
  • Stops paging when a page is shorter than count
  • Implements the whole filter grammar rather than an indexed subset
  • Silently truncates an over-broad result instead of refusing it
  • Treats totalResults as stable across a long walk