How does a client walk a 50,000-entry LDAP search with the simple paged results control?
answer
- size out, cookie back
- empty cookie starts and ends it
- echo it, never read it
- same search every time round
- size zero releases the server's state
basics
~20 sIt attaches the paged results control with a page size and an empty cookie, reads the entries, takes the cookie the server returns on the SearchResultDone, and repeats the identical search with that cookie. The loop ends when the returned cookie is zero-length.
solid answer
~40 sSimple paged results (1.2.840.113556.1.4.319) carries a `realSearchControlValue` of `size` and `cookie`. The first `SearchRequest` sets `size` to the page size and `cookie` to a zero-length string. The server returns up to that many `SearchResultEntry` messages, then a `SearchResultDone` with the control attached carrying an opaque `cookie`. The client sends the **same** search again — same base DN, LDAP search scope, filter and requested attributes, differing only in `messageID`, the `cookie` it echoes back, and possibly `size` — and repeats until the cookie comes back zero-length. The cookie is opaque: echo it byte for byte, never parse it, and do not carry it across connections. It also represents server-side state, so a loop you walk away from leaves that state behind until the server times it out.
code
asn1 · 5 linesrealSearchControlValue ::= SEQUENCE {
size INTEGER (0..maxInt),
-- requested page size on the request
-- an estimate of the result set on the response
cookie OCTET STRING }go deeper
Recall the shape of the loop: send a page size with an empty cookie, read the entries, take the cookie off the final response, send the same search again, and stop when the cookie comes back empty.
Explain why the cookie is opaque, why the repeated request must not drift, and why the size field in the response is an estimate a server may supply rather than a remaining count.
Show that a paged search is server-side state with a cost — release it with a zero page size when you abandon the loop, and design the consumer to tolerate entries added or removed mid-enumeration.
Weigh the estate-level trade: bounded messages and predictable memory against a non-snapshot read, and decide whether downstream consumers can reconcile drift or genuinely need a consistency mechanism instead.
## The problem paging solves A directory server caps what one search may return — its own administrative limit, and the client's `sizeLimit` on the `SearchRequest`. A search over a naming context that holds tens of thousands of directory entries hits that cap and stops, usually with `sizeLimitExceeded (4)` or `adminLimitExceeded (11)`. Neither the client nor the server wants a single message exchange to carry 50,000 entries anyway. Simple paged results, specified in `RFC 2696` and identified by the OID 1.2.840.113556.1.4.319, is the LDAP Control that turns one enormous search into a sequence of bounded ones. ## The control value The control's `controlValue` holds a `realSearchControlValue` with exactly two fields: - `size` — on the request, the number of entries the client wants in this page. On the response, a value the server **may** use as an estimate of the whole result set; many servers simply return zero, so a client must never treat it as a count of what is left. - `cookie` — an `OCTET STRING`. Zero-length on the first request. On each response it identifies the server's position in the result set. ## The loop, precisely 1. Build the `SearchRequest` — base DN, LDAP search scope, filter, requested attributes — and attach the control with `size` set and `cookie` zero-length. 2. Read the `SearchResultEntry` messages that arrive, up to `size` of them. 3. Read the `SearchResultDone`. The paging control is attached to it, carrying the cookie for the next page. 4. If that cookie is zero-length, the result set is exhausted; stop. 5. Otherwise send the search again with a fresh `messageID`, the same base DN, scope, filter and attribute selection, and the cookie you were just handed. Go to step 2. Three rules make or break this loop: - **The cookie is opaque.** It is the server's bookmark in whatever form that server chose — an index position, a key, a handle into a cached result. Parsing it, editing it, logging it as though it were meaningful, or reusing it on a different connection are all mistakes. - **The repeated request must be identical** except for `messageID`, the cookie, and the page size. Change the filter or the base DN and you are asking a different question with someone else's bookmark; a server may reject the request or restart the enumeration. - **The cookie is state on the server.** An in-progress paged search costs the server memory for as long as the client might come back. That is the hidden cost of this control and the reason the next point exists. ## Abandoning a page loop politely A client that stops half way — the importer crashed, the operator cancelled, an error broke the loop — leaves the server holding that state until whatever timeout it applies. The graceful exit is to send the search one more time with the **current cookie and a `size` of zero**, which tells the server the client is done and the state can be released. A long-running booking-system sync that crashes nightly without doing this accumulates abandoned page state on every run, which is how a directory server that was fine for a year starts refusing paged searches on a Monday morning. ## What paging does not give you | assumption | reality | |---|---| | "pages come back in a stable order" | ordering is not promised; pair it with the server-side sorting request if order matters | | "paging removes the server's limits" | administrative limits still apply, and can still end the enumeration | | "the `size` in the response is the total" | it is an estimate a server **may** supply, and is commonly zero | | "the cookie survives a reconnect" | it belongs to that server and that search; a new connection starts over | | "an entry appears exactly once" | the underlying data can change between pages; entries can be missed or repeated | That last row is the one people are surprised by. Paging is not a snapshot. Between page three and page four, a researcher can be added to the group and a laboratory can be closed. A client that needs a consistent picture needs either a mechanism designed to give it one, or a reconciliation step that tolerates the drift. ## What an interviewer is listening for The candidate who has only read the API documentation says "you set a page size and loop". The candidate who has run it says three more things: the cookie is opaque and must be echoed unchanged, the repeated request must not drift, and the loop holds state on the server that somebody has to release.
- A client must stop a paged search half way through. What should it send?The same search once more with the cookie it currently holds and a `size` of zero. That tells the server the enumeration is over and lets it release the state it was holding. Simply dropping the loop leaves that state until the server's own timeout, and a job that crashes nightly accumulates it.
- Can the client change the page size between pages?Yes — `size` is one of the few things allowed to differ between requests in the loop, alongside the `messageID` and the cookie. The base DN, LDAP search scope, filter and requested attributes must stay as they were, because the cookie is a bookmark into the result of that exact search.
- Does paging guarantee every matching directory entry appears exactly once?No. The enumeration is not a snapshot: directory entries can be added, deleted or renamed while the loop runs, so an entry can be missed or seen twice. A caller that needs a consistent picture must either use a mechanism designed for that or reconcile the drift itself.
saying these in an interview costs you the question
- Calls the paged-results cookie an HTTP session cookie
- Parses the cookie to recover the offset it encodes
- Changes the LDAP search filter between pages
- Believes paging removes the server's own size limits
- Assumes pages arrive in a stable order without a sort control
- Abandons the loop without releasing the server's page state