skip to content

How do you design tests that catch tenant crossing and privilege escalation in a multi-tenant system?

level: seniorimportance: should knowfreq 46%

answer

  1. One account can never prove isolation
  2. Two populated scopes, distinguishable data
  3. Replay the happy path as a stranger
  4. Lists, exports and batches leak too
  5. Snapshot state, not just the status

basics

~20 s

Seed two fully populated accounts with distinguishable data, then replay every recorded happy-path request with a session from the other account and with lower-privileged roles. Each must be refused, change no state, and leak nothing about the other account.

solid answer

~50 s

Separate the two shapes: vertical escalation is reaching an operation above your privilege level, horizontal escalation — tenant crossing — is reaching another account's record at your own level. Horizontal is the commoner gap and is invisible to a suite that uses a single account, so the fixture must seed two fully populated scopes with data that cannot be confused. The core technique is identity-substitution replay: record the requests the happy path issues for one scope and re-send each one unchanged except for the session, using every actor from the other scope. Extend past single-record fetches to list and search endpoints, aggregates, exports, batch requests containing one foreign identifier, asynchronous jobs and cached responses. Assert more than the status: snapshot the target scope's state and prove it is unchanged, check the body carries nothing foreign, and check the denied attempt is recorded.

code

pseudocode · 14 lines
pseudocode
school_a = seed_school(name="Northgate", rooms=14, classes=37)
school_b = seed_school(name="Riverbank", rooms=9,  classes=23)

recorded = record_happy_path(as=school_a.planner)

for request in recorded:
    before = snapshot(school_a)
    replayed = request.with_session(sign_in_as(school_b.planner))
    result = send(replayed)

    assert result.status in [403, 404]
    assert not result.body.contains("Northgate")
    assert snapshot(school_a) == before
    assert audit_contains(school_b.planner, request.operation, "denied")

go deeper

for a junior

Know the two words and what separates them: vertical escalation reaches a higher privilege level, horizontal escalation reaches a peer's or another account's data at your own level.

for a middle

Explain the fixture. An interviewer expects you to say why one seeded account cannot prove isolation and how identity-substitution replay converts an existing functional suite into a negative one.

for a senior

Show that you have lived through the leaky surfaces — list endpoints, aggregates, async exports, batch items, caches — and that you assert unchanged state rather than trusting a refusal status.

for a principal

Own isolation as a property with an owner and a budget: which tier these cases run in, how coverage is measured against the operation list rather than a test count, and what the release rule is when a crossing is found.

### Two shapes of the same failure **Vertical escalation** is an actor obtaining an operation reserved for a higher privilege level - a student performing an operation only a planner may perform. **Horizontal escalation**, usually called **tenant crossing** when accounts are the isolation boundary, is an actor at the correct privilege level reaching a record that belongs to a different actor or a different account - a planner at one school reading, or worse editing, another school's timetable. Both are enforcement gaps; they need different fixtures, and the horizontal one is far more common and far more often untested, because every functional test in the suite uses a single account and therefore can never observe it. ### The fixture is the whole trick You cannot test isolation with one populated account. The fixture must seed at least two fully populated scopes with **distinguishable** data: a school with fourteen rooms and one with nine, names that never collide, identifiers that are not adjacent integers so a crossing is not mistaken for an off-by-one. Seed each scope with the same *kinds* of records, so every operation has a valid target in both. Then hold, for each scope, a session for every role in the matrix. That grid of sessions and targets is what the negative cases iterate over. The core technique is **identity substitution replay**. Record the requests a happy path issues for scope A. For every recorded request, re-issue it unchanged except for the session, using each actor from scope B and each lower-privileged actor from scope A. Every one of those must be refused. This converts an existing functional suite into a negative suite at low cost and, crucially, it stays current: when a new operation is added to the happy path, the replay picks it up automatically. ### The surfaces people forget An operation-by-operation replay of the primary read and write endpoints catches the obvious cases. Production incidents come from the surfaces that are not on that list: - **Collection and search endpoints.** The single-record fetch is guarded; the list endpoint filters by a parameter the caller supplies rather than by the session's scope. - **Aggregates and counts.** A total, a chart or an availability calendar computed across scopes discloses the other scope's data without ever returning a record. - **Exports and reports.** Generated asynchronously, often by a job that runs with elevated rights and takes the requested scope from the job payload. - **Batch operations.** A request containing eleven identifiers where ten belong to the caller and one does not. The correct behaviour must be decided - refuse the whole batch, or apply the permitted subset - and then asserted; the common defect is that the check runs on the first item only. - **Asynchronous continuations.** Webhooks, retries, scheduled jobs and message consumers that re-execute work later, frequently without the originating actor's context. - **Caches.** A response cached under a key that omits the scope, returned to the next caller. - **Refusal shape as an oracle.** If a forbidden-but-existing record returns one refusal and a non-existent one returns another, the pair of responses answers a question the caller was not entitled to ask. Assert that the two are indistinguishable where the product's rules require it. ### Asserting the refusal completely A status code alone is a weak assertion for escalation cases, for one specific reason: the check may run *after* the effect. In the timetable planner, booking a room appended the reservation and then evaluated whether the caller owned the class. The refusal was returned correctly, the room was already consumed, and the defect surfaced as phantom unavailability that a four-person team chased for three weeks as a calendar bug. The case that would have caught it snapshots the target scope's state before the attempt and asserts it is identical afterwards - not merely that the response said no. The same case should assert that the response body carries nothing from the other scope, and that the denied attempt appears in the audit trail with the acting identity, since an escalation attempt you cannot see afterwards is one you cannot investigate. ### Where these run and who owns them Run them at the service boundary rather than through the interface. The interface hides the control from the wrong role, so an interface-level case proves only that the button is absent - the whole point of an escalation case is a request the interface would never send. Keep them in the fast tier, not the nightly pack: they are cheap, they are deterministic, and an isolation regression must block a merge rather than be discovered the next morning. Finally, treat coverage of this area as a property of the matrix - every operation replayed cross-scope - rather than a count of tests, because a suite with two hundred escalation cases that all target one endpoint is worse than nineteen that target all nineteen operations.

  • Which surfaces leak across accounts most often after the single-record endpoints are guarded?
    Collection and search endpoints that filter by a caller-supplied parameter rather than the session's scope; aggregates and availability views computed across scopes; asynchronous exports run by a job that takes the scope from its payload; batch requests where the check runs only on the first item; retries and message consumers that lose the originating actor's context; and responses cached under a key that omits the scope.
  • Should a forbidden record and a non-existent record return the same refusal?
    It depends on a rule the product must state, and then the case asserts it. If the two responses differ, the pair answers a question the caller was not entitled to ask — whether that record exists — which turns an enumeration into a disclosure. Where the rule says they must be indistinguishable, assert that explicitly, including timing where the difference is measurable.
  • Why run these at the service boundary rather than through the interface?
    Because the interface hides the control from the wrong role, so an interface-level case proves only that a button is absent. The whole point of an escalation case is a request the interface would never send. Driving the service boundary also makes the cases fast and deterministic enough to sit in the tier that blocks a merge, which is where an isolation regression belongs.

saying these in an interview costs you the question

  • Testing isolation with a single seeded account
  • Asserting the status code without snapshotting state
  • Checking single-record reads but never list or export endpoints
  • Assuming a batch is safe because its first item was checked
  • Driving escalation cases only through the interface
  • Treating a hidden control as proof the operation is refused

context