When should permissions be copied into an index beside the matters table, and what does the delay between raising a wall and the listing obeying it commit you to?
answer
- a copy is fast and behind
- write the propagation budget down
- grants and removals are not symmetric
- the sweep is the only detector
basics
~20 sCopy permissions locally when the permitted set is too large to enumerate and the listing is hot. The copy commits you to a written propagation budget, a lag metric, a reconciliation sweep, and removals propagated faster than grants.
solid answer
~50 sThe copy is for the case where you cannot push identifiers into the query: the permitted set runs to six figures, or it is defined by a rule, or a remote authorization call per page would be the endpoint's whole latency budget. You maintain a table of who-may-see-which-matter in the same database as the matters, indexed, and the listing predicate becomes a local `EXISTS` that costs the same for a lawyer who may see nine matters and one who may see ninety thousand. What you buy it with is a window in which the listing still shows a matter a wall now forbids. Own that window explicitly: write the budget down, export the lag as a metric, and run a sweep that re-derives permitted sets from the authoritative source and counts divergences. Treat removals and grants asymmetrically — a late grant is a support ticket, a late removal is the incident.
code
sql · 12 lines-- The copy, maintained from the conflicts system, living beside the matters.
CREATE TABLE matter_visibility (
principal_id bigint NOT NULL,
matter_id bigint NOT NULL,
source_changed_at timestamp NOT NULL, -- when the authoritative change happened
applied_at timestamp NOT NULL, -- when this row was written here
PRIMARY KEY (principal_id, matter_id)
);
-- Removals are what leak, so they get their own index: raising a wall
-- must delete every principal's row for one matter cheaply.
CREATE INDEX matter_visibility_by_matter ON matter_visibility (matter_id);go deeper
Understand the shape first: a copy of permission data kept next to the rows makes a listing cheap, and a copy is always a little behind whatever it was copied from.
Be able to design the table and its predicate — what the rows are keyed by, which index the listing uses, and which index the wall-raising path needs to clear a matter in one statement.
Show the operating story: a written propagation budget, a lag metric taken from the source's own change timestamp, a reconciliation sweep, and removals applied on the write path rather than behind the same queue as grants.
Own the bet. You are trading a bounded, measured window of wrong visibility for a listing whose cost no longer scales with entitlement, and the trade is only defensible if the detector ships with it and the migration proves both paths agreed.
## When copying permissions earns its place A listing has to page over matters while a wall recorded in another system decides which are visible. If the permitted set is small, fetch its identifiers and push them into the query; nothing below applies. **The copy is for the case where you cannot**: the permitted set runs to six figures, or it is defined by a rule rather than a list, or the listing is hot enough that one remote authorization call per page consumes the endpoint's entire latency budget. What you build is a table in the same database as the matters — one row per principal-and-matter pair, or per group-and-matter where the permission model has an indirection that keeps the table small — indexed so the listing predicate is a local `EXISTS`. The listing then costs what any indexed listing costs, and costs the same for a lawyer who may see nine matters and one who may see ninety thousand. Everything after this follows from one fact: **the table is a copy**, and a copy has a direction and a delay. ## The delay you now own Between the moment a wall goes up in the conflicts system and the moment the matching rows leave your copy, the listing still shows the matter. That window is not an implementation detail to leave unspecified; in a firm with ethical walls it is a commitment with a regulator behind it. Own it in four explicit pieces: - **Budget it.** Write the number down: a wall takes effect in the listing within N seconds at the 99th percentile, within N minutes at worst. A number nobody wrote down is a number nobody can be shown to have violated, which is why it always is. - **Measure it.** Carry the source's change timestamp onto the copied row and export the difference between it and the time the row was written. That is propagation lag, and it is what pages someone — not queue depth, which falls happily to zero when the consumer has died. - **Reconcile it.** Run a sweep that re-derives a sample of principals' permitted sets from the authoritative source and counts differences. Divergence count, not lag, is what catches a dropped change, a bug in the projection, or a backfill that missed a partition — the failures where nothing is pending because nothing knows it is missing. - **Bound the blast radius.** Decide in advance what happens when propagation stops. Once the lag budget is exceeded, the listing can fall back to the authoritative predicate — slower, or narrower — rather than quietly serving stale visibility for an hour. And keep the copy honest about its job: **it decides which matters appear in a list.** Opening one is still decided against the authoritative source. That is a rule deliberately enforced in two places, and the reason to accept the duplication is exactly that the cheap copy is allowed to be a few seconds wrong while the expensive check is not. ## Grants and removals are not the same risk This is the most useful thing to say about the design, and it is routinely missed: | direction | what late propagation means | cost | |---|---|---| | a grant arrives late | a lawyer cannot yet see a matter they are entitled to | a support ticket | | a removal arrives late | a lawyer still sees a matter a wall now forbids | the incident the wall exists to prevent | So do not give the two directions one pipeline and one budget. The usual shape: **apply removals on the write path** — the request that raises the wall deletes the affected rows from the copy, or invalidates that principal's slice of it, before it returns — and let additions arrive on the asynchronous path, where lateness is merely annoying. Where a synchronous delete is impossible, an adequate approximation is a per-principal version marker bumped on any wall change, with listings for a principal whose marker moved in the last few minutes falling back to the authoritative predicate until the projection catches up. ## Migrating onto it while the old path is live 1. Build and backfill the table, and keep it maintained, while nothing reads it. 2. **Shadow it.** For a sampled share of listing requests, compute the page both ways and record every divergence with enough context to diagnose it. Run this longer than the slowest propagation path in the system — a week, not an afternoon. 3. Cut reads over when divergences stay at zero for a sustained window, keeping the old path behind a switch you can flip without a deploy. 4. Remove the old path only after the sweep has survived a full cycle of what breaks projections: a schema change, a bulk import, a failover. ## How you find out you were wrong You do not, unless you built the detector on day one. A stale visibility row produces no error, no failed request and no complaint — **the lawyer who should not have seen the matter is not going to report it**, and the one who cannot see a matter they should will report it within the hour, which is why the harmless direction is the one that gets fixed. The reconciliation sweep IS the detector. Its output belongs on a dashboard someone actually reads, with a threshold that alerts, next to the lag metric. A team that ships the copy without the sweep has not made a trade-off; it has made an unmeasured bet on its own change plumbing.
- The change stream that maintains the copy stops for forty minutes. What does the listing do?Whatever you decided beforehand. The lag metric crosses the budget, and the listing either falls back to the authoritative predicate — slower pages, or a narrower result — or returns an explicit degraded response. What it must not do is keep serving stale visibility silently, because the failure is invisible from the outside: pages stay full and fast while a wall that was raised half an hour ago is still not being honoured.
- A matter is visible to an entire practice group of 400 lawyers. Does the copy really need 400 rows?No, and forcing it to is how the table explodes. Copy the indirection instead: rows of group-and-matter, joined against the principal's group memberships in the listing predicate. The cost is one more join and a second thing that can go stale — membership changes now also need propagation. Fan out to per-principal rows only where the membership side is effectively static, or where the extra join measurably costs more than the storage.
The copy is the printed pull-list at the file-room counter: instant to check, and wrong from the second a wall goes up until somebody reprints it. That is tolerable for deciding what is on the shelf you are shown, and not tolerable as the last word on whether you may open a folder.
saying these in an interview costs you the question
- Once the copy is built it is the source of truth
- The lag is only a few seconds, so it needs no budget
- One pipeline for grants and removals is fine; they are both updates
- Queue depth tells you whether the copy is up to date
- A nightly rebuild catches any drift that matters
- Cut over to the copy and delete the old path the same day