How do you migrate airframe sign-off from a role column to relation tuples while both stay authoritative, and what tells you to cut over?
answer
- behaviour migration, not data
- both stores written, one decides
- compare on live traffic
- two directions, different bugs
- a full business cycle, not a week
basics
~20 sWrite the schema first, dual-write every grant, backfill history, then shadow-read the tuple check on live traffic while the column still decides. Cut over per relation only after disagreements reach zero across a full business cycle, treating over-grants as security findings.
solid answer
~50 sTreat it as a behaviour migration, not a data migration. First prove the schema can express every rule the column implies, including the ones that live in code rather than in the table. Then dual-write: every grant and revocation writes the column and the tuples, with the column still authoritative, and backfill existing state as tuples. Then **shadow-read** — call the tuple check on every real request, compare it with the column's decision, log every disagreement with subject, object and relation, and enforce neither. Two disagreement directions mean different things: tuples say yes where the column said no is an over-grant and a security finding, usually a rewrite that is too broad; tuples say no where the column said yes is a missing backfill. Cut over one relation or one endpoint at a time once the counter has been flat at zero across a full business cycle, not a week.
code
pseudocode · 13 linesdecideSignoff(user, component):
legacy = user.roleColumn in ("inspector", "chief_inspector")
and component.airframe.orgId == user.orgId // authoritative
shadow = tupleStore.check(component, "signoff", user) // enforced nowhere
if shadow != legacy:
record("authz-shadow-disagreement",
user, component, "signoff",
legacy, shadow,
direction: shadow && !legacy ? "OVER_GRANT" : "UNDER_GRANT")
return legacy // the column decides until OVER_GRANT has been flat at zerogo deeper
The key idea is that the new system runs alongside the old one and decides nothing for a while. Its answers are recorded and compared, not enforced.
Be able to lay out the phases in order — schema, dual-write, backfill, shadow-read, narrow cutover — and explain why the column keeps deciding until the comparison is clean.
Bring the operational detail: write ordering and partial-failure behaviour in dual-write, the two disagreement directions and their different urgency, and why denials during shadow must be separable from store failures.
Own the parts nobody assigns: how long a full business cycle really is for this domain, that cutover buys a network dependency on every read path, and that the permanent sampler is what detects a quietly widened rule eighteen months later.
## Why this is not a data migration A role column answers *what is this person*. Relation tuples answer *what path exists between this person and this object*. Those are different questions, and a script that turns each role row into a tuple produces a store that agrees with the column on the cases the column could already express and is silent on everything else — which is, generally, the reason you are moving. So the migration has two halves that must not be confused: **reproducing today's behaviour exactly**, which is the risky part and the one you must prove, and **expressing the rules the column could never hold**, which is new design and should be built only after the first half is proven. Doing both at once means every disagreement is ambiguous: you cannot tell a migration bug from an intended improvement. ## The sequence 1. **Write the schema and prove expressiveness on paper.** Enumerate every rule the current system enforces, including the ones that are not in the table at all — the chief-inspector branch in a service method, the organisation check on the endpoint, the special case for a component under quarantine. Each becomes a relation, a rewrite, or an explicit decision to stop supporting it. 2. **Dual-write.** Every domain event that grants or removes authority writes both stores. Decide the order and the partial-failure behaviour deliberately: write the authoritative store first for a grant so a crash under-grants, and write the *non*-authoritative store first for a revocation so a crash leaves the subject revoked somewhere rather than nowhere. Reconcile asynchronously. 3. **Backfill.** Convert existing state to tuples in one pass, recording the version the backfill reached so you can tell later disagreements from backfill gaps. 4. **Shadow-read.** On every real request, run the tuple check alongside the column decision, return the column's answer, and record disagreements with subject, object, relation and both verdicts. This is the only step that actually produces evidence. 5. **Cut over narrowly.** One relation, or one endpoint, at a time, with the column still written and readable. 6. **Stop writing the column,** then remove it. ## Reading the disagreement log | direction | what it means | how to treat it | |---|---|---| | tuples allow, column denied | the tuple schema grants more widely than today | a security finding, even in shadow mode — you are about to enforce it | | tuples deny, column allowed | a fact was never backfilled, or a rule is missing from the schema | a correctness bug, and the common one early on | | both allow or both deny | agreement | the only thing that may be counted as progress | The first row is the one teams under-react to, because in shadow mode nothing bad happens. But the over-grant direction is exactly what you will ship on cutover day, and a single over-broad rewrite can grant a whole organisation sign-off authority over airframes it does not hold. ## How long to shadow Long enough to see the rules that fire rarely. A maintenance organisation's year contains annual inspections, airworthiness reviews, custody transfers and audits; a rule that only fires during an annual check will not appear in a week of traffic, and the first time it appears will be in production with the tuple store authoritative. A full business cycle is the honest answer, and where that is genuinely too long, the compromise is to enumerate the rare paths and exercise them deliberately rather than to wait less. ## What cutover actually costs On the day you stop consulting the column, three things change at once: - authorization acquires a **network dependency** on the request path, with its own tail latency and its own outages; - a wrong answer is now produced by a **schema rule** rather than by a row someone can read, so the skill needed to debug a denial changes; - **revocation semantics** change, because the store serves reads at a chosen freshness rather than from the same transaction as your application data. Those are the reasons to cut over by relation rather than in one step, and to keep the column readable for a while afterwards. ## What you keep permanently Keep a low-rate sampled comparison running after cutover, against whatever remains of the old rules or against a hand-maintained expectation set. The failure mode this migration has eighteen months later is not a crash — it is a rewrite added for one customer that quietly widened a relation, and the only thing that catches it is something that keeps checking the answer against an independent expectation. That sampler is the cheapest insurance in the whole project, and it is the first thing teams delete.
- Why classify disagreements by direction rather than just counting them?Because the two directions are different bugs with different urgency. An under-grant is a missing fact and shows up as a user complaint on cutover day; an over-grant is a rule that is too broad and shows up as nothing at all. Counting them together lets a falling total hide a rising over-grant rate.
- In dual-write, which store do you write first for a revocation?The non-authoritative one first, so a crash between the two writes leaves the subject revoked in at least one place rather than in neither. Grants take the opposite order for the same reason: a partial failure should under-grant. Then reconcile, because the partial state is real and will happen.
- Can you skip the shadow phase if the tuple backfill was generated from the column itself?No, and that generation is exactly why you cannot. A backfill derived from the column reproduces what the column said, not what the system enforced — the rules living in service code never appear in it, and shadow traffic is the only thing that exercises them.
- What is the strongest reason to cut over relation by relation rather than all at once?Blast radius and attribution. If one relation's rewrite is wrong, only the endpoints using that relation are affected, and the change that caused it is unambiguous. A single cutover makes every subsequent incident a question of which of forty rules was responsible.
saying these in an interview costs you the question
- Calls it a data migration and plans a one-pass conversion script
- Enforces the tuple decision during the comparison phase
- Counts disagreements as one number without a direction
- Assumes rules living in service code are captured by a backfill from the table
- Shadows for a week and treats a quiet period as proof
- Deletes the comparison sampler immediately after cutover