skip to content

A platform team runs an active-active geode deployment across four regions on top of a globally-distributed database that supports writes originating from any region. They need to ship a backward-incompatible schema change to a core, heavily-used table. Walk through why this is riskier than the same change in a single-region deployment, and describe a rollout strategy that avoids taking any geode offline or serving inconsistent responses during the migration.

level: principalimportance: should knowfreq 30%

answer

  1. expand → migrate → contract
  2. additive-first schema change tolerates old code
  3. new code must read old AND new shape during migrate stage
  4. backfill historical rows separately
  5. contract only after all regions confirmed on migrate-stage code

basics

~20 s

In one region, you flip a switch and everyone's on the new version at once. Across four regions, you roll out gradually, so for a while some regions run old code and some run new - meaning the database has to work for both at the same time, or things break for whichever region hasn't updated yet.

solid answer

~50 s

The core risk is the extended window where different geodes run different code versions against the same replicated data - a rollout across four regions inherently can't be atomic, so old code and new code have to coexist correctly for however long the staged rollout takes. The standard strategy is expand/migrate/contract: first ship a schema change that's additive and backward-compatible (new column/table nullable or defaulted, old shape untouched), roll that out everywhere; then ship application code that writes to the new shape while still being able to read the old shape, roll that out region by region, backfilling old data as needed; only once every region is confirmed on code that no longer depends on the old shape do you contract by removing it. Each stage is independently safe for any mix of regions being on old vs. new code.

go deeper

for a junior

Should recognize that rolling a change out to four regions one at a time means there's a period where different regions are running different versions, and that this needs some care.

for a middle

Should be able to name the general idea of making schema changes additive first and cleaning up later, even without using the specific expand/migrate/contract terminology.

for a senior

Should be able to lay out the expand/migrate/contract stages accurately, including that new code must tolerate old-shape data (and old code must be untouched by the additive change) during the migrate stage.

for a principal

Should connect the specific schema-migration mechanics to the general organizational discipline every staged multi-region rollout needs, articulate why replication makes cross-region version mismatches unavoidable during rollout, and reference how real globally-distributed database platforms push this compatibility responsibility onto the application layer.

## Why the geode deployment amplifies the risk A backward-incompatible schema change is risky in any system, but a geode deployment amplifies the risk in a specific way: the rollout can't be atomic across four regions, so there's an extended window - realistically minutes to hours, depending on how cautious the rollout is - during which some geodes are running the old application code and some are running the new code, all against the same underlying replicated dataset. - **In a single-region deployment**, that window is much shorter, a rolling deploy across instances within one region, typically seconds to low minutes, and more importantly there's only one copy of the data being read and written the whole time, so 'old code, new schema' is the only mismatch to worry about. - **In the geode case**, you additionally have 'new code in region A writing data that old code in region B then has to read via replication,' and vice versa, multiplied across however many regions are mid-transition at any given moment. If the schema change is truly backward-incompatible - say, renaming a column, changing its type, or making it required where it used to be optional - then whichever code version encounters data written by the other version will either error out or silently misbehave, and because this is active-active traffic serving real users the whole time, that failure is customer-visible, not caught in a maintenance window. ## The three-stage mitigation The standard mitigation is the **expand/migrate/contract** pattern, which breaks one risky atomic-feeling change into three sequential steps that are each safe to have coexist with either the old or new application code at any point. 1. **Expand**: change the schema in a strictly additive way - add the new column as nullable, or with a safe default, add the new table, without touching or removing anything the old code depends on. This migration is rolled out to all four regions' databases. Because it's purely additive, old application code, which doesn't know the new column exists, keeps working exactly as before, and this stage carries essentially no risk regardless of which regions have received it yet. 2. **Migrate**: deploy new application code that writes to the new shape, populating the new column on every write, while still being able to correctly read both the old shape and the new shape - i.e., the new code tolerates rows that only have the old data, treating a missing new-column value as 'not yet migrated' rather than an error. This code is rolled out region by region, in the team's normal staged rollout order. During this stage, some regions run old code, writing only the old shape which is still fully valid, and some run new code, writing the new shape and populating both, and because of replication, a region running new code will receive replicated rows written by a region still on old code, which is fine, because new code already knows how to handle rows missing the new column. Separately, a backfill job walks existing rows and populates the new column for historical data so that eventually every row has both shapes populated, not just newly-written ones. 3. **Contract**: only once every region has confirmed it's running the migrate-stage code, and the backfill is complete, does the team ship a final code change that stops reading/writing the old shape at all, followed by a schema change that actually removes the old column. Because this stage only starts once no region depends on the old shape anymore, it's safe to remove. ## Why the sequencing is the point The reason this specific sequencing matters is that at every single point in the rollout - no matter which subset of the four regions has received which stage - the schema and the code running against it are mutually compatible, so there's no window where a request can hit a code/data mismatch, regardless of replication timing or how long the staged rollout takes. ## Where it shows up A concrete real-world flavor of this: large-scale multi-region platforms built on globally-distributed databases that support multi-region writes, the kind of database this scenario implies, explicitly document this exact concern, because the database itself replicates asynchronously across regions, application teams are expected to design schema changes to be forward- and backward-compatible during rollout, precisely because 'all regions upgrade instantly and atomically' isn't something the platform can offer without sacrificing the availability the whole geode architecture exists to provide. ## The lesson generalizes The organizational lesson generalizes past just schema changes: in an active-active geode system, any change - config, feature flags, API contracts - has to be designed the same expand/migrate/contract way, because the team is permanently living in a world where different production regions are, briefly but constantly, running different versions of the truth.

  • What would go wrong if the team skipped the 'expand' step and went straight to deploying new code expecting the new column to already exist?
    The new code would immediately error or fail to write correctly in every region until the schema change has actually been applied there, and since regions upgrade code on a staggered schedule while presumably also needing the schema change applied in some order, you'd almost certainly hit a region where new code is live but the column doesn't exist yet, causing failed writes or crashes for real user traffic in that region.
  • Why is a backfill job for historical rows a separate step from the schema and code changes?
    The schema and code changes only guarantee that new writes populate the new column correctly; existing rows written before the migration started still only have the old shape, and if anything later depends on every row having the new field populated, not just new ones, that gap has to be closed explicitly by a backfill process that runs independently of and typically after the code rollout, since it's a data-volume-dependent, potentially long-running job rather than an instantaneous deploy step.
  • How does this expand/migrate/contract discipline generalize beyond schema changes in a geode deployment?
    The same three-stage thinking applies to API contract changes, feature flags, and configuration changes - anything rolled out in stages across regions needs a period where both the old and new behavior can coexist safely, before finally removing the old behavior once every region has confirmed the new one is fully in place, because a geode deployment never has a single instant where all regions are guaranteed to be on the same version.

It's like renovating four identical bridges that all carry live traffic and can never fully close - you can't rip out the old lane markings and repaint new ones in one afternoon per bridge; instead you paint the new markings alongside the old ones first, run both simultaneously while drivers adjust, and only remove the old markings once every bridge and every driver has switched over.

saying these in an interview costs you the question

  • Proposes deploying the schema change and new code to all regions simultaneously as a single atomic step
  • Doesn't mention that new code must still be able to read data written by old code (and vice versa) during rollout
  • Skips backfilling historical rows, assuming the migration only needs to handle newly-written data
  • Removes the old column in the same step that adds the new one
  • Doesn't recognize that replication means one region's new code will receive rows written by another region's old code mid-rollout

context