skip to content

In a deployment-stamp architecture, how does a platform decide which stamp a new tenant is assigned to, and what happens operationally when a stamp approaches its capacity limit?

level: middleimportance: should knowfreq 45%

answer

  1. control plane / tenant directory owns placement
  2. placement policy: most-free-capacity, round robin, dedicated stamp
  3. migration = replicate data + cutover + update directory
  4. proactive rebalancing vs emergency migration
  5. directory is on critical path, holds no tenant data itself

basics

~20 s

A control service picks a stamp for the new tenant — usually whichever stamp has room left, or a dedicated one if required. When a stamp fills up, the platform either stops assigning new tenants to it, provisions a brand-new stamp, or migrates some existing tenants off it to free space.

solid answer

~50 s

Tenant assignment is handled by a lightweight, globally-shared control plane or tenant directory that tracks each stamp's current load/headroom and applies a placement rule — commonly 'first stamp with sufficient free capacity,' round-robin, or a dedicated stamp for tenants needing isolation (compliance, size, or contractual reasons). When a stamp nears its capacity ceiling, options are: stop routing new tenants to it and provision a fresh stamp for future growth; proactively rebalance by migrating some existing tenants to a less-loaded stamp; or vertically grow that stamp's resources if there's still headroom in the underlying infrastructure. Tenant migration between stamps is the hard part — it usually means replicating that tenant's data to the new stamp's database, cutting traffic over with minimal downtime (often via a maintenance window or dual-write/backfill approach), and updating the directory. Getting placement wrong early (e.g., cramming too many large tenants onto one stamp) creates expensive rebalancing work later.

go deeper

for a junior

Should know that some component decides which stamp a new tenant goes to, and that stamps can eventually run out of room.

for a middle

Should describe a placement policy (capacity-based or dedicated-stamp) and name at least one option for handling a full stamp (new stamp, vertical growth, or migration).

for a senior

Should walk through the mechanics of tenant migration (data replication + cutover) and the trade-off between maintenance-window and dual-write approaches.

for a principal

Should discuss proactive rebalancing strategy, the risk of the tenant directory as a critical-path dependency, and how to avoid the 'large tenants clustering on one stamp' failure mode at scale.

## Who owns placement Tenant placement is the piece of glue logic that makes a fleet of otherwise-independent stamps behave like one coherent product. It's owned by what's usually called a **control plane**, **tenant directory**, or **fleet manager** — a small, separately-scaled service (or set of services) that is shared across the whole system, unlike the stamps themselves. Its core job is to maintain a mapping from tenant identity to stamp, and to make the placement decision at signup time: when a new tenant is created, the control plane consults its knowledge of each stamp's current load (number of tenants, database size, request volume, or whatever metric the team uses as a capacity proxy) and applies a placement policy. ## The placement policies Common policies are: - 'Pick the stamp with the most free capacity.' - Round-robin across stamps below a load threshold. - For tenants with special requirements, route straight to a dedicated, isolated stamp reserved for that purpose — a large enterprise customer with contractual isolation requirements, or a customer subject to a specific data-residency rule. This directory then has to be consulted by the routing layer (API gateway, DNS-based routing, or an ingress rule) on every incoming request, so it needs to be fast and highly available — it sits on the critical path for every tenant even though it holds no tenant application data itself. ## Why placement matters so much The reason this placement logic matters so much is that a stamp is a **bounded-capacity unit**, and bad placement decisions compound over time. Stamps are typically sized with some target load in mind, and once a stamp's tenant population pushes it toward its database connection ceiling, storage limit, or latency SLO, the operational choices are limited and none of them are free. ## The three levers when a stamp fills up | Lever | What it does | |---|---| | Stop assigning new tenants | The simplest response is to stop assigning new tenants to that stamp and direct future growth to a freshly-provisioned stamp — cheap to execute, but it does nothing about the stamp that's already full and doesn't help if an existing tenant on that stamp keeps growing | | Vertical growth | The next lever is vertical: grow that stamp's database or app-tier resources, if the underlying infrastructure has headroom — often the first response because it requires no data movement, but it only buys time before hitting the next ceiling | | Tenant migration | The heavier lever is tenant migration: physically moving one or more existing tenants from an over-loaded stamp to a less-loaded one | Migration means replicating the tenant's data (schema, rows, blobs, whatever is tenant-scoped) into the destination stamp's database, then cutting traffic over — commonly done either via a **maintenance window** (simplest, but visible downtime for that tenant) or a **dual-write/backfill approach** where writes go to both stamps during a transition period and reads cut over once the backfill catches up (more complex to build correctly but avoids visible downtime). Either way, the directory has to be updated atomically with the cutover so requests start routing to the new stamp at the right moment, and any in-flight requests during the cutover window need careful handling to avoid split-brain writes landing on the wrong stamp. ## When placement goes wrong Getting placement wrong early is expensive precisely because migration is the expensive escape hatch. - **A common real-world failure mode** is placing several large or fast-growing tenants onto the same stamp early on (because that stamp happened to have room at signup time), only to discover months later that the stamp is now disproportionately loaded relative to its siblings, forcing an unplanned, high-risk migration under time pressure rather than a calm, proactive one. - **Another failure mode is a control plane outage or staleness bug**: if the tenant directory's view of stamp load gets stale (e.g., a metrics pipeline lag), new tenants can keep getting routed onto an already-saturated stamp, degrading service for everyone already there. Teams operating at scale generally build **proactive rebalancing** — periodically identifying stamps trending toward their limits and migrating a few tenants preemptively, during low-traffic windows, rather than waiting for a hard ceiling to force an emergency migration. This tenant-placement-and-rebalancing problem is one of the main reasons the deployment stamp pattern is described as trading raw infrastructure simplicity for operational sophistication: the stamps themselves are conceptually simple to replicate, but the fleet-level logic of who goes where, and how they move, is genuinely hard distributed-systems work.

  • What's the difference between vertically growing a stamp and migrating tenants off it?
    Vertical growth adds more resources (compute, storage, connections) to the existing stamp without moving any data — fast and low-risk, but bounded by the underlying infrastructure's own limits and by whatever the database engine can scale to on one instance. Tenant migration actually moves some tenants' data to a different, less-loaded stamp, which is more work and riskier (requires a cutover) but is the only lever that actually reduces load on the original stamp rather than just delaying the ceiling.
  • Why does the tenant directory need to be highly available even though it stores no application data?
    Because every incoming request needs to resolve which stamp to route to before it can even reach the tenant's actual data, so the directory sits on the critical path of every single request across the whole fleet. If it goes down or serves stale data, requests can be misrouted or fail outright even though every individual stamp is healthy.
  • What's a safer alternative to a hard maintenance-window cutover when migrating a tenant between stamps?
    A dual-write/backfill approach: start replicating the tenant's existing data into the destination stamp in the background, keep writing new changes to both the source and destination during the transition, and only cut reads over to the destination once the backfill has fully caught up — this avoids a visible downtime window but is materially more complex to implement correctly, especially around avoiding lost or conflicting writes during the overlap.

It's like a hotel chain's central reservation system deciding which branch hotel a new guest's booking goes to based on which hotels have vacancy — and when one hotel is fully booked out for months, the chain either stops booking new guests there, builds a new branch, or works to move some existing long-term guests to a less full hotel.

saying these in an interview costs you the question

  • No concept of a control plane/directory deciding placement
  • Thinks stamp capacity limits can always be solved by 'just add more servers' without mentioning the database
  • Assumes tenant migration between stamps is instant or free
  • Doesn't distinguish proactive rebalancing from emergency migration

context