skip to content

You run one in-memory tier shared by a dozen teams; what bounds one tenant's mistake, and when do you split it?

level: principalimportance: should knowfreq 34%

answer

  1. accounting, quota, separation
  2. only the top rung is isolation
  3. the server enforces nothing per owner
  4. split on incomparable worst cases

basics

~20 s

Accounting names an owner and bounds nothing. A quota the platform builds on the path in bounds size and rate. A separate instance is the isolation that holds - split when one tenant's worst case may not become a neighbour's.

solid answer

~50 s

There are three rungs and they are not alternatives - each is built on the one below. Accounting is a per-owner rollup of usage; it changes behaviour and makes an incident nameable, but it bounds nothing. A quota is enforcement the platform builds, because stores of this class generally offer none: admission by owner in a shared client library or a proxy every tenant must pass through, refusing one tenant's writes past its allowance while the others keep serving. It costs a new hop or a bypassable library, and it must be expressed in something measurable per owner - approximate bytes, entry count, write rate. Separation is a separate instance per tenant or per criticality class, and it is the only rung that actually bounds the blast radius. It costs utilisation, because every instance needs its own headroom, and multiplies what you operate and upgrade.

go deeper

for a junior

Carry away the ordering: naming who owns what is reporting, a limit someone enforces on the way in is a bound, and a separate instance is the only true separation.

for a middle

Explain why the middle rung has to be built by the platform - the server has no owner concept - and where it can be enforced so that every caller passes through it.

for a senior

Argue a concrete case: which tenant you would move off the shared instance first, on what evidence, and what the move costs in memory and operational surface.

for a principal

Own the contract across teams and time: what may share, what a mistake is allowed to cost, who changes a single-valued setting for everyone, and how a tenant exits the shared tier before an incident forces it.

## The three rungs The question sounds like a choice and is really a ladder: each rung needs the one below it, and only the top one is isolation. | Rung | What it bounds | What it costs | What it misses | |---|---|---|---| | **Accounting** - per-owner rollup of usage | Nothing | A periodic sampling pass off the serving path, and a convention everyone keeps | A tenant that is slow rather than large; anything happening right now | | **Quota** - admission by owner on the path in | Size and rate, per owner, where the path is enforced | A proxy hop and a new failure domain, or a bypassable library | Anyone who connects directly; damage from a single legal operation | | **Separation** - a separate instance per tenant | The blast radius itself: ceiling, capacity and keyspace are the tenant's own | Utilisation, since each instance needs its own headroom; more to operate and upgrade | Nothing about isolation; it is the honest answer, at a price | ## Accounting: necessary, and not a bound Rolled-up usage per owner turns anonymous growth into someone's growth. That is worth real money in behaviour change and in incident response, and it is the precondition for everything above it: **a quota you cannot measure per owner cannot be applied**. But it is a report. It arrives after the fact, it says nothing about latency imposed on neighbours, and it depends entirely on a naming convention that one team can quietly opt out of. Anybody who offers accounting as the answer to 'what stops team A from hurting team B' has answered a different question. ## Quotas: enforcement the platform has to build Most stores of this class enforce nothing per owner, so the quota has to live on the path in - in a shared client library, or in a proxy every tenant must traverse. Design decisions that matter: - **Pick a measurable dimension.** Approximate bytes per owner, entry count per owner, or write rate per owner. Each catches a different failure, and bytes is the one that is hardest to measure cheaply. - **Decide what exceeding it does.** Refusing that owner's writes while its neighbours keep serving is the whole point; anything softer is a dashboard. - **Accept the bypass.** A library is advisory - anyone can open a connection directly - so a library quota is a guard against accident, not against intent. A proxy is enforceable and is a new hop, a new thing to scale, and a new way for the tier to be down. - **Know what it does not catch.** A tenant within its size and rate allowance can still issue one operation expensive enough to hurt everyone. Quotas bound volume, not cost per call. ## Separation: the rung that actually holds A separate instance gives a tenant its own ceiling, its own capacity, its own keyspace and its own restart schedule. It is the only arrangement in which the honest answer to 'what can my neighbour do to me' is 'nothing', and the answer is not free: each instance carries its own headroom so total utilisation falls, and the number of things to monitor, upgrade and fail over multiplies. Triggers that should move a tenant onto its own instance: - Its entries **exist nowhere else**, while its co-tenants hold replaceable copies. The worst cases are not comparable, so the shared posture cannot be right for both. - Its **latency requirement is tighter** than the worst operation any co-tenant is permitted to run. - Its **growth is unpredictable**, so it will repeatedly consume the shared ceiling as a surprise. - It needs a **different posture** - a different behaviour at the ceiling, a different persistence choice, a different upgrade cadence - and those are single-valued per instance. - Its **cleanups or bulk loads are routine**, which means the procedure will eventually be skipped by someone in a hurry. ## The contract to write down 1. **Who may share at all**, expressed by what the data is rather than by team: replaceable copies may share; state that exists nowhere else gets its own instance. 2. **What one tenant's mistake is allowed to cost**, stated as a number the co-tenants can plan against rather than as an aspiration. 3. **Who changes the instance's settings** - the ceiling, the ceiling posture, the persistence posture, the restart and upgrade schedule - and how the other occupants are told. 4. **Which signals page the platform and which page the owner**: instance-level health is the platform's, per-owner usage against quota is the owner's. 5. **What the exit looks like**, so a tenant that outgrows the shared tier is moved on a plan instead of after an incident. ## What varies, and what to verify A few stores in this class, and several managed offerings, do provide per-owner containers or quotas. Where one exists, establish exactly what it bounds before relying on it: some bound memory and request capacity, others only partition the namespace, which is accounting wearing an enforcement badge. Verify it the same way you would verify any other claim on this tier - by making a tenant exceed it in a test instance and watching whether the neighbour notices.

  • Where do you enforce a quota when the server enforces none?
    In something every tenant must pass through: a shared client library, or a proxy in front of the instance. A library is cheap and bypassable by anyone who opens a connection directly; a proxy is enforceable but is a new hop, a new failure domain and a new thing to scale. Either way the quota must be expressed in something you can measure per owner - approximate bytes, entry count, or write rate.
  • One tenant's entries exist nowhere else while the others hold replaceable copies - what follows for the shared tier?
    That tenant is the first candidate to be split out. Everything the shared tier does routinely - removing entries under pressure, coming back empty after a restart, being cleared by a neighbour's mistake - is survivable for a replaceable copy and an outage for data with no other home. Sharing is only honest between occupants whose worst cases are comparable.
  • What does splitting cost, beyond the instances themselves?
    Utilisation, mostly. Each instance needs its own headroom above its working set, so the same total data needs more memory spread across more processes than it did in one. It also multiplies upgrades, failovers, connection pools and dashboards, and it moves per-tenant capacity planning from one estimate to a dozen.

saying these in an interview costs you the question

  • Offers a per-team naming convention as the isolation answer
  • Assumes the store can enforce a per-owner memory quota itself
  • Believes a client-library quota binds a tenant that connects directly
  • Thinks a size quota also bounds the latency one tenant imposes
  • Treats splitting as free rather than as utilisation traded for isolation
  • Puts data that exists nowhere else on the same instance as replaceable copies