You measure a service's throughput at increasing concurrency levels and find it rises to a peak around 24 in-flight workers, then declines as you push higher. How would you use scalability modelling to decide between buying more hardware, partitioning the system, or redesigning the hot path?
answer
- retrograde curve = adding capacity makes it worse
- cap concurrency at the peak first (hours)
- fit contention vs coherence to choose the fix
- partition/stop sharing changes curve shape
- Little's law L = lambda*W sanity-checks the axis
basics
~20 sA declining curve means coordination cost is growing faster than capacity, so more hardware will make it worse. Fit contention and coherence coefficients to the sweep: contention-dominated means split the bottleneck; coherence-dominated means partition or stop sharing. Meanwhile cap concurrency at the measured peak.
solid answer
~60 sA *retrograde* curve rules out the hardware option immediately: if throughput falls from 24 to 32 workers, more workers on more machines extends the decline. Adding capacity only helps when the curve is still climbing. My sequence: 1. **Cap first.** Enforce a concurrency limit at or just below the measured peak — admission control, a bounded queue, load shedding. This stops the bleeding in hours, before any redesign. 2. **Fit a model** to the throughput-versus-concurrency sweep to separate the two loss terms: contention (serialization on a shared resource, linear cost) from coherence (keeping shared state consistent, quadratic cost). Only the second produces a decline. 3. **Choose by coefficient.** Contention-dominated: split the bottleneck — finer locks, more connections, sharded queues, remove the single coordinator. Coherence-dominated: reduce sharing itself — partition state by key so one owner writes it, batch updates, per-worker accumulators, independent replicas. 4. **Validate** by re-running the sweep; success is a higher peak, not a better single data point. Cross-check with Little's law that measured in-flight concurrency matches rate times latency, so the axis means what you think.
code
text · 9 linescurve shape dominant cost right move
------------------- ---------------- --------------------------
still rising none yet add capacity
flat plateau contention split the shared resource
rises then falls coherence partition / stop sharing
falls from the start retries or fix amplification /
throttling environment first
Always, immediately: cap concurrency at the measured peak.go deeper
Know that a falling throughput curve means more workers hurt, and that the first response is to limit concurrency rather than add capacity.
Distinguish the plateau case from the declining case and match each to its remedy: split the shared resource for a plateau, reduce sharing for a decline.
Validate the measurement, fit the model to separate contention from coherence, cap concurrency immediately, then sequence the fixes by time-to-value and re-measure the whole sweep.
Turn the fitted coefficients into an investment decision with a predicted post-fix peak, argue the hardware option down on evidence rather than principle, and treat partitioning as the only lever that changes the curve's shape rather than its height.
## Read the curve before spending anything A throughput-versus-concurrency curve has three regimes, and the regime dictates the remedy: - **Rising, near-linear**: you are under-provisioned. Adding capacity works. Buy hardware. - **Flat plateau**: something serializes. Extra workers queue behind a shared resource. Hardware buys nothing; splitting the resource does. - **Rising then falling (retrograde)**: coordination cost is growing faster than capacity. Extra workers make throughput *worse*. A peak at 24 followed by decline is the third case, and that single observation eliminates the most commonly proposed fix. Buying machines when the curve is retrograde funds the decline. This is the judgment the question is really testing: recognizing that the shape of the curve, not the size of the incident, chooses the strategy. ## Make the measurement trustworthy first Before acting, verify the axis: - Sweep concurrency, not offered load, and measure at steady state per level with warm caches and pools. - Keep the workload mix constant. A different read/write ratio is a different system with different coefficients. - Cross-check with **Little's law**: `L = lambda x W` — average in-flight requests equals arrival rate times average residence time. If your client says 24 in flight but rate x latency says 60, you are measuring queueing somewhere you did not intend, and the whole curve is suspect. - Confirm the decline is not an artifact: timeouts and retries can manufacture a fake retrograde by amplifying load, and thermal throttling or a noisy neighbour can fake one too. Retries especially create a feedback loop that mimics coherence collapse; rule them out before concluding anything about architecture. ## Separate the two loss terms Fit the sweep to a model with both penalty terms — contention, whose cost grows linearly with worker count, and coherence, whose cost grows quadratically because keeping N workers' view of shared state consistent is a pairwise problem. The relative size of the two coefficients is the diagnosis: **Contention-dominated (plateau).** One resource serializes everything: a global lock, one connection pool, a single-threaded coordinator, one hot partition, one disk. Remedies split the resource: finer-grained locking, per-shard queues, more connections, removing the coordinator from the request path, replacing a global counter with striped counters. These are usually contained changes with predictable payoff. **Coherence-dominated (retrograde).** Workers are reconciling shared mutable state: cache-line ping-pong on a hot object, cross-node invalidation, distributed locking, consensus per request, chatty gossip. Remedies must reduce *sharing itself*: - **Partition by key** so each item has exactly one owner and cross-worker coordination disappears rather than getting faster. - **Batch or amortize** coordination — one reconciliation per hundred operations instead of per operation. - **Per-worker accumulators** merged once at the end, converting continuous crosstalk into a single reduction. - **Relax consistency** where the domain tolerates it — eventual consistency, conflict-free merge, read-your-writes only where required. - **Independent replicas** with no shared state, so capacity scales by replication instead of by concurrency. Note the asymmetry that should drive the decision: contention fixes make the plateau higher, while coherence fixes change the *shape* of the curve. Because the peak concurrency scales with the inverse square root of the coherence coefficient, halving that coefficient buys roughly 40% more peak concurrency — and driving it to zero removes the decline entirely, leaving only a plateau. Structural sharing removal has a qualitatively different payoff from tuning. ## Sequence the work by time-to-value 1. **Hours — cap concurrency.** Enforce an admission limit at or slightly below the measured peak. Bounded queues, a semaphore at the entry point, load shedding with a clear rejection response. This converts a collapsing system into a merely saturated one, and a saturated system with correct backpressure degrades predictably instead of falling over. Do this even while you plan the real fix. 2. **Days — attack contention.** Usually localized, low-risk, measurable. Often raises the peak enough to buy planning time. 3. **Weeks to quarters — remove sharing.** Partitioning is an architectural change with data-migration, routing, and rebalancing consequences. Justify it with the fitted coefficient, not with intuition, and state the expected new peak so the outcome is falsifiable. ## Frame the decision economically Each option has a cost and an expected effect on the curve: - **More hardware**: cheap to try, zero or negative effect on a retrograde curve. Rejected on evidence, not on principle — and worth saying explicitly, because "we measured it and it makes things worse" is a far stronger argument to stakeholders than "scaling out won't help." - **Splitting the bottleneck**: moderate effort, raises the plateau, does not remove the decline if coherence is present. - **Partitioning / removing sharing**: high effort, changes the shape of the curve, and is the only thing that lifts a hard ceiling. - **Reducing per-request work**: always helps, orthogonal to all of the above, and sometimes moves the deadline far enough that no restructuring is needed this year. The strongest answer names the measurement that would change the recommendation, commits to a predicted post-fix peak, and re-runs the same sweep to confirm. Success is a demonstrably higher peak concurrency and a flatter decline beyond it — not a single faster benchmark run.
- Stakeholders want to double the instance count this week. How do you argue against it?Show the measured curve: throughput at 32 workers is below throughput at 24, so more workers extend the decline rather than reverse it. Offer the alternative that helps this week — cap concurrency at the peak and shed excess load, which raises effective throughput immediately and costs nothing — and commit to a fitted model plus a predicted new peak for the structural fix, so the larger spend is judged on evidence.
- How would you know whether the decline is real architecture or a measurement artifact?Rule out amplification and environment first: client retries and timeouts can manufacture a retrograde curve by multiplying offered load exactly when latency rises; thermal throttling, frequency scaling, and noisy neighbours can too. Re-run with retries disabled, verify with Little's law that in-flight concurrency equals rate times residence time, and confirm the peak reproduces across runs and machines before drawing an architectural conclusion.
A kitchen where every cook must confirm each plate with every other: hiring more cooks past a point slows dinner. You either give each cook their own station (partition) or you keep the crew at the size that peaks.
saying these in an interview costs you the question
- Proposing to add nodes or workers when the measured curve is already declining
- Skipping the immediate concurrency cap and going straight to a multi-quarter re-architecture
- Assuming the decline proves a coherence problem without ruling out retry amplification or throttling
- Treating the concurrency axis as trustworthy without a Little's law cross-check
- Declaring success on one faster benchmark run instead of re-running the full sweep to show a higher peak