skip to content

An admin console on a fractional-core tier ran fine for weeks, then crawled during a campaign — why?

level: seniorimportance: should knowfreq 44%

answer

  1. the tier is not just a small machine
  2. baseline share plus a credit balance
  3. idle time accrues, bursts spend
  4. accrual is capped, so idling has a ceiling
  5. the collapse lags the change in demand

basics

~20 s

A fractional-core machine accrues credit while it stays below its baseline share of a processor and spends that credit to burst above it. Weeks of idling built a balance; sustained campaign traffic drained it, and the machine dropped to its baseline share.

solid answer

~50 s

The tier is not simply a small machine — it is a machine with a **baseline share** of a processor plus a **credit balance**. Running below the baseline accrues credit up to a cap; running above it spends credit. For weeks the console used almost nothing and sat at the accrual cap, so every short burst was covered and the machine felt fast. The campaign turned a low, spiky load into a sustained one, the balance drained, and the platform clamped the machine to its baseline share — which is the performance you were paying for all along. The give-aways are processor use pinned flat at the baseline rather than varying with load, and the published credit-balance metric hitting zero at exactly the moment throughput flattened. The fix is either a tier that guarantees the sustained share, or a baseline chosen to match sustained demand.

code

pseudocode · 19 lines
pseudocode
baselineShare = 0.20            # guaranteed fraction of one core
creditCeiling = baselineShare * minutesInOneDay   # cap set by the provider

credit = creditCeiling          # weeks of idling: balance sits at the cap

for each minute:
    used = measuredCoreTimeThisMinute       # 0.0 .. 1.0 of one core

    if used <= baselineShare:
        # accrue the unused baseline, but never past the ceiling
        credit = min(creditCeiling, credit + (baselineShare - used))
    else:
        want = used - baselineShare
        if credit >= want:
            grant(used)                     # burst covered by the balance
            credit = credit - want
        else:
            grant(baselineShare)            # balance empty: clamped to baseline
            credit = 0

go deeper

for a junior

Know that a fractional-core machine guarantees only a small share of a processor and lends you more out of a credit balance built up while it was idle. When that balance empties, it drops to the share it guarantees.

for a middle

Explain the accounting: credit accrues below the baseline up to a cap, is spent above it, and the drop to baseline is the platform behaving exactly as specified rather than a fault.

for a senior

Diagnose it from evidence — the balance series hitting zero, processor use pinned flat at the published baseline, no deployment, self-recovery after a quiet period — and say why a short load test passed anyway.

for a principal

Decide the standard: which classes of workload may sit on this tier at all, what must be proved before one does, and how a team is stopped from qualifying it with a ten-minute test.

## What a fractional-core tier actually sells A fractional-core tier is usually read as "a cheap small machine", and that reading is what makes the failure surprising. What it actually sells is two things bundled together: - a **baseline share** of a processor, guaranteed and continuously available — some fraction of one core, published by the provider for each size in the tier; - a **credit balance** that lets the machine run above that baseline for a while. Credit accrues whenever the machine uses less than its baseline, up to a **cap** the provider sets, and is spent whenever it uses more. That cap matters: idling for a month does not build a month of burst. Once the balance reaches zero, the platform clamps the machine to the baseline share, and the baseline is the performance the price was always for. ## Why the collapse arrives late and looks like a code problem The symptom is delayed by exactly as much credit as had been saved, which is why it lands mid-campaign rather than at the start of it, and why nothing in the deployment history explains it. A low-traffic internal console is the classic victim: for weeks it uses a few percent of a core, sits at the accrual cap, and covers every human click out of credit. Change the load from spiky to sustained and the arithmetic reverses — the balance drains at the rate by which demand exceeds the baseline, then stops covering anything. The distinguishing evidence: - **Processor use goes flat, not high.** A clamped machine reads pinned at a constant low value — the baseline — rather than swinging up to saturation. A team reading that as "the machine is barely busy" concludes wrongly that the machine is not the problem. - **The credit-balance metric hits zero** at the moment throughput flattened. Platforms that sell this tier expose the accrued balance, precisely because the symptom is otherwise unreadable. - **Nothing shipped.** The same build was fine yesterday and is fine again the following morning, once demand has sat below baseline long enough to accrue again. ## Confirming it rather than guessing 1. Line the credit-balance series up against the latency or throughput series, and check that the flattening starts when the balance reaches zero. 2. Check whether processor use is pinned at a constant value equal to the published baseline for the size, rather than varying with load. 3. Confirm that demand changed shape — sustained rather than spiky — while the deployment history did not. 4. Verify recovery: after a quiet period the machine behaves normally again, which no code regression would do by itself. ## The trap in testing A short load test passes. Ten minutes of load is paid for out of an accrued balance, so the test measures the burst rather than the baseline and reports a capacity the machine cannot sustain. A test meant to qualify this tier has to either run long enough to exhaust the balance or start from a deliberately drained one. This is the most common way the failure reaches production after a green test. ## What to do about it | Option | What it changes | When it is right | |---|---|---| | Move to a family with a full, guaranteed share | sustained demand is always served | sustained demand genuinely exceeds the baseline | | Move up within the fractional tier | a larger baseline share and a larger accrual cap | demand grew but is still well under a whole core | | Reduce sustained demand | the workload fits the baseline again | the load is accidental — a chatty poller, an unindexed query | | Accept the baseline | nothing; the console is slow under campaign load | the surface genuinely is not worth a full share | Where providers differ: some offer a mode in which the machine may keep bursting past an empty balance for an extra, metered charge instead of being clamped, and both the accrual rates and the caps are set per provider and per size. The mechanism — a baseline plus a capped, accruing, spendable balance — is common to the tier wherever it is sold. ## When the tier is the right choice It is genuinely good at what it is built for: workloads whose **sustained** demand really does sit below the baseline, with short and rare peaks. Internal tools, development machines, schedulers, low-traffic administrative surfaces. The discipline is to size the baseline against sustained demand and treat the credit as covering spikes only — never to size against the burst and hope the spikes stay short.

  • How would you confirm this is credit exhaustion rather than a code regression?
    Line the credit-balance series up against throughput and check that the flattening begins when the balance reaches zero. Then check that processor use is pinned at a constant value equal to the published baseline rather than swinging with load, confirm nothing shipped, and confirm the machine recovers by itself after a quiet period. No code regression fixes itself when traffic falls.
  • The ten-minute load test passed. Why did it not catch this?
    Because the accrued balance paid for the whole test. Ten minutes of burst measures the burst, not the baseline, and reports a sustainable capacity the machine does not have. A test that qualifies this tier must run long enough to drain the balance, or start from a deliberately drained one.
  • When is a fractional-core tier the right choice rather than a trap?
    When sustained demand genuinely sits below the baseline share and peaks are short and rare — internal tools, development machines, schedulers, low-traffic administrative surfaces. Size the baseline against sustained demand and treat the credit as covering spikes only. The trap is sizing against the burst and assuming the spikes stay short.

A prepaid allowance on a phone plan: you can spend it as fast as you like, and when it runs out the connection still works — at the slow baseline rate you were actually paying for.

saying these in an interview costs you the question

  • Blames the application, since nothing in the code changed that week.
  • Sees processor use flat at a low value and calls the machine idle.
  • Assumes the credit balance refills while the load is still running.
  • Believes the tier is uniformly slow rather than baseline plus spendable credit.
  • Restarts the machine expecting the accrued balance to reset.
  • Accepts a ten-minute load test as proof the tier sustains that load.