skip to content

An EC2 t3 instance running a web service performs well for a while and then its CPU flatlines at a low fixed percentage while requests queue up. Explain the burstable T-instance CPU credit mechanism behind this, and what unlimited mode changes.

level: middleimportance: must knowfreq 62%

answer

  1. baseline plus a bank
  2. credits are vCPU-minutes
  3. balance capped near a day's accrual
  4. standard throttles, unlimited bills
  5. T2 and T3 defaults differ

basics

~20 s

Burstable T-family instances earn CPU credits at a fixed hourly rate and spend them to run above a size-specific baseline. In standard mode, an exhausted balance throttles the instance to that baseline; unlimited mode keeps bursting and bills surplus credits instead.

solid answer

~50 s

T-family instances are sold as a **baseline** percentage of CPU plus the ability to burst above it. Each size accrues CPU credits at a fixed rate per hour; one credit is one vCPU-minute at 100%. Running below baseline banks credits, up to a cap of roughly a day's accrual; running above baseline spends them. In **standard** mode, once `CPUCreditBalance` hits zero the instance is hard-throttled to its baseline — which is exactly the symptom of a flat CPU line with a growing request queue. In **unlimited** mode, the instance keeps bursting and accrues surplus credits, which are paid back by future earnings or billed at a flat rate per vCPU-hour. T3, T3a and T4g default to unlimited; T2 defaults to standard. The fix is either to enable unlimited, or to accept that a sustained-CPU workload does not belong on a burstable family at all — move it to `m` or `c`.

code

bash · 4 lines
bash
aws ec2 describe-instance-credit-specifications --instance-ids i-0123456789abcdef0

aws ec2 modify-instance-credit-specification \
  --instance-credit-specification InstanceId=i-0123456789abcdef0,CpuCredits=unlimited

go deeper

for a junior

Know that T-family instances are burstable rather than full-speed all the time, and that a CPU graph pinned to a flat low line points at exhausted CPU credits.

for a middle

Explain accrual, baseline, the balance cap, and what standard mode does at zero balance versus what unlimited mode does — including that T2 and T3 ship with different defaults.

for a senior

Diagnose it live from CloudWatch, distinguish a throttled instance from a slow dependency, and decide whether the right fix is unlimited mode or moving the workload off the burstable family entirely.

for a principal

Own the guidance about where burstable families are allowed at all. Weigh the fleet-wide surplus-charge exposure of defaulting to unlimited against the outage risk of standard mode, and put the choice in a template rather than in each team's hands.

## What "burstable" actually sells you Every EC2 instance type is a slice of a physical host, but the T family sells a slice that is *smaller than it looks*. A t3.large reports 2 vCPUs to the operating system, and those vCPUs really exist — but your entitlement is only a **baseline** fraction of them over time. The credit system is the accounting mechanism that enforces that entitlement while still letting you exceed it in short bursts. That is the whole idea: most small services idle most of the time and spike occasionally, so AWS can oversubscribe the host and pass the saving on. ## Credits, baseline and the balance cap Three numbers define the behaviour, and all three are per instance **size**: 1. **Accrual rate** — credits earned per hour, fixed for the size. 2. **Baseline utilisation** — the percentage of a vCPU you can sustain forever without spending any net credits. Small sizes have baselines in the low tens of percent; larger T sizes have higher baselines. 3. **Maximum balance** — the cap on banked credits, which corresponds to roughly 24 hours of accrual. One CPU credit equals one vCPU running at 100% for one minute; equivalently, one vCPU at 50% for two minutes. When actual utilisation is below baseline you bank the difference; when it is above, you burn the difference. The 24-hour cap is why the symptom in the question has that particular shape: a freshly launched or long-idle instance carries a full balance, absorbs the traffic beautifully for hours, then falls off a cliff when the bank empties. ## Standard mode: the cliff In standard mode, an empty balance means the hypervisor throttles you to baseline. The signature in CloudWatch is unmistakable and worth being able to describe: - `CPUUtilization` rises, then pins to a flat line at exactly the baseline percentage and stays there. - `CPUCreditBalance` decays to zero just before that flat line begins. - `CPUCreditUsage` shows the burn rate that got you there. - Application latency climbs and queues grow while the CPU graph looks *calm* — which is why teams often misread it as a downstream problem. T2 instances also receive a one-off allowance of **launch credits** so a fresh instance is not throttled during boot and warm-up; those are limited and not replenished, and they mask the problem during exactly the window in which people test. ## Unlimited mode: the bill instead of the cliff Unlimited mode removes the throttle. The instance keeps bursting and accumulates **surplus credits**, tracked as `CPUSurplusCreditBalance`. Surplus is repaid automatically from credits earned in the following hours; anything not repaid within the trailing window, or left outstanding when the instance is stopped or terminated, is charged at a flat published rate per vCPU-hour and surfaces as `CPUSurplusCreditsCharged`. Defaults differ by generation, which is the single most-missed fact here: **T2 launches in standard mode, while T3, T3a and T4g launch in unlimited by default.** So a modern t3 fleet usually does not throttle — it quietly generates a bill line instead. A team that has explicitly set standard mode (or copied an old launch template) gets the throttle behaviour back. ```bash # Check and change the credit mode of a running instance aws ec2 describe-instance-credit-specifications --instance-ids i-0123456789abcdef0 aws ec2 modify-instance-credit-specification \ --instance-credit-specification InstanceId=i-0123456789abcdef0,CpuCredits=unlimited ``` ## Diagnosing it in production The diagnosis is three steps. Confirm the flat CPU line sits at the size's documented baseline rather than at some arbitrary plateau. Overlay `CPUCreditBalance` and check it hit zero first. Then check the credit specification — if it says `standard`, you have found the cause; if it says `unlimited`, look at `CPUSurplusCreditsCharged` on the bill and treat the workload as mis-classified rather than mis-configured. ## The real decision Unlimited mode is a patch, not a design. The credit model is built for **spiky, mostly idle** workloads: dev environments, internal tools, low-traffic sites, bastion hosts. A service with steady CPU demand pays surplus charges every hour it runs and, at that point, a same-size `m` or `c` instance is usually both cheaper and predictable — no cliff, no surprise line item, and a CPU graph that means what it says. Burstable families also compound badly with autoscaling: an autoscaling group that adds T instances under load may add machines whose credit balances are already depleted, so the new capacity is throttled from the moment it becomes healthy. Stating that tradeoff — spiky yes, sustained no — is what an interviewer is listening for, more than the exact accrual numbers.

  • Which CloudWatch metrics would you put on a dashboard to catch this before users do?
    `CPUCreditBalance` with an alarm on a low threshold is the leading indicator — it falls hours before latency moves. Pair it with `CPUCreditUsage` to see the burn rate and `CPUUtilization` to spot the flat line at baseline. On unlimited-mode instances, watch `CPUSurplusCreditBalance` and `CPUSurplusCreditsCharged`, because there the failure mode is financial rather than latency.
  • Why can adding burstable instances to an autoscaling group fail to relieve load?
    Newly launched T3-generation instances do not start with a bank of banked credits the way T2 launch credits imply, so capacity added during a sustained spike can be running near baseline almost immediately. You scale out and the fleet still cannot serve the traffic. For load that stays high long enough to trigger scaling, a non-burstable family is the correct choice.
  • When is a burstable family genuinely the right answer?
    When average utilisation sits well below the baseline and the peaks are short: CI agents, bastion hosts, internal admin tools, staging environments, low-traffic sites. The economics only work if you spend most hours banking credits. If your steady-state utilisation is near or above baseline, you are paying for a discount you never receive.

saying these in an interview costs you the question

  • Thinks a t3.large simply gives you 2 full vCPUs
  • Says credits accumulate forever with no cap
  • Believes unlimited mode is free
  • Assumes T2 and T3 have the same default credit mode
  • Blames the database when CPU flatlines at baseline

context