skip to content

Which Gemma 3 size fits a single 24 GB GPU, and what do you trade away?

level: seniorimportance: should knowfreq 45%

answer

  1. Two bytes per parameter is the anchor
  2. Weights are not the whole budget
  3. Vendor-trained low precision beats DIY
  4. Sliding-window layers exist for a reason
  5. Fits at rest is not fits under load

basics

~20 s

At bfloat16 you need roughly 2 GB per billion parameters, so 24 GB holds 4B or 12B but not 27B's ~54 GB. Google's quantisation-aware-trained int4 checkpoints shrink 27B to roughly 15 GB, which fits — at the cost of some quality and nearly all your KV-cache headroom.

solid answer

~50 s

Start with the arithmetic: weights at bfloat16 cost about **2 bytes per parameter**, so Gemma 3 sizes need roughly 2 GB (1B), 8 GB (4B), 24 GB (12B) and 54 GB (27B) before anything else. On a 24 GB card that means 4B is comfortable, 12B is technically borderline and leaves no room to serve, and 27B is out. Quantisation changes the answer: Google publishes **quantisation-aware-trained (QAT) int4 checkpoints**, and at ~0.5 bytes per parameter 27B lands near 15 GB, so it fits with room to run. What you give up is threefold — some accuracy (QAT recovers most but not all of the naive-quantisation loss), throughput on kernels less optimised than bf16, and **headroom**, because the KV cache for long contexts and concurrent requests competes for the same VRAM. Gemma 3's interleaved local/global attention exists precisely to keep that cache affordable at 128K context, but it does not make it free.

go deeper

for a junior

Learn the anchor: about 2 GB of memory per billion parameters at 16-bit precision, and roughly a quarter of that at 4-bit. That arithmetic already answers most 'will it fit' questions.

for a middle

Explain that weights are only one of three consumers, alongside KV cache and activation overhead, and that quantised checkpoints trade some accuracy for a much smaller footprint.

for a senior

Demonstrate a sizing method: establish a quality floor, walk down the size ladder measuring on your own data, then budget cache from p99 context times concurrency and leave headroom for bursts.

for a principal

Frame it as fleet economics — which sizes you standardise on given your GPU inventory, whether to route cheap traffic to a small model and hard traffic to a large one, and what the quality floor is worth per GPU-hour.

## The three consumers of VRAM Every sizing conversation for a self-hosted model is the same budget with three lines: 1. **Weights** — fixed once you choose a size and a precision. 2. **KV cache** — grows with sequence length multiplied by concurrent requests. This is the line people forget. 3. **Activations and overhead** — workspace for the forward pass, the CUDA context, the vision encoder if you send images, and fragmentation. A card is usable only when all three fit. A model that "fits" with 200 MB spare fits for exactly one short request. ## Weight arithmetic Multiply parameters by bytes per parameter: - **bfloat16 / float16** — 2 bytes. 1B ≈ 2 GB, 4B ≈ 8 GB, 12B ≈ 24 GB, 27B ≈ 54 GB. - **int8** — 1 byte. 27B ≈ 27 GB. - **int4** — ~0.5 bytes plus per-block scales. 27B ≈ 14–16 GB, 12B ≈ 7 GB, 4B ≈ 2.5 GB. On a 24 GB card: 4B at bf16 is easy; 12B at bf16 consumes the whole card and leaves nothing for cache; 27B at bf16 is impossible; **27B at int4 fits with roughly 8–9 GB left over**, which is enough to serve modest contexts at low concurrency. ## Why Gemma's QAT checkpoints matter here Quantising after training is lossy: you round weights to a coarse grid and the model degrades, sometimes sharply on the hardest tasks. **Quantisation-aware training** simulates that rounding *during* training so the model learns weights that survive it. Google publishes int4 QAT checkpoints for Gemma 3 alongside the full-precision ones, which is the reason "27B on one consumer card" is a serious proposition rather than a stunt. The practical rule: prefer a vendor-published QAT build over a quantisation you produced yourself from the bf16 weights, because the vendor spent training compute recovering the loss and you did not. ## The KV cache, and Gemma 3's answer to it The KV cache stores the key and value tensors for every token already in the sequence, for every layer, so generation does not recompute the whole prefix each step. Its size scales with context length, batch size, layer count and head dimension. At a 128K context this term can dwarf the weights in a naive architecture. Gemma 3 addresses this structurally: it **interleaves local sliding-window attention layers with global attention layers**, in a ratio heavily favouring local, with a small window. Local layers only need cache for their window, not for the whole 128K sequence, so the total cache grows far more slowly with context length than an all-global stack would. This is an architecture-level cost decision, and it is the single most useful thing to know about Gemma 3 in a capacity conversation. It still is not free. Long prompts and concurrency multiply. If you fill a 24 GB card with 15 GB of int4 weights and then ask for 128K contexts at batch 8, you will OOM. ## What you actually trade - **Quality.** int4 loses something even with QAT. On easy classification it is invisible; on multi-step reasoning, code, and long-context recall it shows up first. Measure on your own evaluation set, not on a public leaderboard. - **Throughput.** Low-precision kernels are not uniformly faster; on some stacks int4 decode is bandwidth-friendly and wins, on others the dequantisation overhead eats the gain at large batch. Benchmark rather than assume. - **Headroom.** The cost people feel in production is not accuracy, it is that the maximum context and maximum concurrency both shrink, so tail latency spikes under load. ## A defensible sizing method 1. Establish the **quality floor** on your task with the largest size you can rent temporarily. That is your reference. 2. Walk **down** the ladder — 27B int4, 12B bf16, 12B int4, 4B — measuring against the reference on your own data. 3. Compute the budget: weights + KV cache at your **p99 context length** × your target concurrency + ~10–15% overhead. 4. Pick the smallest size that clears the quality floor with the budget satisfied. Smaller models leave headroom, and headroom is what keeps latency stable under bursts. ## Common wrong answers "27B needs 27 GB" confuses parameters with bytes. "Quantisation is free" ignores task-dependent degradation. "It fits, so we're done" ignores the cache. And treating a single-GPU fit as the goal rather than the constraint leads to shipping a model that is one traffic spike from failing. ## Version caution Size line-ups, published quantised builds and architectural details are generation-specific. State the generation you are sizing and re-derive the arithmetic when the family iterates.

  • Your 27B int4 deployment OOMs only under load, never in testing. What is the likely cause?
    The KV cache. Weights are a fixed allocation that testing exercises fully, but cache scales with concurrent sequences and their lengths, so a single-request smoke test never touches the peak. Cap max sequence length and max concurrent sequences explicitly in the serving configuration, size the cache from your p99 context rather than the median, and leave 10–15% of the card unallocated for activations and fragmentation.
  • Would you rather run 12B at bf16 or 27B at int4 on the same 24 GB card?
    27B at int4 usually wins on quality per gigabyte, because parameter count buys more than precision does at these scales, and it also leaves more free VRAM for cache. The exception is when your task is precision-sensitive — tight numerical work, long-chain code generation — or when your serving stack has poor int4 kernels. Decide it by evaluating both on your own task, since the answer is workload-dependent rather than universal.
  • Why prefer a published QAT checkpoint over quantising the bf16 weights yourself?
    Because quantisation-aware training spends real training compute teaching the model to tolerate the rounding, and a post-hoc conversion does not. The published build therefore recovers most of the accuracy a naive int4 conversion loses, at zero cost to you. Rolling your own is worth it only when you need a format or bit-width the vendor did not publish, and even then you should measure against the official build.
  • How does sending images change the memory picture?
    Images pass through a vision encoder and become tokens in the same sequence, so they consume KV cache exactly as text does, and the encoder itself needs weights and activation workspace on the card. A workload with several images per request therefore has a much larger effective context than its text length suggests. Budget from token counts after encoding, not from character counts of the prompt.

saying these in an interview costs you the question

  • Equating parameter count directly with gigabytes
  • Sizing from weights alone and ignoring KV cache
  • Assuming int4 quality loss is always negligible
  • Believing quantisation always increases throughput
  • Testing capacity with one request and calling it sized

context