How do you pick a GPU SKU (L4, A10, A100, H100) for serving an 8B LLM?
answer
- Capacity gates, bandwidth speeds, price decides
- decode re-reads weights every step
- FP8 needs Ada or Hopper
- cheapest card that clears the SLO
- many small replicas versus one flagship
basics
~20 sCapacity gates the choice, bandwidth sets the speed, and price decides between what is left. An 8B model at bf16 needs ~16 GB, so a 24 GB L4 or A10 holds it but leaves thin cache room; bigger cards buy concurrency and far faster decoding.
solid answer
~50 sWork three filters in order. **Capacity**: 8B at bf16 is about 16 GB of weights, so a 24 GB L4 or A10 fits it with only ~6 GB left for KV cache — enough for a few short sessions, not for long-context traffic. **Memory bandwidth**: decoding re-reads the weights every step, so per-stream token rate tracks GB/s. An L4 sits around 300 GB/s, an A10 around 600 GB/s, an A100-80GB near 2 TB/s and an H100 SXM above 3 TB/s — roughly an order of magnitude between the cheapest and the fastest card for the same model. **Kernel support and price**: FP8 needs Ada or Hopper (L4, L40S, H100, H200); Ampere cards such as A10 and A100 do INT8 and 4-bit weight-only instead. Then pick the cheapest card per hour that clears your SLO at your target concurrency — for a small model that is often several mid-tier cards rather than one flagship.
go deeper
Know that a GPU has to hold the weights plus room for cache, and that the card's memory size is the first thing you check before anything else.
Be able to rank cards by memory bandwidth and explain why that predicts decoding speed, and know which generations support FP8 versus INT8.
Justify a fleet: measured tokens/s per dollar on your own prompt mix, replicas versus one large card, and the SLO that made the decision. Include quota availability, not just specs.
Own the fleet-standardization tradeoff. Committing to one SKU generation fixes your precision options, your checkpoint pipeline and your reserved-capacity negotiation for a year or more.
## The decision, in order GPU choice is not "buy the best card". It is three filters applied in sequence, and only the last one is about money. ### Filter 1: does it fit, with room to work? Capacity is a hard gate. An 8B model at bf16 is roughly 16 GB of weights. On a 24 GB card that leaves about 6 GB after CUDA context and workspace — perhaps 15-20k tokens of KV cache for a grouped-query 8B shape. That is fine for chat turns of a few thousand tokens and hopeless for 32k-context document work. Quantizing the same model to FP8 or 4-bit drops weights to ~8 GB or ~4.5 GB and turns the same 24 GB card into a genuinely useful server, because everything you free goes to cache. ### Filter 2: how fast can it read the weights? Token-by-token decoding is dominated by streaming weights out of memory, so single-stream speed tracks memory bandwidth far more than it tracks FLOPs. The rough ceiling is bandwidth divided by the bytes read per step. For an 8B bf16 model (16 GB per step): - L4 (~300 GB/s) -> around 18 tokens/s - A10 (~600 GB/s) -> around 37 tokens/s - A100-80GB (~2 TB/s) -> around 125 tokens/s - H100 SXM (~3.3 TB/s) -> around 200 tokens/s These are ceilings, not measurements, but the ratios are the point: the same model is not "a bit slower" on a cheap card, it is multiples slower per stream. Batching recovers aggregate throughput on all of them, because one weight read serves the whole batch — which is why a slow card can still be economical for offline work while being unusable for an interactive SLO. ### Filter 3: kernels and price Precision support is generational, and it is a deployment constraint you cannot patch around: - **Hopper (H100, H200)** and **Ada (L4, L40S)** have FP8 tensor cores. - **Ampere (A10, A100)** does not; there you use INT8 or 4-bit weight-only checkpoints instead. So "we'll just run the FP8 checkpoint" is a statement about which cards you are allowed to buy. Only after capacity, bandwidth and kernel support have narrowed the field do you compare hourly price — and compare it per unit of *served throughput*, not per card. A card that is three times the price but four times the tokens per second is cheaper. ## Small model, many cheap cards; big model, few big cards For an 8B model the interesting choice is usually horizontal. Several 24 GB cards each running an independent full replica give you fault isolation, simple scaling, and no interconnect traffic at all — no sharding, no all-reduce. One flagship card gives you far better per-stream latency and much more cache room for long contexts. The deciding question is what the workload looks like: many short interactive sessions favour replicas of a cheap card; few long-context sessions with a tight time-per-token target favour the fast card. For models that do not fit on one device, the choice inverts — capacity forces sharding, and then interconnect quality (NVLink between cards in a server versus PCIe) becomes part of the SKU decision rather than an afterthought. ## What else quietly matters - **Availability.** The fastest card you cannot get a quota for is not a plan. Mid-tier SKUs are often available on demand when flagships are not. - **Form factor and power.** L4 is a 72 W single-slot part that fits almost anywhere; that is why it shows up in cheap inference fleets. - **Host memory and disk.** Loading tens of gigabytes of weights on every replica start is a real cost, and the node's disk and network decide how long it takes. - **What you actually measured.** Publish a tokens/s figure from your own model, prompt distribution and concurrency before committing to a fleet. Vendor peak numbers are for a different workload than yours. ## The trap answer "H100, because it is the fastest" is the answer that ends the conversation badly. For an 8B model at modest concurrency it can be several times the cost per token of a smaller card, because you cannot keep it busy. Utilization, not peak specification, is what you are buying.
- Why can a cheaper, slower card still win on cost per token for a batch job?Because batched decoding reads the weights once per step for the whole batch, so aggregate throughput scales with batch size on any card. If nothing is waiting on per-stream latency, you can saturate a cheap card and pay far less per hour for a similar tokens/s total. Interactive traffic cannot do this — a real time-per-token SLO caps how much batching you can hide behind.
- Your 8B model fits on a 24 GB card. When would you still choose an 80 GB one?When you need long contexts or high concurrency: cache room, not weight room, is what runs out. Weights take 16 GB either way, but the 80 GB card leaves roughly ten times the KV headroom, which is the difference between two 32k sessions and twenty. It also helps when you expect to serve a larger model on the same fleet later.
- How does picking an FP8 checkpoint constrain your hardware choices?FP8 tensor-core kernels require Ada or Hopper class GPUs — L4, L40S, H100, H200. On Ampere parts such as A10 and A100 you need an INT8 or 4-bit weight-only checkpoint instead. So the precision decision and the SKU decision are one decision, and mixing generations in a fleet means maintaining two checkpoints.
saying these in an interview costs you the question
- Choosing the flagship card by default
- Comparing GPUs on FLOPs while decoding is bandwidth-bound
- Forgetting KV cache when checking if the model fits
- Assuming any GPU can run an FP8 checkpoint
- Quoting vendor peak throughput instead of measuring your workload