skip to content

Llama 70B won't fit one GPU — how do you choose between tensor parallelism, quantization, and a smaller model?

level: principalimportance: should knowfreq 42%

answer

  1. quality floor and SLO decide, not what fits
  2. shard matrices versus shrink weights versus shrink model
  3. all-reduce every layer wants NVLink
  4. the smallest model that passes is the answer
  5. price per million tokens at real load

basics

~20 s

Decide from the quality floor and the latency SLO, not from what fits. Tensor parallelism keeps full precision but needs multiple GPUs with fast interconnect; quantization fits fewer GPUs at some accuracy cost; a smaller model is cheapest and often good enough once measured on your own evaluations.

solid answer

~60 s

There are three levers and they trade different currencies. **Tensor parallelism** (`--tensor-parallel-size N` in vLLM, `--num-shard` in TGI) shards every layer's weight matrices across N GPUs in one node. Memory divides cleanly — 140 GB of fp16 70B weights becomes ~35 GB per GPU at TP=4 — and latency improves, but every layer ends in an all-reduce, so it wants NVLink or equivalent; over PCIe the collective cost eats the gain. N must divide the head counts, which for Llama's 8 KV heads means 1, 2, 4 or 8. **Quantization** shrinks weights so fewer GPUs are needed and frees VRAM for KV cache, at an accuracy cost you must measure rather than assume. **A smaller model** removes the problem: an 8B on one GPU serves far more traffic per dollar, and for extraction, routing or summarisation it frequently passes the same evaluation. The disciplined order is: define the quality bar with a task evaluation, test the smaller model first, then quantized-large, then multi-GPU full precision — and price each per million tokens at your real load.

code

bash · 5 lines
bash
# Four NVLink-connected GPUs in one node, full precision
vllm serve meta-llama/Llama-3.1-70B-Instruct \
  --tensor-parallel-size 4 \
  --max-model-len 8192 \
  --gpu-memory-utilization 0.92

go deeper

for a junior

Know the three options exist — split the model across GPUs, shrink the weights, or use a smaller model — and that a 70B model does not fit one typical GPU at full precision.

for a middle

Explain the mechanics: tensor parallelism shards each layer's matrices and all-reduces per layer, quantization trades accuracy for memory, and the sharding degree must divide the head counts.

for a senior

Show the operating judgment: keep tensor parallelism inside a fast-interconnect node, validate quantization on your own evaluation set, and load-test before declaring a shape production-ready.

for a principal

Own the decision framing — quality floor, latency SLO and cost per million tokens first; replica shape and workload isolation next; and be willing to conclude that self-hosting this size does not earn its keep.

## Frame it as a decision, not a fitting exercise Llama 3.1 70B at fp16 is roughly 140 GB of weights plus ~320 KiB of KV cache per token. Nothing about that dictates an answer; the answer comes from three constraints you should state before touching a flag: the **quality floor** on your actual tasks, the **latency SLO** (time-to-first-token and inter-token latency), and the **cost per million tokens** at forecast volume. Candidates who jump straight to "shard it across four GPUs" have skipped the only interesting part. ## Lever 1: tensor parallelism Tensor parallelism splits individual weight matrices across GPUs — attention heads and MLP columns are partitioned, each GPU computes its slice, and the results are combined with an all-reduce at the end of each block. In vLLM this is `--tensor-parallel-size`; in TGI, `--num-shard`. What it buys: memory divides almost linearly, and because each GPU reads only its shard of the weights, per-token decode latency also improves — TP is the standard way to get a large model *fast*, not merely to make it fit. What it costs: two all-reduces per transformer layer, on every forward pass. On NVLink-connected GPUs inside one node this is cheap. Over PCIe, or worse across a network, the collective can dominate and you end up slower than a single-GPU quantized deployment. It also introduces failure coupling — one GPU lost takes the whole replica down — and the degree must divide the model's head counts, which for Llama's 8 KV heads means 1, 2, 4 or 8. Crossing nodes is a different regime: **pipeline parallelism** (`--pipeline-parallel-size`) splits layers rather than matrices, sending far less data between stages, so it tolerates slower links but adds pipeline bubbles and does not reduce per-request latency. The usual composition is TP within a node, PP across nodes — and "across nodes" is an admission that the model is too big for your hardware, with all the operational cost that implies. ## Lever 2: quantization Reducing weight precision cuts the weight budget proportionally and, just as importantly, frees VRAM that becomes KV cache — which is what actually sets concurrency. It can turn a two-GPU deployment into one, or make a 70B fit where it otherwise would not. The cost is accuracy, and the honest position is that it is workload-dependent: degradation that is invisible on chat may be fatal on structured extraction or code. Never accept a published benchmark as evidence for your workload; run your own evaluation set. Also confirm the serving stack supports the specific scheme on your hardware and that a quantized kernel path exists, or you will lose the speed you were buying. ## Lever 3: a smaller model The lever most often skipped, and most often correct. Llama 3.1 8B on a single GPU serves several times the throughput per dollar and eliminates the interconnect question entirely. Many production tasks — classification, routing, entity extraction, short summarisation, tool-argument filling — show little measurable gap once you have a real evaluation, especially with a good prompt or a task-specific fine-tune. A frequent architecture is a small model handling the bulk of traffic with escalation to a large one for the minority of hard cases, which usually beats running everything on the large model. ## The order of operations 1. **Build the evaluation first.** Without a task-specific quality bar, every subsequent comparison is opinion. 2. **Try the smaller model.** If it clears the bar, stop; you have removed the problem. 3. **Try quantized-large on the fewest GPUs that fit.** Re-run the evaluation, then load-test. 4. **Then multi-GPU full precision**, sized so tensor parallelism stays inside a node with fast interconnect. 5. **Price each survivor** per million tokens at forecast load, including idle GPU time — a large model at 20% utilisation is often more expensive per token than the hosted alternative. ## The strategic questions behind it Replica shape matters as much as fit: two independent TP=4 replicas give you rolling upgrades and survive a node failure; one TP=8 replica gives lower latency and a single point of failure. Workload isolation matters too — batch summarisation sharing a pool with interactive chat will violate the chat SLO regardless of parallelism strategy, so separate pools are frequently the right answer. And the framing question worth raising explicitly: does self-hosting 70B earn its keep at all? Self-hosting wins on data residency, per-token cost at genuinely high sustained volume, latency control, and freedom from vendor deprecation. It loses on engineering time, idle capacity, and the fact that the frontier moves. A principal-level answer names that tradeoff instead of assuming the fleet.

  • Why does tensor parallelism need a fast interconnect while pipeline parallelism tolerates a slow one?
    Tensor parallelism splits individual matrices, so every layer ends with an all-reduce over full activation tensors — dozens of collectives per token, which only NVLink-class bandwidth absorbs cheaply. Pipeline parallelism splits layers instead, passing one activation tensor per stage boundary, so it survives commodity networking. The price is pipeline bubbles and no reduction in per-request latency, which is why TP is used inside a node and PP across them.
  • How would you decide whether an 8B model is good enough to replace the 70B?
    Build a labelled evaluation from real production traffic covering the tasks you actually run, with a clear pass metric. Run both models, blind-compare, and segment the results — the gap is usually concentrated in a minority of hard cases. If the small model clears the bar on the bulk, deploy it with escalation to the large model for the remainder, which typically costs far less than serving everything large.
  • What is the operational argument for two TP=4 replicas over one TP=8 replica?
    Availability and upgrade safety. Two replicas survive losing a node, allow rolling upgrades without downtime, and let you shift traffic while validating a new version. One wide replica gives lower per-request latency and a slightly larger KV pool per sequence but fails completely if any GPU or link fails. Unless the latency SLO demands the wider shard, two replicas is the more defensible production shape.
  • When is self-hosting 70B the wrong answer entirely?
    When sustained volume is low enough that GPUs sit idle, when the team lacks capacity for on-call GPU operations, or when you would rather ride model improvements than freeze on a checkpoint. Self-hosting earns its keep on data residency, genuinely high steady throughput, latency control, and independence from vendor deprecation. Compute your fully-loaded cost per million tokens at realistic utilisation before committing the fleet.

saying these in an interview costs you the question

  • Choosing the biggest model that fits instead of the smallest that passes
  • Assuming tensor parallelism scales the same over PCIe as over NVLink
  • Treating quantization accuracy loss as negligible without a task evaluation
  • Ignoring that one wide shard is a single point of failure
  • Comparing options on GPU count instead of cost per million tokens

context