skip to content

At a fixed VRAM budget, do you run a bigger Llama at 4-bit or a smaller one at FP16?

level: principalimportance: should knowfreq 40%

answer

  1. Parameters usually beat precision
  2. Big models tolerate quantization better
  3. Four bits is the usual floor
  4. Weights crowd out the KV cache
  5. Task shape decides which axis pays

basics

~20 s

Down to about 4 bits, the larger model usually wins: parameter count buys more capability than precision does, and big models absorb quantization loss better than small ones. Below roughly 3 bits the ordering flips, and throughput and kernel support can override both.

solid answer

~50 s

The default answer is the larger model at 4 bits, and the reasoning is that capability scales strongly with parameter count while the marginal value of precision above 4 bits is small. Large models also carry more redundancy, so they lose proportionally less to the same bit width — a 70B at 4 bits typically beats a 13B at FP16 on reasoning-heavy work. The answer stops holding in three situations. **Below about 3 bits**, degradation becomes steep and unpredictable, especially for precise tasks; a smaller, higher-precision model is then the safer bet. **When latency or throughput binds**, the larger model is slower per token and leaves less memory for KV cache, so it serves fewer concurrent requests at your context length. And **when kernel support is poor** for that format on your GPU, the theoretical memory win does not become a throughput win. So state the default, then say you would settle it by running both through the same task eval and the same load test.

go deeper

for a junior

Know the basic tradeoff: quantizing lets a bigger model fit the same card, and down to about 4 bits the bigger model usually gives better answers than a smaller one at full precision.

for a middle

Explain both mechanisms — capability scales with parameter count, and larger models absorb quantization loss better — and be able to say where the curve falls off, around 3 bits and below.

for a senior

Show you weigh throughput and cache headroom, not just quality: how much memory the weights leave for KV cache at your context length, whether your stack has good kernels for that format, and how you would load-test both candidates.

for a principal

Own the whole decision — cost per served token, a two-tier routing design instead of one compromise model, the operational cost of maintaining extra quantized artifacts, artifact provenance, and headroom for future context commitments.

## The framing A fixed memory budget lets you buy capability along two axes: more parameters, or more bits per parameter. The question is which purchase returns more on your workload. It is a genuine tradeoff question, and an interviewer is checking whether you have a defensible default, know its boundaries, and can name the measurement that settles it. ## The default and why it holds Down to roughly 4 bits, spend the budget on parameters. Two effects compound: 1. **Capability scales with scale.** The gap between model sizes in a family — reasoning depth, instruction following, world knowledge, code quality — is large. The gap between FP16 and a well-made 4-bit build of the same weights is comparatively small. 2. **Bigger models quantize better.** Larger networks carry more redundancy across many more parameters, so the same rounding error perturbs their function proportionally less. Empirically, the accuracy drop from FP16 to 4 bits is mild for 70B-class models and considerably harsher for 1B–8B ones. Scale is precisely what buys you the tolerance. So at 48 GB, a 70B at ~4.8 effective bits (about 40 GB) is usually a stronger model than a 13B at FP16 (about 26 GB) or an 8B at FP16 — on reasoning, long-form synthesis and code. ## Where the default breaks **Below about 3 bits.** The quality curve is not linear. From 8 to 5 bits, almost nothing; from 5 to 4, a little; from 4 to 3, a real step; below 3, often a cliff — with precise behaviours (schema-valid output, tool arguments, arithmetic) going first. If fitting the larger model requires 2-bit, the smaller model at 4 or 8 bits is usually the better product, even before you consider that low-bit builds are less predictable across tasks. **Latency and throughput.** A 70B generates fewer tokens per second than a 13B on the same card, roughly in proportion to the bytes moved per token — and quantization helps that, which is why 4-bit large models are viable at all. But the 70B also leaves far less memory for KV cache. At 48 GB with 40 GB of weights, your cache budget is ~8 GB, which caps context length times concurrency hard. A smaller model leaves room for many more concurrent sequences, which for a high-QPS product may be worth more than per-request quality. **Kernel and stack support.** A memory saving only becomes a throughput saving if your serving stack has a good kernel for that format, bit width, group size and GPU generation. Verify before designing the deployment around it. **Task shape.** Classification, extraction, routing and short summarisation saturate early: a small model at high precision often matches a large one, and then the large model is pure cost. Open-ended reasoning, code and agentic tool loops keep rewarding scale. Match the axis you spend on to the work. **Fine-tuning plans.** If you intend to adapt the model, a smaller base you can train quickly and iterate on may beat a larger frozen one — and the smaller model's higher precision leaves more headroom for adapter merging. ## Second-order considerations a lead owns - **Cost per served token**, not cost per GPU. A larger model that halves your fallback-to-a-hosted-API rate may pay for itself; one that halves your throughput may not. - **Two-tier routing.** Often the right answer is neither: a small high-precision model handling the bulk of traffic, escalating the hard minority to the large 4-bit one. That usually beats a single compromise on both cost and quality. - **Operational surface.** Each additional model in production is another artifact to quantize, validate, version and re-verify when weights update. Two tiers is a real ongoing cost, not just a config line. - **Licensing and provenance** of the quantized artifact — who built it, from which weights, with what calibration data — because a community build with unknown provenance is an unaudited dependency in your serving path. - **Headroom for growth.** Sizing a 70B to fill 40 of 48 GB leaves nothing for longer contexts later. Deployments outgrow their context assumptions faster than their quality ones. ## How to settle it Do not argue it in the abstract. Build both candidates, run them through one task eval drawn from production traffic with deterministic scoring, and run one load test at your real context length and concurrency. You then have two numbers per candidate — quality on your work, and tokens per second at your load — and the decision writes itself. State the default (bigger model at 4-bit), name the conditions that flip it, and name the experiment. That combination is the answer an interviewer is listening for.

  • Why do small models suffer more from the same bit width than large ones?
    Redundancy. A large network spreads its function across far more parameters, so per-weight rounding error perturbs the whole computation proportionally less, and there is more surviving signal for any single damaged channel. Small models have less slack — each weight carries more of the function, so 4-bit rounding removes a larger share of what they encode. That is why a 70B at 4 bits is a mild compromise while a 1B at 3 bits often is not usable for precise work.
  • How does the choice change when your workload is extraction and classification rather than open-ended reasoning?
    It flips toward the smaller, higher-precision model. Narrow tasks saturate early — once a model reliably reads a field or picks a label, extra scale adds little quality but real cost and latency. Spending the budget on precision, throughput and headroom for concurrency serves that product better. Keep the large model, if at all, as an escalation path for the minority of inputs that genuinely need it.
  • What does filling most of your VRAM with weights cost you operationally?
    KV-cache headroom, and therefore the product of context length and concurrency you can serve. With 40 GB of weights on a 48 GB card, roughly 8 GB is left for cache and activations, which caps how many long-context requests can be in flight before you start evicting or queueing. It also leaves no room to grow your context commitment later, which deployments outgrow sooner than they outgrow their quality target.

saying these in an interview costs you the question

  • Claiming a bigger model always wins regardless of bit width
  • Treating 2-bit as just another point on a smooth curve
  • Ignoring KV-cache headroom when weights fill the card
  • Assuming a memory saving automatically becomes a speed-up
  • Deciding from public benchmarks instead of your own eval

context