skip to content

How do you choose between Llama 8B, 70B and 405B for a production workload?

level: principalimportance: should knowfreq 50%

answer

  1. Eval before model, always
  2. Walk up the ladder, not down
  3. Two bytes per parameter at bf16
  4. Bigger weights shrink your batch
  5. The 70B slot kept improving

basics

~20 s

Start from a task evaluation, not a leaderboard. Pick the smallest size that clears your quality bar, then check the footprint: at bf16 the weights alone run roughly 16 GB at 8B, 140 GB at 70B and over 800 GB at 405B. Most teams land on 70B.

solid answer

~50 s

Treat it as an evidence question. Build a task-specific evaluation set first, then walk *up* the ladder until quality clears the bar rather than starting at the top. **8B** is the high-throughput tier: weights are roughly 16 GB at bf16, it fits one accelerator, it fine-tunes cheaply with adapters, and it is a good fit for extraction, classification, routing and grounded RAG answering. **70B** is the default when quality matters — around 140 GB of weights at bf16, so multi-GPU or compression — and the later Llama 3.3 70B closed much of the gap to the flagship, which is why it displaced 405B for most interactive serving. **405B** costs over 800 GB in weights at bf16 and a serving cluster to match; its surviving uses are usually offline, as a synthetic-data generator or judge, not on the request path. Also weigh latency and concurrency, not just accuracy: a bigger model raises time-to-first-token and shrinks how many requests fit in memory at once.

go deeper

for a junior

Know the rough ladder — 8B for cheap high-volume work, 70B when quality matters, 405B rarely — and that weights at bf16 cost about two bytes per parameter.

for a middle

Be able to do the footprint arithmetic out loud and explain why a larger model both slows each request and reduces how many requests fit in memory at once.

for a senior

Show the process: build a task eval, walk up from the smallest model, exhaust retrieval and tuning fixes before escalating, and measure latency and concurrency alongside accuracy.

for a principal

Own the fleet policy — how many size tiers the organization runs, what evidence promotes a workload to a larger one, and how you weigh a marginal quality gain against operational surface and cost per token.

## Frame the question correctly The wrong version of this question is "which model is best". The right one is "what is the smallest variant that clears my quality bar, and what does it cost me in latency, memory and operational complexity?" Everything else follows from that framing, and interviewers at senior and above are listening for it explicitly. ## Step one: an evaluation set before a model You cannot choose without a task-representative eval — a few hundred real examples with a scoring method you trust, whether exact match, a rubric-scored judge, or human review on a sample. Public benchmarks tell you about averages on tasks that are not yours. Once the eval exists, run it bottom-up: start at 8B, and only move up when the smaller model demonstrably fails. Teams that start at the flagship almost never come back down, because nobody wants to argue for reducing quality after the fact. ## Step two: the footprint arithmetic Weights at bf16 cost about two bytes per parameter, so: - **8B** ≈ 16 GB — comfortably one modern accelerator, with room left for KV cache and batching. - **70B** ≈ 140 GB — beyond a single 80 GB device, so either tensor parallelism across multiple GPUs or reduced precision. - **405B** ≈ 810 GB — a multi-node serving cluster, with all the reliability and scheduling burden that brings. Weights are only the floor. Every concurrent request also holds a KV cache proportional to its context length, so long-context traffic pushes the real requirement well above these figures. Compression changes the numbers but is a separate decision with its own accuracy tradeoffs; the parameter-count arithmetic is what tells you which order of magnitude of hardware you are shopping for. ## Step three: latency and concurrency, not just accuracy Decoding is memory-bandwidth-bound: each generated token requires streaming the weights and the caches through memory. Larger weights mean fewer tokens per second per request, all else equal. They also mean less memory left for KV cache, which lowers how many requests you can batch, which raises cost per token and queueing delay under load. A model that is 3% better on your eval but halves your concurrency is often a bad trade for an interactive product and a fine one for a nightly batch job. State that distinction — it is the marker of production judgement. ## Step four: know what displaced what The strongest single fact for this question is that the 70B slot kept improving. Llama 3.3's instruction-tuned 70B reports quality close to the 3.1 405B on many benchmarks while costing roughly a sixth of the weight footprint. That is why 405B largely disappeared from interactive serving: the flagship's remaining jobs are ones where latency does not matter — generating synthetic training data, distilling into smaller students, or acting as an offline judge in an eval harness. Saying this shows you track the family rather than the launch announcement. ## Step five: the alternatives to going bigger Before escalating a size tier, exhaust the cheaper levers, because they often beat it: - **Better retrieval.** Most "the model isn't smart enough" failures in RAG are actually retrieval failures; a bigger model reasoning over the wrong documents still answers wrong. - **Task-specific adaptation.** A small model tuned on your distribution frequently beats a much larger general one on that narrow task, at a fraction of the serving cost. - **Decomposition and routing.** Send the easy majority of traffic to 8B and escalate only hard cases to 70B. A router plus a small model often lands the same aggregate quality at far lower cost — with the complexity cost of running two model tiers, which you must be honest about. - **Prompt and output-format work.** Cheap, fast to iterate, and it changes the eval numbers more often than people expect. ## Step six: the organizational tradeoff Every additional size in production is another checkpoint to evaluate, another set of serving configs, another regression surface at upgrade time, and another thing on call. There is real value in standardising on one tier even when a second would be marginally better for some feature. Conversely, a single tier for everything means either overpaying on easy traffic or underserving hard traffic. Owning that call — how many tiers your organization runs, and what evidence moves a workload between them — is the principal-level content of this question. ## Answering well Give the process: eval first, smallest-that-passes, footprint arithmetic, latency and concurrency effects, cheaper levers before a bigger model, and an explicit fleet policy on how many tiers you run. Anchor it with the concrete numbers — 16 GB, 140 GB, 810 GB at bf16 — and with the observation that a refreshed 70B removed most of the reason to serve 405B interactively.

  • Your eval shows 70B beats 8B by four points. What do you check before approving the upgrade?
    Whether those four points change any user outcome, and what they cost. Check the distribution: if the gain is concentrated on a rare hard slice, route only that slice upward instead of moving all traffic. Then price the change in time-to-first-token, tokens per second, and concurrency at your memory budget. Also test whether retrieval fixes or task-specific tuning on the 8B recover the gap more cheaply.
  • Where does the 405B model still earn its cost?
    Offline work, mostly. Generating synthetic training data for smaller students, serving as a judge in an evaluation harness, and one-off hard analyses where a slow answer is acceptable. All of those tolerate seconds or minutes of latency and run at low concurrency, so the multi-node footprint is amortised over a batch rather than paid per user request.
  • What is the argument against running three different Llama sizes in production?
    Operational surface. Each tier is its own evaluation baseline, serving configuration, capacity plan, upgrade path and on-call burden, and routing between them adds a component that can misroute. The savings must be large and durable to justify it. Many teams settle on two tiers — a cheap workhorse and an escalation target — and treat a third as needing an explicit business case.

saying these in an interview costs you the question

  • Choosing by public leaderboard instead of task eval
  • Starting at the largest model and never re-testing smaller
  • Ignoring that bigger weights shrink batch size
  • Assuming 405B is strictly better for every task
  • Escalating model size before fixing retrieval

context