skip to content

How do Llama 2, Llama 3 and Llama 3.1 differ in sizes and context length?

level: middleimportance: must knowfreq 72%

answer

  1. Three generations, three different windows
  2. Sizes narrow, then widen again
  3. 4K, then 8K, then 128K
  4. 405B arrives with the 3.1 release
  5. 3.3 is a single 70B model

basics

~20 s

Llama 2 shipped 7B, 13B and 70B with a 4K context. Llama 3 narrowed to 8B and 70B at 8K context with a much larger tokenizer vocabulary. Llama 3.1 kept those sizes, added a 405B flagship, and extended all of them to 128K context.

solid answer

~50 s

**Llama 2** (2023) released 7B, 13B and 70B weights with a 4,096-token context and a 32K-entry SentencePiece tokenizer. **Llama 3** (2024) dropped the 13B tier and shipped 8B and 70B, doubling context to 8,192 tokens and moving to a ~128K-entry BPE vocabulary, which makes the same English text cost noticeably fewer tokens. **Llama 3.1** kept 8B and 70B, added the 405B flagship, and extended every size to a 128K-token context via RoPE scaling plus long-sequence continued pretraining. The point releases after it are variant additions rather than new base sizes: 3.2 added on-device 1B/3B text models and 11B/90B vision models, and 3.3 shipped a single instruction-tuned 70B whose quality sits close to 405B. So the practical menu today is roughly 1B/3B for edge, 8B for cheap high-throughput work, 70B as the workhorse, and 405B only when you truly need the ceiling.

go deeper

for a junior

Be able to name the current size ladder — roughly 1B and 3B for edge, 8B, 70B, and a 405B flagship — and say that context grew from 4K in Llama 2 to 128K from Llama 3.1 onward.

for a middle

Explain what changed mechanically between generations: the tokenizer vocabulary jump at Llama 3, grouped-query attention at every size, the 405B addition and the 128K window in 3.1, and why more training data made the 8B beat the older 13B.

for a senior

Show you track generations in production: which one you deployed, what forced the choice, and how a tokenizer or context change altered your cost model, chunking strategy, and cached evaluation baselines.

for a principal

Own the upgrade policy. Decide how often the fleet re-baselines onto a new generation, what evaluation evidence justifies moving, and how you avoid pinning a product to a size tier that the next release quietly makes obsolete.

## Why the generation matters more than the number People say "we're running Llama 70B" as if that identified a model. It does not: a Llama 2 70B and a Llama 3.3 70B have the same parameter count and almost nothing else in common — different tokenizer, different context length, roughly an order of magnitude more training data, and very different instruction-following quality. Interviewers ask this question to check that you track the *generation*, because that is what determines context budget, tokenizer economics, and whether a benchmark number you read still applies. ## Llama 2 (2023) Sizes: 7B, 13B, 70B, each in a base and a chat-tuned variant. (A 34B was described in the paper but never released.) Context window: 4,096 tokens. Tokenizer: SentencePiece BPE with a 32,000-token vocabulary. Attention: standard multi-head attention on 7B and 13B; only the 70B used grouped-query attention. Trained on roughly 2 trillion tokens. The 4K window is the fact that dates this generation hardest — it is too small for document QA or long agent transcripts without chunking, and it is the reason most Llama 2 deployments have been retired. ## Llama 3 (2024) Sizes: 8B and 70B. The 13B tier disappeared; Meta's line since has been that a well-trained 8B covers the small slot and 70B covers the mid slot. Context: 8,192 tokens. Tokenizer: a tiktoken-style BPE with a 128,256-entry vocabulary — about four times larger, which compresses text into fewer tokens. That matters commercially even when you self-host, because fewer tokens per document means less KV cache, faster prefill, and more real text inside a fixed window. Attention: grouped-query attention at *every* size, including 8B, with 8 key/value heads. Training data grew to 15 trillion-plus tokens, which is the main reason an 8B Llama 3 outperforms a 13B Llama 2 on most tasks. ## Llama 3.1 (2024) The generation that introduced two headline facts. First, **405B**: a dense model, the largest openly released Llama, aimed at frontier-class quality and at generating synthetic data to train smaller models. Second, **128K context on every size** — 8B, 70B and 405B all advertise 131,072 positions, achieved by scaling the rotary position embeddings and then continuing pretraining on long sequences in stages. Architecture is otherwise a continuation of Llama 3: dense transformer, GQA, no mixture-of-experts. ## The point releases - **Llama 3.2** added a small tier — 1B and 3B text-only models built by pruning and distilling from the larger Llama 3.1 models, intended for phones and edge devices — and a vision tier, 11B and 90B, which attach an image encoder to a Llama 3.1 text backbone. - **Llama 3.3** shipped exactly one model: an instruction-tuned 70B, text-only, 128K context, whose reported quality lands close to 3.1 405B on many benchmarks. For most teams this is the release that removed the reason to run 405B at all. - The **Llama 4** generation (2025) changed the architectural family, moving the flagships to a mixture-of-experts design where only a fraction of the total parameters is active per token. If you cite Llama 4 numbers, be explicit that total and active parameter counts are different figures — a mistake interviewers listen for. ## How to reason about the menu Think in slots rather than in individual model names, because the names churn: 1. **Edge/on-device (1B–3B).** Runs on a laptop or phone; good for classification, extraction, rewriting, routing. Do not expect multi-step reasoning. 2. **Small server (8B).** The high-throughput workhorse: fits comfortably on a single modern accelerator at bf16, cheap to fine-tune with adapters, good enough for RAG answering and structured extraction. 3. **Mid (70B).** The default when quality matters. Needs multi-GPU or aggressive compression to serve at bf16, since weights alone are around 140 GB at two bytes per parameter. 4. **Flagship (405B and the MoE successors).** Justified by evaluation, not ambition. Frequently used offline — for synthetic data generation and as a judge — rather than on the hot path. ## What to actually say in an interview Give the shape of the progression: sizes narrowed and then extended at both ends; context went 4K → 8K → 128K; the tokenizer changed once, at Llama 3, and that change is why token counts differ between generations for identical text. Then say which generation you deployed and why, and what forced the choice — context length, GPU budget, or an eval you ran. Reciting release dates without that judgement reads as memorisation.

  • Why did the tokenizer change at Llama 3 matter in practice, beyond model quality?
    The vocabulary grew from about 32K to about 128K entries, so the same text encodes into fewer tokens. Fewer tokens means smaller KV cache, faster prefill, and more real content inside a fixed window. It also breaks naive comparisons: a Llama 2 token count and a Llama 3 token count for the same document are not the same number, so cost and context-budget estimates must be recomputed per generation.
  • If Llama 3.3 70B is close to 3.1 405B, is there any reason left to run 405B?
    Occasionally. Reported parity is benchmark-level and average; on the hardest reasoning, long-horizon agentic, and multilingual tails the larger model can still lead. The common surviving use is offline: generating synthetic training data or acting as a judge model, where latency does not matter and quality does. For interactive serving, most teams cannot justify the roughly six-fold weight footprint.
  • Llama 3.2 added 1B and 3B models — how were they produced?
    By pruning and knowledge distillation from the larger Llama 3.1 models rather than training tiny models from scratch. Structured pruning removes capacity from a bigger checkpoint, then distillation trains the smaller student against the larger teacher's outputs. That is why they punch above what their parameter count would suggest, while still being clearly weaker at multi-step reasoning.

saying these in an interview costs you the question

  • Saying Llama 3 shipped a 13B size like Llama 2
  • Claiming Llama 2 supported a 128K context
  • Assuming only 405B got the 128K window in 3.1
  • Treating Llama 2 and Llama 3 token counts as interchangeable
  • Calling 3.2 and 3.3 full new base-model generations

context