skip to content

What distinguishes Llama 3.2's 1B and 3B models from its 11B and 90B ones?

level: juniorimportance: should knowfreq 42%

answer

  1. One release, two very different tiers
  2. Small pair is text-only
  3. Larger pair adds image input
  4. Not trained from scratch
  5. Pruned and distilled from bigger checkpoints

basics

~20 s

Llama 3.2 shipped two different tiers. The 1B and 3B are small text-only models built for on-device and edge use, produced by pruning and distilling larger Llama 3.1 checkpoints. The 11B and 90B are vision models that accept images alongside text.

solid answer

~50 s

The Llama 3.2 release covers two unrelated needs under one version number. The **1B and 3B** are text-only, multilingual, 128K-context models aimed at phones, laptops and edge devices — small enough to run locally, suited to summarisation, rewriting, extraction, classification and routing, but clearly weaker at multi-step reasoning. They were not trained from scratch: Meta produced them by structured pruning plus knowledge distillation from the larger Llama 3.1 models, which is why they outperform their parameter count. The **11B and 90B** are the family's first vision models, pairing an image encoder with a Llama 3.1 text backbone so the model can answer questions about images, charts and documents. Practically: if you need image input, you need the vision tier and its larger footprint; if you need something to run on a device, you take 1B or 3B and design around their reasoning limits.

go deeper

for a junior

Recall that Llama 3.2 has a small text-only pair for on-device work and a larger pair that reads images, and that the small models are for narrow tasks, not deep reasoning.

for a middle

Explain how the small models were made — pruning plus distillation from Llama 3.1 — and how the vision models attach an image encoder to a text backbone, plus what each tier costs to run.

for a senior

Show selection judgement: which parts of a real pipeline you would push onto a 3B on-device model, where you escalate to a server-side text model, and how you evaluate whether the small model's quality holds.

for a principal

Own the tiering strategy — what runs on the device versus the datacentre, how you route between them, and what the privacy, latency and cost arguments are for keeping a small model in the product at all.

## One version number, two product decisions Llama 3.2 is easy to misread because it bundled two things that solve different problems. Nothing about the 1B/3B pair is multimodal, and nothing about the 11B/90B pair is small. Being clear about which tier you mean is most of the answer. ## The small tier: 1B and 3B These are text-only models with the same 128K context as the rest of the 3.x line, positioned for **on-device and edge** deployment: mobile phones, laptops, embedded hardware, or any setting where sending a request to a datacentre is not acceptable for latency, cost or privacy reasons. How they were built matters, because it explains their quality. Rather than pretraining tiny models from scratch, Meta used **structured pruning** — removing capacity from a larger Llama 3.1 checkpoint in a principled way — followed by **knowledge distillation**, where the smaller student is trained against the outputs of a larger teacher rather than only against raw text. Distillation transfers a richer signal than next-token labels alone, so the resulting models land above what their parameter count would naively predict. What they are good for: classification and routing, entity and field extraction, summarising a short document, rewriting and tone adjustment, powering an on-device assistant that hands harder requests upstream. What they are not good for: multi-step reasoning, long tool-use chains, code generation of any complexity, or anything where a subtle factual error is expensive. Treat them as fast, cheap components inside a larger system rather than as a general-purpose brain. ## The vision tier: 11B and 90B These accept images alongside text. Architecturally they attach a separately trained image encoder to an existing Llama 3.1 text model through adapter layers, so the language capability is inherited and the visual capability is added on top. Typical uses are document and chart understanding, image captioning, visual question answering, and grounding a question in a screenshot or photo. Two consequences follow. First, footprint: 11B is meaningfully larger than 8B in weights and larger again in activation memory once image features are in play, and 90B is a multi-accelerator proposition. Second, capability shape: they are image-*understanding* models. They read images; they do not generate them. Anyone answering as if these were image generators has confused two entirely different model categories. ## How the tiers sit against the rest of the family The rest of the 3.x line is text-only and dense: 8B for cheap high-throughput serving, 70B as the quality workhorse (with the later 3.3 release refreshing that slot), 405B as the flagship. Llama 3.2 widened the family at both ends without replacing that middle. So the selection question in an interview usually reduces to three checks: 1. **Does the input include images?** If yes, you are in the vision tier; nothing else in the 3.x line takes an image. 2. **Does it have to run on the device?** If yes, 1B or 3B, and you scope the task accordingly. 3. **Otherwise**, pick from the text ladder by quality need and budget. ## Common traps - Assuming 3.2 replaced 3.1 wholesale. It did not; it added tiers alongside it, and 70B-class text quality came later with 3.3. - Assuming the vision models take video. Video and audio are not part of what this tier offers; treat it as still-image understanding. - Assuming small automatically means fast enough on any device. A 3B model at reasonable precision still needs several gigabytes of memory, and phone-class throughput is modest. - Assuming distillation makes a 3B competitive with a 70B. It closes some of the gap for narrow, well-specified tasks and closes almost none of it for reasoning. ## Answering well Name the two tiers, say what each is *for*, mention that the small pair came from pruning and distillation rather than scratch training, and finish with the selection rule — images force the vision tier, on-device forces the small tier, everything else picks from the text ladder on cost and quality.

  • Can the Llama 3.2 vision models generate images as well as read them?
    No. They are image-understanding models: an image encoder feeds visual features into a Llama text backbone, and the output is text. Captioning, chart and document question answering, and visual grounding all work; producing an image does not. Image generation is a different model category entirely, and conflating the two is a common interview slip.
  • When would you choose Llama 3.2 3B over the 8B model?
    When the deployment target forces it or the task is narrow. On-device or edge settings, strict privacy constraints that rule out a server round trip, or very high-volume simple work such as classification and routing where 8B's extra quality does not change outcomes. For anything involving multi-step reasoning, tool chains or code, the 8B is the safer floor.
  • Why do the 1B and 3B models perform better than their size suggests?
    Because they were derived from stronger models rather than trained from scratch. Structured pruning removes capacity from a larger Llama 3.1 checkpoint, and knowledge distillation then trains the smaller student against the larger teacher's output distribution, which carries far more signal than next-token labels alone. The gap to the larger models narrows most on narrow, well-specified tasks and least on reasoning.

saying these in an interview costs you the question

  • Calling the 11B and 90B image generators
  • Thinking 1B and 3B accept image input
  • Assuming they were pretrained from scratch
  • Saying Llama 3.2 replaced the whole 3.1 line
  • Believing a 3B model runs on any phone unmodified

context