skip to content

Google Gemma

Gemma is Google's open-weight model family you run yourself, now natively multimodal and sized from 1B to 27B. Interviewers ask it when the question is self-hosting rather than calling a hosted API.

on this pageshow

questions

5

What sizes does Gemma 3 come in, and which of them accept image input?

level: juniorimportance: must knowfreq 58%

answer

  1. Four sizes, one of them text-only
  2. Smallest is the odd one out
  3. -it means something in the name
  4. Vision starts above the smallest size
  5. 1B/4B/12B/27B; 4B and up see images

basics

~20 s

Gemma 3 ships at 1B, 4B, 12B and 27B parameters, each in a pretrained and an instruction-tuned variant. The 4B, 12B and 27B checkpoints accept images alongside text; the 1B is text-only and has a much shorter context window.

solid answer

~40 s

Gemma 3 is Google's open-weight family, released in four sizes — 1B, 4B, 12B and 27B parameters — and every size comes as a pretrained base checkpoint plus an instruction-tuned one, marked by the `-it` suffix in the published names (`gemma-3-4b-it`, `gemma-3-27b-it`). Multimodality is not uniform: the 4B, 12B and 27B variants pair the language model with a vision encoder and take interleaved image and text input, while the 1B is text-only. Context length also splits by size — the vision-capable sizes advertise a 128K-token window, the 1B a much smaller one (32K). All of them are weights you download and serve yourself, not a hosted endpoint, so the real selection question is which size your GPU budget supports rather than which API tier you buy.

code

python · 15 lines
python
from transformers import pipeline

pipe = pipeline("image-text-to-text", model="google/gemma-3-4b-it")

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://example.com/chart.png"},
            {"type": "text", "text": "What trend does this chart show?"},
        ],
    }
]

print(pipe(text=messages, max_new_tokens=128))

go deeper

for a junior

Be able to list the four sizes, say that the smallest is text-only, and explain that Gemma is weights you download rather than an API you call. Knowing that -it means instruction-tuned is an easy win.

for a middle

Explain the base versus instruction-tuned split and why you fine-tune from base, and be ready to say how parameter count maps to memory so the size choice becomes an arithmetic question rather than a taste one.

for a senior

Show that you pick a size from measurement: start small, evaluate on your own inputs, and move up only when a real failure mode demands it. Note that images consume context budget and that quantised formats change the sizing calculus.

for a principal

Own the portfolio view: a fleet may want the small size at the edge and a larger one centrally, with routing between them. Frame size selection as a cost-per-quality curve tied to your GPU inventory, not as picking the biggest model available.

## What Gemma is Gemma is Google's family of **open-weight** models: you download the parameter files and run them on your own hardware. That makes it a different kind of product from a hosted API, where you send a request to a vendor's servers and are billed per token. With Gemma there is no endpoint, no API key and no per-token bill — there is a checkpoint, a licence, and whatever GPU or CPU you point at it. ## The four sizes Gemma 3 is published at **1B, 4B, 12B and 27B parameters**. "B" means billions of parameters — the learned numbers inside the model. Parameter count is the single best predictor of two things at once: how capable the model is, and how much memory it needs. Roughly, a model at 16-bit precision needs about 2 GB of memory per billion parameters just to hold its weights, so the four sizes span from something that runs on a laptop to something that needs a serious accelerator. ## Base vs instruction-tuned Every size ships twice: - **Pretrained (base)** — trained only to continue text. It does not reliably follow instructions or hold a conversation. You use it as a starting point for your own fine-tuning. - **Instruction-tuned** — further trained to follow instructions and answer in a chat format. Published names carry an `-it` suffix, e.g. `gemma-3-12b-it`. This is what you want if you are going to prompt the model directly. Picking the base checkpoint by accident and then complaining that the model "ignores instructions" is one of the most common first-day mistakes with any open-weight family. ## Which sizes see images Gemma 3 is **natively multimodal from 4B upward**. The 4B, 12B and 27B checkpoints attach a vision encoder to the language model, so a single prompt can interleave text and images — describe this screenshot, read this chart, compare these two photos. Images are converted into tokens by the vision encoder and consumed in the same sequence as the text tokens, which means images consume context budget just like words do. The **1B is text-only**. It exists for on-device and edge cases where you have a few gigabytes of memory and need text generation, not perception. ## Context windows The vision-capable sizes advertise a **128K-token** context window; the 1B is limited to a much smaller one (32K). Long context is not free at inference time — the key/value cache grows with sequence length and eats GPU memory on top of the weights — but the architectural headroom is there for document-scale prompts on 4B and above. ## How you get the weights Checkpoints are distributed through the usual open-weight channels: model hubs, and prepackaged builds for local runtimes. The same logical model therefore appears in several physical formats — full-precision safetensors for GPU serving, and quantised builds for CPU or small-GPU inference. The size you pick and the format you pick are separate decisions: a 27B in a 4-bit build can occupy less memory than a 12B at full precision. ## Choosing among them A workable default ladder: - **1B** — text-only edge and on-device work, keyword-ish tasks, ultra-low latency, no image input. - **4B** — the smallest genuinely multimodal option; fits comfortably on modest GPUs and is the usual laptop/small-server choice. - **12B** — the middle capability step; noticeably better on reasoning-flavoured tasks, still single-GPU territory when quantised. - **27B** — the family's strongest; needs quantisation or a large accelerator to fit on one card. ## What this leaf is not about The hosted Gemini API is a separate product line with its own request shape and billing; Gemma is the weights you host. Confusing the two is the classic interview stumble — a candidate who answers a Gemma question by describing an API key and an endpoint has misread the question. ## Version caution Open-weight families iterate fast, and size line-ups change between generations. Answer with the generation you are naming, say which generation you mean, and be explicit that the numbers you quote are generation-specific rather than permanent properties of "Gemma".

  • If you only need image understanding, is there any reason to reach past 4B?
    Yes. All the vision-capable sizes accept images, but perception quality and reasoning over what was seen scale with parameter count — reading a dense chart, transcribing messy handwriting, or answering multi-step questions about a screenshot degrade noticeably at 4B. Start at 4B because it is cheap to run, measure on your own images, and only move to 12B or 27B if the failures are real rather than assumed.
  • Why does the published family include base checkpoints at all if the instruction-tuned ones are more useful?
    Because base checkpoints are the right starting point for fine-tuning. An instruction-tuned model already carries a specific chat behaviour and alignment; training on top of it fights that behaviour and can undo it. If you have a domain dataset and want the model to behave your way, you start from base. If you want to prompt it as a chat assistant today, you take the instruction-tuned build.
  • Do images cost context window the way text does?
    Yes. The vision encoder turns an image into a block of tokens that occupy the same sequence as the text, so a prompt with several images has materially less room left for text and for the answer. When you budget a 128K window for a document-plus-screenshots workload, count the image tokens, not just the characters.

saying these in an interview costs you the question

  • Claiming every Gemma 3 size accepts images
  • Describing Gemma as a hosted API with an endpoint and key
  • Assuming base and instruction-tuned checkpoints behave the same
  • Treating the largest size as the only usable one
  • Saying context window is identical across all four sizes

context

open as a page

How do the Gemma Terms of Use differ from an OSI licence like Apache 2.0?

level: middleimportance: must knowfreq 62%

basics

~20 s

Gemma weights ship under Google's own Gemma Terms of Use plus a prohibited-use policy, not an OSI-approved open-source licence. Commercial use is free with no user or revenue threshold, but you accept use restrictions and must pass those terms to anyone you distribute a derivative to.

open as a page

Why can a fine-tuned Gemma 3 degrade at inference if the chat template is applied wrong?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Gemma is trained on an exact turn format — start_of_turn markers with user and model roles, closed by end_of_turn, and a single leading BOS token. Serving prompts that differ from the training format, or that duplicate BOS, puts the model off-distribution and quality drops silently.

open as a page

Which Gemma 3 size fits a single 24 GB GPU, and what do you trade away?

level: seniorimportance: should knowfreq 45%

basics

~20 s

At bfloat16 you need roughly 2 GB per billion parameters, so 24 GB holds 4B or 12B but not 27B's ~54 GB. Google's quantisation-aware-trained int4 checkpoints shrink 27B to roughly 15 GB, which fits — at the cost of some quality and nearly all your KV-cache headroom.

open as a page

How should Gemma's licence terms shape choosing it over an Apache-2.0 open-weight model?

level: principalimportance: should knowfreq 30%

basics

~20 s

Decide by how you deliver the model. If you serve it yourself, Gemma's terms cost you almost nothing. If you ship weights inside a product, its use restrictions and pass-along duties follow every copy, while an Apache-2.0 model imposes only attribution and carries a patent grant.

open as a page