skip to content

In a vision-language model, what do the encoder, connector and decoder each do?

level: middleimportance: must knowfreq 68%

answer

  1. Three boxes in series
  2. Pretrained backbone plus pretrained LLM
  3. Something must translate between their spaces
  4. Encoder frozen, connector trained first
  5. SigLIP-class ViT, projector, text decoder

basics

~20 s

Three parts: a vision encoder (a Vision Transformer, usually SigLIP-class) turns the image into patch embeddings, a connector maps those embeddings into the decoder's embedding space, and the language decoder consumes them alongside ordinary text tokens.

solid answer

~50 s

Nearly every VLM shipped since the LLaVA line follows the same stack. The **vision encoder** is a Vision Transformer pretrained on image-text pairs — SigLIP and SigLIP 2 are the mid-2026 default, having displaced the original CLIP encoders — and it cuts the image into fixed-size patches, emitting one embedding per patch. Those embeddings live in the encoder's own representation space, so a **connector** (an MLP projection, a Q-Former resampler, or inserted cross-attention layers) maps them into vectors the decoder can read as if they were token embeddings. The **decoder** is an ordinary autoregressive text LLM; it attends over image and text vectors together and generates text. Training is staged: freeze encoder and decoder, train only the connector on image-caption pairs, then unfreeze the decoder for visual instruction tuning. That staging is why a small team can build a satellite-imagery VLM without training a vision backbone from scratch.

go deeper

for a junior

Be able to name the three parts in order and say that the image never reaches the language model as pixels — it arrives as vectors produced by a separate vision model.

for a middle

Explain what each box outputs: patch embeddings from the encoder, decoder-space vectors from the connector, text from the decoder. Know that the connector is the piece trained first, with the encoder frozen.

for a senior

Attribute real symptoms to a component — lost fine detail to encoder resolution, cost to patches-per-image, spatial-reasoning gaps to the encoder's pretraining objective — and explain why an encoder swap forces connector retraining.

for a principal

Own the build-versus-buy framing: which components you reuse, where domain adaptation actually pays for imagery your provider has never seen, and what a staged training budget buys you against simply prompting a frontier VLM.

## The canonical stack A vision-language model (VLM) is a text language model that has been given a way to see. In 2026 the dominant architecture is still three components in series: **image → vision encoder → connector → language decoder → text** This is sometimes called "late fusion" or, less kindly, "bolted-on vision", in contrast to models pretrained on interleaved image and text data from the beginning ("natively multimodal" / early-fusion). Even the natively multimodal frontier models still have something that plays the encoder's role; the difference is when in training the modalities were mixed, not that one of the three jobs disappears. ## The vision encoder The encoder is a Vision Transformer (ViT). It splits the image into a grid of non-overlapping square patches (14x14 or 16x16 pixels are typical), linearly embeds each patch, adds positional information, and runs transformer blocks over that sequence. The output is one vector per patch — a spatially ordered set of features, not a single image vector. Crucially, the encoder is *pretrained separately*, usually with a contrastive image-text objective. The original CLIP encoders held this slot for years; SigLIP and then SigLIP 2 replaced them in almost every new open-weights VLM because the sigmoid pairwise loss scales more conveniently and SigLIP 2 adds captioning-style and self-distillation objectives that produce better dense (per-patch) features, plus multilingual text and native-aspect-ratio variants. A better encoder is the cheapest upgrade path for a VLM, which is why encoder swaps are common. ## The connector The encoder's patch vectors have the wrong dimensionality and the wrong semantics for the decoder's input embedding table. The connector — also called the projector, adapter, or resampler — bridges the two. Three families exist: - **Projection / MLP** (LLaVA, Pixtral, NVLM lineage): a one- or two-layer MLP applied per patch. Simple, cheap to train, preserves one token per patch. This family dominates. - **Query-based resampling** (BLIP-2's Q-Former, Flamingo's Perceiver Resampler): a small transformer with a fixed set of learned queries — 32 in BLIP-2 — that cross-attends to the patches and outputs a fixed-length sequence regardless of input size. - **Cross-attention injection** (Flamingo, and later models such as Llama 3.2 Vision): image features never enter the decoder's input sequence; extra gated cross-attention layers inside the decoder attend to them. The connector is the part that is almost always trained, because it is what ties this particular encoder to this particular decoder. Swap either end and it must be retrained. ## The decoder The decoder is a normal text LLM — the same architecture, tokenizer and weights family you would use without images. After the connector runs, image vectors sit in the same sequence as text embeddings (or are attended to from inside, in the cross-attention design), and generation proceeds autoregressively as usual. Nothing about the decoder is inherently visual; its visual competence is inherited from what the connector hands it and from instruction tuning on multimodal data. ## What is trained, and when A typical two-stage recipe, using an in-house satellite and drone imagery VLM as the example: 1. **Alignment / connector pretraining.** Encoder frozen, decoder frozen, connector trained on a large set of (image, caption) pairs — here, tiles paired with analyst captions. Cheap: only a few million parameters move. 2. **Visual instruction tuning.** Connector plus decoder trained (often with LoRA on the decoder) on task-shaped conversations — "how many aircraft are on this apron?", "describe the damage in this frame". The encoder often stays frozen, or only its top blocks unfreeze late, because a frozen encoder is more stable and much cheaper. That is the practical payoff of the three-part split: the expensive, general-purpose parts (a pretrained vision backbone and a pretrained LLM) are reused, and the domain-specific work concentrates in the connector and the instruction data. ## Why the split matters in an interview The architecture explains most VLM behaviour you will be asked about. Fine detail is lost because the encoder resolution and patch size decide what survives. Images are expensive because the connector emits roughly one token per patch. Spatial-reasoning weakness traces back to a contrastively pretrained encoder that was never rewarded for relations. And an encoder swap that "should be a drop-in" degrades output until the connector is retrained. Being able to walk the three boxes and attribute a symptom to one of them is the answer interviewers are listening for.

  • Which of the three components is usually frozen during visual instruction tuning, and why?
    The vision encoder, most often. It was pretrained on far more image-text data than any downstream fine-tune provides, unfreezing it destabilizes the features the connector was aligned to, and it is the most expensive part to train. Many recipes unfreeze only its top blocks, late, with a low learning rate. The connector is always trained; the decoder is usually trained or LoRA-adapted in the instruction stage.
  • What do people mean when they call a model "natively multimodal" rather than a bolted-on VLM?
    That images were part of pretraining from early on — interleaved image-text data with a shared objective — rather than a vision encoder being attached to an already-finished text LLM and aligned afterwards. Early fusion tends to give better cross-modal grounding and lets the model handle interleaved inputs naturally. The three functional roles still exist; the difference is when the modalities were mixed and how jointly the parameters were learned.
  • Your team swaps the vision encoder for a stronger one and output quality collapses. What happened?
    The connector was trained to map the *old* encoder's feature space into the decoder's embedding space. A new encoder produces different features, a possibly different hidden size, and often a different patch count, so the projection is now meaningless. The fix is to redo the alignment stage — retrain the connector against the new encoder — before any instruction tuning; an encoder swap is never a drop-in replacement.

saying these in an interview costs you the question

  • Claiming the LLM itself sees pixels directly
  • Saying the connector performs OCR or captioning
  • Assuming the vision encoder is trained end-to-end from scratch
  • Thinking an encoder swap needs no connector retraining
  • Describing the decoder as a special vision-specific architecture

context