skip to content

When would you choose a Q-Former connector over a simple MLP projection in a VLM?

level: seniorimportance: should knowfreq 38%

answer

  1. It is a compression decision
  2. Fixed output length versus one-per-patch
  3. What you drop, you cannot recover
  4. 32 learned queries versus 4096 patches
  5. Inference-time dial beats baked-in bottleneck

basics

~20 s

Choose query-based resampling when image tokens must be bounded — a Q-Former compresses any patch count into a fixed number of learned queries, capping context cost at the price of discarded detail. MLP projection keeps one token per patch and preserves detail.

solid answer

~50 s

The connector is a **compression decision**. An MLP projection (LLaVA-style) maps each patch embedding into the decoder's space one-for-one: it preserves spatial detail, is trivial to train, and makes image cost scale with image area. A Q-Former (BLIP-2-style) is a small transformer whose fixed set of learned queries — 32 in the original — cross-attends to the patches and emits that fixed number of vectors regardless of input size, so cost is constant but everything the queries did not capture is gone. Gated cross-attention (Flamingo-style) is a third family: image features never join the input sequence at all, and inserted attention layers reach into them, which keeps the text sequence short but changes the decoder's architecture and weights. In mid-2026 projection dominates, because context windows grew, detail-heavy tasks like charts and documents punish compression, and simplicity won. Reach for resampling when many images per request or a hard token budget binds.

go deeper

for a junior

Recognise that more than one kind of connector exists and that the simple one is a small MLP applied per patch. You are not expected to compare families in depth.

for a middle

Describe each family's output shape: one vector per patch for projection, a fixed count for a resampler, none in the input sequence for cross-attention. Tie output length directly to context cost.

for a senior

Make the choice under a stated constraint — token budget, detail sensitivity, training compute, decoder swappability — and explain why per-patch projection plus inference-time tiling has displaced fixed-length resampling.

for a principal

Frame it as where you want the compression knob to live: baked into trained weights or exposed per request. Own the downstream consequence for model upgrades, serving cost curves and how many images a request can carry.

## What a connector must accomplish Between a vision encoder and a language decoder there is a mismatch of dimensionality, of representation space, and of sequence length. The connector resolves all three. The first two are mechanical. The third — how many vectors the decoder ends up paying attention to — is the real design decision, and it is a compression trade: how much of the image's spatial information do you keep, and what does keeping it cost in context? ## Family 1: projection / MLP A one- or two-layer MLP is applied to every patch embedding independently, mapping it into the decoder's embedding dimension. The projected vectors are then prepended (or interleaved) into the decoder's input sequence as if they were token embeddings. - **Sequence length:** one token per patch (or per pooled patch group). - **Parameters:** small — a few million. - **Training:** the alignment stage trains only this, with encoder and decoder frozen. Fast and stable. - **Detail:** maximal. Every patch has its own slot, so the decoder's attention can single out any region. - **Cost:** grows with image area, and with tiling it grows fast. This family — LLaVA, Pixtral, NVLM and most open-weights VLMs — is the mid-2026 default. ## Family 2: query-based resampling (Q-Former, Perceiver Resampler) A small transformer holds N learned query vectors (BLIP-2 used 32). They cross-attend to the encoder's patch outputs and produce exactly N output vectors, which are then projected into the decoder. - **Sequence length:** fixed at N, independent of image size or tile count. - **Parameters:** more than an MLP; the Q-Former is a real transformer with its own pretraining recipe. - **Training:** harder. It has more to learn and is known to be finicky; BLIP-2 used a multi-stage objective to train it. - **Detail:** bounded. Anything the N queries did not attend to is unrecoverable downstream. Dense documents, small text and fine chart labels suffer most. - **Cost:** flat, which is the entire point. ## Family 3: gated cross-attention injection Flamingo inserted new gated cross-attention layers into a frozen decoder; the image features are attended to from inside those layers rather than sitting in the input sequence. Later models such as Llama 3.2 Vision used the same idea. - **Sequence length:** the text sequence is untouched — no image tokens at all in the input. - **Parameters:** substantial; the inserted layers are part of the decoder. - **Training:** you are modifying the decoder, so it is a heavier change and the decoder is no longer a stock text LLM. - **Detail:** can be good, since cross-attention reaches the full patch set every layer. - **Cost:** shifted from context length into per-layer compute, and into architectural coupling. ## How to actually decide Ask four questions: 1. **Does a hard token budget bind?** Many images per request — a batch of drone frames, a multi-page comparison — pushes toward fixed-length resampling. One image per request, with a large context, does not. 2. **How detail-dependent is the task?** Reading dense annotations, small labels, or fine structure argues strongly for one-token-per-patch. Scene-level classification or captioning tolerates heavy compression. 3. **How much training capacity do you have?** An MLP connector aligns in a fraction of the compute a Q-Former needs, and it is far more forgiving. For a small team this often settles it. 4. **Do you want to keep the decoder stock?** Projection and resampling leave the decoder unmodified, so you can swap in a newer LLM and redo only the alignment stage. Cross-attention injection welds vision into the decoder's weights and makes that swap a rebuild. ## Why projection won Three things changed between the BLIP-2 era and now. Context windows grew by orders of magnitude, so a few thousand image tokens stopped being disqualifying. The tasks that matter commercially — documents, charts, screenshots, inspection imagery — are exactly the detail-heavy ones that compression damages. And high-resolution tiling gave a better lever for controlling cost than architectural bottlenecking: you can choose how many tiles to spend on a given image at inference time, whereas a Q-Former's bottleneck is baked in at training time. That is the strongest form of the answer: the resampler's fixed budget is a *training-time* decision, while tiling and pooling give you an *inference-time* dial over the same trade. Prefer the dial you can turn per request. ## Honest caveats This is not a settled hierarchy so much as a current consensus. Query-based compression remains attractive for video and for many-image contexts, where per-patch tokens are simply unaffordable, and hybrid designs (pool aggressively, then project) blur the families. What an interviewer wants is the axis — fixed-length compression versus per-patch fidelity, and what each costs — not a memorized ranking.

  • What does the Flamingo-style cross-attention connector buy you that neither projection nor resampling does?
    It keeps image features entirely out of the decoder's input sequence, so text sequence length is unaffected no matter how many images you condition on, and the inserted layers can attend to the full patch set at every depth. The price is architectural: you are adding trained layers inside the decoder, so it is no longer a stock LLM you can cheaply swap for a newer one, and the parameter and compute cost is much higher than an MLP.
  • If detail loss is the Q-Former's weakness, why did BLIP-2 use one?
    Because in that era the binding constraint was the opposite of today's. Decoders had short context windows, the decoder was kept entirely frozen, and pushing hundreds of per-patch tokens into it was not viable. A fixed 32-query bottleneck made the whole approach affordable. As context windows grew and detail-heavy document and chart tasks became the commercial centre of gravity, the trade flipped toward per-patch projection.
  • How would you empirically test whether your connector is the bottleneck rather than the encoder?
    Hold the encoder and decoder fixed and vary only the connector's output length — for instance compare per-patch projection against the same stack with 2x2 pooling and with aggressive resampling — on a task suite that separates coarse understanding from fine detail. If coarse scores hold while fine-detail scores fall as output length shrinks, the connector's compression is the bottleneck. If both stay flat, look upstream at input resolution and patch size.

saying these in an interview costs you the question

  • Calling the connector just a resize layer with no trade-offs
  • Assuming a fixed query count loses nothing important
  • Believing more connector parameters always means better quality
  • Thinking cross-attention connectors leave the decoder unmodified
  • Treating the choice as settled rather than task-dependent

context