Why did VLMs move from fixed 336x336 resizing to native-resolution tiling?
answer
- What is smaller than a patch is gone
- Squash a big frame, lose the small objects
- Tiles plus one thumbnail
- Aspect ratio matters too
- Detail bought with quadratic token cost
basics
~20 sSquashing a large image into one small square destroys anything smaller than a patch — small objects, fine text, chart labels. Tiling encodes the image as several full-resolution tiles plus a downscaled global view, preserving detail at the cost of far more tokens.
solid answer
~50 sEarly VLMs inherited their encoder's fixed input — CLIP ViT-L/14-336 meant every image, whatever its size or aspect ratio, was squashed to 336x336 and became 576 patches. On a 4000x3000 aerial frame that is roughly a twelvefold linear downscale: a vehicle a few pixels across after resizing falls below one patch and is simply not represented. **Dynamic or native-resolution processing** fixes this by splitting the image into encoder-sized tiles, encoding each at full resolution, and adding one downscaled view of the whole image so global layout is not lost. Newer encoders go further and accept variable resolutions and aspect ratios directly, using 2D rotary position encodings and block-diagonal attention masks so differently shaped images can be packed together. The cost is blunt: tokens scale with area, so a naive tiling of a large frame can run into tens of thousands of tokens, which is why implementations cap the tile count and pool patches.
code
python · 8 linesdef tiled_tokens(width, height, tile=448, patch=14, global_view=True):
cols = -(-width // tile) # ceiling division
rows = -(-height // tile)
per_tile = (tile // patch) ** 2
return (cols * rows + (1 if global_view else 0)) * per_tile
print(tiled_tokens(4000, 3000)) # 65536 - naive, uncapped
print(tiled_tokens(4000, 3000, tile=1024)) # far fewer tiles, coarser detailgo deeper
Know that images used to be squashed to one small fixed square and that modern models split large images into tiles instead, so fine detail survives.
Explain the mechanism: tiles at native resolution plus a downscaled global view, why the global view is needed, and how learned position embeddings limited the older fixed-input encoders.
Reason about the cost curve under a real budget — cap tile counts, crop to regions of interest, pool patches, and pick resolution from the smallest feature that must be resolved rather than from the source file.
Own the policy across a workload: which classes of image get cheap low-resolution passes, which escalate to full tiling, and how that tiering shapes cost, prefill latency and context pressure at production volume.
## The constraint that created the problem A Vision Transformer is trained at a specific input resolution with learned position embeddings for a specific patch grid. Early VLMs used a stock contrastive encoder — CLIP ViT-L/14 at 336x336 was the canonical choice — so the pipeline was: take any image, resize (and often centre-crop) to 336x336, get a 24x24 grid, 576 patches, done. Fixed input, fixed cost, no configuration. That was fine while VLMs were being evaluated on natural-image captioning, where the subject is large and centred. It fell apart the moment the workload became documents, screenshots, charts, and inspection imagery. ## Why fixed resize destroys information Consider a 4000x3000 aerial frame. Forcing it to 336 on the long side is about a 12x linear downscale, 144x in area. A vehicle 40 pixels long in the original is roughly 3 pixels after resizing — smaller than a single 14-pixel patch. The encoder has no representation slot for it. No prompt recovers it, no decoder reasoning recovers it: the information was destroyed before the model saw anything. The same arithmetic explains every classic complaint of that era. Small text in a screenshot, axis labels on a chart, a serial number on a component, hairline rules in a table — all sub-patch after aggressive downscaling. And aspect ratio is a second insult: squashing a wide panorama or a tall document page to a square distorts geometry, while centre-cropping simply throws away the sides. ## Tiling: the first fix Dynamic-resolution or "any-resolution" tiling keeps the encoder unchanged and changes what you feed it: 1. Choose a tile grid whose aspect ratio approximates the image's (a 4:3 frame might become 4x3 tiles). 2. Split the image into tiles of the encoder's native size and encode each one separately at full resolution. 3. Additionally encode one **downscaled view of the entire image** — the "thumbnail" or global view. 4. Concatenate all tile tokens plus the global tokens and hand them to the connector. The global view matters more than it looks. Individual tiles are context-free: a tile containing half a wing has no idea it belongs to an aircraft, and no tile knows the page layout. The thumbnail restores that. ## Native resolution: the second fix Tiling is a workaround around a fixed-input encoder. Newer encoders remove the constraint instead. Two ingredients make variable input practical: - **2D rotary position encodings** in place of learned absolute position embeddings, so any grid shape has well-defined positions without interpolating a table trained for one grid. - **Sequence packing with block-diagonal attention masks**, so images of different shapes and token counts can be batched together while each image only attends within itself. SigLIP 2 shipped native-aspect-ratio variants on this principle, and models in the Qwen-VL line built dynamic-resolution handling into the encoder. The practical effect is that image geometry is preserved and the model chooses how many tokens an image deserves. ## The price, stated plainly Tokens follow area. Take a naive tiling of a 4000x3000 image into 448-pixel tiles at patch size 14: that is 9x7 = 63 tiles, each 32x32 = 1024 patches, plus a global view — over 64,000 tokens for one photograph. That is more than most text prompts by orders of magnitude, and prefill latency scales with it. So real systems layer controls on top: - **Cap the tile count** (a maximum of roughly a dozen tiles is typical) and choose the grid that best fits within the cap. - **Pool or merge patches** — a 2x2 merge cuts tokens fourfold — accepting some detail loss where the task tolerates it. - **Crop before tiling.** If you know the region of interest, tiling the whole frame is waste. - **Match resolution to the smallest feature you must resolve**, rather than defaulting to the original file. ## How the field moved past manual tiling The newest development is that frontier reasoning models do the zooming themselves: during the reasoning trace they crop, zoom and re-inspect a region as an action, rather than requiring the caller to pre-tile. That converts a static preprocessing decision into an adaptive one — full detail is spent only where the model finds it needs it. Manual tiling remains the right tool when you control the pipeline and know in advance which regions matter. ## The judgement an interviewer is listening for Not "tiling is better". The answer they want is the trade: fixed resize buys constant, low cost and loses sub-patch detail irretrievably; tiling and native resolution buy detail at a token cost that grows with area; and the engineering work is choosing where on that curve each task sits, then enforcing it with tile caps, crops and pooling rather than sending whatever the camera produced.
- Why include a downscaled view of the whole image alongside the tiles?Because a tile has no context. Encoded alone, a tile holding part of a runway cannot tell that it is a runway, and no tile carries the overall layout or the relationships between regions. The global view is cheap — one extra tile's worth of tokens — and restores scene-level structure so the decoder can situate what each detailed tile shows. Dropping it is a common cause of locally accurate but globally confused answers.
- What breaks if you feed a fixed-resolution ViT an image at a resolution it was not trained on?Its learned absolute position embeddings were fitted to one specific patch grid, so a different grid has no defined positions; the usual workaround is interpolating that embedding table, which works within a modest range and degrades beyond it. Encoders designed for variable input avoid the issue by using 2D rotary position encodings, which extend to arbitrary grid shapes, together with block-diagonal masks so differently shaped images can still be batched.
- When is a fixed low-resolution resize still the right choice?Whenever the decision depends on gist rather than detail: scene or content classification, coarse routing, moderation triage, duplicate detection, or a first pass that decides which images deserve expensive treatment. A few hundred tokens per image at low resolution keeps throughput and cost sane at volume. Escalate to tiling only for the subset where the task genuinely turns on small structures, text or counts.
saying these in an interview costs you the question
- Assuming higher resolution always improves accuracy
- Forgetting that tiles alone lose global context
- Believing tiling is free or only linearly more expensive
- Thinking a stronger prompt can recover sub-patch detail
- Treating aspect-ratio distortion as harmless