Why does masked-image pretraining mask around 75% of patches when masked text masks only 15%?
answer
- difficulty is the real design knob
- how redundant is the signal?
- neighbouring patches predict each other
- tokens cannot be interpolated
- too high leaves the target undetermined
basics
~20 sImages are highly redundant, so a lightly masked patch can be interpolated from its neighbours and the pretext teaches nothing. Text tokens are far denser in information, so hiding even a small fraction already forces real inference about meaning and structure.
solid answer
~50 sThe masking ratio is how you set the difficulty of the pretext, and the right value depends on how redundant the signal is. Natural images are spatially smooth: at 15% masking nearly every hidden patch is ringed by visible ones, so a network can win by copying local texture and never learns anything about objects. Raising the ratio to roughly three quarters destroys that local escape route — the model has to infer missing regions from distant context and from what the scene plausibly contains. Text is the opposite: tokens are discrete and information-dense, so masking 15% already leaves most predictions genuinely hard, and masking most of a sentence would leave the target underdetermined. The general rule is: mask enough that no low-level, local rule can solve the task, but not so much that the target becomes unguessable noise.
go deeper
Recall that masked pretraining hides part of the input and trains the network to restore it, and that images are masked far more aggressively than text. Be able to say why: neighbouring pixels give the answer away, words do not.
Explain the mechanics: redundancy and spatial autocorrelation make low-ratio image masking solvable by local interpolation, while dense discrete tokens make a small text ratio already hard. Be ready to describe what breaks at each extreme of the ratio.
Show that you would set ratio and mask unit together for a new modality, and that you would validate the choice on downstream results rather than pretext loss. Be ready to diagnose a pretext whose loss looks great but whose features transfer no better than random weights.
Own the framing that pretext difficulty is a design surface with a compute budget attached: a high ratio can make each step cheaper as well as harder. Be ready to argue how much pretraining compute a ratio sweep deserves before it stops paying back.
## What the masking ratio is actually controlling Masked reconstruction is a *pretext*: you deliberately damage the input, ask the network to restore what you removed, and afterwards keep the encoder for a real downstream task. No labels are involved, so the only thing that makes the exercise worthwhile is that **solving the pretext requires the structure you want the encoder to internalise**. That makes the difficulty of the task the central design parameter, and the masking ratio is the coarsest knob for setting it. A pretext that is too easy is not merely slow to learn from — it is actively useless. The loss goes down, the curves look healthy, and the encoder has learned an interpolation rule. ## Why images need a high ratio Natural images are enormously redundant. Pixel values are strongly autocorrelated in space: neighbouring regions share colour, texture and lighting, and edges continue predictably. Hide one small patch out of many and its contents are close to determined by the ring of visible pixels around it. A convolutional or patch-based model can then reach a low reconstruction error with a purely local smoothing-and-texture-continuation strategy that carries no knowledge of what object is in the picture. Push the ratio up to roughly three quarters and that escape route closes. Most hidden patches no longer touch a visible one, so local continuation has nothing to continue from. To fill a large hole the model has to use long-range context and a coarse notion of what the scene contains — the beginnings of the semantics you actually want. A useful concrete case: MAE-style pretraining on a large pile of unlabelled Sentinel-2 satellite tiles. Field boundaries, cloud texture and coastline are exactly the kind of smooth, repetitive content that low masking makes trivial, so the high ratio is not a tuning accident, it is what makes the corpus teach anything at all. It also has a pleasant side effect: when the encoder only processes the visible fraction, each pretraining step is much cheaper, so a high ratio buys both a harder task and more epochs per unit of compute. ## Why text needs a low one Tokens are discrete symbols chosen from a large vocabulary, and each one carries far more information than a small image patch. There is no interpolation shortcut: you cannot average two neighbouring words. At a 15% masking rate, a typical prediction still demands syntax, agreement, world knowledge and sometimes long-range reference. Raising the rate towards image levels removes so much of the sentence that the target stops being determined by what remains — the model is asked to guess among many equally valid completions, and the gradient it receives is dominated by irreducible ambiguity rather than by learnable structure. So the number itself is not the lesson. The transferable rule is: **mask enough that no cheap local rule reconstructs the target, and not so much that the target is no longer inferable from what is left.** ## Mask unit matters as much as mask fraction Ratio is only half the design. What you mask as a unit decides whether the ratio means anything. Consider unlabelled vibration or sensor traces from industrial machines. If you mask individual timesteps, even at a high rate, linear interpolation between the surviving samples reconstructs most of the signal, because the trace is smooth at the sampling scale. The fix is not a higher rate but a bigger unit: mask contiguous **spans** long enough to cover several cycles of the dominant frequency, so that restoring them requires knowing the periodic structure and the operating regime rather than the local slope. The same logic explains why block-shaped image masks are harder than the same fraction scattered as isolated pixels, and why a span-based scheme is the right default for any signal that is smooth in its own domain. ## Choosing the value in practice You cannot pick the ratio by looking at the pretext loss. A harder task has a higher loss by construction, so losses at different ratios are not comparable and the lowest one is usually the most useless. The ratio has to be chosen against downstream behaviour: pretrain a few short runs at different ratios, fine-tune each on the small amount of labelled data you have for the real task, and compare there. Expect a broad plateau rather than a sharp optimum — the failure modes are at the ends, not in the middle. Diagnostically, if raising the ratio substantially does not change downstream results at all, that usually means something else is capping the representation, not that the ratio is irrelevant.
- You are pretraining on unlabelled vibration traces from factory machines. How would you set up the masking?Mask contiguous spans, not individual timesteps. A smooth trace lets linear interpolation recover isolated missing samples at almost any rate, so the pretext stays trivial. Choose a span long enough to cover several cycles of the dominant frequency, so restoring it needs the periodic structure and the machine's operating regime. Then set the fraction of the trace covered by spans the way you would set an image masking ratio.
- What goes wrong if you push the image masking ratio close to 100%?Too little context survives for the target to be determined, so the model hedges and predicts something like the conditional average — blurry, low-information output. The gradient is then dominated by irreducible uncertainty rather than by learnable structure, and sample efficiency drops. You are no longer making the task harder in a useful way, you are making it unanswerable.
- Can you compare masking ratios by their pretext reconstruction losses?No. A higher ratio removes more information, so its loss is higher by construction; the losses live on different tasks and are not comparable. The lowest pretext loss usually belongs to the most trivial setting. Ratios must be compared by downstream results after fine-tuning on the real task with whatever labelled data you have.
Blanking out one word of a sentence versus one brick of a brick wall: the wall you can fill in from its neighbours without understanding anything, so you have to knock out most of it before the task means something.
saying these in an interview costs you the question
- Says 75% is a universal constant for all masked pretraining
- Picks the masking ratio by lowest reconstruction loss
- Thinks a higher ratio always yields a better representation
- Ignores mask unit and only tunes the fraction
- Claims images need more masking because they have more pixels to average