In contrastive self-supervision, how does the augmentation set decide what the encoder ignores?
answer
- the positive pair defines the problem
- you discard whatever the transform changes
- ask whether it can change the label
- hue is diagnostic in dermatology
basics
~20 sA positive pair is one input under two augmentations, so the encoder learns to discard whatever those augmentations change. Colour jitter buys hue invariance on street photos and ruins a dermatology encoder, where hue is the diagnostic signal.
solid answer
~50 sThe objective says one thing: two views of the same input must land in the same place. So every transform in the augmentation set declares the difference it creates irrelevant, and the encoder takes that literally — it learns to throw the information away. Random cropping teaches invariance to scale, position and occlusion; flipping teaches left-right invariance; colour distortion teaches hue and saturation invariance. The test before adding a transform is: **can it ever change the answer to my downstream task?** Recolouring a street photo does not change that a car is a car, so colour jitter is free capacity there. On dermatology images hue and saturation *are* the lesion signal, so an encoder trained to ignore them has been trained to discard the diagnosis. Strength matters too: too weak and matching is trivial, too strong and the two crops no longer share content.
go deeper
Know that a positive pair comes from two random augmentations of the same input, and be able to name a couple of common transforms such as random cropping and colour jitter.
Explain per transform what invariance it induces, and why cropping combined with colour distortion is stronger than either alone. Be able to say why an easy positive pair produces a shallow representation.
Demonstrate the audit: how you screened transforms against a specific downstream task, an ablation you ran, and a case where the standard recipe was wrong for your domain. Concrete examples beat principles here.
Own the fact that one encoder's augmentation set commits every downstream consumer to the same invariances. Be ready to argue whether one shared encoder or several domain-specific ones is the right call for the organisation.
## The invariance contract Contrastive self-supervision has one instruction: bring the two views of an input together, push other inputs apart. There is no other source of semantics. So the augmentation pipeline is not a data-multiplier or a regulariser here — it is the *specification of the learning problem*. Whatever a transform varies, the encoder is being paid to become blind to. Whatever no transform varies, the encoder is free to keep, and will keep, because keeping it helps discriminate against negatives. A useful phrasing: the augmentation set is a list of differences you have declared irrelevant, signed on behalf of every downstream task that will ever use this encoder. ## What the standard transforms actually teach - **Random resized cropping** is the workhorse. It teaches invariance to scale, translation and partial occlusion, and it creates a second, subtler signal: two crops often show *different parts* of the same object, so matching them forces the encoder to represent that those parts co-occur. - **Horizontal flip** teaches left-right invariance. Free for most object recognition; wrong for text, for chirality-sensitive medical views, and for anything where handedness is the label. - **Colour distortion** (jitter of brightness, contrast, saturation and hue, sometimes grayscale conversion) teaches invariance to illumination and to colour statistics. It is unusually important with cropping, because two crops of the same photo share a colour histogram that would otherwise let the encoder match them on low-level statistics alone rather than on content. - **Blur** teaches invariance to focus and to high-frequency texture. - **Rotation** teaches orientation invariance. Sometimes exactly right (microscopy, aerial imagery, where there is no canonical up), sometimes destructive (digits, road-scene understanding). ## The failure that costs a project Take a dermatology encoder pretrained on unlabelled lesion photographs. The default recipe — crop, flip, strong colour jitter — is copied from natural-image work. But the clinical signal in a pigmented lesion lives substantially in hue and saturation: colour variegation, redness, the difference between a brown and a blue-grey area. Strong colour jitter tells the encoder that a lesion recoloured from brown to blue-grey is the *same thing*. Pretraining converges beautifully; the loss curve looks like every published one; and the resulting representation has been explicitly optimised to erase the feature the downstream classifier needs. The same transform on street photography is not merely harmless but valuable, because it stops the encoder from matching crops by their shared colour cast. The transform did not change. The *domain's* answer to 'does this change the label?' changed. ## Choosing an augmentation set 1. **Enumerate the downstream tasks** the encoder is meant to serve, or the label-invariances you are confident about across all of them. 2. **Label-invariance test.** For each candidate transform, ask whether an expert's answer could change under it. If yes, drop it, or weaken its magnitude to a range where the answer provably does not change. 3. **Ablate.** Pretrain short runs with and without each questionable transform and compare a small labelled probe. Augmentation choice usually produces larger downstream differences than architecture choice at this stage. 4. **Calibrate strength.** Too weak: views nearly identical, the objective is satisfied by shallow statistics, downstream gains are small. Too strong: an aggressive crop of a large image may contain none of the object in the other view, so you are training the encoder to unify genuinely unrelated content, which injects noise resembling the false-negative problem from the other direction. ## Beyond images The same reasoning transfers, and the transforms do not. For sensor time series you have amplitude scaling, additive jitter, time warping, permutation of segments and random cropping in time — and the same audit applies. Time warping declares that the *speed* of an activity is irrelevant, which is right if you want to recognise 'walking' regardless of pace and wrong if cadence is the thing being measured. For text and graphs the transform vocabulary is different again, and in every case the question is unchanged: what am I promising is irrelevant? ## How this shows up in an interview The weak answer is 'more augmentation is better' or 'we used the standard recipe'. The strong answer names the invariance each transform buys, names the downstream task it must not destroy, and describes an ablation that decided a specific case. Interviewers ask this because it separates people who ran a pretraining script from people who designed one.
- How would you decide whether a specific augmentation is safe before committing to a full pretraining run?Two checks. First a domain check: show augmented examples to someone who can label them and confirm the label never changes under the transform at the magnitude you plan to use. Then an empirical one: short pretraining runs with and without it, compared on a small labelled probe set. Augmentation ablations at reduced scale are usually predictive enough to make the call cheaply.
- What is the downside of augmentations that are too weak?The two views end up nearly identical, so matching them is easy without understanding content and the encoder settles for low-level statistics that happen to be shared. The loss falls fast and the representation is shallow. Difficulty in the positive pair is what forces semantic features, which is why aggressive cropping combined with colour distortion outperforms either alone.
- Does random cropping teach anything beyond scale and position invariance?Yes. Two crops frequently cover different regions of the same input, so satisfying the objective requires representing that those regions belong together — a part-to-whole co-occurrence signal rather than pure invariance. That is a large part of why cropping is the single most valuable transform in image contrastive pretraining.
The augmentation list is a set of instructions saying 'treat these two things as the same'. The encoder obeys literally, including when you accidentally tell it that a brown lesion and a blue-grey one are the same.
saying these in an interview costs you the question
- Reuses a natural-image augmentation recipe on medical images
- Treats augmentations as mere data multiplication
- Says stronger augmentation is always better
- Assumes the encoder learns which transforms to ignore by itself
- Cannot name the invariance a given transform buys