Why do features from a supervised ImageNet backbone transfer to a task with different classes?
answer
- features are a hierarchy, not a lookup
- edges and textures come first
- the labels only shape the top
- specificity increases with depth
- generic early, source-specific late
basics
~10 sEarly layers learn generic edges, colours and textures that nearly any image task needs, and only the deepest layers specialise to the source categories. Transfer keeps the generic stack and replaces the specialised end.
solid answer
~50 sSupervised pretraining on a large labelled corpus forces the network to build a feature hierarchy, and only part of that hierarchy is really about the source labels. The first layers converge to edge, colour-blob and orientation detectors, the middle layers to textures, motifs and object parts, and only the last block or two encode the combinations that separate the source's 1000 web categories in particular. A 12-class retail-shelf classifier still needs edges, textures and part detectors, so those layers stay useful even though not one of its classes appears in the source label set. What does not transfer is the mapping from parts to the source's specific categories, which is why the classifier layer is thrown away. The practical rule that follows is that transferability decays with depth: the deeper the layer, the more source-specific its content, and the stronger the case for retraining it on the target.
go deeper
Be ready to say in one breath that early layers learn general-purpose edges and textures, later layers learn the source task's own categories, and transfer keeps the first part and replaces the last.
An interviewer expects you to explain the hierarchy layer by layer and to state that transferability falls off with depth, including why disjoint source and target label sets are not an obstacle.
Show you use the gradient of specificity operationally: how far the target domain sits from natural images and how big the target set is should visibly drive how much of the stack you keep versus retrain.
Own the framing that a pretrained backbone is a paid-for representation and a variance reducer. Be able to argue when that investment stops paying — unusual modalities, plentiful target labels — and what the organisation should standardise on.
## The question behind the question Supervised pretraining means training a network end to end on a large labelled corpus for a task you do not actually care about — say a 1000-way web-image classification problem with on the order of a million labelled images — and then reusing the trained weights as the starting point for a target task with a completely different, usually much smaller, labelled dataset. The puzzle an interviewer is probing is: if the source labels are irrelevant to your target, why is the source model worth anything at all? ## A network is a feature hierarchy, not a lookup table A trained deep network does not store the training images. It stores a stack of transformations, each of which converts the representation below it into something a little more abstract. When you inspect what the layers of a supervised image classifier respond to, a consistent pattern appears across architectures, initialisations and even across corpora: - **First layers**: oriented edges, colour blobs, simple frequency and orientation filters — essentially Gabor-like detectors. These emerge whether the source task is 1000 web categories, plant species or scene types. They are a property of natural images, not of the labels. - **Middle layers**: textures, corners, repeated motifs, then object parts — wheels, eyes, text-like strokes, fabric weave. - **Late layers**: configurations of parts that discriminate the *source* classes, and finally a linear classifier over those configurations. The labels drive this. The network only builds parts because parts are what let it separate the source categories under a cross-entropy objective. But the parts it builds are far more general than the categories that motivated them: the texture and part detectors that distinguish one hundred dog breeds are exactly the detectors that also distinguish packaging, fabric or shelf products. ## Why the target's disjoint label set does not matter A 12-class retail-shelf classifier shares zero labels with the source's 1000 web categories. What it shares is the *input distribution* — natural images with edges, textures, lighting variation, occlusion and scale variation — and therefore the intermediate abstractions that any classifier over such images needs. Transfer works on the strength of shared input structure and shared useful intermediate abstractions, not shared classes. Concretely, the penultimate feature vector of the pretrained network is a description of the image in terms of parts and textures; a small target dataset is usually enough to learn a new linear map from that description to 12 outputs, even though it would be nowhere near enough to learn the description itself from scratch. ## Specificity increases with depth The useful way to hold this is as a gradient rather than a binary split. Layer 1 is almost purely generic. The last block is almost purely source-specific. Everything between shades from one to the other, and where the shading happens depends on how far the target domain sits from the source domain. For natural photographs of retail shelves, quite deep layers stay useful. For X-rays, satellite imagery or spectrograms, the input statistics diverge earlier, and the useful reuse typically stops lower in the stack — you keep the early filters and retrain more of the top. This is also why the standard recipes differ by target dataset size. A small target set can only afford to fit a shallow map on top of mostly frozen features; a large target set can afford to retrain deep layers and let them re-specialise to the target's own distinctions. ## What supervised pretraining buys you, mechanically Three things, and it is worth separating them: 1. **A better starting point in parameter space.** Optimisation begins somewhere that already computes useful abstractions, so the target run converges faster and from a region that generalises well. 2. **A regulariser.** With a small target set, initialising from pretrained weights and taking modest steps constrains the solution to stay near a representation learned from far more data, which reduces variance on the target task. 3. **Sample efficiency.** The expensive part — learning what an edge, a texture and a part are — was paid for once, using someone else's labels. ## Getting the boundaries right Two statements a candidate should avoid. First, transfer is *not* about class overlap; a source class list that happens to include something similar to a target class is a bonus, not the mechanism. Second, the layers are not equally transferable, so 'we reused the pretrained weights' is an incomplete answer — the interesting engineering is deciding how much of the stack to keep and how much to retrain, and that decision is driven by target dataset size and by how far the target inputs are from the source inputs.
- Does the same generic-to-specific split hold when the target is medical or satellite imagery?The ordering holds, but the useful depth shrinks. X-rays, overhead imagery and spectrograms share low-level structure with natural photographs — edges, gradients, texture — so the first layers still help, but mid and late features encode object parts and configurations that do not exist in the target domain. In practice you keep less of the stack and retrain more of it, and the gain over training from scratch is smaller than on natural images.
- How does the granularity of the source label set change what the middle layers learn?Fine-grained labels force finer features. A source task that must separate hundreds of visually similar categories cannot succeed on coarse shape alone, so it has to build discriminative texture and part detectors. Coarse labels can be solved with cruder features and tend to produce a less useful mid-level representation. Label diversity and granularity in the source corpus are therefore part of what makes a backbone transferable, not just the raw image count.
- If the early filters are so generic, why not hand-design them instead of pretraining?You can — hand-designed oriented filters and gradient-histogram descriptors predate learned ones and are close cousins of what layer 1 converges to. The reason not to is that the layers above them are what matter, and those are jointly optimised with the early filters. Learning the whole stack lets the early layers adapt to whatever the deeper layers need, and it scales to modalities where nobody knows the right hand-designed filter.
saying these in an interview costs you the question
- Says the network memorises source images and looks them up
- Claims transfer only works if source and target classes overlap
- Treats every layer as equally transferable
- Thinks the classifier layer is the only source-specific part
- Believes pretraining helps only by saving training time