Designing an augmentation policy for chest radiographs, how do you decide a transform is label-preserving?
answer
- would the labeller answer the same?
- nuisance variation versus label-carrying variation
- laterality is clinically meaningful
- does deployment vary that way?
- check per-class recall, not the aggregate
basics
~20 sAsk whether an expert labeller would give the transformed image the same label. Brightness and contrast jitter mimic exposure differences and are safe on chest radiographs; a horizontal flip is not, since mirroring destroys laterality.
solid answer
~50 sThe test is mechanical: apply the transform and ask whether the labelling function - in practice the expert who annotates it - returns the same answer. On chest radiographs that splits the catalogue cleanly. Exposure and detector settings vary between machines, so brightness, contrast and gamma jitter, mild noise and slight blur are label-preserving. Positioning varies a little, so small rotations, translations and mild scaling are fine. A horizontal flip is not, and it is the trap: many findings are lateral - a left versus right pleural effusion, dextrocardia - so mirroring yields an image whose correct label differs from the one you carried over, and it silently teaches the model to ignore side. The rule is to augment variation that genuinely occurs in deployment and never variation the label depends on, then have a domain expert review sampled augmented images.
go deeper
Know the rule itself: an augmentation may change the image but must never change the correct answer. Be able to give one transform that is safe on photographs and one that is not safe on medical images.
Explain the check operationally - would the annotator return the same label - and sort a concrete transform list into safe and unsafe for a named domain, with the reason for each.
Show that you validate rather than assume: sampled review with a domain expert, per-class metrics after a policy change, bounded ranges, and separating label-breaking from merely useless transforms.
Own the fact that the policy is a documented claim about which invariances the system asserts, tied to the label definitions. Be ready to say who signs off on it and what happens to it when the label taxonomy changes.
## The criterion, stated precisely Let `f` be the labelling function - the process, usually human, that assigns the ground-truth label. A transform `T` is label-preserving for a domain if `f(T(x)) = f(x)` for essentially every `x` in that domain. Nothing about the model enters this definition. It is a statement about the data and the annotation protocol, and it must be settled before training. When the criterion fails you are not regularising. You are appending mislabelled examples to the training set at whatever rate the transform fires, which puts a ceiling on achievable accuracy that no amount of training removes. ## Applying it to chest radiographs Start by sorting the sources of variation into two bins. **Nuisance variation that really occurs.** Different machines and exposure settings produce different overall brightness, contrast and gamma. Detectors add noise. Patients are not positioned identically, so images differ by a few degrees of rotation, small translations and a little scale. Augmenting these is the whole point - you want a model that does not care which room the study was taken in. **Variation the label depends on.** Laterality is the big one in this domain. Many findings are side-specific: a pleural effusion is reported as left or right; dextrocardia and situs inversus are defined by organs sitting on the wrong side. A horizontal flip maps a normal study onto something that reads as an abnormal one and vice versa, so `f(T(x)) != f(x)`. Radiographs also usually carry a burned-in laterality marker, so the flip additionally creates an image that contradicts itself - a marker that says one side over anatomy on the other. The conclusion is not a rule about flips in general. Horizontal flip is the safest augmentation in existence for natural photographs of objects. It is domain-specific, and that is the lesson. ## The second question: does the invariance exist at deployment? Label preservation is necessary, not sufficient. A 180-degree rotation of a chest radiograph is arguably label-preserving in the narrow sense - a radiologist could still read the finding - but studies always arrive upright. Training on upside-down images spends capacity buying an invariance the deployment distribution never asks for, and can cost accuracy by making the task harder for no return. So the policy has two filters: the transform must preserve the label, and the variation it simulates should be variation you actually expect to see. ## Transforms that are label-preserving for some classes only The awkward middle case. A transform can be safe for most of the label space and destructive for a slice of it. Aggressive cropping on radiographs is an example: crops that clip the costophrenic angles can remove exactly the evidence for a small effusion, so the image now shows nothing and the label still says effusion. The class-conditional damage does not show up in the overall validation number if that class is rare - it shows up as a per-class recall collapse. Always check per-class metrics when a policy changes, not just the aggregate. ## Moving the reasoning to another modality The same discipline transfers. For a spoken-command recognizer trained on 40 hours of speech, the inputs are log-mel spectrograms - time on one axis, frequency on the other. Masking a bounded span of time steps or a bounded band of frequency channels leaves the command word recoverable from the surrounding context, so the label survives, and it simulates dropout of information a real microphone environment can cause. Two constraints keep it label-preserving: the mask width has to be bounded relative to the utterance length, or a short command disappears entirely; and the flip that is free on photographs is fatal here, because reversing the time axis turns speech into something no human would transcribe the same way. Frequency-axis flipping is equally destructive - it inverts the formant structure that identifies the phonemes. ## How to validate a policy before it ships 1. **Sample and look.** Render a grid of augmented examples at the strength you plan to use. Most broken policies are visible in thirty seconds. 2. **Have the labeller confirm.** Show augmented images to a domain expert and ask for the label without telling them the original. Disagreement is your answer. 3. **Ablate.** Add one transform family at a time and watch clean validation metrics, per class. 4. **Write the policy down.** The list of transforms and their ranges is a statement of which invariances the system asserts. It belongs in the model documentation next to the label definitions, because a later change to either can invalidate the other. ## The one-line summary for an interview Augmentation encodes invariances. Every transform you add is a claim that the label does not depend on that factor. Make the claim deliberately, domain by domain, and verify it with someone who assigns the labels - not with a generic list of transforms copied from another dataset.
- How does the same reasoning transfer to log-mel spectrograms for a spoken-command recognizer?Masking a bounded span of time steps or a bounded band of frequency channels is label-preserving: the word stays recoverable from the surrounding context, and it simulates real information loss. The bounds matter - a mask wide relative to a short utterance can erase the command. Reversing the time axis or flipping the frequency axis destroys the label outright, so the photo-style flip has no analogue here.
- A transform preserves the label but the invariance never occurs at deployment. Should you still use it?Usually not. Chest studies always arrive upright, so training on 180-degree rotations buys an invariance nothing will ever exercise, makes the fitting problem harder and can cost clean accuracy. Label preservation is the necessary filter; matching real deployment variation is the second one that decides whether the transform earns its place.
- What if a transform is safe for most classes but destroys evidence for one rare class?Aggressive crops that clip the costophrenic angles can remove the only evidence for a small effusion while the label still says effusion. Aggregate validation accuracy will not move if the class is rare, so watch per-class recall whenever the policy changes, and bound the transform's range so it cannot remove the region a class is defined by.
Rotating a photograph of a cat still shows a cat. Mirroring a floor plan still shows a building, but now the fire exit is on the wrong side - and the label was about the exit.
saying these in an interview costs you the question
- Copies a generic transform list without checking the domain
- Flips medical images horizontally because flips are standard
- Treats label preservation as a property of the model, not the data
- Judges a policy change on aggregate accuracy only
- Adds invariances the deployment distribution never exhibits