skip to content

How does mixup combine two training images and their one-hot labels into one example?

level: middleimportance: should knowfreq 52%

answer

  1. one weight, two things blended
  2. the target moves with the pixels
  3. cross entropy is linear in the target
  4. Beta below one is U-shaped
  5. label tracking, not label preservation

basics

~20 s

mixup draws a mixing weight lam from a Beta distribution and emits one blended example: pixels lam*x_i + (1-lam)x_j, with the target lamy_i + (1-lam)*y_j. The label is interpolated by the same weight, never rounded back to a single class.

solid answer

~50 s

For a pair of training examples you sample lam from Beta(a, a) - a common setting is a = 0.2 - then train on the single point `x = lam*x_i + (1-lam)*x_j` with the soft target `y = lam*y_i + (1-lam)*y_j`, where both labels are one-hot vectors. Because cross entropy is linear in the target, the loss equals `lam*CE(p, y_i) + (1-lam)*CE(p, y_j)`, so no soft vector need be materialised. With a = 0.2 the Beta density is U-shaped, so most draws land near 0 or 1 and most mixes are mild. What it teaches is linear behaviour between examples: the predicted probabilities should move in proportion to how far the input moved. The usual payoffs are better generalisation and less overconfidence on clean data; the price is that training accuracy stops being interpretable and the mixed image is not a natural image.

go deeper

for a junior

Know that mixup blends two training images and that the label is blended by the same amount, so the target is a mixture rather than a single class.

for a middle

Write the construction out: a weight drawn from a Beta distribution, applied to both pixels and one-hot targets, and explain why cross entropy lets you compute it as a weighted sum of two ordinary losses.

for a senior

Discuss the strength dial and the consequences: what the Beta parameter does to the blend distribution, the effect on confidence and calibration, and why training accuracy stops being readable.

for a principal

Argue when a label-mixing regulariser is the right spend at all, given that it trades label preservation for label tracking and produces inputs that never occur at inference.

## The construction Given two training examples `(x_i, y_i)` and `(x_j, y_j)` with one-hot label vectors, mixup samples `lam ~ Beta(a, a)` and forms a single training point: ``` x = lam*x_i + (1 - lam)*x_j y = lam*y_i + (1 - lam)*y_j ``` The pixel blend is an elementwise convex combination - a ghosted superposition of the two pictures. The crucial half is the second line: the target is interpolated with the *same* weight. If lam is 0.7 and the two classes are A and B, the target is 0.7 on A and 0.3 on B. It is never rounded to the majority class; rounding would turn mixup into plain input noise with a wrong label attached. In practice the pairing is done inside a batch by shuffling it against itself, which costs nothing. ## Why the loss is easy Cross entropy against a soft target is `-sum_c y_c * log p_c`, which is linear in `y`. Substituting the interpolated target gives ``` L = lam * CE(p, y_i) + (1 - lam) * CE(p, y_j) ``` So you compute the ordinary loss twice against the two original labels and take a weighted average. No soft label vector needs to exist explicitly. This linearity is specific to cross entropy against one-hot targets; a loss that is not linear in the target does not decompose this way. ## What Beta(0.2, 0.2) actually looks like This trips people up. Beta(a, a) with a = 1 is uniform on [0, 1]. With a *below* 1 the density is U-shaped: mass piles up at both ends, so lam is usually close to 0 or 1 and the blended image is usually dominated by one of the two sources with a faint ghost of the other. Strong 50/50 mixes are the tail, not the norm. With a above 1 the density peaks at 0.5 and near-even blends become typical, which is a much more aggressive setting. So `a` is the strength dial, and small values are the mild end - the opposite of the intuition that a smaller number means less of something. ## What it teaches, and what it costs The stated motivation is to encourage linear behaviour between training examples: as the input moves a fraction of the way from one example to another, the predicted distribution should move the same fraction. Standard empirical risk minimisation says nothing about the space between training points, and a high-capacity network is free to behave arbitrarily there - which is where sharp, overconfident decision regions live. mixup fills that space with supervised targets. Observed consequences: - **Generalisation**: usually a modest but real improvement on clean test data, most visible when the dataset is small relative to the model. - **Confidence**: models trained with mixup are typically less overconfident on clean inputs, because they are constantly trained toward non-degenerate targets. Pushed hard - a large `a` - the same mechanism can leave a model underconfident, so the calibration effect is a dial, not a guarantee. - **Robustness to label noise**: a mislabelled example rarely dominates a blend, which softens its influence. The costs are real too. The mixed image is off the natural image manifold, so you are training partly on inputs that will never appear at inference. Training accuracy against mixed targets is no longer a quantity you can read - report clean-data metrics. And mixup interacts with anything else that softens targets, so stacking it with an aggressive label-smoothing coefficient can over-soften. ## The relationship to the label-preserving rule Every other augmentation on this topic obeys a hard constraint: do not change the label. mixup does not obey it - it changes the input in a way that genuinely changes the correct answer, and then it changes the label to match. That is the whole trick, and it is the cleanest way to describe mixup in an interview: it trades *label preservation* for *label tracking*. cutmix is the same bargain with a different geometry. Instead of blending intensities everywhere, it cuts a rectangle out of one image and pastes the corresponding rectangle from another, then mixes the labels by the pasted area fraction: if the patch covers 30 percent of the frame, the target is 0.7 on the source class and 0.3 on the pasted class. The pixels stay locally natural - no ghosting - and the model is pushed to use the whole frame rather than one discriminative spot. Its weakness is that area is only a proxy for evidence: a patch can land on empty background and contribute none of the pasted class while still claiming 30 percent of the label. ## When to reach for it mixup is cheap, needs no domain knowledge, and is a reasonable default on classification with soft-target-compatible losses when you are short of data. It is a poor fit where a blended input is meaningless or the target is not a class distribution - dense localisation tasks are the common example, where cutmix-style pasting with adjusted targets is the better-behaved relative.

  • How does cutmix differ from mixup, and how is its label decided?
    cutmix cuts a rectangle from one image and pastes the same region from another instead of blending intensities, so local pixel statistics stay natural. The target is mixed by pasted area: a patch covering 30 percent of two satellite land-cover tiles gives 0.7 to the host class and 0.3 to the pasted one. The weakness is that area only proxies evidence - the patch may land on empty background.
  • Why does the mixing weight for mixup usually come from a Beta with a parameter below one?
    Beta(a, a) with a below 1 is U-shaped, so the weight lands near 0 or 1 most of the time and most blends are mild, with strong mixes as the tail. A parameter above 1 peaks at 0.5 and makes near-even blends typical, which is a far more aggressive regulariser. The parameter is the strength dial.
  • How do you evaluate a model trained with mixup?
    On clean, unmixed data - mixup is a training-time construction only. Training accuracy against interpolated targets is not comparable to anything, so track a clean validation metric for selection. If you care about confidence, measure calibration explicitly, since mixup shifts it and a heavy setting can leave the model underconfident.

saying these in an interview costs you the question

  • Blends the pixels but keeps the dominant image's hard label
  • Thinks a Beta parameter of 0.2 gives mostly even blends
  • Applies mixup to inputs at inference time
  • Cannot state what target a mixed image receives
  • Reads training accuracy on mixed batches as real accuracy

context