skip to content

What is the difference between the functional margin and the geometric margin of a separating hyperplane?

level: middleimportance: should knowfreq 45%

answer

  1. one is a score, one is a distance
  2. divide by the length of the normal vector
  3. multiply w and b by ten
  4. one quantity moves, the other does not
  5. pin the closest point at exactly one

basics

~20 s

The functional margin of a point is the signed score y times (w.x + b); the geometric margin divides that by the length of w, giving an actual perpendicular distance. Rescaling w and b inflates the functional margin but leaves the geometric one unchanged.

solid answer

~50 s

For a hyperplane `w.x + b = 0` and a label `y` in {-1, +1}, the functional margin of a training point is `y * (w.x + b)` - positive when the point is on the correct side, larger when the score is more confident. The geometric margin is `y * (w.x + b) / ||w||`, the true perpendicular distance from the point to the surface. The difference matters because the parameters are not identified by the hyperplane: multiply `w` and `b` by 10 and you get exactly the same decision surface, but every functional margin is now ten times bigger while no point has moved. So maximising the functional margin is meaningless - you can drive it to infinity by rescaling. The standard fix is to pin the scale by requiring the closest points to have functional margin exactly 1. Then the geometric margin is `1 / ||w||`, the street between the two supporting hyperplanes is `2 / ||w||` wide, and maximising the margin becomes minimising `||w||^2`.

go deeper

for a junior

Recall that one margin is a raw signed score and the other is a genuine distance obtained by dividing by the length of the weight vector. Knowing which one is scale-dependent is the point.

for a middle

Walk through the rescaling argument out loud: multiply the weights and offset by a constant, show the surface and the predictions are identical, and note which margin changed. Then derive the width of the street as two over the weight norm.

for a senior

Connect the distinction to the objective actually solved: explain why the naive maximisation is unbounded, why the unit-margin constraint is a normalisation rather than a modelling choice, and why raw scores are not comparable across fitted models.

for a principal

Be ready to generalise the lesson - identifiability. Argue why any objective invariant to a reparameterisation must be stated in invariant quantities, and how that discipline shows up elsewhere when teams compare scores across independently fitted models.

## Two quantities that look similar Write a linear decision surface as `w.x + b = 0` and let labels be `y` in {-1, +1}. For a training point `(x_i, y_i)`: - **Functional margin**: `f_i = y_i * (w.x_i + b)`. It is positive exactly when the point is on the correct side, and it grows with the raw score. - **Geometric margin**: `g_i = y_i * (w.x_i + b) / ||w|| = f_i / ||w||`, where `||w||` is the Euclidean length of the normal vector. This is the actual perpendicular distance from the point to the surface, measured in the units of the feature space. For the dataset as a whole, each margin is defined as the minimum over all training points: the closest point sets the number. ## The rescaling test The crucial fact is that `(w, b)` and `(10w, 10b)` describe the *same* hyperplane: `10 * (w.x + b) = 0` has exactly the same solution set as `w.x + b = 0`, and the predicted sign is unchanged for every input. Now compare margins under that rescaling: - every functional margin is multiplied by 10; - every geometric margin is unchanged, because both the numerator and `||w||` scale by 10 and the factor cancels. That is the whole distinction. The functional margin is a property of the *parameterisation*; the geometric margin is a property of the *geometry*. No point moved, no prediction changed, so any quantity that changed cannot be a real measure of separation. ## Why this sinks the naive objective A first attempt at a maximum-margin rule might be "choose `w, b` to maximise the smallest functional margin". That problem has no solution: take any separating hyperplane and scale its parameters up without bound, and the objective goes to infinity while the classifier is literally unchanged. The objective is unbounded for a degenerate reason. Maximising the smallest *geometric* margin is the well-posed version, since it is invariant to that rescaling. ## Canonical scaling The invariance also gives a free choice: since scaling does not change the classifier, you may fix it however is convenient. The standard convention is the **canonical form** - normalise so the closest points have functional margin exactly 1, i.e. impose ``` y_i * (w.x_i + b) >= 1 for every i ``` with equality for the closest points. Under this convention: - the geometric margin of the dataset is `1 / ||w||`; - the two supporting hyperplanes `w.x + b = +1` and `w.x + b = -1` sit one on each side, and the perpendicular gap between them - the width of the street - is `2 / ||w||`; - maximising `2 / ||w||` is the same as minimising `||w||`, and by convention `||w||^2 / 2`, which is smooth and convex. So the familiar objective, "minimise the squared norm of the weight vector subject to every point having functional margin at least one", is not an arbitrary formula: it is the geometric-margin maximisation problem after the scale ambiguity has been removed. ## Reading the two numbers in practice Because the functional margin depends on the parameterisation, its absolute value tells you nothing on its own - a score of 4 from one fitted model and a score of 0.4 from another may describe identical boundaries with identical confidence. Comparisons are only meaningful after normalising by `||w||`, or when comparing points under one fixed model. Under the canonical convention the two coincide in a useful way: a point with functional margin 1 sits exactly on the kerb, functional margin greater than 1 means it is off the street entirely, and in the strictly separable case no point has functional margin below 1. ## Common confusions A frequent error is treating a large raw score as evidence of a confident classifier without dividing by `||w||`. Another is thinking the geometric margin is unitless: it is a distance in feature space, so it carries whatever units the features are measured in, and comparing margins across differently-measured feature sets is not meaningful. A third is assuming shrinking `||w||` makes the boundary move - it does not move the surface, it widens the street, which is precisely why the optimisation targets it.

  • Why is the constraint written as y times (w.x + b) at least 1 rather than at least 0?
    At least 0 only asks for correct classification and admits a boundary touching the data. Fixing the threshold at 1 both demands a strict gap and removes the scale ambiguity, so that the geometric margin becomes `1 / ||w||` and the objective is well posed. The value 1 itself is arbitrary - any positive constant gives the same hyperplane.
  • Under the canonical scaling, what does a smaller norm of the weight vector mean geometrically?
    A wider street. The gap between the two supporting hyperplanes is `2 / ||w||`, so shrinking the norm widens the separation while the constraints keep every point off the street. That is why the maximisation of a distance turns into the minimisation of a squared norm.
  • Can you compare functional margins from two separately fitted linear models?
    Not directly. Each model has its own arbitrary parameter scale, so one may report scores ten times larger purely because its weight vector is longer. Divide each score by its own weight-vector norm to get perpendicular distances, and even then the comparison only holds if both models use the same feature units.

saying these in an interview costs you the question

  • Treats a large raw score as a large distance
  • Says the two margins differ only by a constant across models
  • Thinks rescaling the weights moves the decision boundary
  • Cannot explain why maximising the functional margin is ill-posed
  • Believes the choice of 1 in the constraint is theoretically special

context