skip to content

In an RBF-kernel SVM, what does gamma control, and what breaks when it is large?

level: middleimportance: must knowfreq 66%

answer

  1. inverse neighbourhood radius
  2. how fast similarity decays with distance
  3. each support vector gets a bump of influence
  4. gamma sets the bump's width
  5. huge gamma: an island per training point

basics

~20 s

Gamma sets how fast the RBF kernel's similarity decays with distance, so it is an inverse neighbourhood radius. Large gamma makes each training point influential only in a tiny region, producing an islands-around-points boundary that memorises the training set.

solid answer

~50 s

The RBF kernel is `k(x, z) = exp(-gamma * ||x - z||^2)`. It equals 1 when the two points coincide and falls towards 0 as they separate, and gamma sets how fast that fall happens — large gamma means a narrow neighbourhood, small gamma a wide one. Because the decision function is a weighted sum of kernel values against the support vectors, gamma directly sets how local the boundary is. Push gamma high and every training point influences only its immediate surroundings: the boundary degenerates into a bubble around each point, training accuracy goes to 100% and held-out accuracy collapses — the classic high-variance failure. Push gamma towards zero and every pair of points looks equally similar, the decision function flattens out, and the model underfits. Sweeping gamma over several orders of magnitude and watching the gap between training and validation scores is how you find the usable band.

go deeper

for a junior

Recall the direction: in the RBF kernel exp(-gamma times squared distance), a bigger gamma means a more wiggly, more local boundary and a higher risk of overfitting. Know that gamma is tuned, never left at a value copied from another dataset.

for a middle

Explain the mechanism, not just the direction. Each support vector contributes a bump of influence whose width gamma sets, so large gamma yields isolated islands and a kernel matrix approaching the identity, while small gamma flattens everything towards a single smooth surface.

for a senior

Demonstrate the diagnosis: read training-versus-validation curves across a logarithmic gamma sweep, name the underfit band and the memorisation band from their signatures, and explain why collecting more data does not rescue a gamma that is far too large.

for a principal

Own the search strategy. Argue for log-scale ranges anchored to the typical squared distance in the data, for tuning the kernel width jointly with the model's other dials rather than sequentially, and for a validation protocol that will not be fooled by a memorised training score.

## The function gamma lives in The RBF (radial basis function, also called Gaussian) kernel is ``` k(x, z) = exp(-gamma * ||x - z||^2) ``` where `||x - z||` is the Euclidean distance between the two points and `gamma > 0`. Read the formula as a similarity score: it returns 1 when `x` and `z` are the same point, and decays monotonically towards 0 as they move apart. Nothing else in the expression varies, so gamma alone decides how quickly that decay happens. It is often written in the equivalent bandwidth form `exp(-||x - z||^2 / (2*sigma^2))`, which makes the relationship explicit: `gamma = 1 / (2*sigma^2)`. Large gamma is a small sigma is a narrow bell. The most useful mental label is **inverse neighbourhood radius**: gamma answers "how far away does a point stop mattering?" ## Why that becomes the shape of the boundary A trained kernel SVM scores a new point as a weighted sum of kernel values against the support vectors it kept, plus an offset. Each support vector therefore contributes a bump of influence centred on itself, and gamma is the width of that bump. - **Large gamma — narrow bumps.** A support vector's influence dies out before it reaches its neighbours. The boundary stops being a smooth surface and becomes a collection of small islands, one wrapped around each awkward training point. Every training point can be given its own island, so the training set is classified perfectly; nothing has been learned about the space between the points. The matrix of pairwise kernel values approaches the identity matrix — every point is similar only to itself — and the model has effectively memorised its data. - **Small gamma — wide bumps.** The exponent goes to zero for every pair, so every kernel value approaches 1 and all points look alike. The weighted sum becomes almost the same number everywhere, the boundary becomes very smooth and eventually almost featureless, and the model underfits: it cannot express even the structure that is genuinely there. Between those extremes is a band where the bumps overlap enough to interpolate between training points but not so much that all local structure is smeared away. ## Diagnosing it in practice A worked case: 60 spectrometer channels per sample, the task is identifying which material produced the spectrum, and you sweep gamma from 0.001 up to 100 with everything else fixed. - At `gamma = 0.001` training and validation accuracy are both mediocre and nearly equal. Both curves being low and together is the signature of underfitting — the model is too smooth. - Somewhere in the middle both rise, validation peaks, and the gap between the two stays modest. That is the band you want. - By `gamma = 100` training accuracy is 100% and validation has fallen off a cliff. A large train/validation gap with perfect training performance is the signature of the islands regime, and no amount of extra data collection fixes it while gamma stays there — the setting itself is wrong. Two practical notes follow. First, sweep gamma **logarithmically**, not linearly: the interesting range spans orders of magnitude, and a linear grid from 0.001 to 100 wastes almost every point on the high end. Second, since gamma multiplies a squared Euclidean distance, the sensible starting scale depends on the typical squared distance between points in your data, which depends on the number of features and their units — a value that worked on one dataset carries no meaning on another. ## What large gamma is not It is not a bias problem, so do not respond to the symptom by adding capacity elsewhere. It is variance: the model is too flexible and is fitting the sample rather than the population. It is also not a data-quantity problem in the usual sense — with gamma huge enough, more data simply means more islands. And gamma is a property of the kernel, not of the classifier's tolerance for misclassification; those are separate dials that must be searched together rather than one after the other, because a very local kernel and a very forgiving fit can partially mask each other's symptoms on the training set. ## The reasoning to show A good answer connects three levels in order: the formula (gamma scales the squared distance inside an exponential), the geometry (it is the width of each support vector's zone of influence), and the diagnosis (large gamma yields perfect training accuracy with collapsed validation accuracy). Anyone who can only recite "large gamma overfits" has memorised the symptom without the mechanism, and the follow-up about what the kernel matrix looks like in that regime will expose it.

  • What does the matrix of pairwise RBF kernel values look like as gamma becomes very large?
    It approaches the identity matrix. Off-diagonal entries are `exp(-gamma * d^2)` for a non-zero distance `d`, which goes to 0 as gamma grows, while every diagonal entry stays exactly 1 because a point's distance to itself is zero. Every point is similar only to itself, which is precisely why the model can do nothing but memorise.
  • What does an RBF SVM behave like as gamma approaches zero?
    Increasingly like a model with almost no flexibility. Every kernel value approaches 1, so the weighted sum that produces the score varies barely at all across the input space and the boundary flattens towards something a straight separator could have drawn. You see it as training and validation scores that are both poor and nearly identical.
  • Why should a gamma search be on a logarithmic grid?
    Because gamma acts multiplicatively inside an exponential, so what matters is its order of magnitude, not its absolute step. The behavioural difference between 0.001 and 0.01 is enormous while the difference between 50 and 51 is nil. A linear grid over 0.001 to 100 spends nearly all its evaluations in the useless high-gamma region.

Each support vector is a lamp and gamma is how tightly its beam is focused. Focus every beam to a pinprick and each lamp lights only the spot it stands on, leaving the rest of the room dark and unclassified.

saying these in an interview costs you the question

  • Says gamma controls the tolerance for misclassified training points
  • Thinks large gamma smooths the boundary rather than sharpening it
  • Answers only large gamma overfits with no mechanism
  • Claims more training data fixes an over-large gamma
  • Searches gamma on a linear grid over several orders of magnitude

context