skip to content

What does the bandwidth of a kernel density estimate control, and how does it mislead?

level: seniorimportance: nice to knowfreq 26%

answer

  1. One smooth bump per observation
  2. Width of the bump is the dial
  3. Bias-variance in a new coat
  4. Two humps flattened into one bell
  5. Default rules assume a single mode

basics

~10 s

Bandwidth sets how wide the smooth bump placed at each observation is, so it controls smoothness. Too large flattens two genuine humps into one bell; too small turns individual observations into spurious peaks.

solid answer

~50 s

A kernel density estimate replaces each observation with a small smooth bump, usually a narrow bell, and adds them up, dividing by the sample size so the curve integrates to 1. The bandwidth is the width of those bumps. A large bandwidth spreads each observation's mass widely, so neighbouring clusters overlap and a genuinely two-humped distribution is drawn as a single smooth bell, which is exactly the structure you were plotting to find. A small bandwidth leaves each observation as its own spike, so sampling noise reads as many modes. The kernel shape matters far less than the bandwidth. Automatic rules such as Silverman's rule of thumb scale the bandwidth with the sample spread and n^(-1/5) under a near-normal assumption, so they deliberately over-smooth multimodal data. Vary the bandwidth and see which features survive before believing any of them.

go deeper

for a junior

Be ready to say that a density curve is a smoothed picture of the data and that the amount of smoothing is a setting someone chose, so the curve is not simply a fact about the sample.

for a middle

Explain the construction, one kernel per observation summed and normalised, and describe both failure directions: over-smoothing merges modes, under-smoothing manufactures them. Note that bandwidth matters far more than kernel shape.

for a senior

Show the diagnostic habit. Vary the bandwidth before believing a mode, recognise density spilling past a hard boundary as an artifact, and know that default rules derived under a normality assumption systematically hide a second hump.

for a principal

Own what your organisation puts in front of decision-makers. A smooth curve carries an air of authority a small sample does not earn, so decide when a density plot is appropriate at all and what must be shown alongside it.

## What a kernel density estimate is A kernel density estimate, or KDE, draws a smooth curve intended to show the shape of a distribution from a sample. The construction is simple: put a small symmetric bump, called a kernel, centred on each observation; add all the bumps together; divide by the number of observations so the total area under the curve is 1. Written out, the estimate at a point x is the average over observations of K((x - xi) / h) / h, where K is the kernel shape and h is the bandwidth. Common kernel shapes are the Gaussian bell and the Epanechnikov parabola. The choice among reasonable kernels has only a small effect on the result; the bandwidth h dominates. ## What the bandwidth does The bandwidth is the width of each bump: how far one observation's influence spreads along the value axis. **Large bandwidth (over-smoothing).** Each observation's mass is spread thinly across a wide range, so contributions from distant observations blend together. Structure is erased. The canonical failure is a distribution with two genuine clusters, say a population of quick operations and a population of slow ones, being drawn as one broad symmetric bell. Nothing in the picture hints that two populations exist. You went to the plot to find shape and the smoothing parameter deleted the shape. **Small bandwidth (under-smoothing).** Each bump is narrow, so the curve rises sharply at each observation and drops between them. With a modest sample you get a spiky curve with one local peak per observation, or per small cluster of observations. Every one of those peaks is sampling noise. Draw a fresh sample from the same source and the peaks move. This is the standard bias-variance tradeoff. Large h means low variance and high bias: stable across samples, systematically wrong about shape. Small h means high variance and low bias: faithful on average but dominated by noise in any single sample. ## The smoothing choice does not disappear, it moves A KDE is often reached for as the grown-up alternative to a histogram, on the grounds that it has no arbitrary bin edges. That is half true. The dependence on bin *placement* really does go away, since the curve does not depend on where a grid starts. But the dependence on a smoothing *width* is exactly the same problem in a new coat: bin width becomes bandwidth. If you would not trust a histogram drawn at one arbitrary bin width, do not trust a KDE drawn at one arbitrary bandwidth. ## Automatic bandwidth rules and their assumption Default bandwidths usually come from a normal-reference rule. Silverman's rule of thumb sets the bandwidth proportional to a measure of the sample's spread times n^(-1/5), derived by asking what bandwidth would be best *if the data were normal*. That derivation is the catch. A normal distribution has one mode, so the rule is tuned to reproduce one mode. Applied to genuinely bimodal data it reliably over-smooths and hides the second hump. Cross-validation-based selectors are less biased in this respect but noisier and more expensive. The operational conclusion is that a default bandwidth is a hypothesis, not a setting. Redraw at roughly half and twice the default. A feature that survives all three is probably real; one that appears only at the smallest bandwidth is probably noise; one that vanishes at the default may still be real and merely smoothed away. ## Boundary spill A second, frequently missed failure is at hard limits. Durations, counts, prices and latencies cannot be negative, but the kernel centred on an observation near zero puts half its mass below zero. The estimate then shows positive density on impossible values and, worse, is biased downward near the boundary because there are no observations on the other side to contribute mass back. If a density plot of response times shows a visible curve to the left of zero, that is the artifact, not a data problem. Remedies include reflecting the data at the boundary, transforming to a scale where the constraint disappears, or using a bounded estimator. ## Practical guidance - Use a KDE to communicate a shape you have already verified, not to discover shape from a single draw. - Always state or show the bandwidth. A density plot without it is unreproducible. - Cross-check any claimed mode against a plain view of the data, such as a histogram at several widths or an empirical CDF, whose steep-then-flat-then-steep pattern reveals separate clusters with no smoothing parameter at all. - Beware of small samples. With a few dozen observations a KDE looks confident and authoritative while carrying almost no information about shape; the smooth curve conveys a precision the sample does not have. - Respect hard boundaries explicitly rather than letting the curve spill past them. ## The interview answer Name the mechanism (a kernel per observation, summed and normalised), identify the bandwidth as the width of those kernels, place it on the bias-variance tradeoff with a concrete failure in each direction, note that the kernel shape barely matters while the bandwidth dominates, and finish with the professional habit of varying the bandwidth and cross-checking against an unsmoothed view.

  • Does a kernel density estimate remove the arbitrary choice a histogram forces on you?
    Only partly. It removes the dependence on where bin edges are placed, since no grid is involved. It does not remove the smoothing width: bin width simply becomes bandwidth, and the same bias-variance tradeoff applies. The honest practice is identical in both cases, which is to redraw at several smoothing levels and trust only features that persist.
  • A density plot of response times shows positive density below zero milliseconds. What happened?
    Boundary spill. Each kernel is symmetric around its observation, so observations close to zero place part of their mass on the impossible side of the boundary. It also biases the estimate downward just inside the boundary, because no observations exist beyond it to contribute mass back. Fix it by reflecting the data at the boundary, transforming the scale, or using a bounded estimator.
  • How would you decide whether a second hump in a density plot is real?
    Redraw at several bandwidths, including well below the default, and see whether the hump persists. Check the raw observations in that region: a real second population has substantial mass, not three points. Then look for an explanatory variable, split the data by it, and see whether the two groups separate. A hump you can attribute to a mechanism is far more convincing than one you can only see.

Bandwidth is the blur radius in a photo editor: turn it up far enough and two faces in the crowd merge into one smudge.

saying these in an interview costs you the question

  • Thinks a density curve has no tuning parameter
  • Claims kernel shape matters more than bandwidth
  • Accepts the default bandwidth as objectively correct
  • Reads every wiggle of an under-smoothed curve as a mode
  • Ignores density drawn below a hard zero boundary

context