How do you choose DBSCAN's eps parameter from a sorted k-distance curve?
answer
- You have no labels to score against
- Look at each point's neighbour distances
- Distance to the minPts-th nearest neighbour, sorted
- The radius where the curve turns upward
basics
~20 sCompute every point's distance to its minPts-th nearest neighbour, sort those distances, and plot them. The curve stays flat where points have close neighbours and turns sharply upward where they do not. Set eps at that knee.
solid answer
~50 sFirst fix `minPts`, because the curve is built for a specific one — a common heuristic is roughly twice the number of features. Scale the features, since `eps` is one distance across all of them. Then, for every row, compute the distance to its `minPts`-th nearest neighbour, sort those distances descending, and plot them. On a 100k-row network-flow table the plot is flat and low for most of its length — the bulk of rows sit in dense traffic — then bends sharply upward for the tail of isolated rows. That knee is the distance at which points stop having close company, so it is the natural boundary between "dense" and "sparse". Read `eps` off the knee. Treat it as a starting range rather than an answer: run two or three values around it, check the noise fraction and cluster count, and have a domain expert sanity-check the result.
go deeper
Be ready to say that eps is a distance and that you have no labels to tune it against, so you inspect the data's own neighbour distances. Knowing that a plot of sorted neighbour distances is the standard starting point is enough at this level.
Explain the construction step by step: fix minPts, scale, compute each point's distance to its minPts-th nearest neighbour, sort, plot, read the bend. Say why the bend is meaningful — it is where points stop having close company.
Demonstrate the follow-through: two or three eps values across the knee, then judging them on noise fraction, cluster size distribution and stability under a small nudge, plus a domain expert reviewing sampled members. Say out loud that a knee-free curve is a finding about the data.
Own the decision of how much unlabelled tuning is worth doing at all before committing engineering time — whether the density gap is stable enough across data refreshes to build a pipeline on, and what the team does when the chosen radius silently stops fitting next quarter's data.
## Why this method exists DBSCAN's `eps` is a distance threshold, and you are choosing it without labels. There is nothing to hold out and score against. So instead of optimising, you look at the distance structure of the data itself and find the value where the data's own notion of "close" runs out. The sorted k-distance curve is the standard way to do that. ## Building the curve 1. **Fix minPts first.** The curve is defined relative to a chosen k, and you should use the same k you will pass as `minPts`. Common heuristics: at least the number of features plus one, often around twice the number of features, and never below 3 — higher for noisy data, because a larger count demands more evidence before calling a region dense. 2. **Scale the features.** `eps` is a single radius in the joint feature space. If one column is in bytes and another in seconds, the byte column swamps the distance and `eps` ends up measuring only that column. Standardise, or use a distance the domain actually justifies. 3. **For each point, find its k-th nearest neighbour and record that distance.** One number per row. 4. **Sort those numbers and plot them** — descending on the y-axis against rank on the x-axis is the usual presentation, though ascending works identically with the bend mirrored. ## Reading it For a 100k-row network-flow table, most rows are ordinary traffic packed tightly together, so most k-distances are small and nearly identical: the curve is a long flat stretch. A minority of rows sit off on their own, and their k-th neighbour is far away, so the curve climbs — gently at first, then steeply for the truly isolated rows. The **knee** is where the flat stretch gives way to the climb. Below it, points have `minPts` neighbours within a small radius, which is exactly DBSCAN's definition of dense. Above it, they do not. Setting `eps` at the knee therefore says: everything that behaves like the bulk of the data is dense, everything past the bend is a candidate for noise. The knee is rarely a single crisp pixel. Read the range it spans, take two or three values across it, and compare outcomes. ## Validating the choice without labels - **Noise fraction.** Look at what share of rows come back unassigned. There is no universal target; there is a target *for your problem*. If you are hunting rare anomalous flows, a few percent noise may be exactly what you want. If you are segmenting all traffic into behaviour groups, discarding a third of it is a signal that `eps` is too tight or `minPts` too high. - **Cluster count and size distribution.** One giant cluster plus a scattering of singletons usually means `eps` is above the knee and everything has merged. Dozens of tiny fragments usually means it is below. - **Stability.** Nudge `eps` by ten percent either way. If the partition reorganises completely, the density gap you thought you found is not really there. - **Domain review.** Pull a sample of members from each cluster and from the noise bucket, and have someone who knows the data say whether the grouping means anything. Unsupervised results earn trust this way, not by a number. ## When the curve has no knee A curve that rises smoothly from end to end is telling you something real: there is no density gap. Possible causes, in the order worth checking: - **Unscaled features**, so the plotted distance is really one column's spread. Fix and re-plot. - **Genuinely uniform data** — one diffuse blob with no separable structure, in which case density clustering has nothing to find and a different formulation is needed. - **Too many dimensions.** As dimensionality grows, distances between points concentrate: the nearest and farthest neighbour of a point become nearly equally far away. The k-distance curve flattens into a straight line and every `eps` behaves the same. Reducing dimensionality first, or choosing a distance that stays meaningful in that space, is the fix — not squinting harder at the plot. - **Multiple densities.** Sometimes there are two knees, not none, because two parts of the data are dense at different scales. A single `eps` cannot serve both, and that is a structural limit of the algorithm rather than a plotting problem. ## The interview point The method is short to describe, and the discriminating part is what you say afterwards: that scaling comes first, that the knee gives a range rather than an answer, that you check the noise fraction and cluster sizes before committing, and that a missing knee is information about the data rather than a reason to pick the midpoint.
- Your k-distance curve rises smoothly with no visible knee — what does that tell you?That there is no density gap to exploit. Check scaling first, since one unscaled column can flatten the curve into that column's spread. If scaling does not help, the data may be one diffuse blob, or dimensionality may be high enough that distances concentrate and every point's nearest and farthest neighbours are nearly equally far. Either way, density clustering is likely the wrong tool here.
- Why must features be scaled before you read eps off the plot?Because eps is a single radius in the combined feature space. A column measured in the thousands contributes far more to the distance than one measured in tenths, so the k-distance curve — and therefore eps — ends up describing that one column. Scaling puts the features on comparable footing so the radius means something about the whole record.
- DBSCAN labels 45 percent of your rows as noise. How do you react?First check scaling, then treat it as an eps-too-small or minPts-too-high signal and move eps toward the upper end of the knee range. But do not assume it is a defect: if the data genuinely is mostly sparse, a high noise share is the honest answer. Look at what the noise rows have in common — if they form a coherent group, that is a finding, not leftovers.
saying these in an interview costs you the question
- Tunes eps to whichever value yields the most clusters
- Picks eps before scaling the features
- Treats the knee as an exact value rather than a starting range
- Says eps is a percentage of the data or a fraction of points
- Sets minPts to 1, making every point its own cluster
- Reads a knee-free curve as a plotting problem rather than a finding