skip to content

How do you compute distance for a record mixing salary, tenure, department and a remote flag?

level: seniorimportance: nice to knowfreq 30%

answer

  1. no single metric spans euros and codes
  2. score each feature separately, then average
  3. numeric gap over the feature's range
  4. match or mismatch gives zero or one
  5. all-binary case collapses to Hamming

basics

~20 s

Build a per-feature dissimilarity that already lives on a 0-to-1 scale and average those, which is what Gower's coefficient does: absolute difference over the feature's range for numbers, zero or one for a categorical match or mismatch.

solid answer

~50 s

No single classical metric covers euros, years, a department code and a boolean at once, so you either encode everything numerically or use a mixed-type dissimilarity. Gower's approach computes a contribution per feature, each already in [0, 1], then averages them: for a numeric feature it is `|x_j - y_j| / range_j`; for a nominal or binary feature it is 0 if the two records match and 1 if they differ; ordinal features are compared on ranks. Averaging means every feature carries equal weight unless you supply weights, and missing values are handled by skipping that feature for that pair and averaging over the rest. All-binary features reduce it to Hamming distance over the item count: the fraction of positions where two records disagree. The main caveat: dividing by the observed range lets one extreme value shrink everyone else's contribution.

go deeper

for a junior

Know that you cannot subtract two category labels, so mixed records need either numeric encoding of the categories or a metric that scores each feature type separately.

for a middle

Describe the per-feature 0-to-1 contributions and the average, including the match/mismatch rule for categories and the range division for numbers, and name Hamming as the all-binary case.

for a senior

Diagnose the failure modes in a real table: rare one-hot levels after standardisation, one extreme salary flattening the salary term, and pairs compared on different feature sets under missingness.

for a principal

Own the weighting decision. Averaging silently declares every column equally important; decide whether encode-then-Euclidean or a mixed dissimilarity fits the pipeline, and who signs off on the implied importance of each feature.

## Why the usual metrics do not apply Euclidean, Manhattan and cosine distance all assume every feature is a number whose differences mean something. An HR record breaks that assumption: salary in euros, tenure in years, a department code with a dozen unordered levels, and a remote-working flag. There is no meaningful subtraction between `Finance` and `Logistics`, and no shared unit joining euros to years. ## The naive route and where it hurts The common move is to encode everything numerically — one indicator column per department level, 0/1 for the flag, standardised salary and tenure — and use Euclidean distance. This works, and it is often the right pragmatic choice, but three things happen quietly. First, a department mismatch has a fixed cost: two records in different departments differ by 1 in two indicator columns, contributing 2 to the squared distance, whether the departments are adjacent teams or unrelated functions. Second, **standardising indicator columns is dangerous**: a rare level's column has a tiny standard deviation, so dividing by it can make a mismatch on that rare level swamp every other feature in the record. Third, the implicit weighting between the categorical block and the numeric block is an accident of how many levels each categorical had, not a decision anyone made. ## Gower's construction Gower's coefficient tackles the problem differently: instead of mapping everything into one space and then measuring, it measures each feature *in its own terms* and forces the result onto a common 0-to-1 scale before combining. For a pair of records `i` and `j`, per feature `k`: - **Quantitative** (salary, tenure): `d_k = |x_ik - x_jk| / R_k`, where `R_k` is the feature's range across the data. Identical values give 0, the two extremes of the feature give 1. - **Nominal** (department): `d_k = 0` if the codes match, `1` if they do not. - **Binary** (remote flag): the same match/mismatch rule for a symmetric flag. For an *asymmetric* binary attribute — where presence is informative and absence is not, such as a rare qualification — Gower's original scheme drops the pair from the average when both records are absent, so two people who both lack a rare certificate are not credited with a similarity for it. - **Ordinal**: converted to ranks first, then treated like a quantitative feature on the rank scale. The overall dissimilarity is the weighted average of the per-feature terms, `sum_k w_k d_k / sum_k w_k`, with weights defaulting to 1. Because each term is bounded by 1, so is the result. ## The special cases worth naming If every feature is a symmetric binary indicator — a 20-item symptom checklist, say — every per-feature term is 0 or 1 and the average is simply the count of disagreeing items divided by 20. That count is **Hamming distance**, the number of positions at which two equal-length sequences differ; Gower is its normalised form. Finding the most similar prior patient on such a checklist is a pure Hamming query, and no scaling question arises at all, because every feature already contributes on the same 0/1 scale by construction. If every feature is quantitative, Gower is Manhattan distance on range-normalised features, divided by the feature count. ## Missing values, weights and the caveats Missing data is handled per pair: if either record lacks feature `k`, that feature is excluded from both the numerator and the denominator for that pair. This is convenient but not free — different pairs are then compared on different feature sets, so a pair that agrees on the three features they share can outrank a pair compared on all twelve. Watch it when missingness is heavy or systematic. Three more caveats belong in a senior answer. **Range normalisation is outlier-sensitive**: one executive salary of ten million stretches `R_salary` so that every ordinary salary difference collapses towards zero and salary silently stops mattering. **Equal default weights are an assumption**, and with eight numeric columns and one flag, the flag contributes one ninth of the decision no matter how important it is to the business — set weights deliberately if the domain has an opinion. **The output is a pairwise dissimilarity, not a set of coordinates**, so methods that need points in a vector space rather than a distance matrix need a different treatment. ## Choosing between the two routes Encode-then-Euclidean is faster, composes with anything that expects numeric features, and is a reasonable default when categoricals are few and low-cardinality. A Gower-style mixed dissimilarity is worth it when the categorical and numeric blocks are both substantial, when missingness is common, or when you want the per-feature contributions to be interpretable — being able to say "this pair scored 0.7 because department and remote status both differ" is a real advantage when a human has to sanity-check the neighbours. The strong answer names both routes, the failure mode of each, and the data property that decides.

  • What goes wrong if you standardise one-hot indicator columns before a distance?
    Rare levels blow up. An indicator that is 1 in half a percent of rows has a very small standard deviation, so dividing by it turns a single mismatch on that rare level into a huge contribution that outweighs every other feature. Records then group by rare-category membership rather than by overall similarity. Either leave indicators on their natural 0/1 scale or weight the categorical block deliberately.
  • How does a Gower-style dissimilarity handle a record with a missing tenure?
    It drops that feature from the comparison for pairs where it is missing and averages over the features both records have. That avoids imputing a value you do not have, but it means different pairs are scored on different feature sets, so a pair compared on three features is not strictly comparable to one compared on twelve. With heavy or systematic missingness, that distortion is worth measuring before trusting the neighbour lists.
  • For a 20-item binary symptom checklist, what does the mixed-type dissimilarity reduce to?
    Hamming distance, normalised. Every feature contributes 0 for a match and 1 for a mismatch, so the average is the number of items on which the two records disagree divided by 20. Finding the most similar prior patient is then just counting disagreements. If some items are asymmetric — a rare finding where joint absence says nothing — those both-absent pairs should be excluded rather than counted as agreement.

saying these in an interview costs you the question

  • Claims Euclidean distance works fine on raw category codes
  • Standardises one-hot indicator columns without noticing rare-level blowup
  • Never mentions that averaging assumes equal feature weights
  • Ignores that range normalisation is distorted by one extreme value
  • Treats Hamming and Gower as unrelated rather than a special case

context