Which measure of centre works for a nominal column such as payment method?
answer
- which operations does a label support
- you cannot add or sort labels meaningfully
- counting is all that survives
- most frequent, not most of them
- the frequency table is the real summary
basics
~20 sThe mode — the most frequently occurring category. Payment methods have no order and no arithmetic, so mean and median cannot be computed at all; only how often each value occurs is meaningful. Report the full frequency table alongside it.
solid answer
~50 sOnly the mode. The mean requires adding and dividing values, and card plus wallet plus transfer is not a quantity. The median requires sorting, and unordered labels have no meaningful sort — any ordering you impose is an artefact of encoding, so a median computed over category codes changes when you relabel them. The mode is the value with the highest frequency, and it survives because counting is the only operation an unordered label supports. Two caveats matter in practice. The mode can tie or be effectively ambiguous when two categories are close, in which case say the distribution is bimodal rather than picking one. And a mode is only a plurality, not a majority: if the top payment method holds 22% share, the sentence the typical customer pays by that method is wrong. The frequency table is the real summary; the mode is its headline.
go deeper
Know that unordered labels support only counting, so the mode is the only centre available, and that a mode is the most common value rather than a majority.
Explain the failure precisely: the mean does arithmetic on arbitrary codes and the median depends on an arbitrary sort, so both change when categories are relabelled while counts do not.
Show reporting discipline — lead with the frequency table and shares, flag near-ties as bimodal instead of declaring a winner, and refuse claims a plurality cannot support.
Own how categorical dimensions are summarised and governed across reporting: stable category definitions, an explicit rule for the other bucket, and wording standards that stop a plurality being sold as a majority.
## What each centre needs from the data The three classic centres make escalating demands on the values they are given. - **Mode**: needs only equality testing — can I tell whether two observations are the same value? Every column supports this. - **Median**: needs an ordering — can I say one value comes before another? Available for anything ordered. - **Mean**: needs arithmetic — addition and division on the values themselves. A nominal column such as payment method, browser, country or referral source is a set of labels with no inherent order. Equality is the only relation it supports, so the mode is the only one of the three that is defined. ## Why the mean and median genuinely fail The mean fails first and most obviously. Card, wallet and bank transfer cannot be summed. If the column is stored as integer codes — 1, 2, 3 — a mean *can* be computed, and this is where careless work slips through: the number 1.87 is arithmetic performed on identifiers, and it changes if someone renumbers the categories. A statistic that depends on an arbitrary encoding is not describing the data. The median fails for a subtler reason. Sorting requires a meaningful order, and label order is arbitrary. Sort the categories alphabetically and one value sits in the middle; sort by code and a different one does. Because the answer depends on a choice nobody made on statistical grounds, the median carries no information about the column. The mode is immune to both problems. It asks only which label appears most often, and relabelling or reordering the categories changes nothing about the counts. ## Using the mode honestly The mode is a weak summary, and knowing its limits is most of what separates a good answer. **A mode is a plurality, not a majority.** If the leading payment method holds 22% of transactions, 78% of customers used something else. Saying customers typically pay that way misrepresents the data. Quote the share alongside the label — leading method at 22% — and the sentence becomes true. **A mode can tie or be unstable.** Two categories at 21% and 22% are, for reporting purposes, indistinguishable, and next month's sample may swap them. Report both and call the distribution bimodal rather than presenting a coin flip as a finding. A genuine tie has the same treatment: the mode is a set, not necessarily a single value. **A mode ignores everything except the winner.** It says nothing about how concentrated or spread out the remaining categories are — one dominant option with a long tail of rare ones and a near-uniform spread across five options can share the same modal label. The **frequency table**, listing every category with its count and share, is the actual summary of a nominal column; the mode is the one-line headline you put on top of it. ## Beyond nominal data The mode is defined for every kind of column, but its usefulness varies. - **On ordered categories** it is still meaningful and the median becomes available too, giving you a real choice. - **On discrete counts** the mode is well behaved and can be informative, especially when one count value dominates — for example the number of items per order being overwhelmingly one. - **On continuous measurements** the raw mode is close to useless: every recorded value tends to be unique, so the most frequent value is an artefact of rounding or of measurement precision. Talking about the mode of a continuous variable really means talking about the location of the peak of its distribution, which is estimated from binned counts or a smoothed density rather than read off the raw values. Saying that plainly is a good sign in an interview. ## Practical reporting advice For a nominal column, produce the frequency table first: category, count, share, sorted by count descending, with rare categories optionally collapsed into an other bucket once you have checked that bucket is not hiding something interesting. Then lead with the mode and its share. If two shares are within sampling noise of each other, say so rather than declaring a winner. And watch the direction of the claim. The mode tells you which value is most common; it does not tell you what most users do, it does not tell you what a typical user does in any average sense, and it does not license comparison of one nominal column's centre to another's magnitude. Counting is the only operation the data supports, so counting is the only claim the summary can make.
- Someone computed a mean of 1.87 over payment methods stored as codes 1, 2 and 3. What is wrong?The arithmetic was performed on identifiers, not quantities. The codes are arbitrary labels, so renumbering the categories changes the mean without changing a single transaction. Any statistic that moves when you relabel is describing the encoding rather than the data. Replace it with a frequency table and the modal category with its share.
- Two payment methods hold 21% and 22% share. What do you report as the mode?Report both and describe the distribution as effectively bimodal. A one-point gap is within ordinary sampling variation, so declaring a single winner presents noise as a finding and is likely to reverse next period. Give the shares explicitly so the reader sees how close the contest is.
- Is the mode useful for a continuous measurement such as response time?Not from raw values, where nearly every observation is unique and the most frequent value reflects rounding rather than the data. The meaningful version is the location of the peak of the distribution, estimated from binned counts or a smoothed density. Say which binning or smoothing produced it, because the peak's position depends on that choice.
Asking for the mean of payment methods is like asking for the average colour of the cars in a car park by adding the names. You can count how many are red, and nothing else.
saying these in an interview costs you the question
- Computes a mean over category codes and reports it
- Claims the modal category is what most users do
- Picks one winner when two shares are effectively tied
- Says the median works if you sort the labels alphabetically
- Reports the mode of a continuous column read off raw values
- Reports only the mode and never the frequency table