skip to content

Classification and Regression

Binary, multiclass and multilabel classification versus predicting a number, plus ordinal and count targets that fit neither cleanly. Interviewers probe when a task is framed as the wrong one.

on this pageshow

questions

4

Is a 1-to-5 star satisfaction rating a classification or a regression target?

level: juniorimportance: must knowfreq 72%

answer

  1. ask what the numbers really represent
  2. ordered - but equally spaced?
  3. regression assumes 1-to-2 equals 4-to-5
  4. five nominal classes throw away order
  5. cumulative at-least-k cutpoints keep both

basics

~20 s

A 1-to-5 star rating is an ordinal target: the values are ordered, but the gaps between them are not guaranteed equal. Regression uses the order and assumes equal spacing; five-class classification keeps the classes but discards the order.

solid answer

~50 s

It is an ordinal target, which sits between the two standard task types. Regression on the numbers 1 to 5 exploits the ordering - predicting 4 when the truth is 5 is a smaller error than predicting 1 - but it assumes the step from 1 to 2 star equals the step from 4 to 5, which is rarely true of human scales, and it produces fractional, unbounded predictions you then have to round and clip. Five unordered classes keep the discreteness and gives a probability for each star, but the model is penalised identically for guessing 4 or 1 when the truth is 5. Ordinal framings cut the scale at a series of cutpoints and model cumulative `at least k stars` decisions, keeping the order without asserting equal spacing. I choose by what the downstream decision consumes: a number to average, or a distribution over stars.

go deeper

for a junior

Be ready to name the three target shapes - unordered categories, ordered categories, and evenly spaced numbers - and place a star rating in the middle one. Say out loud that ordinal targets are the reason this question has no one-word answer.

for a middle

Explain the mechanics of each framing: what regression assumes about spacing, what unbounded output forces you to clip and round, and exactly which information a five-class model discards. Mention the cumulative at-least-k ordinal framing as the compromise.

for a senior

Show that you pick the framing from the downstream decision, not from habit, and that you would check where a nominal model's errors land before committing. Note that averaging a bimodal rating distribution produces a number describing no real customer.

for a principal

Own the argument that target granularity is a product decision with a cost: every extra level you model is a distinction someone must act on, label consistently and monitor. Be ready to defend collapsing a rating scale even though it looks like losing information.

## The three shapes a target can have The *target* (or label) is what the model must output for each example. Before picking an algorithm, classify the target itself: - **Nominal** - unordered categories. Car colour recorded as red, blue or white; the country a shipment came from. There is no sense in which blue is *between* red and white. - **Ordinal** - ordered categories whose spacing is unknown or unequal. A 1-to-5 star rating, a survey scale from strongly disagree to strongly agree, a credit grade from A to E. You can say 4 is more than 3, but you cannot say the distance from 1 to 2 equals the distance from 4 to 5. - **Continuous (interval or ratio)** - ordered *and* evenly spaced, so arithmetic is meaningful. Minutes, kilograms, currency. A star rating is squarely ordinal. That is why the question has no one-word answer: classification and regression are the two standard task types, and ordinal targets have a foot in each. ## Framing 1: regression on the numbers 1 to 5 You fit a model that outputs a real number and train it to be close to the observed star. **What you gain.** The ordering is used automatically - a prediction of 4.6 against a true 5 is scored as a near miss, while a prediction of 1.2 is scored as a bad one. That is usually the behaviour you want, and it makes the model far more data-efficient than one that treats the five stars as unrelated symbols. **What you assume.** That the scale is evenly spaced. On real rating scales it is not: the latent satisfaction gap between 1 and 2 stars is typically much smaller than between 4 and 5, because ratings pile up at the extremes. Human ratings on marketplaces are famously J-shaped - mostly 5s, a lump of 1s, and very little in the middle. **What you must clean up.** The output is unbounded, so the model can predict 5.4 or 0.7; you have to clip. And if the product shows whole stars, you have to convert - and rounding at the .5 boundaries is a convention, not an optimum. The cutoffs that maximise agreement with the true star are usually not the halfway points. ## Framing 2: five unordered classes You fit a model that outputs a probability for each of the five stars. **What you gain.** Discreteness is preserved, and you get a full distribution rather than a single number. That matters more than people expect: a polarising product with half its ratings at 1 and half at 5 has a mean near 3, which describes nobody. The class distribution shows the bimodality; the mean hides it. **What you lose.** The order. The training signal treats a predicted 4 and a predicted 1 as equally wrong when the truth is 5, so the model has to rediscover the ordering from data instead of being handed it. You also need enough examples of each star, and the middle classes are usually the sparse ones. ## Framing 3: ordinal modelling The compromise fits a single underlying score plus a set of increasing cutpoints, and models the cumulative decisions `at least 2 stars`, `at least 3 stars`, `at least 4 stars`, `at least 5 stars`. Subtracting adjacent cumulative probabilities gives a probability per star that is guaranteed to respect the ordering, and the cutpoints are learned rather than assumed evenly spaced. This keeps the order without the equal-spacing claim - exactly what an ordinal target needs. ## Framing 4: collapse it If the decision downstream is binary - route anything at 1 or 2 stars to a support agent - then the honest target is binary, and modelling five classes to then threshold them adds variance for no benefit. Match the granularity of the target to the granularity of the decision. ## How to choose in an interview answer Ask what consumes the prediction: - A single number that will be averaged across many items (an expected rating shown on a listing) - regression, or the expectation taken over a classification model's distribution. - A distribution shown to a human, or a decision that depends on the shape - classification or ordinal. - A yes/no action - collapse to binary. A useful diagnostic: fit the five-class nominal model and look at where its mistakes land. If nearly all errors are adjacent to the correct star, the ordering is real and an ordinal or regression framing will use your data better. If the errors scatter uniformly across classes, the numeric labels may not be carrying order at all - which is itself worth knowing before you assume they do.

  • Your model outputs 3.7 stars but the interface only shows whole stars - how do you turn that into a prediction?
    Decide first whether you need the expected rating or the most likely rating. If you need a displayed star, pick cutoffs that maximise agreement with the observed labels rather than rounding at the halfway points - on a J-shaped scale the boundary that separates 4 from 5 is usually not at 4.5. If you need an average across many items, keep the fractional value and never round at all.
  • You fit five unordered classes and almost every error is one star away from the truth - what does that tell you?
    That the ordering is genuinely present in the data and the model is recovering it the hard way, from scratch. That is evidence to switch to an ordinal or regression framing, which hands the model the ordering for free and so needs less data to reach the same accuracy - particularly in the sparse middle classes.
  • When would you deliberately collapse a 5-star rating into a binary target?
    When the downstream action is binary - escalate unhappy customers, or flag a listing for review. Collapsing concentrates all the data into two classes, removes the sparse middle categories, and stops you optimising distinctions nobody acts on. The cost is that you can no longer report a granular rating, so keep the raw stars stored even if the model does not use them.

Nominal labels are boxes on a shelf, continuous targets are marks on a ruler. An ordinal scale is a staircase: you know which step is higher, but not that every step is the same height.

saying these in an interview costs you the question

  • Says any target with few distinct values must be classification
  • Assumes the gap from 1 to 2 stars equals the gap from 4 to 5
  • Rounds regression output at the halfway points without checking
  • Claims five-way classification uses the ordering of the stars
  • Confuses the mean predicted rating with the most likely rating

context

open as a page

An article can carry any of 20 topic tags at once - how should you frame that target?

level: middleimportance: must knowfreq 66%

basics

~20 s

That is a multilabel target: the tags are not mutually exclusive, so an article can carry zero, one or several. Frame it as 20 independent yes/no decisions with a probability per tag, not one distribution over 20 tags summing to one.

open as a page

Should you predict delivery ETA in minutes and threshold it, or classify 'late by 30+ minutes' directly?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Model the number when you need the ETA itself or the cutoff may move; classify the flag directly when the cutoff is fixed and only the decision matters. A regressor optimises accuracy everywhere, not around the 30-minute boundary.

open as a page

Hourly bike-dock rental counts are your target - what breaks if you treat them as plain continuous regression?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A count is a non-negative integer, usually right-skewed with many zeros and a spread that grows with its own level. Plain squared-error regression predicts unbounded numbers and assumes constant spread, so it emits negative counts and lets busy hours dominate.

open as a page