skip to content

Task Types and Targets

What the model must output - a class, a number, an ordering, an outlier flag, or a future value - and how that choice fixes the loss and the metric downstream. Interviewers start here.

on this pageshow

explore

questions

13

Is a 1-to-5 star satisfaction rating a classification or a regression target?

level: juniorimportance: must knowfreq 72%

answer

  1. ask what the numbers really represent
  2. ordered - but equally spaced?
  3. regression assumes 1-to-2 equals 4-to-5
  4. five nominal classes throw away order
  5. cumulative at-least-k cutpoints keep both

basics

~20 s

A 1-to-5 star rating is an ordinal target: the values are ordered, but the gaps between them are not guaranteed equal. Regression uses the order and assumes equal spacing; five-class classification keeps the classes but discards the order.

solid answer

~50 s

It is an ordinal target, which sits between the two standard task types. Regression on the numbers 1 to 5 exploits the ordering - predicting 4 when the truth is 5 is a smaller error than predicting 1 - but it assumes the step from 1 to 2 star equals the step from 4 to 5, which is rarely true of human scales, and it produces fractional, unbounded predictions you then have to round and clip. Five unordered classes keep the discreteness and gives a probability for each star, but the model is penalised identically for guessing 4 or 1 when the truth is 5. Ordinal framings cut the scale at a series of cutpoints and model cumulative `at least k stars` decisions, keeping the order without asserting equal spacing. I choose by what the downstream decision consumes: a number to average, or a distribution over stars.

go deeper

for a junior

Be ready to name the three target shapes - unordered categories, ordered categories, and evenly spaced numbers - and place a star rating in the middle one. Say out loud that ordinal targets are the reason this question has no one-word answer.

for a middle

Explain the mechanics of each framing: what regression assumes about spacing, what unbounded output forces you to clip and round, and exactly which information a five-class model discards. Mention the cumulative at-least-k ordinal framing as the compromise.

for a senior

Show that you pick the framing from the downstream decision, not from habit, and that you would check where a nominal model's errors land before committing. Note that averaging a bimodal rating distribution produces a number describing no real customer.

for a principal

Own the argument that target granularity is a product decision with a cost: every extra level you model is a distinction someone must act on, label consistently and monitor. Be ready to defend collapsing a rating scale even though it looks like losing information.

## The three shapes a target can have The *target* (or label) is what the model must output for each example. Before picking an algorithm, classify the target itself: - **Nominal** - unordered categories. Car colour recorded as red, blue or white; the country a shipment came from. There is no sense in which blue is *between* red and white. - **Ordinal** - ordered categories whose spacing is unknown or unequal. A 1-to-5 star rating, a survey scale from strongly disagree to strongly agree, a credit grade from A to E. You can say 4 is more than 3, but you cannot say the distance from 1 to 2 equals the distance from 4 to 5. - **Continuous (interval or ratio)** - ordered *and* evenly spaced, so arithmetic is meaningful. Minutes, kilograms, currency. A star rating is squarely ordinal. That is why the question has no one-word answer: classification and regression are the two standard task types, and ordinal targets have a foot in each. ## Framing 1: regression on the numbers 1 to 5 You fit a model that outputs a real number and train it to be close to the observed star. **What you gain.** The ordering is used automatically - a prediction of 4.6 against a true 5 is scored as a near miss, while a prediction of 1.2 is scored as a bad one. That is usually the behaviour you want, and it makes the model far more data-efficient than one that treats the five stars as unrelated symbols. **What you assume.** That the scale is evenly spaced. On real rating scales it is not: the latent satisfaction gap between 1 and 2 stars is typically much smaller than between 4 and 5, because ratings pile up at the extremes. Human ratings on marketplaces are famously J-shaped - mostly 5s, a lump of 1s, and very little in the middle. **What you must clean up.** The output is unbounded, so the model can predict 5.4 or 0.7; you have to clip. And if the product shows whole stars, you have to convert - and rounding at the .5 boundaries is a convention, not an optimum. The cutoffs that maximise agreement with the true star are usually not the halfway points. ## Framing 2: five unordered classes You fit a model that outputs a probability for each of the five stars. **What you gain.** Discreteness is preserved, and you get a full distribution rather than a single number. That matters more than people expect: a polarising product with half its ratings at 1 and half at 5 has a mean near 3, which describes nobody. The class distribution shows the bimodality; the mean hides it. **What you lose.** The order. The training signal treats a predicted 4 and a predicted 1 as equally wrong when the truth is 5, so the model has to rediscover the ordering from data instead of being handed it. You also need enough examples of each star, and the middle classes are usually the sparse ones. ## Framing 3: ordinal modelling The compromise fits a single underlying score plus a set of increasing cutpoints, and models the cumulative decisions `at least 2 stars`, `at least 3 stars`, `at least 4 stars`, `at least 5 stars`. Subtracting adjacent cumulative probabilities gives a probability per star that is guaranteed to respect the ordering, and the cutpoints are learned rather than assumed evenly spaced. This keeps the order without the equal-spacing claim - exactly what an ordinal target needs. ## Framing 4: collapse it If the decision downstream is binary - route anything at 1 or 2 stars to a support agent - then the honest target is binary, and modelling five classes to then threshold them adds variance for no benefit. Match the granularity of the target to the granularity of the decision. ## How to choose in an interview answer Ask what consumes the prediction: - A single number that will be averaged across many items (an expected rating shown on a listing) - regression, or the expectation taken over a classification model's distribution. - A distribution shown to a human, or a decision that depends on the shape - classification or ordinal. - A yes/no action - collapse to binary. A useful diagnostic: fit the five-class nominal model and look at where its mistakes land. If nearly all errors are adjacent to the correct star, the ordering is real and an ordinal or regression framing will use your data better. If the errors scatter uniformly across classes, the numeric labels may not be carrying order at all - which is itself worth knowing before you assume they do.

  • Your model outputs 3.7 stars but the interface only shows whole stars - how do you turn that into a prediction?
    Decide first whether you need the expected rating or the most likely rating. If you need a displayed star, pick cutoffs that maximise agreement with the observed labels rather than rounding at the halfway points - on a J-shaped scale the boundary that separates 4 from 5 is usually not at 4.5. If you need an average across many items, keep the fractional value and never round at all.
  • You fit five unordered classes and almost every error is one star away from the truth - what does that tell you?
    That the ordering is genuinely present in the data and the model is recovering it the hard way, from scratch. That is evidence to switch to an ordinal or regression framing, which hands the model the ordering for free and so needs less data to reach the same accuracy - particularly in the sparse middle classes.
  • When would you deliberately collapse a 5-star rating into a binary target?
    When the downstream action is binary - escalate unhappy customers, or flag a listing for review. Collapsing concentrates all the data into two classes, removes the sparse middle categories, and stops you optimising distinctions nobody acts on. The cost is that you can no longer report a granular rating, so keep the raw stars stored even if the model does not use them.

Nominal labels are boxes on a shelf, continuous targets are marks on a ruler. An ordinal scale is a staircase: you know which step is higher, but not that every step is the same height.

saying these in an interview costs you the question

  • Says any target with few distinct values must be classification
  • Assumes the gap from 1 to 2 stars equals the gap from 4 to 5
  • Rounds regression output at the halfway points without checking
  • Claims five-way classification uses the ordering of the stars
  • Confuses the mean predicted rating with the most likely rating

context

open as a page

How do you turn a raw time series into a supervised table of lag and rolling-window features?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Build one row per time step whose features use only past values — lags such as 1, 7 and 14 steps back, plus rolling means or spreads over trailing windows — and whose label is the value at the forecast horizon.

open as a page

Why is seasonal-naive the first baseline you fit for a 48-hour electricity-load forecast?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Because it sets the bar almost for free. Seasonal-naive predicts each future hour with the load observed at the same hour one week earlier, capturing the daily and weekly cycles without any training; a model that cannot beat it is adding nothing.

open as a page

An article can carry any of 20 topic tags at once - how should you frame that target?

level: middleimportance: must knowfreq 66%

basics

~20 s

That is a multilabel target: the tags are not mutually exclusive, so an article can carry zero, one or several. Frame it as 20 independent yes/no decisions with a probability per tag, not one distribution over 20 tags summing to one.

open as a page

When do you pose a triage task as ranking a queue rather than classifying each case?

level: middleimportance: must knowfreq 62%

basics

~20 s

Pose it as ranking when a fixed review capacity, not a decision rule, consumes the output. Only the relative order near the top then changes outcomes, so absolute score values and a class boundary buy nothing.

open as a page

When do you pose anomaly detection as one-class modelling of normal data instead of supervised classification?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Choose one-class modelling when the anomalies you will face are not represented by the anomalies you have labelled, so no boundary learned from past examples transfers. Choose supervised classification when labelled anomalies are plentiful and representative of what comes next.

open as a page

When would you use direct multi-step forecasting instead of feeding predictions back in recursively?

level: middleimportance: should knowfreq 52%

basics

~20 s

Recursive forecasting iterates one one-step model, feeding its predictions back as inputs, so errors compound. Direct forecasting fits a separate model per horizon on real observed history: no compounding, but many models and thinner data.

open as a page

Should you predict delivery ETA in minutes and threshold it, or classify 'late by 30+ minutes' directly?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Model the number when you need the ETA itself or the cutoff may move; classify the flag directly when the cutoff is fixed and only the decision matters. A regressor optimises accuracy everywhere, not around the 30-minute boundary.

open as a page

Your 06:00 taxi-demand forecast lacks the last three hours of counts — how do you frame the task?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Define features against the last landed observation, not the clock. With a three-hour delay the shortest usable lag is four hours, and training rows must be rebuilt at that cutoff so the model never relies on values serving lacks.

open as a page

As reviewed labels accumulate from an anomaly-alert queue, should the task be re-posed as supervised classification?

level: principalimportance: should knowfreq 28%

basics

~20 s

Only with a correction for how those labels were produced. Reviewers judged only what the incumbent detector surfaced, so a classifier trained on them learns to imitate that detector and stays blind wherever it never looked.

open as a page

Hourly bike-dock rental counts are your target - what breaks if you treat them as plain continuous regression?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A count is a non-negative integer, usually right-skewed with many zeros and a spread that grows with its own level. Plain squared-error regression predicts unbounded numbers and assumes constant spread, so it emits negative counts and lets busy hours dominate.

open as a page

How does graded relevance change a search-ranking task compared with binary relevant-or-not labels?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Graded labels make the target an ordering among relevant items rather than a set of hits. With perfect, good, marginal and bad grades, ranking a marginal result above a perfect one is an error binary labels cannot express.

open as a page

Would you fit one global model across 4,000 SKUs in 60 stores, or one model per series?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Almost always one global model. Pooling 240,000 store-SKU series into a single training table, with store and product identity as features, shares patterns across short and sparse histories, handles new products, and leaves one artifact to operate instead of 240,000.

open as a page