How do supervised and unsupervised learning differ in what the training data must contain?
answer
- one setup has an answer key
- look for a target column
- predicting versus describing
- which one can report accuracy
basics
~20 sSupervised learning trains on examples that each carry a known target value and learns to predict it. Unsupervised learning has no target; it describes structure in the inputs alone - groups, directions of variation, or where the data is dense.
solid answer
~40 sThe split is defined by one thing: whether every training example carries a target value. In supervised learning each row is a pair of inputs and a known answer, and the model learns a mapping from inputs to that answer, so you can hold out rows and measure how often the prediction matches the truth. In unsupervised learning there is no answer column, so there is nothing to predict; the goal is to describe the inputs — partition them into groups, express them in fewer coordinates, or model where they are dense. The practical consequence is evaluation. Supervised work gets a held-out score. Unsupervised work has no ground truth to score against, so its output has to be judged by stability, expert review, or the value of the decision it feeds.
go deeper
Be ready to state the split in one sentence — target column present or absent — and to name clustering and dimensionality reduction as unsupervised, classification and regression as supervised.
Explain why the split is about the data rather than the algorithm, using a pair like k-nearest-neighbours and k-means, and spell out why only one of the two regimes has a held-out score.
Show that you treat the target column as something someone has to fund, define and maintain, and that you know an unsupervised result still needs a validation story before anyone acts on it.
Own the framing call: whether a business question should be answered predictively at all, what a labelling programme would cost, and what the organisation gives up in measurability by staying descriptive.
## The line is drawn by the target, not by the algorithm Every training dataset can be thought of as a table. The columns you are allowed to look at when a new case arrives are the **features** (inputs). A column holding the thing you wish you could know in advance — will this pump fail, is this slide malignant, what will this house sell for — is the **target** (also called the label, the response, or `y`). **Supervised learning** is learning from a table where the target column is present and filled in. The learner sees pairs `(x, y)` and searches for a function that maps `x` to `y` well enough to work on cases it has never seen. Supervised problems split further by the type of the target: a categorical target makes it classification, a continuous one makes it regression. **Unsupervised learning** is learning from a table with no target column at all. Nothing says which rows are 'right'. The learner's job is descriptive: say something true about the shape of the inputs. Three families are standard: - **Clustering** — partition rows into groups whose members resemble each other more than they resemble members of other groups. - **Dimensionality reduction** — re-express many correlated columns as a few new coordinates that keep most of the variation. - **Density estimation** — model where in feature space examples tend to fall, so you can say a new point is typical or unusual. Notice that all three describe the inputs. None of them predicts a quantity that somebody could have written down in advance and checked. ## The distinction is about the data, not the model family A common confusion is to attach the labels 'supervised' and 'unsupervised' to algorithms or to sophistication. They belong to the *problem setup*. The clearest illustration is two methods that both rest on distance and both have a number called k, and land on opposite sides of the line: k-nearest-neighbours is supervised — it looks up the labelled neighbours of a new point and votes on their known targets — while k-means is unsupervised, grouping points with no targets anywhere in sight. Same geometry, different regimes, because one has an answer column and one does not. The same raw data can also support both, at different times. A 40,000-hour archive of pump-vibration telemetry in which no engineer ever tagged which runs preceded a failure has no target column, so only unsupervised questions are askable of it today: which operating regimes exist, how many distinct vibration signatures there are. The moment maintenance starts logging failure events against timestamps, the same feature columns acquire a target and the same archive supports a supervised question. ## Why the distinction matters in practice **Evaluation.** This is the consequence that bites hardest. Supervised learning comes with an honest scoreboard: hide some labelled rows, predict them, count the errors. Unsupervised learning has no such scoreboard, because there is no truth to compare against. You cannot report the 'accuracy' of a clustering of unlabelled data — the phrase has no referent. Judgment has to come from elsewhere: does the structure survive resampling, does a domain expert recognise it, does acting on it improve a measurable outcome. **Cost.** Labels are an asset that is bought, not found. In digital pathology one label can cost a board-certified pathologist twenty minutes of reading, which is why an archive of 90,000 slides may have only 400 labelled ones. Framing a problem as supervised commits somebody to producing that column, at that price, with a written definition of what counts. Framing it as unsupervised avoids the bill but changes the question you are able to answer. **Assumptions.** Unsupervised does not mean assumption-free or human-free. You still choose which features go in, how they are scaled — which silently sets how much each feature contributes to any notion of similarity — and what 'similar' means at all. Those choices determine the structure that comes out. What you lose relative to supervised learning is not the need for judgment; it is the ability to be measurably wrong. ## The boundary is a spectrum, not a wall Between the two extremes sit regimes with partial signal: a few labels plus a large unlabelled pool, labels invented from the data's own structure, and reward-only feedback from acting. Those are separate topics with their own methods. For framing purposes the useful reflex remains the simple one: before choosing any algorithm, ask what the target column is, whether it exists, and who is going to fill it in.
- Which side of the line do clustering, dimensionality reduction and density estimation fall on, and why?All three are unsupervised: none of them consumes a target column. Clustering partitions rows by similarity, dimensionality reduction re-expresses many correlated columns as a few coordinates that retain most of the variation, and density estimation models where in feature space examples tend to fall. Each describes the inputs rather than predicting an answer somebody could have written down beforehand.
- Can the same dataset support both a supervised and an unsupervised task?Yes, and the trigger is whether a target has been recorded. A vibration archive with no failure tags supports only descriptive questions — which operating regimes exist, how many distinct signatures. Once failures are logged against timestamps, the very same feature columns gain a target and support a supervised question. The features did not change; the availability of an answer column did.
- Does unsupervised learning mean the method makes no human assumptions?No. You still choose which features enter, how they are scaled, and what counts as similar — and scaling in particular silently decides how much each feature contributes. Different choices produce different structure from identical data. What you give up compared with supervised learning is not judgment but the ability to be measurably wrong against a held-out truth.
Supervised learning is studying with the answer key printed at the back of the book. Unsupervised learning is being handed the questions only and asked what they have in common.
saying these in an interview costs you the question
- Says unsupervised data has no features, only rows
- Calls any non-neural model unsupervised
- Thinks unsupervised learning needs no human choices
- Claims regression is unsupervised because the output is continuous
- Reports an accuracy figure for a clustering of unlabelled data