A stakeholder asks for a model that flags 'high performers' but nobody has defined the term — how do you proceed?
answer
- the algorithm is not the blocker
- the target does not exist yet
- find an outcome already recorded
- ask two reviewers to label independently
- their disagreement caps the model
basics
~20 sAn undefined target is a definition problem, not a modelling problem. Find a recorded outcome that can stand in for 'high performer', name the bias that proxy encodes, test whether two reviewers agree, price the labelling, and consider a descriptive deliverable.
solid answer
~50 sThe ask has no target, so there is nothing supervised to build yet - the first work is definitional, not algorithmic. I look for an outcome already recorded that could stand in: promotion within eighteen months, retention, quota attainment. Then I interrogate the proxy, because a model trained on manager ratings learns to imitate managers, biases included, while being presented as if it measured performance. Next I check whether the concept is learnable: give two experienced reviewers the same fifty employees and see how often they agree. Weak agreement caps what any model can achieve on those labels and is itself worth reporting. In parallel I price the labelling. And I hold open the option that the honest deliverable is descriptive - here is how the workforce actually varies - rather than a prediction dressed up as ground truth.
go deeper
Recall that supervised learning cannot start without a target column, and that inventing one silently is the mistake. Say out loud that the definition has to come before any model choice.
Explain proxy targets concretely: name two or three recorded outcomes and state what a model trained on each would actually be learning to reproduce.
Demonstrate the cheap experiment — an inter-reviewer agreement study — and explain why the disagreement rate bounds achievable performance, plus how you would price the labelling before committing.
Own the call on whether the organisation should build this at all: what the proxy optimises, who is affected by the output, and when a descriptive deliverable is the responsible answer.
## The request is not yet a machine-learning problem Supervised learning needs a target column. An HR performance dataset where 'high performer' has never been agreed does not have one, and no choice of algorithm creates it. So the first phase of this work is definitional. Skipping it is the most common way this kind of project fails: a target gets picked silently — often whichever column was convenient — and every downstream number then quietly means something nobody agreed to. ## Step 1: find candidate proxies that were actually recorded Ask what outcome the organisation already writes down that correlates with what they mean. Typical candidates: promotion within a defined window, retention past a tenure threshold, attainment against a quota, a manager's annual rating. Each is a **proxy** — a recorded quantity standing in for an unrecorded concept. Write the candidates down and state, for each, what a model trained on it would actually learn. That sentence is the deliverable of this step, because it is what stakeholders will otherwise never hear: - Train on manager ratings and the model learns to *imitate managers*, inheriting whatever leniency, recency effects or group bias the ratings carry. - Train on promotion and it learns *who gets promoted*, which mixes performance with visibility, tenure, and headcount available in each department. - Train on retention and it learns *who stayed*, which is partly about the labour market and partly about who was not being managed out. None of these is 'performance'. Choosing one is a decision about what the organisation is willing to optimise, and it belongs to the stakeholder, made explicitly, not to the modeller, made by default. ## Step 2: test whether the concept is learnable at all Before funding anything, run a small agreement study. Hand two experienced reviewers the same fifty employee records with a written definition and have them label independently. Then measure how often they agree. This is the most informative cheap experiment available, because **inter-annotator agreement bounds achievable model performance**. If two humans applying the same definition disagree on 40% of cases, the labels contain that much irreducible disagreement, and a model scored against one reviewer's labels cannot do better than the other reviewer does. Low agreement is not a reason to give up quietly; it is a finding to report, and usually it means the written definition is under-specified. Sharpen the guidelines, rerun the study, and see whether agreement moves. If it does not move, the concept is not a target — it is an opinion. ## Step 3: price the labels A definition nobody will fund is not a definition. Estimate how many labelled cases the problem needs and what each costs in expert time. This is where supervised framing gets expensive: in digital pathology one label can cost a board-certified specialist twenty minutes, which is exactly why an archive of 90,000 slides carries only 400 labelled ones. HR labels are cheaper per unit but scarcer in a different way, because the population is small and each label consumes a senior reviewer's judgment. Put the number in front of the stakeholder alongside the proxy options. Very often the conversation ends with a better-scoped question, because the cost makes people say out loud what they actually wanted. ## Step 4: consider that the honest deliverable may be descriptive If no defensible target exists and none will be funded, the alternative is not to fake one. It is to change what you deliver: a description of how the workforce actually varies along recorded, uncontroversial dimensions, presented as structure rather than as a verdict on individuals. That answers 'what does our workforce look like' honestly, and it makes clear it is not answering 'who is good'. The failure to avoid here is inventing ground truth: grouping employees, declaring the most flattering group to be the high performers, and then training a classifier to reproduce that assignment. The result looks like a supervised model with an impressive score, but the score only says the classifier can reproduce the earlier partition. Circularity of that kind is very hard to spot once it has been through two pipeline stages and one slide deck. ## Step 5: ask whether the decision should be automated at all With people-affecting targets there is a legitimacy question that sits above the modelling one. Who sees the score, what action it triggers, whether the subject can contest it, and what the proxy's known biases would do to particular groups. A senior answer raises this without being asked, because it is the part a stakeholder cannot be expected to raise on their own. ## What a good answer sounds like in the room A strong candidate reframes the request within a minute — 'there is no target yet, so let me tell you what we would have to agree on' — offers two or three concrete recorded proxies with their distortions named, proposes the agreement study as the cheap next experiment, and states plainly that a descriptive deliverable is a legitimate outcome rather than a failure. The same pattern generalises past HR. Whenever an ask arrives with no recorded outcome behind it — a 40,000-hour archive of pump-vibration telemetry where nobody ever tagged which runs preceded a failure is the same situation with different vocabulary — the options are identical: find a recorded proxy, start instrumenting so the outcome exists in future data, or deliver structure instead of prediction.
- Two reviewers labelling the same fifty employees agree on only sixty percent. What do you do with that?Report it as the headline finding. That level of disagreement caps what any model scored against those labels can achieve, so promising high accuracy would be dishonest. Usually it means the written definition is under-specified, so I sharpen the guidelines, rerun the study on fresh cases, and see whether agreement improves. If it does not, the concept is an opinion, not a target.
- What is wrong with grouping the employees first and then treating the best group as the label?It invents ground truth. Training a classifier to reproduce a partition you created yourself yields a strong-looking score that measures only how reproducible your own grouping was, not whether it identifies anything real. Worse, the circularity becomes invisible once the pipeline has two stages, and the output is then presented as evidence about people.
- The stakeholder insists on manager ratings as the target because they are already in the warehouse. What do you say?I would accept it only with the meaning stated in writing: the model predicts manager ratings, so it reproduces manager behaviour including leniency, recency effects and any group bias in the historic ratings. Then I would check the rating distribution across managers and groups, and agree in advance how the output may be used given what it actually measures.
- How does this change when the target could exist but nobody recorded it, as with untagged equipment failures?The definitional problem shrinks and the data-collection problem grows. Everyone agrees what a failure is, so the fix is instrumentation: start logging failure events with timestamps so future data carries the target, and meanwhile deliver descriptive structure. You are buying a supervised capability that begins accruing today rather than pretending one exists already.
saying these in an interview costs you the question
- Picks a convenient column as the target without saying so
- Jumps straight to model choice and tuning
- Treats manager ratings as objective performance
- Clusters the data and calls one group the positives
- Never asks what labelling would cost or who does it
- Ignores that the subjects are people who can be harmed