A clustering sweep suggests k=9 but operations can staff only 4 playbooks - how do you set k?
answer
- the sweep advises, operations constrains
- k is a design parameter
- cluster fine, map many-to-one
- price the cap in business units
- a mapping is cheaper to change than a model
basics
~20 sThe staffing ceiling is a hard constraint and the sweep is advisory, so k is a design decision. Usually the best move is to cluster at the finer k and define a many-to-one mapping from those groups onto the four playbooks, keeping the structure while satisfying operations.
solid answer
~50 sStart from the fact that k is a design parameter, not an estimate: no curve knows the organisation can run four handling playbooks. Two workable routes. Force k=4 directly, which is simple to explain but discards the finer structure and produces groups that may be internally mixed on exactly the variables the playbook keys on. Or cluster at 9 and define an explicit many-to-one mapping onto the four playbooks - this keeps the finer analysis alive, lets you re-map without re-clustering when capacity changes, and makes the compromise visible instead of silent. I would usually take the second, and measure what the cap costs: compare within-group homogeneity on the decision-relevant variables under both, and where possible run the downstream metric. That measurement is also the negotiation material - if the fifth playbook is worth a large share of the gap, the ceiling itself is the thing to challenge.
go deeper
Understand that the number of groups shipped is a business decision, and that a group nobody can act on differently is not a useful group however clean the maths looks.
Be able to describe both routes - clustering at the ceiling versus clustering finer and mapping onto it - and state what each gains and gives up.
Show how you would price the cap: within-group spread on decision-relevant variables plus a downstream replay or test, and who owns the mapping once it is deployed.
Own the negotiation. Turn the granularity question into a marginal-value-per-playbook curve in business units, and set the governance for revisiting the cap as capacity changes.
## The framing that earns the question Curves suggest; they do not decide. A clustering exists to change what someone does, and if only four distinct things can be done, then four is a hard constraint on the *deployment*. The senior mistake is to argue with the constraint using statistics. The principal move is to make the constraint explicit, choose a structure that respects it, and quantify what it costs so the constraint can be renegotiated with evidence later. ## Route A: cluster directly at the ceiling Set k=4 and ship. Advantages: one artefact, one label per record, nothing to explain twice, and the groups are optimised for the number of treatments that actually exist. Disadvantages: you have deliberately thrown away structure you know is there, and the four groups will be internally heterogeneous - possibly on exactly the variables the playbook depends on, which is where a coarse segmentation quietly fails. It is the right choice when the finer structure has no analytical use of its own and nobody will ever want to look under the four. ## Route B: cluster fine, then map Cluster at the k the data supports (9 here), then define an explicit mapping from the nine groups onto the four playbooks. This is usually the stronger design: - The mapping is a **business artefact you can change in an afternoon**; re-clustering is a project. When capacity rises to five playbooks, or one playbook is retired, you re-map. - It keeps the fine groups available for analysis, forecasting and capacity planning even though only four treatments exist today. - It makes the compromise visible. Anyone can see which two fine groups were merged and challenge that specific merge, instead of arguing about a number. The conditions it requires: the nine must group *naturally* into four - ideally three or four clear families with a couple of borderline cases. Check by cross-tabulating the nine-way labels against a four-way solution, and by looking at which fine groups sit closest to each other. If the nine do not fall into four families at all, the mapping will be arbitrary and you are carrying two artefacts for no benefit; take Route A instead. A second real risk: two artefacts drift. The nine-way model and the mapping have to be versioned and re-validated together, and someone has to own the mapping. Without that ownership, the mapping quietly becomes stale while the model is retrained. ## Measuring the cost of the cap Whatever route you take, quantify what four costs relative to nine so the ceiling can be revisited on evidence: - **Homogeneity on decision-relevant variables.** Not overall inertia - the spread within each shipped group on the two or three variables the playbook actually keys on. If a playbook assumes short handle times and its group spans the full range, the cap is hurting. - **Downstream metric.** Where you can run it - offline replay or a small live test - compare the outcome the clustering feeds (routing accuracy, resolution time, conversion) under the coarse and fine groupings. This is the only measurement a sponsor really weighs. - **Marginal value of one more playbook.** Present it as a curve of value against the number of playbooks, which is the same diminishing-returns argument the elbow makes, but in units the business owns. That is what turns "we want more clusters" into a staffing decision someone can approve. ## Other constraints that cap k, beyond staffing - **Minimum viable group size.** A group too small to staff, to measure, or to reach a decision-grade sample in a test is not a usable segment however clean it looks. - **Explainability.** People must be able to hold the groups in their head and act on them consistently; past roughly a handful, adherence drops regardless of statistics. - **Maintenance and governance.** Each group carries content, rules, monitoring and review. Two more groups is real recurring cost, not a one-off. - **Fairness and regulatory exposure.** More granular treatment differentiation means more surface area to justify - some distinctions you can measure you may not be permitted to act on. ## What to say in the interview Lead with "k is a decision, and the operational ceiling is part of the decision, not an obstacle to it". Then give both routes, name the condition that selects between them (do the nine fall naturally into four families?), and finish on measurement - the cost of the cap in the business's own units, which is what makes the ceiling revisable rather than permanent.
- How would you measure what forcing k from 9 down to 4 actually costs?Two ways. Internally, compare the spread within each shipped group on the specific variables the playbook keys on - not overall inertia, which is not in business units. Externally, replay or A/B the downstream outcome under both groupings. The second is what a sponsor weighs, and it converts 'we want more clusters' into a priced staffing decision.
- When is clustering at 9 and mapping to 4 worse than clustering at 4 directly?When the nine do not fall into four natural families, so the mapping is arbitrary - you then carry two artefacts, a model and a mapping, with no analytical gain. It is also worse when nobody owns the mapping: it drifts out of date while the model is retrained, and the deployed grouping stops matching the documented one.
- How would you argue for raising the ceiling from four playbooks to five?With the marginal-value curve, not the clustering curve. Show what the fifth group would receive that differs materially, estimate the outcome lift from the replay or test, and set it against the staffing, content and governance cost of one more playbook. A sponsor approves a priced increment; they do not approve a silhouette peak.
A menu can be printed with four dishes while the kitchen still tracks nine recipes; the grouping shown to customers is a business choice, and it can be regrouped without rewriting the recipes.
saying these in an interview costs you the question
- Insists the statistical k must be shipped
- Treats the staffing cap as an obstacle to argue away
- Merges fine groups with no stated mapping rule
- Ignores minimum group size for staffing and testing
- Prices the compromise only in inertia, not outcomes