skip to content

After mining 2 million basket rules, how do you decide which few are worth acting on?

level: seniorimportance: should knowfreq 46%

answer

  1. counts, not percentages
  2. millions of rules select their own extremes
  3. shorter rule, same confidence, drop the longer
  4. stability on a later window
  5. co-occurrence is not an intervention

basics

~20 s

Filter on absolute joint basket counts, not percentages; rank by lift with a leverage floor; drop rules that are redundant given a shorter rule; then test the survivors, because a co-occurrence rule does not predict what happens when you intervene.

solid answer

~50 s

Two million rules is a search result, not a finding. First cut on **volume**: convert support to an absolute basket count and drop anything backed by too few baskets, because with millions of rules extreme lift values arise by chance from tiny counts. Next cut on **redundancy**: if `{A, B} -> {C}` has the same confidence as `{A} -> {C}`, the extra condition adds nothing and only the shorter rule survives. Then cut on **actionability**: the antecedent must be something you can trigger on, the consequent something with margin, and the pairing must not be already obvious to every category manager. Finally, validate. Re-mine on a later time window and see whether the rule survives, then run a live experiment, because a rule says the items co-occur - it does not say that moving them together, or discounting one, changes anything.

go deeper

for a junior

Know that a mining run produces far more rules than anyone can use, and that filtering them is the real work. Absolute basket counts matter more than impressive-looking percentages.

for a middle

Be able to name concrete filters and apply them in order: volume floor, redundancy against shorter rules, then business actionability. Explain why lift alone ranks badly on a large rule set.

for a senior

Show the intervention gap explicitly. A rule describes baskets under the layout and promotions in force when they were recorded, so a shortlist is a set of hypotheses that needs a stability check and then a live test.

for a principal

Own the pipeline that turns mining into decisions: who sets thresholds, what the volume floor is, which rules get experiments, and what evidence a rule needs before it changes a planogram. Without that, rule mining generates decks rather than value.

## Why the output volume is the problem A mining run's rule count is set by the thresholds, not by the data's information content. Lower the minimum support and the count grows by orders of magnitude. Two million rules cannot be read, so the only question that matters is what filter turns them into a shortlist somebody can act on. Work through four cuts in order: statistical reliability, redundancy, actionability, and finally causal validity. ## Cut 1 - reliability, measured in baskets Support is a fraction, and fractions hide sample size. Convert every rule to the absolute number of baskets behind it before judging anything. Take a rare rule: `{gluten-free flour} -> {xanthan gum}` at 0.02% support with lift 400. The instinct is to dismiss it as noise. But over ten million baskets, 0.02% is two thousand baskets - a large, stable count, and a lift of 400 on that volume is a real niche relationship worth a shelf decision. The same rule over five thousand baskets is a single basket, and its lift of 400 is arithmetic on a coincidence. Identical metrics, opposite conclusions, and only the absolute count distinguishes them. This matters more than usual here because of multiplicity. When you evaluate millions of candidate rules, the most extreme lift values in the output are selected *because* they are extreme, and at small counts the extremes are dominated by chance. A minimum joint-basket floor - some teams also hold out a slice of transactions and require the rule to clear its thresholds there too - is the cheapest defence. ## Cut 2 - redundancy Rule sets are massively redundant, because every frequent itemset produces rules for every split and every subset of a frequent itemset is itself frequent. Two standard tests: - **Longer antecedent, no gain.** If `confidence({A, B} -> {C})` is not meaningfully above `confidence({A} -> {C})`, then B is a passenger. Keep the shorter rule. - **Consequent already implied.** If the extra item on the right is nearly ubiquitous among the antecedent's baskets anyway, the rule restates the item's popularity. Mining closed itemsets rather than all frequent itemsets removes a large share of this before rules are ever formed, and it is cheaper than filtering two million rules afterwards. ## Cut 3 - actionability A statistically fine rule can still be useless. - **Can you trigger on the antecedent?** A rule whose left-hand side is a basket you only observe at the till cannot drive an online recommendation. - **Is there margin in the consequent?** Promoting a low-margin staple because it pairs with something is often value-destroying even when the rule is real. - **Is it already known?** Category managers know that pasta pairs with sauce. A rule set that surfaces only what the business already believes is producing no decisions, however good the metrics look. - **Is the antecedent something you can change?** Rules whose left-hand side is a season or a store format are descriptive, not levers. ## Cut 4 - the intervention gap This is the one that separates a senior answer. The beer-and-nappies story - the retail legend that a chain discovered the two items sold together and moved them adjacent - gets repeated as proof of the technique's value, usually with no leverage figure, no basket count and no account of whether sales actually moved afterwards. That is the shape of the trap: a rule is a statement about baskets that already happened, under the store layout, assortment and promotions that were in force when they happened. Acting on a rule is an **intervention**, and the rule does not predict its effect. Three specific ways it goes wrong: 1. **The co-purchase was already happening.** Shoppers who want both already buy both, so moving the products adjacent or bundling them changes nothing except perhaps the discount you now pay on purchases you would have got anyway. 2. **A common cause drives both.** A promotion, a season, a store format or a weekly shop pattern puts both items in the basket. Change one and the association evaporates because it never ran through the other item. 3. **The rule reflects the current layout.** Items placed near each other co-occur partly because they are near each other. Rules mined from that data then recommend the layout that produced them. The answer is to treat the shortlist as hypotheses. Re-mine on a later window and keep only rules that persist - a stability check costs nothing and kills most spurious survivors. Then run the shortlist as a proper experiment: a store or user holdout, a pre-agreed metric (incremental basket value, not attach rate, which will rise by construction), and a decision rule fixed before the test. Rules that survive both are the handful you ship. ## What to report For each surviving rule, report the absolute joint basket count, support, confidence, lift and leverage together, plus the stability check and the expected value of acting. Never hand over a top-ten by lift alone; on two million rules that list is close to a list of accidents.

  • How do you choose the minimum support before the run rather than after?
    Set it from the decision, not the algorithm. Decide the smallest number of baskets that would justify acting - say a thousand - and divide by the log size to get the fraction. Then start above that, look at how many frequent itemsets and rules come out, and walk the threshold down in steps, because runtime and output volume grow non-linearly as it falls. A threshold that produces an unreadable rule set is too low regardless of what the data supports.
  • A rule has lift 400 at 0.02% support - do you act on it?
    Only after converting the fraction to a count. Over ten million baskets that is two thousand baskets, which is plenty of evidence and probably a genuine niche pairing worth an adjacency or a bundle. Over a few thousand baskets it is one or two transactions and the lift is an artefact of dividing by a tiny denominator. Same metrics, and the log size decides.
  • How would you validate a shortlisted rule before rolling it out?
    Two stages. First a cheap stability check: re-mine a later, unseen window of transactions and require the rule to clear its thresholds again, which removes most chance findings. Then an experiment - a store or user holdout with the change applied to one arm, a metric agreed in advance such as incremental basket value rather than attach rate, and a decision rule fixed before you look. Without the second stage you are betting that a correlation survives an intervention.
  • What makes one rule redundant given another?
    A longer antecedent that does not buy confidence. If adding B to the left-hand side leaves confidence essentially where `{A} -> {C}` had it, B is along for the ride and the shorter rule dominates - it fires more often for the same reliability. Mining closed itemsets, which keep only itemsets with no equally-supported superset, removes much of this redundancy before rule generation instead of after.

saying these in an interview costs you the question

  • Ships the top ten rules ranked by lift with no volume check
  • Reads a mined rule as evidence that one purchase causes another
  • Judges a rule from its support percentage without the basket count
  • Repeats the beer-and-nappies story as proof the method works
  • Never re-checks whether a rule holds on a later time window
  • Measures the rollout by attach rate rather than incremental value

context