skip to content

Why can a basket rule with 40% support and 67% confidence still be worthless?

level: middleimportance: must knowfreq 66%

answer

  1. look at the right-hand item alone
  2. a popular consequent inflates confidence
  3. confidence never sees the base rate
  4. ratio near one, excess near zero
  5. strip carrier bags before mining

basics

~20 s

Because both numbers can come entirely from the right-hand item being popular. If milk sits in 70% of baskets, any rule predicting milk reaches high confidence by default. Lift near 1 and leverage near 0 expose it.

solid answer

~40 s

Support and confidence are both blind to how common the consequent is on its own. Take weekly supermarket baskets where bread is in 60% and milk in 70%: `{bread} -> {milk}` shows support 0.40 and confidence 0.67, which reads as commanding, but milk is in 70% of *all* baskets, so the rule does slightly worse than guessing milk with no information at all. Lift is 0.67/0.70 = 0.95 and leverage is 0.40 - 0.60*0.70 = -0.02. The extreme form of the same trap is `{plastic carrier bag} -> {anything}`: the bag is in most baskets, so it appears on the right of thousands of high-confidence rules that carry no signal. The fixes are to rank by lift or leverage rather than confidence, and to strip near-ubiquitous items before mining at all.

code

python · 19 lines
python
baskets = [
    {"bread", "milk"}, {"bread", "milk", "eggs"}, {"bread", "jam"},
    {"milk", "eggs"}, {"bread", "milk"}, {"milk"}, {"bread", "milk", "jam"},
    {"eggs"}, {"bread"}, {"milk", "jam"},
]
n = len(baskets)


def support(items):
    return sum(items <= b for b in baskets) / n


left, right = {"bread"}, {"milk"}
sup = support(left | right)
conf = sup / support(left)
lift = conf / support(right)
lev = sup - support(left) * support(right)
print(round(sup, 2), round(conf, 2), round(lift, 2), round(lev, 3))
# 0.4 0.67 0.95 -0.02  ->  frequent, reliable, and not worth acting on

go deeper

for a junior

Remember that a rule needs a third number beyond support and confidence, and that lift near 1 means the pairing is unremarkable however impressive the first two look.

for a middle

Be able to compute lift and leverage from the same counts and to say precisely what confidence leaves out. Explaining why a popular consequent inflates confidence is the core of the answer at this level.

for a senior

Demonstrate the workflow, not just the metric: blacklist non-merchandise lines, apply a maximum-support cut, then rank by lift with a floor on absolute joint baskets before anything reaches a stakeholder.

for a principal

Frame this as a metric-governance problem. Decide once which score the organisation ranks rules by and what the volume floor is, or every analyst will bring a different top-ten list and each one will be defensible.

## The blind spot in confidence `confidence(A -> B) = support(A and B) / support(A)`. Look at what is missing: `support(B)` appears nowhere. Confidence measures how reliably B follows A without ever asking how easy it is to find B in the first place. When the consequent is common, confidence is high for almost any antecedent you attach to it, and a rule list ranked by confidence turns into a list of the store's most popular items with random things bolted onto the left. ## The bread-and-milk case Weekly supermarket checkout data. Bread is in 60% of baskets, milk in 70%, and both together in 40%. ``` support({bread, milk}) = 0.40 confidence(bread -> milk)= 0.40 / 0.60 = 0.67 lift = 0.67 / 0.70 = 0.95 leverage = 0.40 - 0.60 * 0.70 = -0.02 ``` Support of 0.40 is enormous - the rule covers four baskets in ten. Confidence of 0.67 says two thirds of bread buyers take milk. Both look like a headline. But milk is in 70% of baskets full stop, so *knowing about the bread makes you slightly worse at predicting milk*. Lift 0.95 and negative leverage say the same thing twice. The pairing is frequent because both items are staples, not because the items have anything to do with each other. ## Leverage, and why it is not just lift again `leverage(A -> B) = support(A and B) - support(A) * support(B)` Lift is a **ratio**, leverage a **difference**, and the difference is measured in baskets. That distinction decides which rules survive contact with a business. - A staple pairing can have huge support and leverage near zero, as above: enormous volume, no excess. - A rare pairing can have spectacular lift and leverage near zero too, because the excess is a ratio over almost nothing. Leverage 0.0001 on ten million baskets is a thousand baskets; on five thousand baskets it is half of one. So lift ranks rules by strength of association and leverage ranks them by how much co-occurrence actually exists beyond chance. Reading both keeps you from shipping either a trivial staple rule or a lottery ticket. **Conviction** is a third correction sometimes seen: `(1 - support(B)) / (1 - confidence(A -> B))`. It also equals 1 under independence and grows without bound as the rule approaches being exception-free, which makes it sensitive to the rare counter-examples that lift ignores. ## Ubiquitous items Every real transaction log has items that are in most baskets and mean nothing: the plastic carrier bag, the loyalty-card line, the deposit on a returnable bottle, bananas in some chains. A bag at 60% frequency lands on the right of a large fraction of all high-confidence rules, because 60% is the floor for any antecedent. It also lands on the *left* of rules with near-total confidence for other staples. The practical handling: 1. Blacklist non-merchandise lines (bags, fees, deposits) before the mining run, not after. 2. Apply a maximum-support cut alongside the minimum, so items present in almost every basket are excluded from candidate generation. 3. Mine the remaining catalogue, then rank by lift with a leverage floor. That sequence typically removes more junk than any amount of post-hoc rule filtering, and it makes the run faster, because ubiquitous items are exactly the ones that generate the most frequent itemsets. ## Why frequent does not mean informative The underlying point generalises past shopping. Support and confidence are unconditional and conditional *frequencies*; neither is a comparison against a baseline. Any metric that never looks at the base rate of the thing being predicted will reward predicting whatever is common. That is the same shape as an accuracy score on an imbalanced problem: high, correct, and uninformative. Lift and leverage restore the baseline, one as a ratio, one as an absolute excess. ## Interpreting lift 0.95 Slightly below 1 is not 'almost strong' - it is on the wrong side of neutral. On a very large basket count it hints at mild substitution or at two items competing for the same trip, which can be worth a look as a cannibalisation question. On modest counts it sits inside the noise. Either way it is not a cross-sell, and presenting it as one because the support and confidence look impressive is the mistake this whole question is about.

  • When do lift and leverage disagree, and which do you trust?
    They disagree at the extremes of frequency. A staple pairing has volume but almost no excess, so lift sits near 1 while support is huge; a two-basket coincidence has enormous lift and leverage close to zero. Trust neither alone: use lift to judge whether an association exists and leverage, or the raw joint basket count, to judge whether enough of it exists to be worth a decision.
  • How do you handle items that appear in most baskets?
    Exclude them before the run with a maximum-support cut and an explicit blacklist for non-merchandise lines such as bags, deposits and fees. Left in, they dominate frequent itemset generation, slow the mining down and populate the consequent of thousands of meaningless high-confidence rules. If a ubiquitous item genuinely matters to the business, analyse it on its own rather than letting it flood the rule set.
  • Is a rule with lift 0.95 ever useful?
    Only as a negative signal, and only with volume behind it. Across millions of baskets a consistent sub-1 lift between two items in the same category can flag substitution - shoppers pick one or the other - which matters for range and promotion planning. On thin data it is noise. What it is never is a cross-sell recommendation, however good the support and confidence look.

Predicting that a shopper takes a carrier bag is like predicting rain in a rainforest: you will be right most days, and the forecast still tells nobody anything they did not know.

saying these in an interview costs you the question

  • Ranks mined rules by confidence and ships the top of the list
  • Reads high support as evidence of a strong association
  • Says lift 0.95 is close to 1 so the rule is nearly strong
  • Never checks how frequent the consequent is by itself
  • Leaves carrier bags and loyalty lines in the transaction data

context