The object being scored
Association rule mining runs over a transaction log: each row is a basket, an unordered set of items bought in one trip. A rule is written {A} -> {B} and read as baskets containing A tend to also contain B. The left-hand side is the antecedent, the right-hand side the consequent; either side may hold several items. Nothing about a rule is causal and nothing is ordered in time - it is a statement about items landing in the same set.
Three numbers score every rule, and all three come from counting baskets in that one log.
Support
support(X) = baskets containing every item of X / total baskets
For the rule {A} -> {B}, the rule's support is support(A and B): the share of baskets containing the entire itemset. Support answers how often does this situation even arise. A rule at 0.00001 support can be perfectly reliable and still not worth a line of code, because it fires a handful of times a year.
Support is symmetric: {A} -> {B} and {B} -> {A} have identical support, because set membership has no direction. It is also the quantity every mining algorithm prunes on, which is why the minimum-support threshold is the single knob that decides how long a run takes.
Confidence
confidence(A -> B) = support(A and B) / support(A)
Of the baskets that contain A, what fraction also contain B. This is the rule's reliability given that the antecedent fired, and it is what a trigger-based cross-sell cares about: the shopper has A in the cart, how often is B there too. Dividing by support(A) rather than support(B) is what makes confidence directional. A rule from a rare item to a common one scores high confidence; reverse it and the number collapses.
Lift
lift(A -> B) = confidence(A -> B) / support(B)
B appears in support(B) of all baskets. Among the A baskets it appears in confidence of them. Lift is the ratio of those two rates: how much more often B turns up when A is present than it turns up in the population at large.
- lift = 1 - the A baskets look exactly like every other basket with respect to B; the rule adds nothing.
- lift > 1 - positive association; the items co-occur more than their individual frequencies alone would produce.
- lift < 1 - negative association; seeing A goes with not seeing B, which is what substitutes look like.
Lift, like support, is symmetric: swapping the sides leaves it unchanged.
A worked example
10,000 baskets. Coffee in 1,200. Filters in 400. Both in 300.
support = 300 / 10000 = 0.03
confidence = 300 / 1200 = 0.25
support(filters)= 400 / 10000 = 0.04
lift = 0.25 / 0.04 = 6.25
Read it out loud: the rule fires on 3% of trips; when coffee is in the basket, filters are there a quarter of the time; and a quarter is more than six times the 4% rate at which filters appear generally. Now reverse it: confidence({filters} -> {coffee}) = 300/400 = 0.75, three times the other direction, while support and lift do not move at all. That asymmetry is the most useful single fact about confidence.
Why you need all three
Each metric answers a different question and each is useless alone.
- Support alone finds the obvious: the most frequent itemsets in a supermarket are the things everyone buys, and a rule between two staples is common without being informative.
- Confidence alone is inflated by a popular consequent. If an item is in most baskets, almost any antecedent reaches high confidence for it.
- Lift alone rewards rarity: two obscure items that happened to land in the same three baskets can score a spectacular ratio off almost no data.
So mining is usually run as filter by minimum support, then rank by lift, then sanity-check the absolute basket count behind each rule.
What counts as a transaction
Before any of these numbers mean anything, someone decided what a row is: one checkout, one online session, one customer-month. Widen the window and almost every pair of items co-occurs somewhere, so support and confidence inflate while the rule stops describing a single shopping decision. The definition of a transaction moves every metric on this page and it is a modelling choice, not a data-engineering detail.
Common mistakes
Reporting support as a raw count rather than a fraction makes thresholds meaningless across datasets of different sizes. Treating confidence as symmetric produces recommendations pointed the wrong way. And treating lift near 1 as nearly strong inverts the scale: 1 is the neutral point, not the floor.