In a weekly acquisition cohort table, why is averaging down a column misleading?
answer
- two different time axes in one grid
- the grid is a triangle, not a rectangle
- recent cohorts cannot reach the far columns
- diagonals carry the calendar shocks
- unweighted mean is not the pooled rate
basics
~20 sA cohort table is triangular: recent cohorts have not lived long enough to appear in later columns. Averaging a column therefore describes only the older cohorts that reached that age, blending different acquisition weeks and channels.
solid answer
~50 sLay the table out the standard way: one row per acquisition week, one column per week since acquisition, each cell a retention rate. Reading **across a row** gives one cohort's lifecycle curve. Reading **down a column** compares cohorts at equal age, which is the fair comparison — but only among cohorts old enough to have a value there, so column 8 contains only cohorts acquired at least eight weeks ago. Averaging that column silently describes the older half of the business, and if acquisition mix shifted, those cohorts are not representative. The **diagonals** are calendar weeks: an outage, a holiday or a tracking break shows up as a diagonal streak, not a row or a column. A second trap is weighting: an unweighted mean of cohort rates is not the pooled rate when cohort sizes differ.
go deeper
Be ready to say what a row, a column and a cell mean in an acquisition cohort table, and to point out that the recent rows are short because those weeks have not happened yet.
Explain why the triangular shape makes a column average describe only mature cohorts, and why an unweighted mean of cell rates differs from the pooled rate when cohort sizes vary.
Show you check diagonals for calendar shocks, blank immature cells rather than filling them, and split by acquisition channel before letting anyone read a column movement as a behaviour change.
Own the reporting standard: which comparison the company treats as authoritative, how maturity cut-offs are enforced in tooling, and how you stop a mix shift from being reported as a retention win.
## The layout The standard acquisition cohort table has: - **one row per acquisition period** — say each week's signups; - **one column per period since acquisition** — week 0, week 1, week 2, and so on; - **each cell** holding the share of that row's cohort that was active in that column's week. The row labels are calendar; the column labels are age. That single fact — two different time axes in one grid — is what makes the table powerful and what makes it easy to misread. ## Three ways to read it **Across a row** is one cohort ageing. This is the retention curve for people acquired in that week, and it is the only reading where nothing is being mixed. **Down a column** compares different cohorts at the **same age**. This is the right way to ask "are newer cohorts better than older ones?", because it holds age fixed. It is the comparison teams actually want when they ask whether a product change improved retention. **Along a diagonal** is a fixed **calendar week** across cohorts of increasing age. Anything that hit everyone at once — an outage, a holiday, a pricing change, a broken analytics SDK release — appears as a diagonal streak. If you only ever read rows and columns, calendar shocks look like a mysterious dip that starts at different ages in different cohorts. ## Why the column average misleads **The table is triangular.** A cohort acquired three weeks ago has values for weeks 0-3 and nothing beyond, because those weeks have not happened yet. So column 8 is populated only by cohorts at least eight weeks old. The average of that column is therefore the average over a **non-random, systematically older subset** of cohorts — and "older cohorts" usually means a different acquisition mix, different pricing, and a different product. This produces a specific illusion: plotting the column averages as if they were a retention curve makes the curve look better or worse than any real cohort's, because each point is computed on a different set of cohorts. The later the column, the more selective the set. A team that improved retention recently will see the improvement in early columns only, and the later columns will keep reporting the old regime for months. **Unweighted versus pooled.** A plain mean of cell values treats a 200-user cohort and a 50,000-user cohort equally. The pooled rate — total retained users in that column divided by total cohort members contributing to it — is a different number, and the gap is largest exactly when acquisition volume is volatile. Say which one you are showing. **Acquisition mix.** Rows differ by more than date. A week with a paid-install campaign brings in users who behave nothing like organic signups, and its whole row sits lower. Averaging the column blends channels in whatever proportion they happened to be acquired in, so a movement in the average can be pure mix shift with no behavioural change at all. ## What to do instead - **Enforce equal maturity.** Compare only cohorts that have all reached the age in question, and say the cut-off out loud: "week-8 retention across the twelve cohorts acquired eight or more weeks ago". - **Grey out or blank the immature cells** rather than filling them with a partial value. Half-elapsed weeks reported as if complete are the single most common cohort-table defect. - **Show the cohort sizes** next to the rates so a reader can see which rows carry weight and which are small-sample noise. - **Segment before averaging** when the acquisition mix is unstable — separate tables per channel beat one blended table with a caveat nobody reads. - **Scan the diagonals** as a routine check before attributing anything to a cohort or an age. ## The summary sentence A cohort table has three readings — lifecycle along rows, cohort quality down columns, calendar along diagonals — and the column reading is valid only among cohorts of equal maturity, weighted the way you say you weighted them.
- A retention dip appears at week 3 in one cohort, week 5 in another and week 6 in a third. What should you check?The diagonal. If those cells all fall in the same calendar week, the cause is an event that hit every cohort at once — an outage, a holiday, a release that broke event tracking — rather than something about product age. Row and column readings will never surface it, because the dip sits at a different age in every cohort. Overlaying calendar dates on the grid makes it obvious.
- How do you show the current week's cohort in a table without it looking terrible?Blank the cells whose period has not fully elapsed rather than reporting a partial value, and mark the row as in progress. A cohort three days into week 0 has genuinely not had the chance to be active for the full period, so a partial cell is not a low number, it is an incomplete one. Reporting it beside complete cells invites a false alarm about a collapse in new-user quality.
It is like averaging exam scores in the column marked 'fourth year' at a university: only students who have been there four years appear, so the average describes the survivors of an older intake, not the school today.
saying these in an interview costs you the question
- Plots column averages and calls it the retention curve
- Reports partly elapsed periods as finished cells
- Ignores that acquisition mix differs between rows
- Averages cell rates without regard to cohort size
- Never looks along diagonals for calendar events