Monthly totals from a re-spaced daily series show February well below January — is that a real decline?
answer
- the step is a calendar rule
- 28 against 31
- a total scales with the bucket's width
- divide by each bucket's own width
- the end bucket is partial, not low
basics
~20 sA calendar month is 28 to 31 days, so a per-bucket sum scales with the bucket's width and February sits below its neighbours even at a flat rate. Divide each total by its own bucket's width before reading a decline.
solid answer
~50 sA calendar step is a rule, not a fixed length. Months hold 28 to 31 days, quarters 90 to 92, and a year 365 or 366, so a sum or a count over a monthly grid is roughly proportional to how wide each bucket happens to be. A perfectly flat daily rate produces a February total about a tenth below a 31-day month with nothing in the world having changed. Three repairs work: divide each bucket's total by its own width to get a rate, use an equal-width step such as a fixed number of days when comparability matters more than calendar meaning, or compare only like with like — the same calendar month year over year. Watch the end buckets separately: a bucket that is partial because the run happened mid-period is short for a different reason entirely.
go deeper
Know that monthly buckets are not the same width, so a monthly total is partly a measure of how long the month was.
Explain which aggregates are width-sensitive and which are not, and why dividing by each bucket's own width is the repair while dividing by an average is not.
Show that you separate the three reasons a bucket total is low — unequal calendar width, a partial end bucket, and a genuine change — before anyone is allowed to draw a conclusion from it.
The call is what the team publishes as the canonical figure. A rate is comparable, a total is what the business states, and committing to only one of them guarantees an argument later.
## A calendar step is a rule, not a length When a grid's step is expressed in calendar units, the buckets it produces are not the same width: - a month covers 28, 29, 30 or 31 days; - a quarter covers 90, 91 or 92 days; - a year covers 365 or 366; - and a step built from any of these inherits the same unevenness. Nothing about this is a defect — a calendar month is the thing people actually want to report on. The defect is comparing the resulting totals as if the buckets were interchangeable. February is about **10% shorter** than a 31-day month, so a completely flat underlying rate produces a February figure about 10% lower, every single year, in every such series. ## What unequal width does to each kind of aggregate | what the bucket holds | sensitivity to bucket width | |---|---| | a sum of a per-day quantity | roughly proportional to width — the full effect | | a count of records | roughly proportional to width, for a steady arrival rate | | a mean over the records in the bucket | largely insensitive to width; it moves with what the records say, not how many days they span | | a maximum or a minimum | weakly sensitive — a shorter bucket has fewer opportunities to reach an extreme | | the last value in the bucket | insensitive, but stamped at a different position | The common mistake is to assume the mean is exposed the same way the sum is. A per-record mean is not a per-day quantity, so widening the bucket does not scale it; that is precisely why a rate is the repair for a sum. ## Three ways to make buckets comparable 1. **Normalise to a rate.** Divide each bucket's total by that bucket's own width — its own day count, not an average of 30. An average divisor leaves most of the distortion in place and adds a new one. 2. **Use an equal-width step.** A step of a fixed number of days produces buckets of identical width, at the cost that they no longer line up with calendar names anyone recognises, and they drift against month ends year on year. 3. **Compare like with like.** Hold the calendar unit fixed and compare across years — this February against last February — so width is constant by construction. Each costs something. A rate is comparable but is no longer the figure a report or an invoice states. An equal-width step is clean arithmetic but is answering a question nobody asked in those terms. Comparing year over year is honest but throws away the most recent comparison a reader wants. ## The end buckets are short for a different reason A bucket at either end of the span can be **partial**: the run happened mid-period, or the span started mid-period, so that bucket covers less time than the step declares. In a chart, a partial final bucket and a genuinely low one look identical, and the partial one is the more common of the two. Separate them explicitly: - compare the last stamp in the input with the closing edge of the final bucket; if the input stops first, the bucket is partial; - do the same at the leading edge for the first bucket; - mark partial buckets in the output, or exclude them, rather than leaving a reader to treat them as completed periods. Note that this is a *third* reason a bucket total can be low, alongside unequal calendar width and a genuine change. The three need separating before anyone draws a conclusion, and none of them raises an error. ## What to publish The safest habit is to publish the total as the fact and the rate as the comparison, side by side, with the bucket's width available as a column. That way the monthly figure stays the monthly figure, the comparison is done on something comparable, and nobody has to reconstruct the divisor from the stamps to check your work. A last caution on vocabulary: when you say "daily", "monthly" or "quarterly", you are naming a rule in a calendar, not a fixed quantity of elapsed time. The arithmetic above is the monthly case because it is the loudest, but the same caution applies at any size built from calendar units.
- You normalise monthly totals to a per-day rate. What have you made harder to read?The figure most people asked for. A per-day rate is comparable across buckets but is no longer the monthly number a report or an invoice states, and it is noisier on short buckets because the divisor is smaller. The usual resolution is to publish both: the total as the fact, the rate as the comparison.
- How do you distinguish a partial final bucket from a genuinely low one?Compare the last stamp in the input with the closing edge of the final bucket. If the input stops before that edge, the bucket covers less time than the step declares and its total is not comparable. Mark it or exclude it rather than letting a reader treat it as a completed period.
saying these in an interview costs you the question
- Comparing monthly totals as if months were the same length
- Reading a February dip as a change in behaviour
- Treating a partial final bucket as a genuine drop
- Assuming a per-record mean scales with bucket width the way a sum does
- Normalising by a fixed thirty days instead of each bucket's own width