Hourly bike-dock rental counts are your target - what breaks if you treat them as plain continuous regression?
answer
- look at the target's histogram first
- non-negative integers, long right tail
- spread grows with the level
- zeros may mean closed, not quiet
- log link keeps predictions positive
basics
~20 sA count is a non-negative integer, usually right-skewed with many zeros and a spread that grows with its own level. Plain squared-error regression predicts unbounded numbers and assumes constant spread, so it emits negative counts and lets busy hours dominate.
solid answer
~50 sCounts are a distinct target shape, not continuous numbers that happen to be whole. They have a hard floor at zero, no upper bound, a long right tail, and variance that rises with the mean - a station averaging 2 rentals an hour varies by a rental or two, one averaging 60 varies by tens. Fitting an additive squared-error model on the raw counts gives you predictions below zero, treats a 5-rental miss at a quiet station as equal to a 5-rental miss at a peak one, and fits the tail at the expense of the many small values. The usual repairs are framing choices: a multiplicative model with a log link, so predictions stay positive and errors scale with the level, or a rate target with an exposure offset when stations differ in size or opening hours. Separately, ask whether a zero means quiet or closed - those are different targets.
go deeper
Be ready to say why a count is not just any number: it cannot go below zero, it is usually skewed with lots of small values, and a model that predicts negative rentals is telling you the framing is wrong.
Explain the mechanics - squared error assumes constant spread while a count's spread grows with its level, and effects on counts are multiplicative rather than additive. Know that a log link keeps predictions positive and scales errors to the level.
Show you separate structural zeros from genuine low demand, and that you notice when an empty station censors the recorded count so the target is a lower bound on demand rather than demand itself. Choose the aggregation window from the decision cadence.
Own the question of what the business actually wants predicted - observed rentals or latent demand - and be ready to argue for the instrumentation or the modelling investment needed to close that gap, since every downstream decision inherits the difference.
## What makes a count its own kind of target A count is the number of events observed in a window: rentals per station-hour, support tickets per day, defects per batch. It is ordered and evenly spaced - three rentals really is one more than two, unlike a 3-star rating - so it is closer to continuous regression than to classification. But it has structure that an ordinary continuous target does not: - **A hard floor at zero**, and no ceiling in principle. - **Integer values**, though this is usually the least important property. - **Right skew.** Most hours are quiet, a few are enormous. - **Variance tied to the level.** Under the Poisson model, the variance equals the mean, so a station averaging 2 rentals an hour has a spread of roughly 1.4, and one averaging 100 has a spread of roughly 10. Real count data is often *overdispersed* - variance larger than the mean - because events cluster on weather, commuting patterns and events. - **Excess zeros.** Overnight hours, bad weather, and out-of-service docks pile up at zero far beyond what a simple count model expects. ## What an additive squared-error model assumes, and where it collides Plain linear regression fit by least squares assumes an unbounded real output, effects that add, and errors of roughly constant spread across the whole range. Each assumption collides with a count: - **Unbounded output.** The model will happily predict -1.4 rentals at 3am. A negative count is not a small numerical annoyance; it is proof the framing does not know the target's support. - **Constant spread.** Squared error implicitly treats a 5-rental miss the same whether the truth is 3 or 300. In reality the first is a disaster and the second is noise. The fit is therefore dominated by the busy hours, and the many quiet ones are fitted badly - which matters if the decision is about restocking small stations. - **Additive effects.** Effects on counts are usually multiplicative: rain does not remove a fixed 12 rentals, it removes roughly 40% of them, whatever the base level. ## Framing repairs **Log-transform the target.** Fitting on `log(1 + y)` tames the skew and makes effects multiplicative. The catch: exponentiating the prediction back does not recover the mean - the exponential of an average log is closer to a median than a mean, so a naive back-transform is systematically low. If the consumer needs an expected count (fleet sizing sums over stations), that bias matters. **Use a count-shaped model instead of transforming.** A multiplicative model with a log link on the mean - the Poisson family - keeps predictions strictly positive, makes effects multiplicative, and weights errors relative to the level rather than absolutely. Gradient-boosted trees accept a count objective of the same shape. Nothing about this requires the data to be exactly Poisson; it requires the *shape* of the target to be respected. **Model a rate, not a raw count.** If stations have 8 docks or 40, or are open different hours, the raw count confounds demand with capacity. Predict rentals per dock-hour, or keep the count as the target with the exposure entering as an offset. This is a target-definition decision and it usually explains more variance than any feature you could add. ## The zeros deserve their own paragraph Not all zeros mean the same thing: - **Structural zeros.** The station was closed, or the dock was out of service. Demand was not zero; it was unobservable. Including these hours teaches the model that certain conditions produce no demand, which is false. - **Sampled zeros.** The station was open and nobody rented. This is genuine low demand and belongs in the data. If structural zeros are a large share of rows, either exclude them or model the target in two parts: first whether any rental occurred, then how many given at least one. Conflating the two is one of the most common framing errors on count data. There is a related and subtler issue: **censoring by supply**. An empty station records zero rentals even when demand was high, because there were no bikes to rent. The observed count is then a lower bound on demand, not demand itself. If the decision is a rebalancing one, training on observed rentals will systematically under-serve exactly the stations that run dry - the model learns that empty stations have no demand. Recognising that the recorded target is not the quantity you actually want is a strong answer here. ## Choosing the aggregation window Hourly, daily and weekly counts of the same events are different targets with different shapes. Aggregating up reduces zeros, shrinks skew and makes plain regression more defensible; aggregating down gives finer decisions but a spikier, zero-heavy target. Pick the window the decision runs at - if trucks rebalance hourly, hourly is the target, and you take the harder distribution that comes with it. ## Summary of the sanity checks Plot the target's histogram before choosing anything. Check the floor, the tail, the fraction of zeros, and whether the spread grows with the level (group rows by predicted level and compare their variance). Those four observations decide the framing, and they take minutes.
- When would you model rentals per dock-hour instead of the raw rental count?Whenever exposure varies across the rows. A 40-dock station open all day and an 8-dock station open at peak hours produce very different counts for the same underlying demand, so the raw count confounds demand with capacity. Modelling a rate, or keeping the count with exposure entering as an offset, removes that confound and usually explains more variance than any extra feature would.
- Half your station-hours are zero because the station was closed - how does that change the target?Those are structural zeros: demand existed but was unobservable, so they are not evidence of low demand. Either drop them, or split the target in two - whether any rental occurred, and how many given at least one. Leaving them in teaches the model that the conditions surrounding closures suppress demand, which is a property of the operation, not of riders.
- You take logs of the count, fit, then exponentiate back - what is the catch?The back-transform is biased for the mean. Exponentiating an average of logs lands nearer the median than the expected value, so summed predictions come out low - which matters when the consumer adds counts across stations for fleet sizing. Either apply a correction for the residual spread, or fit a model with a log link directly on the mean and avoid the round trip.
saying these in an interview costs you the question
- Treats a skewed integer count exactly like a symmetric continuous target
- Reports negative predicted counts without questioning the framing
- Assumes error spread is constant across quiet and peak hours
- Treats a closed station's zero as genuine zero demand
- Exponentiates a log-scale prediction and calls it the expected count