skip to content

Why encode a pickup hour as a sine-cosine pair instead of the integer 0-23?

level: middleimportance: nice to knowfreq 40%

answer

  1. the ordering breaks at exactly one point
  2. the numeric gap between 23 and 0
  3. put the twenty-four hours on a circle
  4. two coordinates, sine and cosine

basics

~20 s

Hour 23 and hour 0 are one hour apart but 23 units apart as integers. Mapping the hour onto a circle with sin(2pih/24) and cos(2pih/24) makes them neighbours again, which matters to any distance- or magnitude-based model.

solid answer

~50 s

Hour of day wraps around, but the integer encoding does not: 23 and 0 are adjacent in time and maximally far apart numerically. Any model that treats the column as a magnitude inherits that lie — a linear model gets one monotone coefficient across the whole day, and a distance-based model like k-nearest-neighbours thinks 23:00 and midnight are the two most dissimilar hours there are. The fix is to place the hour on a circle: `h_sin = sin(2*pi*h/24)` and `h_cos = cos(2*pi*h/24)`. You need both, because a single sinusoid maps two different hours to the same value — 03:00 and 09:00 both give a sine of about 0.71 — while the pair identifies a unique point. The same trick with period 12 handles month, so December and January end up adjacent. Divide by the period, not by the maximum value.

code

python · 12 lines
python
import math

def cyc(h, period=24):
    a = 2 * math.pi * h / period
    return (math.sin(a), math.cos(a))

def gap(p, q):
    return math.hypot(p[0] - q[0], p[1] - q[1])

print(abs(23 - 0), abs(23 - 22))          # 23 1
print(round(gap(cyc(23), cyc(0)), 3))     # 0.261
print(round(gap(cyc(23), cyc(22)), 3))    # 0.261

go deeper

for a junior

Be ready to say what goes wrong with the plain integer hour: 23 and 0 are one hour apart in reality but 23 units apart in the column, and some models read that gap literally.

for a middle

Explain the mechanics — the angle 2pih/period, why two coordinates are required, and why the divisor is the period rather than the largest value in the column.

for a senior

Show you know when it pays: distance-based and linear models benefit most, trees least, and the smoothness assumption is wrong when the domain has a hard boundary at a specific hour.

for a principal

Frame it as a prior you are choosing — two columns and smoothness versus twenty-four free parameters — and tie that choice to how many rows you actually have and how sharp the effect is.

## The problem: an ordering that wraps Hour of day is cyclical. 23:00 is one hour from midnight, and midnight is one hour from 01:00. An integer encoding 0-23 represents the ordering correctly *within* the day and then breaks at exactly one point: the distance from 23 to 0 is 23, the largest gap in the column, when the true gap is the smallest one there is. Whether that matters depends on how the model reads the column. - A **linear model** multiplies the hour by one coefficient, so it can only say "later in the day means more" or "means less". It cannot represent a bump at 08:00 and another at 18:00, and it certainly cannot make 23 and 0 behave alike. - A **distance-based** model — k-nearest-neighbours, k-means, anything using Euclidean distance — computes gaps directly. With raw hours, a 23:00 ride and a midnight ride are the least similar pair of rides in the dataset on that dimension. - A **tree** is less affected, because it only ever asks "is hour < t", which is invariant to the scale. But it still cannot express a window that crosses midnight in one split: 22:00-02:00 requires two branches, `hour >= 22` and `hour <= 2`, and it must find both. ## The fix: put the clock on a circle Map the hour to an angle and take its two coordinates: ``` angle = 2 * pi * h / 24 h_sin = sin(angle) h_cos = cos(angle) ``` Now every hour is a point on the unit circle, evenly spaced, and Euclidean distance between consecutive hours is identical everywhere — including between 23 and 0. The chord between any two adjacent hours is `2 * sin(pi / 24)`, about 0.261, whether the pair is 22 and 23 or 23 and 0. The wrap-around has disappeared into the geometry. ## Why both coordinates Sine alone is not one-to-one over a cycle. With period 24, `sin(2*pi*3/24)` and `sin(2*pi*9/24)` are both about 0.707 — 03:00 and 09:00 collapse onto the same value, and any model reading only that column treats them as identical. Cosine distinguishes them (about +0.707 versus -0.707). Two coordinates are the minimum needed to name a point on a circle, and dropping either one folds the day in half. ## Getting the period right Divide by the **period** — the number of distinct positions in the cycle — not by the maximum observed value. For hour of day the period is 24, not 23. Dividing by 23 sends hour 23 to an angle of 2*pi, which is the same point as hour 0, so two genuinely different hours become indistinguishable. The same rule gives period 12 for calendar month (so December and January are neighbours), period 7 for day of week, and period 365 or 366 for day of year. ## What the encoding assumes Sine-cosine encoding buys smoothness and periodicity, and those are assumptions. For a linear model, one sine-cosine pair spans exactly one smooth wave per cycle — a single peak and a single trough per day. If demand has a morning peak *and* an evening peak, one pair cannot express both; adding a second pair at half the period gives the extra shape. It also assumes neighbouring hours behave similarly. If your domain has a hard discontinuity — a tariff that switches at 07:00 sharp — smoothness is the wrong prior, and giving each hour its own indicator column will fit the step better, at the cost of 24 columns and no sharing between neighbours. With plenty of rows, the indicator version usually wins on flexibility; with few rows, the two-column circular version is far cheaper. For tree ensembles the gain is usually modest, because a tree can already isolate arbitrary hour intervals given enough splits. It is worth trying — the wrap-around window becomes a single contiguous region rather than two branches — but it is a candidate to validate, not a rule. ## Where it does not apply Only use it on genuinely cyclical quantities. Wind direction in degrees (period 360) and compass bearings qualify. An unordered category such as product type does not — there is no cycle, and imposing one invents adjacency that is not there. And an ordinary bounded numeric column, like age, is not cyclical either: the oldest customer is not adjacent to the youngest. ## Summary The integer hour is right about ordering and wrong about the one place the clock wraps. Two columns, sine and cosine of `2 * pi * h / period`, make every adjacent pair equidistant, keep every hour uniquely identified, and cost two columns instead of twenty-four — as long as you divide by the period and are willing to assume the effect varies smoothly around the cycle.

  • Why do you need both sine and cosine rather than the sine term alone?
    Because a single sinusoid is not one-to-one over a cycle. With period 24, 03:00 and 09:00 both produce a sine of about 0.71, so the model cannot tell morning from mid-morning. Cosine separates them, roughly +0.71 versus -0.71. Two coordinates are the minimum needed to identify a unique point on a circle; keeping one folds the day in half.
  • Does a gradient-boosted tree gain anything from the sine-cosine pair?
    Less than a linear or distance-based model. A tree splits on thresholds, so it can already carve out any hour interval and is indifferent to the scale of the column. What it cannot do cheaply is express a window that crosses midnight — 22:00 to 02:00 needs two separate branches. The circular encoding makes that one contiguous region, so it is worth testing, not assuming.
  • When would you give each hour its own indicator column instead of using sine and cosine?
    When adjacent hours genuinely behave differently — a tariff or shift change at a sharp boundary. Sine-cosine imposes smoothness around the cycle, so it cannot fit a step. Indicator columns give every hour its own parameter and fit the discontinuity, but cost 24 columns, share nothing between neighbours and need far more rows to estimate well.
  • How would you apply the same idea to the calendar month?
    Use period 12: sin(2*pi*m/12) and cos(2*pi*m/12) with m running 1 to 12, or 0 to 11 as long as you are consistent. December and January then sit one step apart on the circle instead of eleven units apart as integers. Day of week uses period 7, and day of year 365 or 366.

Numbering the hours is like cutting a clock face at midnight and laying it out as a straight ruler: the two ends are next to each other on the clock but at opposite ends of the ruler. Sine and cosine glue the ruler back into a circle.

saying these in an interview costs you the question

  • Claiming hour is numeric so the raw integer is fine everywhere
  • Keeping only the sine column and losing half the day
  • Dividing by 23, the maximum, so hour 23 collides with midnight
  • Applying sine-cosine encoding to unordered categories
  • Assuming the encoding always improves a tree ensemble

context