How would you decide whether the hotel pricing model refits nightly or weekly?
answer
- price both sides, do not prefer one
- average age, not worst-case age
- 3.5 days weekly against 0.5 nightly
- six extra runs, not one
- wall-clock and signal accrual can veto
basics
~20 sPrice both sides. Compare the revenue lost to average model age at each cadence against the cost of the extra runs, then check the non-money bounds: whether a run fits the nightly window, and whether a day of bookings adds enough new signal to matter.
solid answer
~40 sCadence is a trade between **staleness cost** and **run cost**, and both are measurable. Estimate staleness from a backtest: fit as of an earlier date, score the following days, and read how quality falls per day of model age. A weekly refit serves at an average age of 3.5 days, a nightly one at 0.5, so nightly buys back three days of decay every week; against that sits six extra runs. Then apply the bounds money does not capture: the refit's wall-clock must fit inside the publish window, a single day must add enough new bookings to change the fit, and every unattended run publishes a model version no human read. The answer is a number, not a preference.
code
pseudocode · 19 linesrun_cost = cost of one refresh run // compute + pipeline time
weekly_revenue = booking revenue per week
// decay measured by a staleness backtest:
// fit as of day 0, score days 1..7, read loss against model age
loss_per_day = 0.004 * weekly_revenue // 0.4% per day of age
mean_age_weekly = 3.5 // refit day 0, serves 7 days
mean_age_nightly = 0.5 // refit day 0, serves 1 day
extra_loss_of_weekly = loss_per_day * (3.5 - 0.5) // = 0.012 * weekly_revenue
extra_cost_of_nightly = 6 * run_cost // 7 runs vs 1
if extra_loss_of_weekly > extra_cost_of_nightly:
choose nightly
else:
choose weekly
// break-even: run_cost = 0.002 * weekly_revenuego deeper
Know that refreshing more often costs more runs and leaves the model less stale, and that the decision compares those two costs rather than following a convention.
Do the arithmetic: average model age at each cadence, a measured decay rate per day, and the number of extra runs the faster cadence actually buys.
Add the bounds money misses — the run's wall-clock against the publish window, how much signal a day accrues, and the standard an unattended publish has to clear.
Own the cadence as a reviewable parameter: who re-measures the decay rate, what market change re-opens the choice, and how much unattended publishing the organisation accepts.
## The two costs you are trading Every cadence choice is the same comparison, and a design round expects you to name both sides: - **Staleness cost** — what the business loses because the live model version is, on average, some days old. It is not the loss on the last day before a refit; it is the average over the interval, because rates are served every day of it. - **Refresh cost** — compute for the run, pipeline wall-clock, contention with other jobs, and the operational load of publishing model versions more often. Staleness is the side people wave at. It is measurable: refit the model as of an earlier cut, score each of the following days with it, and plot quality against days since the fit. That gives a decay rate — say **0.4% of booking revenue lost per day of model age** — which is the only number that makes the comparison real. ## A worked comparison, nightly against weekly With a refit at day 0, a weekly cadence serves the model at ages 0 through 7, so **mean age is 3.5 days**. A nightly cadence serves ages 0 through 1, so **mean age is 0.5 days**. Nightly therefore removes **3.0 days** of average age. At 0.4% of booking revenue per day, that is **1.2% of weekly booking revenue** recovered. The extra spend is **six additional runs per week** — seven against one. So: - nightly wins when `6 x run_cost < 0.012 x weekly_revenue`; - **break-even run cost is 0.2% of weekly booking revenue**; - a run costing 0.3% of weekly revenue makes nightly cost 1.8% to recover 1.2% — weekly wins. The arithmetic is worth doing out loud in the interview, because the common error is comparing one run's cost against the full 1.2%, rather than the six extra runs a week that nightly actually buys. ## What bounds the cadence besides money Even when the arithmetic favours nightly, three hard bounds can veto it: 1. **Wall-clock.** If the refresh must publish by 06:00 and the refit takes nine hours, nightly is not available at that scope. Either the scope shrinks (a narrower training window, fewer stages rerun) or the cadence lengthens. 2. **Signal accrual.** One day of bookings may move the fit hardly at all for a small portfolio, in which case nightly runs spend money to produce near-identical model versions. The question to ask is how many new labelled examples a day contributes relative to the training window's size. 3. **Unattended publishing.** Every scheduled refresh publishes a model version that prices rooms without anyone reading it. Raising the cadence raises how often that happens, which raises the standard the automatic gate in front of promotion has to meet. ## How the decay curve changes the answer The comparison above assumes decay is roughly linear in model age. Two common shapes change the conclusion: | decay shape | what it means | cadence implication | |---|---|---| | linear | quality falls steadily with age | the arithmetic above applies directly | | flat then cliff | stable for days, then a break at a season or event boundary | a clock is the wrong instrument — use an event trigger | | fast then flat | most decay in the first day or two | a shorter interval pays much more than the linear model suggests | A flat-then-cliff curve is the tell that cadence is not your real problem: no interval both avoids the cliff and avoids paying for runs that change nothing. ## Reading the answer off the numbers A defensible answer in a design round sounds like this: *"I would measure the decay rate from a backtest, compute the average-age difference between the candidate cadences, price the extra runs, and pick the cheaper side — then check the refit fits the publish window and that a day adds enough new bookings to be worth fitting. On these numbers weekly wins, and I would keep a drift trigger on top so a mid-week change does not wait out the interval."* The last clause matters: choosing a cadence never means choosing *only* a cadence. The cadence is the ceiling on staleness; the evidence triggers are what stop the ceiling from being the whole story. ## What changes the answer later Re-open the decision when the decay rate changes (a more volatile market, a new competitive segment), when run cost changes materially (a cheaper refresh scope, a smaller snapshot), or when the portfolio grows enough that a day of bookings is now a meaningful share of the training window. Cadence is a parameter with an owner and a review date, not a constant fixed at launch.
- How do you measure the decay rate the comparison depends on?Run a staleness backtest. Fit the model as of an earlier training cut, then score each of the following days with that frozen model version and record the quality metric against days since the fit. The slope is the decay rate. It must be measured on days after the cut, never on held-out rows from inside the training window, or it measures generalisation rather than ageing.
- Does moving from weekly to nightly shorten the training window?No — they are independent choices. The cadence sets how often a fit happens; the window sets how much history each fit reads. A nightly refresh can still train on two years of bookings. Confusing the two produces the false worry that frequent refits starve the model of data.
saying these in an interview costs you the question
- Compares one run's cost against a whole week of recovered loss
- Uses worst-case model age instead of average age over the interval
- Thinks a shorter cadence shortens the training window
- Picks a cadence by habit without measuring a decay rate
- Ignores whether the refit's wall-clock fits the publish window