Your 06:00 taxi-demand forecast lacks the last three hours of counts — how do you frame the task?
answer
- when it runs is not when data lands
- the information cutoff sets the origin
- shortest usable lag is the latency
- train rows must simulate the cutoff
- worst-case delay, or cutoff age as feature
basics
~20 sDefine features against the last landed observation, not the clock. With a three-hour delay the shortest usable lag is four hours, and training rows must be rebuilt at that cutoff so the model never relies on values serving lacks.
solid answer
~50 sThe forecast runs at 06:00 but the information cutoff is 03:00, and those are two different things. If the training table was built from the complete history, it holds a lag-1 feature containing the 05:00 count, a value the serving path will never have. That is train-serve skew: the offline score is inflated and production quietly underperforms. The fix is to rebuild each training row as it would have looked at its own origin, dropping the lags that fall inside the latency window so the closest usable lag is four hours, and ending every rolling window at the cutoff. It also changes the horizon arithmetic: forecasting through the end of the day from a 03:00 cutoff means horizons of four to twenty-one hours. If the delay varies, I design for the worst case or feed the age of the newest observation in as a feature.
go deeper
Know that a forecast can only use data that has actually arrived, and that the newest hour is often not there yet. Being able to say the shortest lag depends on the delay is already a good answer here.
Explain the mechanism of train-serve skew: the training table was built from complete history, so the model learns to rely on a feature the serving path cannot supply. Say how you would rebuild the rows to match the cutoff.
Demonstrate that you would ask what time the data lands before designing features, redo the horizon arithmetic from the cutoff, handle variable delay explicitly, and monitor serving-time feature availability rather than trusting it.
Own the contract with the upstream data owners: what latency is guaranteed, what happens when it is breached, and whether the forecast schedule should move rather than the model absorbing the gap.
## Two origins, not one Every scheduled forecast has two timestamps that beginners collapse into one. - The **run time** is when the job executes: 06:00. - The **information cutoff** — the effective forecast origin — is the timestamp of the newest observation actually available to it: 03:00, if counts land three hours late. Everything the model may use is defined by the second, not the first. Data pipelines, upstream batch jobs, late-arriving records and settlement delays all push the two apart, and the gap is rarely zero in production. ## The failure this creates Training tables are usually built from a complete historical dataset, where every timestamp is present. Build a lag-1 feature there and it holds the 05:00 count for the 06:00 origin. Serve that model and lag-1 is missing, stale, or filled with whatever the imputation default happens to be. The model has learned to lean hard on its most informative feature and is then denied it. Symptoms: - Offline error looks strong; production error is materially worse, and the gap is largest at short horizons where recent counts matter most. - The gap does not close with tuning, because it is not a modelling problem. - A feature-importance view shows the model relying most on the feature that is least reliable at serving time. This is the specific train-serve skew that data latency produces, and it is one of the most common ways a well-built forecasting model disappoints after launch. ## Framing the task correctly **1. Define features relative to the cutoff.** With a three-hour delay, no feature may reference the interval `(cutoff, run time]`. The shortest lag becomes `lag_4` relative to the run hour; rolling windows must end at the cutoff, not at the run time. Everything a longer lag can offer — same hour yesterday, same hour last week — is unaffected and becomes proportionally more valuable. **2. Rebuild training rows as of their own origin.** The safest construction simulates the pipeline: for each historical origin, assemble only what would have been visible then. Where the source system records when a row *arrived* as well as when it *happened*, use the arrival timestamp to reconstruct the visible history exactly. Where it does not, applying a fixed offset is a reasonable approximation. **3. Redo the horizon arithmetic.** "Forecast the rest of today" sounds like eighteen hourly steps. Measured from a 03:00 cutoff it is horizons four through twenty-one. That matters because forecast difficulty is a function of distance from the information cutoff, and because the baseline you compare against must be given the same handicap — a seasonal-naive rule computed from the full history is not a fair reference for a model that cannot see the last three hours. **4. Handle variable latency deliberately.** If the delay is usually three hours but occasionally six, designing for the average guarantees intermittent failure. Two workable options: - **Build for the worst case.** Use only lags of six hours or more. Simple, robust, and costs accuracy on the majority of runs. - **Make the cutoff a feature.** Supply the age of the newest observation as an input and generate training rows across a range of simulated cutoffs. The model then learns to lean on recent lags when they exist and on seasonal structure when they do not. This costs more engineering and requires that the serving path reports the true age. **5. Resist imputing the gap with your own predictions.** Filling 04:00 and 05:00 with model estimates and then forecasting forward is recursion wearing a disguise: you inherit the imputation error, and the model was trained on true values, so it treats the estimates as if they were observed. If you do it anyway, train on the same imputed inputs so training and serving at least agree. ## Detecting it when nobody warns you The tell is a comparison between the feature values seen in training and those seen in scoring: a lag that is always populated offline and frequently null or stale online. Logging the distribution of each feature at serving time and diffing it against training is the standard defence, and the age of the newest observation is worth logging as a first-class field. A performance gap that is concentrated at short horizons and absent at long ones points at exactly this cause. ## Why interviewers like this question It separates candidates who have only built forecasting models offline from those who have shipped one. The offline builder answers with features and models; the shipper asks what time the data lands, and designs the task around the answer.
- How would you detect this problem if nobody told you about the delay?Log feature values at scoring time and diff their distributions against training. A lag that is always populated offline but frequently null or stale online is the signature. A second tell is a production shortfall concentrated at short horizons and absent at long ones, since only the near horizons depended on the missing recent counts.
- The delay is usually three hours but sometimes six — how do you handle that?Do not design for the mean. Either build to the worst case and use only lags of six hours or more, or make the age of the newest observation an explicit feature and generate training rows across a range of simulated cutoffs, so the model learns to fall back on seasonal structure when recent values are absent.
- Is it acceptable to backfill the missing hours with the model's own estimates?It is recursion in disguise: you inherit the imputation's error and the model was trained on real observations, so it trusts the estimates as if they were measured. If you do it, at minimum train on the same imputed inputs so training and serving agree. Usually dropping those lags is simpler and safer.
saying these in an interview costs you the question
- Builds a lag-1 feature unavailable at scoring time
- Assumes the newest observation is always available at run time
- Fills late-arriving hours with zeros only in serving
- Designs for the average delay instead of the worst case
- Confuses when the job runs with when data lands
- Compares against a baseline given the full history