skip to content

Your hosted scheduled collection run is failing while your own runs of the same collection pass. How do you triage that?

level: seniorimportance: should knowfreq 45%

answer

  1. Only the requests are shared; everything else swapped
  2. Copy, environment, identity, origin, driver
  3. Run the hosted copy, not the repository copy
  4. Rotation missed a store you do not hold
  5. Failing only from outside is a real finding

basics

~20 s

Compare what your own run never exercises: the copy the schedule holds, the environment stored beside it, the credential in that account, and the public path in from outside your network. Only then suspect the service.

solid answer

~50 s

Do not start with the service. Start with everything the hosted arrangement substitutes. **The copy** — the hosted collection may be an older document asserting different things. **The environment** — the values it uses are stored in that account, and may point somewhere else. **The identity** — the credential lives there too, so a rotation on your side does not reach it. **The origin** — the run enters from the public internet, so it passes through name resolution, your edge and whatever filtering sits in front of the service, none of which your in-network run touches. **The driver** — a run may have started late, been retried, or not started at all. Work inward: take the hosted copy, run it yourself against the same environment, then run it from outside your network. Distinguishing a genuine intermittent defect from an unreliable case is a separate testing discipline.

go deeper

for a junior

Remember that the two runs share only the requests. The copy, the environment, the credential and where the request comes from are all different, so a disagreement does not automatically mean the service is fine.

for a middle

Be able to list the substitutions — copy, environment, identity, origin, driver — and explain why reproducing with the hosted copy rather than your repository copy is the first real step.

for a senior

Show the elimination order and what each outcome implies about process: an incomplete rotation runbook, an unbounded drift window, or a genuine defect on the public path that your pipeline could never have found.

for a principal

Own the systemic reading: how long drift survived, whether rotation covers stores you do not hold, and how much of your feedback loop quietly depends on another party's clock being up.

## Why the naive conclusion is wrong The tempting inference is: the hosted run is red, my run is green, therefore the hosted run is unreliable. That inference is unsafe in both directions. The hosted run may be seeing something real that your run structurally cannot see, or it may be failing for a reason that has nothing to do with your service at all. The reason you cannot tell yet is that **the two runs differ in more places than the collection**. So triage here is not about the requests. It is about enumerating the substitutions the arrangement made and eliminating them in order. ## The five substitutions, and what each one breaks | Substitution | What differs | The failure it produces | |---|---|---| | **The copy** | the hosted account holds its own collection | asserts against requests or expectations you have since changed | | **The environment** | values stored beside the hosted copy | points at a different target, or carries a stale value | | **The identity** | credentials kept in that account | authentication failures after a rotation you performed elsewhere | | **The origin** | the request enters from the public side | name resolution, edge rules, filtering and throttling your in-network run never meets | | **The driver** | their clock and their machines | a run that started late, was retried, or did not start | A candidate who can produce that table has answered the question, because triage is then simply walking it. ## The order to walk it in 1. **Read the failure itself first.** Which request, which assertion, and is it the same request failing every time or a different one? A single request failing consistently points at that path; scattered failures point at identity, origin or the driver. 2. **Take the hosted copy and run it yourself.** Not your repository copy — the copy the schedule actually holds. If it fails for you too, this is copy drift or a real defect, and the origin is innocent. 3. **Check the environment it used.** Confirm the target it names is the target you think, and that the values it carries are current. A hosted check quietly pointed at a different environment produces failures that make no sense against the one you are looking at. 4. **Check the identity.** If the failure is an authentication or authorisation failure, the credential stored in that account is the first suspect, particularly if anything was rotated recently on your side. This is the single most common cause of a schedule that goes red without any deploy. 5. **Reproduce from outside your network.** If the hosted copy passes from inside and fails from outside, the difference is the path in. That is precisely the vantage point you bought by putting the check there, and the finding is a real one about how your system answers strangers. 6. **Only now suspect the service.** And when you do, remember that a run entering from outside may reach a different instance, a different edge, or a different set of rules than a run started inside. ## What to do with the answer - **If it is the copy:** republish the hosted copy from the authoritative repository copy, and note how long the drift had gone unnoticed — that number tells you how much to trust every other hosted check. - **If it is the identity:** the rotation process is incomplete, because it does not include a store you do not hold. That is a process defect, not a one-off fix. - **If it is the origin:** you have found something your own pipeline could not have found. Treat it as a genuine defect in the path, and be precise about which hop it lives in. - **If it is the driver:** you have learned that your feedback loop has a dependency on somebody else's availability, and that its silence needs its own detection. - **If it is the service:** the hosted check did its job, and the only remaining question is why the same failure was invisible from inside. ## Boundaries worth naming out loud Saying where a question stops is itself a senior signal. Two neighbouring subjects are not this one: - **Deciding whether an intermittent failure is a real defect or an unreliable case** is a testing discipline with its own criteria, and "it is flaky" is not a triage step — it is a conclusion you may only reach after the five substitutions have been eliminated. - **Whether this check is a good indicator of service health**, and what it misses compared with real traffic, is a service-level measurement question owned elsewhere. What belongs here is the discipline of remembering that the hosted run and your run share only the requests, and everything around them was swapped.

  • The hosted run fails on authentication but your local run passes. What do you check first?
    The credential stored in that account. A rotation performed in your own secret store does not reach a copy held on the vendor's side, so the hosted run keeps presenting the old one. Treat it as a process defect rather than a one-off: any rotation runbook has to include every store you do not hold, or the same failure returns on the next rotation.
  • The hosted copy passes when you run it from inside your network but fails on their clock. What have you learned?
    That the difference is the path in, not the collection or the service logic. The hosted run enters through public name resolution, your edge, and whatever filtering or throttling sits in front of the service. That is exactly the vantage point the arrangement exists to provide, so the finding is real and the next step is identifying which hop rejects it.
  • When is it legitimate to conclude the hosted run is simply unreliable?
    Only after eliminating the copy, the environment, the identity, the origin and the driver — and even then, unreliability is a conclusion with its own discipline behind it, not a triage shortcut. Reaching for it first is how a genuine outside-only defect stays invisible for weeks while everyone assures each other the scheduled check is noisy.

saying these in an interview costs you the question

  • Declares the scheduled run unreliable before eliminating anything
  • Reproduces with the repository copy rather than the hosted copy
  • Forgets the credential lives in an account they do not hold
  • Ignores that the request enters from the public side
  • Assumes both runs target the same environment
  • Never considers that the run may not have started at all