Your green environment passed every smoke test, but seconds after the blue-green cutover sent it 100% of production traffic, latency spiked and errors appeared for several minutes before settling. What causes this, and what does keeping a warm standby actually involve?
answer
- passed the test at zero load
- caches and pools start empty
- runtime has not optimised anything yet
- scaled for the traffic it was getting
- the rollback target is cold too
basics
~20 sA standby that passed smoke tests at near-zero traffic is still cold: empty caches, unestablished connection pools, uncompiled hot paths and an autoscaler sized for no load. Warming means giving it real load — shadow traffic, pre-scaling and pre-opened pools — before the switch.
solid answer
~60 sSmoke tests prove the code works; they say nothing about the environment being conditioned for production load. At the moment of the flip, green has empty application and page caches so nearly every request becomes a downstream call, cold connection pools so a few thousand clients trigger a connect-and-handshake storm against the database, hot paths that have not been optimised yet by the runtime, lazily-loaded code and configuration still being resolved, and an autoscaler whose instance count reflects the near-zero load it has been serving. Downstream systems feel it too — they briefly see a second full fleet's worth of connections. Warming means conditioning the standby before it takes the traffic: pre-scale it to blue's capacity, set minimum pool sizes so connections exist at startup, mirror or replay a slice of real traffic into it, and prefer ramping the weight over an instantaneous flip. And note the symmetry — if you scaled blue down to save money after the switch, your emergency rollback flips onto an equally cold environment.
go deeper
Know that an idle environment is not the same as a ready one: caches are empty and connections have not been opened, so the first burst of real traffic is much more expensive to serve than the steady state.
Enumerate the specific cold resources — application and page caches, connection pools, runtime optimisation of hot paths, lazily-loaded code, and the instance count the autoscaler settled on — and explain why smoke tests at near-zero request rates exercise none of them.
Show how you would condition the standby in practice: pre-scale to match, set pool minimums, mirror a slice of production traffic, ramp the weight rather than flip it, and hold the old environment at full size for a bake period so the rollback is genuinely fast.
Own the economics. The whole cost of blue-green is the second environment, and scaling it down is exactly what destroys the fast-rollback property you paid for. Set the bake-period policy and decide which services justify a permanently warm standby.
## Why the smoke tests told you nothing A smoke test is a correctness check executed at a request rate near zero: a handful of calls to prove routing, configuration and dependencies are wired up. Production is a load condition. Almost everything that hurts at the moment of a cutover is a property of the environment's *state under load*, and an idle environment has none of that state. The green environment is not broken; it is cold. ## What is actually cold **Caches, at every layer.** In-process caches are empty, so requests that normally hit memory now hit a database. A shared cache tier holds nothing for this fleet. Operating-system page cache on the new hosts has not pulled the working set off disk. The result is a multiple — sometimes an order of magnitude — of the steady-state load against every downstream dependency, precisely at the moment you switched. **Connection pools.** Pools typically start empty or at a small minimum and grow on demand. Flipping thousands of concurrent clients onto green means every worker simultaneously discovers it has no connection and opens one: TCP handshake, TLS handshake, authentication, session setup, against a database or dependency that has never seen this fleet before. That connect storm is slow, it is bursty, and it can hit a connection ceiling that produces outright errors rather than latency. **Runtime optimisation.** Managed runtimes execute cold code slowly and optimise the hot paths only after observing them. Lazily-initialised singletons, dependency-injection graphs, class or module loading, template compilation, regex compilation and configuration fetches all happen on the first request that needs them. The first few thousand requests are genuinely more expensive than the millionth. **Capacity.** If green has been running idle, its autoscaler has scaled it to whatever near-zero traffic warranted. The flip delivers full production load to a fleet sized for nothing, and scale-up is not instantaneous — instance provisioning, image pull and startup take minutes, during which the undersized fleet is being crushed. Ironically, cold instances also serve fewer requests per second than warm ones, so you need *more* capacity during the transition than in steady state, not less. **Downstream blast.** For the overlap window, dependencies see two fleets: blue's connections draining and green's being established. Connection limits, rate limits and connection-tracking tables are all sized for one. The visible failure often surfaces in a dependency, which makes it easy to misdiagnose. ## What warming actually means "Warm standby" is not a checkbox; it is a set of deliberate, individually cheap moves: - **Pre-scale.** Bring green to blue's instance count *before* the switch, and hold it there. Never flip onto a fleet sized by idle traffic. - **Set pool minimums.** Configure connection pools with a non-zero minimum so connections are established at startup, spread out, instead of all at once under load. - **Prime what you can.** Run a startup routine that loads the obvious hot dataset, compiles templates, and touches the code paths that matter, so the first user request is not the first execution. - **Send it real traffic.** Mirroring or shadowing a copy of production requests into green — responses discarded — fills caches and pools and optimises hot paths with a realistic distribution, which synthetic load rarely reproduces. Be careful that mirrored requests do not perform writes or trigger side effects such as emails or payments. - **Ramp instead of flip.** Move the weight in steps rather than in one move. This is the highest-value change and the cheapest to implement, because your load balancer already supports weights. It costs you the clean instantaneous cutover and gives you both warming and a small blast radius. ## Diagnosing it after the fact The signature is distinctive: errors and latency spike at the exact second of the switch, are worst on cache-heavy or dependency-heavy endpoints, and *recover on their own* over a few minutes with no intervention. Compare cache hit ratios, pool acquisition wait times and downstream connection counts across the cutover; the shape of a self-healing curve tells you this was conditioning, not a code regression. The mistake to avoid is rolling back on the strength of the spike alone — the rollback flips onto blue, which is fine, but if you conclude the release was bad you will chase a defect that does not exist. ## The symmetric trap The headline benefit of blue-green is that rollback is a traffic switch. That is only true while blue is still running at full size. Teams scale blue to zero minutes after the cutover to stop paying for two environments, and in doing so convert their instant rollback into a cold-start of exactly the kind that just hurt them — this time during an incident, with a hostile audience. Define a bake period during which the old environment stays at full capacity, decommission it only afterwards, and treat the cost of that window as the price of the property you bought.
- What is the cheapest way to warm a standby that does not involve building traffic-mirroring infrastructure?Ramp instead of flip. Move the weight from blue to green in steps — 1%, 10%, 50%, 100% — pausing between them. Real traffic fills the caches and pools, the runtime optimises under genuine load, and the autoscaler scales against real demand, all while the blast radius is still small. It costs you the instantaneous cutover, which is usually the cheaper thing to give up.
- How does the cold-start problem change your rollback plan for blue-green?It removes the assumption that rollback is instant. If blue was scaled to zero or torn down after the cutover, flipping back lands on the same cold environment you just fought, during an incident. Keep the old environment running at full size for a defined bake period, then decommission it — the storage and compute cost of that window is what buys you an actually fast undo.
- Why can a cold standby cause errors in a downstream service rather than in itself?Because the flip briefly doubles what downstreams see. Green opens a full set of new connections while blue's are still draining, so the database hits its connection ceiling, a partner API hits its rate limit, or a connection-tracking table fills. The symptom appears as failures in a dependency and gets misdiagnosed as an unrelated outage.
A restaurant kitchen can pass a health inspection with the lights on and nobody eating. That tells you nothing about whether the ovens are hot and the prep is done when two hundred covers walk in at once.
saying these in an interview costs you the question
- Smoke tests passing means the environment is ready for production load
- Blue-green cutover is instant because it is just a traffic switch
- Cold caches only affect latency, never error rates
- Autoscaling handles it, so pre-scaling is unnecessary
- Rollback is free because the old environment is still there