skip to content

Maintenance & Version Lifecycle

Managed services patch, restart and upgrade on the provider's schedule, and announce end of support with a date. Asked because a deprecation notice becomes an unplanned engineering project.

on this pageshow

questions

5

Your managed database has a weekly maintenance window - what may the provider do inside it, and what does your application see?

level: juniorimportance: must knowfreq 68%

answer

  1. when, not whether
  2. a slot you nominate
  3. routine work, deferrable work
  4. expect dropped connections
  5. standby promoted, endpoint unchanged

basics

~20 s

A maintenance window is a recurring slot you nominate in which the provider may patch and restart your managed instance. The application normally sees dropped connections and a short interruption, or a failover to the standby replica, not a seamless change.

solid answer

~40 s

A managed service still runs on software and machines someone has to patch: the engine, the operating system under it, and the firmware below that. The **maintenance window** is the recurring slot - a day, a start time in a fixed time zone and a duration - in which the provider is allowed to do that routine work on your instance. You are choosing *when* deferrable maintenance is attempted, not *whether* the instance is ever patched or restarted. What the application sees is a short interruption: open connections are severed and in-flight transactions roll back; on a replicated deployment the standby is promoted and the same endpoint name now resolves to a different machine. Urgent work can still land outside the window, so the window is a scheduling control rather than a shield.

go deeper

for a junior

Recall the shape: a recurring slot, chosen by you or assigned by the platform, in which the provider may patch and restart the instance. Expect dropped connections during it rather than a silent, seamless update.

for a middle

Explain the mechanics: which layers get patched, why a replicated deployment turns a restart into a promotion, and why the endpoint name stays the same while the machine behind it changes.

for a senior

Show the operational judgment: choose the slot against your own batch calendar, align windows across dependent managed services, and prove the application reconnects rather than assuming it does.

for a principal

Frame it as an estate standard: one window convention, recorded in the change calendar, and a client-side contract that every service meets, so provider-initiated restarts stop generating incidents.

## What a maintenance window is A managed service is software the provider operates on machines the provider owns, and both layers need routine work: fixes to the service engine, patches to the operating system underneath it, and firmware or hypervisor updates below that. A **maintenance window** is the recurring slot - typically expressed as a day of the week, a start time in a fixed time zone, and a duration in minutes - during which the provider is permitted to carry out that routine work on your instance. The window is a scheduling control, not an exemption. You are choosing **when** deferrable maintenance is attempted, not **whether** the service is ever patched or restarted. Providers let you nominate the slot, and if you do not nominate one the platform assigns one on your behalf - which is precisely why picking it deliberately is worth the five minutes it takes. The same idea appears on every managed product an estate runs, not only on databases: a managed cache, a managed message broker and a managed search cluster each carry a window of their own, and they are rarely aligned unless someone aligned them. ## What the provider may do inside it - Apply a **minor patch** to the service engine - a fix inside the same major version line. - Patch or reboot the **host** underneath the instance, or migrate the instance onto an already-patched host. - Update the layers you never see: firmware, the hypervisor, the storage fabric. - Carry out a change **you** requested and chose to defer, such as a resize or a setting that only takes effect after a restart. - On a replicated deployment, **promote the standby replica**, patch what used to be the primary, and leave it as the new standby. ## What your application sees The honest answer is a short interruption, not an invisible change. A typical sequence: 1. Open connections to the instance are severed, and transactions in flight roll back rather than completing. 2. The connection endpoint name stays the same, but on a replicated deployment it now resolves to the promoted replica's address. 3. New connections succeed as soon as the new primary accepts them. Clients holding pooled dead connections, or caching the previously resolved address, keep failing until they re-resolve and reconnect. 4. The old primary is patched in the background and rejoins as the standby, usually with no further client-visible event. | Deployment shape | What maintenance does | What the client sees | |---|---|---| | Single instance | Patch and restart in place | Connections dropped for the length of the restart | | Primary with a standby | Promote the standby, patch the old primary after | A brief failover, then service from a different machine | | Extra read replicas | Each replica patched on its own pass | Staggered short interruptions, one replica at a time | ## Choosing the window well - Pick **your** traffic trough, not the region's. An overnight slot in the region's local time may sit exactly on top of a nightly batch run. - Look at scheduled work, not only interactive traffic: reporting jobs, exports and reconciliation runs are the things that notice a dropped connection most. - In a regulated business, record the window in the change calendar, so an unattended restart is recognised as planned maintenance instead of being triaged as an incident in the middle of the night. - Keep one convention across the estate so whoever is on call knows when platform-generated noise is expected, and so two dependent services are not patched in the same minute. ## Where the window stops protecting you Three limits are worth stating plainly. - **Urgent work can land outside it.** Providers reserve the right to apply a critical fix, or an upgrade whose deadline has passed, outside the nominated slot. The window governs deferrable work. - **Deferral is finite.** A pending action you keep postponing is normally applied by a stated date, and deferred items stack: the window you have skipped three times becomes the window that does three things at once. - **The duration is not a downtime budget.** A one-hour window does not mean an hour of downtime, and it does not entitle you to an hour of it either. It bounds when work may start, not what that work costs you, and what you are owed if maintenance breaks the service is a separate question about the availability commitment. The practical consequence is that a maintenance window only helps if the application survives a restart at an arbitrary moment inside it. Connection pools that validate before handing a connection out, bounded connection lifetimes, and jobs that can be re-run without double-applying their effects turn a planned promotion into a log line. Without them, a sixty-second failover turns into a morning of manual recovery.

  • What happens if you never nominate a maintenance window?
    You still have one. The platform assigns a slot when the instance is created, usually derived from quiet hours in the region rather than quiet hours for your workload. That is the argument for setting it yourself: an assigned window is as likely to land on your nightly batch run as anywhere else.
  • Can maintenance ever run outside the window you nominated?
    Yes. The window governs routine, deferrable work. A critical security fix, or an upgrade whose end-of-support date has already passed, can be applied outside it, and a pending change you have deferred repeatedly is eventually applied by a stated deadline. Treat the window as scheduling, not as a veto.
  • Does a replicated deployment make maintenance invisible to clients?
    It shortens the interruption, it does not remove it. Promoting the standby cuts the outage to the time the failover takes, but open connections are still severed and in-flight transactions still roll back. Clients need to reconnect and re-resolve the endpoint, so the client side decides whether the event is visible.

Like the notified overnight maintenance slot in an office building: you get to say which night the lift is out, not whether it is ever serviced, and anyone working late still takes the stairs.

saying these in an interview costs you the question

  • Thinks a maintenance window means the service is never restarted
  • Assumes the provider can never act outside the nominated window
  • Reads the window duration as the downtime the provider is allowed
  • Picks the slot by convenience without checking batch and reporting schedules
  • Believes a standby replica makes the restart invisible with no client changes
  • Assumes deferring maintenance is something you can keep doing indefinitely
open as a page

Why does a managed database service apply minor patches for you but wait for you to start a major version upgrade?

level: middleimportance: must knowfreq 56%

basics

~20 s

A minor patch stays inside one major version line and is meant to preserve behaviour, so the provider can apply it fleet-wide. A major upgrade changes behaviour the provider cannot test against your workload, so you schedule it, test it and own its rollback.

open as a page

A deprecation notice gives your managed engine's major version an end-of-support date nine months out - what work does that create?

level: seniorimportance: should knowfreq 50%

basics

~20 s

It converts a date into an engineering project: inventory every instance on that version, find what the new version breaks, rehearse the upgrade on restored data, and schedule a cutover with a way back. Let the date pass and the provider upgrades you on its own schedule.

open as a page

Maintenance promoted the standby and the database was back within a minute, yet a nightly job kept failing for an hour - why?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The platform recovered; the client did not. Pooled connections opened before the promotion stayed in the pool, so every borrow handed the job a dead session, and a cached address plus no reconnect logic kept it failing until the process was restarted.

open as a page

Your provider is retiring the instance family your managed instances run on - how does that differ from an engine version reaching end of support?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

The lifecycle is the same - a notice, a date, a forced action - but the change is underneath rather than inside. A family retirement replaces the machine, so the work is a move plus performance and cost re-checking, with no query compatibility to test.

open as a page