skip to content

How does Elastic Cloud handle snapshots and version upgrades, and what still needs your action?

level: middleimportance: should knowfreq 44%

answer

  1. backups exist before you configure any
  2. incremental at the segment level
  3. the platform restarts, you prepare
  4. deprecations and old index formats first
  5. own bucket for retention and exit

basics

~20 s

Hosted deployments get a managed snapshot repository and an automatic snapshot policy, plus one-click orchestrated rolling upgrades. You still resolve deprecations, reindex indices from two majors back, verify client and extension compatibility, and own long-term backup retention.

solid answer

~50 s

Elastic Cloud registers a managed snapshot repository (`found-snapshots`) backed by object storage and runs an automatic snapshot lifecycle policy against it, so a baseline backup exists from day one. You can restore into the same deployment or spin up a new deployment from another deployment's snapshot. For anything beyond that — long retention, your own bucket, cross-account copies, or an exit path off the service — you register an **additional repository** and run your own SLM policy. Upgrades are orchestrated: pick a version and the platform performs a rolling restart in the right order. The **preparation is still yours**. Run the Upgrade Assistant, clear deprecation warnings, reindex or delete indices created two major versions back (Elasticsearch reads only the previous major's index format), check that any custom extensions and client libraries support the target version, and pick a window — a rolling upgrade still degrades capacity node by node.

code

bash · 7 lines
bash
# Inspect the platform-managed repository and its lifecycle policy
GET _snapshot/found-snapshots
GET _slm/policy
GET _slm/stats

# Confirm the most recent snapshot actually succeeded before an upgrade
GET _snapshot/found-snapshots/_current

go deeper

for a junior

Know that hosted deployments are snapshotted automatically and that upgrades are started from the console. Being able to say a backup exists by default, and that you still check deprecations, is enough here.

for a middle

Explain incremental segment-level snapshots, the managed repository, and the concrete pre-upgrade checklist: deprecations, indices from two majors back, extension and client compatibility.

for a senior

Show that you rehearse upgrades on a deployment restored from a snapshot, monitor SLM success, keep a customer-owned repository for retention and exit, and plan for the capacity dip during a rolling restart.

for a principal

Own the backup and upgrade policy across environments: retention tiers and where they live, who controls the bucket, an upgrade cadence that never lets indices fall two majors behind, and a documented exit path off the service.

## Snapshots on a hosted deployment Every hosted Elastic Cloud deployment comes with a snapshot repository already registered — it appears as `found-snapshots` — pointing at object storage managed by Elastic, together with an automatically created snapshot lifecycle management policy. This means a newly created deployment is being backed up before anyone has thought about backups, which is one of the more genuinely valuable things the managed service does. Snapshots in Elasticsearch are incremental at the **segment** level: a snapshot references segment files already in the repository and only uploads segments that are new since the last snapshot. That is why frequent snapshots of a large index are far cheaper than the index size suggests, and it is also why deleting an old snapshot rarely frees the space you expect — shared segments stay until the last snapshot referencing them is gone. Restoring from the managed repository works inside the deployment, and the console additionally lets you create a **new deployment restored from another deployment's snapshot**, which is the normal way to clone production into a staging environment or to recover from a catastrophic misconfiguration. ## Why the automatic repository is not the whole backup story The managed repository is tied to the deployment and to Elastic's storage account. Teams with regulatory retention requirements, a need for backups under their own cloud account and IAM control, cross-region copies, or a desire to keep a clean migration path to a self-managed cluster register a **second, customer-owned repository** — an S3, GCS or Azure bucket — and drive it with their own SLM policy and retention rules. That customer-owned repository doubles as the exit strategy: snapshot out, restore into a cluster you run. A further caveat: a snapshot can only be restored into a cluster of the same major version or the next one, and a restored index still carries the Lucene format it was written with. Snapshots are not a way to escape the version-compatibility rules described below. ## Upgrades: what the platform does Selecting a newer stack version in the console triggers a plan change. The platform performs a rolling upgrade — taking nodes out one at a time, restarting them on the new version, waiting for health to recover before proceeding, and upgrading Kibana and other components in the correct order. You do not write the runbook, sequence the restarts, or manage the risk of two nodes down at once. ## Upgrades: what remains yours **Deprecations.** Kibana's Upgrade Assistant lists deprecated settings, mappings and API usage for the target major. Clearing these is application work — renaming settings, changing mapping constructs, adjusting queries — and it must happen before the upgrade, not after. **Old index formats.** Elasticsearch reads indices created by the **previous major version only**. Before a major upgrade, any index created two majors back must be reindexed into a current-format index or deleted. On long-lived clusters with old time-series indices this is the single most common upgrade blocker, and it takes real time and disk headroom. **Extensions and plugins.** Custom bundles and plugins are version-pinned. A synonym bundle or a plugin that has not been rebuilt for the target version will block or break the upgrade, so bundle compatibility is checked and updated as part of preparation. **Clients and integrations.** Language clients, Logstash pipelines and Beats have their own compatibility matrices. A major upgrade that leaves a client behind produces application errors that look like cluster problems. **Capacity and timing.** A rolling upgrade removes one node at a time from service, so indexing and search capacity dip for the duration and shard recovery consumes I/O afterward. On a tight cluster this is a real availability event; schedule it, and make sure indices have replicas so that a node leaving does not turn shards unavailable. **A verified snapshot.** Even with automatic snapshots, confirm a recent successful snapshot exists before a major upgrade rather than assuming it. ## Rehearsing the upgrade The strongest answer names the rehearsal: create a temporary deployment restored from a production snapshot, upgrade that, run the application's query suite and check relevance and aggregation results, then repeat on production. Managed hosting makes this cheap — the clone is a console action and the deployment can be deleted afterward — which turns "we hope it works" into a tested change. ## Failure modes to be able to name An upgrade that stalls because a shard cannot be allocated; an index in the old format discovered mid-upgrade; a query that silently changes behaviour because a deprecated construct was removed; an SLM policy quietly failing for days because nobody watches it; and a restore that fails because the target cluster is on an older major than the snapshot's source. Each of these is squarely on the customer's side of the line.

  • Why does deleting old snapshots often free far less storage than expected?
    Snapshots are incremental over segment files: many snapshots reference the same underlying segments. Deleting one snapshot only removes segment files that no remaining snapshot still references, so a long chain of snapshots over a slowly changing index shares nearly all its data. Space is reclaimed in steps as whole generations age out, not proportionally to snapshot count.
  • How would you rehearse a major-version upgrade of a production Elastic Cloud deployment?
    Create a temporary deployment restored from a recent production snapshot, upgrade that clone to the target version, then run the application's query and aggregation suite against it and compare results and latencies. Fix deprecations and reindex old-format indices in production based on what the rehearsal surfaced, then upgrade production and delete the clone.
  • An upgrade prepares fine but the deployment reports an index it cannot read. What happened?
    Almost certainly an index created two major versions before the target. Elasticsearch reads only the previous major's index format, so such indices must be reindexed into a current-format index or deleted before the upgrade proceeds. Long-retained time-series indices are the usual culprits, which is why lifecycle policies that delete or roll data forward prevent this class of blocker.

saying these in an interview costs you the question

  • Assumes automatic snapshots satisfy every retention requirement
  • Thinks an upgrade needs no preparation because it is managed
  • Believes a snapshot restore upgrades old index formats
  • Says a rolling upgrade has no capacity impact
  • Never checks that the snapshot policy is actually succeeding

context