skip to content

What does Elastic Cloud manage for you, and what stays your responsibility?

level: middleimportance: must knowfreq 62%

answer

  1. someone else's hardware, still your data model
  2. infrastructure bought, design decisions not
  3. think mappings, shards, ILM, queries
  4. no shell, no arbitrary JVM flags
  5. the bill is still yours to control

basics

~10 s

Elastic Cloud handles provisioning, node placement, TLS endpoints, backups, hardware replacement and orchestrated upgrades. Data modelling stays yours: mappings and analysis, shard counts, ILM policies, queries and relevance, role design, and cost control.

solid answer

~40 s

The managed service owns the **infrastructure layer**: it provisions instances per data tier, spreads them across availability zones, terminates TLS on a stable endpoint, replaces failed hardware, takes automatic snapshots to Elastic-managed object storage, and orchestrates rolling version upgrades from a button. It also removes shell access and arbitrary JVM tuning — heap is derived from the node size you pick, and `elasticsearch.yml` is only editable through an allow-listed user-settings box. Everything that determines whether the cluster performs well stays yours: **mapping and analyzer design, shard count and shard sizing, ILM policy authorship, query shape and relevance tuning, role and API-key design, client retry behaviour, and the bill**. A managed service will happily run a badly sharded index into the ground. The interview point is that Elastic Cloud buys away operations, not data modelling.

go deeper

for a junior

Be able to say that Elastic runs the servers, backups and upgrades while you still design indices and queries. Naming two concrete items on each side is enough at this stage.

for a middle

Explain the mechanics behind the split: per-tier instance sizing, orchestrated rolling upgrades, a managed snapshot repository, and the restrictions — no shell, no JVM flags, allow-listed user settings, plugins as uploaded extensions.

for a senior

Demonstrate that you know which failures the platform will not save you from: mapping explosions, oversharding, expensive aggregations, missing ILM. Be ready to describe traffic filters and API-key scoping as your responsibility too.

for a principal

Own the model choice — hosted, Elastic Cloud Enterprise, Kubernetes operator, or Serverless — and the governance around it: retention and cost policy, exit strategy through your own snapshot repository, and where compliance obligations land.

## Why this question gets asked Interviewers use the buy-vs-run framing to find out whether a candidate has actually operated a cluster or has only used one. The tell is whether you can name the seam accurately in both directions: what genuinely disappears from your plate, and what stubbornly does not. This answer assumes hosted Elastic Cloud deployments running Elasticsearch 8.x/9.x; the Serverless offering moves the line further toward the provider and is discussed at the end. ## What Elastic runs for you **Provisioning and topology.** You describe a deployment as sizes per component — hot tier, warm tier, cold and frozen tiers, Kibana, machine-learning nodes, an integrations server — and choose how many availability zones each spans. The platform creates the instances, configures discovery, and adds dedicated master-eligible nodes as the deployment grows rather than making you plan them. **Endpoints and TLS.** Each deployment gets a stable HTTPS endpoint with a valid certificate. You never manage node certificates, keystores, or the transport-layer security configuration between nodes. **Hardware failure.** When an instance dies, the platform replaces it and the cluster recovers shard copies onto the new node. This is why you address the deployment endpoint and never a node. **Backups.** A hosted deployment gets a managed snapshot repository backed by object storage (named `found-snapshots`) and an automatic snapshot lifecycle policy, so a baseline backup exists without you configuring one. **Version upgrades.** Selecting a new version triggers an orchestrated rolling upgrade — node ordering, restarts, and health checks are the platform's job. **Scaling mechanics.** Resizing a tier is a plan change: the platform grows instances or adds them and rebalances. Autoscaling can raise data-tier capacity as storage demand grows, and can move machine-learning capacity in both directions. ## What stays yours **Mappings and analysis.** Field types, multi-fields, analyzers, `dynamic` behaviour and the field-count budget are entirely your design. A mapping explosion from unbounded dynamic fields is your outage, not the provider's. **Shard strategy.** Primary shard counts, index sizing, rollover thresholds and the resulting shard-per-node density are yours. Oversharding is the single most common self-inflicted problem on managed clusters, and no amount of paid infrastructure fixes it. **Lifecycle policy.** ILM runs inside Elasticsearch, but the policy — when to roll over, when to move to warm, when to force-merge, when to delete — is authored by you, and it is what makes tiered storage cost-effective. **Queries and relevance.** Query shape, filter-context usage, aggregation cost, pagination strategy and relevance tuning are application concerns. Slow searches on a managed cluster are almost always query or mapping problems. **Security inside the cluster.** Roles, document- and field-level security, API keys and their scoping and rotation are yours. So is deciding on traffic filters — a hosted endpoint is internet-reachable until you restrict it with IP allow-lists or private-link connectivity. **Client behaviour.** Bulk sizing, retry-and-backoff on `429`, and circuit-breaking in the application belong to you. **Cost.** The service makes capacity a slider, which makes overspending easy. Right-sizing tiers and retention is an ongoing task with no automated substitute. ## Where the boundary bites The restrictions are the part candidates forget. There is no shell on the nodes. JVM options are not user-settable — heap follows the instance size you choose. `elasticsearch.yml` is editable only through the user-settings field, and only for allow-listed settings that vary by version. Plugins cannot be dropped onto the filesystem: supported plugins are selected from a list, and custom artifacts such as synonym or stopword dictionaries are uploaded as extensions and referenced from analyzers. If your design depends on a bespoke native plugin or a kernel-level tuning knob, the hosted service is the wrong shape and you should say so. ## Middle options and Serverless Between "fully hosted" and "you run everything" sit Elastic Cloud Enterprise, which brings the orchestration layer onto your own hardware, and Elastic Cloud on Kubernetes, an operator that manages clusters as custom resources in your own Kubernetes estate. Both hand back infrastructure control while keeping some of the orchestration benefit. Elastic Cloud Serverless goes the other way: it removes node and shard sizing from the user entirely, so the responsibility line sits further toward the provider than in a hosted deployment. Be explicit about which model you are describing, because the answer differs. ## How to answer Name three or four items on each side, then state the principle: managed hosting removes the operations that scale with node count, and removes none of the design decisions that scale with data and query volume.

  • How do you add a custom synonym file to a hosted Elastic Cloud deployment?
    You cannot copy files onto the nodes — there is no shell. Instead you upload the dictionary as a deployment **extension** (a bundle) through the console or API, apply it to the deployment, and reference the file by its bundle-relative path from the analyzer's synonym filter. Updating the file means uploading a new bundle version and applying it, which triggers a plan change.
  • If Elastic Cloud takes automatic snapshots, why would you still configure your own repository?
    The automatic repository is managed by Elastic and tied to the deployment's lifecycle, so it is a convenience backup rather than an independently owned one. Organisations that need long retention, their own object-store control, cross-region or cross-account copies, or a clean exit path to a self-managed cluster register an additional custom repository and run their own SLM policy against it.
  • Why can't you just set a larger heap when a node hits its circuit breaker on Elastic Cloud?
    JVM options are not user-configurable on hosted deployments; heap is derived from the instance size for that tier. The lever is the plan — pick a larger node size — or, better, fix the cause: expensive aggregations, high field-data usage, too many shards per node, or unbounded bucket counts in a query.

It is like leasing a fully serviced apartment: the landlord handles the roof, the wiring and the boiler, but nobody else decides how you arrange the furniture — and a bad layout is still a bad layout.

saying these in an interview costs you the question

  • Assumes the provider tunes shards and mappings for you
  • Thinks a managed cluster cannot go red
  • Believes you can SSH in to edit elasticsearch.yml
  • Says autoscaling removes the need for capacity planning
  • Treats the hosted endpoint as private by default

context