skip to content

Cores, Collections & SolrCloud

A standalone Solr serves cores; SolrCloud turns them into sharded, replicated collections coordinated by ZooKeeper. Interviewers ask about replica types and leader election because that is where SolrCloud's behaviour diverges most sharply from Elasticsearch.

on this pageshow

questions

6

In Solr, what is the difference between a core and a collection?

level: juniorimportance: must knowfreq 72%

answer

  1. One is physical, one is logical
  2. Think about what lives on a single node
  3. Shards and replicas sit between the two
  4. A replica is implemented as one of them
  5. Collections exist only in cloud mode

basics

~20 s

A Solr core is one physical Lucene index living on one node, with its own configuration and data directory. A collection is a SolrCloud-level logical index spread across shards and replicas; every replica is physically a core on some node.

solid answer

~50 s

A **core** is Solr's unit of physical storage: one Lucene index, one `core.properties` file, one config set (`solrconfig.xml` plus the schema), one data directory, all on a single node. Standalone Solr serves cores directly — you talk to `/solr/<coreName>/select`. A **collection** only exists in SolrCloud. It is a logical index that is split into *shards* (disjoint slices of the document space), and each shard is materialised as one or more *replicas*. Each replica is, physically, a core on some node — Solr even names them that way, e.g. `products_shard1_replica_n1`. So the relationship is: collection → shards → replicas → cores. In SolrCloud you create and manage collections through the Collections API and let Solr create the underlying cores; using the CoreAdmin API to hand-make cores in a cloud cluster is how people corrupt cluster state.

code

bash · 4 lines
bash
# Standalone Solr: a core is a directory with a core.properties marker
# $SOLR_HOME/products/core.properties
name=products
configSet=_default

go deeper

for a junior

Be able to say the one-liner: a core is one index on one node, a collection is the cloud-level index made of shards and replicas. Knowing that you query /solr/<collection>/select in cloud mode is enough at this level.

for a middle

Explain the full chain — collection to shards to replicas to cores — and where the topology and configuration actually live (ZooKeeper state.json and configsets), not on local disk.

for a senior

Show the operational consequences: why config changes go through upconfig plus RELOAD, why CoreAdmin is off-limits in cloud mode, and how aliases let you swap a collection under a stable name.

for a principal

Own the layout decision — how many collections versus how many shards, whether tenants get their own collection, and how that choice affects ZooKeeper state size, config governance and reindex strategy.

## The core: Solr's physical unit A **core** is a single Lucene index plus everything needed to run it: a `solrconfig.xml` (request handlers, caches, update chain, commit settings), a schema (`managed-schema` or `schema.xml`), and a data directory holding the Lucene segments and the transaction log. On disk a core is announced by a small `core.properties` marker file; Solr walks `SOLR_HOME` at startup, finds those files, and loads a core for each one. That discovery mechanism is why a core is fundamentally a *node-local* thing — it cannot span two machines, and nothing about a core knows that other cores exist. In standalone (non-cloud) Solr, cores are the whole story. You address them directly: `/solr/products/select?q=laptop`. You create, reload, unload, swap or split them with the CoreAdmin API at `/solr/admin/cores`. If you want more than one index — say `products` and `logs` — you run two cores. ## The collection: SolrCloud's logical unit When Solr runs in cloud mode (`bin/solr start -c`, with a ZooKeeper ensemble), the addressable unit becomes a **collection**. A collection is a logical index that Solr may spread over many nodes: - The collection is divided into **shards**. With the default `compositeId` router, each shard owns a disjoint slice of a 32-bit hash space, so every document belongs to exactly one shard. - Each shard has one or more **replicas** — full copies of that shard's data. One replica per shard is the **leader** and coordinates indexing for the shard. - Every replica is physically a core on a node. Solr's generated core names make the mapping visible: `products_shard1_replica_n1`, `products_shard2_replica_t3` — collection, shard, replica type and number. So: one collection = N shards = N × R replicas = N × R cores, distributed over the cluster. ## Who owns what The collection's definition — its shards, their hash ranges, the list of replicas, their states and which one is leader — is not on disk next to the index. It lives in ZooKeeper, in `/collections/<name>/state.json`. The configuration a collection uses is also in ZooKeeper, as a **configset** under `/configs/<name>`, shared by every replica of the collection (and possibly by several collections). A standalone core, by contrast, reads its config from its own local directory. That difference explains the operational rule: **in SolrCloud, use the Collections API, not the CoreAdmin API.** `action=CREATE` on `/solr/admin/collections` computes shard ranges, picks nodes, creates the cores and writes the state; hand-creating a core with CoreAdmin produces an index Solr's cluster state does not know about, or worse, one it half-knows about. ## What this changes in practice - **Addressing.** In cloud mode you query `/solr/<collection>/select` on *any* node; that node becomes the aggregator and fans the query out to one replica of every shard. Querying a specific core (`/solr/products_shard1_replica_n1/select?distrib=false`) is a debugging tool, not a normal access path. - **Config changes.** Editing a file on one node does nothing useful in cloud mode. You upload the configset to ZooKeeper (`bin/solr zk upconfig`) and then RELOAD the collection so every replica picks it up. - **Scaling.** Adding capacity to a standalone core means a bigger machine. Adding capacity to a collection means more shards (for index size and indexing throughput) or more replicas (for query throughput and failure tolerance). - **Aliases.** Collections can sit behind an alias, so `products` can be repointed from `products_v1` to `products_v2` atomically. Cores have no such indirection. ## The sentence to say in an interview "A core is one Lucene index on one node with its own config and data directory; a collection is a SolrCloud logical index made of shards, and each shard replica *is* a core. Standalone Solr has only cores; SolrCloud manages cores for you through collections, with the topology stored in ZooKeeper."

  • If a replica is just a core, why shouldn't I create cores directly with the CoreAdmin API in a SolrCloud cluster?
    Because the authoritative topology lives in ZooKeeper, not on disk. The Collections API assigns the shard's hash range, picks a node, creates the core, and registers the replica in `state.json` in one coordinated step. A hand-made core either never appears in cluster state — so it receives no routed documents and no queries — or appears inconsistently and confuses routing and leader election. CoreAdmin remains a legitimate debugging and standalone-mode tool.
  • Where does a SolrCloud collection get its solrconfig.xml and schema from?
    From a configset stored in ZooKeeper under `/configs/<name>`, referenced by the collection at creation time via `collection.configName`. Every replica of the collection loads the same configset, and several collections may share one. You change it by uploading a new version (`bin/solr zk upconfig` or the Configsets API) and then issuing a RELOAD for the collection so each replica reopens with the new config.
  • Can two collections share the same underlying core?
    No. A core belongs to exactly one replica of exactly one shard of exactly one collection; that ownership is baked into its `core.properties` and into cluster state. Two collections can share a *configset*, and an alias can make several collections queryable under one name, but the index files themselves are never shared.

saying these in an interview costs you the question

  • Says a collection is just a bigger core
  • Claims a single core can span multiple nodes
  • Creates cores with CoreAdmin inside a SolrCloud cluster
  • Thinks collections also exist in standalone mode
  • Believes each shard is one core regardless of replicas

context

open as a page

How do SolrCloud's NRT, TLOG and PULL replica types differ?

level: middleimportance: must knowfreq 60%

basics

~20 s

NRT replicas index documents locally and keep a transaction log, so they support near-real-time search and can become leader. TLOG replicas keep a transaction log but copy the leader's index instead of indexing, and can become leader. PULL replicas only copy the index and can never lead.

open as a page

How does SolrCloud's compositeId router decide which shard a document lands on?

level: middleimportance: should knowfreq 50%

basics

~20 s

It hashes the document's uniqueKey into a 32-bit value and sends the document to whichever shard owns that value's hash range. If the key contains a routing prefix such as tenant!docId, the high bits come from the prefix, so every document sharing that prefix lands on the same shard.

open as a page

What are the two phases of a distributed SolrCloud query across shards?

level: middleimportance: should knowfreq 45%

basics

~20 s

First, the receiving node asks one replica of every shard for the top matching document ids and sort values, and merges them into a global top-N. Second, it fetches the stored fields and highlighting for just those winning ids from the shards that hold them.

open as a page

How does a SolrCloud shard elect its leader, and why can a shard end up with no leader?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Leader-eligible replicas of a shard queue up on ephemeral sequential ZooKeeper nodes; the lowest sequence wins and publishes itself as leader. A shard is left leaderless when no eligible replica is available or up to date — for example only PULL replicas survive, or every replica is down.

open as a page

What happens to a SolrCloud cluster when its ZooKeeper ensemble loses quorum?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Solr keeps serving queries from its cached cluster state, but indexing stops: without a writable ZooKeeper it cannot update collection state, elect leaders or register recovering replicas. Solr treats ZooKeeper as the authority, so writes fail rather than proceeding blindly.

open as a page