In Solr, what is the difference between a core and a collection?
answer
- One is physical, one is logical
- Think about what lives on a single node
- Shards and replicas sit between the two
- A replica is implemented as one of them
- Collections exist only in cloud mode
basics
~20 sA Solr core is one physical Lucene index living on one node, with its own configuration and data directory. A collection is a SolrCloud-level logical index spread across shards and replicas; every replica is physically a core on some node.
solid answer
~50 sA **core** is Solr's unit of physical storage: one Lucene index, one `core.properties` file, one config set (`solrconfig.xml` plus the schema), one data directory, all on a single node. Standalone Solr serves cores directly — you talk to `/solr/<coreName>/select`. A **collection** only exists in SolrCloud. It is a logical index that is split into *shards* (disjoint slices of the document space), and each shard is materialised as one or more *replicas*. Each replica is, physically, a core on some node — Solr even names them that way, e.g. `products_shard1_replica_n1`. So the relationship is: collection → shards → replicas → cores. In SolrCloud you create and manage collections through the Collections API and let Solr create the underlying cores; using the CoreAdmin API to hand-make cores in a cloud cluster is how people corrupt cluster state.
code
bash · 4 lines# Standalone Solr: a core is a directory with a core.properties marker
# $SOLR_HOME/products/core.properties
name=products
configSet=_defaultgo deeper
Be able to say the one-liner: a core is one index on one node, a collection is the cloud-level index made of shards and replicas. Knowing that you query /solr/<collection>/select in cloud mode is enough at this level.
Explain the full chain — collection to shards to replicas to cores — and where the topology and configuration actually live (ZooKeeper state.json and configsets), not on local disk.
Show the operational consequences: why config changes go through upconfig plus RELOAD, why CoreAdmin is off-limits in cloud mode, and how aliases let you swap a collection under a stable name.
Own the layout decision — how many collections versus how many shards, whether tenants get their own collection, and how that choice affects ZooKeeper state size, config governance and reindex strategy.
## The core: Solr's physical unit A **core** is a single Lucene index plus everything needed to run it: a `solrconfig.xml` (request handlers, caches, update chain, commit settings), a schema (`managed-schema` or `schema.xml`), and a data directory holding the Lucene segments and the transaction log. On disk a core is announced by a small `core.properties` marker file; Solr walks `SOLR_HOME` at startup, finds those files, and loads a core for each one. That discovery mechanism is why a core is fundamentally a *node-local* thing — it cannot span two machines, and nothing about a core knows that other cores exist. In standalone (non-cloud) Solr, cores are the whole story. You address them directly: `/solr/products/select?q=laptop`. You create, reload, unload, swap or split them with the CoreAdmin API at `/solr/admin/cores`. If you want more than one index — say `products` and `logs` — you run two cores. ## The collection: SolrCloud's logical unit When Solr runs in cloud mode (`bin/solr start -c`, with a ZooKeeper ensemble), the addressable unit becomes a **collection**. A collection is a logical index that Solr may spread over many nodes: - The collection is divided into **shards**. With the default `compositeId` router, each shard owns a disjoint slice of a 32-bit hash space, so every document belongs to exactly one shard. - Each shard has one or more **replicas** — full copies of that shard's data. One replica per shard is the **leader** and coordinates indexing for the shard. - Every replica is physically a core on a node. Solr's generated core names make the mapping visible: `products_shard1_replica_n1`, `products_shard2_replica_t3` — collection, shard, replica type and number. So: one collection = N shards = N × R replicas = N × R cores, distributed over the cluster. ## Who owns what The collection's definition — its shards, their hash ranges, the list of replicas, their states and which one is leader — is not on disk next to the index. It lives in ZooKeeper, in `/collections/<name>/state.json`. The configuration a collection uses is also in ZooKeeper, as a **configset** under `/configs/<name>`, shared by every replica of the collection (and possibly by several collections). A standalone core, by contrast, reads its config from its own local directory. That difference explains the operational rule: **in SolrCloud, use the Collections API, not the CoreAdmin API.** `action=CREATE` on `/solr/admin/collections` computes shard ranges, picks nodes, creates the cores and writes the state; hand-creating a core with CoreAdmin produces an index Solr's cluster state does not know about, or worse, one it half-knows about. ## What this changes in practice - **Addressing.** In cloud mode you query `/solr/<collection>/select` on *any* node; that node becomes the aggregator and fans the query out to one replica of every shard. Querying a specific core (`/solr/products_shard1_replica_n1/select?distrib=false`) is a debugging tool, not a normal access path. - **Config changes.** Editing a file on one node does nothing useful in cloud mode. You upload the configset to ZooKeeper (`bin/solr zk upconfig`) and then RELOAD the collection so every replica picks it up. - **Scaling.** Adding capacity to a standalone core means a bigger machine. Adding capacity to a collection means more shards (for index size and indexing throughput) or more replicas (for query throughput and failure tolerance). - **Aliases.** Collections can sit behind an alias, so `products` can be repointed from `products_v1` to `products_v2` atomically. Cores have no such indirection. ## The sentence to say in an interview "A core is one Lucene index on one node with its own config and data directory; a collection is a SolrCloud logical index made of shards, and each shard replica *is* a core. Standalone Solr has only cores; SolrCloud manages cores for you through collections, with the topology stored in ZooKeeper."
- If a replica is just a core, why shouldn't I create cores directly with the CoreAdmin API in a SolrCloud cluster?Because the authoritative topology lives in ZooKeeper, not on disk. The Collections API assigns the shard's hash range, picks a node, creates the core, and registers the replica in `state.json` in one coordinated step. A hand-made core either never appears in cluster state — so it receives no routed documents and no queries — or appears inconsistently and confuses routing and leader election. CoreAdmin remains a legitimate debugging and standalone-mode tool.
- Where does a SolrCloud collection get its solrconfig.xml and schema from?From a configset stored in ZooKeeper under `/configs/<name>`, referenced by the collection at creation time via `collection.configName`. Every replica of the collection loads the same configset, and several collections may share one. You change it by uploading a new version (`bin/solr zk upconfig` or the Configsets API) and then issuing a RELOAD for the collection so each replica reopens with the new config.
- Can two collections share the same underlying core?No. A core belongs to exactly one replica of exactly one shard of exactly one collection; that ownership is baked into its `core.properties` and into cluster state. Two collections can share a *configset*, and an alias can make several collections queryable under one name, but the index files themselves are never shared.
saying these in an interview costs you the question
- Says a collection is just a bigger core
- Claims a single core can span multiple nodes
- Creates cores with CoreAdmin inside a SolrCloud cluster
- Thinks collections also exist in standalone mode
- Believes each shard is one core regardless of replicas