skip to content

Solr

The other mature Lucene search server, still common in enterprise and library systems: schema-driven fields, strong faceting, and SolrCloud for distribution. Interviews usually want the comparison — where Solr and Elasticsearch differ in schema handling, APIs, and operational model.

on this pageshow

explore

questions

12

In Solr, what is the difference between a core and a collection?

level: juniorimportance: must knowfreq 72%

answer

  1. One is physical, one is logical
  2. Think about what lives on a single node
  3. Shards and replicas sit between the two
  4. A replica is implemented as one of them
  5. Collections exist only in cloud mode

basics

~20 s

A Solr core is one physical Lucene index living on one node, with its own configuration and data directory. A collection is a SolrCloud-level logical index spread across shards and replicas; every replica is physically a core on some node.

solid answer

~50 s

A **core** is Solr's unit of physical storage: one Lucene index, one `core.properties` file, one config set (`solrconfig.xml` plus the schema), one data directory, all on a single node. Standalone Solr serves cores directly — you talk to `/solr/<coreName>/select`. A **collection** only exists in SolrCloud. It is a logical index that is split into *shards* (disjoint slices of the document space), and each shard is materialised as one or more *replicas*. Each replica is, physically, a core on some node — Solr even names them that way, e.g. `products_shard1_replica_n1`. So the relationship is: collection → shards → replicas → cores. In SolrCloud you create and manage collections through the Collections API and let Solr create the underlying cores; using the CoreAdmin API to hand-make cores in a cloud cluster is how people corrupt cluster state.

code

bash · 4 lines
bash
# Standalone Solr: a core is a directory with a core.properties marker
# $SOLR_HOME/products/core.properties
name=products
configSet=_default

go deeper

for a junior

Be able to say the one-liner: a core is one index on one node, a collection is the cloud-level index made of shards and replicas. Knowing that you query /solr/<collection>/select in cloud mode is enough at this level.

for a middle

Explain the full chain — collection to shards to replicas to cores — and where the topology and configuration actually live (ZooKeeper state.json and configsets), not on local disk.

for a senior

Show the operational consequences: why config changes go through upconfig plus RELOAD, why CoreAdmin is off-limits in cloud mode, and how aliases let you swap a collection under a stable name.

for a principal

Own the layout decision — how many collections versus how many shards, whether tenants get their own collection, and how that choice affects ZooKeeper state size, config governance and reindex strategy.

## The core: Solr's physical unit A **core** is a single Lucene index plus everything needed to run it: a `solrconfig.xml` (request handlers, caches, update chain, commit settings), a schema (`managed-schema` or `schema.xml`), and a data directory holding the Lucene segments and the transaction log. On disk a core is announced by a small `core.properties` marker file; Solr walks `SOLR_HOME` at startup, finds those files, and loads a core for each one. That discovery mechanism is why a core is fundamentally a *node-local* thing — it cannot span two machines, and nothing about a core knows that other cores exist. In standalone (non-cloud) Solr, cores are the whole story. You address them directly: `/solr/products/select?q=laptop`. You create, reload, unload, swap or split them with the CoreAdmin API at `/solr/admin/cores`. If you want more than one index — say `products` and `logs` — you run two cores. ## The collection: SolrCloud's logical unit When Solr runs in cloud mode (`bin/solr start -c`, with a ZooKeeper ensemble), the addressable unit becomes a **collection**. A collection is a logical index that Solr may spread over many nodes: - The collection is divided into **shards**. With the default `compositeId` router, each shard owns a disjoint slice of a 32-bit hash space, so every document belongs to exactly one shard. - Each shard has one or more **replicas** — full copies of that shard's data. One replica per shard is the **leader** and coordinates indexing for the shard. - Every replica is physically a core on a node. Solr's generated core names make the mapping visible: `products_shard1_replica_n1`, `products_shard2_replica_t3` — collection, shard, replica type and number. So: one collection = N shards = N × R replicas = N × R cores, distributed over the cluster. ## Who owns what The collection's definition — its shards, their hash ranges, the list of replicas, their states and which one is leader — is not on disk next to the index. It lives in ZooKeeper, in `/collections/<name>/state.json`. The configuration a collection uses is also in ZooKeeper, as a **configset** under `/configs/<name>`, shared by every replica of the collection (and possibly by several collections). A standalone core, by contrast, reads its config from its own local directory. That difference explains the operational rule: **in SolrCloud, use the Collections API, not the CoreAdmin API.** `action=CREATE` on `/solr/admin/collections` computes shard ranges, picks nodes, creates the cores and writes the state; hand-creating a core with CoreAdmin produces an index Solr's cluster state does not know about, or worse, one it half-knows about. ## What this changes in practice - **Addressing.** In cloud mode you query `/solr/<collection>/select` on *any* node; that node becomes the aggregator and fans the query out to one replica of every shard. Querying a specific core (`/solr/products_shard1_replica_n1/select?distrib=false`) is a debugging tool, not a normal access path. - **Config changes.** Editing a file on one node does nothing useful in cloud mode. You upload the configset to ZooKeeper (`bin/solr zk upconfig`) and then RELOAD the collection so every replica picks it up. - **Scaling.** Adding capacity to a standalone core means a bigger machine. Adding capacity to a collection means more shards (for index size and indexing throughput) or more replicas (for query throughput and failure tolerance). - **Aliases.** Collections can sit behind an alias, so `products` can be repointed from `products_v1` to `products_v2` atomically. Cores have no such indirection. ## The sentence to say in an interview "A core is one Lucene index on one node with its own config and data directory; a collection is a SolrCloud logical index made of shards, and each shard replica *is* a core. Standalone Solr has only cores; SolrCloud manages cores for you through collections, with the topology stored in ZooKeeper."

  • If a replica is just a core, why shouldn't I create cores directly with the CoreAdmin API in a SolrCloud cluster?
    Because the authoritative topology lives in ZooKeeper, not on disk. The Collections API assigns the shard's hash range, picks a node, creates the core, and registers the replica in `state.json` in one coordinated step. A hand-made core either never appears in cluster state — so it receives no routed documents and no queries — or appears inconsistently and confuses routing and leader election. CoreAdmin remains a legitimate debugging and standalone-mode tool.
  • Where does a SolrCloud collection get its solrconfig.xml and schema from?
    From a configset stored in ZooKeeper under `/configs/<name>`, referenced by the collection at creation time via `collection.configName`. Every replica of the collection loads the same configset, and several collections may share one. You change it by uploading a new version (`bin/solr zk upconfig` or the Configsets API) and then issuing a RELOAD for the collection so each replica reopens with the new config.
  • Can two collections share the same underlying core?
    No. A core belongs to exactly one replica of exactly one shard of exactly one collection; that ownership is baked into its `core.properties` and into cluster state. Two collections can share a *configset*, and an alias can make several collections queryable under one name, but the index files themselves are never shared.

saying these in an interview costs you the question

  • Says a collection is just a bigger core
  • Claims a single core can span multiple nodes
  • Creates cores with CoreAdmin inside a SolrCloud cluster
  • Thinks collections also exist in standalone mode
  • Believes each shard is one core regardless of replicas

context

open as a page

In a Solr schema, what is the difference between the solr.StrField and solr.TextField field types?

level: juniorimportance: must knowfreq 72%

basics

~20 s

solr.StrField stores the value verbatim with no analysis, so it only matches, sorts and facets on the whole string. solr.TextField runs an analysis chain that tokenizes and normalizes text, enabling full-text matching on individual words.

open as a page

How do SolrCloud's NRT, TLOG and PULL replica types differ?

level: middleimportance: must knowfreq 60%

basics

~20 s

NRT replicas index documents locally and keep a transaction log, so they support near-real-time search and can become leader. TLOG replicas keep a transaction log but copy the leader's index instead of indexing, and can become leader. PULL replicas only copy the index and can never lead.

open as a page

Why do Solr applications send user queries through the edismax parser instead of the default lucene parser?

level: middleimportance: must knowfreq 74%

basics

~20 s

The default lucene parser expects strict query syntax and errors on stray characters, and it searches one default field. edismax accepts raw user text, searches many weighted fields via qf, and adds relevance controls such as mm, pf, tie and boost.

open as a page

How does SolrCloud's compositeId router decide which shard a document lands on?

level: middleimportance: should knowfreq 50%

basics

~20 s

It hashes the document's uniqueKey into a 32-bit value and sends the document to whichever shard owns that value's hash range. If the key contains a routing prefix such as tenant!docId, the high bits come from the prefix, so every document sharing that prefix lands on the same shard.

open as a page

What are the two phases of a distributed SolrCloud query across shards?

level: middleimportance: should knowfreq 45%

basics

~20 s

First, the receiving node asks one replica of every shard for the top matching document ids and sort values, and merges them into a global top-N. Second, it fetches the stored fields and highlighting for just those winning ids from the shards that hold them.

open as a page

In a Solr schema, how does a dynamicField pattern decide which incoming field names it matches, and what breaks a tie?

level: middleimportance: should knowfreq 52%

basics

~20 s

A dynamicField name is a glob with a wildcard only at the start or the end, such as i or attr. An explicitly declared field always wins over any pattern, and among competing patterns the longest one wins.

open as a page

What does Solr's schemaless mode do when it indexes a document containing an undeclared field?

level: middleimportance: should knowfreq 54%

basics

~20 s

It guesses a type from the value in the first document that carries the field, then permanently adds an explicit field declaration to the managed schema through an update processor chain. The guess is never revisited for later documents.

open as a page

How does a SolrCloud shard elect its leader, and why can a shard end up with no leader?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Leader-eligible replicas of a shard queue up on ephemeral sequential ZooKeeper nodes; the lowest sequence wins and publishes itself as leader. A shard is left leaderless when no eligible replica is available or up to date — for example only PULL replicas survive, or every replica is down.

open as a page

What happens to a SolrCloud cluster when its ZooKeeper ensemble loses quorum?

level: seniorimportance: should knowfreq 55%

basics

~20 s

Solr keeps serving queries from its cached cluster state, but indexing stops: without a writable ZooKeeper it cannot update collection state, elect leaders or register recovering replicas. Solr treats ZooKeeper as the authority, so writes fail rather than proceeding blindly.

open as a page

After adding fq=category:books, a Solr facet on category returns only that one bucket — why, and how do you keep the others?

level: seniorimportance: should knowfreq 56%

basics

~20 s

Facets are computed over the documents remaining after all filter queries, so filtering on category leaves only that value to count. Tag the filter and exclude it from that facet, with fq={!tag=cat} plus facet.field={!ex=cat}, to restore the full list.

open as a page

How does Solr's filterCache store an entry, and why can an fq clause using NOW make it useless?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Each filterCache entry maps one fq to a set of internal Lucene document ids matching it across the whole index, roughly one bit per document when dense. An fq containing NOW resolves to the current millisecond, producing a unique key per request and a permanent cache miss.

open as a page