How do SolrCloud's NRT, TLOG and PULL replica types differ?
answer
- Three letters: N, T, P
- One of them can never lead
- Two of them keep a transaction log
- Ask who does the Lucene indexing work
- Freshness versus duplicated CPU
basics
~20 sNRT replicas index documents locally and keep a transaction log, so they support near-real-time search and can become leader. TLOG replicas keep a transaction log but copy the leader's index instead of indexing, and can become leader. PULL replicas only copy the index and can never lead.
solid answer
~60 sAll three are full copies of a shard, but they differ in **how they get their data** and **whether they can be elected leader**. - **NRT** (the default) indexes every document locally into its own Lucene index and writes to its own transaction log. Because it indexes locally it can soft-commit and open a new searcher quickly, which is what gives near-real-time search. It is leader-eligible. - **TLOG** keeps a transaction log but does *not* index documents itself. It replicates the leader's index files after the leader commits. It is leader-eligible: if elected, it starts indexing locally and behaves like an NRT leader. - **PULL** keeps no transaction log and does no indexing — it only pulls index files from the leader. It can **never** become leader, so a shard whose only survivors are PULL replicas has no leader and cannot accept writes. The trade-off is CPU versus freshness: TLOG and PULL replicas do the analysis and merge work once on the leader rather than N times, but their searchers only advance when replication runs, so they lag behind the leader.
code
bash · 2 lines# Two leader-eligible TLOG replicas per shard plus two read-only PULL replicas
curl "http://localhost:8983/solr/admin/collections?action=CREATE&name=products&numShards=2&tlogReplicas=2&pullReplicas=2&collection.configName=products_config"go deeper
Recall the three names and the headline difference: NRT indexes locally, TLOG and PULL copy the leader's index, and PULL can never become leader. NRT is the default if nobody chose.
Explain the mechanics — who keeps a transaction log, who runs the analysis chain, and why replication-based replicas only see new data after the leader hard-commits.
Show you would size the mix in production: at least two leader-eligible replicas per shard, PULL copies for read fan-out, shards.preference to route traffic, and commit settings tuned for the resulting lag.
Own the trade-off across the platform — indexing CPU and merge I/O versus query capacity and freshness SLA — and decide whether a lag of seconds is acceptable to the product before the topology is chosen.
## Why more than one replica type exists In the original SolrCloud model every replica of a shard did the same work: each one received every document, ran the full analysis chain, built its own segments and merged them. That is excellent for freshness — a soft commit on any replica opens a searcher over data it has already indexed — but it multiplies indexing CPU and merge I/O by the replication factor. Solr 7 added replica *types* so you can decide, per collection, how much of that duplicated work you want to pay for. ## NRT (near real time) An NRT replica is the classic behaviour and remains the default (`nrtReplicas`, core names ending `_replica_n1`). It receives every update from the shard leader, appends it to its own transaction log, and indexes it into its own Lucene index. Because the data is genuinely in this replica's index, a **soft commit** can open a new searcher immediately, so documents become visible within roughly the soft-commit interval. NRT replicas are leader-eligible. Cost: every replica repeats the analysis, indexing and merging work, and every replica's merge activity competes with its own query traffic. ## TLOG A TLOG replica (`tlogReplicas`, `_replica_t1`) maintains a transaction log — it durably records the updates it is sent — but it does **not** index them into a local Lucene index. Instead, it uses index replication to copy the leader's segment files after the leader has done a hard commit. Its searcher therefore advances in steps, when replication completes, not continuously. The transaction log is what makes it leader-eligible. If the leader is lost and a TLOG replica is elected, it replays what it needs from its log, switches into indexing mode, and from then on behaves as an NRT-style leader for the shard. ## PULL A PULL replica (`pullReplicas`, `_replica_p1`) is the cheapest: no transaction log, no local indexing, only index replication from the leader. It is a pure read-scaling copy. Because it has no transaction log, it cannot guarantee it holds every acknowledged update, so Solr excludes it from leader election entirely. That exclusion is the operational landmine: a shard configured with one NRT leader and several PULL replicas has **no failover for writes**. Lose the NRT replica and the shard is leaderless — queries may still be served from the PULL copies, but indexing to that shard fails until the leader-eligible replica returns. ## Choosing a mix - **All NRT** — the default, and correct when you need sub-second visibility of new documents and your indexing rate is moderate. - **All TLOG** — write-heavy workloads where a few seconds (or minutes) of search lag is acceptable. Indexing and merging happen once, on the leader; every replica can still take over as leader. - **TLOG leaders + PULL replicas** — heavy read scaling. Keep at least two TLOG replicas per shard so leadership can move, and add PULL replicas for query fan-out. - **NRT + PULL** — workable, but keep at least two NRT replicas per shard, for the leaderless reason above. You can steer queries toward a particular type with `shards.preference`, for example `shards.preference=replica.type:PULL` to keep analytics traffic off leader-eligible replicas, optionally combined with `replica.location:local` to prefer a same-node replica. ## Consistency consequences With all-NRT replicas, two consecutive queries hitting different replicas can still disagree briefly (their soft commits are not synchronised), but the window is small. With TLOG or PULL replicas the window is a whole replication cycle: a document indexed on the leader is invisible on the followers until the leader hard-commits and the followers pull the new segments. If your application does an index-then-immediately-search flow, that lag is a functional bug waiting to happen — either stay all-NRT, or route those reads to the leader (`shards.preference=replica.leader:true`). Also note the interaction with commit settings: TLOG and PULL replicas depend on the leader's *hard* commits to produce new segment files to replicate, so a very long `autoCommit` interval on the leader directly becomes search lag on the followers. ## Naming and inspection Replica type shows up in the core name (`_replica_n`, `_replica_t`, `_replica_p`) and in the CLUSTERSTATUS output, so you can always confirm what a running cluster actually has rather than what the create command intended.
- What happens to a shard whose only leader-eligible replica goes down, leaving just PULL replicas?The shard is left with no leader. PULL replicas cannot be elected because they keep no transaction log and therefore cannot prove they hold every acknowledged update. Queries can still be served from the stale PULL copies, but any update routed to that shard fails until a TLOG or NRT replica for the shard comes back or is added. This is why a PULL-heavy layout still needs at least two leader-eligible replicas per shard.
- Why do TLOG replicas lag behind the leader in search results?Because they never index locally. They only receive new segment files through index replication, which the leader can offer only after a hard commit produces them. Their searcher therefore advances in discrete jumps rather than continuously, so visibility lag is roughly the leader's hard-commit interval plus the replication and searcher-warming time — not the soft-commit interval that governs NRT replicas.
- How do you keep expensive analytics queries off your leader-eligible replicas?Use `shards.preference` on the request, for example `shards.preference=replica.type:PULL`, so the aggregator prefers PULL replicas when choosing which replica of each shard to query. You can chain preferences — `replica.location:local` first to avoid a network hop, then `replica.type:PULL`. It is a preference, not a hard restriction: if no matching replica is available, Solr falls back to another one.
- When a TLOG replica is elected leader, what does it do differently?It stops being a passive replication target and starts indexing. It replays from its transaction log as needed to catch up to the last acknowledged updates, then accepts updates for the shard, assigns versions, indexes them into its own Lucene index and distributes them to the other replicas. In other words a TLOG replica behaves like an NRT replica for exactly as long as it holds leadership.
saying these in an interview costs you the question
- Says PULL replicas can be elected leader
- Thinks TLOG replicas index documents locally
- Claims all replica types give identical search freshness
- Runs one NRT plus many PULL replicas with no failover
- Assumes replica type is set per node rather than per collection