skip to content

In a partitioned datastore where the primary key determines which partition a row lives on, what's the difference between a local secondary index and a global secondary index on some other (non-partition-key) attribute, and what does each cost you on reads versus writes?

level: seniorimportance: must knowfreq 55%

answer

  1. local = same partition, write cheap read fans out
  2. global = own partitioned index, read direct write propagates
  3. GSI eventually consistent, LSI strongly consistent
  4. local index query cost scales with cluster size
  5. DynamoDB LSI vs GSI

basics

~20 s

A local secondary index lives on the same partition as the data it indexes, so writes stay on one partition but a query has to check every partition to be complete. A global secondary index is its own separate partitioned structure, so a query can hit just the index's partitions directly, but every write now has to update two different partitions.

solid answer

~50 s

When you partition a table by its primary key, queries on a different attribute (e.g., find all orders for status=shipped when the table is partitioned by order_id) don't know which partition to check. A local secondary index solves this by building the index co-located within each partition, indexed only over that partition's local data; a write only touches one partition, keeping writes single-partition and atomic, but a query on the indexed attribute must be sent to every partition and results merged (scatter-gather). A global secondary index instead builds a separate, independently partitioned index structure — partitioned by the indexed attribute — so a query goes directly to the relevant index partition(s) without fanning out. The cost shifts to writes: every write must also update the global index, which usually lives on a different partition/node, making writes cross-partition and typically asynchronous/eventually consistent. DynamoDB's GSIs are eventually consistent for exactly this reason; its local secondary indexes are strongly consistent because they share a partition with the base data.

go deeper

for a junior

Should grasp that a secondary index on a non-partition-key attribute needs extra machinery, and that there's a basic write-vs-read cost trade-off, without necessarily naming local/global by name.

for a middle

Should correctly describe the mechanism difference (co-located vs separately-partitioned index) and know which side of the write/read trade-off each falls on.

for a senior

Should explain the consistency implication (why GSIs are typically eventually consistent, why LSIs can be strongly consistent) and reason about which to choose for a given query/write pattern, citing a concrete system like DynamoDB.

for a principal

Should evaluate this trade-off at the schema/system-design level — e.g., denormalizing into purpose-built query tables instead of relying on secondary indexes, or routing to an external index (Elasticsearch/OpenSearch) when neither meets the requirements — and reason about monitoring needed to catch each failure mode.

## The problem a secondary index solves Once a table is partitioned by a primary (partition) key, the system knows exactly which partition to check for any query that includes that key. The problem arises for queries on a different attribute that isn't the partition key: 'find all orders with `status=shipped`' in a table partitioned by `order_id` gives you no way to know which partition holds matching rows. A secondary index exists to answer exactly this kind of query efficiently, and the two dominant designs — **local** and **global** — differ in where the index data physically lives relative to the base data, driving a direct trade-off between write cost and read fan-out cost. ## The local secondary index A local secondary index is built and stored **within each partition**, indexing only the rows that already live on that partition. Concretely, partition 3 maintains its own local index mapping `status` to the `order_ids` on partition 3 with that status, and partition 7 maintains a separate, independent local index for its own rows. - **Writes are cheap.** Because the index entry for a row always lives on the exact same partition as the row itself, writing or updating a row only ever touches one partition — the base row and its local index entry update together, atomically, without cross-partition coordination. Single-partition writes are cheap and low-latency, and don't need distributed transaction machinery. - **The cost lands on reads.** Since there's no global knowledge of which partitions contain matching rows, a query on the index has to be sent to every partition (scatter-gather), each partition searches its own local index, and the coordinator merges partial results. Query cost scales with the number of partitions regardless of how selective the query actually is. ## The global secondary index A global secondary index takes the opposite approach: it's built as an entirely separate index structure, **partitioned independently** — typically partitioned by the indexed attribute itself rather than by the base table's primary key. - **Reads go straight to the answer.** Now a query like '`status=shipped`' can hash or range on `status` directly and go straight to the specific index partition(s) holding matching entries. This gives read performance that scales with query selectivity rather than cluster size. - **The cost moves to the write path.** Any write to the base table must also propagate an update to the global index, which — because it's partitioned differently — usually lives on a different physical partition/node. Keeping both in sync synchronously would require a distributed transaction on every write, hurting write latency and availability. In practice, most systems make this propagation asynchronous, accepting **eventual consistency**: the global index catches up shortly after the write, but a read against the index immediately after a write can miss it or see stale data. | Design | Where the index lives | Write path | Read path | |---|---|---|---| | local | within each partition | one partition, together, atomically | sent to every partition (scatter-gather) | | global | an entirely separate structure, partitioned independently | propagate an update to a different physical partition/node | straight to the specific index partition(s) | ## Where you meet this: DynamoDB The concrete, named example most engineers meet this in is **DynamoDB**: - its **Local Secondary Indexes** are constrained to share the same partition key as the base table (only the sort key differs) and are strongly consistent, because index and data are physically co-located and updated together; - its **Global Secondary Indexes** can be built on any attribute, are independently partitioned, support querying efficiently across the whole table's keyspace, but are explicitly documented as eventually consistent — a write to the base table might not be visible in a GSI query for a short window afterward. This isn't a DynamoDB quirk; it's the direct structural consequence of the local-vs-global design. The same trade-off shows up in other partitioned systems (e.g., Cassandra's now-deprecated built-in secondary indexes were local-index-style with the same scatter-gather read cost, part of why Cassandra users are steered toward denormalized tables or external indexing like Elasticsearch instead). ## The two failure modes in production - **The local-index failure mode** shows up as query latency and load that scales linearly with cluster size no matter how selective the filter is — adding nodes to handle more data actually makes local-secondary-index queries slower, since more partitions must be scanned, a counter-intuitive and easy-to-miss operational trap. - **The global-index failure mode** shows up as stale-read bugs: an application writes a row, immediately queries the global index expecting to find it, and doesn't — a class of bug invisible in local testing (near-zero replication lag) but appearing intermittently in production under load. ## The decision rule The practical decision rule: 1. use local secondary indexes when write cost and consistency matter more and the cluster is small enough (or queries selective enough within a partition) that fan-out is tolerable; 2. use global secondary indexes when read selectivity across the whole dataset matters and the application can tolerate eventual consistency on the indexed view.

  • Why can't a global secondary index just be updated synchronously with the base table write to avoid eventual consistency?
    Because the index lives on a different partition/node than the base row, a synchronous update would require a distributed transaction across both, which adds latency to every write and reduces availability — if the index partition is unreachable, the base write would have to block or fail too. Most systems trade that away for asynchronous propagation, accepting a brief staleness window for fast, available single-partition writes.
  • If a table's local secondary index query has to scatter-gather across every partition, why is a local secondary index ever preferable to a global one?
    When the query is naturally scoped within a single partition already (e.g., DynamoDB's LSI requires the same partition key, just a different sort key), there's no fan-out at all — you get the strong-consistency and low-write-cost benefits with none of the scatter-gather cost, because the query only ever needs one partition to begin with.
  • What operational symptom would tip you off that an application is relying on a local (scatter-gather) secondary index at a scale where it shouldn't?
    Query latency and load on a secondary-index lookup that grows as the cluster adds partitions/nodes, even when the actual result set size and selectivity haven't changed — that's the signature of a fan-out cost scaling with node count rather than with the number of matching rows.

A local secondary index is like each branch of a library keeping its own card catalog of just the books on its shelves — checking out or shelving a book only touches that branch's catalog, but finding 'every mystery novel in the whole library system' means calling every branch and asking them to check their own catalog. A global secondary index is like a single citywide card catalog organized by genre — you can look up 'mystery novels' in one place, but every branch has to phone in an update to that central catalog whenever a book arrives or leaves, and there's a delay before the citywide catalog reflects it.

saying these in an interview costs you the question

  • Thinks 'local' and 'global' refer to geographic data-center location rather than index co-location with the base partition
  • Doesn't know that global secondary indexes are typically eventually consistent and why
  • Assumes secondary indexes are free on the write path regardless of type
  • Can't explain why a local-index query has to touch every partition
  • Believes adding nodes always improves secondary-index query performance

context