skip to content

How do secondary indexes (@Indexed) enable derived queries on Redis repositories, and what are the costs?

level: seniorimportance: should knowfreq 35%

answer

  1. @Indexed -> SET keyspace:field:value of ids
  2. findByX reads/SINTERs those SETs then HGETALL
  3. phantom/idx bookkeeping keeps indexes correct on update/delete
  4. equality only; no ranges/LIKE
  5. TTL cleanup needs keyspace notifications enabled

basics

~20 s

Add @Indexed to a field; Spring keeps a Redis SET of ids per field value (like people:firstname:Ann). Then a derived finder such as findByFirstname reads that set. Without @Indexed you can't query by that field.

solid answer

~40 s

Redis has no query engine, so Spring Data emulates secondary indexes with extra SETs. Annotating a field @Indexed makes Spring maintain, on every save, a SET keyspace:field:value (e.g. people:firstname:Ann) containing the ids of matching entities, plus per-entity 'phantom' index bookkeeping so old entries are cleaned up on update/delete. A derived query method like List<Person> findByFirstname(String n) intersects the relevant index SETs (SINTER for ANDed conditions) to get ids, then loads each hash. You can index nested paths and use composite finders (findByFirstnameAndCity). Costs: every write updates all index SETs (more commands, more memory), only equality/simple matching is supported (no ranges/LIKE without extra work), stale indexes if TTL expiry cleanup relies on keyspace notifications being enabled, and no ad-hoc queries — you must index a field before you can query it.

code

java · 16 lines
java
@RedisHash("people")
public class Person {
    @Id private String id;
    @Indexed private String firstname;      // SADD people:firstname:<value> <id>
    @Indexed private String city;
}

public interface PersonRepository extends CrudRepository<Person, String> {
    List<Person> findByFirstname(String firstname);            // SMEMBERS people:firstname:Ann
    List<Person> findByFirstnameAndCity(String fn, String c);  // SINTER of two index sets
}

// Enable TTL-driven index cleanup via keyspace events:
@Configuration
@EnableRedisRepositories(enableKeyspaceEvents = RedisKeyValueAdapter.EnableKeyspaceEvents.ON_STARTUP)
public class RedisRepoConfig {}

go deeper

for a junior

Should know @Indexed is required to query a field via a derived finder.

for a middle

Should explain the per-value SET-of-ids mechanism and that only equality is supported.

for a senior

Should cover SINTER for composite queries, index bookkeeping on update/delete, and the keyspace-notification TTL gotcha.

for a principal

Should reason about write amplification, memory/cardinality tradeoffs, consistency limits, and when to abandon Redis repositories for RediSearch or a query store.

**The problem** — plain Redis lets you look up a value only by its key. There's no `SELECT ... WHERE firstname = 'Ann'`. So how does `PersonRepository.findByFirstname("Ann")` work? Spring Data Redis builds and maintains **secondary indexes** by hand out of ordinary Redis SETs. **@Indexed** — put `@Indexed` on a field of a `@RedisHash` entity: ```java @RedisHash("people") class Person { @Id String id; @Indexed String firstname; Address address; // @Indexed can go on address.city too } ``` Now, whenever you `save` a Person with `firstname = "Ann"` and id `42`, Spring executes (besides the `HSET people:42 ...`) a `SADD people:firstname:Ann 42`. That SET — key `keyspace:field:value` — is the index: it names every id whose firstname is Ann. **Phantom / bookkeeping keys** — to keep indexes correct on update and delete, Spring stores a helper structure (a per-entity SET, historically `people:42:idx`, and an internal 'phantom' copy of the aggregate) recording which index entries this entity currently belongs to. On `save`, Spring diffs old vs new index membership and issues `SREM`/`SADD` so a changed firstname removes `42` from the old value's SET and adds it to the new one. On `delete`, it removes the entity from all its index SETs and the id-keyspace SET. **How a derived query runs** — `findByFirstname("Ann")`: 1. Compute the index key `people:firstname:Ann`. 2. Read the id members (`SMEMBERS`). 3. Load each hash (`HGETALL people:<id>`) and map back to Person. For composite conditions like `findByFirstnameAndAddress_City("Ann","Berlin")`, Spring **intersects** the two index SETs with `SINTER people:firstname:Ann people:address.city:Berlin`, then loads the surviving ids. `Or` conditions use union semantics. **@GeoIndexed** — a special index type storing coordinates in a Redis GEO structure, enabling `findByLocationNear(Point, Distance)`. Beyond plain equality. **Capabilities and limits:** - **Only equality / simple matching** out of the box. No `LIKE`, no `<`/`>` ranges, no sorting by arbitrary fields — those aren't derivable from equality SETs. (You'd model ranges yourself with ZSETs via the template.) - **You must index up front.** A field without `@Indexed` cannot be queried by a derived method; there's no full-scan fallback that filters in Redis. - **Write amplification.** Every save/update/delete touches the id-keyspace SET plus one SET per indexed field plus the bookkeeping keys — several extra round-trips and more memory. Heavily-indexed, write-hot entities pay for it. - **Memory cost.** Each index value is a SET; high-cardinality fields (e.g., email) create many small SETs. - **TTL + expiry correctness.** If entities have TTL, Redis expiring the main hash does NOT automatically prune index SETs unless Redis **keyspace notifications** are enabled and the RedisKeyValueAdapter listens for expiry events (config `EnableKeyspaceEvents`). Otherwise indexes can point at ids whose hash is gone (stale index). This is the classic operational gotcha. - **Consistency** — index updates aren't a single atomic transaction with the hash write by default; a crash mid-save can leave indexes slightly out of sync. **When to use** — fine for a handful of low-cardinality lookup fields on aggregates you fetch by id most of the time (e.g., find sessions by userId). When you need rich querying, ranges, full-text, or aggregation, Redis repositories are the wrong tool — use a real query store or RediSearch, or drive ZSET/custom structures via RedisTemplate. **Relationship to the rest** — indexes are a repository feature (RedisMappingContext + IndexResolver), unrelated to your RedisTemplate serializer configuration. Enable with `@EnableRedisRepositories`; enable expiry cleanup with keyspace events.

  • Why can findByAgeBetween(...) not be satisfied by a plain @Indexed field?
    Equality indexes are SETs keyed by exact value; they can't answer range queries. Ranges need an ordered structure like a Redis ZSET, which @Indexed doesn't create — you'd model it manually via RedisTemplate or use RediSearch.
  • An entity has a TTL and an @Indexed field. After it expires, findByThatField still returns its id. Why?
    Redis expired the main hash but the index SET still lists the id because keyspace-notification-driven cleanup wasn't enabled. Enable EnableKeyspaceEvents so the adapter prunes indexes on expiry.

saying these in an interview costs you the question

  • Thinking you can query any field without @Indexed (there's no Redis-side full scan filter)
  • Expecting range/LIKE/sorting from @Indexed equality indexes
  • Ignoring write amplification and memory cost of indexes on write-hot entities
  • Assuming TTL automatically cleans index SETs without keyspace notifications

context