How does RedisHealthIndicator work, and how does one built-in indicator's status roll up into the overall /health status and HTTP code?
answer
- redis: INFO → redis_version
- StatusAggregator worst-wins: DOWN>OUT_OF_SERVICE>UP>UNKNOWN
- HttpCodeStatusMapper DOWN→503, UP→200
- non-critical dep can eject instance
- fix: disable or use readiness group
basics
~20 sRedisHealthIndicator asks Redis for its INFO (server version) via the connection factory; success is UP, a connection error is DOWN. The overall /health status is the worst of all indicators, and DOWN/OUT_OF_SERVICE map to HTTP 503 by default.
solid answer
~40 sRedisHealthIndicator uses the RedisConnectionFactory to open a connection and run the INFO command, reading redis_version (and cluster info for clustered setups) into the details; a connection or command failure yields DOWN. Rolling up: Actuator's StatusAggregator combines every contributor into one status by taking the worst, with the default severity order DOWN > OUT_OF_SERVICE > UP > UNKNOWN. So a single DOWN indicator (Redis, db, disk) makes the aggregate DOWN. That aggregate is mapped to an HTTP code by HttpCodeStatusMapper: UP → 200, and DOWN/OUT_OF_SERVICE → 503 by default. Because a non-critical cache like Redis going DOWN can eject the whole instance from a load balancer, teams often either disable it (management.health.redis.enabled=false) or move it out of the readiness health group so only truly critical dependencies gate traffic.
code
java · 13 lines// Keep a flaky cache from marking the whole app DOWN / 503.
// Option A: drop it entirely
management.health.redis.enabled=false
// Option B (better): keep it visible on /health but exclude it from readiness,
// so only critical deps gate traffic.
management.endpoint.health.probes.enabled=true
management.endpoint.health.group.readiness.include=db,readinessState
// /actuator/health/readiness now ignores redis;
// /actuator/health still shows it for monitoring.
// Override the HTTP mapping if you insist DOWN should still be 200 (rare):
// management.endpoint.health.status.http-mapping.DOWN=200go deeper
Know redis check hits Redis and one DOWN makes /health DOWN → 503.
Name the severity order and the UP→200 / DOWN→503 mapping.
Use health groups to keep non-critical deps out of readiness; know StatusAggregator/HttpCodeStatusMapper are configurable.
Define which dependencies are traffic-gating vs merely observable, and codify that via readiness groups and custom status ordering across services.
## RedisHealthIndicator Registered as **`redis`** when Spring Data Redis and a `RedisConnectionFactory` bean are present. Each scrape: 1. Obtains a connection via the `RedisConnectionFactory` (Lettuce by default, or Jedis). 2. Executes the Redis **`INFO`** command and reads server metadata — primarily `redis_version`. For a clustered `RedisConnectionFactory` it uses cluster info (e.g. `CLUSTER INFO`, reporting slot/cluster size). 3. Success → **`Status.UP`** with details like `version`. A connection failure or command error → **`Status.DOWN`** with the exception. ## Status aggregation — the rollup The `/actuator/health` endpoint gathers all contributors and must produce a single top-level status. That's the job of the **`StatusAggregator`**. The default (`SimpleStatusAggregator`) sorts by a severity order: ``` DOWN > OUT_OF_SERVICE > UP > UNKNOWN ``` The aggregate is the **worst** (highest-severity) status present. So if `redis` is `DOWN` while `db`, `ping`, and `diskSpace` are `UP`, the overall status is **`DOWN`**. You can customize the order with `management.endpoint.health.status.order`. ## HTTP status mapping A separate **`HttpCodeStatusMapper`** turns the aggregate `Status` into an HTTP response code. Defaults: - `UP`, `UNKNOWN` → **200** - `DOWN`, `OUT_OF_SERVICE` → **503** You can override via `management.endpoint.health.status.http-mapping.<STATUS>=<code>`. ## The operational problem — and the fix Because aggregation takes the worst status, a **non-critical** dependency dragging the whole endpoint to `DOWN`/503 will get the instance pulled from a load balancer or marked not-ready — even though the app can still serve most requests from, say, the database. Two standard remedies: 1. **Disable it** if it's truly non-critical: `management.health.redis.enabled=false`. 2. **Move it out of the readiness group** using health groups, so only critical dependencies gate traffic: ``` management.endpoint.health.group.readiness.include=db,readinessState ``` The main `/health` can still surface Redis for monitoring, while `/health/readiness` ignores it. ## Gotchas - **Lettuce vs Jedis** behavior on timeouts differs; a hung Redis can slow the scrape. Configure client timeouts. - **Cluster mode** reports differently and a partial cluster failure may still read as UP depending on cluster state. - Don't conflate 'the indicator ran' with 'the cache is usable' — INFO succeeding doesn't guarantee your keyspace/commands work under memory pressure (e.g. OOM eviction), but it's a good reachability signal.
- Redis is a best-effort cache, but when it's down your pods drop out of the load balancer. How do you decouple the two?Move the redis indicator out of the readiness group (management.endpoint.health.group.readiness.include lists only critical deps), or disable it with management.health.redis.enabled=false. Aggregation takes the worst status, so any DOWN indicator inside readiness causes a 503; excluding it removes that coupling.
- What component decides the overall status is DOWN, and what maps DOWN to 503?StatusAggregator (default SimpleStatusAggregator) picks the worst status among contributors; HttpCodeStatusMapper then maps DOWN/OUT_OF_SERVICE to HTTP 503 (UP/UNKNOWN to 200). Both are customizable via management.endpoint.health.status.* properties.
saying these in an interview costs you the question
- Thinking overall status is UP unless ALL indicators fail (it's the worst single status)
- Believing DOWN returns HTTP 500 (it's 503 by default)
- Assuming you can't stop a non-critical dependency from ejecting the instance (health groups solve exactly this)