Design a production-grade defense-in-depth strategy for locking down Actuator across exposure, authorization, data masking, and network.
answer
- four orthogonal controls: exposure / authz / masking / network
- every layer fails closed & independent
- never exposure=*; shutdown off
- integration-test each layer (401 / masked / full)
- internal port != skip auth; least-privilege actuator role
basics
~20 sCombine independent layers: expose only needed endpoints, guard them with a SecurityFilterChain requiring an admin role via EndpointRequest.toAnyEndpoint(), keep /env and /configprops values masked (show-values NEVER or WHEN_AUTHORIZED), disable dangerous ones like /shutdown, and isolate the management port to an internal network.
solid answer
~40 sI treat actuator security as four orthogonal controls so any single misconfiguration isn't catastrophic. (1) Minimize attack surface: management.endpoints.web.exposure.include lists only what's needed (health, info, metrics, prometheus), never '*', and /shutdown stays disabled. (2) Authorize: a dedicated ordered SecurityFilterChain scoped by EndpointRequest.toAnyEndpoint().excluding("health") requires ROLE_ACTUATOR_ADMIN with HTTP Basic for machine callers; health stays permitAll for probes with show-details=when-authorized. (3) Data minimization: show-values=NEVER (or WHEN_AUTHORIZED + roles) on env/configprops plus custom SanitizingFunction beans. (4) Network: management.server.port bound to an internal interface, scraped only from the monitoring subnet. I verify each layer with tests, and treat the internal port as isolation, not a reason to drop auth. I also watch that exposure, security, and masking are independent — all three must be right.
code
java · 17 lines@Bean
@Order(SecurityProperties.BASIC_AUTH_ORDER - 1)
SecurityFilterChain actuatorChain(HttpSecurity http) throws Exception {
http.securityMatcher(EndpointRequest.toAnyEndpoint())
.authorizeHttpRequests(a -> a
.requestMatchers(EndpointRequest.to("health", "info")).permitAll()
.anyRequest().hasRole("ACTUATOR_ADMIN"))
.httpBasic(Customizer.withDefaults())
.csrf(CsrfConfigurer::disable);
return http.build();
}
// Regression test that keeps the guard honest
@Test
void anonymousCannotReadEnv() throws Exception {
mockMvc.perform(get("/actuator/env")).andExpect(status().isUnauthorized());
}go deeper
Not expected to design this; should recognize the individual pieces.
Should implement exposure + filter chain + masking correctly.
Should combine all layers and test them, understanding each is independent.
Should architect fail-closed independent controls, enforce a baseline across services, and reason about trade-offs and regression prevention.
**Principle: independent, overlapping controls.** The recurring actuator breach is a single point of failure — someone sets `exposure.include=*`, or a filter chain regresses, or an internal port is assumed 'safe'. A principal-level design makes each control fail closed and independent, so one mistake doesn't expose secrets. **Layer 1 — Attack surface minimization.** - `management.endpoints.web.exposure.include` = explicit allowlist (`health,info,metrics,prometheus`). Never `*`. Consider `exclude` as a backstop. - `management.endpoint.shutdown.enabled=false` (default) — never enable over HTTP. - Disable endpoints you don't operate (`heapdump`, `threaddump`) unless needed; each is a distinct data-leak vector. **Layer 2 — Authentication & authorization.** - Dedicated `SecurityFilterChain` with `@Order` above the app chain, `securityMatcher(EndpointRequest.toAnyEndpoint())`. - `permitAll()` on `to("health","info")` for K8s probes; `hasRole("ACTUATOR_ADMIN")` for the rest. - HTTP Basic or mTLS for machine callers; CSRF disabled on the stateless actuator chain. - `management.endpoint.health.show-details=when-authorized` and `health.roles` so anonymous probes see only UP/DOWN. **Layer 3 — Data minimization (masking).** - `management.endpoint.env.show-values` / `configprops.show-values` = `NEVER` by default; `WHEN_AUTHORIZED` + `env.roles` only if operators truly need values. - Additive `SanitizingFunction` beans for custom secret keys; never blow away defaults via `keys-to-sanitize`. - Rationale: even an authenticated admin shouldn't casually see raw secrets (screenshots, logs, over-broad roles). **Layer 4 — Network isolation.** - `management.server.port` on a separate listener, `management.server.address` bound to a private/loopback interface; firewall/NetworkPolicy so only the monitoring subnet scrapes it. - Independent TLS via `management.server.ssl.*`. - Explicitly test that auth still applies on the management context (separate context caveat). **Cross-cutting concerns:** - **Independence check:** exposure, security, and masking are separate — verify all three in an integration test (`@SpringBootTest` hitting `/actuator/env` anonymously expects 401; authenticated-but-under-privileged expects masked; correct role expects data). Regression tests prevent the classic `exposure=*` slip. - **Observability of the guard itself:** audit-log access to sensitive endpoints; alert on unexpected `/env` or `/heapdump` hits. - **Least privilege on roles:** actuator admin role distinct from application admin. - **Supply chain:** keep Boot patched — actuator/data-leak CVEs recur. - **Config drift:** enforce via a shared parent config / platform baseline so every service inherits the hardened defaults rather than re-deriving them. **Trade-offs:** more layers = more operational surface (probe repointing, cert management, role plumbing). For low-risk internal-only services you might collapse layer 4. The principle stands: never depend on a single control, and make the default deny.
- Why keep four separate controls instead of relying on the SecurityFilterChain alone?Because a single control is a single point of failure. If the chain regresses or an ordering bug lets requests through, minimal exposure, value masking, and network isolation each still limit the blast radius. Independent overlapping controls make the system fail closed.
- How do you prevent the classic 'someone set exposure.include=*' regression?Enforce a hardened baseline via a shared parent config/platform, and add integration tests that assert sensitive endpoints require auth and return masked values. A test hitting /actuator/env anonymously expecting 401 catches the regression in CI.
- When would you deliberately drop the separate management port?For a low-risk, purely internal service where the operational cost (probe repointing, separate TLS, NetworkPolicy) outweighs the marginal isolation, relying on exposure minimization + Spring Security + masking as the remaining independent layers.
saying these in an interview costs you the question
- Relying on a single control (only the filter chain, or only the internal port)
- exposure.include=* 'because it's behind auth anyway'
- No regression test, so exposure/masking drift silently
- Reusing the application admin role for actuator (no least privilege)
- Enabling /shutdown or /heapdump without a clear need