How do you distinguish an observable side effect from a benign one when deciding whether a method still counts as a query under CQS?
answer
- benign = no observer can detect it
- observable = later call / thread / monitor / external system sees it
- memoize + lazy hash = benign; audit + counter + cursor = observable
- concurrency, externalization, resource use flip benign → observable
- returning a live mutable reference leaks mutation through a query
basics
~20 sA side effect is benign if no caller can tell it happened — like filling an internal cache. It is observable if any later call, another thread, or an external system can detect it. Only benign effects still count as a query.
solid answer
~50 sThe test is **detectability**, not intent. Ask: after calling this method, can any subsequent query, another thread, a monitoring system, or an external observer distinguish a program where it ran from one where it didn't? If not, the effect is benign and the method is still "observably pure" — a legitimate query under CQS. Classic benign cases: memoization, lazy initialization of a derived field, populating a read-through cache, internal statistics used only for heuristics. Classic observable ones: writing audit rows, incrementing a business counter, advancing a cursor, consuming a token or rate-limit quota, sending anything over the network, mutating a returned collection. Three things flip benign into observable: **concurrency** (an unsynchronized lazy write is a data race, and it makes results non-deterministic under eviction), **externalization** (a metric or log line someone alerts on), and **resource consumption** (billed API calls, connection-pool leases). So classify from the observer's point of view, and document the exception where you take it.
go deeper
Say that if nobody can tell the method did anything extra — like remembering an answer to be faster — it still counts as a query; if something outside can notice, it's a command.
Give the detectability test and the two lists: memoization and lazy caching on the benign side, audit writes, counters, and cursor advances on the observable side.
Add the three forces that flip the classification (concurrency, externalization, resource consumption), note that returning a mutable reference leaks mutation, and tie the classification to retry safety, prefetching, and lock-free reads.
Turn it into policy: strict by default, documented exceptions with an invalidation condition, mandatory re-audit when a component goes concurrent or distributed, and awareness that HTTP intermediaries will physically repeat anything you expose as a safe method.
## The question behind the question **Command-Query Separation (CQS)** says a query returns a value and changes no state. Taken literally, almost no real query qualifies: reading a value warms a CPU cache, allocates memory, may compile a query plan, advances a profiler counter. Clearly that reading is too strict. The usable version says a query must not change **observable** state, which forces you to define "observable". ## The definition An effect is **observable** if there exists any observer — a later call in the same program, a concurrent thread, another process, a monitoring or billing system, a human reading a log — that can distinguish an execution where the effect happened from one where it did not. An effect is **benign** (the literature also says *observably pure*, or in some type systems *effectively pure*) if no such observer exists. The method then behaves, from outside, exactly like a pure query, and every CQS benefit still applies: safe to repeat, safe to log, safe to skip, safe to reorder. Note the framing: this is not about whether the effect is *small* or *well-intentioned*, but about whether anyone can *see* it. ## Canonical benign effects 1. **Memoization.** `fibonacci(n)` storing computed values in a private table. Repeat calls return the same answer, faster. Nothing else changes. 2. **Lazy initialization of a derived field.** `hashCode()` computing once and caching (the classic pattern in immutable string types). The value is a deterministic function of already-fixed state. 3. **Read-through cache population.** `repository.findById(id)` filling a local cache on a miss — provided the cache is authoritative-consistent and invisible in results. 4. **Internal heuristics.** A collection that records access patterns to choose a growth strategy; a query planner recording statistics. 5. **Debug counters that nobody consumes.** ## Canonical observable effects (still commands, however they're named) 1. **Audit trail writes** — another query (`findAuditEntries`) can see them; a compliance system depends on them. 2. **Business counters** — a view counter, download counter, or "last accessed" timestamp that is part of the domain. 3. **Cursor / position advance** — `Iterator.next()`, `Reader.read()`. The next call returns something different, which is the definition of observable. 4. **Token consumption** — a rate-limit budget, a nonce, a one-time password check that burns the code, a paginated API cursor. 5. **Network sends** — publishing an event, calling a paid third-party endpoint, sending an email. 6. **Returning a live mutable reference** — `getItems()` handing back the internal list. The call didn't mutate, but it *transferred the capability to mutate*, so callers can now change state through what looks like a query. Return a copy or an immutable view. ## Three forces that turn benign into observable ### 1. Concurrency The single biggest trap. A "harmless" lazy write from a read path is a **data race** if unsynchronized: two threads may compute and publish different objects, or one may see a partially constructed value. Even with correct synchronization, a shared cache makes results **timing-dependent** — under eviction, memory pressure, or a cache stampede, behaviour that was invisible single-threaded becomes visible. Worse, if reads now need a lock to protect the cache, you have destroyed the main CQS dividend (lock-free parallel reads). ### 2. Externalization The moment an effect leaves the process, someone can observe it. A metric increment on a read path becomes observable as soon as a dashboard or alert depends on it; a log line becomes observable to an auditor. It's the same write; what changed is that an observer appeared. ### 3. Resource consumption Billed API calls, connection-pool leases, file handles, disk space in a cache directory, memory that can push the process into GC pressure or OOM. Effects that consume a bounded, shared resource are observable through failure modes even if never directly read. ## Why this classification actually matters It drives four concrete decisions: - **Retry safety.** A lost response to a truly benign read can be resent. A read that consumed a token or wrote an audit row cannot be retried blindly. - **Caching and prefetch.** HTTP proxies, browsers, and clients speculatively repeat `GET` requests. If your `GET` handler has an observable effect, prefetching corrupts data — a real historical incident class (crawlers following "delete" links exposed as GETs). - **Debuggability.** A benign-effect query can be called from an assertion or a debugger watch expression. An observable-effect one cannot, and you get Heisenbugs. - **Concurrency design.** Only truly read-only paths can be served from a replica, a snapshot, or without a lock. ## Practical policy 1. **Default to strict.** Treat any write as observable unless you have argued otherwise. 2. **Make the argument explicit.** Document the exception at the method: what is written, why no observer sees it, and what would invalidate that. 3. **Re-audit on concurrency changes.** "Benign single-threaded" is not "benign". 4. **Name honestly when in doubt.** `findAndTouch`, `getOrCreate`, `consumeToken` are better than a getter that surprises. 5. **Don't leak mutable internals** from queries; the guarantee has to hold transitively through returned references.
- A read-only endpoint increments a per-user quota counter. Benign or observable?Observable. The quota is externally visible state — the user can be throttled, another query can report remaining quota, and the call is no longer safe to retry or prefetch. It should be treated as a command-plus-query mix and named or documented as such.
- Why can a lazily computed hash code be considered benign while a lazily loaded association often cannot?A hash code is a deterministic function of already-immutable state, computed identically every time, so no observer can distinguish cached from recomputed. A lazily loaded association hits the database on read: it can fail, block, throw outside a session, take unbounded time, and returns data that may differ from what a later load would return — all observable, and famously fragile under concurrency and detached-entity access.
A librarian who reshelves a book more efficiently after you ask for it has had a benign effect — you'd never know. A librarian who stamps your card and decrements your remaining loans every time you ask whether a book exists has had an observable one: the act of asking cost you something, and asking twice is not the same as asking once.