A new broker capability is enabled cluster-wide, yet one application never gets it and logs no error - why?
answer
- on is not the same as reachable
- delivered per connection
- the agreement has no field for it
- nothing failed, so nothing logged
- the pinned library is the suspect
basics
~20 sThat application's connections agreed an older wire version, which cannot express the capability. The server answers in the older shape and simply withholds what that shape has no room for. Nothing failed, so nothing is logged.
solid answer
~50 sEnabling a capability on the cluster is only one of the gates it has to pass. The other is the wire version that each connection agreed at its connect step: if that agreement predates the capability, the request and response shapes in use have nowhere to carry it, so the server answers the older way and the feature is absent for that connection. This is not an error condition - the older exchange is valid and completed successfully - so there is nothing for the client or the server to report. The usual cause is an old client library that cannot advertise anything newer, and the usual tell is that the capability works for every other application against the same cluster. Chasing the cluster configuration here wastes the outage; the constraint sits on the application's side of the connection.
go deeper
Recall that turning a feature on for the cluster does not mean every connected application can use it, and that an application can be missing a feature with nothing broken and nothing logged.
Explain the gates: the nodes implement it, the cluster permits it, the connection's agreed wire version can express it, and the library actually asks for it. Say which one produces silence.
Demonstrate the diagnosis order: compare applications before re-reading cluster settings, check client library versions and when connections last re-established, and separate what the agreement can carry from what the library requests.
Treat it as an estate signal. A capability that lands for nine teams and not the tenth is telling you the client inventory is unmanaged, and the same team will block the next one too.
## The symptom A capability is turned on for the whole cluster. Most applications pick it up. One does not. Its connections are healthy, its throughput is normal, its error rate is flat, and no log line anywhere mentions the capability. The operator re-checks the cluster setting, finds it correct, and re-checks it again. This is the most confusing shape in client compatibility, because **the absence is not a failure**. Every request that application sends is answered successfully. What it never sees is an answer that carries something the older conversation has no field for. ## The gates a capability has to pass For an application to actually use a newer capability, several independent conditions all have to hold. Missing any one of them produces the same silent nothing: 1. **The broker nodes run a release that implements it.** Assumed here: the roll finished. 2. **The cluster's own agreement permits it.** The agreed internal version that fixes how nodes speak to each other is normally held below the running binaries until every node is new, and a capability can stay inert until it is raised. Again, assumed here: done. 3. **The connection's agreed wire version can express it.** This is the gate that fails in this scenario. The agreement was settled at the connect step from what both peers advertised, and it predates the capability. 4. **The client library knows to ask for it.** Even inside a capable agreement, a library that has no concept of the feature will never request it, and its own built-in default - which is whatever that library shipped with, not the cluster-wide default - governs its behaviour. Note how different gate 4 is from gate 3, because candidates merge them. Gate 3 is about what the connection *can carry*. Gate 4 is about what the library *chooses to ask for*. ## Why silence is the correct behaviour A server serving an older agreement is doing exactly what it promised. It received a well-formed request in a shape it still supports, and it answered in that shape. There is no failure to attribute and no side that would naturally raise one: the client did not ask for something and get refused; it never had a way to ask. The loud failure exists too, and it is worth contrasting: a client that *requires* something the server cannot offer fails at the connect step, at startup, with an unambiguous message. Compatibility problems therefore split cleanly into two very different experiences: | Direction | Where it shows | How it is noticed | |---|---|---| | Client older than the capability | Nowhere - the older exchange succeeds | Only by comparison with other applications, or by auditing agreements | | Client requires more than the server offers | At the connect step | Immediately, as a refused connection at startup | ## What varies by platform Do not carry one platform's shape into the answer. Where the contract is a set of negotiated per-request revisions, the absence is exactly as described. Where a platform versions a remote endpoint instead, an old caller keeps calling the old endpoint and gets its old behaviour until that endpoint is retired. Where the surface is a rented cluster on a hosted tier, the agreement may not be visible to you at all, and the provider's supported-client statement is the only handle you have. Some platforms also expose, per connection, what was agreed - and where that exists, this whole investigation is one query instead of a week. ## How to diagnose it - **Compare applications, not settings.** If one application lacks the capability and its neighbours have it, the cluster is not the variable. - **Read the client library version that application ships,** and the span of releases the current server release serves. An application whose dependency was pinned years ago is the default suspect. - **Check when its connections last re-established.** A library upgrade that has not reconnected has not taken effect. - **Separate can-express from will-ask.** If the library is new enough for the agreement but still does nothing, the gap is that the library has no support for the feature, which no cluster change will fix. - **Record what you found as an inventory item,** because the same application will be the one that blocks the next capability too. ## The lesson to carry The cluster-wide switch is a necessary condition, never a sufficient one. A capability is delivered per connection, through an agreement neither side can exceed, and the quietest outcome in this whole area - a feature that is on everywhere and present nowhere for one team - is produced by an old library and no error at all.
- How would you prove the client library is the constraint rather than the cluster?Show the capability working for another application on the same cluster, then compare client library versions. If the platform reports what each connection agreed, read that directly. Confirming the cluster setting twice proves nothing, because the setting is not the gate that failed.
- The library is new enough, the agreement supports the capability, and the application still does not use it. What now?Then the gap is not what the connection can carry but what the library asks for. A library whose own built-in default leaves the feature off, or which has no code for it, will never request it inside a perfectly capable agreement. That is an application change, not a cluster change.
saying these in an interview costs you the question
- Assumes a cluster-wide switch reaches every application at once.
- Expects an error whenever a connection cannot express a capability.
- Blames the cluster configuration when the client library is the constraint.
- Thinks restarting the broker nodes will make the feature appear.
- Confuses the client library's own default with the cluster-wide default.
- Believes monitoring would have caught it, since dashboards stay green.