A NETCONF client completes SSH login and the hello exchange, sends its first <rpc>, then waits forever for a reply; what framing mismatch explains it?
answer
- the reader never sees an end
- decided from one hello, not two
- which side is still waiting
- look at the bytes after the hellos
basics
~20 sFraming must follow both hellos: chunked only if both list :base:1.1, otherwise ]]>]]>. A client that switches to chunked on the server's hello alone, or never switches, sends or expects a message end the other side never produces, so one reader waits.
solid answer
~50 sIn NETCONF over SSH the framing after the hellos is fixed by **both** hellos: chunked if both advertise `:base:1.1`, `]]>]]>` otherwise. A framing hang means one reader is waiting for an end of message that never arrives. Classic cases: the client's own hello listed only `base:1.0` but it switched to chunked because the server listed `base:1.1`, so the server waits for `]]>]]>`; a base:1.0 client forgot the trailing `]]>]]>`; or the client's reader expects the wrong format and never recognises a reply that did arrive. The reverse mismatch, a client still sending `]]>]]>` in a chunked session, is a decoding error on which RFC 6242 makes the server close the channel, so it shows up as a dropped session rather than a hang. Diagnose at an endpoint, since SSH encrypts the wire: compare the two hellos, then look at the first bytes each side wrote after them.
go deeper
Recall that NETCONF over SSH needs framing, and that the two sides must agree on it after the hellos or one of them never sees a message end.
Explain the selection rule from both hellos and why the hello itself always ends with ]]>]]>, then predict which side waits in each mismatch.
Diagnose at the endpoints: compare both hellos, read the raw bytes written and received, use the in-rpcs and in-bad-rpcs counters, and tell a hang from a channel closed on a decoding error.
Turn the failure into client design rules: one framing decision from both hellos, reply timeouts on every RPC, and logging of the negotiated capabilities on each session.
## Why a framing mistake looks like a hang NETCONF over SSH (RFC 6242) carries XML documents on a byte stream, and the receiver finds where each message ends only through **framing**. In **end-of-message (EOM)** framing a message ends at `]]>]]>`; in **chunked** framing it ends at the `\n##\n` that follows one or more length-prefixed chunks (`\n#<size>\n` plus that many octets). If a sender writes a format the receiver is not reading, the receiver either never finds an end, and keeps reading, or finds bytes it cannot decode. The first is a **hang**. The second is a **decoding error**, on which RFC 6242 requires the peer to close the SSH channel, which looks like a **dropped session**. ## The rule both sides must apply 1. Each peer sends its `<hello>` and ends it with `]]>]]>`, always. 2. Each peer reads the other's hello. 3. If **both** hellos advertise `:base:1.1`, every later message, in both directions, uses chunked framing. 4. Otherwise every later message uses `]]>]]>`. The decision needs **both** lists. Client code that looks only at the server's hello, or that hard-codes one framing, is the usual root cause. ## The cases, side by side | What happened | Framing the session should use | What the client did | Symptom | |---|---|---|---| | Client hello listed only `base:1.0`; server listed both | EOM | Switched to chunked because the server offered 1.1 | Server waits for `]]>]]>`: **hang** | | Client listed `base:1.1`; server, an RFC 4741-era implementation, listed only `base:1.0` | EOM | Sent chunked anyway | Server waits for `]]>]]>`: **hang** | | Both used EOM | EOM | Sent the `<rpc>` without the trailing `]]>]]>`, or never flushed it | Server waits: **hang** | | Both listed `base:1.1` | chunked | Server replied chunked, but the client's reader looks for `]]>]]>` | Reply arrived, client never recognises its end: **hang** | | Client sent its hello chunked | EOM for the hello | Framed the hello in the new format | Server never finds the end of the hello: **hang** before any RPC | | Both listed `base:1.1` | chunked | Kept writing `]]>]]>` | Server reads `<` where `\n#` must be: **decoding error, channel closed** | The last row is the instructive one. RFC 6242 says a peer **MUST** terminate the session by closing the channel on any decoding error. A server that instead keeps buffering, looking for a chunk header, turns that case into a hang too, but that is the implementation's behaviour, not the RFC's. ## How to diagnose it SSH encrypts the channel, so capturing packets between the hosts shows only encrypted SSH traffic. The evidence lives at the endpoints: - **Log both hellos** at the client and check whether each lists `urn:ietf:params:netconf:base:1.1`. Remember that the element namespace `urn:ietf:params:xml:ns:netconf:base:1.0` says nothing about the version. - **Log the raw bytes** the client writes after the hellos. Chunked framing starts with a line feed and `#`; EOM framing starts with the XML and ends with `]]>]]>`. - **Log the raw bytes the client receives.** If a reply arrived and the client is still waiting, the client's reader is the side using the wrong framing. - **Read the server's counters** in the RFC 6022 monitoring data: `in-rpcs` counts correct RPCs received and `in-bad-rpcs` the rejected ones, so neither moving while the client waits says the server never delimited the request at all. `in-bad-hellos` rising points one step earlier, at the hello. ## Fixing it properly - Decide framing from **both** hellos, in one place, and use it for both writing and reading. - Always end the **hello** with `]]>]]>`, whatever the client supports. - Flush after each message; a framed message sitting in a client buffer is as invisible to the server as a missing marker. - Set a client-side **reply timeout**, so a framing bug fails loudly instead of hanging a job. ## What this is not - Not an authentication problem: SSH login and the hellos already succeeded. - Not an XML error: a malformed document that the server *can* delimit earns a `malformed-message` error reply, which is the opposite of silence. - Not something the server reports: a server that cannot find the end of a message has nothing to reply to.
- Why does a client still sending ]]>]]> in a chunked session usually produce a closed channel rather than a hang?In a chunked session the server expects each message to begin with a line feed and #. Seeing < instead is a decoding error, and RFC 6242 says the peer MUST then close the SSH channel. Only a server that ignores that rule and keeps buffering turns this case into a hang.
- How do you tell whether the client or the server is the side that is waiting?Log the raw bytes the client receives. If reply bytes arrived after the RPC and the client is still blocked, its reader is using the wrong framing. If nothing arrived and neither the server's in-rpcs nor in-bad-rpcs counter has moved, the server never found the end of the request.
- Why can't a client simply pick chunked framing whenever the server's hello lists base:1.1?Because the rule needs both hellos. If the client's own hello did not list base:1.1, the server will keep using ]]>]]> and wait for that marker, while the client sends chunks it will never recognise as complete.
saying these in an interview costs you the question
- If the server lists base:1.1, the client should switch to chunked framing.
- A framing mistake always comes back as an rpc-error from the server.
- Capturing traffic on port 830 between the hosts shows the bad framing.
- The base:1.0 namespace on the hello means the session uses old framing.
- Framing is chosen per message, so a stuck RPC can just be resent differently.