An encrypted internet tunnel carries synchronous replication that stalls every afternoon while its average round-trip time looks fine — which property of the path is failing?
answer
- the mean is the wrong statistic
- lock-step waits for the slowest
- look at the tail, per direction
- your own uplink queues first
- a circuit narrows, never zeroes
basics
~20 sLatency consistency, not average latency. A lock-step protocol waits for each round trip, so its throughput is set by the slow tail, and a tunnel over a shared internet path has a tail that widens whenever anything upstream or on your own uplink is busy.
solid answer
~50 sThe mean is the wrong statistic for this workload. Synchronous replication is lock-step: every commit waits for a round trip, so the number of commits per second is governed by the *slowest* round trips, not the typical one. A path can hold a perfectly good average while its high percentiles double in the afternoon, and that is exactly what a shared internet path does when your own uplink is carrying a bulk copy, when upstream links congest at peak, or when the path changes underneath a live session. So measure the tail — high percentiles per direction, plus loss, because one lost packet costs a lock-step protocol a full timeout. Then fix what you own: get the bulk traffic off that uplink or shape it. A dedicated circuit is the structural answer, because it narrows the band rather than removing variance entirely.
go deeper
Know that latency has a distribution, not a single value, and that a workload waiting for each reply cares about the slow end of it. A good average can hide a bad tail.
Explain why a lock-step protocol turns tail latency into a throughput limit, and name the causes in order: your own uplink queueing, upstream congestion, a path change mid-session, and loss paying a full timeout.
Demonstrate the diagnosis: percentiles per direction at fine resolution, loss beside latency, the batch schedule overlaid, replication lag from the application's side. Then say which fix you own today and which needs a circuit.
Decide whether the workload should be crossing a shared path at all. Weigh re-shaping the application's consistency requirement against the cost and lead time of a dedicated path, and set the rule for which traffic classes are allowed on tunnels.
## Why the average hides the failure Synchronous replication is a lock-step protocol: the writer sends, waits for an acknowledgement, and only then proceeds. Its maximum rate is therefore one round trip per operation, and the operations that matter are the slow ones. If ninety-nine round trips in a hundred take a few milliseconds and one takes several hundred, the mean barely moves — but every transaction that lands on that one waits, and if the writer has a bounded queue behind it, the whole queue waits with it. A monitoring page showing a flat average round-trip time is therefore perfectly consistent with a ledger that stalls every afternoon. The property that is failing has a name: **latency consistency**, the width of the distribution rather than its centre. It is the thing a dedicated circuit is actually bought for, and the thing a tunnel over the public internet is structurally worst at. ## Where the variation comes from - **Your own uplink.** The tunnel shares the building's internet connection with everything else that leaves it. A nightly-turned-afternoon bulk copy fills the queue, and the latency-sensitive traffic waits behind it. This is the most common cause and the only one you fully own. - **Upstream congestion.** The internet path crosses networks nobody in the conversation operates, and their busy hour may be your afternoon. - **Path changes.** The route between the two ends can change mid-session. The new path may be fine, better, or noticeably longer, and nothing announces it to your application. - **The tunnel endpoint itself.** Encrypting and decrypting costs work on a device that is also doing other things; when it saturates, latency climbs before throughput visibly falls. - **Loss amplifying everything.** A lock-step protocol pays a full retransmission timeout for a single lost packet, so a small loss rate produces a very large tail. ## What to measure, in order 1. **Percentiles, not the mean** — and at a resolution fine enough to see an afternoon. A daily average is useless here. 2. **Each direction separately.** The path out and the path back can differ, and the congested one may not be the one you assumed. 3. **Loss alongside latency.** Latency percentiles without a loss figure cannot distinguish queueing from retransmission. 4. **The batch schedule next to the latency chart.** If the tail rises when a known job starts, you have your cause and you do not need a circuit to prove it. 5. **Replication lag as the application sees it**, so you can tell whether the path or the store is the constraint. ## What each path can honestly promise | Property | Encrypted tunnel over the internet | Dedicated private circuit | |---|---|---| | Typical round trip | Usually fine | Usually fine, and similar | | High percentiles | Wide, and time-of-day dependent | Narrow, and stable across the day | | Path stability | Can change mid-session | Fixed for the term | | Who else shares it | Everyone, on every hop | Nobody, on the circuit itself | | What you can do about it | Shape your own traffic, add tunnels | Size the port; the band is already narrow | Note the direction of the claim carefully: a circuit **narrows the band**, it does not make variation zero. Queueing still happens on the devices at each end, and a saturated circuit queues exactly like anything else. ## Fixes short of ordering a circuit - Move the bulk copy off that uplink, or shape it so it cannot fill the queue in front of the replication traffic. - Separate the flows across distinct tunnels so the noisy one and the sensitive one are not sharing a single termination device — remembering that this helps *different* flows, not one heavy flow. - Reduce how often the application pays a round trip: batch smaller writes, or take the deliberate decision to replicate asynchronously and accept a defined amount of loss instead of a stall. - If the traffic genuinely cannot tolerate a variable path and cannot be made less chatty, the tail is telling you the workload needs a dedicated circuit — and, until it arrives, that this workload should not be the one running across the tunnel.
- Why does a small packet-loss rate hurt this workload far more than the percentage suggests?Because a lock-step protocol cannot proceed past the operation whose packet was lost. It waits for a retransmission timeout, which is orders of magnitude longer than a normal round trip, and every waiting transaction behind it inherits that delay. A loss rate that would be invisible to a bulk transfer therefore shows up as a visible stall in replication.
- Would a second tunnel across the same uplink fix the afternoon stall?Usually not. If the cause is queueing on the shared uplink, both tunnels sit behind the same queue and both inherit the same tail. A second tunnel helps when the constraint is the termination device or when you can put the noisy flow and the sensitive flow on genuinely different paths — not when the bottleneck is upstream of both.
saying these in an interview costs you the question
- Reads a flat average round-trip time as proof the path is healthy
- Assumes only the provider can do anything about variable latency
- Thinks a second tunnel over the same uplink removes the queueing
- Describes a dedicated circuit as having no latency variation at all
- Blames the data store before checking what shares the uplink