After enrolling a service in a sidecar-based service mesh, the application's first outbound calls at pod startup fail with connection refused or an immediate 503, and then everything works normally. What is happening, and how do you fix it?
answer
- redirection is installed before the proxy serves
- only the first few seconds fail
- the app log shows nothing at all
- start order, not timeouts
- the same race runs backwards at shutdown
basics
~20 sTraffic redirection is in place before the sidecar proxy is serving, so the application's earliest calls are steered into a proxy that has no listener or no configuration yet and are refused. Fix it by making the proxy start and become ready before the application container runs.
solid answer
~50 sEnrolling a workload installs redirection rules in the pod's network namespace so that all outbound traffic goes through the local proxy. Those rules apply from the moment the pod's network is set up, but the proxy still has to start, fetch its configuration from the control plane and open its listeners. If the application container is racing it — a client that connects to a database or a dependency in its constructor is the classic case — its packets are redirected to a socket nobody is listening on, and it sees a connection reset, or a proxy-generated 503 that never appears in the application's log because the request never reached the app. The durable fix is ordering: on Kubernetes, run the proxy as a native sidecar container so it starts and passes its startup probe before application containers begin. Meshes also offer a setting that holds the application until the proxy reports ready, and application-side retry with backoff on startup dependencies is worth having regardless.
code
yaml · 20 linesapiVersion: v1
kind: Pod
metadata:
name: checkout
spec:
initContainers:
- name: mesh-proxy
image: example/mesh-proxy:1.0
restartPolicy: Always
startupProbe:
httpGet:
path: /ready
port: 15021
periodSeconds: 1
failureThreshold: 30
containers:
- name: app
image: example/checkout:2.4
ports:
- containerPort: 8080go deeper
Know that the sidecar is a separate container and that a pod's containers do not automatically start in a useful order.
Explain why redirection being active before the proxy is serving produces a connection refused, and why the failing requests never reach the application's log.
Diagnose it from the timestamp pattern and the missing callee logs, then fix the ordering at both ends — proxy ready before the app, proxy stopped after it.
Make the correct injection mode a platform default so no team re-discovers this, and decide whether a data-plane model without per-pod lifecycle coupling is worth adopting for that reason alone.
## Why the failure is invisible in the application log The defining symptom is asymmetry: the caller records failures, the callee records nothing, and the whole thing clears up after a few seconds. That pattern says a component *between* the two produced the response. In a mesh, the component is the local proxy. Enrolment works by making the proxy unavoidable. Something — an init container, or a CNI-level component — programs redirection rules in the pod's network namespace so that outbound connections from any process in the pod are sent to the proxy's outbound port, and inbound connections are sent to its inbound port. The application is not modified and does not know. That is the whole value proposition, and it is also the trap: **redirection is in force before the proxy is able to serve.** ## The startup race, step by step 1. The pod's network namespace is created and the redirection rules are installed. 2. The proxy container starts. It must connect to the control plane, receive listeners, routes, upstream clusters and endpoints, and its identity certificate, then open its listeners. This takes anywhere from a fraction of a second to several seconds on a loaded node or a busy control plane. 3. Meanwhile the application container starts. If it opens a connection during initialisation — warming a connection pool, reading configuration from a remote store, registering with a discovery service, running a migration — that connection is redirected to a proxy port that is not yet accepting. 4. The application sees connection refused or a reset. If the proxy is up but has no route yet, it can accept and then return a 503 of its own instead. 5. Frameworks that fail fast on startup dependencies then crash the container, Kubernetes restarts it, and by the second attempt the proxy is ready — which is why the incident looks intermittent and "fixes itself". The mirror image happens at shutdown. On pod termination every container gets SIGTERM at roughly the same time. If the proxy exits while the application is still finishing in-flight work, those last outbound calls fail. And for a Job or CronJob the opposite hurts: the application exits successfully but the proxy keeps running, so the pod never reaches Completed and the Job hangs forever. ## Fixing it properly **Order the startup.** Kubernetes native sidecar containers — an entry in `initContainers` with `restartPolicy: Always` — start before the regular containers, keep running alongside them, and are terminated only after the last regular container exits. That single mechanism fixes both ends: the proxy is up and past its startup probe before the app runs, and Job pods complete because the sidecar is stopped once the workload finishes. It is on by default from Kubernetes 1.29; meshes adopted it as the recommended injection mode. ```yaml spec: initContainers: - name: proxy image: mesh-proxy:1.0 restartPolicy: Always # native sidecar: starts first, stops last startupProbe: httpGet: { path: /ready, port: 15021 } containers: - name: app image: app:1.0 # only starts once the proxy is ready ``` **Use the mesh's hold-the-app setting.** Where native sidecars are not available, most meshes expose an option that blocks the application container until the proxy signals readiness. It solves startup but not the Job-completion case, which needs a separate mechanism to tell the proxy to exit. **Make the application tolerant.** Retrying startup dependencies with backoff instead of crashing on the first failure is good practice independent of the mesh — the same code protects you against a dependency that is merely slow to come up. Deferring connection pool warm-up until after the app is serving also removes the race. **Drain in the right order at shutdown.** The application should stop accepting new work, finish in-flight requests, and only then let the proxy exit; native sidecar ordering gives you that for free. **Or remove the per-pod proxy.** A node-level data plane has no per-pod startup race at all, because the shared proxy is already running before your pod is scheduled. That is one of the strongest practical arguments for that model. ## How to confirm the diagnosis quickly Compare timestamps: the failures cluster in the first seconds of the pod's life and never recur. Check the proxy's own access log or readiness endpoint on the calling pod for the corresponding entries. Test the hypothesis by excluding the port from redirection, or by starting the application with a deliberate delay — if the failures disappear, the race is confirmed. What you should not do is silence it by widening the application's timeouts; that hides the ordering bug and leaves the shutdown half of the problem in place.
- Why do the team's batch Job pods stay Running forever after the same service was enrolled in the mesh?The application container exits when the job finishes, but the sidecar proxy is a long-running process that never exits on its own, so the pod never reaches Completed. Native sidecar containers fix it, because Kubernetes terminates them once the last regular container has exited. Otherwise the job has to signal the proxy to shut down explicitly.
- Why is widening the application's startup timeout a poor fix?It hides the ordering bug rather than removing it. A slower control plane or a loaded node re-creates the race at a longer delay, the failure comes back as a rare flake, and it does nothing for the shutdown side, where the proxy exiting first still kills in-flight outbound calls. Fix the ordering, not the patience.
- How would you confirm within minutes that the proxy startup race is the cause and not the dependency itself?Check that the failures are confined to the first seconds of pod life and never recur, and that the callee logged nothing. Then look at the calling pod's proxy access log for those requests. A decisive test is delaying the application's first call, or temporarily excluding that port from redirection — if the errors vanish, the race is confirmed.
saying these in an interview costs you the question
- Blames the downstream service that logged nothing
- Fixes it by increasing the application's connect timeout
- Assumes containers in a pod start in the listed order
- Thinks redirection begins only once the proxy is ready
- Ignores the mirrored failure during pod shutdown