A busy Linux NAT gateway starts refusing new connections while existing ones keep working, and the kernel log shows "nf_conntrack: table full, dropping packet". What is filling up, what sets its ceiling, and what are your realistic options?
answer
- old connections fine, new ones fail
- a fixed-size kernel table, sized at load
- entries retire on timers, not on close
- the default established timeout is days
- some traffic need not be tracked at all
basics
~20 sThe kernel's connection-tracking table is full, so packets that would create a new flow entry are dropped while tracked flows continue. The ceiling is net.netfilter.nf_conntrack_max. Options: raise the limit and hash size, shorten timeouts, or exempt bulk traffic from tracking.
solid answer
~50 sThat message means the connection-tracking table has hit `net.netfilter.nf_conntrack_max`, so any packet that would need a new entry is dropped. The symptom is characteristic: established flows are unaffected because their entries already exist, while new connections fail seemingly at random. Diagnosis is comparing the current count against the maximum and looking at what is consuming entries — usually a high connection rate, a flood of short-lived UDP, or long-lived TCP entries held open by the five-day established timeout after peers vanished. The fixes, in order: raise `nf_conntrack_max` together with the hash-bucket count so lookups stay short, shorten the timeouts that are actually holding entries, and on high-rate paths use the raw table to exempt traffic from tracking entirely. Sizing matters: each entry costs a few hundred bytes of unswappable kernel memory, so the ceiling should be a deliberate number, not an arbitrarily large one.
go deeper
Recognise that the kernel keeps a finite table of tracked flows and that filling it makes new connections fail while existing ones keep working. Knowing where to look for the count and the limit is enough here.
Explain what an entry is, that it is retired by a timer rather than by a connection closing, and why the entry limit and the hash bucket count have to be raised together.
Diagnose from the symptom shape, identify which traffic is consuming entries, and choose between enlarging the table, shortening the timeouts that actually hold entries, and exempting traffic from tracking — with the memory cost stated.
Set the standard: what conntrack sizing and timeout profile the fleet ships with, what the alerting threshold on utilisation is, and where stateful inspection is worth its per-flow memory and CPU cost at all.
## What is actually full Connection tracking maintains one entry per flow, holding the original and reply tuples, the state, and any NAT mapping. Those entries live in a fixed-size hash table sized at module load. When the number of entries reaches `net.netfilter.nf_conntrack_max`, the kernel cannot record a new flow, so it drops the packet that would have created one and logs the message — rate-limited, so the log line count says nothing about the scale of the loss. The resulting failure mode is diagnostic in itself: **existing connections are fine, new ones fail**. Anything already in the table has its entry and continues to be forwarded and translated; every new connection attempt has a chance of being discarded. Users describe it as "the internet is flaky" while a ping to an already-open destination looks perfect and application logs show nothing but timeouts. ## Confirming it Compare the live count with the ceiling: ``` sysctl net.netfilter.nf_conntrack_count net.netfilter.nf_conntrack_max ``` If the count is pinned at the maximum, that is the answer. It is worth alerting on this ratio permanently on any gateway, because the condition is invisible until it is total. ## What sets the ceiling Two knobs work together. `net.netfilter.nf_conntrack_max` is the entry limit. `net.netfilter.nf_conntrack_buckets` is the number of hash buckets, and it is what determines lookup cost: entries that hash to the same bucket form a chain that must be walked. The defaults are derived from system memory at module load, so a small VM and a large router start from very different numbers. Raising `nf_conntrack_max` without raising the bucket count is the classic half-fix: the table accepts more flows, but each lookup walks longer chains, and lookups happen for every packet. Keep the conventional ratio — roughly four entries per bucket — when scaling up. Memory is the real constraint. Each entry costs a few hundred bytes of kernel memory that cannot be swapped. Millions of entries is a decision about RAM, not a free parameter. ## What is consuming entries Three patterns account for most cases. **Sheer connection rate.** A busy proxy or gateway opening many short-lived outbound connections generates entries faster than timeouts retire them. **Long timeouts on dead flows.** The established-TCP timeout defaults to five days. A connection whose peer disappeared without a FIN or RST — a laptop closing its lid, a container removed — holds its entry for that whole time. On a gateway with high churn this is often the single largest contributor, and reducing it to hours is a common and safe change. **UDP and scan traffic.** Every distinct UDP tuple creates an entry, so DNS volume, VoIP, telemetry, or a port scan against the gateway can each burn entries fast. Their timeouts are short — tens of seconds — but the arrival rate can outrun them. ## The remedies, in order 1. **Right-size the table.** Raise `nf_conntrack_max` and `nf_conntrack_buckets` together, and persist them so they survive reboot and module reload. 2. **Shorten the timeouts that are actually holding entries.** The established-TCP timeout is the usual target; the transitional TCP and UDP timers are already short. 3. **Stop tracking what does not need tracking.** Rules in the `raw` table run before conntrack and can exempt traffic from it entirely. High-rate flows that need no stateful matching and no translation — a monitoring feed, bulk traffic between two known endpoints — cost nothing once untracked. Note the tradeoff: untracked traffic cannot be matched by state, so it must be permitted explicitly. 4. **Reduce what passes through the box.** If the gateway is tracking traffic that has no reason to be translated or inspected, routing it around the NAT path removes the load rather than accommodating it. 5. **Clear stale entries as an emergency measure.** Flushing entries frees the table immediately and breaks the flows that were using them; it buys time, it is not a fix. ## The judgment being tested The interviewer is not looking for a sysctl name. They are looking for whether you recognise the *shape* of the failure — old connections fine, new ones failing, nothing wrong in the application — attribute it to a finite kernel resource, and then reason about whether to enlarge the resource, retire entries faster, or stop consuming it. And whether you notice that a machine which does no NAT and uses no stateful rules need not track connections at all.
- Why does raising nf_conntrack_max alone sometimes make performance worse instead of better?Because the hash bucket count is a separate setting. Entries hashing to the same bucket form a chain that is walked on every packet's lookup, so multiplying the entry limit without multiplying buckets lengthens those chains and raises per-packet cost across all traffic. Scale both together, keeping roughly the conventional four entries per bucket.
- Which timeout is usually worth changing first on a gateway under entry pressure, and why?The established-TCP timeout, which defaults to five days. Connections whose peer vanished without a FIN or RST keep their entry for that entire period, so on a host with high client churn dead flows can dominate the table. Cutting it to a few hours retires them quickly and rarely harms genuine long-lived sessions, which usually have keepalives anyway.
- How would you stop tracking traffic you do not need state for?Rules in the `raw` table run before connection tracking, so traffic matched there can be exempted from it and consumes no entry. It suits high-rate flows that need neither stateful matching nor translation. The tradeoff is that untracked packets can never match an established state, so they must be permitted by explicit rules in both directions.
- Does every Linux host pay this cost?No. Connection tracking is engaged when something needs it — a stateful rule, a NAT rule, or a helper. A host doing pure routing with stateless rules, or none at all, need not track anything, and on high-throughput forwarding boxes deliberately avoiding stateful matching is a legitimate design choice rather than an oversight.
saying these in an interview costs you the question
- Blames the application because existing connections work
- Assumes the port range, not the tracking table, is exhausted
- Raises the entry limit without raising hash buckets
- Thinks entries are freed when a connection closes
- Treats flushing the table as a permanent fix