skip to content

What does the Atlas IP access list control, and why does a blocked client see a timeout?

level: middleimportance: should knowfreq 64%

answer

  1. An allowlist, evaluated at the edge
  2. It runs before any credential check
  3. That is why the error is a timeout
  4. Matches the public egress address, not the pod IP
  5. 0.0.0.0/0 is the classic anti-pattern

basics

~20 s

The Atlas IP access list is a per-project allowlist of public source addresses, CIDR blocks, or AWS security groups permitted to open a connection. Traffic from anything else is dropped before authentication, so the driver reports a server-selection timeout rather than an authentication error.

solid answer

~50 s

Every Atlas project carries an IP access list, and it is evaluated at the network edge before any credential is examined. Entries are single addresses, CIDR ranges, or — when AWS VPC peering is in place — a peer security group; each entry can carry a comment and an optional expiry so temporary access cleans itself up. The address Atlas matches on is the **public source address it observes**, which for containers and serverless functions is the NAT gateway or egress address, not the pod or task IP. Because rejected packets never reach the authentication handshake, the failure looks like `ServerSelectionTimeoutError` or a connection timeout — a useful diagnostic tell, since a wrong password would instead produce an explicit authentication failure. The anti-pattern is `0.0.0.0/0`, which leaves the cluster reachable from the whole internet with only credentials and TLS in front of it. Manage entries through the Atlas Admin API or Terraform so they stay reviewable.

go deeper

for a junior

Know that Atlas refuses connections from addresses you have not added, where that list lives, and that opening it to the whole internet is not an acceptable fix.

for a middle

Explain that admission happens before authentication, so a blocked client times out rather than failing to authenticate, and that Atlas matches the public egress address it observes.

for a senior

Demonstrate the diagnostic split between timeout and auth errors, handle NAT and autoscaling egress properly, and manage entries as code with comments and expiries rather than console clicks.

for a principal

Own the perimeter strategy: private endpoints for application traffic, a minimal audited access list for tooling, project-per-environment isolation, and a policy that catches an 0.0.0.0/0 entry before it reaches production.

## What the list is Every Atlas project has an IP access list (formerly called the whitelist). It is an allowlist of source addresses permitted to open a connection to any cluster in that project. It is a **project-scoped** control, not a per-cluster one: adding an entry opens that source to every cluster in the project, which is another reason to keep environments in separate projects. An entry can be: - a single IP address, - a CIDR block, or - an AWS security group ID, usable when the project has AWS VPC peering configured — this is much more robust than chasing instance IPs, because membership follows the security group. Each entry can carry a free-text comment (use it — an unexplained address is impossible to safely remove later) and an optional expiry, after which Atlas deletes it automatically. Temporary entries are the right shape for a support engineer debugging from a laptop. ## Why the failure looks like a timeout This is the detail interviewers probe. Admission control happens **before** the MongoDB handshake. A packet from an address that is not on the list is dropped; the driver never negotiates TLS, never runs `hello`, and never presents credentials. What the application sees is the driver exhausting its server-selection timeout and reporting that no suitable server was found — often with a message about all hosts being unreachable. Contrast that with credential problems, which happen *after* a connection is established: - Wrong username or password → an explicit authentication failure. - Correct credentials, missing role → an authorization error, `not authorized on <db> to execute command ...`, on the first operation. So the error class alone tells you which control failed. "Timeout with no auth error" points at the network layer: the access list, egress rules, DNS resolution of the SRV record, or a proxy. "Auth failed" points at the database user. Being able to make that split quickly is most of the value of the question. ## Which address must be listed Atlas matches the public source address of the connection *as Atlas sees it*. That routinely surprises people: - **Kubernetes**: the pod IP is private and meaningless to Atlas. What must be listed is the egress address — the NAT gateway, the node's public IP, or an egress-controller address. If pods can leave through several NAT gateways, all of them need entries. - **Serverless functions**: without a fixed egress path, addresses are unstable. Pin the function to a VPC with a NAT gateway, or use a private endpoint instead. - **Office and VPN users**: home ISP addresses change. Prefer VPN egress, and use short-lived entries for exceptions. - **Autoscaling fleets**: list the CIDR or, on AWS with peering, a security group rather than individual instances. ## Interaction with private connectivity The access list governs the public path. If the project uses **VPC peering**, the traffic arrives over private addressing but the access list still applies — you add the peer VPC's CIDR or the peer security group. If the project uses **private endpoints** (AWS PrivateLink, Azure Private Link, GCP Private Service Connect), the endpoint approval is itself the admission control for that path, so you are not maintaining public IP entries for it. A hardened deployment often ends up with a private endpoint carrying application traffic and a very short access list covering only administrative tooling. ## Operating it Treat the list as configuration, not as console clicks. The Atlas Admin API and the Terraform provider both manage entries, which gives you review, history, and a reproducible environment. In CI, a job that needs cluster access from ephemeral runners can add an entry with a short expiry via the API and let Atlas clean up. The worst common shortcut is `0.0.0.0/0`. It does not disable authentication or TLS, so it is not an instant breach, but it discards an entire independent layer of defence and exposes the cluster to internet-wide credential-stuffing and scanning. It shows up constantly in tutorials and then survives into production; finding and removing it is a standard hardening task. If you genuinely cannot enumerate egress addresses, that is an argument for private endpoints, not for opening the world. Finally, remember the access list is only admission. It never grants privileges: a listed address with no valid database user gets nowhere, and a valid database user from an unlisted address gets nowhere either. Both gates are independent, and production needs both to be tight.

  • A pod in Kubernetes cannot reach Atlas. Which address do you add, and how do you find it?
    Add the cluster's egress address — the NAT gateway or node public IP that Atlas actually observes — not the pod IP. Find it by curling an echo service from inside the pod, or by reading the NAT gateway's elastic IP. If pods egress through several gateways, every one needs an entry, which is usually the moment to switch to a private endpoint instead.
  • How do you give a CI job temporary access without leaving a permanent hole?
    Have the job add its runner's address through the Atlas Admin API with a short expiry, so Atlas removes it automatically, and delete it explicitly at the end of the run. Better still, run CI from a fixed egress or a private endpoint so no dynamic entry is needed at all.
  • Does the IP access list replace the need for TLS and authentication?
    No. It is admission control only — it decides who may open a connection, never what they may do. Every connection still negotiates TLS and still authenticates as a database user with explicit roles. Treating the access list as sufficient is exactly how an over-permissive entry turns into a breach.

saying these in an interview costs you the question

  • Adds 0.0.0.0/0 to make the connection work and leaves it
  • Thinks a blocked source produces an authentication error
  • Tries to list private pod or container IPs
  • Believes the access list is per cluster rather than per project
  • Treats network admission as a substitute for least-privilege roles

context