You inherit an nginx deployment where 60 virtual hosts are hand-edited into a single 3,000-line nginx.conf. How would you restructure that configuration, and how would you make changing it safe?
answer
- one file per vhost, included by mask
- shared defaults once at http level
- snippets for the repeated fragments
- validate rendered config in CI
- templating trades greppability for consistency
basics
~20 sSplit it into one file per virtual host pulled in by include, lift genuinely shared settings once into the http context so inheritance does the repetition, factor recurring fragments into snippets, and gate every change on a rendered nginx -t in CI with a scripted reload.
solid answer
~60 sThree moves, in order. First, **split by ownership**: one file per vhost under `conf.d`, included by mask, so a change to one site is a change to one file and reviewers can see whose it is. Second, **lift the shared settings**: anything identical across all 60 hosts belongs once in `http` and is inherited, and repeated fragments that are not fleet-wide — TLS parameters, header sets, proxy defaults — become snippet files each vhost includes. That is what turns 3,000 lines into a few hundred plus data. Third, **make change safe**: keep the whole tree in version control, render and validate it in CI with `nginx -t` inside an image matching production, deploy by writing files and reloading in one scripted step, and verify afterwards with a request rather than the reload's exit code. The trade to state honestly: heavy templating buys consistency and costs greppability, since the file on disk is no longer what a human wrote — so keep the include depth shallow and treat `nginx -T` as the readable form.
go deeper
Know the shape of a tidy config: one file per site pulled in by include, common settings declared once higher up, and nginx -t run before any reload.
Explain how inheritance and snippet includes remove duplication, why the include mask and sorted wildcard order matter, and what a per-vhost file buys during review.
Describe the change pipeline end to end: version control, render, validate in a production-matching image, scripted write-and-reload, then verify with a real request per changed host.
Own the trade: how far to templatise before the config stops being readable, where per-site escape hatches are allowed, who owns fleet-wide defaults, and how a captured nginx -T dump per deploy serves as the audit trail.
## What actually makes a big nginx config unmanageable Size alone is not the problem. Three things are: 1. **Repetition without a source of truth.** The same TLS block, header set and proxy defaults copied 60 times, drifting slowly, so no one can say what the intended value is. 2. **No ownership boundary.** One file means every change touches the same blob, every review is a diff against unrelated sites, and one bad edit takes down all 60 hosts at once because the config is parsed as a whole. 3. **No safe way to test.** Changes are made on the host and validated by reloading production. Each has a different fix, and doing them in the wrong order wastes effort — splitting files without deduplicating just distributes the duplication. ## Split by ownership One file per virtual host under `/etc/nginx/conf.d/`, or `sites-available` with `sites-enabled` symlinks on Debian-family systems. The mask matters: `include /etc/nginx/conf.d/*.conf;` silently ignores anything not ending in `.conf`, which is a trap for backups and a feature for staging a file you have not enabled yet. Wildcard includes are expanded in sorted order, so use a numeric prefix (`00-defaults.conf`) when order matters — for the default server, for a catch-all vhost, or where first-wins behaviour applies. Keep the ordering dependency to a minimum; a config where site 40 depends on being parsed after site 12 is a config nobody can safely add to. ## Lift the shared settings, and trust inheritance This is where the line count actually goes. Anything identical across all hosts — `log_format`, `sendfile`, `keepalive_timeout`, gzip settings, default `client_max_body_size`, default proxy timeouts — is declared once in `http` and inherited by every server and location beneath it. A vhost then contains only what makes it different: `listen`, `server_name`, its certificate paths, its locations. For fragments that recur but are not fleet-wide, use snippet files: ```nginx # conf.d/shop.example.com.conf server { listen 443 ssl; server_name shop.example.com; include snippets/tls-modern.conf; include snippets/proxy-defaults.conf; location / { proxy_pass http://shop_app; } } ``` One caveat worth designing around: the list-valued directives (`add_header`, `proxy_set_header`) are inherited only if the level declares none of them, so a snippet is often the *correct* mechanism rather than a convenience — it lets a deeper level restate the full set from a single source. ## Decide how far to templatise Once vhosts are uniform, generating them from data is tempting: a per-site record plus a template rendered by your configuration-management or build tooling. It is right when sites are genuinely homogeneous and numerous, and when onboarding a site should be a data change rather than a config change. The cost is real and worth saying out loud in an interview: the file on the host is no longer what a human wrote, so debugging gains an indirection step, and a template bug is a fleet-wide bug rather than a one-site bug. A reasonable middle ground is templating the repetitive 90% while allowing a per-site escape hatch include for the genuinely odd host — with the escape hatch reviewed, because it is where drift returns. ## Make change safe - **Version control the whole tree**, including snippets, and review changes like code. - **Validate in CI**: render the config and run `nginx -t` inside a container built from the same nginx version and module set as production. A syntax check against a different build can pass on a directive the production binary does not have. - **Deploy atomically**: write files, `nginx -t`, then reload; never hand-edit on a host. Reload is graceful — new workers pick up the new config while old ones drain — so a correct config applies without dropping connections. - **Verify by behaviour**, not by exit code: reload success only means the config parsed. Follow it with a request against a canary path per changed vhost. - **Keep `nginx -T` as the readable artifact.** However many includes and templates you add, that dump is the config as nginx understands it. Capturing it on each deploy gives you a diffable record of what really changed — often the fastest way to explain an incident. ## What good looks like afterwards A reviewer can see which site a change affects; a fleet-wide value has exactly one place to change; a broken edit fails in CI rather than on reload; and the number of lines is proportional to the number of genuinely distinct decisions, not to the number of hosts.
- What is the risk of deep include nesting, and how do you keep it in check?Every layer of indirection makes it harder to answer "what value is in effect for this host" by reading files, and reviewers start approving diffs whose effect they cannot see. Keep nesting to about two levels — vhost file plus snippet — forbid snippets that include other snippets, and treat the `nginx -T` dump as the artifact people actually read and diff.
- How would you validate a configuration change before it reaches production?Render it and run `nginx -t` inside a container built from the production nginx version and module set, since a directive from a module you do not ship will pass elsewhere and fail on the host. Then deploy to one node, request a canary path per changed vhost, and only roll on when the responses match expectations.
- When is templating nginx config the wrong call?When the sites are genuinely heterogeneous, or when there are few of them. A dozen hand-written, greppable vhosts are easier to reason about than a dozen renderings of a template full of conditionals — and a template bug hits every site at once. Templating pays off with homogeneity and volume, where onboarding should be a data change.
saying these in an interview costs you the question
- Splits into many files but leaves the duplication in place
- Repeats fleet-wide settings in every server block instead of using http
- Validates changes by reloading production and watching
- Templates everything, leaving no readable config on the host
- Treats a successful reload as proof the intended change is live