You inherit SSH access for a few hundred Linux servers where each account's `authorized_keys` file is maintained by hand. What actually breaks at that scale, and how would you choose between config-management-pushed key files, an `AuthorizedKeysCommand` lookup, and SSH certificates?
answer
- adding is loud, removing is silent
- who can log into this box?
- convergence speed vs runtime dependency
- access that expires by itself
- test the break-glass path
basics
~20 sRevocation and inventory break first: a departing engineer's key survives on whichever hosts were missed, and nobody can answer who can reach what. The three options trade convergence speed against a runtime dependency, and short-lived certificates remove per-host key files entirely.
solid answer
~60 sHand-maintained key files fail on removal, not on addition. Adding a key is visible and someone chases it; removing one is invisible, so a leaver's key persists on any host that was down, out of the inventory, or built from a stale image — and there is no authoritative answer to "who can log into this box". The three mechanisms differ in where the truth lives. Config management renders `authorized_keys` from a reviewed source of truth: simple, auditable in version control, but revocation is only as fast as convergence and only covers hosts the agent reaches. `AuthorizedKeysCommand` has sshd call a helper at login time to fetch the account's keys from a directory service: revocation is immediate, at the cost of putting a live dependency on the authentication path, which must fail closed and fast. Short-lived certificates signed by a CA remove the problem class — hosts trust one CA key, access expires by itself, and host certificates also end `known_hosts` churn — but you now operate and protect a signing service. Whichever you pick, keep one break-glass local key and tested console access.
code
bash · 2 linesssh-keygen -s /etc/ssh/ca/user_ca -I [email protected] -n deploy,ops -V +8h alice.pub
ssh-keygen -Lf alice-cert.pubgo deeper
Understand that authorised keys are per-account files, so access removal means touching every account on every host — and that this is exactly why larger fleets stop managing them by hand.
Be able to describe rendering authorized_keys from a reviewed source of truth, and explain why revocation then takes effect only at the next convergence rather than instantly.
Compare the mechanisms on revocation latency, availability coupling and auditability, and insist on fingerprint-level logging so a session on a shared account still attributes to a person.
Own the trade explicitly: name your revocation SLA, decide whether the authentication path may carry a runtime dependency, argue for expiry over revocation, and keep a break-glass route that is exercised on a schedule.
## What actually breaks At three hosts, hand-edited `authorized_keys` is fine. At three hundred, four things fail, and they fail quietly. **Revocation.** Granting access is self-correcting — the person complains until it works. Revoking is not: nobody notices a key that was left behind. When someone leaves, you must touch every account on every host, and the hosts that were powered off, absent from the inventory, or rebuilt from an old image keep the key. Access outlives employment, which is the finding that ends up in an audit report. **Inventory.** "Who can log into this server?" requires reading a file on that server, per account. There is no central answer, so the question is usually answered from memory, and the memory is wrong. **Drift.** Files edited by hand diverge. Two hosts that should be identical are not, and the difference is discovered during an incident. **Attribution.** Shared accounts with shared keys mean the log says `deploy` and nothing more. Per-person keys plus fingerprint logging is the minimum bar for saying who did something. ## Option 1 — render the file from config management The source of truth is a repository: a group-to-key mapping reviewed like code, rendered into each account's `authorized_keys` on every convergence run. *Buys you:* a real inventory (query the repo, not the fleet), change review, easy rollback, and no new runtime dependency — if the config-management server is down, logins still work off the last rendered file. *Costs you:* revocation is eventual. It takes effect on the next successful convergence, so your removal SLA is your convergence interval plus whatever fraction of the fleet the agent failed to reach. You also need a report of hosts that have not converged recently, or the gap is invisible again. This is the pragmatic answer for most fleets and is genuinely defensible. ## Option 2 — look the keys up at login time sshd can be configured to call a helper program that prints the authorised keys for the user being authenticated, running it under a dedicated unprivileged account. The helper queries a directory service or an internal API instead of reading a file. *Buys you:* immediate revocation — disable the person centrally and the next login attempt anywhere fails — plus one place to ask who has access. *Costs you:* you have put a network service on the authentication path of every host. It must be fast, because it runs inside every login; it must fail closed, or an outage becomes an authentication bypass; and if it fails closed, an outage of that service means nobody can log in anywhere — precisely when you need to. That is a real availability trade, and it is why the break-glass path below stops being optional. ## Option 3 — short-lived certificates A CA key signs a user's public key into a certificate carrying principals (which accounts it may access) and a validity window. Hosts are configured to trust the CA, and the per-host key files disappear. ``` ssh-keygen -s /path/to/user_ca -I [email protected] -n deploy,ops -V +8h alice.pub ``` *Buys you:* access that expires on its own — the strongest property in the list, because it removes the need to remember to revoke. Onboarding is issuing a certificate; offboarding is not issuing another. Hosts hold one trusted CA public key instead of hundreds of key lines. The same mechanism run in the other direction — signing *host* keys and giving clients one certificate-authority entry — eliminates the `known_hosts` warning after every rebuild, which is a second, often-underrated win. *Costs you:* a CA to run and protect. The signing key is now the crown jewel, so it wants hardware protection and a tightly controlled issuance service tied to your identity provider. Certificate revocation before expiry is awkward, which is why validity windows are kept short — hours, not months. And you need a plan for the day the issuance service is down. ## How to choose Ask three questions, in order: 1. **How fast must revocation be?** Hours is fine → config management. Minutes, contractually → lookup or certificates. 2. **Can the authentication path carry a runtime dependency?** If a login must work when the control plane is down, prefer rendered files or certificates, whose validation is local, over a live lookup. 3. **Do you have someone to operate a CA?** Certificates are the best end state and the highest floor of operational maturity. A poorly run CA is worse than well-run config management. A common and honest path is to run config management now, add fingerprint-level logging so attribution works, and move to certificates when there is a team to own the issuance service. ## The part nobody should skip Whatever you choose, keep a break-glass route that does not depend on the mechanism you just built: one locally installed emergency key, or console access, held under change control, alarmed when used, and — critically — *tested on a schedule*. Every one of these designs has a failure mode where the normal path stops issuing access, and the difference between an incident and a disaster is whether the escape hatch was exercised recently enough to still work.
- Why is expiry a stronger control than revocation?Revocation depends on someone remembering to act and on the removal reaching every host; expiry is the default outcome of doing nothing. Short-lived certificates fail closed by design, so a forgotten offboarding self-corrects within the validity window. Revocation lists still matter for the compromised-key case, but they are the exception path rather than the mechanism you rely on daily.
- What is the danger of putting a directory lookup on the authentication path of every host?You have coupled login availability to a service. If it fails open, an outage becomes an authentication bypass; if it fails closed — the correct choice — an outage locks everyone out of everything at the worst moment. It must be fast, cached carefully if at all, and paired with a break-glass route whose validity does not depend on it.
- Your incident response needs to know which human ran a command on a shared `deploy` account. What has to be true?Per-person key material and logging that records it. The daemon logs the key fingerprint on an accepted public-key login, or the certificate identity when certificates are used, so the session can be tied to a person. A shared private key or a shared password destroys that link permanently — no amount of downstream logging recovers it.
saying these in an interview costs you the question
- Treats offboarding as a one-host chmod
- Keeps one shared key for the whole team
- Puts a lookup service on the auth path with no fallback
- Issues certificates with year-long validity
- Never tests the emergency console path