skip to content

You own a fleet of Linux servers where several teams must operate their own services without being given full root. How would you design the account, group and sudo policy?

level: principalimportance: should knowfreq 36%

answer

  1. central identity, deliberate local exceptions
  2. grant to groups, review membership
  3. allow-list absolute paths and arguments
  4. policy as code, validated before install
  5. break-glass defined before the emergency

basics

~20 s

Back identity centrally, grant privilege to groups rather than named users, express every grant as a narrow reviewed sudoers drop-in shipped by configuration management, reserve passwordless rules for dedicated automation accounts, and keep an audited break-glass path for the day the directory is unreachable.

solid answer

~50 s

Four decisions carry the design. **Where identity lives**: a central directory reached through NSS and PAM, so joiners and leavers are one change rather than a change per host — with a small, deliberate set of local accounts that survive the directory being down. **What the unit of grant is**: always a group, never a username, so access review becomes membership review and the sudoers files stay stable. **What each grant permits**: an allow-list of absolute paths with exact arguments, reviewed for shell escapes and arbitrary-write primitives, ideally pointing at root-owned wrapper scripts that validate their own input rather than at general-purpose binaries. **How grants are shipped**: `/etc/sudoers.d` drop-ins in version control, syntax-checked with `visudo -cf` in CI before a host ever sees them. Around that, reserve NOPASSWD for non-interactive service accounts, forward sudo logs off the host so a root-capable user cannot edit their own trail, and define a break-glass route through the console before you need it.

go deeper

for a junior

Understand that privilege on shared servers is granted through group membership and a policy file, not by handing out the root password, and that changes are reviewed before they ship.

for a middle

Be able to write the narrow rule: a group, an absolute path, exact arguments, installed as a validated drop-in — and explain why each of those parts is there.

for a senior

Show the operational judgment: which grants are safe given what the binary can do, how automation gets its own account, and how you roll a policy change out without locking a fleet out.

for a principal

Own the whole model — identity source and its outage behaviour, group design tied to services rather than teams, the review cadence, off-host audit, and a documented break-glass path with alerting.

## Start from the failure you are preventing "Teams need to operate their services" is a request for a *capability*, not for root. The design goal is that the ordinary day requires no privilege escalation at all, the occasional day requires a narrow one, and full root is a rare, noisy, accountable event. Every decision below trades convenience against how much of the fleet one compromised session reaches. ## 1. Where identity lives Local accounts on each host do not scale past a handful of machines: offboarding becomes a fleet-wide sweep you can never prove finished. Back identity in one directory — LDAP/AD via SSSD, or your cloud provider's equivalent — reached through NSS for `passwd`/`group` and PAM for authentication. That gives one place to disable a person and one place to answer "who can reach production?". The cost is a hard dependency. Decide explicitly what happens when the directory is unreachable: SSSD's credential cache covers users who logged in recently, and you keep a small set of deliberately local accounts — typically one break-glass account per host or per role — so the fleet is recoverable. Write down which accounts those are; a fleet with an undocumented local account set has a shadow access-control system. Also standardise UID and GID **numbers** across hosts. When home directories or data volumes move between machines, ownership travels as numbers, and a `deploy` group that is GID 2001 here and 3007 there produces access failures that look like permission bugs. ## 2. The unit of grant is a group Never put a username in a sudoers rule. Grant to `%svc-payments-ops`, and let directory membership decide who that is. Three benefits: the sudoers files become stable artefacts that rarely change; adding or removing a person is a directory operation with its own audit trail; and access review turns into "who is in this group", a question a service owner can answer without reading sudoers syntax. Size the groups around a *service and an action*, not around a team or a host. `%app-restart` beats `%backend-team`, because the second grows silently as the team's remit grows. ## 3. What each grant permits Allow-lists only. For every rule, ask what the named binary can be made to do as root: does it offer a shell escape or a pager, can it write an attacker-chosen path, does it read configuration the caller controls? Any of those collapses the grant to full root, so the review is about the program's capabilities, not the intent of the sentence you wrote. Two sudoers specifics shape the rules: a command listed with no argument list permits every argument, and wildcards routinely match more than intended. So pin absolute paths and exact arguments. Where the argument genuinely varies, the durable answer is a small root-owned wrapper script that validates its input and does one thing, with sudo granted on the wrapper. That moves the policy into code you can test, and it keeps the sudoers line trivially reviewable. Where the platform offers a purpose-built mechanism, prefer it: `sudoedit` for editing root-owned files, so no editor ever runs as root; polkit rules for controlling specific systemd units over D-Bus; or an operations API that removes the need for a shell on the host entirely. The best sudo policy is the one most teams never need to invoke. ## 4. Passwords, NOPASSWD and sessions Humans re-authenticate. The cost is one prompt per timestamp window, and the benefit is that a hijacked terminal or a stolen SSH key is not immediately root. NOPASSWD belongs to non-interactive automation: a dedicated account, one narrow rule, key-based access, and no interactive login. Auditing "which NOPASSWD rules exist and which account each belongs to" should be a standing report, because that list is where privilege quietly accumulates. ## 5. Ship it as code Sudoers changes go through version control and configuration management as `/etc/sudoers.d` drop-ins, mode `0440`, root-owned, with names free of dots (a dot in the filename makes sudo ignore the file — a silent way to "apply" a policy that never loads). CI runs `visudo -cf` on every file, because an unparseable policy makes `sudo` refuse to run at all; on hosts where root has no password that is a lockout needing console access. Roll out to a canary host first and keep a second root session open while it lands. ## 6. Audit, and its honest limits Sudo logs the real user, the tty, the working directory and the argv. Forward that off the host in real time, because a user who successfully escalates to root can edit local logs. Add session capture (`log_output`, replayed with `sudoreplay`) for the highest-privilege roles only — it is expensive and it captures secrets typed into the terminal. Be candid about the ceiling: the moment a rule yields an interactive root shell, the log records the shell and nothing after it. That is an argument for narrow rules, not for more logging. ## 7. Break-glass and review Define the emergency path before the emergency: who can reach the console or BMC, how the local break-glass credential is stored and checked out, and what alert fires when it is used. Then set a review cadence — group membership quarterly, NOPASSWD rules and any rule granting a shell every cycle, and re-review a rule whenever the binary it names is upgraded, since a new subcommand can widen an old grant without anyone touching sudoers.

  • How do you keep hosts usable when the central directory is unreachable?
    Two layers. The directory client caches credentials for users who have logged in recently, which covers short outages for people already active. Underneath that, keep a documented, small set of local accounts — break-glass, and where needed one operational account per role — with their own sudo grants, stored credentials checked out under audit, and an alert on use. The point is that the fallback is designed and inventoried rather than discovered during the outage.
  • A team asks for NOPASSWD sudo so their deployment script stops prompting. What do you offer instead?
    Move the automation off the human's identity. Give the pipeline a dedicated service account with key-based access, one narrow NOPASSWD rule naming an absolute path and exact arguments — ideally a root-owned wrapper that validates its input — and no interactive login. The human keeps a passworded grant. That way the passwordless privilege belongs to something inventoried and reviewable, not to whichever person happened to be on call.
  • What are the real limits of sudo logging as an audit control?
    Sudo records the command it was asked to run, not what that command did. Once a rule yields an interactive root shell, everything after that line is invisible. A root-capable user can also edit local logs, so anything not forwarded off the host in real time is untrustworthy. Session capture helps for a few high-privilege roles but records secrets typed at the terminal. Narrow rules buy more auditability than any amount of logging.

saying these in an interview costs you the question

  • Naming individual users in sudoers instead of groups
  • Granting ALL with a few commands subtracted
  • Treating sudo logs as tamper-proof on the host itself
  • Letting NOPASSWD rules attach to human accounts
  • Shipping a sudoers change fleet-wide without validation

context