skip to content

An `apt install` on an Ubuntu server is interrupted partway through. Afterwards apt refuses to run at all, first with "Could not get lock /var/lib/dpkg/lock-frontend" and later reporting that a package is only half-configured. How do you diagnose and recover the box?

level: seniorimportance: should knowfreq 50%

answer

  1. two problems, fixed in order
  2. identify the holder before removing anything
  3. unattended-upgrades and the daily timers
  4. dpkg -l second column tells the state
  5. reconfigure fails until the cause is fixed

basics

~20 s

Treat them as two separate problems. Find who holds the lock with lsof or fuser — usually unattended-upgrades — and wait or stop that service rather than deleting the lock file. Then fix the underlying failure and run dpkg --configure -a, followed by apt-get install -f if dependencies are still unsatisfied.

solid answer

~50 s

The lock message means another process is mid-transaction; identify it with `sudo lsof /var/lib/dpkg/lock-frontend` or `sudo fuser -v /var/lib/dpkg/lock-frontend` before touching anything. On Ubuntu it is almost always unattended-upgrades fired by `apt-daily-upgrade.timer`, so the right move is to wait for it, or stop the service during provisioning — never `rm` the lock while dpkg is live, because two writers on the dpkg status database is how you turn a delay into corruption. Once the lock clears, `dpkg -l | grep -v '^ii'` and `sudo dpkg --audit` show what is in a bad state: `iU` for unpacked-but-unconfigured, `iF` for half-configured. `sudo dpkg --configure -a` re-runs the pending maintainer scripts and `sudo apt-get install -f` fixes leftover unmet dependencies. Crucially, find out *why* the postinst failed first — usually a service that will not start, or a full `/var` — because otherwise the reconfigure fails identically.

code

bash · 11 lines
bash
# 1. Who is holding the package system?
sudo lsof /var/lib/dpkg/lock-frontend
systemctl list-timers 'apt-*'

# 2. What state was it left in?
dpkg -l | grep -v '^ii'
sudo dpkg --audit

# 3. Recover, once the underlying failure is fixed
sudo dpkg --configure -a
sudo apt-get install -f

go deeper

for a junior

Know that only one apt or dpkg transaction can run at a time, that the lock usually means an automatic update is in progress, and that waiting is the first response rather than deleting anything.

for a middle

Name the lock files, identify the holder with lsof or fuser, and explain what dpkg --configure -a and apt-get install -f each do. Read the two-letter package states in dpkg -l.

for a senior

Demonstrate root-cause discipline: the reconfigure only succeeds once the failing maintainer script's real cause is fixed, so go to journalctl and disk usage before retrying, and keep the forceful dpkg options as a recorded last resort.

for a principal

Own the prevention: provisioning that masks the daily apt timers, terminates package operations gracefully, and checks capacity before patching, plus a documented recovery runbook so this is a five-minute known incident rather than a rebuild.

## Two problems that arrive together The symptoms look like one incident, but they are separate and are fixed in order: something else is holding the package system, and separately the package database has been left mid-transaction. Fix the lock first, because you cannot diagnose the second while you cannot run anything. ## Who holds the lock apt and dpkg serialise on several files: * `/var/lib/dpkg/lock-frontend` — held for the whole apt-level transaction, and the one usually named in the error. * `/var/lib/dpkg/lock` — the dpkg database lock itself. * `/var/cache/apt/archives/lock` — the download cache. * `/var/lib/apt/lists/lock` — held during `update`. Identify the holder rather than guessing: ```bash sudo lsof /var/lib/dpkg/lock-frontend sudo fuser -v /var/lib/dpkg/lock-frontend ps -ef | grep -E 'apt|dpkg|unattended' ``` On Ubuntu the answer is nearly always `unattended-upgrades`, started by `apt-daily.timer` (the index refresh) and `apt-daily-upgrade.timer` (the install pass). A first-boot provisioning run racing those timers is the single most common cause of this error in automation. The correct responses, in order of preference: wait for the run to finish; stop the service for the maintenance window; or, for images and golden-provisioning where you want determinism, mask the two timers for the duration and re-enable them at the end. `systemctl list-timers 'apt-*'` shows when they will next fire. What you must not do is delete the lock file while its owner is alive. The lock is what prevents two processes writing `/var/lib/dpkg/status` at once, and removing it converts a wait into genuine database damage. Deleting the file is only defensible when you have positively established that no process holds it — a stale lock left by a killed process or an unclean shutdown — and even then, checking with `lsof` first is the whole safeguard. ## Reading the damaged state Once the lock clears, ask dpkg what it thinks: ```bash dpkg -l | grep -v '^ii' sudo dpkg --audit ``` In `dpkg -l` the first column is the desired state and the second the current one, so: * `iU` — install wanted, package unpacked but not configured. * `iF` — install wanted, package half-configured; its `postinst` ran and failed. * `iH` — half-installed; unpacking itself was interrupted. * `rc` — removed, but its configuration files remain (usually benign). `dpkg --audit` prints the same set as prose, which is easier to hand to a colleague. ## Recovering ```bash sudo dpkg --configure -a sudo apt-get install -f ``` `dpkg --configure -a` walks every package left unconfigured and re-runs its `postinst`. `apt-get install -f` (equivalently `apt --fix-broken install`) asks apt's resolver to complete a transaction whose dependencies are unsatisfied, typically fetching packages that were unpacked but never pulled in. Both are useless — and will fail identically — if the original cause is still present. Debian and Ubuntu maintainer scripts commonly start or restart the service as their last act, so a `postinst` failure usually means the service itself refused to start. Read the actual failure: ```bash journalctl -u nginx --since '30 min ago' -p err df -h /var df -i /var ``` The recurring causes are a port already bound by something else, a configuration file the package upgrade rendered invalid, a full `/var` (or exhausted inodes, which `df -h` alone will not show), and a service whose dependency is itself broken. Fix that, then reconfigure. ## When it still will not clear Occasionally a package is stuck in a state where dpkg refuses to proceed until it is reinstalled. `sudo apt-get install --reinstall <pkg>` is the ordinary escape. The forceful options — `dpkg --remove --force-remove-reinstreq` and friends — do exist and do work, but they tell dpkg to ignore its own safety checks, and they should be a last resort on a box you are willing to rebuild, with the reason recorded. ## The preventable version Most of these incidents are avoidable. Provisioning that stops or masks the daily apt timers before it starts, does not kill package operations on timeout, and checks `/var` free space up front will simply not produce this failure. And when an install genuinely must be interrupted, sending `SIGTERM` and letting dpkg finish the current package is far kinder than `SIGKILL` in the middle of an unpack.

  • Is it ever right to delete /var/lib/dpkg/lock-frontend?
    Only after positively establishing that nothing holds it — `lsof` and `fuser` both silent, and no apt or dpkg process running — which is the genuinely stale case after an unclean shutdown or a killed process. Removing it while its owner is alive allows two writers on the dpkg status database, turning a wait you would have won by doing nothing into real corruption.
  • `dpkg --configure -a` fails with the same error every time you run it. What now?
    Stop re-running it and read why the maintainer script fails. It usually ends by starting the service, so `journalctl -u <unit>` shows the real cause — a bound port, a config the upgrade invalidated, a missing dependency. Also check `df -h /var` and `df -i /var`, since a full filesystem or exhausted inodes produces the same failure and is invisible in apt's own output.
  • How do you stop a provisioning run from racing unattended-upgrades in the first place?
    Mask `apt-daily.timer` and `apt-daily-upgrade.timer` at the start of the run and restore them at the end, or stop `unattended-upgrades.service` for the window. `systemctl list-timers 'apt-*'` shows when they would next fire. Retrying on the lock error is a weaker fallback: it works, but it makes provisioning duration depend on someone else's transaction.

saying these in an interview costs you the question

  • Deletes the lock file as the first step
  • Assumes the lock means a crashed process rather than checking
  • Reruns dpkg --configure -a without reading why postinst failed
  • Forgets a full /var or exhausted inodes causes identical symptoms
  • Reaches straight for --force-remove-reinstreq to make it go away

context