An `apt install` on an Ubuntu server is interrupted partway through. Afterwards apt refuses to run at all, first with "Could not get lock /var/lib/dpkg/lock-frontend" and later reporting that a package is only half-configured. How do you diagnose and recover the box?
answer
- two problems, fixed in order
- identify the holder before removing anything
- unattended-upgrades and the daily timers
- dpkg -l second column tells the state
- reconfigure fails until the cause is fixed
basics
~20 sTreat them as two separate problems. Find who holds the lock with lsof or fuser — usually unattended-upgrades — and wait or stop that service rather than deleting the lock file. Then fix the underlying failure and run dpkg --configure -a, followed by apt-get install -f if dependencies are still unsatisfied.
solid answer
~50 sThe lock message means another process is mid-transaction; identify it with `sudo lsof /var/lib/dpkg/lock-frontend` or `sudo fuser -v /var/lib/dpkg/lock-frontend` before touching anything. On Ubuntu it is almost always unattended-upgrades fired by `apt-daily-upgrade.timer`, so the right move is to wait for it, or stop the service during provisioning — never `rm` the lock while dpkg is live, because two writers on the dpkg status database is how you turn a delay into corruption. Once the lock clears, `dpkg -l | grep -v '^ii'` and `sudo dpkg --audit` show what is in a bad state: `iU` for unpacked-but-unconfigured, `iF` for half-configured. `sudo dpkg --configure -a` re-runs the pending maintainer scripts and `sudo apt-get install -f` fixes leftover unmet dependencies. Crucially, find out *why* the postinst failed first — usually a service that will not start, or a full `/var` — because otherwise the reconfigure fails identically.
code
bash · 11 lines# 1. Who is holding the package system?
sudo lsof /var/lib/dpkg/lock-frontend
systemctl list-timers 'apt-*'
# 2. What state was it left in?
dpkg -l | grep -v '^ii'
sudo dpkg --audit
# 3. Recover, once the underlying failure is fixed
sudo dpkg --configure -a
sudo apt-get install -fgo deeper
Know that only one apt or dpkg transaction can run at a time, that the lock usually means an automatic update is in progress, and that waiting is the first response rather than deleting anything.
Name the lock files, identify the holder with lsof or fuser, and explain what dpkg --configure -a and apt-get install -f each do. Read the two-letter package states in dpkg -l.
Demonstrate root-cause discipline: the reconfigure only succeeds once the failing maintainer script's real cause is fixed, so go to journalctl and disk usage before retrying, and keep the forceful dpkg options as a recorded last resort.
Own the prevention: provisioning that masks the daily apt timers, terminates package operations gracefully, and checks capacity before patching, plus a documented recovery runbook so this is a five-minute known incident rather than a rebuild.
## Two problems that arrive together The symptoms look like one incident, but they are separate and are fixed in order: something else is holding the package system, and separately the package database has been left mid-transaction. Fix the lock first, because you cannot diagnose the second while you cannot run anything. ## Who holds the lock apt and dpkg serialise on several files: * `/var/lib/dpkg/lock-frontend` — held for the whole apt-level transaction, and the one usually named in the error. * `/var/lib/dpkg/lock` — the dpkg database lock itself. * `/var/cache/apt/archives/lock` — the download cache. * `/var/lib/apt/lists/lock` — held during `update`. Identify the holder rather than guessing: ```bash sudo lsof /var/lib/dpkg/lock-frontend sudo fuser -v /var/lib/dpkg/lock-frontend ps -ef | grep -E 'apt|dpkg|unattended' ``` On Ubuntu the answer is nearly always `unattended-upgrades`, started by `apt-daily.timer` (the index refresh) and `apt-daily-upgrade.timer` (the install pass). A first-boot provisioning run racing those timers is the single most common cause of this error in automation. The correct responses, in order of preference: wait for the run to finish; stop the service for the maintenance window; or, for images and golden-provisioning where you want determinism, mask the two timers for the duration and re-enable them at the end. `systemctl list-timers 'apt-*'` shows when they will next fire. What you must not do is delete the lock file while its owner is alive. The lock is what prevents two processes writing `/var/lib/dpkg/status` at once, and removing it converts a wait into genuine database damage. Deleting the file is only defensible when you have positively established that no process holds it — a stale lock left by a killed process or an unclean shutdown — and even then, checking with `lsof` first is the whole safeguard. ## Reading the damaged state Once the lock clears, ask dpkg what it thinks: ```bash dpkg -l | grep -v '^ii' sudo dpkg --audit ``` In `dpkg -l` the first column is the desired state and the second the current one, so: * `iU` — install wanted, package unpacked but not configured. * `iF` — install wanted, package half-configured; its `postinst` ran and failed. * `iH` — half-installed; unpacking itself was interrupted. * `rc` — removed, but its configuration files remain (usually benign). `dpkg --audit` prints the same set as prose, which is easier to hand to a colleague. ## Recovering ```bash sudo dpkg --configure -a sudo apt-get install -f ``` `dpkg --configure -a` walks every package left unconfigured and re-runs its `postinst`. `apt-get install -f` (equivalently `apt --fix-broken install`) asks apt's resolver to complete a transaction whose dependencies are unsatisfied, typically fetching packages that were unpacked but never pulled in. Both are useless — and will fail identically — if the original cause is still present. Debian and Ubuntu maintainer scripts commonly start or restart the service as their last act, so a `postinst` failure usually means the service itself refused to start. Read the actual failure: ```bash journalctl -u nginx --since '30 min ago' -p err df -h /var df -i /var ``` The recurring causes are a port already bound by something else, a configuration file the package upgrade rendered invalid, a full `/var` (or exhausted inodes, which `df -h` alone will not show), and a service whose dependency is itself broken. Fix that, then reconfigure. ## When it still will not clear Occasionally a package is stuck in a state where dpkg refuses to proceed until it is reinstalled. `sudo apt-get install --reinstall <pkg>` is the ordinary escape. The forceful options — `dpkg --remove --force-remove-reinstreq` and friends — do exist and do work, but they tell dpkg to ignore its own safety checks, and they should be a last resort on a box you are willing to rebuild, with the reason recorded. ## The preventable version Most of these incidents are avoidable. Provisioning that stops or masks the daily apt timers before it starts, does not kill package operations on timeout, and checks `/var` free space up front will simply not produce this failure. And when an install genuinely must be interrupted, sending `SIGTERM` and letting dpkg finish the current package is far kinder than `SIGKILL` in the middle of an unpack.
- Is it ever right to delete /var/lib/dpkg/lock-frontend?Only after positively establishing that nothing holds it — `lsof` and `fuser` both silent, and no apt or dpkg process running — which is the genuinely stale case after an unclean shutdown or a killed process. Removing it while its owner is alive allows two writers on the dpkg status database, turning a wait you would have won by doing nothing into real corruption.
- `dpkg --configure -a` fails with the same error every time you run it. What now?Stop re-running it and read why the maintainer script fails. It usually ends by starting the service, so `journalctl -u <unit>` shows the real cause — a bound port, a config the upgrade invalidated, a missing dependency. Also check `df -h /var` and `df -i /var`, since a full filesystem or exhausted inodes produces the same failure and is invisible in apt's own output.
- How do you stop a provisioning run from racing unattended-upgrades in the first place?Mask `apt-daily.timer` and `apt-daily-upgrade.timer` at the start of the run and restore them at the end, or stop `unattended-upgrades.service` for the window. `systemctl list-timers 'apt-*'` shows when they would next fire. Retrying on the lock error is a weaker fallback: it works, but it makes provisioning duration depend on someone else's transaction.
saying these in an interview costs you the question
- Deletes the lock file as the first step
- Assumes the lock means a crashed process rather than checking
- Reruns dpkg --configure -a without reading why postinst failed
- Forgets a full /var or exhausted inodes causes identical symptoms
- Reaches straight for --force-remove-reinstreq to make it go away