A nightly Django management command that expires shop coupons catches every exception and writes it to self.stderr; why does cron never alert, and how would you harden it?
answer
- what the scheduler actually watches
- handle() returned normally
- CommandError and its returncode
- re-runnable after a partial run
- autocommit, not ATOMIC_REQUESTS
basics
~20 sCatching every exception lets handle() return normally, so the process exits 0 and cron sees success. Raise CommandError or let exceptions propagate for a non-zero status, and make the work idempotent so a re-run finishes what a partial run left.
solid answer
~40 sA scheduler judges a job by its **exit status**. If `handle()` catches everything and writes to `self.stderr`, it returns normally and the process exits 0, so cron and any monitor see success while coupons stay active. Hardening has three parts. First, let failures fail: re-raise unexpected exceptions, or wrap them in `CommandError(..., returncode=...)` with a clear message, so the command line prints it to stderr and exits non-zero. Second, make the work **idempotent**: compute one cutoff with `timezone.now()`, select only rows still `is_active=True`, and update in set-based batches so a re-run completes what a crashed run left. Third, remember a command runs in autocommit — `ATOMIC_REQUESTS` covers views only — so each committed batch stays committed; design for partial progress instead of assuming rollback.
code
python · 38 lines# shop/management/commands/expire_coupons.py (hardened)
import logging
from django.core.management.base import BaseCommand, CommandError
from django.db import DatabaseError
from django.utils import timezone
from shop.models import Coupon
logger = logging.getLogger(__name__)
BATCH_SIZE = 500
class Command(BaseCommand):
help = "Deactivate expired coupons in batches; safe to re-run."
def handle(self, *args, **options):
cutoff = timezone.now()
total = 0
try:
while True:
ids = list(
Coupon.objects.filter(is_active=True, expires_at__lt=cutoff)
.order_by("pk")
.values_list("pk", flat=True)[:BATCH_SIZE]
)
if not ids:
break
total += Coupon.objects.filter(pk__in=ids, is_active=True).update(
is_active=False
)
except DatabaseError as exc:
logger.exception("coupon expiry stopped after %d coupon(s)", total)
raise CommandError(
f"Coupon expiry stopped after {total} coupon(s): {exc}"
) from exc
if options["verbosity"] >= 1:
self.stdout.write(self.style.SUCCESS(f"Expired {total} coupon(s)."))go deeper
Recall that a command which swallows exceptions exits 0, and that raising CommandError makes Django print the error and exit with status 1.
Explain run_from_argv's three outcomes (normal return, CommandError, other exception) and why commands run in autocommit rather than inside a request transaction.
Demonstrate the full hardening: honest exit status, idempotent set-based batches, a single cutoff, dry-run and verbosity support, and safe overlap between runs.
Define a house contract for scheduled commands — idempotent, re-runnable, exit-status truthful, monitored on both status and output — so every team's jobs fail loudly and recover alike.
## Why the failure is invisible Schedulers such as cron, a container job runner or a CI pipeline do not read your messages; they read the **exit status** of the process. A management command started through `manage.py` goes through `BaseCommand.run_from_argv()`: - If `handle()` returns normally, the process exits with status **0**. - If `handle()` raises `CommandError`, Django writes `CommandError: <message>` to stderr and calls `sys.exit(returncode)` — **1** by default. - If `handle()` raises anything else, the exception escapes and Python prints a traceback and exits non-zero. A `try: ... except Exception as exc: self.stderr.write(str(exc))` around the body turns every one of those failures into the first case. The text lands in a log nobody reads and the job is recorded as successful. ## Step 1 — let failures change the exit status 1. Delete the blanket `except Exception`. Catch only what you can explain. 2. For an expected failure (a missing table, a locked database, a bad option), raise **`CommandError`** with a message an on-call engineer can act on; pass `returncode=` if your scheduler distinguishes failure kinds. 3. Chain the original exception with `raise CommandError(...) from exc` so `--traceback` shows the cause when you debug by hand. 4. Keep human-readable progress on `self.stdout` and send structured events to the `logging` module, so a monitor can alert on the log as well as on the status. ## Step 2 — make a partial run safe The second half of the incident is "half the coupons stay active". That is a design property, not bad luck: - **Management commands run in autocommit.** Each `save()` or `update()` commits on its own. `ATOMIC_REQUESTS` wraps views only; nothing wraps a command. - A loop that saves coupons one by one therefore leaves every coupon before the crash committed and every one after it untouched. Rather than fighting that, design for it: | Property | How the command gets it | |---|---| | **Idempotent** | Filter on the state you are changing (`is_active=True`) so already-expired coupons are never touched twice | | **Deterministic cutoff** | Compute `timezone.now()` once at the start, not per row | | **Set-based** | `QuerySet.update()` issues one `UPDATE` per batch instead of a `save()` per object | | **Bounded** | Batches of primary keys keep each statement and lock short | | **Re-runnable** | A crashed run leaves a smaller remainder that the next run finishes | If a batch must be all-or-nothing together with other writes, that is where `transaction.atomic()` comes in — a question about transaction design rather than about the command itself. ## Step 3 — make it observable and operable - Add `--dry-run` so an operator can see what would change before running it by hand. - Respect `options["verbosity"]`: quiet at `-v 0` for cron, detailed at `-v 2` when debugging. - End with one summary line (`Expired 1,204 coupon(s) in 3 batches.`) so the log answers "did it do anything?". - Remember that `manage.py` runs the **system checks** before `handle()`. After a deploy that introduces a check error, every nightly run fails with a `SystemCheckError`, which is itself a `CommandError` — a non-zero exit you now want to see. ## Overlapping runs If a slow run is still going when the next one starts, or someone runs the command by hand, an idempotent set-based update makes the overlap harmless: both runs only flip rows that are still active. Side effects per coupon, such as emailing the owner, are different — they need the row claimed exactly once, and that is a locking decision to make deliberately rather than an accident of timing. ## A review checklist for scheduled commands 1. Does every failure path end in a raised exception, so the exit status is non-zero? 2. Is the selection keyed on the state being changed, so a re-run is harmless? 3. Is the time window computed once, from an aware `timezone.now()`? 4. Are writes set-based and bounded, rather than one `save()` per row in one long loop? 5. Is there a `--dry-run`, and does the command stay quiet at `-v 0`? 6. Is there a test that calls it with `call_command()` and asserts both the output and the database state? ## Where the scheduling lives The command is the unit of work. Whether cron, a platform scheduler or a task queue's periodic scheduler starts it at 02:00 is a separate choice; the command's contract to all of them is the same: do the work idempotently and report the truth through the exit status.
- Why not wrap the whole expiry loop in one transaction.atomic() block so a crash rolls everything back?It would make the run all-or-nothing, but one long transaction holds row locks on every expired coupon until the end and loses all progress on any failure, so a slow night can block checkout writes and never finish. Short set-based batches under autocommit keep locks brief, and idempotency makes partial progress safe to resume.
- The command must email each coupon owner as it expires; what changes?Idempotency of the flag no longer covers the side effect: an overlapping or retried run could email twice. Each coupon must be claimed exactly once, for example by recording the send in the same write that flips the flag and only emailing rows that write claimed, and the email should be sent after that write commits.
saying these in an interview costs you the question
- Writing the error to self.stderr is enough for cron to report a failure
- Management commands run inside a transaction like ATOMIC_REQUESTS views
- A CommandError from manage.py exits with status 0 because Django handled it
- Saving coupons one by one is safer than a set-based update
- Catching Exception around handle() is the robust default for scheduled jobs