Why is git pull risky in an unattended deploy script, and what should it run instead?
answer
- Only one half of pull is deterministic
- Think about what the script cannot recover from
- Conflict markers land in served files
- Separate transfer from update
- fetch, then --ff-only or a hard reset
basics
~20 sBecause pull's integration step can fail in ways a script cannot handle: it can create a merge commit, stop mid-conflict with markers in deployed files, or abort on divergence. A fetch followed by an explicit, deterministic update to the fetched tip is predictable.
solid answer
~50 s`git pull` is fetch plus merge or rebase, and only the fetch half is deterministic. On a server checkout the branch can diverge — a local hotfix, a file edited in place, a previous half-finished run — and then pull either creates a merge commit nobody reviewed, or stops with conflict markers written into files the application is about to serve, or refuses outright. None of those are states an unattended script recovers from, and the failure surfaces as broken content rather than a clean non-zero exit. The safer pattern is to separate transfer from update: `git fetch origin` always succeeds or fails cleanly, and then move the checkout to a known object — `git merge --ff-only origin/main` if you want the run to fail loudly when the checkout has diverged, or `git reset --hard origin/<branch>` if the checkout is disposable and must simply match the remote.
code
bash · 7 linesset -e
git fetch origin
# fail loudly if the checkout has diverged:
git merge --ff-only origin/main
# or, for a disposable checkout that must equal the remote:
# git reset --hard origin/main
git rev-parse HEADgo deeper
Know that pull can stop mid-way with conflicts or create a merge commit, and that a script cannot resolve either — so automation fetches first and updates explicitly.
Explain each failure mode in terms of pull's second half: fast-forward versus diverged, conflict markers written to the working tree, and a rebasing pull refusing a dirty tree.
Show the operating judgment: fetch is the only deterministic step, choose --ff-only to fail loudly or reset --hard for a disposable checkout, and log the deployed commit ID rather than a branch name.
Own the model behind the choice — is this checkout a cache of the remote or a machine humans touch? — and set the convention so every deploy path makes the same, stated assumption.
## The core objection Automation needs operations whose outcome is a function of their input. `git fetch` is one: it downloads objects and moves remote-tracking refs, and its only realistic failure is a network or authentication error, which exits non-zero and changes nothing about the working tree. `git pull` bolts a second command onto that — `git merge` or `git rebase` — whose outcome depends on the local state of a machine no one has looked at in months. ## The failure modes, concretely - **A merge commit nobody made.** If the checkout has any local commit and the remote has advanced, a merging pull commits a merge. The deployed tree is now content that exists nowhere in the reviewed history, and the machine has quietly become a fork. - **Conflict markers in served files.** A failed merge leaves the working tree half-updated with `<<<<<<<` markers written into real files. The application may keep serving them. The script sees a non-zero exit, but the damage is already on disk, and a naive retry loop re-runs into the same conflicted state. - **A refusal.** Modern Git aborts a diverging pull entirely when no pull mode is configured, and a rebasing pull refuses to start with a dirty working tree. Both are safer than the alternatives, but they mean the deploy silently did not happen unless the script checks. - **Local edits.** Someone debugging on the box edited a file in place. A merging pull may tolerate it, a rebasing one will not, and either way the deployed tree no longer matches any commit. - **Which branch?** `git pull` with no arguments resolves the remote and branch from the current branch's configuration. On a checkout in an unexpected state — detached HEAD, a leftover branch — it may pull something other than what the script intended. ## The pattern that works Split the two halves and make the second one explicit: First `git fetch origin`. This is the only network step, it is idempotent, and after it `origin/<branch>` names exactly the commit you intend to deploy. Then choose the update semantics deliberately: - **`git merge --ff-only origin/main`** when the checkout is supposed to be a clean follower. If it has diverged, the command fails without touching anything, and the deploy fails loudly with the checkout intact — which is what you want, because divergence on a deploy box is a real anomaly. - **`git reset --hard origin/main`** when the checkout is disposable and the contract is "this directory must equal the remote branch". This discards local commits and local edits by design; pair it with a clean step if untracked files also matter. Never use it where someone might legitimately keep state in the tree. Either way, record the exact object ID you deployed. `git rev-parse HEAD` after the update gives you the identifier that turns "the deploy ran" into "this commit is live". ## The deeper point This question is really the fetch-versus-pull distinction applied under pressure. Pull is a convenience for a human who is present, can read a conflict, and can decide what to do. Automation has none of those properties, so it should use the half of pull that is deterministic and then state its own integration policy explicitly. The same reasoning is why many engineers set `pull.ff=only` for their own interactive use: it converts an implicit decision into an explicit one. ## What a strong answer adds A strong candidate also notes that the choice between the fast-forward-only and the reset variant is a statement about what the checkout *is*. If it is a cache of the remote, reset is correct and divergence is meaningless. If it is a machine where humans occasionally and legitimately intervene, a hard reset destroys evidence, and failing loudly on divergence is the right behaviour. Picking one without deciding which model applies is the actual mistake — not the choice of command.
- When would you choose merge --ff-only over reset --hard for a deploy checkout?Choose `--ff-only` when divergence should be treated as an anomaly worth stopping for — the command fails without touching the tree, so the evidence survives for someone to look at. Choose `reset --hard` when the checkout is explicitly disposable and its contract is simply to equal the remote branch; then divergence is meaningless and discarding it is correct.
- How does the script know afterwards exactly what it deployed?Resolve and log the object ID: `git rev-parse HEAD` after the update. A branch name is not an answer — it moves. Recording the commit ID is what lets you correlate an incident with a specific tree, and it also lets a later run detect that the checkout is not where the previous deploy left it.
saying these in an interview costs you the question
- Says pull is fine because it usually fast-forwards
- Ignores that conflicts write markers into deployed files
- Retries the pull in a loop on failure
- Uses reset --hard on a box holding real local state
- Cannot separate the fetch half from the integration half