You lead platform engineering at a company shipping to Kubernetes clusters, VM fleets and managed serverless. Would you mandate pull-based GitOps delivery everywhere, and how would you decide?
answer
- mandate properties, not mechanism
- does a reconciler exist for this API
- blast radius scales with environment count
- agents are a second control plane
- push is fine if the credential is short-lived and scoped
basics
~20 sNo. Pull-based delivery needs a reconciling agent for the target's API, which is mature for Kubernetes and rare elsewhere. Mandate it for cluster workloads and keep push with short-lived, narrowly scoped credentials for VMs and serverless.
solid answer
~50 sA blanket mandate fails on the first target that has no reconciler. Pull delivery presumes an agent that can continuously compare declared state to a live API and converge it; Kubernetes has well-established ones, VM fleets and managed serverless mostly do not, so mandating pull there means writing and operating agents yourself. I would decide per target class on four axes: does a mature reconciler exist for this API; how large is the credential blast radius you are removing, which scales with the number of environments; do you need synchronous gating or ordered cross-service releases at deploy time; and who will operate, upgrade and debug the agents. That usually lands on pull for Kubernetes workloads and configuration, and push elsewhere — but with the properties that actually mattered kept uniform: desired state in Git, changes made by reviewed commits, short-lived scoped credentials wherever push remains, and one audit story across both.
go deeper
Know that pull-based delivery depends on an agent that can reconcile the target, and that this exists for Kubernetes but not for most other deployment targets.
Be able to name what each model needs — a reconciler and an agent to operate for pull, a credential and a network path for push — and give a concrete target where pull does not fit.
Argue the decision on evidence: environment count driving credential blast radius, whether releases need deploy-time gating, and the real operational cost of running agents across a fleet.
Separate the properties from the mechanism and mandate the properties — declared state in version control, short-lived scoped credentials, build-once promotion, one deployment record — then choose pull or push per target class and sequence the migration where the payoff is concentrated.
## Reject the framing, then give a rule The question is a mandate question, and the interviewer is watching for whether you can separate the *properties* GitOps delivers from the *mechanism* that delivers them. The properties — desired state declared in version control, changes made by reviewed commit, continuous convergence, an audit trail that is the repository history — are worth standardising. The pull mechanism is one implementation, and it has a hard prerequisite. ## The prerequisite: a reconciler for the target API Pull delivery needs software that can (a) read declared state, (b) query the live state of the target, (c) compute and apply a difference, and (d) do it repeatedly and safely. Kubernetes is unusually good at this because the API is declarative, the objects carry both spec and status, and mature agents exist. Most other targets are not: - **VM fleets** have configuration-management agents that converge host state, but rolling out new machine images is typically an imperative orchestration. - **Managed serverless and cloud resources** are reconciled by infrastructure-as-code tooling, usually run from a pipeline, sometimes by a controller — but the operating model differs enough that pretending it is the same thing hides real risk. - **Database schema changes** are inherently sequenced and irreversible; no continuous reconciler should own them. Mandating pull for these means either building agents you now maintain, or forcing a bad fit and getting a half-implementation. ## The four questions I would ask per target class **1. Does a mature reconciler exist for this API?** If not, the mandate cost is a bespoke control plane. Say no. **2. How much credential blast radius am I actually removing?** This scales with environments. One team with one cluster gains little — the credential exists either way, and you have added an agent to operate. Forty clusters across three business units, with production credentials currently sitting in a shared CI system, is a different conversation entirely, and the pull model is worth real cost there. **3. Do releases need deploy-time gating or ordering?** A reconciler has no run and no exit code. If a target's releases require synchronous approval, a coordinated sequence across services, or verification that must block promotion, either express that declaratively with health gating and dependencies or accept that a pipeline still owns it. **4. Who operates the agents?** Every agent is software with a version, CVEs, upgrade risk, its own privileged footprint, and a debugging surface engineers must learn. A platform team that mandates pull owns a second control plane. If you cannot staff that, the mandate degrades into unmaintained agents drifting behind, which is worse than a well-run push pipeline. ## What I would standardise regardless of mechanism This is the part that distinguishes a lead's answer: - **Desired state in version control for every target**, including the ones deployed by push. The audit and review properties do not depend on who initiates the connection. - **No long-lived deployment credentials anywhere.** Where push remains, use short-lived federated credentials scoped to one environment and one action, never a shared admin identity across environments. - **Build once, deploy many.** The artifact promoted between environments is the same one, whichever mechanism applies it. - **One deployment-event record.** Whether an agent or a pipeline performed the change, it should land in the same timeline so incident response does not need two tools. - **Merge rights treated as production access**, since under pull that is exactly what they are, and under push it is what they are becoming. ## The answer I would give Mandate the properties, choose the mechanism per target. Pull for Kubernetes workloads and cluster configuration, where the reconciler is mature and the credential win is largest — and roll it out where the environment count makes it worth the operational cost, not as an org-wide edict on day one. Push for everything else, hardened with short-lived scoped credentials. Revisit when a target class gets a reconciler worth adopting. What I would not accept is a team keeping long-lived production admin credentials in a CI system because "we do not do GitOps here" — that is the risk the whole discussion exists to address, and it has a fix in either model.
- A team with a single cluster asks to keep pushing with kubectl from CI. What do you require of them?I would accept it and set conditions: no long-lived kubeconfig, use short-lived federated credentials scoped to that one cluster and to the namespaces they own, keep manifests in version control so the desired state is reviewable, and deploy the same artifact they built rather than rebuilding. The credential properties matter more than which side initiates the connection.
- How would you sequence a migration to pull-based delivery across many teams?Start where the payoff is concentrated — the clusters whose credentials currently sit in a shared CI system and cover production. Prove the operating model on one team including alerting, break-glass and upgrades, publish it as a paved path, then migrate by demand rather than decree. A mandate ahead of a working support model produces unmaintained agents.
- What signal would tell you the pull model was not paying off for a team?If they have rebuilt a synchronous deploy-and-verify pipeline on top of the agent — polling for sync, gating on it, running verification inline — they now operate two control planes for one outcome. Either the deployment genuinely needs pipeline-time orchestration, or the agent's feedback story is inadequate; both are worth fixing rather than tolerating.
saying these in an interview costs you the question
- Mandating pull everywhere without asking whether a reconciler exists
- Treating GitOps as a tool choice rather than a set of properties
- Ignoring the cost of operating and upgrading agents per cluster
- Assuming push delivery must mean long-lived admin credentials
- Applying the same model to database schema changes as to workloads