Your team has adopted a declarative IaC tool. Where does imperative scripting against a cloud SDK or CLI still legitimately win, and how do you keep those scripts from becoming invisible operations?
answer
- durable resource versus one-shot action
- no 'already true' condition to assert
- bootstrapping and bulk work
- dry-run mode is the poor engineer's plan
- reviewed and parameterised, not edited
basics
~20 sScripting still wins for one-shot actions rather than durable resources: operational tasks, bulk data work, bootstrapping the tooling itself, and anything with no resource model. Keep those scripts in the repository, reviewed, re-runnable and logged like any other change.
solid answer
~60 sThe dividing line is durable resource versus one-shot action. Anything that should still exist tomorrow — a network, a queue, a role — belongs in the declarative description, because you want it reviewed, diffable and destroyable by deleting a line. Actions that happen once and have no "already true" condition — trigger a restore, drain a node, rotate a credential, backfill a table, force a failover — have no natural declarative form; wrapping them in a resource only fakes it. Scripting also wins for investigation, for bulk operations over hundreds of existing objects, for gluing the declarative runs together in a pipeline, and for the bootstrap chicken-and-egg where the tooling's own prerequisites do not yet exist. The discipline is that those scripts get treated as production code: in the repository, reviewed, idempotent or explicitly guarded, parameterised rather than edited before each run, and producing a log of who ran what. An unversioned script on somebody's laptop is not less risky than a console click — it is the same risk with more reach.
go deeper
Know that not everything belongs in the declarative code, and give one example of each side: a network belongs in the configuration, a one-off data migration does not.
State the boundary clearly — durable resources are declared, one-shot actions are scripted — and explain why an event has no desired state for the engine to converge on.
Show the operational discipline you would attach: scripts in the repository, reviewed, parameterised, dry-run capable, run through the pipeline identity. Name the bootstrap chicken-and-egg and the provider-gap case as legitimate exceptions with an expiry.
Own where the organisation draws and enforces the line: which classes of action are allowed outside the declarative flow at all, who may run them, what audit evidence is required, and how you stop the exception list from quietly growing into a second, unreviewed control plane.
## The line: durable resource versus one-shot action The question behind "should this be declared or scripted?" is not "which tool do I like" but **does this thing have a desired state that persists?** A VPC, a queue, a role, a DNS record — these should still exist tomorrow, in a known shape. They belong in the declarative description, because you want three specific properties: someone can read the file and know what exists; a change is previewable and reviewable; and deleting the declaration removes the thing. A restore, a node drain, a credential rotation, a data backfill, a forced failover — these are *events*. There is no "already true" condition to assert, so there is nothing for a declarative engine to converge on. Modelling them as resources produces the well-known anti-pattern: a resource that exists only to record that a script was run once, which then has to be tricked into re-running with a fake input change. ## Where scripting genuinely wins **Operational, day-two actions.** Everything in the events list above. These are procedures, and procedures are imperative by nature. **Investigation.** Answering "how many of these do we have, and which ones are misconfigured?" across an estate is a read-only loop. Nobody should be writing a resource declaration to find something out. **Bulk operations over existing objects.** Tagging ten thousand objects, migrating data between buckets, walking every account to check one setting. The declarative tool would need each object modelled; a loop needs none. **Bootstrapping the tooling itself.** The classic chicken-and-egg: the declarative tool needs somewhere to keep its records and an identity to run as, and neither exists yet. Something has to create those prerequisites, and that something is usually a small script run once. **Gluing the pipeline together.** The wrapper that decides which directories to run in, assembles credentials, passes an approved change through to the apply step — that is orchestration, and it is imperative code by definition. **Anything the provider has not modelled.** Providers lag their own cloud's APIs. When a service is a month old, or a setting only exists on the API and not as a resource, a script may be the only route. Treat this one as temporary: revisit it, because the reason it is a script is expected to expire. ## Where scripting is the wrong answer The common failure is using a script because the declarative route is *awkward right now*: a stubborn resource, a change nobody wants to preview, a hurry. That trades a permanent loss of visibility for a temporary convenience. Once a durable resource is created imperatively, it is invisible to the tool — never reviewed, never diffed, never cleaned up with the rest of the estate — and the next person to read the repository will believe it does not exist. The other failure is the escape hatch that shells out *from inside* a declarative apply. It defeats the preview, since the engine cannot show what an arbitrary command will do; it defeats removal, since deleting the declaration cannot undo the command's effects; and it usually defeats idempotency too. Sometimes it is still the pragmatic choice, but say out loud what it costs. ## Keeping the scripts from becoming invisible operations The risk with imperative operational tooling is not that it is imperative — it is that it tends to live outside every control the team put around infrastructure change. Treat it as production code: - **In the repository, and reviewed.** A script that touches production gets the same review as the declarative code that does. Being short is not an exemption. - **Parameterised, not edited.** A script whose account or resource id is edited in place before each run is a script that will one day be run against the wrong environment because someone forgot to change a line back. - **Idempotent or explicitly guarded.** If it can be safely re-run, say so and make it true. If it truly cannot, make it refuse to run twice — confirmation prompt, existence check, or a recorded marker. - **Dry-run first.** Provide a mode that prints what it would do. This is the poor engineer's plan step, and it recovers most of the review value the declarative flow gives you for free. - **Least privilege and an audit trail.** Run it through the same identity path as the pipeline where you can, so there is a record of who did what. A long-lived personal admin key used to run a helper script is a bigger finding than the script itself. - **Revisit the reason.** Keep a note of *why* it is a script. Provider gaps close; when the resource lands, migrate it and delete the script. ## How to say this in an interview The answer interviewers want is not "never script" — that reads as dogma from someone who has not run anything. It is the boundary plus the discipline: **declare the things that should exist, script the things that happen, and hold the scripts to the same review, repeatability and audit standard as the declarations.** Then name one thing you would genuinely never declare, and one thing you would genuinely never script, so it is clear the line comes from experience rather than a blog post.
- How do you keep a necessary operational script from drifting into unreviewed shadow tooling?Hold it to the same bar as the declarative code: committed and reviewed, parameterised rather than edited before each run, offering a dry-run mode, and executed through the pipeline's identity so there is an audit trail. Add a note recording why it is a script — provider gap, one-shot action — and revisit it, because gaps close and the script should then be deleted.
- Someone proposes modelling a data backfill as a resource so it runs through the normal apply. What do you say?That it fakes a desired state which does not exist. A backfill has no "already true" condition, so the resource ends up existing only to record that a script ran, and re-running it requires tricking the tool with a fake input change. Worse, the engine cannot preview what the command will do, so the apply's change summary quietly stops being trustworthy.
- What is the risk of creating a durable resource with a quick script because the declarative route is awkward?It becomes invisible. The tool does not know it exists, so it is never previewed, never reviewed, never updated with the rest of the estate, and never removed when the surrounding stack is torn down. Anyone reading the repository will conclude it does not exist. If you must do it, plan to bring it under management rather than leaving it orphaned.
saying these in an interview costs you the question
- Everything must be declarative; scripts are always a smell
- A script is fine if it is short and only run manually
- One-off actions should be modelled as resources so they are tracked
- If the tool has no resource for it, it cannot be automated
- Glue scripts do not need review because they only call other tools