In Amazon ECS, explain what a task definition, a task, and a service each are, and when you would use the RunTask API directly instead of putting the workload behind a service.
answer
- blueprint, instance, supervisor
- family plus revision number
- desired count and reconciliation
- finite work versus long-lived work
- RunTask returns and forgets
basics
~20 sAn ECS task definition is the immutable blueprint: image, CPU and memory, IAM roles, logging. A task is one running copy of it. A service keeps a desired count of tasks alive and replaces failures; RunTask launches one-off jobs.
solid answer
~50 sECS separates the *what* from the *how many*. A **task definition** is a versioned, immutable document (a family plus a revision number) that lists container definitions — image, CPU and memory, port mappings, environment and `secrets`, the `taskRoleArn` and `executionRoleArn`, the `networkMode`, and `logConfiguration`. A **task** is one running instantiation of a specific revision; when it stops, it is gone. A **service** is a controller that holds a desired count of tasks from one task definition, replaces any that stop, optionally registers them into an ALB or NLB target group, and performs rolling deployments when you point it at a new revision. You use `RunTask` directly when the work is finite and should not be restarted — a nightly ETL job, a database migration, a one-off admin command — typically fired by EventBridge Scheduler or Step Functions. Long-lived request-serving processes belong in a service.
go deeper
Be able to say the three nouns in one breath: definition is the blueprint, task is a running copy, service keeps copies running. Know that revisions are immutable and numbered.
Explain the service reconciliation loop, what happens when an essential container exits, and how a service registers tasks into a load balancer target group. Name the fields you would actually set in a task definition.
Show judgment about which workloads belong in a service versus RunTask, how you schedule and observe one-off tasks, and how immutable revisions make rollback a configuration change rather than a rebuild.
Own the platform convention: how teams version task definition families, where the deploy boundary sits, and whether batch work should be ECS tasks, Step Functions, or a different compute model entirely.
## The three objects ECS has an unusually small object model, and most confusion comes from collapsing its three core nouns into one. They are: the task definition (a document), the task (a running thing), and the service (a controller that keeps running things running). ## The task definition A task definition is registered with `RegisterTaskDefinition` and belongs to a **family**. Every registration produces a new **revision**, so `my-api:7` names one exact, immutable document. You never edit revision 7; you register revision 8. That immutability is what makes ECS deployments and rollbacks trivial — a rollback is just pointing the service back at the older revision. Inside it, the important fields are: - `containerDefinitions` — an array. Each entry has `name`, `image`, `portMappings`, `environment`, `secrets` (values pulled from Secrets Manager or SSM Parameter Store at start time), `essential`, `dependsOn`, `healthCheck`, `stopTimeout`, and `logConfiguration`. - `cpu` and `memory` at the task level — required for Fargate, which only accepts specific valid CPU/memory combinations. - `networkMode` — `awsvpc`, `bridge`, `host` or `none`. Fargate requires `awsvpc`. - `requiresCompatibilities` — `["FARGATE"]`, `["EC2"]`, or both. - `taskRoleArn` and `executionRoleArn` — two different IAM roles with two different consumers. - `logConfiguration` — usually `{"logDriver": "awslogs", "options": {"awslogs-group": ..., "awslogs-region": ..., "awslogs-stream-prefix": ...}}`. Note what is *not* in it: how many copies to run, which subnets, which load balancer. Those are properties of the service or of the `RunTask` call, not of the blueprint. That split is deliberate — the same revision can run as one debug task and as forty service tasks. ## A task A task is one instantiation of one revision. All containers in a task are placed together on the same host (or, on Fargate, in the same isolated micro-VM), share the task's network namespace under `awsvpc`, and can share volumes. A task moves through `PROVISIONING`, `PENDING`, `RUNNING`, `DEACTIVATING`, `STOPPING`, `DEPROVISIONING`, `STOPPED`. When a container marked `essential: true` exits, the whole task stops. Nothing restarts a bare task — that is the service's job. ## The service `CreateService` builds a controller with a `desiredCount`. It reconciles continuously: if a task stops for any reason — the process crashed, the host died, the container health check failed — the service starts a replacement. It also owns: - **Load balancer registration.** The `loadBalancers` block names a `targetGroupArn`, `containerName` and `containerPort`; ECS registers and deregisters task IPs (or host ports) in that target group as tasks come and go. - **Deployments.** Update the service to a new task definition revision and ECS runs a rolling replacement bounded by `minimumHealthyPercent` and `maximumPercent`. - **Scaling.** Application Auto Scaling attaches to the service's `desiredCount` and can target-track on predefined metrics such as `ECSServiceAverageCPUUtilization` or `ALBRequestCountPerTarget`. - **Scheduling strategy.** `REPLICA` (run N copies, spread by placement strategy) is the default. `DAEMON` runs exactly one task per eligible container instance and is only available on the EC2 launch type. ## RunTask: the one-off path ```bash aws ecs run-task \ --cluster batch \ --task-definition nightly-etl:12 \ --launch-type FARGATE \ --network-configuration 'awsvpcConfiguration={subnets=[subnet-abc],securityGroups=[sg-123],assignPublicIp=DISABLED}' ``` `RunTask` starts a task and returns. Nothing supervises it; if it exits non-zero, it stays stopped and you read the reason from `DescribeTasks` and the logs. That is exactly what you want for finite work: a schema migration, a report generation, a data backfill. Wire it to a schedule with EventBridge Scheduler, or into a workflow with Step Functions, which can wait for the task to finish and branch on its exit status. The rule of thumb: if the correct response to the process exiting is "start it again", use a service. If the correct response is "record the outcome", use `RunTask`. ## Where candidates go wrong The classic mistake is treating the task definition as mutable configuration and trying to "change the image" in place — there is no in-place change, only a new revision. The second is running batch work as a service with `desiredCount: 1`, which turns every successful completion into a restart loop. The third is expecting a bare `RunTask` to be restarted after a host failure; only a service reconciles.
- Your service is running revision 12 and the new revision 13 is broken. What is the fastest way back?Update the service to point at revision 12 again — revisions are immutable, so the old blueprint is still there byte-for-byte and a normal rolling deployment brings it back. Better still, enable the deployment circuit breaker with `rollback` set to true so ECS detects a deployment whose tasks never stabilise and reverts automatically without a human in the loop.
- What does marking a container `essential: false` in a task definition change?A non-essential container can exit without stopping the task; the other containers keep running. It is how you model a helper that legitimately finishes — a config fetcher or a one-shot sidecar. At least one container in the task must be essential, and when an essential container exits ECS stops every other container in that task.
- When would you choose the DAEMON scheduling strategy for a service?For node-level agents that must run exactly one copy per container instance — a log or metrics collector, a security agent. ECS then starts a task on every instance that satisfies the placement constraints and on any new instance that joins. It is only available on the EC2 launch type, because Fargate has no instances you own to place one per.
saying these in an interview costs you the question
- Says a task definition can be edited in place
- Runs finite batch jobs as a service with desiredCount 1
- Thinks a bare RunTask task is restarted after a crash
- Confuses a task with a single container
- Believes the subnet list lives in the task definition