skip to content

AWS App Runner takes a container image or a source repository and hands back an HTTPS URL. What does it create on your behalf, what does it scale on, and how are you billed when no requests are arriving?

level: middleimportance: should knowfreq 42%

answer

  1. no cluster, no load balancer, no certificate
  2. scales on concurrent requests, not CPU
  3. three numbers: concurrency, min, max
  4. warm instances still cost memory
  5. private resources need a VPC connector

basics

~20 s

App Runner runs your container on AWS-managed capacity behind a managed HTTPS endpoint, with no cluster, load balancer or certificate of yours. It scales on concurrent requests per instance, and bills memory continuously for warm instances but CPU only while requests are being served.

solid answer

~50 s

You give App Runner a source — a container image in ECR, or a connected GitHub or Bitbucket repository that App Runner builds for you using a managed runtime and an optional `apprunner.yaml` — plus CPU, memory, the port to listen on, environment variables and a health check. It creates nothing in your VPC: the endpoint, TLS certificate and routing are AWS-managed, and you get a URL on `awsapprunner.com`. Scaling is driven by an auto scaling configuration of three numbers — max concurrency (requests in flight per instance), min size and max size. When in-flight requests per instance exceed max concurrency, App Runner adds instances up to max size. Min-size instances stay warm so requests avoid a cold start, and you pay the **memory** rate for them all the time; **CPU** is billed only while an instance is actively processing a request. Reaching private resources such as an RDS database requires a VPC connector.

code

yaml · 13 lines
yaml
version: 1.0
runtime: nodejs18
build:
  commands:
    build:
      - npm ci
run:
  command: node server.js
  network:
    port: 8080
  env:
    - name: APP_ENV
      value: production

go deeper

for a junior

Know that App Runner takes an image or a source repo and returns a working HTTPS URL, with no cluster or load balancer for you to create.

for a middle

Explain the concurrency-based scaling knobs — max concurrency, min size, max size — and the split bill where memory is charged for warm instances and CPU only during requests.

for a senior

Show the operating consequences: choose max concurrency from the app's threading model, add a VPC connector for private dependencies and remember NAT for public egress, and move background work out of the container.

for a principal

Own the boundary question — which services are simple enough that giving up load-balancer and sidecar control is a good trade, and what the exit looks like when a team outgrows a single-container HTTP model.

## What App Runner replaces App Runner is the most opinionated container service AWS sells. On a typical container deployment you would create a cluster or capacity provider, a task or service definition, a load balancer, a target group, a listener, a certificate, security groups, and scaling policies. App Runner replaces all of that with one resource: a *service*. There is no cluster to size, no load balancer in your account, and no certificate to renew — the HTTPS endpoint on `*.awsapprunner.com` is provisioned with an AWS-managed certificate, and a custom domain can be attached with validation records. ## Two source types A service is created from either: - **an image**, from a private or public ECR repository. App Runner pulls it using an *access role* you grant, and can redeploy automatically when a new image is pushed to the tracked tag; - **source code**, from a GitHub or Bitbucket connection. App Runner builds it with a managed runtime (Python, Node.js, Java, .NET, Go, PHP, Ruby) and either the build/start commands you supply in the console or an `apprunner.yaml` in the repository root, and can auto-deploy on every push to the tracked branch. The container itself gets an **instance role** for the AWS API calls your application makes — the same separation ECS draws between pulling the image and running the code. ## The scaling model, and why it is not CPU-based An Auto Scaling group scales on a metric such as average CPU. App Runner scales on **concurrency**: the auto scaling configuration sets *max concurrency* — how many requests one instance handles simultaneously — plus *min size* and *max size*. When the number of in-flight requests would exceed max concurrency on the existing instances, App Runner starts another, up to max size; when load falls, it retires instances back down toward min size. This is the request-driven model, and it fits HTTP services whose cost is dominated by concurrent request handling. It fits badly when work is not request-shaped: a service whose load is a background queue drain looks idle to App Runner no matter how busy it is. Max concurrency is a real tuning knob. Set it to 1 and you get one request per instance — strong isolation for a thread-unsafe or heavily CPU-bound app, at the price of many instances and a much higher chance of hitting max size under a burst. Set it high and you pack more requests per instance, using CPU better but risking queueing inside your own process. ## The billing model, and what "idle" means App Runner splits the bill. **Provisioned memory** is billed for every instance that exists, including warm min-size instances with nothing to do. **CPU** is billed only while an instance is actively processing requests. So a service at zero traffic with min size 1 is not free, but it is cheap — you pay to keep a warm container ready, and warm is exactly what buys you the absence of cold starts. The non-obvious consequence is behavioural rather than financial: between requests, an instance's CPU is throttled. **Background threads, in-process schedulers and async work kicked off after the response is written are unreliable** — they may make almost no progress until the next request arrives. Anything periodic belongs outside the request path (for example a scheduled event that calls an endpoint), not in a timer thread inside the container. ```yaml # apprunner.yaml — source-code services version: 1.0 runtime: nodejs18 build: commands: build: - npm ci run: command: node server.js network: port: 8080 ``` ## Networking By default outbound traffic leaves over AWS-managed networking, which is fine for calling public AWS APIs but cannot reach anything in a private subnet. To talk to an RDS instance, an ElastiCache cluster or an internal service, you attach a **VPC connector**: you nominate subnets and security groups, App Runner creates network interfaces there, and the database's security group allows the connector's. Once egress is routed through the VPC, calls to the public internet need the usual route through a NAT gateway — the identical trap that catches functions placed in a VPC. Inbound can also be made private through an interface VPC endpoint, so the service is not reachable from the internet at all. ## Where it stops One container per service, one HTTP port, no sidecars, no direct control of the load balancer, no non-HTTP protocols. Those limits are the whole product: App Runner is for a stateless HTTP service you want online without operating a platform, and the moment you need a second container next to it, the abstraction is the wrong shape.

  • Your App Runner service must query an RDS database in a private subnet. What do you configure?
    A VPC connector: set the service's egress to VPC and nominate subnets and a security group, so App Runner places network interfaces in your VPC and the database's security group can allow that group. Two consequences follow — the subnets you choose need a NAT route if the app also calls the public internet, and the connector's interfaces consume IP addresses in those subnets.
  • What does setting App Runner's max concurrency to 1 actually do?
    Each instance handles one request at a time, so App Runner adds an instance for every additional concurrent request. That gives strong isolation for thread-unsafe or CPU-bound code, but it multiplies instance count, raises cost, and makes hitting max size under a burst far more likely. Higher values pack more requests per instance and use each container's CPU better.
  • Why is an in-process scheduler a bad idea inside an App Runner container?
    Because CPU is only guaranteed while the instance is handling a request; between requests it is throttled, so timers and background threads may barely advance. Move periodic work out of the container — have a scheduled event invoke an endpoint, or run the job on a compute model that is billed and scheduled for continuous work.

saying these in an interview costs you the question

  • Says App Runner scales on CPU utilisation like an Auto Scaling group
  • Assumes a service with no traffic costs nothing
  • Expects background threads to run between requests
  • Thinks it can reach a private database with no VPC connector
  • Believes you must supply your own load balancer and certificate

context