skip to content

A nightly ETL job runs for about 40 minutes and needs roughly 8 GB of memory. A colleague proposes running it on AWS Lambda. What is your objection, and which AWS compute option would you choose instead?

level: middleimportance: must knowfreq 68%

answer

  1. memory fits, duration does not
  2. hard ceiling, not a soft quota
  3. killed mid-run leaves work half-applied
  4. run-to-completion task instead
  5. or partition the work into chunks

basics

~20 s

A Lambda function cannot run for 40 minutes: the maximum timeout is 15 minutes, so the job would be killed mid-run. Memory is fine, duration is not. Run it as a scheduled container task, on AWS Batch, or split it into shorter chunks.

solid answer

~50 s

The memory is not the problem — Lambda functions can be configured up to 10,240 MB, so 8 GB fits. The duration is the problem: a function's maximum timeout is 15 minutes (as of 2025), and a 40-minute run would simply be terminated part-way through, most likely leaving the ETL half-applied. I would either move the job to a compute model with no such ceiling — a scheduled ECS task on Fargate, an AWS Batch job, or an EC2 instance that runs and shuts down — or, if there is a good reason to stay serverless, restructure the work into chunks that each finish well inside the limit and drive them with a state machine or a queue. The first option is usually the cheaper decision, because rewriting a working batch job to fit a platform limit is real engineering effort spent on nothing the business asked for.

go deeper

for a junior

Know that a Lambda function has a maximum runtime measured in minutes and that long batch work belongs somewhere else; be able to say the memory was never the blocker.

for a middle

State the 15-minute ceiling precisely, explain that it is a hard product limit, and name a concrete run-to-completion alternative such as a scheduled container task or AWS Batch.

for a senior

Bring up what happens when a long job dies part-way, and design for resumability or atomic output so the compute choice is no longer a correctness risk.

for a principal

Weigh a rewrite-to-fit against adding another supported compute model to the platform, and be explicit that engineering time spent contorting a job around a limit is a cost the business is paying for uniformity.

## Separate the two limits The proposal fails on exactly one dimension, and saying which one shows you know the platform rather than a slogan. **Memory: fine.** A Lambda function's memory is configurable from 128 MB up to 10,240 MB (as of 2025), and CPU is allocated in proportion, so an 8 GB configuration is legal and would also give the job a healthy CPU share. **Duration: fatal.** The maximum configurable function timeout is 15 minutes. This is not a soft limit you can raise with a support ticket; it is the shape of the product. At the timeout the invocation is terminated. For a batch job, being killed 15 minutes into a 40-minute run is worse than not starting, because the job is now half-applied and the next night's run inherits whatever mess it left. ## Choosing the replacement Once Lambda is out, the question is which long-running model fits. **A scheduled container task** is the common answer. You already need a container image, or can build one trivially; the task runs to completion and stops, so you pay only for the ~40 minutes it runs, and there is no instance to keep alive between nights. This keeps the serverless *billing* shape — pay for what runs — while removing the duration ceiling. EventBridge Scheduler can start the task on a cron expression. **AWS Batch** is the answer when the job is one of many, or when it needs queuing, retry, dependency ordering, array jobs, or a mix of instance types. Batch can run on Fargate or on EC2 capacity, and it exists precisely for the "submit work, let something else find capacity" pattern. If the nightly ETL is really twenty nightly ETLs with dependencies, this is the better fit. **EC2** earns the job when it needs something only a host provides — a GPU, very large memory, local NVMe scratch space, or licensed software. A plain instance that boots, runs the job, and terminates is a perfectly respectable batch platform. ## The other path: make the work fit Sometimes staying on Lambda is right — for example the organisation has one paved road, and it is functions. Then you do not fight the timeout, you decompose the job: - **Partition the data.** If the 40 minutes is a loop over a million rows, run N functions over N slices in parallel. Total wall-clock time drops as well as fitting each invocation inside the limit. This works only if the slices are independent. - **Checkpoint and continue.** Each invocation processes as much as it can, records where it stopped, and hands the cursor to the next invocation. A state machine driving a loop with the cursor in its state is the clean form of this. - **Move the heavy lifting into a service that does the work for you.** A lot of "40 minutes of ETL" is really "a large query and a large write", which the data store can often do without your process staying alive. All three cost engineering time and add failure modes — partial progress, retries, idempotency. That is the real tradeoff: you are trading a rewrite for platform uniformity. ## The idempotency point that earns credit Whichever way you go, an interviewer will be pleased if you volunteer that a batch job must be safe to re-run. Long jobs get killed for reasons other than timeouts — a deploy, a Spot reclamation, a host failure — so the design question is not only "can it run for 40 minutes" but "what happens when it dies at minute 30?" Writing to a staging location and swapping atomically at the end, or making each unit of work idempotent, is what makes the choice of compute model less consequential. ## How to answer this out loud Name the disqualifier precisely (duration, not memory), name a concrete replacement with a reason (scheduled container task for a single job, Batch for a fleet of them, EC2 when the host itself matters), and then name the condition under which you would keep Lambda anyway (the work partitions cleanly, or your organisation only supports functions). That structure — disqualify, replace, caveat — is what the question is testing.

  • The team says they will just retry the function until it finishes. Why does that not work?
    Retrying restarts the invocation from the beginning, so a job that needs 40 minutes of work never gets past the first 15 unless it resumes from a checkpoint. Without stored progress, every attempt does the same first chunk again and the job never completes.
  • How would you make a 40-minute job safe against being killed at minute 30, whatever compute you run it on?
    Make it resumable or atomic: write output to a staging location and swap it in only on success, or checkpoint progress so a rerun continues rather than restarts. Then a kill is merely wasted time rather than corrupted data, and the compute choice stops being a correctness question.
  • When would AWS Batch beat a single scheduled container task?
    When there are many jobs rather than one — you want a queue, dependencies between jobs, array jobs over a partitioned dataset, automatic retries, or a mix of instance sizes chosen per job. For one nightly task on a fixed schedule, Batch is more machinery than the problem needs.

saying these in an interview costs you the question

  • Says the 15-minute Lambda timeout can be raised via a support request
  • Claims 8 GB exceeds what a Lambda function can be configured with
  • Proposes keeping the function alive by pinging it repeatedly
  • Ignores what happens to data when the job is killed part-way
  • Reaches for a permanently running EC2 instance for a nightly job

context