skip to content

A long-running batch job assumes an IAM role and then fails with an ExpiredToken error about an hour in, even though the role's MaxSessionDuration is set to 12 hours. What would you check?

level: seniorimportance: nice to knowfreq 32%

answer

  1. a suspiciously round one hour
  2. ceiling is not the default
  3. chaining caps it regardless
  4. snapshot credentials cannot renew
  5. let the SDK re-assume

basics

~20 s

Check whether the job requested a longer DurationSeconds at all, whether it is role chaining — assuming a second role from the first role's credentials caps the session at one hour — and whether the credentials were captured once into environment variables so nothing can refresh them.

solid answer

~50 s

Three causes cover almost every case. First, `MaxSessionDuration` is a ceiling, not a default: unless the caller passes `DurationSeconds`, the session is one hour regardless of the role's 12-hour setting. Second, **role chaining** — using one role's temporary credentials to assume another role — caps the resulting session at one hour no matter what either role's maximum says, and asking for more than an hour on a chained call is rejected outright. Third, and most common in practice, the job grabbed credentials once at startup and exported them into environment variables or a file; the SDK then has no way to re-assume, so it rides the same session until it dies. The right fix is usually not a longer session but a refreshable one: configure the profile with `role_arn` and `source_profile`, or let the SDK's own provider handle the assumption, so it re-assumes before expiry. For work that genuinely outlives any session, make the job checkpoint and re-assume rather than betting on a long token.

code

ini · 9 lines
ini
[profile source]
region = eu-west-1

; The SDK owns the assumption and re-assumes as expiry approaches
[profile batch]
role_arn = arn:aws:iam::111122223333:role/Batch
source_profile = source
role_session_name = nightly-report
duration_seconds = 3600

go deeper

for a junior

Recognise ExpiredToken as "the temporary credentials ran out" and know that a role session has a limited life which somebody has to renew.

for a middle

Explain the mechanics: MaxSessionDuration is a ceiling, DurationSeconds is the request, the permitted range starts at 15 minutes, and a snapshot of credentials in environment variables cannot refresh itself.

for a senior

Diagnose from the one-hour signature — spot role chaining in a profile or wrapper, distinguish it from a missing duration request, and argue for refreshable credentials and restartable jobs over longer sessions.

for a principal

Own the standard: what session lengths are permitted and why, when a 12-hour role is justified, and how long-running work is architected so credential expiry is never the thing that ends a job.

## The symptom `ExpiredToken` (you may also see `ExpiredTokenException`, depending on the service) means the session token presented is past its expiration. The clue in this scenario is the *timing*: almost exactly one hour, which is the STS default and also the role-chaining ceiling. When a failure lands on a round number, chase the round number. ## Cause 1: nobody asked for longer `MaxSessionDuration` is an attribute of the role, settable between 1 and 12 hours, and it is a **maximum**, not a default. If the caller does not pass `DurationSeconds` on `sts:AssumeRole`, it gets a one-hour session. Raising the role's maximum to 12 hours and changing nothing else changes nothing at all — a very common false fix. ```bash aws iam update-role --role-name Batch --max-session-duration 43200 # the ceiling aws sts assume-role --role-arn arn:aws:iam::111122223333:role/Batch \ --role-session-name nightly --duration-seconds 43200 # the request ``` Both halves are required. The permitted request range starts at 900 seconds and runs up to the role's own maximum; ask for more than the role allows and the call fails rather than silently truncating. ## Cause 2: role chaining **Role chaining** is using one role's temporary credentials to assume a second role. Common shapes: a pipeline assumes a hub role and then a spoke role per account; a profile with `role_arn` whose `source_profile` is itself a role profile; a job that assumes a cross-account role and then a narrower one inside it. A chained session is limited to **one hour**, whatever `MaxSessionDuration` says on either role, and a chained `AssumeRole` call requesting more than an hour is rejected. This is the classic explanation for exactly-one-hour failures in a system where somebody has already, confidently, raised the maximum to twelve. The chain is often invisible in the code, hiding in a `~/.aws/config` layout or in a wrapper script someone else wrote — reading the profile chain is the fastest way to spot it. ## Cause 3: credentials frozen at start Even when a session is long enough in principle, a job dies at the boundary if nothing can renew it. The pattern that causes it: ```bash # anti-pattern: a snapshot that can never refresh eval $(aws sts assume-role --role-arn ... --role-session-name job | jq -r '...export...') run-the-eight-hour-job ``` Those exported environment variables are a snapshot. The SDK inside the job sees static credentials and has no idea they came from an assumption, no role ARN to re-assume, and no reason to try. It signs happily until the expiry passes and then every call fails at once. The fix is to let the SDK own the assumption. In `~/.aws/config`, a profile with `role_arn` plus `source_profile` (or a web-identity token file, or the environment's own provider) makes the SDK re-assume automatically as expiry approaches. Then a long job is fine even with one-hour sessions, because it is issued a fresh one every hour. ## Cause 4: federated session ceilings If the credentials came from federation rather than a plain `AssumeRole`, other ceilings apply. A SAML session created by `sts:AssumeRoleWithSAML` is additionally bounded by the duration the identity provider asserts, so a 12-hour role can still yield a short session because the IdP said so. Sessions obtained through a federated login portal have their own limits set by the identity source. Check where the session actually came from before blaming the role. ## The design answer The instinct to reach for a longer session is usually the wrong one, and interviewers listen for that. Long sessions widen the window a stolen credential is useful for, and they still fail eventually — a 12-hour ceiling only moves the cliff. The durable answers are: - **Make refresh work.** Configure credentials so the SDK can re-assume, rather than passing a snapshot into the process. - **Make the job restartable.** Checkpoint progress so a long run can resume; a batch job that cannot survive losing its credentials usually cannot survive losing its instance either, and that is the deeper problem. - **Break up the work.** Long single-process jobs are a poor fit for temporary credentials by design; chunking the work into tasks that each assume fresh credentials removes the class of failure. - **Reserve the 12-hour maximum** for cases with a real justification, and treat routine requests to raise it as a smell that something is holding credentials it should be refreshing.

  • How would you confirm quickly that role chaining is what is capping the session?
    Trace how the credentials are produced: read the profile chain in `~/.aws/config` for a `role_arn` whose `source_profile` is itself a role profile, or look for a second `AssumeRole` call in the job's own code or wrapper. A direct test is to request more than an hour on that call — a chained assumption is rejected rather than truncated.
  • Why is raising MaxSessionDuration to 12 hours a poor default response to this failure?
    It often fixes nothing, because the caller still has to request the longer duration and a chained session ignores the ceiling anyway. Even when it works it only moves the cliff and widens the window in which a stolen credential is usable. The better fix is making refresh work and making the job restartable.
  • The job runs in a container that assumes a role. What is the cleanest way to keep its credentials fresh?
    Do not hand the process a credential snapshot. Let the SDK resolve credentials itself from the environment it is given so it can re-fetch and re-assume as expiry approaches, and keep the assumption configuration — role ARN and source — visible to the SDK rather than resolved once by a wrapper script that exports three variables.

saying these in an interview costs you the question

  • Assuming MaxSessionDuration is the default session length
  • Not knowing chained sessions are capped at one hour
  • Exporting assumed credentials once and expecting them to renew
  • Answering only "request a longer duration"
  • Blaming clock skew before checking how credentials are obtained

context