skip to content

questions

4

How does a nightly job prove which workload it is to a data service outside the cluster, with no stored password?

level: middleimportance: must knowfreq 58%

answer

  1. no shared secret exists to steal
  2. the platform vouches, the workload cannot
  3. identity comes from the workload spec
  4. one audience, expiry in minutes
  5. presented, verified, then traded for access

basics

~20 s

The platform mints the job a short-lived token that names the workload itself, addressed to one audience and valid for minutes. The job presents it to the far service, which verifies the issuer's signature and grants only what that workload's grant allows.

solid answer

~50 s

Nothing is baked into the image. When the job starts, the platform hands it a freshly minted token — normally a file it rewrites underneath the workload before expiry — saying *which workload this is*: the identity declared in the workload spec, the boundary it runs in, the issuer that signed it, an `audience` naming the one service it is addressed to, and an `expiry` a few minutes out. The job re-reads that file at the point of use and presents the token to the data service, or to a broker in front of it. The far side verifies the signature against the issuer's published verification keys, checks the audience names itself and the token is still valid, maps the identity onto a grant it holds, and either serves the call or returns a short-lived credential of its own kind. No shared secret ever exists to leak.

code

pseudocode · 13 lines
pseudocode
on workload start:
    token_path = platform_issued_token(identity = "nightly-reconciler",
                                      audience = "ledger-store",
                                      lifetime = 10 minutes)
    # the platform rewrites this file well before each expiry

function read_ledger(query):
    token = read(token_path)              # re-read, never cache for the process lifetime
    grant = ledger_broker.exchange(token) # broker checks signature, audience, expiry
    if grant.rejected == "expired":
        token = read(token_path)          # a refreshed copy is already on disk
        grant = ledger_broker.exchange(token)
    return ledger.read(query, grant.credential)   # this credential expires too

go deeper

for a junior

Recall that the job holds no stored password. The platform hands it a token at start-up that names the workload, is addressed to one service, and stops being valid after a few minutes.

for a middle

Explain what is inside the token — who signed it, which workload it names, which single service may accept it, when it dies — and who checks each of those. Say clearly that identity is not permission.

for a senior

Show you have operated it: the token is a file refreshed underneath a long-running process, so read it at the point of use, retry once on an expiry rejection, and expect the far side to trade it for a credential of its own with a different clock.

for a principal

Weigh what moves when identity becomes platform-issued. Grants at the far side are now written against names your platform hands out, so who may declare a workload with a given identity name becomes an access-control decision you own.

## The credential that is not there A workload that must reach a service outside its own cluster has to prove who it is. The old answer is a **static key**: a long-lived value created once by a person, pasted into a configuration store or a build, and handed to the workload. It has no natural end, it is identical everywhere it is copied, and nothing in it says which workload is using it. Most incident reports that begin *the key was in the repository* begin here. **Workload identity replaces the secret with a statement.** Rather than the workload holding something only it should know, the platform — which already knows exactly what it started, because it started it — attests to the workload for whoever is asking. ## What the platform issues At start-up, and repeatedly afterwards, the platform mints a token and puts it where that one container can read it, usually as a file mounted into it alone. The token carries: - **an identity name** — the identity declared for this workload in its spec, not the machine and not the person who deployed it; - **the boundary it runs in**, so the same name in two tenancy boundaries is two different identities; - **the issuer**, which names the component that signed it and therefore whose keys a verifier must check; - **an audience** — the single service the token is addressed to; - **an expiry**, typically minutes rather than hours; - **a signature** over all of it, made with a key the platform holds and never gives to the workload. Two negatives matter as much as the contents. The workload cannot mint its own token, because it has no signing key. And the token is not a permission: it says *who*, never *what may be done*. ## Presenting it, and trading it 1. The job reads the token at the point of use, not once at start-up — the platform rewrites the file ahead of expiry, so a copy held in memory for hours is a dead token. 2. It presents the token to the far service, or to a credential broker sitting in front of it. 3. The far side verifies the signature against verification keys published by the issuer it has been configured to trust, checks that the audience names itself, and checks the token has not expired. 4. It maps the identity name onto a **grant** it holds — *this identity may read these tables* — and either serves the call directly or returns a short-lived credential of its own kind, which the job then uses. Step 4 is the trade, and it is the step most descriptions skip. The platform's token is proof of identity issued by a system the far side does not operate; many far services will not act on a foreign token directly, so they exchange it for their own credential and everything after that looks like an ordinary authenticated caller to them. What comes back expires too, on its own clock. ## Static key against issued identity | | static key | platform-issued workload token | |---|---|---| | where it comes from | created once, by a person | minted per workload, continuously | | what it names | nothing; it is only a value | the workload's own identity | | lifetime | until somebody rotates it | minutes | | where copies rest | a store, a spec, an image layer, a log | one file inside one container | | rotation | a scheduled human task | happens by construction | | what a stolen copy is worth | everything it grants, indefinitely | the remaining minutes, at one audience | ## Where designs differ Some platforms issue a signed token; others issue a short-lived certificate and let the far side authenticate the connection itself. Some far services accept the platform's token on the call; others insist on an exchange first. The shape underneath is the same everywhere: a component that knows what it started attests to it, on a clock, for one destination. ## What this does not fix - **It is not authorization.** The grant at the far side still has to be narrow. An identity with a sweeping grant is a sweeping credential with better logging. - **It is not credential delivery.** A partner's key, or a password for a service that cannot federate at all, still has to reach the process some other way; that is a different mechanism. - **It does not survive a weak identity model.** If any team can declare a workload carrying any identity name, the far side's grants are written against names your platform lets anyone claim.

  • The data service will not act on the platform's token directly — what happens instead?
    The job presents the platform token to a broker on the far side that has been configured to trust the issuer. The broker verifies the signature, the audience and the expiry, maps the identity name onto a grant, and returns a short-lived credential of its own kind. The platform token is the proof of identity; what comes back is what actually opens the door, and it carries its own expiry.
  • The job runs for six hours — why must it re-read the token file rather than hold the first copy?
    Because the token lives for minutes. The platform refreshes the file in place ahead of expiry, so a process that read it once at start-up is presenting a dead token by its second call. Read at the point of use, and treat an expiry rejection as a signal to reload and retry rather than to fail the job.
  • Why does the token name a single audience rather than being usable anywhere?
    Because a token works for whoever holds it. Addressed to one service, it is useless at a second one that trusts the same issuer — that verifier rejects it as not addressed here, even though the signature is good and the token is fresh. Without that, the first service you call can turn round and present your token to another as you, for the rest of its window.

A courier without a company card, carrying instead a note the dispatcher rewrites every ten minutes: this is our courier, for the north depot, until quarter past.

saying these in an interview costs you the question

  • Thinks the platform stores a password for the workload somewhere
  • Says the workload generates and signs its own identity token
  • Treats the token as proof of permission rather than identity
  • Assumes any service that trusts the issuer may accept the token
  • Reads the token once at start-up and caches it for hours
  • Believes a short expiry means the token needs no protection in transit
open as a page

Why does a credential attached to the node grant far more than one issued to a single workload?

level: seniorimportance: must knowfreq 50%

basics

~20 s

A node credential belongs to the machine, so its grant must cover everything any workload placed there needs, and anything running on that host can reach it. A per-workload token names one workload and carries only that workload's access.

open as a page

An attacker captures a workload's ten-minute token — what does the short lifetime actually limit, and what does it not?

level: seniorimportance: should knowfreq 36%

basics

~20 s

A short lifetime kills the whole class of credentials found later in an image, a log or a backup, because a captured copy is already dead. It does not limit use inside the window, anything traded for it, or an attacker still executing in the workload and reading each refreshed token.

open as a page

Making every outward call depend on a platform token issuer adds a failure domain — how do you design for it?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Design around the one property that makes it survivable: held tokens stay valid to their own expiry, so an issuer outage degrades over one lifetime and hits newly started workloads first. Token lifetime is the lever, traded directly against a stolen token's window.

open as a page