You operate one private container registry (Harbor or a cloud registry) for about a dozen product teams. Design the access model: who may push, who may pull, and how build pipelines and production hosts obtain credentials.
answer
- Namespace per team = permission boundary
- Only pipelines push; humans read
- Robot accounts / IAM with expiry
- Release namespace: immutable tags, deploy by digest
- Workload identity > stored secret
basics
~20 sNamespace repositories per team, grant humans read-only plus break-glass, and give push rights only to machine identities scoped to their own namespace. Production pulls read-only, ideally via workload identity rather than stored secrets. Separate promotion namespaces, make release tags immutable, and set credential expiry and audit.
solid answer
~60 sI start from the principle that **only pipelines push, everything else pulls**. - **Namespacing.** One project/namespace per team (`team-payments/*`), so every grant is expressible as a repository prefix. - **Identities.** Machine identities distinct from humans: Harbor robot accounts scoped to a project with an expiry, or cloud IAM principals (ECR repository policies plus IAM, Artifact Registry per-repository IAM roles). Humans get pull on their own namespace and on shared bases; nobody pushes interactively to a release namespace. - **Promotion.** Separate `staging/` and `release/` (or separate registries). Only the promotion pipeline can push to release, and release tags are immutable so a tag cannot be re-pointed after approval. - **Credential delivery.** No static registry passwords on long-lived machines: credential helpers backed by workload identity on nodes and runners; short-lived job tokens in ephemeral CI. - **Guardrails.** Retention and garbage collection, vulnerability scanning gates, a pull-through cache for upstream images, and audit logs on push and delete. The test I apply: if one runner is compromised, what can it overwrite? The answer should be "one team's staging namespace, until the token expires."
code
json · 23 lines{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "CIPush",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::111122223333:role/ci-payments" },
"Action": [
"ecr:InitiateLayerUpload",
"ecr:UploadLayerPart",
"ecr:CompleteLayerUpload",
"ecr:PutImage",
"ecr:BatchCheckLayerAvailability"
]
},
{
"Sid": "RuntimePull",
"Effect": "Allow",
"Principal": { "AWS": "arn:aws:iam::444455556666:root" },
"Action": ["ecr:BatchGetImage", "ecr:GetDownloadUrlForLayer"]
}
]
}go deeper
Know the basic split: pipelines push, everything else pulls, and credentials differ per registry host.
Describe namespacing per team, separate machine identities with expiry, and read-only credentials on runtime hosts.
Add the promotion boundary, tag immutability, deployment by digest, and credential helpers backed by workload identity rather than stored secrets.
Lead with blast-radius reasoning and supply-chain provenance, trade off a single registry with prefixes against separate build/production registries, and cover break-glass, audit, retention and quota governance.
## Framing A registry access model is a supply-chain control. The artifacts that production runs enter the world here, so the questions are: which identity can introduce an artifact, which can *replace* one, and how are those identities proven. ## 1. Structure the namespace first Authorization is expressed over repository paths, so the path layout *is* the permission model. A workable default: ``` team-payments/api team-payments/worker platform/base-jdk21 release/payments-api cache/docker.io/... (pull-through mirror) ``` Harbor calls the top level a *project* and attaches members, robot accounts, quotas, scanning policy and immutability rules to it. ECR has flat repository names but supports prefix-based IAM resource patterns (`arn:aws:ecr:...:repository/team-payments/*`). Artifact Registry binds IAM per repository. All three let you say "this identity, this prefix, these actions" — provided the prefixes were designed for it. ## 2. Separate human and machine identities **Humans** get read access to their team's namespace and to shared base images, plus whatever the registry UI needs for triage. They should not hold push rights to anything production consumes; if a human can push a tag that a deployment resolves, your provenance story is "someone's laptop". **Machines** get purpose-built identities: - Harbor **robot accounts**, scoped to a project (or a set of repositories) with explicit actions (`push`, `pull`, `delete`) and an **expiry date**. Their secret is shown once and is not tied to a person who may leave. - **ECR**: the identity is an IAM principal. Pull needs `ecr:BatchGetImage` and `ecr:GetDownloadUrlForLayer`; push additionally needs `ecr:InitiateLayerUpload`, `ecr:UploadLayerPart`, `ecr:CompleteLayerUpload` and `ecr:PutImage`. Note that `ecr:GetAuthorizationToken` cannot be scoped to a repository — it is account-level — so per-repository control comes from the other actions and from repository policies. - **Artifact Registry / GCR**: `roles/artifactregistry.reader` versus `writer`, bound per repository to a service account or a workload identity. The asymmetry to insist on: **write is rare and narrow, read is broad and boring.** Production nodes and most CI jobs need pull only. ## 3. Model the promotion path Builds land in a team's own namespace. Promotion to something production may run is a distinct, audited action performed by a pipeline identity that is the *only* principal with push rights to the release namespace. Combine with: - **Tag immutability** on the release namespace, so an approved tag cannot be silently re-pointed at different content. - **Deploying by digest**, so what was approved is what runs. - **Retention plus garbage collection** on non-release namespaces, so storage growth does not force ad-hoc deletions that need broad delete rights. A blunt but effective variant is two registries: an internal build registry that many identities can write, and a production registry that only the promotion pipeline can write and that production can read. The trust boundary becomes a network and identity boundary, not just a path prefix. ## 4. Deliver credentials without distributing secrets Ranked best to worst: 1. **Workload identity + credential helper.** The node, task, or CI job authenticates as itself; a `docker-credential-*` helper mints a short-lived registry token per request. No secret at rest, rotation is automatic, revocation is a policy edit. 2. **Short-lived job tokens.** An ephemeral runner receives a token valid for the job's lifetime and scoped to the repositories it touches. 3. **Robot accounts with expiry**, delivered through a secret manager, for platforms without ambient identity. Rotation must be automated or the expiry becomes an outage. 4. **A shared static password in a config file.** This is the state to eliminate; it is unrotatable in practice and appears in image layers and logs. For clusters, image pull credentials are the orchestrator's own mechanism, but the same rule holds: prefer node/workload identity over a stored secret, and grant pull only. ## 5. Guardrails around the model - **Pull-through cache / mirror** for upstream public images, so build reliability does not depend on an external registry and every dependency is recorded internally. - **Vulnerability scanning** on push with policy gates on the promotion step, not on every developer push (which trains people to ignore it). - **Audit logging** of push, delete and permission changes, retained where it can be queried during an incident. - **Quotas** per project so one team cannot exhaust shared storage. - **Break-glass** path: a documented, time-boxed elevation with an audit trail, because the answer to "nobody can push manually" must not be "so we share the pipeline's credential". ## 6. How I would justify it The evaluation question is blast radius. Compromise a build runner: with the model above the attacker can publish into one team's staging namespace, cannot touch release, cannot alter an approved tag, and loses access when the token expires. Compromise a production node: it holds a pull-only workload identity for one namespace and can exfiltrate images it could already run. Those are acceptable failure modes; a single shared read-write account for everyone is not.
- A team asks for push rights so developers can hotfix an image directly. How do you respond?I would decline standing push rights to any namespace production reads, because it destroys the provenance guarantee that every running artifact came from a build pipeline. The need behind the request is speed, so I would make the pipeline fast and offer an expedited path with the same audit trail. If a genuine emergency case exists, it becomes a time-boxed break-glass grant that is logged and expires by itself.
- How do credentials for pulling upstream public base images fit into this model?I front them with a pull-through cache or mirror in the same registry, so builds resolve `cache/docker.io/...` internally. That removes a runtime dependency on an external service, gives one place to record and scan every third-party image, and keeps upstream account credentials in one configured location rather than spread across runners. It also sidesteps anonymous-pull throttling on shared egress addresses.
Think of the release namespace as a bonded warehouse: goods arrive freely in the yard, but only one licensed carrier moves them inside, and once sealed the crate label cannot be swapped.
saying these in an interview costs you the question
- One shared read-write account used by every pipeline and every node
- Giving production hosts push rights "in case we need to retag"
- Relying on mutable tags for release identity instead of digests
- Assuming ECR's GetAuthorizationToken can be scoped per repository
- Treating scanning as the control while anyone can still push into the release namespace