Writing / Infra
Keyless deploys: OIDC from GitHub Actions to AWS, and the subject-claim gotcha
Every deploy pipeline has to answer one question before it does anything useful: how does a build runner prove to a cloud account that it belongs there. For most of the last decade the answer was an IAM user's access key pasted into the CI provider's secret store, and it's still one checkbox away in most tutorials.
Why static keys were the default
An access key is a password that never expires. That's the appeal — it works everywhere and needs
no federation setup — and it's also the problem. The credential outlives the job it was issued
for, so a leak through a build log or a stale .env stays exploitable until somebody notices.
Rotation is manual, so in practice it doesn't happen. And one key usually serves every workflow in
the repo, so its permissions drift toward the union of everything any pipeline ever needed.
What OIDC federation changed
The current standard stores no secret at all. GitHub Actions runs its own OpenID Connect provider,
minting a short-lived token per job that describes which repository, ref, and environment asked
for it. AWS trusts that issuer, and sts:AssumeRoleWithWebIdentity exchanges the token for
temporary credentials. GCP calls it Workload Identity Federation and Azure a federated credential;
the shape is the same.
The security then lives entirely in the trust policy's conditions, because a role that trusts GitHub's issuer and nothing else will hand credentials to any repository on GitHub.
How this site is wired
Two roles, split by blast radius. A CI role, assumable from the application repos on any ref,
authenticates to ECR and pushes to specific repository ARNs. A deploy role, assumable only from
the infrastructure repo on refs/heads/main, calls ssm:SendCommand against instances carrying
one Name tag. A build can't deploy, and a feature branch can't deploy at all. Both pin the
audience claim to sts.amazonaws.com.
The claim isn't what the examples show
Both policies were written against the format every quick-start prints:
repo:OWNER/REPO:ref:refs/heads/main
AssumeRoleWithWebIdentity was denied — not with a complaint about a malformed policy, just a
denial, with the role present, the provider registered, and every value in the condition correct
when read on its own.
With GitHub's immutable-ID subject claims enabled, the sub carries numeric IDs alongside the
names:
repo:OWNER@owner_id/REPO@repo_id:ref:refs/heads/main
IDs survive renames and transfers, which is the point of the feature. But both forms exist
depending on whether a repository has it switched on, and StringEquals against the plain form
cannot match the ID form at all. The only way to see the mismatch from outside is to decode a
token GitHub actually issued and compare its literal sub against what the policy expects.
The fix
StringLike, matching both forms explicitly:
"StringLike": {
"token.actions.githubusercontent.com:sub": [
"repo:OWNER/REPO:ref:refs/heads/main",
"repo:OWNER@*/REPO@*:ref:refs/heads/main"
]
}Two shortcuts are worth refusing. repo:*/*:... matches every repository on GitHub owned by
anyone, turning a deploy role into one a stranger can assume by pushing a workflow. And
repo:OWNER*/REPO*:... looks pinned but isn't — OWNER* matches any account name beginning
with those characters, and GitHub account names are first-come. Wildcarding after the @ is safe:
neither account nor repository names may contain one, so it can only ever stand in for the numeric
ID, which stays out of version control.
Where a different shape fits
That suits this pipeline — two repositories, one account, deploys gated on a branch. Change any of those and the answer moves:
- A human in the loop. Scope on a GitHub Environment (
repo:OWNER/REPO:environment:prod) for required reviewers and wait timers a branch condition can't express. - Shared reusable workflows. The
job_workflow_refclaim pins the workflow file itself, which is a tighter statement than "some job in this repo." - Several accounts or environments. A provider and role per account, rather than one role whose conditions try to encode the distinction.
- Self-hosted runners, or workloads already inside AWS. Skip federation entirely; an instance profile or an EKS service-account role hands credentials over directly.
Keyless authentication relocates the risk rather than removing it. There's no secret left to leak, so the trust policy's conditions are the credential — and a wildcard written to clear an error message is worth about as much as a key committed to the repository.