Roles, STS and the Credential Chain

Where the credentials in your terminal actually came from, why a role is not a user with extra steps, and how IMDSv2 turned an SSRF bug into a non-event.

beginner 20 min lesson hands-on task included

An IAM user is a permanent identity with permanent credentials. An IAM role is a set of permissions that anything can temporarily become. Almost every credential incident in the wild comes from using the first where the second belonged.

This lesson is about where credentials come from, which is a question most people cannot answer about their own laptop.


Topic 1: Roles Are Temporary Identities

A role has two policies, and confusing them is the most common source of “why can’t I assume this role”:

  • Trust policy (AssumeRolePolicyDocument) — who may become this role. This is a resource policy on the role itself.
  • Permissions policy — what the role may do once assumed.
// Trust policy: this EC2 service, and only this, may assume the role
{
  "Version": "2012-10-17",
  "Statement": [{
    "Effect": "Allow",
    "Principal": { "Service": "ec2.amazonaws.com" },
    "Action": "sts:AssumeRole"
  }]
}
// Trust policy for cross-account: their role, and only with an ExternalId
{
  "Effect": "Allow",
  "Principal": { "AWS": "arn:aws:iam::444455556666:role/their-ci" },
  "Action": "sts:AssumeRole",
  "Condition": { "StringEquals": { "sts:ExternalId": "a-secret-you-generated" } }
}

sts:ExternalId exists to stop the confused deputy problem: a third party who legitimately holds a role ARN for one customer cannot use it to reach another customer’s account, because each has a distinct external ID that the caller must supply. If you are handing a role to a vendor, require one; if a vendor does not offer one, ask why.

Assuming a role gives you a session, not a login:

aws sts assume-role \
  --role-arn arn:aws:iam::111122223333:role/deploy \
  --role-session-name alice-deploy-2026-08-17 \
  --duration-seconds 3600

Three things about the session, all of which show up in operations:

  • The session ARN is arn:aws:sts::111122223333:assumed-role/deploy/alice-deploy-2026-08-17. The session name lands in CloudTrail, so making it meaningful is the difference between an audit trail and a shrug. Use the human or the CI job, not default.
  • Default duration is 1 hour, maximum is the role’s MaxSessionDuration (up to 12 hours). Role chaining — assuming a role from another assumed role — caps at 1 hour regardless, which is why long CI pipelines fail 60 minutes in.
  • You can pass a session policy to further reduce permissions for that session only. Useful for a job that should hold a subset of what the role can do.

Topic 2: The Credential Chain

SDK CREDENTIAL CHAIN — FIRST HIT WINS, THE REST ARE NEVER CONSULTED 1. Explicit parameters in code the worst option, and it is first 2. Environment variables AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY / AWS_SESSION_TOKEN 3. Shared config file ~/.aws/credentials, ~/.aws/config — AWS_PROFILE picks the section 4. Web identity token AWS_WEB_IDENTITY_TOKEN_FILE — this is how EKS IRSA arrives 5. Container credentials ECS / EKS Pod Identity agent on 169.254.170.2 or .23 6. Instance metadata (IMDS) the EC2 instance profile, rotated for you IMDSv2 IS TWO CALLS PUT /latest/api/token → X-aws-ec2-metadata-token GET /latest/meta-data/... + token A session token an SSRF cannot mint, and hop-limit 1 keeps it off the network. WHAT LEAKS INSTEAD A long-lived access key in: · a committed .env or Dockerfile · a CI variable nobody rotates · a laptop that left the company A role's credentials expire. A key does not. THE CHAIN EXPLAINS MOST "WRONG ACCOUNT" INCIDENTS A stale AWS_ACCESS_KEY_ID in the shell beats the profile you carefully selected. Run sts get-caller-identity before you believe anything.
First hit wins. Nothing later in the chain is consulted, which is exactly why a forgotten environment variable can send a carefully configured deployment into the wrong account.

Every AWS SDK and the CLI walk the same ordered list. Committing it to memory saves hours over a career:

  1. Explicit credentials in code
  2. Environment variables — AWS_ACCESS_KEY_ID, AWS_SECRET_ACCESS_KEY, AWS_SESSION_TOKEN
  3. Shared config — ~/.aws/credentials and ~/.aws/config, section chosen by AWS_PROFILE
  4. Web identity token file — AWS_WEB_IDENTITY_TOKEN_FILE (this is how EKS IRSA arrives)
  5. Container credentials — the ECS task role or the EKS Pod Identity agent
  6. Instance metadata — the EC2 instance profile

The habit that prevents a whole category of mistakes:

$ aws sts get-caller-identity
{
    "UserId": "AROA3EXAMPLE:alice-deploy-2026-08-17",
    "Account": "111122223333",
    "Arn": "arn:aws:sts::111122223333:assumed-role/deploy/alice-deploy-2026-08-17"
}

Run it before anything destructive. It answers “which account am I in and as whom” in one call, needs no permissions at all, and it is the only reliable answer — the profile you think is active is not evidence.

For human access, the modern answer is IAM Identity Center with aws sso login: short-lived credentials from your identity provider, no access keys anywhere, one config file mapping profiles to accounts and roles.

# ~/.aws/config
[profile prod]
sso_session = corp
sso_account_id = 111122223333
sso_role_name = PowerUser
region = eu-west-1

[profile prod-admin]
role_arn = arn:aws:iam::111122223333:role/admin
source_profile = prod          # chained: assumes admin from prod

Topic 3: Instance Metadata and IMDSv2

An EC2 instance with an instance profile gets credentials from a link-local address, 169.254.169.254. IMDSv1 answered a plain GET. That meant any server-side request forgery bug in your application — a URL-fetching feature, a misconfigured proxy, an SSRF in a dependency — could ask the application to fetch its own credentials and print them.

IMDSv2 requires a session token obtained with a PUT:

TOKEN=$(curl -sX PUT "http://169.254.169.254/latest/api/token" \
  -H "X-aws-ec2-metadata-token-ttl-seconds: 21600")

curl -s -H "X-aws-ec2-metadata-token: $TOKEN" \
  http://169.254.169.254/latest/meta-data/iam/security-credentials/

Why that closes the hole: an SSRF can usually make a GET to a URL it was tricked into. Making it issue a PUT with a custom header is a much higher bar, and the response to the PUT has to be read and re-sent — which a blind SSRF cannot do. Additionally the token response has a TTL and a hop limit: with http-put-response-hop-limit: 1, the response’s IP TTL is 1, so it cannot cross the instance boundary. A container on the host bridged network, or a misconfigured reverse proxy, gets nothing.

Enforce it, and prefer setting it in the launch template so every new instance is born correct:

aws ec2 modify-instance-metadata-options \
  --instance-id i-0abc123 \
  --http-tokens required \
  --http-put-response-hop-limit 1 \
  --http-endpoint enabled

Set --http-endpoint disabled for instances that need no AWS API access at all. Then confirm nothing in your fleet still uses v1 — CloudWatch publishes a per-instance MetadataNoToken metric, which is exactly the “who will break when I enforce this” report you need before enforcing it.

Note for containers on EC2: a hop limit of 1 blocks pods from reaching IMDS at all, which is desirable — pods should get credentials from IRSA or Pod Identity, not from the node’s role. If a pod needs the node’s permissions, that is usually a design smell worth fixing rather than a hop limit worth raising.


Topic 4: Where Each Workload Should Get Credentials

WorkloadCorrect mechanismNever
EC2 instanceInstance profileKeys in user data or on disk
ECS taskTask role (per task, not per instance)The EC2 instance role
EKS podIRSA or EKS Pod IdentityThe node instance role
LambdaExecution roleKeys in environment variables
GitHub ActionsOIDC → AssumeRoleWithWebIdentityA long-lived key in repo secrets
A humanIAM Identity Center / SSOAn IAM user with an access key
An on-prem serverIAM Roles Anywhere, or SSM hybrid activationA key that lives forever

The pattern behind every row: the workload proves its identity with something it already has — an instance identity document, a Kubernetes service account token, a GitHub OIDC token — and trades it for temporary credentials. No secret to store, rotate, or leak.

GitHub Actions is worth a concrete example, because it is where long-lived keys most often persist out of habit:

permissions:
  id-token: write            # required for OIDC
  contents: read
steps:
  - uses: aws-actions/configure-aws-credentials@v4
    with:
      role-to-assume: arn:aws:iam::111122223333:role/gha-deploy
      aws-region: eu-west-1

And the trust policy that makes it safe — note the sub condition, which is the part people leave as * and thereby let any repository on GitHub assume the role:

{
  "Effect": "Allow",
  "Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/token.actions.githubusercontent.com" },
  "Action": "sts:AssumeRoleWithWebIdentity",
  "Condition": {
    "StringEquals": { "token.actions.githubusercontent.com:aud": "sts.amazonaws.com" },
    "StringLike": { "token.actions.githubusercontent.com:sub": "repo:my-org/my-repo:ref:refs/heads/main" }
  }
}

Topic 5: When a Long-Lived Key Is Unavoidable

Sometimes a third-party tool supports nothing else. Then:

  • One key per integration, never shared, named so its owner is obvious.
  • Rotate on a schedule you actually execute — create the second key, deploy it, verify, deactivate the first, wait, delete. Two active keys per user exist precisely to make this zero-downtime.
  • Alarm on the key’s use from an unexpected source: aws:SourceIp conditions, or a CloudWatch alarm on a CloudTrail metric filter for that access key ID.
  • Find the ones you have forgotten:
aws iam generate-credential-report >/dev/null
aws iam get-credential-report --query Content --output text | base64 -d | \
  awk -F, 'NR==1 || ($9=="true" && $10<"2026-05")' | cut -d, -f1,9,10
# column 9 = access_key_1_active, column 10 = last used date

An access key that has never been used, or was last used a year ago, is pure liability. Delete it. If something breaks, that is the audit finding you wanted.

Try it yourself: put a stale AWS_ACCESS_KEY_ID in your shell and run aws sts get-caller-identity — then unset it and run again. That two-line experiment is the whole lesson, and it is the answer to most “it works locally but not in CI” reports.

Common mistake: granting a role to a whole EC2 instance because one process on it needs one permission. Everything on that host — every dependency, every sidecar, every shell someone opens — inherits it. Split the workload, or move to per-task and per-pod identity where the boundary is the process rather than the machine.