There are two identity questions on every EKS cluster, and they are frequently confused. How does a pod get AWS permissions is IRSA or Pod Identity. How does a human or a pipeline get Kubernetes permissions is access entries. Different mechanisms, different failure modes, and both are worth being able to debug from memory.
Topic 1: Why Not the Node Role
The lazy path is to attach permissions to the node’s instance profile. Every pod on that node can then use them, because every pod can reach the instance metadata service.
That means the least-trusted container on the node holds the same AWS permissions as the most-trusted one. A vulnerable dependency in a low-importance service becomes a path to whatever the node role can do — which, on the average cluster, is more than anyone intended.
Block it explicitly. Set the hop limit to 1 so pods cannot reach IMDS at all, and keep the node role down to what the kubelet and the CNI actually require:
aws ec2 modify-instance-metadata-options --instance-id i-0abc \
--http-put-response-hop-limit 1 --http-tokens required
The node role should hold AmazonEKSWorkerNodePolicy, AmazonEC2ContainerRegistryReadOnly, AmazonEKS_CNI_Policy — and nothing about your application’s buckets, queues or databases.
Topic 2: IRSA, Step by Step
1. The cluster has an OIDC provider. EKS publishes a signed OIDC discovery document; you register it as an IAM identity provider once per cluster:
eksctl utils associate-iam-oidc-provider --cluster prod --approve
# or read it and create the provider by hand:
aws eks describe-cluster --name prod --query 'cluster.identity.oidc.issuer' --output text
2. The role trusts that provider, for one specific service account:
{
"Effect": "Allow",
"Principal": { "Federated": "arn:aws:iam::111122223333:oidc-provider/oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE" },
"Action": "sts:AssumeRoleWithWebIdentity",
"Condition": {
"StringEquals": {
"oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:aud": "sts.amazonaws.com",
"oidc.eks.eu-west-1.amazonaws.com/id/EXAMPLED539D4633E53DE1B71EXAMPLE:sub": "system:serviceaccount:payments:api"
}
}
}
3. The ServiceAccount points at the role:
apiVersion: v1
kind: ServiceAccount
metadata:
name: api
namespace: payments
annotations:
eks.amazonaws.com/role-arn: arn:aws:iam::111122223333:role/payments-api
4. A mutating webhook injects the plumbing into every pod using that service account — a projected token at /var/run/secrets/eks.amazonaws.com/serviceaccount/token, plus AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE. The SDK finds them at step 4 of the credential chain and calls AssumeRoleWithWebIdentity. Credentials last about an hour and the SDK refreshes them.
The trust policy is where this breaks, and it breaks silently. Three specific mistakes:
StringLikewith a wildcard.system:serviceaccount:*:apilets any namespace with a service account namedapiassume the role. On a multi-tenant cluster that is a privilege escalation any team can perform by naming a service account.- Omitting the
audcondition. The audience check is what binds the token to STS. Without it the trust is materially weaker. - Wrong namespace or service account name. Produces
AccessDenied ... Not authorized to perform sts:AssumeRoleWithWebIdentityfrom inside the pod, with no indication of which half is wrong.
Debugging is fortunately mechanical:
kubectl exec -it deploy/api -n payments -- env | grep AWS_
# expect AWS_ROLE_ARN and AWS_WEB_IDENTITY_TOKEN_FILE
kubectl exec -it deploy/api -n payments -- aws sts get-caller-identity
# expect assumed-role/payments-api/botocore-session-...
Missing environment variables mean the webhook did not fire — usually the pod predates the annotation (restart it) or the service account name on the pod does not match. Variables present but get-caller-identity failing means the trust policy.
Topic 3: EKS Pod Identity
Pod Identity is the newer mechanism and removes most of the wiring. Instead of an OIDC provider per cluster and a trust policy per role, you install the Pod Identity Agent add-on and create an association:
aws eks create-pod-identity-association --cluster-name prod \
--namespace payments --service-account api \
--role-arn arn:aws:iam::111122223333:role/payments-api
The role’s trust policy becomes uniform and boring:
{
"Effect": "Allow",
"Principal": { "Service": "pods.eks.amazonaws.com" },
"Action": ["sts:AssumeRole", "sts:TagSession"]
}
| IRSA | Pod Identity | |
|---|---|---|
| Setup | OIDC provider per cluster | One add-on |
| Role reuse across clusters | New trust entry per cluster | Same role, new association |
| Where the mapping lives | The role’s trust policy | The EKS API |
| Works outside EKS | Yes — any OIDC-capable Kubernetes | EKS only |
| Session tags | No | Yes |
Pod Identity is the better default for new EKS clusters. IRSA remains necessary for non-EKS Kubernetes and is still what most existing clusters and most Helm charts assume, so you will read and debug both for years.
Topic 4: Who Can Talk to the Cluster
Authentication to the Kubernetes API is always IAM; authorisation is always Kubernetes RBAC. The mapping between them used to live in a ConfigMap called aws-auth, which had two properties that caused real incidents: it was a single object every cluster admin edited by hand, and a malformed edit locked everyone out with no way back in.
Access entries replaced it with an API:
aws eks create-access-entry --cluster-name prod \
--principal-arn arn:aws:iam::111122223333:role/platform-admin \
--type STANDARD
aws eks associate-access-policy --cluster-name prod \
--principal-arn arn:aws:iam::111122223333:role/platform-admin \
--policy-arn arn:aws:eks::aws:cluster-access-policy/AmazonEKSClusterAdminPolicy \
--access-scope type=cluster
Why this is better: it is an API call with validation, it is auditable in CloudTrail, it can be managed in Terraform without a kubectl dependency, and a mistake is reversible without cluster access. The authentication mode can be API, API_AND_CONFIG_MAP, or CONFIG_MAP; migrate to API_AND_CONFIG_MAP, move entries across, then to API.
The cluster creator’s implicit admin rights are worth knowing about: the IAM principal that created the cluster historically held system:masters implicitly and invisibly. If that was a personal role or a CI role nobody tracks, you have an administrator that appears in no list. Newer clusters make this explicit and controllable — check yours, and make sure a durable role, not a person, holds cluster admin.
Topic 5: Security Groups for Pods, and What It Costs
The VPC CNI can attach a security group to individual pods with a SecurityGroupPolicy, which lets a pod be an AWS network principal — an RDS security group can allow sg-payments-pods rather than the whole node CIDR.
The catch is capacity: pods with security groups use branch ENIs, and the number of branch ENIs per instance is limited and much smaller than the ordinary IP budget. Enabling it for everything reduces pod density sharply. Use it where the boundary is genuinely needed — the pods that reach a regulated database — and use NetworkPolicy for everything else.
Topic 6: The Scanning and Policy Layer
Three complementary things, which people often conflate:
- Image scanning — ECR scanning (basic, or enhanced via Inspector) finds known CVEs in images. Run it on push, and fail the pipeline on critical findings rather than emailing them.
- Cluster configuration scanning — kube-bench for CIS benchmarks, kubescape or Trivy for misconfigurations. These answer “is this cluster configured to a known standard”.
- Admission policy — Pod Security Admission, or OPA Gatekeeper / Kyverno, enforcing rules at deploy time so a violating workload is rejected rather than reported.
Scanning tells you what is wrong; admission control stops it recurring. A cluster with scanning and no admission policy accumulates the same finding indefinitely, and the report becomes noise everyone has learned to skip.
Try it yourself: create an IRSA role whose trust policy uses StringLike with system:serviceaccount:*:api. Then create a service account named api in a completely unrelated namespace and assume the role from a pod there. Watching that succeed is the fastest possible way to learn why the condition must be exact.
Common mistake: granting the IRSA role broad permissions “for now” because the pod’s exact API calls are unknown. The pod runs, the permissions stay, and a year later nobody can safely reduce them. Instead run wide in a non-production cluster, read the CloudTrail events for that role, and write the policy from what the workload actually called — the same technique from the IAM lesson, applied where it matters most.