IAM is not complicated, but it is unforgiving: it does exactly what the policies say, in a fixed order, with no inference. Almost every “IAM is confusing” moment is actually a request being evaluated against a policy the engineer did not know applied.
Learn the evaluation order once and IAM stops being mysterious.
Topic 1: The Evaluation Order
Every API call carries a principal (who), an action (s3:GetObject), a resource (which bucket and key), and a request context (source IP, whether it used TLS, MFA, tags, time). AWS evaluates:
- Explicit deny anywhere → denied. Final. No policy, no seniority, no root account overrides it. This is a feature: it is how a guardrail becomes unconditional.
- Organizations SCPs must permit the action, at every level of the OU tree above the account.
- Permissions boundary, if the principal has one, must permit it.
- An identity policy or a resource policy must explicitly allow it.
- Otherwise → implicit deny. The default answer is no.
The two consequences worth memorising:
- Nothing is allowed until something allows it. A brand-new IAM user with no policies attached can do exactly nothing, including reading its own username.
- An explicit Deny beats every Allow, always. So a guardrail policy is written as a Deny, never as “just don’t grant it” — because someone will grant it.
Topic 2: Anatomy of a Policy Statement
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "ReadAppConfigOnly",
"Effect": "Allow",
"Action": ["s3:GetObject", "s3:ListBucket"],
"Resource": [
"arn:aws:s3:::app-config",
"arn:aws:s3:::app-config/*"
],
"Condition": {
"Bool": { "aws:SecureTransport": "true" },
"StringEquals": { "aws:PrincipalTag/team": "platform" }
}
}
]
}
"Version": "2012-10-17" is a policy language version, not a date you should update. Omit it and you get the 2008 language, which silently drops policy variables.
The two-ARN pattern for S3 is not redundant. Bucket-level actions (ListBucket, GetBucketLocation) take the bucket ARN; object-level actions (GetObject, PutObject) take bucket/*. Listing a bucket with only bucket/* fails, and the error message does not say why. This single asymmetry accounts for a large share of S3 permission tickets.
Condition keys are the difference between least privilege and theatre. The ones that pay for themselves immediately:
| Key | Use |
|---|---|
aws:SecureTransport | Deny anything not over TLS |
aws:MultiFactorAuthPresent | Require MFA for destructive actions |
aws:PrincipalOrgID | Allow only principals from your own organization |
aws:SourceIp / aws:VpcSourceIp | Restrict to your network — careful, breaks via NAT and VPC endpoints differently |
aws:RequestTag / aws:ResourceTag | Tag-based access control |
ec2:InstanceType | Stop a sandbox account launching a p5 |
A statement with multiple conditions requires all of them (AND). Multiple values inside one condition operator are OR. Getting that backwards produces a policy that appears to work and enforces nothing.
Topic 3: Identity Policies vs Resource Policies
Two places a permission can live, and the difference matters most across account boundaries.
Identity policy — attached to a user, group or role: “this principal may do X”. Resource policy — attached to the resource: “these principals may do X to me”. S3 bucket policies, KMS key policies, SQS queue policies, Lambda resource policies, SNS topic policies.
SAME ACCOUNT identity policy OR resource policy → allowed
CROSS ACCOUNT identity policy AND resource policy → both required
Cross-account access needs both sides to agree, which is a deliberate design: no one can grant themselves access to your bucket, and you cannot grant access to a principal whose own admin has not permitted it. When cross-account access “silently fails”, one of the two sides is missing — and the error is returned from the resource’s side.
KMS is the exception worth flagging. For a KMS key, the key policy is authoritative. If the key policy does not delegate to IAM ("Principal": {"AWS": "arn:aws:iam::111122223333:root"} with an Allow), then IAM policies granting kms:Decrypt do nothing at all. An account administrator with iam:* and kms:* can be completely locked out of a key, and the only fix is a support case. Read the key policy before you write the IAM policy.
Topic 4: Managed, Inline and Boundaries
| Kind | Reuse | When to use |
|---|---|---|
| AWS managed policy | Everywhere, maintained by AWS | Convenience and coarse; ReadOnlyAccess grants more than most people expect |
| Customer managed policy | Attach to many principals, versioned, rollback | The default choice |
| Inline policy | One principal only, deleted with it | A permission that must never be reused or accidentally attached elsewhere |
| Permissions boundary | A ceiling, not a grant | Letting developers create roles without letting them create admin roles |
A permissions boundary grants nothing. It caps. Effective permissions are the intersection of the identity policy and the boundary. The pattern it exists for: let a team create IAM roles for their own workloads, while a boundary guarantees no role they create can exceed what the team itself has. Without it, iam:CreateRole + iam:AttachRolePolicy is a privilege escalation path to administrator, and it is a short one.
Quotas that shape your design: 10 managed policies per principal, 6,144 characters per managed policy, 2,048 for an inline user policy. Teams hit the 10-policy limit and start merging policies badly; the correct answer is usually fewer, better-scoped policies rather than one enormous one.
Topic 5: Reading an AccessDenied
The message is more informative than its reputation:
An error occurred (AccessDenied) when calling the GetObject operation:
User: arn:aws:sts::111122223333:assumed-role/app-role/i-0abc123
is not authorized to perform: s3:GetObject
on resource: arn:aws:s3:::prod-data/report.csv
because no identity-based policy allows the s3:GetObject action
Four facts to extract, in order:
- The principal.
assumed-role/app-role/...— is that the role you thought the code was using? Surprisingly often it is not. Check the credential chain. - The exact action.
s3:GetObject, not “S3 access”. Your policy may grants3:Get*and the SDK may be callings3:GetObjectAclors3:GetObjectTaggingunderneath. - The exact resource. Compare character by character with your policy’s
Resource. A missing/*is the usual answer. - The reason clause. Modern messages name the policy type: no identity-based policy allows, with an explicit deny in a service control policy, with an explicit deny in a resource-based policy. That clause maps directly onto the five evaluation stages, and it tells you which file to open.
When the message is not enough, stop guessing and simulate:
aws iam simulate-principal-policy \
--policy-source-arn arn:aws:iam::111122223333:role/app-role \
--action-names s3:GetObject \
--resource-arns arn:aws:s3:::prod-data/report.csv \
--query 'EvaluationResults[].[EvalDecision,MatchedStatements[].SourcePolicyId]'
The simulator evaluates without performing the action, and it reports which statement matched. For the same job in the console, IAM Access Analyzer’s policy checks and the “last accessed” data on a role are how you find over-granted permissions that nobody has used in 90 days.
Topic 6: Writing Policies That Age Well
Start from CloudTrail, not from the docs. Deploy with a broad policy in a non-production account, run the real workload, then read what it actually called:
aws cloudtrail lookup-events --max-results 200 \
--lookup-attributes AttributeKey=Username,AttributeValue=app-role \
--query 'Events[].[EventName,EventSource]' --output text | sort | uniq -c | sort -rn
That list is the policy. IAM Access Analyzer will generate one from CloudTrail directly.
Never use "Action": "*" with "Resource": "*" outside a break-glass role, and if you have a break-glass role, alarm on its use rather than pretending it does not exist.
Prefer conditions over more statements. Deny * unless aws:PrincipalOrgID matches is one statement that closes an entire class of exposure, and it keeps working as you add services.
Deny by default at the organization level, allow at the account level. Guardrails belong in SCPs, grants belong in IAM. Mixing them produces policies nobody dares change.
Try it yourself: attach a policy that allows s3:* on one bucket, then add this Deny and watch it win:
{
"Effect": "Deny",
"Action": "s3:*",
"Resource": "*",
"Condition": { "Bool": { "aws:SecureTransport": "false" } }
}
Common mistake: debugging a permission problem by adding permissions until it works. You end up with a role holding four unnecessary grants and no idea which one mattered — and the one that mattered may have been the resource ARN, not the action. Read the denial, fix the one thing it names, and re-run. If you have added a grant and the error did not change, remove it again.