The Inherited Permission
Workloads Lose Object-Storage Access the Moment Their Identity Is Upgraded
The situation you’re stepping into
A managed Kubernetes cluster where workloads have historically obtained cloud credentials from the node’s own instance role. Every pod on a node inherits whatever that node can do — which is convenient, and is also the reason the migration is happening: it means the least-privileged pod on the box holds the most-privileged permission set attached to it.
The platform team migrates to per-workload identity: each ServiceAccount is federated to its own cloud role, so a pod receives exactly the permissions its workload needs. This is unambiguously the right direction, and it is a security improvement.
The migration completes. Nodes are healthy. Pods are Running. And every job that reads or writes object storage begins failing with access denied.
What the team observed
- Data pipeline pods start, run, and fail on their first storage call with a permissions error.
- Pods are
RunningandReady. No restarts, no crash loops. - Nodes are
Ready. Cluster metrics normal. - The new roles exist, the trust relationships are correct, and the identity provider is configured.
- Applications were not redeployed with new code. Nothing about the application changed.
- Rolling back the migration on one node pool immediately restores access.
The rollback working is a strong clue and a dangerous temptation. It tells you the change caused it — but rolling back permanently means abandoning a real security improvement because of what will turn out to be a small, fixable gap.
?
Decision Point 1 The application fails on a storage call but the pod is healthy and the role exists. What is the first thing to establish, and why is reading the role definition not it?
There is a difference between the identity you configured and the identity the process actually ends up using at runtime.
Commit to your answer, then reveal the responder’s move
→
The application fails on a storage call but the pod is healthy and the role exists. What is the first thing to establish, and why is reading the role definition not it?
There is a difference between the identity you configured and the identity the process actually ends up using at runtime.
Commit to your answer, then reveal the responder’s move →Establish which identity the process is actually using at runtime, from inside the pod.
Reading the role definition tells you what you intended. It does not tell you which credential the SDK inside the container actually resolved — and that resolution is done by a provider chain with a defined precedence, not by your intent.
kubectl exec -it <pod> -- sh
# Ask the cloud API who it thinks you are:
aws sts get-caller-identity
The answer splits the investigation cleanly in two:
| Result | Meaning |
|---|---|
| Returns the new per-workload role | Identity federation is working. This is a permissions problem. |
| Returns the node instance role | Federation is not taking effect. This is a wiring problem. |
| Fails entirely | Token projection or trust relationship is broken. |
In this incident it returns the new role. The migration worked exactly as designed — the pod is correctly assuming its own identity. So the failure is not that the new identity is absent; it is that the new identity cannot do something the old one could.
?
Decision Point 2 The pod is correctly assuming its new role, and that role's policy looks right to the team who wrote it. Why did access work yesterday and not today, when nobody removed any permission?
Nobody removed a permission. Something stopped being inherited.
Commit to your answer, then reveal the responder’s move
→
The pod is correctly assuming its new role, and that role's policy looks right to the team who wrote it. Why did access work yesterday and not today, when nobody removed any permission?
Nobody removed a permission. Something stopped being inherited.
Commit to your answer, then reveal the responder’s move →Because the workload was never using the permission you think it was using. It was using the node’s.
Under the old model the credential chain resolved to the node instance role, and that role had accumulated permissions over years — added by different people, for different workloads, for reasons rarely documented. Every pod on the node silently inherited all of them. The object-storage permission the data pipeline depended on was one of these: granted to the node, never to the workload, and never written down as a dependency of that workload.
Per-workload identity replaces that inheritance with an explicit grant. The new role contains what somebody remembered to put in it — reconstructed from the application’s documentation, or from reading its code, or from asking the team. Anything the workload relied on implicitly, and nobody remembered, is simply gone.
This is the defining property of the migration, and it is a feature rather than a bug: it converts an implicit, invisible permission set into an explicit one. The incident is the moment that conversion reveals what was undocumented.
# What could the node do?
aws iam list-attached-role-policies --role-name <node-instance-role>
# What can the workload do now?
aws iam list-attached-role-policies --role-name <workload-role>
# The difference is your list of undocumented dependencies.
?
Decision Point 3 Why did the SDK use the new role rather than falling back to the node role that still had the permission? A fallback would have avoided the outage entirely.
The credential provider chain is ordered, and it stops at the first credential it can obtain — not the first one that works.
Commit to your answer, then reveal the responder’s move
→
Why did the SDK use the new role rather than falling back to the node role that still had the permission? A fallback would have avoided the outage entirely.
The credential provider chain is ordered, and it stops at the first credential it can obtain — not the first one that works.
Commit to your answer, then reveal the responder’s move →Because the provider chain resolves by precedence, not by success.
Cloud SDKs walk an ordered list of credential sources and use the first one that yields a credential. Projected workload identity sits above the node instance metadata endpoint in that order. So once the ServiceAccount is federated and the token is projected into the pod, the SDK finds a valid credential at the higher-precedence source and stops looking.
The critical detail: the chain stops at the first credential it can obtain, not the first one that is sufficient. There is no fallback on authorisation failure. An access-denied response is a perfectly valid answer from a perfectly valid identity — the SDK has no reason to try a different one, and doing so would be a security hole rather than a convenience.
This is why the failure is total and immediate rather than partial. Every call goes to the new identity, and every call that needed the inherited permission fails.
The same precedence logic explains the reverse symptom, which is worth recognising: if get-caller-identity had returned the node role, the token projection would have failed and the SDK would have fallen through to the lower-precedence source — presenting as “the migration silently did nothing”.
?
Decision Point 4 You could restore service in a minute by attaching the old policy wholesale to the new role. Why is that the wrong fix, and what do you do instead?
The whole point of the migration was to stop granting every workload everything the node could do.
Commit to your answer, then reveal the responder’s move
→
You could restore service in a minute by attaching the old policy wholesale to the new role. Why is that the wrong fix, and what do you do instead?
The whole point of the migration was to stop granting every workload everything the node could do.
Commit to your answer, then reveal the responder’s move →Because copying the node’s policy set onto the workload role reproduces exactly the problem the migration existed to solve, with more moving parts. You would end the incident with per-workload identity that grants node-wide permissions — the same over-privilege, now harder to see.
The correct fix grants only the permission this workload actually needs, scoped to the resources it actually touches:
- Determine the real requirement from the failure itself. The error names the operation and the resource. Cloud audit logs give you the full set of calls the workload made under the old identity — the authoritative answer, better than anyone’s memory.
- Write a scoped policy: those actions, on those resource paths, and nothing else.
- Apply and verify from inside the pod before declaring it fixed:
kubectl exec -it <pod> -- aws sts get-caller-identity # correct role
kubectl exec -it <pod> -- aws s3 ls s3://<bucket>/<prefix>/ # correct access
- Then reduce the node role to the minimal baseline it genuinely needs — image pulling, logging, and the node’s own operation. Until you do this, the old inherited permissions are still available to anything not yet migrated, and the security benefit has not actually been realised.
Step 4 is the one that gets skipped, and skipping it means the migration delivered complexity without the improvement.
Root cause
The workload depended on a permission it inherited from the node, and that dependency was never documented anywhere. Under node-level credentials, every pod silently received the node instance role’s full permission set. The per-workload role was built from what the team could reconstruct about the application’s requirements, and the object-storage grant — added to the node role at some earlier point, for some earlier reason — was not among them.
Because projected workload identity takes precedence over node instance metadata in the SDK’s credential chain, the pod correctly assumed its new, narrower role and never fell back. Every storage call was made with an identity that had never been granted access, and the SDK had no reason to try anything else.
The infrastructure was healthy throughout. The migration succeeded. The failure was a permission gap between an implicit grant and an explicit one, and it was invisible until the implicit grant was removed.
Resolution and prevention
# 1. Establish what the workload actually calls (audit trail, not memory)
# Filter cloud audit logs by the node role, for the pod's activity window.
# 2. Grant exactly that, scoped to the resources in question
aws iam put-role-policy --role-name <workload-role> \
--policy-name storage-access --policy-document file://scoped-policy.json
# 3. Verify from inside the pod, not from your laptop
kubectl exec -it <pod> -- aws sts get-caller-identity
kubectl exec -it <pod> -- aws s3 ls s3://<bucket>/<prefix>/
# 4. Reduce the node role to its minimal baseline
Prevention:
- Add a policy-diff gate to the pipeline. On any role change, diff the effective permissions before and after and fail the build on an unreviewed reduction. This is the control that turns a silent runtime failure into a blocked pull request, and it is the single highest-value output of this incident.
- Derive the new policy from audit logs, not from documentation. What a workload actually called over the last ninety days is evidence. What it is documented to need is a hypothesis.
- Migrate one workload at a time, verifying from inside the pod before proceeding. A big-bang identity migration converts one small permission gap into a platform-wide outage.
- Keep an explicit inventory of node-role permissions and which workloads rely on each, produced during migration. The absence of this inventory is what made the gap invisible.
- Verify with
get-caller-identityfrom inside the pod as a standard post-migration check. It distinguishes “wrong identity” from “right identity, wrong permissions” in one command, and those have completely different fixes.
Telling this story to a recruiter
Situation. After migrating Kubernetes workloads from node-level cloud credentials to per-workload federated identity, applications silently lost object-storage access. Pods were Running, nodes were healthy, the new roles existed and were being assumed correctly — but every data pipeline job failed with permission errors, and the symptom looked like an application bug rather than an infrastructure change.
Task. Restore application access without abandoning the security migration, and without simply re-granting the broad permissions the migration was designed to remove.
Action. Ran a caller-identity check from inside the affected pods, which confirmed the workload identity was being assumed correctly — narrowing the problem from wiring to permissions in a single command. Traced the root cause to credential precedence: projected workload identity takes priority over the node instance role in the SDK’s provider chain, so the pod used the new role exclusively with no fallback, and the object-storage permission it had previously inherited from the node had never been replicated to the new role. Diffed the node role’s attached policies against the new role’s to produce the list of undocumented inherited dependencies, then wrote a scoped policy granting only the specific actions on the specific resource paths the workload used, verified from inside the pod, and reduced the node role to a minimal baseline.
Result. Application access fully restored with no further data pipeline interruptions, and the security improvement retained rather than rolled back. Added a policy-diff validation stage to the CI/CD pipeline so any future role change surfaces permission reductions at review time instead of at runtime. Node role permissions were reduced to the minimum required for node operation, which delivered the actual security benefit the migration had been intended to produce.
What this demonstrates. Debugging identity and access across the boundary between a cluster and its cloud provider; understanding credential resolution well enough to predict which identity a process will use; and resisting the fast fix that would have restored service while discarding the reason for the change.
Interview deep-dive: the full case study
Why this failure is structurally guaranteed, not bad luck. Node-level credentials create implicit permission grants. Every pod inherits the node’s full permission set without any declaration, so the set of permissions a workload actually depends on is never written down — it is discovered only when something removes it. Any migration from implicit to explicit identity will surface exactly these gaps. The question is never whether they exist, only whether you find them in a controlled migration or in production.
Why get-caller-identity is the highest-value first command. It splits the problem space in half immediately. Wrong identity means token projection, trust policy or ServiceAccount annotation — a wiring problem. Right identity means the role exists and is assumable but lacks a grant — a permissions problem. These have entirely different investigation paths, and running one command inside the pod tells you which one you are in. Inspecting the role definition from your own terminal answers a different and less useful question: what you configured, rather than what the process resolved.
Why precedence rather than fallback is the correct design. It would be convenient here if the SDK had retried with the node role. It would also be a serious security flaw: a workload denied by its own identity could silently escalate to a broader one, and least-privilege would be unenforceable. The chain stopping at the first obtainable credential — not the first sufficient one — is what makes per-workload identity meaningful. The behaviour that caused this outage is the behaviour that makes the migration worth doing.
Why audit logs beat documentation for building the new policy. Documentation records what someone believed the workload needed at the time they wrote it. Audit logs record what it actually called. On a workload of any age these diverge — features were added, integrations were bolted on, and a permission was granted at the node level as a quick fix nobody recorded. Deriving the policy from ninety days of observed calls produces a policy that fits the workload rather than the story about the workload.
Why reducing the node role is the step that must not be skipped. Until the node role is trimmed, every unmigrated pod still inherits the old broad permissions, and any workload that fails over or is scheduled before its identity is configured picks them up again. The migration has added a layer of configuration without removing the risk it targeted. The incident is only genuinely closed when the inherited path no longer grants anything worth having.
The control that generalises. A policy-diff gate in CI is cheap and catches this entire class of problem before it reaches a cluster. Permission additions are reviewed as a matter of course because someone requested them; permission reductions frequently arrive as a side effect of a refactor, a role replacement, or a migration, and nobody reviews the delta. Making the diff explicit at review time is what converts a silent runtime failure into a conversation.