Blast Radius Containment Through Resource-Level IAM Scoping
Permission models set incident ceilings long before attackers arrive.
Summary
Permission models set incident ceilings long before attackers arrive.
Blast radius gets treated as something you measure after an attacker is already inside: how far did they get, what did they touch, how bad is it. That framing is backwards. By the time anyone is asking "how far can this go," the answer was already fixed, weeks or months earlier, when someone provisioned a role, attached a policy, or granted an exception that never got revoked. The permission model is the blast radius. Everything that happens during an incident is just that model playing out in real time.
Why blast radius is determined before compromise, not during it
The instinct to "contain the breach" treats blast radius as a live variable, something a security team can shrink mid-incident through fast detection and quick isolation. Speed matters, but it doesn't change the ceiling. A compromised credential authenticates, then the attacker moves to whatever that credential was already allowed to reach. Nothing about that second step is improvised. It was decided the day the role was created.
Most large blast radii trace back to the same structural habit: a permission set built for the most demanding use case anyone could imagine, then applied uniformly to everyone who needed any version of that access. Add in the one-time exception granted for a deadline and never scoped back, and accounts end up holding far more reach than their actual job requires. None of that requires a sophisticated attacker. It just requires someone to find the account.
Reducing blast radius is an access design problem, solved or failed at provisioning time, long before anyone signs in with bad intentions. Everything else downstream, logging, alerting, isolation, is damage control around a ceiling that was already set.
How the permission evaluation stack sets the ceiling
Cloud IAM doesn't grant access through one check. It runs a stack of layers, and the identity's actual reach is set by whichever layer in that stack is most restrictive, not by whichever layer is most generous. In AWS, evaluation runs from Organization-level Service Control Policies, which cap maximum permissions for every principal in the account and can only take access away, never hand it out, through resource-based and identity-based policies, down to Permission Boundaries, which cap what a specific principal is allowed to hold even if its own policy claims more.
It has two distinct sides, and they don't automatically cover each other. SCPs govern identities: what a user, role, or service is permitted to attempt. They say nothing about what a resource is willing to accept from whoever shows up asking. An S3 bucket's own resource-based policy can still make that bucket public, and an SCP can't stop it, because SCPs don't apply directly to resource-based policies and can't restrict access from an external principal even when the bucket's own policy allows it. There's also a known blind spot in the model itself: SCPs don't reliably constrain AWS service principals acting on an account's behalf, so a chunk of inter-service activity sits outside the identity-side ceiling no matter how tight the SCP is written.
Resource Control Policies exist to close that specific gap. Instead of restricting what an identity can ask for, an RCP sets a ceiling directly on the resource, so it rejects requests regardless of which identity, internal or external, is making them. An organization running SCPs without RCPs has built half a fence. Identities are boxed in. Resources are still standing open to whatever external or service principal manages to reach them. Least privilege at the identity layer doesn't fix that, because the gap is about what resources are willing to give, not about what identities are allowed to want.
What September 2025 SCP changes enable that was not possible before
In September 2025, AWS expanded SCPs to support the full IAM policy language, and the change is worth understanding precisely because of what it fixes and what it doesn't. Before the update, SCPs couldn't use conditions inside Allow statements, couldn't target individual resource ARNs inside Allow statements, couldn't pair NotAction with Allow, couldn't use NotResource in either the Allow or Deny branch, and only allowed wildcards in Action strings if the wildcard sat alone or at the end of the string. The practical effect was that teams writing SCPs had to either accept a broader policy than they actually wanted or stack multiple overlapping policies to fake the precision the language itself didn't support.
After the update, an SCP can target one specific resource ARN directly. It can apply a condition inside an Allow statement, time of day, source IP, whether MFA was present, the same conditional logic that's long been available in identity policies. It can use NotAction alongside Allow to express "everything except these specific actions" instead of enumerating every permitted action by hand. None of this breaks existing SCPs; the change is additive, so organizations can migrate policies one at a time rather than rewriting the whole guardrail layer at once.
The blast radius consequence is concrete: a guardrail that used to have to permit a wider action space, because the policy language simply couldn't express the narrower intent, now leaves a smaller reachable surface even if the underlying identity policy hasn't changed at all. That's a real gain in precision. It doesn't touch the service-principal blind spot described above, and it doesn't replace the need for resource-side controls like RCPs. It raises the ceiling's precision on the identity side. The resource side still needs its own enforcement.
Storm-2949 as a proof of the ceiling argument, in both directions
The Storm-2949 attack is a useful case study precisely because nothing about it required custom malware. The intrusion crossed three cloud service layers, SaaS into PaaS into IaaS, and reached into endpoint environments, using entirely legitimate administrative tools: Graph API calls, Azure RBAC role assignments, Run Command, VMAccess extensions, publishing profiles. Every step was something an admin does on a normal Tuesday. The attack surface was the permission model, not a software flaw.
The breach point shows the ceiling argument at its clearest. The compromised user held the Owner role over a specific Azure Key Vault, and Owner is about as broad as a role gets. In four minutes, the attacker reconfigured the vault's access policies, extracted dozens of secrets, including database connection strings and identity credentials, and used them to log into the production web application. None of that required cleverness. It required the role already in place, which happened to be Owner on a vault instead of a scoped reader on a handful of specific secrets.
The same attack also demonstrates that a narrowly scoped identity can hold where a broad one fails. When the attacker tried to abuse a separate VM's managed identity, requesting a token from the Azure Instance Metadata Service to pull more Key Vault secrets, the attempt failed, because that managed identity's permissions didn't include the access needed to retrieve them. One narrowly scoped workload identity closed a door that an Owner-level identity had left wide open two steps earlier, inside the same incident. The post-incident recommendation from the case is specific: audit Azure RBAC role assignments and Key Vault access policies on a regular cadence to catch over-permissioned identities, and Microsoft's own guidance favors RBAC over Key Vault access policies for exactly this reason, because RBAC gives a cleaner audit trail of who can do what.
Capital One's 2019 breach runs the same pattern through less technical complexity. The IAM role attached to the compromised workload had access well beyond what that workload needed to function, and the scope of that role, not the existence of the underlying vulnerability, set the size of the incident. Public reporting put the total cost above $300 million. Different cloud, different year, identical lesson: the vulnerability gets you in, the role's scope decides how much that's worth.
Why machine identities expand the problem surface faster than human-identity programs can track
Machine identities, service accounts, API keys, workload tokens, automation roles, run on the same permission models built for human users, and they accumulate access through the same provisioning habits. A template gets copied, a broad grant gets reused, a one-time integration gets permanent credentials. Nobody reviews a service account's access the way they review a departing employee's. Over-privileged service accounts are the baseline in most enterprise environments. A majority of organizations have at least one IAM role that's gone unused for a long stretch of time but is still active and still carries full-strength permissions, a standing door nobody remembers leaving unlocked.
One broadly permissive role attached to many instances at once compounds the risk. A single compromised credential in that setup doesn't open one door. It opens every door that role was attached to, simultaneously, and it does it at whatever speed the automation runs, which is faster than any human lateral-movement sequence. Third-party integrations make the same problem worse from the outside: a meaningful share of third-party AWS integrations carry access far beyond what the integration actually needs, sometimes reaching all data in the account, sometimes enough to take the account over entirely.
Lifecycle is the deeper issue here, beyond scope alone. Human identities get reviewed when something changes: a new manager, a new role, an exit. Machine identities mostly don't have an equivalent trigger. They sit there, fully permissioned, until someone runs an audit, and most programs don't run that audit often. Resource-level scoping matters here precisely because the machine-identity population is large and growing. A hard ceiling enforced at the resource layer limits what any single compromised machine identity can do, without requiring someone to manually re-review every entitlement on every service account before the next audit cycle comes around.
How persistent AI agents break the assumptions that human-scale IAM was built on
A persistent AI agent isn't a single request that comes in, gets processed, and disappears. It holds context across sessions, carries delegated credentials forward, and calls tools on its own initiative, often without a human signing off on each individual step. That changes what the permission model has to answer for. It's bounding the cumulative reach of an ongoing, self-directed process that keeps running after the human who kicked it off has moved on to something else.
A standard stateless AI application handles one request and throws the session context away when it's done. A persistent agent keeps that context: an instruction injected by an adversary early on can keep influencing the agent's reasoning in every session that follows, not just the one where it landed. Agents often inherit the same access level as the human or service account that spun them up, and they act on that access at machine speed, turning what would be a slow, visible, human-paced lateral movement into something that finishes before anyone notices it started.
Two principles govern this: least agency, which restricts an agent's permissions to the minimum capability its stated task actually requires, and blast radius, which caps the maximum damage any single agent can do before some containment mechanism kicks in. The concrete technique that puts both into practice is just-in-time provisioning: credentials and tool permissions get minted at the start of a task, scoped tightly to that task, and revoked the moment the task finishes. The agent requests access through an identity gateway that evaluates the context, the stated intent, and the applicable policy, then issues a short-lived token with the smallest time-to-live the task can run on. Standing permissions don't have a safe version in this model. There isn't one.
What makes agents the sharpest version of the blast-radius argument is timing. Security leaders broadly agree that governing AI agents is a critical priority, but most organizations haven't put adequate, agent-specific policy in place yet. That gap between stated urgency and actual control makes skipping tight design-time permissioning costly fast, because an agent that's over-permissioned doesn't wait for a human attacker to exploit it. It just keeps acting on the permissions it already has, at whatever speed its next task calls for.
Where least privilege alone fails to set the ceiling
Least privilege is the right instinct, but it answers only half the question. It constrains what an identity can request. It says nothing about what a resource will accept once a request arrives, and nothing about what happens when the number of admin accounts stays high or an admin role's scope stays wide by default. A program can enforce least privilege rigorously at the identity layer and still leave a compromised admin account with "everything in the account" as its effective reach, because the admin role itself was never narrowed.
Infrastructure-as-Code sharpens the same point. A single pull request can rebuild an entire environment: the potential blast radius of one merged change is enormous, regardless of how carefully any individual identity's permissions were scoped. The review workflow wrapped around the tool is the fix: plan output surfaced on every pull request, state files scoped per environment instead of shared across all of them, separate cloud accounts per environment so a mistake in staging can't touch production, and a policy engine that flags destructive diffs before anything applies. The tool doesn't reduce the radius. The governance process built around the tool does.
Resource-level controls close what identity-scoped policies structurally can't reach, because identity policies were never designed to govern what the resource itself accepts, which makes them a necessary complement to least privilege. A policy that exists only in a wiki page or an onboarding doc depends on every developer reading it and following it under deadline pressure, and that dependency is a hope, not a system-level control. Enforcement that evaluates every call before it executes is a different category of protection than a log that shows, after the fact, what somebody did.
Operationalizing resource-level scoping: the design decisions that set the ceiling in practice
None of this comes from a single toggle in a console. It comes from a set of decisions made at the moment something gets provisioned, each one narrowing what a future compromised identity, human or machine, would be able to reach. Scope roles to specific resource ARNs instead of account-wide wildcards, so a compromised credential inherits only a short list of resources. Layer Resource Control Policies alongside SCPs so the resource side of the evaluation stack enforces its own ceiling instead of trusting the identity side to do all the work. Use the new SCP capabilities from the September 2025 update, conditions inside Allow statements, targeted resource ARNs, NotAction logic, to write guardrails shaped like the workload's actual behavior. Treat service accounts and workload identities with the same review cadence as human accounts, since an idle but still-active role is exactly the kind of standing door that compounding incidents exploit. For anything involving a persistent AI agent, default to just-in-time credential issuance scoped to the task at hand, rather than a standing grant that outlives the job it was meant for. Wrap provisioning-tool changes in a review workflow, plan output on every pull request, environment-scoped state, a policy engine that catches destructive diffs, since the blast radius of a single merged change can outstrip anything a single compromised identity could do on its own.
Each of those is a provisioning-time decision, not an incident-time reaction. That's the whole argument. By the time an attacker is inside, the ceiling has already been set.
Sources
- Blast Radius and Cloud Threat Detection
- AWS Organizations supports full IAM policy language for service control policies (SCPs) - AWS
- Unlock new possibilities: AWS Organizations service control policy now supports full IAM language
- Get more out of service control policies in a multi-account environment
- How Storm-2949 turned a compromised identity into a cloud-wide breach
- Agent Security is a Systems Problem