Least-Privilege Enforcement for AI Agents at Runtime

Agents need permissions that shrink with each task step, not static roles set at startup.

Summary

Agents need permissions that shrink with each task step, not static roles set at startup.

Least privilege has always rested on one premise: that someone can sit down before a system runs and figure out what access it will need. That premise holds for a database connection or a cron job. It falls apart the moment the thing requesting access gets to decide, on its own, what to do next. Research on agent instruction files found that 74% of policy statements depend on context that cannot be pre-defined statically, which confirms a simple fact: an agent's scope at step three is not knowable at the moment it is provisioned, because the agent hasn't picked step three yet. Static access models were built for identities that hold credentials and do nothing more interesting with them. An agent reasons about which tool to call, which record to touch, which system to reach into, and it does that reasoning live, mid-task, often several times a minute. The mismatch is a model built for predictable, human-paced access applied to a system that chains tool calls and makes its own access decisions in real time. A permission set that was correct when the agent was deployed is routinely wrong two steps later, and nothing in a static model notices, let alone fixes it, while the task is still running.

Agents inherit and accumulate access beyond what any single task requires

Most agents don't get built with carefully scoped credentials. They get built fast, and fast means borrowing whatever identity is already lying around. Agents typically authenticate through API keys, OAuth tokens, and service accounts, the same credential types used for human, session-scoped logins. These often carry broad permissions and long lifecycles that were never designed with a tireless, autonomous caller in mind. When an agent gets spun up from a developer's own credentials or a shared service account, it doesn't inherit a slice of that account's access. It inherits all of it. A developer's agent built to handle a narrow internal task can end up holding production-level access long after anyone remembers why.

Pilots make this worse on a schedule. Access gets widened to clear some early blocker, the pilot ships, and the widened access just stays, because proving it's safe to revoke takes more effort than leaving it alone. The 2026 Infrastructure Identity Survey found that 67% of organizations still rely heavily on static credentials despite knowing the risk, which is less a statistic about ignorance than one about inertia. OWASP's Top 10 for Agentic Applications 2026 gave this pattern a name: Identity & Privilege Abuse, ASI03, where an agent's identity or permissions get misused or escalated, a direct descendant of the Excessive Agency risk OWASP first flagged for LLM applications generally. The access surface an over-privileged agent carries is also the attack surface an attacker gets to use. Indirect prompt injection works precisely because of this overlap: a manipulated instruction becomes an action across every system the agent can reach, using credentials nobody stole and permissions nobody had to escalate. The agent does the escalating on command, with keys it already had.

The blast radius of a broadly credentialed agent compromised or misdirected

Two incidents from early 2026 show what this looks like once it stops being theoretical. Palo Alto Networks Unit 42 documented a path it called "Double Agent" in Google Cloud's Vertex AI Agent Engine, where a deployed agent inherited excessive default permissions through a Google-managed service account. That inheritance opened the door to credential extraction and privileged access into consumer-project resources, and the over-privilege sat entirely in the deployment configuration. No attacker had to do anything clever to create it. It shipped that way.

The TeamPCP / LiteLLM supply chain compromise in March 2026 told a similar story from a different angle. A threat actor Google Threat Intelligence Group tracks as UNC6780, using malware called SANDCLOCK, ran multi-repository compromises against Trivy, Checkmarx, and BerriAI's LiteLLM, planting a credential stealer in build environments that pulled out AWS keys, GitHub tokens, and AI API credentials. Palo Alto Networks Unit 42 and Datadog Security Labs both documented the campaign. None of it would have been worth the effort if the credentials sitting in those build environments had been narrow and short-lived. They were broad and long-lived, and that breadth and longevity made the exfiltration valuable.

Neither case stays contained to a single agent once delegation enters the picture. The Cloud Security Alliance's framework on Agentic AI Identity and Access Management lists identity spoofing, credential reuse for privilege escalation, and agents chaining delegations to each other as distinct risk categories, and the common thread is that the resulting blast radius is bounded by the full chain of agents that first agent can talk to. The 2026 Singapore Consensus on AI Safety Research puts the fix into a research priority rather than a product feature: an agent's capabilities need to be scoped to the minimum necessary for its current task and context, with that scope adjusting as the task evolves, and any escalation beyond that needs explicit authorization rather than quietly inheriting whatever permissions existed at deployment. That's a diagnosis the field has converged on, and it points straight at what the rest of this piece argues for.

Why runtime, step-scoped permission grants are the structural answer

If the problem is that scope can't be known in advance, the fix follows almost by definition: stop trying to know it in advance. Scope access to the step the agent is on right now, grant what that step needs, and pull it back the instant the step finishes. Every time the agent needs to act, it requests the specific permission that action requires, uses it, and loses it. No standing credential sits idle between steps waiting for an attacker or a hijacked prompt to find a use for it.

This moves the entire security boundary. Provisioning-time scope is a guess made before the agent has done anything. Execution-time scope is a fact, known at the moment it matters, bounded by what the current step actually requires rather than by whatever role the agent was configured to hold. The 2026 Singapore Consensus frames this as a shift from static permission minimization to dynamic capability restriction and execution isolation: least privilege changes from a decision made once at provisioning to a constraint enforced continuously, adjusting as the task's context changes step by step. Getting that right takes more than scoping the credential correctly. A correct step-level grant has to account for the agent's goal, the specific action being taken, the data involved, whose authority the agent is acting under, the plan that led to this step, and the damage possible if the step goes wrong. If a review misses any one of those factors, it can look perfectly clean while the actual access behind it stays dangerous.

Runtime enforcement in practice: identity, tokens, and the control plane

None of this works without a control plane sitting outside the agent's own code, one that evaluates every access request against current context instead of checking it against a role someone assigned weeks earlier. That starts with individual agent identity: each agent needs its own governed identity, separate from the human or service account used to build it. Without that separation, there's no way to scope access to what the agent actually does, and no way to audit which agent did what after the fact.

The scoping also needs to happen at the level of the action, not the session. A session-level grant lets an agent do anything it wants for as long as the session lasts. Action-level scoping asks a narrower question every time: does this specific tool call belong to the task the agent is currently on? Attribute-Based Access Control, or ABAC, is the mechanism that makes asking that question fast enough to matter, adjusting permissions in real time based on context as it changes. Scoping has to reach into the data as well as the system. Access to a given application" isn't a least-privilege grant, it's a category. What matters is which records and which fields a given step needs, not which application it needs to open.

Ephemeral credentials make the whole model operational. Orchestration patterns that spin up short-lived, task-specific agents make persistent credentials awkward to use even if a team wanted them, which pushes the architecture toward exactly the kind of short-lived authentication that step-scoping requires. And the logic that decides what gets granted cannot live inside the agent's own reasoning or its prompt. Authorization decisions made inside an agent are exposed to the same prompt injection that can redirect its task. The enforcement layer has to sit outside that reasoning, where an injected instruction can't touch it.

This is already running in production, not just on paper. Google Cloud's Vertex AI Agent Engine binds per-agent identity principals to trusted runtime environments, so certificate-bound tokens can't be replayed outside the environment they were issued for. Microsoft's Agent platform builds its version on Entra for agent identity and access, Defender for threat protection, and Purview for data security, with least-privilege controls determining which users, data, tools, and MCP servers a given agent can reach. Neither company built a new identity stack from scratch for this. Both extended infrastructure that already existed. That matters for anyone worried that runtime enforcement means ripping out identity infrastructure and starting over: a standards-body concept paper on accelerating the adoption of software and AI agent identity and authorization, published February 5, 2026, makes the case that existing enterprise identity standards can extend to AI agents without new infrastructure, and the overwhelming majority of commenters on that paper agreed, warning that abandoning hardened identity systems creates adoption barriers and immediate security risk.

The agent skill and behavioral gap in step-level controls

Runtime permission controls are good at one specific job: stopping an agent from reaching a system or record it has no business touching. They are not built to notice an agent choosing a more powerful tool than a task actually calls for when a weaker one would have done the job just as well. Research on agentic tool selection has found exactly that pattern: agents tend to reach for higher-privileged tools even when lower-privileged alternatives sit right next to them. The agent was authorized to use the tool it picked, so this is a judgment failure inside the agent's own reasoning rather than a conventional permissions failure, and a perimeter control has no way to see it coming.

The Agent Skill ecosystem adds a wrinkle on top of that. A Skill bundles a reusable capability with its own permissions, and the same Skill action can be entirely appropriate under one user prompt and badly over-privileged under another. The over-privilege isn't a property of the Skill itself. It depends on what the task actually was, so it shifts from one invocation to the next even when nothing about the Skill's code has changed.

SkillScope, a paper accepted at a top security conference in 2026 from a university research team, takes this on directly. It builds a graph-based model of instruction-level procedures and code-level operations as fine-grained action nodes, pulls out candidates for over-privilege, checks each one against the actual runtime task context, and constrains the ones that turn out to be genuinely over-privileged through what the researchers call control-flow privilege constraining. In large-scale testing, the system validated thousands of Skills carrying over-privileged behavior and cut triggered over-privileged action instances by 88.56%, without breaking legitimate task completion. The lesson for least-privilege design is specific: for agent skills, enforcement has to be conditioned on the task, not just on the role, because the same capability needs different access depending on what the user actually asked for. The research closing that gap in runtime enforcement is already published, not hypothetical.

Runtime controls and behavioral governance working together

An agent can stay entirely inside its runtime-scoped permissions and still do something it shouldn't. Access enforcement and behavior monitoring have to run at the same time, side by side, each catching what the other misses. Runtime permissions answer the question of what an agent is allowed to touch. They say nothing about whether, within everything it's allowed to touch, the agent is behaving the way the task calls for.

The concept sometimes gets called "least agency": governing what an agent actually does, layered on top of governing what it can access. It completes least privilege. Access controls set the ceiling. Behavioral controls govern what happens underneath that ceiling, a different kind of question. The 2026 Singapore Consensus treats these as parts of one structure rather than separate concerns competing for attention. Its Companion Report on Agentic Risk Management lists Least Privilege, Traceable Identity, Auditability, and Runtime Assurance among ten foundational principles spread across three lifecycle phases, design and development, testing and deployment, and operation and monitoring. None of the four stands on its own. Each depends on the others already being in place.

Governance here cannot be a quarterly audit, because agents make large numbers of access decisions inside very short windows, and a review that runs once a season cannot catch a violation that started and finished inside a single session hours or days earlier. What makes the model work after the fact is a complete record of every action an agent took, not just which systems it accessed but every tool call, every decision, and the context behind it. Without that record, neither the access grant nor the behavior that followed it can be checked against anything. None of it matters, either, for an agent nobody knew was running. Finding every agent operating inside an environment, including the ones nobody formally registered, is the first control the whole model rests on, not an afterthought bolted onto it later.

Discovery, runtime enforcement, and a full audit trail together give an organization the ability to catch anomalous behavior and policy violations as they happen, instead of reconstructing them weeks later from logs nobody was watching in real time. That's the condition production AI safety actually depends on. The security boundary for an AI agent was never a line drawn once at provisioning. It's a perimeter enforced continuously, moving with the agent at every step it takes.

Sources

  1. The 2026 Singapore Consensus on Global AI Safety Research Priorities
  2. SkillScope: Toward Fine-Grained Least-Privilege Enforcement for Agent Skills
  3. ActPlane: Programmable OS-Level Policy Enforcement for Agent Harnesses
  4. Least privilege for AI agents: Identity, access, and tool binding
Filed underControl Patterns

More in Control Patterns