Safer AI agents: design permissions to limit blast radius

Safer AI agents: design permissions to limit blast radius

Prompt filtering helps, but it is not a trust boundary. Design agents so one failure does not become a production incident.

An AI agent can read data, call tools, and recommend or perform actions. That ability to act is what changes the security problem. Input filtering remains valuable for reducing noise, but it should not be treated as the only wall between a model and important systems.

Start with a permission map

For every agent, list four groups: data it can read, tools it can invoke, actions it can create, and destinations to which it can send data. An agent that reads sensitive data, processes untrusted content, and can send results outside the organization should be treated as a high-risk priority.

Place four controls in the right locations

  1. Separate identity: the agent does not borrow a long-lived user or administrator token. Grant short-lived, task-scoped, revocable permissions.
  2. An action broker: the model proposes structured intent; an independent component validates schema, allowlists, and quotas before execution.
  3. Risk-based approval: deletion, production change, data export, or an action that creates an outside obligation requires human confirmation.
  4. Observability and reversal: retain inputs, tool calls, effective permissions, and results, with a kill switch and a rollback path.

Move through maturity stages

Begin with read-and-summarize mode. Once quality is proven, move to recommendations that people approve; only then consider automation for low-impact, bounded, reversible actions. Avoid jumping directly from a chatbot to infrastructure-changing permissions.

Six production-readiness tests

  • Can the agent access a secret through a prompt or log?
  • Is data from the web, email, or a file labelled as untrusted?
  • Does each tool limit parameters and network destinations?
  • How does the system respond to an expired, revoked, or invalid token?
  • Can an operator reconstruct why an action occurred?
  • Can an erroneous action be stopped and reversed within the agreed time?

The goal is not to make an agent incapable of mistakes. The goal is to keep mistakes contained, visible, and recoverable before they become material harm.

Published ; updated

Related pages