For most of the chatbot era, the security boundary was relatively easy to understand.
A user sent some text.
A model returned some text.
The output could still be wrong, unsafe, or sensitive. But in most cases, the model itself was not directly changing production infrastructure, deleting a database, modifying a repository, or running a shell command on a developer's laptop.
Agents change that.
An AI agent can decide to:
rm -rf ./migrations
Or:
database.drop_table("customers")
Or:
aws.iam.delete_role(...)
Or call an MCP tool, modify a repository, retrieve a secret, open a network connection, trigger a deployment, or invoke another agent.
The security problem is no longer limited to what the model says.
We now need to care about what the model is actually allowed to do.
That is where AI agent runtime security begins.
The new security boundary is the action
Think about a typical agent loop.
The model receives some context, reasons about what to do next, selects a tool, constructs its arguments, receives the result, and continues.
Somewhere inside that loop, something important happens:
Intent turns into execution.
Before that moment, the agent is mostly reasoning.
After that moment, something in the real world may have changed.
A file may be gone.
A production configuration may have been modified.
A secret may have been retrieved.
An API request may have been sent.
A deployment may have started.
That moment is what we can think of as the action boundary.
For security teams, it is becoming one of the most important boundaries in an agentic system.
A simplified flow looks like this:
User / Environment
↓
Model
↓
Agent decides an action
↓
Structured tool call
↓
POLICY CHECK
↓
Allow / Warn / Deny / Approve
↓
Tool
↓
Real-world effect
The closer a security decision happens to this boundary, the better chance it has of preventing the side effect rather than merely explaining it afterward.
A request moves from the user and environment through the model, an agent decision, and a tool call. Runtime policy is the checkpoint: allow, warn, or deny before a tool, API, shell, or MCP call can produce a real-world side effect.

Model safety and action security are different problems
A lot of AI security work starts around the model.
Teams scan prompts.
They look for jailbreak attempts.
They detect prompt injection.
They redact PII.
They moderate output.
They filter unsafe content.
Those controls are useful.
But they answer a different question.
A model guardrail might ask:
Is this input or output suspicious?
Runtime authorization asks:
Is this action actually allowed?
Those questions are related, but they are not interchangeable.
Consider the earlier command:
rm -rf ./migrations
Nothing about the command is inherently malformed.
There may even be a perfectly reasonable chain of reasoning behind it. Perhaps the agent believes those migrations are obsolete. Perhaps the user asked it to clean up the project.
But the important security questions are contextual:
- Which agent is running this command?
- Which repository is it operating in?
- Is this a development laptop or a production environment?
- Is
./migrationsconsidered a protected path? - Is deletion permitted?
- Does this particular action require approval?
That decision should not depend entirely on whether another language model thinks the command looks dangerous.
It should depend on policy.
OWASP makes a similar distinction in its guidance around Excessive Agency. The risk appears when an LLM-based system has more functionality, permissions, or autonomy than it needs. Its mitigation guidance includes limiting available tools and permissions, requiring approval for high-impact operations, and enforcing authorization in downstream systems instead of relying on an LLM to decide whether an action should be permitted. OWASP's Agentic Security Initiative treats that class of risk as its own problem.
In other words:
Do not ask the agent to police its own authority.
A model guardrail asks whether the content is safe: prompt injection, PII, jailbreaks, unsafe output. Runtime policy asks whether the action is allowed: file.delete, shell.execute, database.drop, cloud.modify.

Security after execution is already too late
Observability matters.
You want to know:
- what an agent attempted,
- what tool it selected,
- which arguments it supplied,
- who initiated the task,
- what resource was affected,
- and what eventually happened.
But an audit log and a security control are not the same thing.
Suppose an agent deletes a production table and your monitoring system records:
14:32:08
agent=deploy-agent
tool=database
operation=drop_table
resource=customers
status=success
That is excellent evidence.
It is not prevention.
The table is still gone.
For actions with meaningful side effects, security needs an opportunity to intervene before the underlying tool executes.
The flow becomes:
Agent proposes action
↓
Action + context captured
↓
Policy evaluated
↓
Decision made
↓
Only allowed actions execute
This distinction becomes increasingly important as agents move from assistants that recommend changes to systems that autonomously perform them.
An audit log records a destructive command after the database is already gone. A runtime policy checkpoint can block that same command before execution, so the database stays intact.

A tool name is not enough context
Runtime security also needs to operate at a finer level than:
Allow database
Deny shell
Allow MCP
That quickly becomes impractical.
A shell is not inherently unsafe.
A database is not inherently safe.
An MCP server is not inherently trustworthy or untrustworthy.
What matters is the concrete operation being attempted.
For example:
database.read
resource: production/orders
may be allowed.
While:
database.drop
resource: production/orders
may be denied.
The same applies to files.
An engineering agent might legitimately need:
read ./migrations/**
while having no reason to perform:
delete ./migrations/**
A useful runtime decision may therefore incorporate several pieces of context:
Identity
Which agent, workload, developer, service account, or delegated user is acting?
Action
What operation is actually being requested?
Arguments
What parameters are being supplied?
Resource
Which file, database, API endpoint, cloud resource, repository, or secret is affected?
Environment
Is the action happening in local development, staging, or production?
Data sensitivity
Does the action touch credentials, PII, financial data, source code, or other protected information?
Approval state
Does this operation require a human before it can continue?
A policy engine can use these signals together.
That is much closer to traditional authorization than to content moderation.
Probabilistic systems need deterministic boundaries
There is an interesting tension at the center of agent security.
The agent is probabilistic by design.
Security policy often should not be.
An LLM is useful when there is ambiguity.
It can interpret instructions, generate code, choose tools, summarize information, or reason about an unfamiliar problem.
But many organizational security rules are not ambiguous at all.
For example:
Coding agents must not read ~/.ssh/**
or:
Agents cannot delete production databases.
or:
Changes to IAM policies require human approval.
or:
Customer PII cannot be sent to an unapproved model provider.
You do not necessarily need another LLM to decide whether those rules apply.
The organization already made the decision.
The runtime system just needs to enforce it.
That produces a clean division of responsibility:
AI:
"What should I do next?"
Policy:
"Are you allowed to do it?"
This separation also improves explainability.
Instead of receiving:
BLOCKED
Because an AI security classifier believes this request may be unsafe.
a developer can receive something much more concrete:
BLOCKED
Rule: protected_paths
Action: file.delete
Resource: ./migrations/**
Reason: Migration files are release-critical.
A reason is more useful than a mystery.
Runtime security is a stack, not a feature
There is no single control called "agent security" that solves everything.
Several layers work together.
Identity
Who is performing the action?
An agent operating for one developer should not automatically inherit the authority of an administrator.
Authentication
Can the system prove the identity making the request?
Authorization
What resources and operations can that identity access?
Runtime action policy
Even if the identity can access a system, is this particular operation acceptable under the current conditions?
Data controls
Can sensitive information cross this boundary?
Should a secret be masked?
Can customer data be sent to this model?
Human approval
Should a person explicitly approve a high-impact operation?
Auditability
What action was proposed?
Which policy matched?
What decision was returned?
Did execution happen?
These layers solve different problems.
Strong authentication does not make every authenticated operation safe.
Perfect logging does not stop destructive behavior.
A system prompt telling an agent to "never delete production resources" is not the same thing as authorization.
Each layer has a job.
Coding agents make the problem obvious
Coding agents are a useful example because they sit surprisingly close to high-value resources.
Depending on their configuration, they may be able to:
- read an entire repository,
- modify source code,
- execute shell commands,
- inspect environment variables,
- call MCP servers,
- access GitHub,
- invoke cloud APIs,
- interact with CI/CD systems,
- or create additional agents.
Most of these capabilities are useful.
The objective is not to remove the agent's capabilities.
It is to constrain them.
Imagine a policy for a repository containing release-critical migration files.
You might allow:
action: file.read
resource: ./migrations/**
while denying:
action: file.write
resource: ./migrations/**
and:
action: file.delete
resource: ./migrations/**
The coding agent can still understand the migrations.
It can inspect them.
It can reason about them.
It simply cannot silently rewrite or delete them.
The agent proposes rm -rf ./migrations. Policy stops it at the boundary: deny, rule protected_paths, because migration files are release-critical.

Another organization might permit production database reads while requiring explicit approval for writes.
Another may block access to:
~/.ssh/**
.env
**/credentials.json
The important part is that these limits exist independently of what the model decides.
That reduces the blast radius of hallucinations, prompt injection, compromised context, poorly scoped instructions, and simple mistakes.
This also changes how we think about least privilege
Historically, least privilege has often meant giving an application the minimum infrastructure permissions it needs.
Agents make this more complicated.
One agent process may perform dozens of different kinds of work.
Giving the process access to a credential can sometimes implicitly give every action produced by that process access to the same authority.
Runtime policy lets us introduce another level of restriction.
Instead of only asking:
Can this process access GitHub?
we can ask:
Can this agent call GitHub's merge operation
on this repository
from this environment
without human approval?
That is a much stronger security boundary.
Where Control Zero fits
This is the problem Control Zero is designed around.
Control Zero evaluates actions against policy at runtime and supports decisions such as allow, warn, or deny before execution on supported enforcement surfaces. Its current integrations include coding-agent hooks, SDK-governed actions, gateway-based governance, and additional surfaces such as browser data controls.
Coding-agent hooks, SDK actions, the AI gateway, and MCP tools all propose actions to the policy engine. Identity, action, resource, and context decide allow, warn, or deny, and every decision is audited. Only a permitted action executes.

For coding assistants, Control Zero attaches to pre-execution hooks exposed by the underlying agent host.
Where the host provides an enforceable hook, a denied action can be stopped before the tool executes. Control Zero explicitly records enforcement coverage per event because different coding-agent hosts expose different hook capabilities.
That distinction is important.
A security product should not claim that an action can be blocked when the underlying platform never provides that action to the security layer.
Coverage is part of the security model.
Control Zero policies can also run locally. In its local-only mode, policies are loaded from a local YAML or JSON file, evaluated in-process, and do not require network access to the Control Zero backend.
A rule might conceptually look like:
- effect: deny
action: "file:delete"
resource: "./migrations/**"
reason: "Migration files are release-critical"
The exact policy will vary by environment, but the architecture remains the same:
Agent proposes something
↓
Control Zero sees the action
↓
Policy is evaluated
↓
ALLOW / WARN / DENY
↓
Allowed action executes
The model can remain flexible.
The security boundary does not have to be.
Start with the rules you already know
Organizations do not need hundreds of agent-specific policies on day one.
Start with the rules that are already obvious.
For example:
Never allow agents to read private SSH keys.
Never allow an autonomous agent to delete
production databases.
Require approval before modifying IAM.
Protect release-critical repository paths.
Prevent secrets from being sent to
unapproved external services.
These rules are not model problems.
They are authorization problems.
Once those boundaries are enforced, teams can gradually add more context and finer-grained policies as agents gain additional capabilities.
A useful mental model
When thinking about the security architecture around agents, this model is helpful:
Prompts
Shape what the agent is likely to do.
Identity
Establishes who is acting.
Permissions
Define what systems the identity can access.
Runtime policy
Decides whether this specific action is allowed.
Audit
Records what was proposed and what happened.
Each layer matters.
But as agents gain the ability to affect real systems, runtime action control becomes increasingly important.
Because eventually every autonomous system reaches the same moment:
The model has finished reasoning.
It has selected its next action.
It is about to turn intent into reality.
That is the moment security needs to be there.
Your AI can be probabilistic. Your security policy doesn't have to be.
Agents will continue getting more capable.
They will receive access to more tools, more APIs, more infrastructure, more data, and longer-running workflows.
Trying to eliminate every bad decision at the model layer is unlikely to be enough.
A safer architecture assumes that the agent will occasionally make a bad decision.
Then it asks a more practical question:
What is the agent allowed to do when that happens?
Define the boundary.
Evaluate the action.
Enforce the decision before execution.
The agent can decide what it wants to do.
It should not get to decide what it is allowed to do.
