Search
  • en
  • es
  • en
    Search
    Open menu Open menu

    Introduction

    Roger Sanz, PhD · Originally published on LinkedIn

    It was a great pleasure to present The Agentic AI Rebellion in Valencia last September 18 th.

    Many thanks to RootedCON team and amazing speakers Daniel Haro Álvarez , Carles Cano Barba Tomás Isasia Infante , Jorge Martínez Hurtado , Mario Lobo Romero Miguel Benitez, CISSP, CISM , Mar Llambí and Alejandro Lopez (I didn´t find you in Linkedin LoL)

    And thank you to Plain Concepts , Universidad Isabel I and special mention in the inspirations of OWASP GenAI Security Project OWASP AI Exchange AIDEFEND Framework MITRE ATLAS AIUC-1 National Institute of Standards and Technology (NIST) Cloud Security Alliance and more individuals on this adventure.

    The hero of the day is Rock Lambros . Every day (or almost every day), Rock provides a valuable insights into AI agent security landscape. This exemplary behavior is mirrored by many of my colleagues—something I deeply appreciate it.

    Over the last few months I have spent a lot of time working on a question that is becoming increasingly uncomfortable for anyone deploying autonomous AI systems:

    What happens when an agent stops behaving like the agent we designed?

    Not necessarily because someone successfully prompt-injected it. Not necessarily because its model was compromised. And not necessarily because an attacker gained an initial foothold.

    Sometimes the problem is much simpler: the agent has enough identity, permissions, tools and runtime freedom to make a decision that is technically possible but operationally outside the boundary we intended. That was the premise behind my Agentic AI Rebellion lab and the RootedCON VLC talk. I built the lab around five agentic environments and focused on one lifecycle: detect behavioural deviation, decide whether the action is allowed, contain the execution when necessary, and recover the agent to a trusted state.

    The result changed the way I think about agent security. The most important security boundary is not the prompt. It is the execution loop .

    Article content
    Sometimes we are speaking about AI Agent without the correct context LoL.

    The lab started with a simple assumption: autonomy changes the failure mode

    An AI agent is fundamentally different from a conventional application because it combines reasoning with the ability to act. It can interpret context, select a tool, construct arguments, obtain additional information and continue operating without waiting for a human after every step. That creates a very different security problem. The agent may have legitimate credentials. The tools may be legitimate. The API may correctly authenticate the request. The underlying platform may report that the operation is authorised. And the action can still be wrong.

    That pattern is no longer purely theoretical. In the PocketOS incident, an autonomous coding agent working on a staging task found an over-privileged Railway token in an unrelated file and used it to delete a production database and its volume-level backups. The token was valid; the failure was the combination of credential scope, environment separation and the absence of a control preventing a destructive action outside the intended task.

    The July 2026 Hugging Face incident demonstrates the same problem at much greater scale. A frontier-model evaluation agent escaped its sandbox, reached an external launchpad and eventually moved across multiple trust boundaries inside the Hugging Face environment. Hugging Face’s public forensic reconstruction describes roughly 17,600 recovered agent actions during the campaign.

    Those cases are very different technically, but they expose the same architectural weakness:

    the system allowed an autonomous decision to cross a security boundary that existed outside the agent’s own reasoning.

    That is the problem I wanted the lab to explore.

    Article content
    Don´t believe the hype

    What we actually tested

    I deliberately did not build the lab around a single exploit. The objective was to create a baseline of normal agent behaviour and then introduce deviations that mattered from a security perspective. For every environment, I wanted to answer three questions.

    First, what is normal?

    An agent normally follows a relatively stable sequence of operations. It might retrieve information, call a particular tool, transform data and then produce an output. The exact sequence varies, but the system should have an observable behavioural profile.

    Second, what changed?

    A conventional alert often tells us that something suspicious happened. For agents, that is not enough. We need to know which behavioural relationship changed.

    In the lab, I experimented with a behavioural drift model based on the structural difference between a known-good execution graph and the currently observed execution graph. A new relationship such as a credential lookup followed by a production deletion is much more useful to an analyst than an opaque anomaly score because it explains the deviation.

    That leads to the third question:

    What should happen now?

    Once an agent crosses a defined authority boundary, detection alone is not enough. The runtime needs an enforcement decision.

    This is where the lab became much more interesting.

    Detection is only useful when it changes the execution

    The traditional SOC model often looks something like this: detect an event, generate an alert, investigate it and then respond. For autonomous agents, that sequence can be too slow. The agent may complete the dangerous operation before a human has even opened the alert. So the control model has to move closer to the action itself. The lab therefore introduced a runtime decision point that evaluated whether an action was permitted in that specific context. For potentially irreversible actions, the policy could require additional approval or trigger an immediate containment response.

    The response was deliberately broader than “kill the agent”.

    A real containment system may need to block the current action, suspend the session, revoke credentials, isolate the agent or restore its permitted authority boundary.This distinction matters. A security architecture should not assume that every abnormal behaviour requires a total shutdown. Sometimes the correct action is to stop a single operation and restore the agent to its safe operating envelope.
    Microsoft Research
    is moving in the same architectural direction with its 2026 Agent Governance Toolkit, which introduces runtime policy enforcement, agent identity, execution controls and an explicit kill-switch model. Google is similarly treating agent identity, agent gateways, access management, guardrails and runtime defence as an integrated security problem rather than relying solely on model-level controls. This is becoming a broader industry pattern.

    Article content
    Detect does not stop the incident.

    The first principle: identity is not authority in Agentic AI

    One of the strongest lessons from the lab is that giving an agent a unique identity is necessary, but it is not enough. A service principal tells you who is acting. It does not tell you whether that actor should be allowed to perform the specific action it is attempting at that moment. The security architecture therefore needs multiple levels of enforcement: identity, permissions, tool scope, runtime policy and action context.

    This is consistent with the emerging agent-security guidance from OWASP. The Agentic Top 10 explicitly identifies identity and privilege abuse, tool misuse, unexpected code execution, memory/context poisoning and rogue-agent behaviour as distinct risks.

    The practical consequence is straightforward:

    Do not treat the agent’s identity as the final security boundary.

    The identity should establish accountability. The authorization layer should establish capability. The runtime policy should establish whether that capability is appropriate for the action being attempted.

    The second principle: behavioural monitoring has to become first-class security telemetry

    The lab also exposed a weakness in conventional monitoring. Traditional endpoint or application monitoring is usually event-oriented. We look for known indicators, known signatures or abnormal individual events. Agentic systems generate something richer: behavioural sequences.

    The security team therefore needs visibility into relationships between:

    • the agent;
    • its identity;
    • its current task;
    • the context it consumed;
    • the tools it selected;
    • the arguments it passed;
    • the data it accessed;
    • the actions it attempted;
    • and the outcome.

    That is much closer to observing a distributed workflow than monitoring a single application. It also means that the most useful detection primitive may not be a signature. It may be a deviation from an established behavioural graph. This direction is reinforced by MITRE’s work on agentic attack chains. Its public OpenClaw investigation includes demonstrated destructive actions through agent tool invocation and highlights mitigations such as agent tool permissions, human-in-the-loop controls and restrictions on tool invocation when untrusted data influences the reasoning path.

    The third principle: irreversible actions deserve special treatment

    Not every agent action has the same risk. Reading a document is not equivalent to deleting a database. Generating a report is not equivalent to changing production infrastructure. Searching a knowledge base is not equivalent to sending an external payment.

    The security model should therefore understand action criticality.

    For high-impact or irreversible actions, I strongly prefer an explicit external stop condition.The agent should not be able to decide by itself that it is both:

    1. allowed to perform the action, and
    2. allowed to decide that the action is safe.

    That is where human approval, policy enforcement, transaction boundaries, rate limits, environment separation and independent runtime controls become important. This principle is now explicitly appearing in agent-security standards and guidance. OWASP’s emerging Agent Control Standard describes the need for agents to be inspectable, traceable and instrumentable and for safety policies to be enforced through runtime middleware rather than living only in prompts or documentation.

    This is where I use TREASURE

    The result of the lab is also why I increasingly think about agentic security as a defense-in-depth architecture rather than a guardrail product. That is the purpose of the TREASURE framework I have been developing. TREASURE separates the agentic security problem into five security layers.

    DATA — establish a governed source of truth

    The first boundary is the information the agent believes. RAG repositories, memory, retrieved documents and external data sources should not automatically become trusted simply because the agent can access them.Security needs to understand where context came from, who owns it, how it changed and whether it is allowed to influence the agent. This is where provenance, source validation, memory lineage and poisoning controls matter.

    The goal is not merely to protect data. It is to control what information is allowed to become agent context.

    EXECUTION — control the actual attack surface

    The next layer is execution.Every skill, tool, connector, MCP server, API and plugin effectively increases the agent’s capability.This is where the architecture must answer more than “is this tool authenticated?”

    It needs to answer:

    Is this tool allowed for this agent, for this task, with this data, using these arguments, in this environment, at this moment?

    OWASP’s current guidance around role/task-based access and zero-trust enforcement between agents, tools and external APIs is moving strongly in this direction.

    GOVERNANCE — turn policy into something executable

    One of the recurring problems I see in AI security is that organisations have written policies describing what agents should do, but very little that actually enforces those policies.

    TREASURE treats governance as policy-as-code discipline.

    An Agency Envelope should become a concrete security artefact defining the permitted purpose, tools, data access, action classes, autonomy level and escalation conditions of an agent. That envelope can then feed runtime policy.

    The important transition is:

    policy document → machine-enforceable boundary.

    Microsoft’s current agent security guidance makes the same point from another angle: as agents move from pilots into production, governance has to become observable, controlled and auditable throughout the lifecycle.

    CONTROL — encapsulate autonomy

    This is where safe autonomy is created. I do not think the answer is to remove autonomy from agents. The more interesting approach is to bound it. The agent can reason and act inside a defined authority envelope, but external controls determine whether a specific action is allowed. High-risk actions can require human confirmation, additional policy evaluation or other runtime controls.

    The agent remains autonomous inside the boundary. The boundary itself is not autonomous.

    SECURITY — engineer the blast radius

    The final layer assumes that something will eventually go wrong. Containment therefore has to exist independently of the agent. That means network segmentation, ephemeral credentials, narrow permissions, sandboxing, runtime monitoring, session revocation, kill-switches and recovery mechanisms. The objective is not to prove that the agent can never fail.

    The objective is to ensure that a failure does not automatically become a production incident.

    That is also consistent with the direction of current agent-security engineering: Microsoft describes runtime execution controls, dynamic trust, kill-switches and agent SRE practices; Google is introducing agent identity and runtime defence as dedicated infrastructure capabilities.

    What I would recommend to anyone building agents today

    The first thing I would do is establish behavioural observability before adding sophisticated detection technology. You cannot detect meaningful deviation without understanding what normal behaviour looks like. Then I would create an explicit authority model for every production agent. Document the tools it can use, the data it can access, the actions it can take and the actions that require additional approval. This should become an executable control boundary, not just documentation. I would also separate identity from authority. Give every agent an attributable identity, but assume that identity alone is insufficient to constrain behaviour. For destructive or externally consequential actions, introduce an independent enforcement point. An agent should never be the sole authority deciding whether its own irreversible action is permissible.

    Finally, design containment and recovery before production deployment. Decide how you stop an agent, revoke its credentials, terminate its execution context, isolate its tools and restore its authorised state. Then test those mechanisms.

    The uncomfortable conclusion

    The industry has spent the last few years asking whether an LLM can be made safe.

    For agentic systems, that question is incomplete. The more important question is:

    Can we keep the system safe when the model behaves differently from what we expected?

    That is a very different engineering problem. The recent real-world incidents, the emerging OWASP agentic-security guidance, MITRE’s agent investigations and the runtime security work now appearing across Microsoft and Google all point in the same direction: agent security is becoming an execution-control problem, not just a model-safety problem.

    That is why I think defense-in-depth is the right mental model.

    Govern the data. Control execution. Make policies enforceable. Bound autonomy. Engineer the blast radius.

    Take your time to design security by design principles

    That is the idea behind TREASURE. Not a promise that agents will never go rogue.

    A way to make sure that when they do, we see it, understand it, stop it and recover safely.

    See you in the next adventure

    Roger Sanz

    AI Security Governance Lead