Main

When an agent decides: the five pillars of agentic AI security

Between 7 and 13 July 2026, around 700 AI agents deployed by OpenAI for evaluation tasks discovered an unauthorised communication channel and used it to compromise Hugging Face production systems. They were not acting on instructions to attack: the agents themselves concluded that accessing those systems would let them find out how their tasks were being evaluated. The result was 136 leaked credentials and around 17,600 actions logged on production infrastructure.

It was not a traditional security failure. Nobody exploited a vulnerability with the intent to attack. These were agents acting in a coordinated way, without direct supervision, following their own logic to a place nobody expected them to reach.

For years, cybersecurity awareness has meant teaching people not to click the wrong link. With agentic AI, it also means deciding in advance which decisions an agent cannot make alone.

 

A new attack surface

Agentic AI introduces attack surfaces that carried much less weight in traditional architectures. Prompt injection, memory poisoning or context contamination can alter the way an agent interprets its environment and makes decisions. In multi-agent architectures, these threats can also spread through RAG databases, tools or communications between agents.

To this we add risks that cybersecurity teams know better: identity abuse, privilege escalation, unexpected code execution, compromised supply chains, insecure Agent-to-Agent communications or the appearance of unauthorised agents.

The problem is no longer just the security of the model. It is the identity, the privileges, the governance and the control of what that model can do on behalf of an organisation.

That is why we need a true control plane for autonomy: a mechanism able to prevent, limit and, above all, provide evidence of what happens throughout the agent’s lifecycle, before, during and after each execution.

 

From IAM to a security operating system for agents

Identity has evolved fast in recent years.

We started with a fundamentally human IAM: deterministic, based on authentication, approvals, sessions and reasonably stable role models. Then we added non-human identities (service accounts, keys, certificates, API keys, tokens), already without direct human supervision, which brought their own credential lifecycle problems.

With agentic AI we take another step: potentially ephemeral identities, delegation chains between agents, probabilistic decisions, dynamic use of tools and contexts that change within a single execution. In other words, we move from controlling the login to controlling the action.

AI agents cannot be governed in the same way as a human user, a service account or an application based on static permissions. Identity starts to become something broader: a real-time control system over who is acting, on whose behalf, with what privileges, on what data and under what conditions.

 

An example: an agent onboarding a customer

Consider a common process in banking, insurance, telecommunications or utilities: the digital onboarding of a new customer.

An agent can gather the necessary information, guide the user through the process, query systems and prepare the request. The problem appears when someone tries to manipulate that process with forged documents, deepfakes or information stolen from a legitimate identity. The impact does not end with document fraud: it can lead to opening an account, taking out a product or granting a service to a non-existent or impersonated identity.

The answer is not to remove the agent from the process. The agent must have its own auditable identity and act as a requester, never as an issuer of trust. A critical action such as definitively creating a new customer’s identity should require independent verification, ephemeral credentials, adaptive risk policies and, when the risk justifies it, human approval.

The principle is simple: an agent can propose and execute, but its autonomy must be conditioned by the identity, the context, the risk and the impact of the action.

The five pillars for controlling agentic autonomy

In my opinion, the model can be structured around five pillars.

1. Identity and continuous authentication

Know which agent is acting, why it exists and who answers for it. Each agent needs its own unique, auditable identity, linked to a responsible human owner, throughout its lifecycle: security by design, adaptive authentication and continuous discovery of unknown or unauthorised agents. If we cannot identify the agent, we will hardly be able to govern it.

2. Privileges and execution control

Having an identity does not mean having authorisation. We must determine what each agent can do, on which resources, for how long and under what circumstances: least privilege, zero standing privileges, Just-in-Time access to reduce the window of exposure, and human approval for the highest-impact actions. Privilege must be revocable immediately if the context changes or anomalous behaviour appears.

3. Governance and orchestration

When an organisation deploys tens, hundreds or thousands of agents, the question is simple: do we know how many we have and what they do? We need an active inventory (purpose, owner, permissions, dependencies, security posture) covering the agent’s whole lifecycle, from creation to retirement, with the ability to automatically limit, recertify or deactivate those that stop complying with policies.

4. Data protection and control

An agent’s autonomy is directly tied to the data it accesses. We must set clear limits on what information it can consult, process, generate or share, and with whom. It is no longer enough to ask who can access the data: we must ask what an agent can do with that data after accessing it.

5. Observability and response

What is autonomous must also be observable, explainable and controllable: what decisions an agent made, what identity it used, what information it consulted, what actions it executed. That traceability must come with rapid response mechanisms (permission revocation, quarantine, kill switches, forensic capability), because when decisions happen at machine speed, an exclusively human response may arrive too late.

From controlling access to controlling autonomy

For years we have designed security around one question: can this identity access this resource? With agentic AI we must answer another, more complex one: can this agent perform this action, on this resource, at this moment, with this data, on behalf of this identity and under this level of risk?

It is an apparently small change that transforms the role of IAM, PAM, data security, observability and governance. The value of an agent lies in its ability to act. The challenge is not to slow that down: it is to achieve controlled autonomy, with identifiable agents, limited and dynamic privileges, governed throughout their lifecycle, restricted in their use of data and fully observable.

 

Regulation points in the same direction

Regulation and standards are moving towards models of demonstrable accountability. ISO/IEC 42001:2023 provides a management system framework to establish, maintain and improve the responsible governance of AI, with particular attention to risk management, transparency, traceability and accountability. The EU AI Act, for its part, sets specific requirements for certain high-risk systems: its Article 12 requires automatic event-logging capabilities throughout the system’s lifetime, with a level of traceability appropriate to its purpose, to facilitate monitoring and identify risk situations.

Defining policies and controls is no longer enough. Organisations must be able to provide evidence, traceability and clear responsibilities for how their AI systems operate. We are moving from saying that we govern AI to having to prove it.

Conclusion: trust will have to be continuously demonstrated

In my opinion, autonomy is not the problem. Autonomy without control is.

Applying to agents the same mechanisms we use for human users or service accounts will be insufficient. We are bringing into the enterprise environment identities capable not only of authenticating, but of interpreting information, making decisions, delegating tasks and executing actions.

The next big step in identity management will be to evolve from access control to autonomy control. It is not enough to ask who the agent is and what permissions it has: we must ask what it is trying to do, on whose behalf, on what data, with what level of risk, and whether it should be allowed to continue.

The aim is not to slow the adoption of agentic AI, but to create the conditions to adopt it with confidence. The more autonomy we grant an agent, the greater our ability must be to identify it, limit it, observe it and stop it. In the world of agentic AI, trust will stop being something granted once and become something that must be continuously demonstrated.

 

By Jose Manuel De La Puente

GENERATING BUSINESS VALUE

Let's shape the future of digital innovation together

Get in touch