Ibrahim Mukherjee
London based entrepreneur, cybersecurity analyst doing a PhD in AI

Risk Guidance for AI Agents – Swiss Cheese Model.

NVidia has released a new model for AI Agent safety. 

It doesn’t go far enough. So let’s analyse and find pathways for safer AI deployment in critical conditions. I will present a model I call the Swiss Cheese model for AI Agent safety. Calling on diverse disciplines like systems engineering, redundancy from computer engineering, as well as how we govern humans in the workplace – the mission is to find a sustainable architecture for deploying AI agents in enterprise.

Before we decide how to control AI agents, we need to understand what an agent actually is.

The problem is the terminology makes this sound much more complicated than it needs to be.

There are really three steps in the story: software, Large Language Models, and AI agents.

The easiest way to understand the difference is this.

Software does what it says on the tin.

An LLM is a thinking layer without inherent access.

An AI agent is that thinking layer connected to tools that let it retrieve information and do things.

That final step changes the risk.

Because the moment an AI can do something, rather than merely tell us something, we have given it authority.

And that leads to the central question of agentic AI:

What can this AI do without asking?

1. Software: Does What It Says on the Tin

Start with ordinary software.

Imagine a calculator.

You press:

2 + 2

It returns:

4

Or imagine payroll software.

A programmer might write a rule saying:

If salary = £50,000, calculate monthly gross salary as £50,000 ÷ 12.

The computer executes the rule.

But its basic nature is relatively easy to understand.

We write instructions. The computer executes them.

That also gives us familiar ways of testing it.

Does the button work?

Does this input produce the expected output?

Does a £5,001 transaction trigger the approval rule?

Engineers can write unit tests and integration tests. They can inspect the code and compare the result against the specification.

Then came the LLM.

2. The LLM: A Thinking Layer Without Access

A Large Language Model works differently.

Instead of telling the computer exactly how to respond to every possible question, we give the model information and ask it to work something out.

Suppose you paste an email into an LLM:

“ABC Ltd has not paid invoice 3287. It is now 35 days overdue. What should I do?”

The model might answer:

“Send the customer a reminder, ask for payment within seven days and escalate the invoice if payment is not received.”

Nobody explicitly programmed that exact answer.

The model interpreted the information and generated a response.

That is enormously useful.

But notice something important.

Nothing has actually happened.

The customer has not received an email.

The accounting database has not changed.

No payment has moved.

The LLM has produced information.

It needs something else.

It needs tools.

3. The Agent: Give the Thinking Layer Hands

This is the easiest way to understand an AI agent.

An AI agent is essentially a thinking layer given tools.

The tools become its digital hands.

Give it an email tool and it may be able to read or send email.

Give it a database tool and it may retrieve or modify records.

Give it a browser and it can potentially search websites.

Give it access to a CRM and it can potentially update customers.

Give it a terminal and it may run commands.

Give it cloud-management tools and it may modify infrastructure.

Give it an interface to a robot and eventually its actions can reach the physical world.

4. So What Exactly Is a “Tool”?

The word tool can make the process sound mysterious.

It isn’t..

Technically, these tools are usually ordinary software interfaces.

Often they are APIs: structured ways for one computer system to ask another computer system to perform an operation.

The important point is this:

The LLM normally does not directly perform the external action.

It asks software to perform the action.

That distinction is crucial for understanding both agents and their safety.

5. What Actually Happens During a Tool Call?

Suppose the agent needs to find overdue invoices.

The process can be reduced to a few simple steps.

The organisation first tells the model which tools exist.

For example:

Tool: SearchInvoices

Purpose: Find invoices matching specified conditions.

The user then says:

“Deal with invoices more than 30 days overdue.”

The LLM reasons that it needs information from the accounting system.

So instead of answering the user, it effectively says:

I want to use SearchInvoices.

And it supplies the necessary parameters:

Days overdue: greater than 30.

At this point something extremely important happens.

The model has requested an action. It has not necessarily performed it.

The application surrounding the model receives the request.

That application can check:

Is this agent allowed to use SearchInvoices?

Is this user allowed to ask the agent to do this?

Are these parameters permitted?

Does this action require approval?

If the checks pass, ordinary software calls the accounting system.

The accounting system returns:

Invoice 3287 — ABC Ltd — £4,200 — 35 days overdue.

Invoice 4192 — XYZ Ltd — £1,750 — 43 days overdue.

Those results are then given back to the LLM.

The model can now reason again.

OpenAI’s current function-calling documentation describes almost exactly this cycle: provide the model with available tools, receive a tool-call request, execute the relevant code outside the model, return the tool output, and allow the model to respond or request further tools. (OpenAI Developers⁠)

So:

Model asks → software checks → tool acts → result returns → model decides again.

That is the basic machinery of an AI agent..

6. Read and Write Are Very Different Things

This immediately reveals where the risk appears.

Suppose an agent has an email tool.

That statement tells us almost nothing.

Can it search email?

Can it read email?

Can it draft email?

Can it send email?

Can it forward attachments?

Can it delete email?

Can it change mailbox rules?

Those are radically different levels of authority.

Likewise, “database access” could mean:

Read one customer record.

Or:

Read every customer record.

Or:

Modify records.

Or:

Delete the database.

The model might be identical in every case.

The difference is the authority attached to its tools.

That leads to the defining operational variable.

7. Agent Risk = Delegated Authority

Imagine exactly the same model in three environments.

The first can read five documents.

The second can read and send corporate email.

The third can deploy software, alter infrastructure and initiate financial transactions.

The intelligence has not changed.

But clearly the risk has.

What changed was the authority surrounding the intelligence.

So the simplest useful question for agentic AI is:

What can it do without asking?

That is its delegated authority.

AI has many other risks: privacy, bias, misinformation, cybersecurity, reliability and more.

But the distinctive operational change from an LLM to an agent is that we have connected intelligence to authority to act.

This is why capability and authority must never be confused.

Capability is what the AI knows how to do.

Authority is what we permit it to do.

A model may know how to transfer £10 million.

Its credentials do not need to permit a £10 million transfer.

A model may know how to delete a database.

Its database account does not need DELETE permission.

A model may understand cloud administration.

It does not therefore need administrator credentials.

Humans understand this already.

A junior accountant might perfectly understand a £20 million transaction.

That does not mean the company gives them authority to approve it.

8. AWS: Control the Tool Call

This is why current AWS guidance concentrates so heavily on what happens between the agent and its tools.

AWS says agents interact with external systems through APIs, databases, files and other services. It recommends that every tool invocation be authorised against policy before execution, with high-risk actions stopped for human approval.

AWS also recommends dedicated agent identities, explicit tool scopes, input validation, rate limits and controls enforced outside the agent’s reasoning loop. (AWS Documentation⁠)

That last point matters enormously.

Do not merely tell the model:

“Please don’t delete anything important.”

Remove the DELETE permission.

The model can request the action all day.

The surrounding system should still say:

No.

9. Microsoft: Give the Agent an Identity

Microsoft approaches the same problem through identity.

Its current Entra guidance makes a simple but important shift: once agents can plan workflows, call tools and access enterprise information, organisations need to ask not simply whether an agent can complete a task, but whether it should be permitted to perform each action, against which resource, and under whose authority. (Microsoft Learn⁠)

Microsoft recommends unique agent identities, named owners or sponsors, narrowly scoped permissions, tool allowlists, logging and tested revocation mechanisms.

This points towards something I believe every consequential agent will eventually need:

An AI Passport.

When a human employee enters a corporation, they receive an identity.

The company knows:

Who are you?

What is your role?

Who is your manager?

Which systems may you enter?

What can you change?

What requires approval?

When does your access expire?

Agents need the same.

An AI Passport should identify the agent, its human owner, purpose, tools, information access, delegated authority, approval requirements and expiry.

If something possesses authority, it needs identity.

10. OpenAI: Keep Approval Between Thinking and Acting

OpenAI’s current agent-safety guidance highlights another crucial point.

External information can manipulate the thinking layer.

Imagine an agent reading an email containing hidden or malicious instructions intended to make the AI behave differently.

That is prompt injection.

If the AI merely summarises the email, the damage may be limited.

But if the same model can subsequently invoke powerful tools, manipulated information can potentially become manipulated action.

OpenAI therefore recommends techniques including structured data boundaries, guardrails, tool approvals and human confirmation of operations. Its current guidance says tool approvals should remain enabled for MCP operations. (OpenAI Developers⁠)

Again, notice where the safety boundary sits:

LLM thinks.

↓

LLM requests tool.

↓

Independent check.

↓

Tool acts.

The space between request and execution is one of the most important places in agent safety.

11. NVIDIA: Put a Wall Around the Whole Thing

And this is why NVIDIA’s announcement on September 28 matters.

NVIDIA has unveiled the Open Agent Safety Platform, consisting principally of OpenShell and the Sentry reference architecture.

OpenShell creates a secure runtime boundary around the agent.

NVIDIA says it can govern what the agent can access and change, network communication, runtime credentials and interactions with outside systems. Its supervisor launches the agent as a restricted process and enforces policy around process identity, filesystem access, network egress and credentials. (NVIDIA Investor Relations⁠)

In plain English:

The AI can ask. OpenShell decides whether the outside world will let it.

That is a fundamental distinction.

A model-level safeguard influences what an agent tries to do.

A runtime boundary controls what the agent is actually allowed to do.

NVIDIA explicitly makes this distinction in its current documentation. (NVIDIA⁠)

But NVIDIA has added another layer.

12. Sentry Watches From Outside

NVIDIA Sentry is designed as an out-of-band watchdog running on BlueField-4 DPUs.

That means the monitoring and enforcement can sit outside both the agent and its host software.

NVIDIA says Sentry can continuously inspect agent activity, verify identity, enforce granular access policies and quarantine an agent in milliseconds if it attempts to move outside its boundaries. (NVIDIA Investor Relations⁠)

This is significant because it creates another independent control.

Imagine:

Agent

↓

requests tool

↓

OpenShell

checks runtime policy

↓

tool accesses system

while independently:

Sentry watches the activity from outside.

If something goes wrong inside the agent, OpenShell remains.

If the host itself is compromised, Sentry is intended to provide another isolated boundary.

NVIDIA has effectively taken the Swiss Cheese Model and pushed it down into the computing stack.

13. Why Swiss Cheese?

James Reason’s Swiss Cheese Model gives us a remarkably useful way to think about agent safety.

Imagine several slices of Swiss cheese stacked together.

Every slice represents a safety control.

Every slice has holes.

No safeguard is perfect.

But the holes normally occur in different places.

A failure passing through one layer gets caught by another.

A catastrophe requires enough holes to line up.

Agentic AI needs exactly this philosophy.

The model’s instructions are one slice.

Tool permissions are another.

The AI Passport is another.

Human approval is another.

Transaction limits are another.

OpenShell can be another.

Sentry can be another.

Monitoring is another.

Rate limiting is another.

A kill switch is another.

The objective is not one perfect safeguard.

The objective is that one failure is not enough.

14. Protect Against the Risk You Haven’t Imagined

This becomes particularly important because agents are useful precisely because humans do not specify every intermediate step.

We give an agent an objective.

It works out a route.

Usually that flexibility is the benefit.

But occasionally it may find a route we did not expect.

Security teams can test known risks such as prompt injection, excessive permissions and privilege escalation.

But no checklist can contain every future failure.

Call the remainder unprepared risk:

the failure we did not know enough to prepare for specifically.

Swiss Cheese provides a defence against that uncertainty.

Perhaps the model safeguard fails.

But the tool lacks permission.

Perhaps the permission is wrong.

But the transaction limit stops the action.

Perhaps that fails.

Human approval intervenes.

Perhaps the human makes a mistake.

Independent monitoring detects abnormal activity.

Perhaps something still goes wrong.

Rate limiting stops one mistake becoming one million mistakes.

Multiple controls at multiple points can cancel out failures we never predicted individually.

15. Recursive Self-Learning Creates the Final Problem

Now add one more development.

Agents may increasingly modify their own software, strategies or workflows, test the modification and retain successful improvements.

This is often discussed as recursive self-improvement or recursive self-learning.

That does not automatically make an agent dangerous.

An accounting agent might become better at accounting.

A coding agent might improve its coding workflow.

A research agent might improve how it searches.

The critical boundary is different:

Self-improvement must never silently become self-authorisation.

Imagine an agent allowed to spend £100.

Fine.

Now imagine it can modify the configuration containing the £100 limit.

The problem has fundamentally changed.

The agent no longer merely possesses authority.

It possesses authority over its own authority.

16. Never Let the Agent Control the Controls

That gives us the meta-rule.

An agent may propose a new tool.

It may request additional access.

It may propose increasing a spending limit.

It may even propose a better version of itself.

But another independent authority must decide whether that change becomes operational.

The agent should not control its identity system.

It should not control its final permission boundary.

It should not control the independent monitor watching it.

It should not be able to disable its own kill switch.

Think of these as the agent’s constitution.

The agent can operate under the constitution.

It cannot unilaterally rewrite the constitution.

This is why NVIDIA’s approach matters so much.

The company is deliberately moving important enforcement outside the agent process, and Sentry moves another layer outside the host software. (NVIDIA⁠)

17. Where Erasys ClearFrame Fits

We can now see the whole architecture.

OpenAI helps define how models interact with tools and where approvals and safeguards can sit.

AWS provides infrastructure for controlling tool calls, identities, permissions and runtime access.

Microsoft provides identity and authorization structures so agents can be governed as first-class actors inside organisations.

NVIDIA is now adding external runtime and hardware-isolated enforcement around agent execution.

These solve important pieces of the engineering problem.

But somebody still has to answer the organisational question:

How much authority should this particular agent receive?

That is where Erasys ClearFrame can sit.

Not as another agent.

Not as another model.

But as the governance frame above the technical controls.

Give every agent an AI Passport.

Record its identity.

Record its human owner.

Record its purpose.

Record its tools.

Record what information those tools can retrieve.

Record what those tools can change.

Record what requires approval.

Record when the authority expires.

And classify the agent according to its authority.

An information-only agent sits at the bottom.

An agent making small reversible changes sits above it.

An agent capable of modifying production systems, moving significant money or controlling physical infrastructure sits much higher.

More authority means more independent slices of cheese.

ClearFrame defines where the boundaries should be.

Microsoft, AWS, NVIDIA and other infrastructure can help enforce them.

18. The Whole Agent in One Minute

Strip away all the terminology and an AI agent is surprisingly understandable.

A human gives it a goal.

The LLM thinks about the goal.

If it needs information, it requests a tool.

Software checks whether that tool call is permitted.

A connector retrieves information from email, a database, a file, a browser or another system.

The information returns to the model.

The model thinks again.

If something needs changing, it requests another tool.

The permission layer checks again.

The tool performs the approved action.

The result returns.

The agent observes what happened.

Then it decides what to do next.

And the loop continues.

That is agentic AI:

Think → Request → Check → Act → Observe → Think again.

The safety architecture should therefore surround every part of that loop.

The Agents Amongst Us

We probably cannot predict every strange thing increasingly capable agents will eventually try. The permutations and combinations are too many. But we can narrow the subset of “catastrophic failures” considerably using this model.

We can do this by doing something much simpler.

We can control what happens when they try it.

That is what aviation, banking, cybersecurity and industrial safety learned long ago.

Do not rely on one perfect layer.

Build several independent ones.

Know who possesses authority.

Know where that authority ends.

And keep the final authority over those limits somewhere outside the thing being governed.

Because in the agentic era, the defining safety question may ultimately be remarkably simple:

What can it actually do without asking?

About the Author
Ibrahim Mukherjee is a London-based entrepreneur, PhD researcher in AI at Brunel, University of London, and founder of the UK's first 'Sovereign AI' initiative Fahm.uk. Voted Outstanding Innovator of the Year 2025 by the AI Journal, he runs Erasys (behavioural biometrics) and SanRa (cybersecurity), holding an MSc in Psychology and CISO qualification.
Related Topics
Related Posts
Sign in or Register
Please use the following structure: example@domain.com
Or Continue with
By registering you agree to the terms and conditions
Register to continue
Or Continue with
Log in to continue
Sign in or Register
Or Continue with
check your email
Check your email
We sent an email to you at .
It has a link that will sign you in.