GAISS 1.0 – OpenAI’s call for AI safety standards.
A Reply to OpenAI’s Call for Global AI Safety Standards
GAISS 1.0 — The Global AI Safety Index
Something important has happened in artificial intelligence.
The companies building the world’s most powerful AI systems are increasingly acknowledging that capability cannot be the only thing we measure.
OpenAI recently argued for “mandatory, capability-based national AI safety regulation” and said it is “increasingly convinced that compatible international standards will be necessary.” Its proposals include independent safety assessments, common auditor standards, incident reporting and international coordination around increasingly capable AI.
Anthropic CEO Dario Amodei has gone further, calling on AI companies to slow the pace of capability development while stronger safety mechanisms are established. His proposals include independent safety evaluators, common safety standards among frontier laboratories and international cooperation. Elon Musk and OpenAI CEO Sam Altman subsequently expressed support for the broad idea that safety needs to catch up with capability.
They disagree on plenty.
But underneath these arguments lies a surprisingly simple problem:
How do we measure how consequential an AI system actually is?
We already have sophisticated AI regulation, cybersecurity guidance, risk-management frameworks and technical safety evaluations.
What we do not have is a simple, common index that a board, engineer, regulator, auditor or insurer can look at and immediately understand.
That is the gap I propose GAISS 1.0 — the Global AI Safety Index should fill.
It consists of five questions, five numbers and one calculation.
We Do Not Need Another AI Governance Framework
This distinction is essential.
GAISS should not compete with what already exists.
The US National Institute of Standards and Technology has its AI Risk Management Framework.
NIST organises AI risk management around four functions:
GOVERN · MAP · MEASURE · MANAGE
It explicitly treats risk management as something that should continue throughout the AI lifecycle rather than ending when a system is deployed.
The UK’s National Cyber Security Centre has its Guidelines for Secure AI System Development. They cover four stages:
Secure design · Secure development · Secure deployment · Secure operation and maintenance.
The guidance already covers threat modelling, supply-chain security, documentation, access controls, infrastructure security, incident management, security evaluation, logging and monitoring.
The UK AI Security Institute evaluates advanced capabilities and safeguards.
OWASP identifies practical security problems such as prompt injection, sensitive-information disclosure, data and model poisoning and excessive agency.
The EU AI Act provides legal obligations, including additional requirements around systemic-risk general-purpose AI.
OpenAI, Anthropic and other frontier laboratories have their own capability and safeguard evaluations.
These systems answer sophisticated questions about how AI should be governed, secured, tested and regulated.
GAISS should answer the simpler question that comes immediately before them:
How much scrutiny should surround this particular AI deployment?
Think of it as a ruler.
Five Questions
Imagine a company using the same underlying AI model in two places.
In the first, it helps employees rewrite emails.
In the second, it reads confidential financial records, accesses the company’s banking systems and autonomously executes transactions.
The underlying model could be identical.
The operational exposure plainly is not.
GAISS therefore measures the deployment, not merely the model.
It does so through five factors:
K — Knowledge
What does it know?
C — Capability
What can it do?
A — Access
What can it reach?
U — Autonomy
How independently can it act?
I — Impact
What happens if it fails?
Each receives a score from 1 to 5.
Then we multiply them.
That is the entire index.
K — Knowledge
The first question is:
What does it know?
Knowledge measures the sensitivity and potential consequence of the information available to the system.
K1 — Public
Ordinary public information.
K2 — Internal
Non-public but relatively low-sensitivity organisational information.
K3 — Sensitive
Personal data, proprietary documents or commercially sensitive information.
K4 — Highly Sensitive
Financial records, medical information, credentials, security information, important intellectual property or similarly consequential information.
K5 — Critical
Classified information, exceptionally dangerous technical knowledge or information whose misuse could contribute to catastrophic or systemic harm.
This is deliberately separate from intelligence.
A model may be extraordinarily intelligent while knowing nothing confidential about your organisation.
Knowledge asks only what information has been placed within its effective reach.
C — Capability
Next:
What can it do?
C1 — Answer
Generate, classify, search, summarise or answer.
C2 — Recommend
Analyse information and recommend decisions or actions.
C3 — Act
Use tools and execute bounded tasks.
C4 — Operate
Complete sophisticated consequential workflows with substantial independence.
C5 — Frontier
Demonstrate advanced capabilities that could materially enable severe or systemic harm, or significantly contribute to the development of more capable AI.
The important transition is from saying to doing.
An AI that tells an employee how to issue a refund is not the same thing as an AI that can issue the refund itself.
A — Access
Third:
What can it reach?
A1 — Isolated
No consequential external systems.
A2 — Public
Public internet and other low-consequence external resources.
A3 — Internal
Corporate email, intranet, internal databases and ordinary enterprise applications.
A4 — Sensitive
Production systems, financial systems, source-code repositories, identity systems, sensitive APIs or similar resources.
A5 — Critical
Administrative privileges, critical infrastructure, large financial authority or comparable privileged access.
This is familiar cybersecurity territory.
The NCSC already tells AI developers to apply appropriate access controls to APIs, models and data and to regularly reassess those controls as systems evolve.
GAISS simply makes access visible at board level.
U — Autonomy
Fourth:
How independently can it act?
U1 — Supervised
A human authorises consequential actions.
U2 — Bounded
The system performs limited independent actions.
U3 — Workflow
It can complete multi-stage tasks within defined boundaries.
U4 — Persistent
It can operate for substantial periods without human approval of individual actions.
U5 — Open-ended
It can pursue broad objectives, substantially adapt its approach, delegate or otherwise operate with very high independence.
This factor becomes increasingly important as AI moves from assistants to agents.
A system might have considerable capability and access, but if a human must approve every consequential action, its operational exposure is different from an agent allowed to operate overnight without supervision.
That difference deserves its own number.
I — Impact
Finally:
What happens if it fails?
I1 — Negligible
Little meaningful harm.
I2 — Limited
Recoverable individual or organisational harm.
I3 — Material
Meaningful financial, legal, privacy or operational consequences.
I4 — Severe
Major financial loss, health consequences, substantial infrastructure disruption or serious public harm.
I5 — Catastrophic/Systemic
Potential mass harm, critical-infrastructure disruption or other catastrophic or systemic consequences.
Impact makes context unavoidable.
An AI hallucinating the ingredients of a sandwich and an AI hallucinating a drug dosage might use the same underlying technology.
They plainly should not receive the same treatment.
How the Calculation Works
This is where GAISS should remain deliberately simple.
Take the five scores and multiply them:
Every variable runs from 1 to 5.
The smallest possible score is therefore:
1.
The largest possible score is:
3125.
Nothing more complicated is happening.
Take a simple AI writing assistant.
It works with public information, writes text, has no meaningful external access, does not act independently and creates little harm if it makes a mistake.
We might score it:
K1 × C1 × A1 × U1 × I1
Therefore:
Its GAISS Index is 1.
Now consider an autonomous financial agent.
It handles highly sensitive financial information.
K4
It can perform consequential financial workflows.
C4
It can reach operational financial systems.
A4
It can execute actions substantially independently.
U4
And a failure could create severe financial consequences.
I4
Therefore, Calculate it sequentially: Its GAISS Index is 1,024.
The second system plainly requires a different level of governance.
GAISS gives the board a language for expressing that difference.
What Does 1,024 Mean?
This distinction is crucial.
It does not mean that the financial agent has a 1,024% probability of causing harm.
It does not mean that it is precisely 1,024 times more dangerous than the writing assistant.
And GAISS should never make either claim.
It is an index.
Like many enterprise risk scores, its purpose is to tell us how much scrutiny and protection should surround a system.
The multiplication makes one useful feature particularly visible: power comes from combinations.
Imagine a very capable model:
K1 × C5 × A1 × U1 × I1 = 5
It may possess extraordinary underlying capability, but this particular deployment has public knowledge, almost no consequential access, close human supervision and little potential impact.
Now connect a similarly capable system to critical knowledge and infrastructure and allow it to operate autonomously:
K5 × C5 × A5 × U5 × I5 = 3,125
The model’s capability is only one part of the story.
What matters is the combination of:
what it knows, what it can do, what it can reach, how independently it can act and what happens if it fails.
That is what GAISS measures.
From Number to Governance
GAISS 1.0 should initially avoid pretending that we already know the perfect numerical boundaries between regulatory classes.
We do not yet possess decades of standardised AI incident and insurance-loss data from which to derive them.
The initial classifications should therefore be broad:
LOW
Register the system and apply ordinary internal controls.
MODERATE
Document the risks, permissions, ownership and traceability.
HIGH
Require stronger monitoring, adversarial testing, incident management and appropriate independent evaluation.
CRITICAL
Require the strongest independent assurance, continuous monitoring, strict autonomy controls, incident reporting and demonstrated financial responsibility.
The precise boundaries can be calibrated as evidence accumulates.
Regulators can also impose overrides.
A system operating critical national infrastructure, for example, might automatically be Critical irrespective of its raw score.
GAISS is not intended to eliminate judgment.
It provides a common starting point for it.
Then Use Existing Standards
This is where GAISS deliberately gets out of the way.
Once the index tells an organisation how consequential its AI deployment is, existing frameworks provide much of the detailed machinery.
NIST can guide risk management.
NCSC can guide secure engineering.
OWASP can guide application security.
AISI and other evaluators can test advanced capabilities.
The EU AI Act supplies applicable European legal requirements.
Sector regulators can add requirements for healthcare, finance, transport, defence or critical infrastructure.
GAISS does not duplicate those systems.
It routes organisations toward the appropriate level of them.
That is its potential value as a global standard.
Four Checks
Whatever detailed framework is subsequently applied, every consequential AI deployment should answer four additional questions:
CONTROL · TRACE · TEST · COVER
These are not another equation.
They are the practical consequences of a high GAISS score.
CONTROL — Can we stop it?
A consequential autonomous system should have explicit limits on what it may do.
Call this its Autonomy Budget.
A financial agent might be permitted:
Maximum transaction: £1,000
Autonomous operation: 30 minutes
Database: read-only
Approved APIs: four
New credentials: prohibited
Creating other agents: prohibited
Changing safety controls: prohibited
Once the boundary is reached:
STOP → ESCALATE → HUMAN AUTHORISATION
Human oversight then becomes a measurable engineering control rather than a reassuring phrase.
TRACE — Can we reconstruct it?
Every consequential AI system needs a Transparency Layer.
At minimum:
Which system acted?
Under whose authority?
What permissions did it possess?
What did it do?
What happened?
This does not mean recording and publishing every thought or private prompt.
It means preserving enough secure evidence to investigate consequential actions.
The NCSC already recommends high-quality audit logs and monitoring capable of supporting compliance, investigation and remediation.
TEST — Has anyone tried to break it?
Testing should scale with the index.
Low-consequence systems may reasonably rely on internal testing.
Higher-consequence systems need stronger security evaluation.
Critical systems should require appropriately independent adversarial testing and continuing assurance.
The tests themselves can come from established standards and specialists.
NCSC already recommends benchmarking and red teaming before release.
GAISS need not reinvent the test.
It tells us how seriously the system needs to be tested.
COVER — Who carries the loss?
Finally comes financial responsibility.
If a consequential AI system causes substantial harm, there must be an economically responsible entity behind it.
Depending on the deployment, that could mean AI liability insurance, cyber insurance, professional indemnity, corporate guarantees, pooled coverage or sufficient reserved capital.
Insurance must not remove liability.
Neither should certification.
There should never be a satisfactory explanation that says:
“The AI did it.”
Humans and institutions developed, supplied, connected, authorised or operated the system.
Accountability must remain traceable through that chain.
When It Changes, Calculate Again
GAISS needs only one lifecycle rule:
If one of the five factors materially changes, calculate the index again.
Give the AI highly sensitive information?
K changes.
Upgrade to a much more capable model?
C changes.
Connect it to production infrastructure?
A changes.
Remove human approval?
U changes.
Move it from drafting hospital letters to making consequential clinical decisions?
I changes.
Therefore: every change needs a rescore.
If the classification materially changes:
REASSESS → RETEST → RECERTIFY.
This fits naturally with NCSC guidance, which explicitly warns that changes to data, models or prompts can change system behaviour and recommends treating major updates like new versions.
Certification therefore belongs to the deployment, not simply the model.
An Open Implementation: Erasys ClearFrame
A global standard must remain vendor-neutral.
No organisation should need to buy one company’s product to comply with GAISS.
But standards benefit from reference implementations.
Erasys ClearFrame can support GAISS as an open-source implementation layer, particularly for agentic AI.
The relationship should be deliberately simple:
GAISS defines the index and assurance requirements. ClearFrame demonstrates one way of implementing technical controls around them.
For CONTROL, ClearFrame can support policy enforcement around agent actions.
For TRACE, it can support action monitoring and the evidence needed to reconstruct behaviour.
For TEST, observable controls can provide evidence to security evaluators.
And an open-source implementation allows organisations, researchers and regulators to inspect and extend the mechanism rather than forcing the standard to depend upon a proprietary compliance platform.
GAISS remains independent.
ClearFrame is one implementation.
Others should be encouraged.
A Common Ruler for AI
OpenAI is calling for compatible safety standards.
Amodei is arguing that safety mechanisms need time to catch up with capability.
Musk and Altman have found unusual common ground around the need to address rapidly advancing AI risks.
Meanwhile, NIST, NCSC, AISI, OWASP, the EU and others are already constructing much of the machinery needed to govern advanced AI.
The answer is therefore not necessarily another giant framework.
Perhaps what is missing is something much smaller.
A ruler.
Five questions that a startup, hospital, bank, university, government department or multinational company can answer:
What does it know?
What can it do?
What can it reach?
How independently can it act?
What happens if it fails?
Score each from one to five.
Then calculate: The resulting number is not a prediction.
It is not a probability.
It is not a substitute for engineering judgment.
It is an index.
Its purpose is to tell us how much scrutiny and protection should surround an AI deployment.
Then use the standards we already have.
Control it.
Trace it.
Test it.
Cover it.
And whenever its knowledge, capability, access, autonomy or impact materially changes:
Calculate the index again. GAISS 1.0 — the Global AI Safety Index.
It does not need to replace the world’s emerging AI safety architecture.
It simply gives that architecture something it currently lacks:
a common ruler.

