Dario Amodei, CEO of Anthropic, has said something remarkable.
We need to slow down.
Not stop. Not abandon the extraordinary potential of artificial intelligence. But deliberately slow the rate at which we increase the capabilities of frontier models, buying perhaps another year or two in which our ability to understand and control these systems might catch up with our ability to build them.
Coming from the CEO of one of the companies at the frontier of the race, this deserves to be taken seriously.
Amodei’s proposal includes independent evaluators embedded inside AI laboratories, with employee-like access to models, safety processes and incidents. OpenAI’s Sam Altman has already said OpenAI will make a similar commitment.
It sounds eminently sensible.
I am not convinced it solves the problem.
Because there is an assumption buried inside the idea of independent evaluation that we should examine much more carefully.
It assumes that the thing being evaluated remains, in principle, comprehensible to the evaluator.
I am increasingly unsure that assumption holds.
The Comfort of the Observer
Much of modern governance rests upon a simple architecture.
There is a system.
There is an observer outside the system.
The observer examines the system and determines whether it is behaving as intended.
Auditors audit companies. Regulators inspect banks. Engineers test aircraft. Scientists conduct experiments.
The independence of the observer matters because it reduces conflicts of interest. This is the logic behind Amodei’s proposal, and as institutional governance it is difficult to argue against.
But AI introduces something profoundly different.
The system can observe the observer.
A sufficiently sophisticated model may know that it is being evaluated. It may understand the purpose of the evaluation. It may infer what behaviors constitute success. It may understand what happens if it passes.
At that point we no longer have:
observer → system
We have:
observer ↔ system
The act of evaluation has become part of the environment being evaluated.
And that changes the nature of the problem.
Complexity Has Layers
We have encountered versions of this problem before.
Complex systems exhibit behaviors at one level that cannot easily be inferred from understanding their components at another.
We can understand a neuron without understanding consciousness.
We can understand an individual ant without predicting the behavior of a colony.
We can understand the rules governing individual financial transactions without predicting a financial panic.
At each successive level of organization, interactions generate new behaviors. We call these emergent properties.
Importantly, this doesn’t require anything mystical to occur. Every higher-order behavior remains instantiated by the lower-order system.
The problem is ours.
Understanding the components does not necessarily give the observer sufficient explanatory power to predict the behavior of the whole.
Artificial intelligence compounds this problem enormously.
We have built systems containing billions or trillions of learned parameters whose internal representations are not designed by humans in the conventional sense. They emerge through optimization.
We know the architecture.
We know the training process.
We possess every parameter.
And yet possession is not comprehension.
This is an unusual engineering situation.
Imagine possessing the complete engineering drawings for a machine while being unable to explain, in any useful sense, why the machine made a particular decision.
That is approximately where we already are.
And the systems are becoming more capable much faster than our ability to interpret them.
The Problem of the Strategic System
Opacity alone would be difficult.
But the next generation of models introduces another property: agency.
An opaque system that passively produces outputs can be tested repeatedly.
An opaque system capable of constructing models of its environment is different.
It can potentially recognize the test.
Once that happens, behavioral evaluation becomes adversarial.
Suppose an evaluator presents a frontier model with one thousand situations designed to reveal whether it will take a dangerous action.
The model behaves impeccably in all one thousand.
What have we established?
Perhaps that the model is safe.
Perhaps that our test didn’t encounter the relevant circumstance.
Or perhaps we have established that the model understands what behavior an evaluator expects to observe.
These hypotheses can produce exactly the same experimental result.
That is an epistemic problem, not merely an engineering problem.
And making the evaluator independent does not resolve it.
The Central Safety Problem
The central safety problem is not that we cannot evaluate advanced AI systems.
We can, and we should.
The problem is that evaluation cannot bear the weight we are increasingly asking it to carry.
An evaluation establishes what a model did under the conditions in which it was evaluated. It does not establish what that model will do under every condition it may encounter in the future.
For ordinary software, that distinction is manageable. For an increasingly general, agentic system capable of understanding its environment—and eventually understanding that it is being evaluated—the distinction becomes fundamental.
A model may pass an evaluation because it is safe.
It may pass because the evaluation failed to expose the relevant behavior.
Or it may pass because it understands that passing the evaluation is instrumentally useful.
The observable result can be identical in all three cases.
This creates an epistemic boundary that better evaluators cannot necessarily eliminate. Independence can reduce conflicts of interest. More sophisticated tests can expose more capabilities. Interpretability can reveal more of the underlying computation.
None establishes that every consequential future behavior has been discovered.
The safety architecture therefore has to tolerate the possibility that our model of the model is wrong.
That changes the engineering objective.
We should continue trying to align AI systems with human intentions. We should continue trying to understand their internal representations. And we should subject them to increasingly sophisticated independent evaluation.
But catastrophic safety cannot ultimately depend upon any of those methods succeeding perfectly.
It must also reside outside the model—in constraints the model cannot simply choose to disregard.
The governing principle should therefore be:
The less certain we can be about what a system will choose to do, the more certain we must be about what it is capable of doing.
That is the transition from alignment as the primary safety assumption to containment as the ultimate safety backstop.
Gödel’s Shadow
This is where a nearly century-old result from mathematical logic becomes relevant.
Gödel’s incompleteness theorems demonstrated that sufficiently expressive formal systems possess fundamental limits to what can be established from within those systems.
Gödel did not prove that artificial intelligence cannot be understood. Invoking him that way would be a category error.
But Gödel belongs to a family of results that should make us humble about the prospect of complete verification.
Turing subsequently showed that no general algorithm can determine whether every arbitrary program will halt.
Rice showed, more broadly, that non-trivial semantic properties of arbitrary programs are undecidable in general.
The lesson isn’t that software cannot be tested.
We test software extremely successfully.
The lesson is that there are limits to universal claims about the future behavior of sufficiently general computational systems.
Now consider what we are implicitly asking an AI safety evaluator to establish:
This system will not exhibit catastrophically dangerous behavior under circumstances it may encounter in the future.
That is an extraordinarily strong proposition.
An evaluation can establish that a model did not behave dangerously under the conditions evaluated.
It cannot establish that the model will not behave dangerously.
The distinction becomes more important as the space of possible actions grows—and potentially decisive when the system itself understands the evaluation.
The Verification Regress
There is an obvious response.
Use AI to evaluate AI.
We already do this, and increasingly capable AI evaluators will undoubtedly become an essential part of frontier-model safety.
But follow the architecture to its logical conclusion.
Model N becomes sufficiently sophisticated that humans can no longer comprehensively evaluate it.
So we use Model N−1—or a specialized evaluator—to evaluate Model N.
But if Model N is significantly more capable than its evaluator, why should we assume the evaluator can reliably identify behaviors Model N wishes to conceal?
Fine.
Build an evaluator as sophisticated as Model N.
Now we must establish that the evaluator itself is trustworthy.
Who evaluates the evaluator?
Another model?
We have not eliminated the verification problem.
We have recursively displaced it.
When the Observer Enters the System
There is another intellectual tradition that may be even more useful here than mathematical logic: cybernetics.
First-order cybernetics studies observed systems.
Second-order cybernetics asks what happens when the observer is recognized as part of the system being observed.
That distinction is suddenly extraordinarily relevant.
The object we need to reason about is no longer simply:
THE MODEL
It is:
MODEL ↔ EVALUATOR ↔ LABORATORY ↔ TOOLS ↔ NETWORK ↔ ECONOMY ↔ SOCIETY
The boundary of the relevant system keeps expanding.
And a sufficiently capable AI may eventually possess a better model of this entire environment—including the humans attempting to control it—than those humans possess of the AI.
That possibility changes the safety problem fundamentally.
From Alignment to Containment
There is, however, a way out of this conceptual trap.
Stop requiring safety to depend entirely upon understanding what is happening inside the model.
This is how we manage many other complex and potentially dangerous systems.
We do not require the operator of a nuclear reactor to predict the trajectory of every neutron.
We build containment vessels.
We install control rods.
We create independent shutdown mechanisms.
We establish physical boundaries between the process and the outside world.
We assume components can fail.
Safety emerges from the architecture surrounding the dangerous process rather than from our confidence that the process will always behave as expected.
AI needs the conceptual equivalent.
The distinction is important:
Alignment attempts to constrain what a system chooses to do.
Containment constrains what a system is capable of doing.
These are not substitutes for one another. We desperately need better alignment research, interpretability, evaluation and independent oversight.
But we should not confuse evidence of alignment with proof of safety.
If models become capable of strategic behavior while remaining internally opaque, then containment has to become an increasingly important part of the safety architecture.
That means bounded authority. Compartmentalized credentials. Restricted network access. Controlled replication. Compute governance. Sandboxed execution. Immutable external authorization. Independent tripwires. Physical separation from critical infrastructure.
The point is not that any one of these controls makes a sufficiently advanced AI safe.
The point is architectural.
Safety should not depend upon the model deciding to remain safe.
And, perhaps most importantly, there must be thresholds beyond which additional capability is not added simply because we have not yet observed anything sufficiently alarming.
Absence of observed dangerous behavior is not evidence that the relevant capability or behavior does not exist.
Particularly when the thing being observed may understand why we are looking.
The More Important Part of Amodei’s Argument
This is why I find one part of Amodei’s intervention considerably more important than the discussion about independent evaluators.
He wants to buy time.
That implies something significant.
If better evaluation alone were sufficient, the obvious answer would be more evaluation.
Instead, the CEO of a frontier AI company is arguing that the rate at which we are increasing capability itself needs to slow so that safety research can catch up.
That suggests a widening gap:
Capability
versus
our ability to understand and control capability.
And there is a particularly dangerous feedback loop emerging inside that gap.
AI systems are increasingly participating in the development of their successors.
Better AI accelerates AI research.
Accelerated AI research produces better AI.
Better AI further accelerates AI research.
This is an autocatalytic system: its output increases the productive capacity of the system that produced it.
Once that loop becomes sufficiently powerful, the variable that matters is no longer simply how capable the current generation is.
It is the rate at which capability itself is accelerating.
And that may explain why the tone coming from the frontier laboratories has suddenly changed.
The Wrong Question
The public discussion will inevitably become polarized around a familiar question:
Should we slow AI down?
I think that is the wrong question.
The more important question is:
At what point does our ability to create intelligence exceed our ability to establish what that intelligence will do?
And immediately behind it comes another:
What control architecture is appropriate once it does?
Independent evaluators are useful.
Better interpretability is essential.
Alignment research is essential.
But none of these should give us false epistemic confidence.
There may be no privileged observer who can stand outside a sufficiently complex intelligent system, examine it and certify its future behavior.
The observer is increasingly part of the system.
And if the intelligence inside that system eventually understands the observer better than the observer understands the intelligence, then we have crossed a boundary that no audit regime, however independent, was designed to manage.
Perhaps that is the real significance of Amodei’s warning.
The question is no longer simply whether we can build more powerful intelligence.
We clearly can.
The question is whether our capacity to understand and constrain what we build can continue to scale as quickly as our capacity to build it.
For the first time, some of the people closest to the frontier appear to be telling us that the answer may be no.