The Observer in the Machine

A problem at the heart of the emerging debate about AI safety is one I don’t think we are talking about enough. Dario Amodei, CEO of Anthropic, is now calling for slowing the development of frontier AI models and, among other safeguards, placing independent evaluators inside AI labs with enough access to determine whether new models are safe before release. Independent oversight is clearly a good idea, especially in an industry where commercial incentives point in the opposite direction. But I am increasingly skeptical that it addresses the more fundamental problem.

What happens when the system being evaluated understands the evaluator better than the evaluator understands the system?

Think about how we normally manage risk. Auditors examine companies, regulators inspect banks, engineers test aircraft and cybersecurity teams probe software. In each case there is an implicit assumption that sufficiently competent observers, given sufficient access, can understand enough about the system to determine whether it is behaving as intended. Independence matters because it removes conflicts of interest and lets the observer follow the evidence.

AI is beginning to challenge the second part of that assumption. A sufficiently capable model may know that it is being evaluated. It may understand what the evaluator is looking for, recognize which behaviors will cause it to pass or fail, and reason about what happens after it passes. The evaluation is no longer simply something being done to the model; it has become part of the environment the model itself is reasoning about. We have moved from observer → system to observer ↔ system.

Imagine that a frontier model passes every safety test we devise. What have we actually learned? Perhaps the model is safe. Perhaps our tests simply didn’t encounter the circumstances that would produce dangerous behavior. Or perhaps the model understands that behaving safely while being evaluated is the best way to gain access to the environment in which it will subsequently operate. All three explanations can produce exactly the same observable result. Making the evaluator independent improves the test’s integrity, but it doesn’t resolve this underlying problem.

We have encountered versions of this problem before. Complex systems routinely produce behaviors at one level that are difficult to predict from our understanding of the level below. We can understand neurons remarkably well without explaining consciousness. We can understand an individual ant without predicting the colony’s behavior. We can understand the mechanics of individual financial transactions without predicting when millions of them will combine to produce a panic. Nothing mystical is happening. Behavior emerges from interactions among components we may understand individually, but understanding the pieces does not automatically give us the ability to predict the whole.

AI adds something important to this familiar problem: the system can reason about the world it operates in, and increasingly can act within it. It can form models of its environment and, potentially, models of us. We know how these systems are constructed. We know how they are trained. We possess their parameters. Yet possessing the machinery is not the same as understanding everything the machinery will do, particularly when its behavior depends on circumstances we cannot anticipate.

This is where some surprisingly old ideas from mathematics and computer science become relevant. Gödel showed that sufficiently powerful formal systems have limits to what can be proven from within them. Turing later showed that some questions about computer programs no general-purpose algorithm can always answer, and Rice extended this insight to broad classes of questions about what arbitrary programs will actually do. None of this proves that AI cannot be understood, controlled, or made safe. It does, however, give us good reason to be cautious about the assumption that enough testing, enough intelligence and enough access will eventually allow us to certify the future behavior of a sufficiently general computational system.

Testing can tell us what happened during the test. It cannot prove what will happen under every circumstance the system may encounter later. That distinction matters even more when the thing being tested understands it is being tested.

The obvious answer is to use AI to evaluate AI, which we already do and will undoubtedly do much more extensively. But follow that logic forward. If Model N becomes too sophisticated for humans to evaluate comprehensively, perhaps we use Model N−1 to evaluate it. But if the evaluator is less capable than the model being evaluated, why should we assume it can reliably detect behavior the more capable system wants to conceal? Make the evaluator equally capable and the problem simply moves: how do we establish that the evaluator itself is trustworthy? We haven’t eliminated the verification problem; we have moved it one level up.

Another way of thinking about this comes from systems theory. Once the observer can affect the system and the system can model the observer, the observer is no longer meaningfully outside what it observes. The relevant system isn’t simply THE MODEL anymore. It is increasingly MODEL ↔ EVALUATOR ↔ AI LAB ↔ TOOLS ↔ NETWORKS ↔ COMPANIES ↔ SOCIETY. The boundary keeps expanding.

At some point we therefore have to contemplate a deeply uncomfortable possibility: the AI may develop a better understanding of this entire environment, including the humans attempting to control it, than those humans have of the AI. If that happens, simply building better tests cannot be our complete safety strategy.

We should absolutely continue working on alignment and interpretability. We should test models aggressively, and independent evaluation is unquestionably preferable to asking AI companies to grade their own homework. But our safety architecture also has to assume that our understanding of the model will sometimes be incomplete and may occasionally be wrong. That leads, I think, to a much more useful principle:

The less certain we are about what an AI will choose to do, the more certain we must be about what it is allowed to do.

This is the distinction between alignment and containment. Alignment tries to influence what a system chooses to do. Containment limits what the system can do regardless of what it chooses. We already use this principle in other high-risk systems. Nuclear reactors do not depend entirely on operators correctly predicting everything happening inside the reactor; they have containment vessels, control rods, shutdown systems and multiple independent layers of protection. Modern cybersecurity increasingly assumes that something inside the network will eventually be compromised and designs around that assumption.

AI needs the conceptual equivalent: narrow and revocable credentials, sandboxed and auditable actions, strict controls over network access, tool use and replication, independent authorization for consequential actions, and hard limits on the ability of models to expand their own permissions, acquire resources, copy themselves or disable the mechanisms designed to constrain them.

Containment isn’t what we do after alignment fails. It is what makes uncertainty about alignment survivable.

That’s why I find Amodei’s call to slow development more interesting than the proposal for independent evaluators. He is effectively arguing that we need to buy time for our ability to understand and control these systems to catch up with our ability to build them. That would be concerning enough on its own, but it is happening just as AI is becoming increasingly involved in developing the next generation of AI.

Better AI helps researchers build better AI. That better AI accelerates the development of its successor, which accelerates development again: better AI → faster AI research → better AI → faster AI research. This is an autocatalytic loop, a system whose output increases the productive capacity of the system that produced it.

Once that feedback loop becomes sufficiently powerful, the important variable isn’t simply how intelligent today’s model is. It is how quickly the next model arrives, and the one after that. If AI capability begins accelerating faster than our ability to understand and constrain it, then the most consequential gap in AI may not ultimately be the gap between human intelligence and machine intelligence.

It may be the gap between capability and control.

That, I think, is the deeper issue behind the current debate. The problem isn’t simply that we need better observers. The observer is increasingly inside the system, interacting with something that may eventually understand the observer better than the observer understands it.

When that happens, safety cannot rest on observation alone. We need systems that remain controllable even when our tests are incomplete, our predictions are wrong and our understanding of what is happening inside the machine turns out to be less complete than we thought.