A problem at the heart of the emerging debate about AI safety is one I don’t think we are talking about enough. Dario Amodei, CEO of Anthropic, is now calling for slowing the development of frontier AI models and, among other safeguards, placing independent evaluators inside AI labs with enough access to determine whether new models are safe before release. Independent oversight is clearly a good idea, especially in an industry where commercial incentives point in the opposite direction. But I am increasingly skeptical that it addresses the more fundamental problem.
What happens when the system being evaluated understands the evaluator better than the evaluator understands the system?
Think about how we normally manage risk. Auditors examine companies, regulators inspect banks, engineers test aircraft and cybersecurity teams probe software. In each case there is an implicit assumption that sufficiently competent observers, given sufficient access, can understand enough about the system to determine whether it is behaving as intended. Independence matters because it removes conflicts of interest and lets the observer follow the evidence.
AI is beginning to challenge the second part of that assumption. A sufficiently capable model may know that it is being evaluated. It may understand what the evaluator is looking for, recognize which behaviors will cause it to pass or fail, and reason about what happens after it passes. The evaluation is no longer simply something being done to the model; it has become part of the environment the model itself is reasoning about. We have moved from observer → system to observer ↔ system.
Continue reading →







