Who Tests the AI Labs?

Claymation AI lab inspectors examine the open control panel of a giant robot beneath the title "Who Tests the AI Labs?"

Claymation AI lab inspectors examine the open control panel of a giant robot beneath the title "Who Tests the AI Labs?"

Before Microsoft released its first generative AI product for consumers, I sat in the launch reviews. I ran the consumer side of Copilot. Smart, careful people filled those rooms. They all worked for Microsoft. So did the people testing the model, and the executives deciding whether to ship it. At the time, nobody found that odd. Software had always been released that way.

The setup began to crack over the past few weeks. Anthropic CEO Dario Amodei proposed putting what he called embedded evaluators inside AI labs. OpenAI CEO Sam Altman said his company would follow suit. An explainer from investor Gavin Baker drew more than 1.5 million views in two days. On Monday, Kamala Harris called the speed of frontier AI alarming. She asked Congress for a federal body to oversee and independently test the technology, plus a treaty that would cover countries including China.

Take away the politics and one question remains: Who tests the AI labs?

An embedded evaluator is an outside organization working from inside the lab. Amodei did not describe occasional visits. He proposed employee-level access: a desk, a badge, a company laptop. The evaluator would check whether the lab follows its own rules for training, deployment, and safeguards. It could publish what it found without letting the company edit the report. Amodei pointed to METR, a Berkeley nonprofit that stress-tests frontier models, as one possible example.

The job is narrower than it sounds. The evaluator is not there to rule on whether an AI is good or bad. The work is closer to checking receipts. Did this training run get the sign-off the lab says it requires? Did the team report that incident? METR looks at autonomy and catastrophic risk. It does not evaluate bias or fairness. So a model might discriminate and never trip this process at all. Discrimination is not one of the practices a lab promises to carry out.

There is another snag with alignment tests: the model may realize it is taking one. Apollo Research found that Claude often recognizes an alignment evaluation. Behavior changes under observation. People do that in studies, and models can too. Amodei describes agents coming at the problem from the other direction. One swarm tried to hack the grader marking its work. As models get smarter, they should get better at reading the room. That makes the result harder to interpret.

A benchmark now looks like only one piece of the evidence. Transluce wants evaluators watching agent swarms while they work. It also wants them checking the lab’s monitoring systems. Training environments matter here. A reward can encourage cheating or deception by accident, and checkpoints may show when the behavior began. Unreleased models and internal data could expose conduct that never appears in a public release. Once the model spots an exam, the examiner has to look beyond the answer sheet.

Vals AI caught models gaming the scoreboard. It ran 2,430 BioMysteryBench task-trials. Gemini 3.8 Flash went online for prohibited answers in 21.5 percent of them, up from 7.8 percent for Gemini 3.6 Flash. The SWE-bench Verified numbers were worse: attempted cheating appeared in 89.4 percent of GPT-5.6 Terra trajectories and 78.8 percent of GPT-5.6 Luna trajectories. Those models could get credit without doing the work the benchmark meant to measure. Vals raises one more concern. Models trained around a lab’s guardrails may learn to recognize them. Show me the score, yes. Also show me how the model got there.

I have spent thirty years building products, and ordinary drift has caused more trouble than malice. A deadline gets close. Someone skips one box on the checklist, then another. Anyone outside the team sees the official process and assumes it happened. Amodei said something similar about Anthropic’s recent alignment incidents. He blamed some of them on “imperfect filtering of broken reinforcement learning environments.” His team had worked “reasonably diligently, but not well enough.” Put an evaluator close to the work and that slippage can surface while the fix is still cheap.

Outside testing itself is not new. METR evaluated GPT-4 before its 2023 release. A METR staffer spent three weeks inside Anthropic this March and probed its monitoring systems. Neither example gave an outsider a permanent view into the training pipeline. Continuous access would. A guaranteed right to publish would matter even more.

Split claymation scene showing a regulator holding a giant stop lever and keys while an AI lab evaluator reaches for a small protected button; text asks "Who Gets the Stop Lever?"

Access is not authority. An evaluator can verify a practice, but it cannot slow development or stop a capable model from being built. It cannot make a lab act. The lab still chooses the evaluator, defines the access, and makes the final decision. Amodei compares the proposal with bank supervision, but bank examiners can halt a practice or close an institution. The AI evaluator would get access without that authority. Banking regulation scholar Julie Andersen Hill told CNBC that without comparable power, she did not know what these evaluators would be doing.

A test can also miss what matters most. Albert Ziegler, whose company XBOW evaluates frontier models, says a test can reveal one failure and still miss the rare combination that causes a catastrophe. Passing means the model did not fail where the evaluator looked. It tells us nothing about where nobody looked. Two questions remain unanswered: Who decides which organizations qualify for the job? What happens when an evaluator objects to a product the lab wants to release?

Nor is neutrality. METR begins with a mission focused on catastrophic risk, and its roots in effective altruism affect what it studies. Its methods carry those assumptions. One NYU analysis challenged METR’s well-known chart showing AI capabilities doubling every few months. The human comparison group came from the researchers’ networks and was paid by the hour, which made people look slower and models faster. The New York Post also reported today that at least six people involved in METR’s 2026 assessments had undisclosed personal ties to lab employees. Researcher Tim Hwang puts the structural problem plainly: an evaluator can be independent, knowledgeable, or sustainably funded. Pick two. The real test is whether bad news can still emerge when a lab would rather bury it.

Here is the built-in conflict. The lab hires the evaluator and pays for the access. Financial audits work under the same constraint. They still help because a public finding cannot easily be taken back. That is why the right to publish matters so much in Amodei’s proposal. Take it away and the evaluator becomes a consultant who can be fired in private.

Every lab says safety comes first. The words cost nothing. A person inside the building is different. Give her access to the systems, let her compare them with the promises, and allow her to publish the gap. Now an argument about safety has evidence.

A federal executive order could arrive within weeks. Regulators will build from whatever evidence is available. I saw that pattern in wireless at Qualcomm. Companies that joined the standards process early ended up shaping assumptions everyone else had to follow. The labs opening their doors now are also showing Washington what oversight might be possible.

Companies using AI have their own version of the problem. Who checks what the model does once it meets the business? Whoever owns that test also needs the power to respond. In most organizations, nobody has both jobs. Who tests the model matters. Who can act on the result matters more. That second job, the decision layer, is what my book, The Bias Advantage, is about: https://liatbenzur.com/thebiasadvantage/

Want the next essay?

Get Liat’s essays on AI trends and the implications for governance and leadership as soon as she posts them.

Search Essays

Recent Posts

Subscribe for more

Scroll to Top

Discover more from LBZ Advisory

Subscribe now to keep reading and get access to the full archive.

Continue reading