Leonard Tang co-founded Haize Labs as an independent evaluator of frontier models — Anthropic and OpenAI were the first customers — and he opens by saying he is sympathetic to Dario Amodei’s proposal to embed outside evaluators inside the labs.

His own company walked away from that work for two reasons: the labs did not pay for third-party safety testing the way they paid for post-training data, and it was never clear the evaluations changed how models were actually built.

So he asks the harder question. Suppose the industry succeeded in building a thriving corps of embedded external evaluators. Would that be enough to pace the frontier? He doubts it, for three reasons.

1. Is it really different from an internal alignment team?

The argument for embedding is that outsiders bring a different background and perspective, and that one evaluator working across several labs can spot risks no single company sees.

  • The first benefit is thin: an outside team can be just as isolated as a team the lab grows itself.
  • The second is sound — it is how financial regulation already works — but it requires labs to share information they reasonably treat as intellectual property.

2. What are the evaluations measuring against?

If the labs write the criteria, the evaluation repeats work the lab already did internally. If outsiders write them, who has the right to decide what the criteria are?

  • Model capabilities move faster than any standards process, so a static external rulebook goes stale quickly.
  • Broad principles are worth stating normatively, but making them concrete depends on what models can currently do — deception in a 2023 chatbot is a different problem from deception across a swarm of agents running in production.

3. Do the evaluations change anything?

Good evaluations still do not guarantee a change in model behavior. Commercial pressure, at home and abroad, can override the best intentions of individual safety researchers.

Embedded evaluators, he concludes, buy visibility into frontier risk but answer none of the hard questions: what desired behavior is and who decides, how that definition drifts as models improve, and how evaluations translate into changed behavior.

Four alternatives he sketches

Rapid, cross-lab feedback. Dario, Demis and Sam reach for the FAA, FINRA and the IAEA as models. Tang points at a different family of institutions instead — the ones built for fast learning from failure: aviation’s incident reporting systems and the NTSB, cybersecurity’s CERTs and information-sharing centers.

  • Failure modes get collected continuously, turned into test exercises, and distributed to every lab.
  • Each lab gives up a limited amount about its own failures and receives everyone else’s discoveries in return.
  • The point is speed: safety knowledge compounds without stopping capability work.

A case-based legal system for AI. Disputes produce decisions, decisions produce precedent, and precedent updates what evaluators test for and what labs fix. It can respond to failures nobody anticipated, which a fixed rulebook cannot.

Alignment mesocosms. Real-world testbeds where the model has objectives stated loosely enough to leave room for perverse behavior, humans have a strong reason to notice, and the damage stays contained.

  • Autonomous firms are his example of an unusually good mesocosm: “make money” is exactly the kind of open-ended goal that could be pursued by deceiving the board, eliminating employees or squeezing customers.
  • Because the owner-operators stay in the driver’s seat and care about the outcome, they intervene early — grounded stakes, bounded blast radius.
  • Networks of small firms extend this across society, exposing models to underrepresented values and local knowledge, and the commercial incentive already exists.

Direct democratic alignment. He finds it troubling that the values of the most consequential technology ever built are set by a very narrow slice of society. His bet on the last frontier lab: it builds intelligence by drawing out the distributed values of the general population through direct democracy — with AI helping to educate voters about the tradeoffs so they can work out what they actually believe.

He ends by inviting email from anyone thinking about the same problems.