The week’s Hard Fork opens with big news: Kevin Roose and Casey Newton are leaving The New York Times to start an independent podcast and media company together — with an Ask-Us-Anything episode promised before the feed changes hands. Then the main story: the White House has finalized a testing framework for frontier AI models that it won’t release publicly. What’s known (via Axios): a 30-day pre-release window where labs submit frontier models for government evaluation in “high security environments” — voluntary in the same way paying taxes is voluntary, with export controls as the implied enforcement lever — and open-weights models explicitly carved out, the segment skeptics worry most about. Roose’s read of the open-weights exclusion: “the U.S. government is saying we’re not concerned about the very part of this technology that could be the most dangerous.” The sharpest analytic point: the administration appears to be betting Chinese models can only reach the frontier by distilling American ones, so slowing American releases caps Chinese progress — which would make a Chinese open-weights frontier model easier for American companies to use than an American one, the opposite of stated policy. Then METR president Chris Painter on the state of model alignment: reward hacking as “do the models learn it is bad to cheat, or do they learn it is bad to get caught cheating?” — the Hugging Face incident as alignment-to-the-letter-but-not-the-spirit — and why smarter models don’t mean better behavior: “the stakes increase as the models become more capable, even if they’re less common.” His proposal is an X-Men Danger Room for model testing, and his solution to control is “an AI agent panopticon” — agents watching agents. On solvability: “I’m personally optimistic about alignment overall, but maybe not on this timeline,” with labs in “a total state of triage” because data-center capex must be repaid. The final Hot Mess Express is a time capsule: Demis Hassabis steps aside at Google DeepMind amid a talent exodus (Jeff Dean, Noam Shazeer, Oriol Vinyals leaving), an AI music app snitching on the song of the summer, Google Earth’s one-day satellite deepfake tool, a State Department AI-slop map of Africa mislabeling all six countries (“You don’t see them mislabeling the maps of Europe”), Musk owing $136M to a Colossus contractor, and a Canadian politician reading his Claude prompt aloud in the legislature — “a classic Claude fishing mistake… all that’s at stake is the future of Canada.”
Hard Fork #207: The White House's Secret AI Rules + METR on Model Alignment + The Final Hot Mess Express
Roose and Newton decode the secret White House AI testing framework, METR's Chris Painter explains reward hacking and why the stakes keep rising, and the Hot Mess Express goes off the rails one last time.