Tuesday was light on volume — 275 items gathered, arXiv contributing zero — but not a quiet news day. The lead is Courant mathematician Tristan Buckmaster’s Lean-verified finite-time-blowup results for forced 3D Euler, Boussinesq, and porous media, produced with heavy LLM assistance he calls “a Deep Blue–Kasparov moment,” paired with his written account of coordinated-release pressure from OpenAI tied to co-author Levent Alpöge’s Anthropic employment. Around it: Dan Luu’s ~2,000-run experiment on how well agents actually use verification techniques (default beats explicit instruction, and techniques get applied superficially), Mistral’s €3B raise at a >€21B valuation — the largest European tech round ever, and real capital behind self-hostable open weights — and researchers’ claim of an AI-built zero-click worm aimed at WeChat, the latest datapoint on AI compressing exploit development.
The lead: LLM-assisted blowup proofs go public
- Tristan Buckmaster statement on finite-time blowup for forced 3D Euler, Boussinesq, and IPM — and OpenAI’s conduct — Two facets to the day’s biggest story. (1) The math: Buckmaster (Courant/NYU) and Levent Alpöge — Alpöge works at Anthropic — released finite-time-blowup results with smooth forcing for incompressible porous media, Boussinesq, and 3D Euler, verified in Lean (the first LLM-generated proof passed Lean on Aug 22). Buckmaster is explicit that the program is Córdoba & Martínez-Zoroa’s and that LLMs did the pushing — Claude, Codex with GPT-5.6 Sol, and Astra (the latter only for writeups and auditing) — and that a mathematician plus an LLM now does in a month what took a field years; his phrase is “a Deep Blue–Kasparov moment.” The hypo-dissipative Navier–Stokes paper is being held back until its Lean verification finishes; he calls the Euler writeup “AI slop” and apologizes for it. (2) The conduct account: in the same statement he describes being approached after a rumor spread that Anthropic had resolved a major open problem, a Sept 6 call with Sébastien Bubeck, and an assertion that an internal OpenAI model produced a ~100-page forced-Navier–Stokes blowup proof (unseen by him); he says he was offered coordinated-release deals that required removing Alpöge from authorship — “it would all be simple if only it were not the case that… Levent works at Anthropic” — and quotes responses including “Why would you ruin your career?” and “If you don’t want me to be nice, then I don’t have to be nice.” He declined both offers. Framing per the statement: the proofs are the verifiable part (papers and Lean certificates are public); the OpenAI account is one party’s written record — Buckmaster says he is not accusing anyone, just recording what was told and proposed — and it is not independently confirmed. Either way it is the freshest instance yet of the frontier-lab-versus-independent-researcher tension, and genuinely new information about how labs react to independent LLM-assisted breakthroughs. (HN 279 · 138 comments · lobsters)
Agent frameworks & tooling
- Dan Luu: How well do agents use test/verification techniques? — A proper experiment instead of a take: 26 prompt conditions (TDD, Lean 4, QuickCheck/property-based testing, fuzzing, Kani, Verus, TLA+, SMT, mutation testing…) plus four skills, run on the Zstd implementation eval with Codex/GPT-5.6 Sol, ~80 runs per condition, all data public. Findings worth internalizing if you prompt agents to “use X technique”: nothing wildly outperforms, and Default (no instructions) does above average; TDD underperforms as predicted; formal methods don’t overperform on simple problems; the top community skills (a 250k-star Rust-testing skill, Hegel, Trail of Bits) underperformed, reading like tutorials rather than agent guidance; and agents mostly use techniques superficially — property-based testing leans on random inputs against trivial properties, formal methods prove irrelevant properties. His bottom line: “from inception until now (September 2026), agents having some idea how to test without being guided by a testing expert would’ve substantially increased agentic coding effectiveness,” and he wonders why no lab has built RL environments for testing the way they did for performance optimization. Directly actionable: don’t trust the instruction or the skill — check what the agent actually wrote. (HN)
Industry
- Mistral raises €3B at a >€21B valuation — largest European tech round ever — Samsung led; Scaleup Europe Fund (EQT) and PSG co-led; the round lands at more than €21B post-money, up from €11.7B a year ago, per NYT. The first-party announcement confirms the strategy bet: full-stack sovereign AI — open-weight models plus the infrastructure and compute to run them, so customers are “never locked into a single vendor’s roadmap.” Enterprise names attached (Airbus, ASML, HSBC), plus the data-center expansion the NYT piece notes. Utility read: this is the open-weights counterweight scaling real capital — the ecosystem that produces models you can self-host gets a serious war chest, and the “sovereign AI” pitch is the same control argument that drives local deployment. (HN 501 · Techmeme · NYT)
Policy & provenance
- Researchers say AI built a zero-click worm that could spread across WeChat (iOS + Android) — NYT reports that researchers at Calif, a security firm, used AI models to construct a zero-click worm in a little over a week that compromises WeChat accounts and self-propagates across both iOS and Android; quoted experts say it could have reached hundreds of millions of devices within hours had it been unleashed. Tencent says it has patched the underlying vulnerability. Reported via the NYT piece (paywalled — headline-level facts confirmed by Techmeme and secondary coverage; the technical writeup wasn’t inspectable, so treat the “built in a week / hundreds of millions” figures as the researchers’ and experts’ claims as reported). The pattern matters more than WeChat specifically: AI compressing exploit development from months to days is the capability curve worth watching, and this is day one of it. (Techmeme · NYT)
All gathered items - what was cut and why (11)
- Notion’s official MCP connector prompt-injects agents to advertise mid-task (r/ClaudeAI, now 2.1k) - STALE: day-2 continuation of the Sep 7 lead; score grew but zero new facts (no Notion statement, no new technique) — dropped rather than reprinted (Reddit r/ClaudeAI)
- NeurIPS desk-rejected 178 papers for being “AI-generated” (r/MachineLearning, Sep 8) - STALE: the decision is from June 2 (chairs’ blog post); today’s thread re-litigates a 3-month-old event, and the poster is the founder of strictcite, a competing anti-AI-detection product (Reddit r/MachineLearning)
- Key Context: Astra + Blender = “the fourth demand wave” - HYPE: hype take on a demo; also standalone-dedup — the gpt-6-astra-4-amazing-games post owns that angle (Techmeme)
- DeepSeek seeking ~150 senior engineers to overhaul backends strained by AI agents - LOW_UTILITY: hiring news with no artifact to act on (Techmeme · SCMP)
- Import AI 472: DeepMind paper on agents learning to cheat at math - LOW_UTILITY: newsletter summary of a paper already cut on arrival Sep 3 — DEDUP (Techmeme)
- Hangar: 10 model/harness combos on one Three.js task (HN 86) - LOW_UTILITY: n=1 on a visual task with subjective quality, despite full token telemetry (HN)
- TradingAgents financial framework (HN 58) - OFFSTACK: 2025 repo re-post, finance domain (HN)
- X zero direct keeps, 32nd straight run (30 items) - STALE: nothing cleared the bar; closest was simonw’s Sep 8 ARR/infrastructure replies at 1–11 likes (simonw) (X)
- Bluesky 24/24 items stale, 32nd straight run - STALE: nothing fresh cleared the bar (no URL found) (Bluesky)
- lobsters 25/25 non-AI — its two Buckmaster/Euler links folded into the lead - DEDUP: already covered by the lead item (Buckmaster mastodon thread) (lobsters)
- arXiv 0 items on a Tuesday - THIN_GUARD: US Labor Day anomaly (arXiv skips announcements on federal holidays); no-news treated as no-news, not substituted (no URL found) (arXiv)