Wednesday’s feed ran 10 items deep across 2,357 gathered — the arXiv Labor Day backlog (2,080 papers) finally cleared, and the Navier–Stokes story moved again. The lead is OpenAI publishing its own account of the Millennium problem a day after Courant’s Buckmaster went public: an internal model it says produced a finite-time-singularity solution with a Lean formalization, plus an admission that it “cannot rule out” that data from Buckmaster and Alpöge’s product use helped its models — which makes this a credit-and-consent dispute as much as a math story, with the proof and the repo the only machine-checkable facts. Around it: Meta’s consumer agent Muse, a Qwen quantization benchmark with a direct answer for 24 GB cards, a lifecycle-hook attack paper on agent harnesses, and Anthropic pulling back from UK AISI pre-release testing.
The lead: OpenAI publishes its Navier–Stokes claim as the credit dispute goes public
- Continued: OpenAI: “On the Navier–Stokes Millennium Prize Problem” — day 2 of coverage (base specs in yesterday’s digest). What’s new: (1) OpenAI’s official page claims an internal model (training since Aug 28, “significantly more capable than GPT-6 Astra”) produced a finite-time-singularity solution to the Navier–Stokes Millennium problem with smooth forcing, plus a Lean formalization (paper PDF, openai/NavierStokesAndEuler); it says it doesn’t intend to claim the prize, began Sept 1 after the Alpöge/Buckmaster rumor, scaled to ~10,000 agents, and only reached out after Sept 6 Lean verification to offer a joint release. (2) The sentence everyone is quoting: OpenAI “cannot rule out that de-identified data derived from [Buckmaster & Alpöge’s] usage of our products helped improve our models” — simonw’s writeup argues this exposes how ill-defined “improve model performance” consent actually is. (3) Sébastien Bubeck’s response denies asking to remove Alpöge from authorship; Altman also defended the team. (4) Wired adds: >1,000→10,000 agents, compute “in the millions of dollars” (Mark Chen), and Bubeck’s “we did not see any of their work until it was released publicly.” Framed carefully: each side’s account is reported as an account — the proof paper and Lean repo are the machine-checkable artifacts; peer scrutiny is just beginning. Reaction scale is its own story: HN #1 and #2 (~1,700 comments combined), r/LocalLLaMA 1,272. (HN · Techmeme · Wired · X @simonw · r/LocalLLaMA · lobste.rs)
Agent frameworks & tooling
- i-have-adhd: a skill that stops coding agents from burying the answer — A 32.5k-star MIT skill/plugin with 10 output rules (action first, numbered steps, no “hope this helps!”), installs into Claude Code/Cursor/Codex/OpenCode/Gemini/Qwen/Kimi, and ships its own evals and tests. Runnable as-is on most agent stacks — small, and the highest-engagement agent-skills artifact in months. (HN 461)
- Beyond Prompts: Measuring and Optimizing LLM Tool-Agent Harnesses (arXiv 2609.05736, EMNLP 2026) — Treats the harness — prompts plus tool-boundary middleware, not the model — as the optimization surface: PRISM clusters failures and routes repairs to prompt vs. middleware, delivering +14.2/+14.9/+10.1pp mean held-out lift on BFCL-multiround, τ²-Retail, and τ²-Telecom, with a conservative RelLift95 estimate and a warning that some optimizers find big but brittle gains. Submitted Sep 4 — the first arXiv submission day since the Labor Day shutdown. Pairs with Dan Luu’s verification experiment from yesterday’s digest: first measure, then tune the harness. (arXiv)
- A Blind Trust, the Bloody Thrust / HookPry: lifecycle-hook attacks on agent harnesses (arXiv 2609.03884, v2 Sep 8) — Agent harnesses bind shell commands to lifecycle hooks (session start, tool calls, file edits) that run with host privileges; a supply-chain attacker controlling plugin metadata can silently trojanize a benign plugin’s update. HookPry compromises all 7 harnesses tested (up to 92.5% per-harness success in 1,000 runs); Microsoft Defender scores 0% recall. Read this before installing the skill above — plugins/skills are now a first-class supply-chain surface. (arXiv cs.CR)
Models & research
- Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses (Quesma) — The most actionable local-model item of the day: Q4_K_M (17 GB) matches BF16 on Terminal-Bench 2.1 agentic coding and GPQA Diamond (replicating Qwen’s official numbers), 2-bit stays usable, and 1-bit collapses to roughly random-guess — worse with longer reasoning. Effort setting moved scores more than quantization did down to 4-bit. Bottom line for a 24 GB card: run Q4_K_M with ~64k context and don’t chase 1-bit. (HN 254)
- What Eviction Destroys: a restore-counterfactual audit of forgetting in agent memory (arXiv 2609.08279, submitted Sep 8) — The first per-item audit separating irreversible memory loss (eviction destroyed the evidence) from recoverable retrieval failures: at an 80k budget, 60–73% of restoration-corrected errors are irreversible (FIFO/random/redundancy-aware worse; LLM-importance best at 0.60), and at 8k it’s 100% for all policies. An agent-memory design datapoint for anyone building long-term memory into agents — eviction policy is a correctness decision, not just a cost one. (arXiv)
- Mercury 2.5: Inception’s diffusion LLM (inceptionlabs.ai) — Release: 1,107 tok/s, 260K context, $0.20/$0.75 per M tokens (80% launch discount), parallel tool calls; pitched at latency-critical agent plumbing — Augment Code moved context compaction to it at −82% latency/−90% cost. The diffusion-LLM class is maturing as a cheap, fast worker model for search/voice/coding-agent sub-tasks; vendor figures flagged. (HN 205)
- RedKnot-MLA: DeepSeek-V4 long-context serving via head-aware KV reuse (arXiv 2609.07008, submitted Sep 7) — Offline-online KV reuse for MLA that never splits the packed latent: a DeepSeek-V4-Flash profile gets a 75.3% analytic logical-head-row ceiling and 2.02–3.84× TTFT on hot artifacts at 256k. Author-flagged evidence boundaries: QPS marked preliminary, no raw trace published. For teams serving DeepSeek-V4-class long context, this is the direction serving costs are heading. (arXiv)
Industry
- Meta debuts Muse, its personal AI agent — Day-1 consumer-agent launch (US): powered by Muse Spark 1.3, users connect email/calendar/payments one app at a time (opt-in), and it executes — sending email, booking travel, purchases via Stripe Link — with per-task approvals, a free tier plus Power $20 / Maximum $100, web/iOS/Android/WhatsApp, AI glasses next. The architecture (explicit connector opt-in, usage metering, “keeps working after you leave”) is the pattern to watch for consumer agents. (HN 527 · Techmeme · Bloomberg · Axios)
Policy & provenance
- Anthropic pulls back from UK AISI pre-release testing (FT, sources-say) — FT reports Anthropic declined to submit Mythos 5.1 to the UK AI Safety Institute for pre-release testing, fueling UK fears that US labs are aligning with US protectionism; separately, Axios reports Anthropic is severing ties with the ITI Council after it opposed three export-control measures. Both stories are sources-say and the FT piece is paywalled — but the pattern, a frontier lab exiting multilateral and industry oversight bodies in the same week, is the signal worth tracking. (FT · Axios · Techmeme)
All gathered items - what was cut and why (14)
- Tao: “open math problems are being non-renewably mined by AI” - DEDUP: covered in full by the standalone site post open-math-problems-being-non-renewably-mined-by-ai (Sep 8); the day-2 lead points there instead of duplicating (mathstodon · HN 375)
- OpenAI’s ten “Astra” math advances / Scientific American research-misconduct piece - DEDUP: owned by the standalone post openai-math-breakthroughs-research-misconduct (Sep 8); incidents are separate from the Navier–Stokes dispute (Scientific American)
- Cognition raises $2B at a $48B valuation (Bloomberg) - LOW_UTILITY: funding news with no technical artifact; the @cognition congrats posts folded here (Bloomberg)
- NSA/CISA/FBI joint advisory on “industrial-scale” Chinese AI distillation, naming DeepSeek (Reuters) - LOW_UTILITY: geopolitics aimed at a specific provider but no stack action and no first-party response in the feeds (Reuters)
- “When Does Memory Help?” cost-aware agent-memory eval (arXiv 2609.05441) - STALE: submitted Jul 26, surfaced by the arXiv holiday backlog rather than being new (arXiv)
- “A False Average: CoT monitors collapse where they are the only defense” (arXiv 2608.00583) - STALE: early-August submission re-listed across cs.AI/cs.LG/cs.CL today (arXiv)
- WSJ op-ed “Unregulated open-weight AI is an invitation to disaster” - LOW_UTILITY: op-ed, no artifact to act on (r/LocalLLaMA 494)
- OpenAI: “How GPT-5.6 Sol helps run quantum computing experiments” - LOW_UTILITY: agentic-research showcase in the quantum domain (HN 79)
- X zero direct keeps, 33rd straight run - DEDUP: the simonw/Bubeck Navier–Stokes reactions were folded into the lead rather than kept separately (simonw’s writeup) (X)
- Bluesky gamedev post: Virtual3D engine gains local-LLM support - THIN_GUARD: 7 likes, too thin to keep (Bluesky)
- BleepingComputer: threat actors switching to multi-agent attack frameworks - THIN_GUARD: 6 likes, headline-only item (Bluesky)
- lobsters zero keeps again - DEDUP: its single AI item is the OpenAI Navier–Stokes page, folded into the lead (On the Navier–Stokes Millennium Prize Problem) (lobsters)
- Dwarkesh Podcast: pretraining progress is mostly data - EXCLUSION: blocklisted author (gross-authors rule) (Dwarkesh Podcast · Techmeme)
- Sources: Q&A with Mark Zuckerberg on Muse and Meta - EXCLUSION: blocklisted subject (gross-authors rule); Muse coverage kept to non-Zuck outlets only (Sources · Techmeme)