Monday’s pacing fight went global: China’s Foreign Ministry called the lab-CEO slowdown warnings fearmongering, its Ministry of State Security issued its first statement on AI, and AI-linked Asian stocks fell 5%+ — while Microsoft answered with a self-authored model code of conduct rather than any deceleration commitment. Around it, arXiv came back from the weekend with two harness papers worth reading: the first honest measurement of the SKILL.md pattern (real gains on some repositories, indistinguishable from run-to-run variance on others) and a same-model test of whether vendor harnesses actually win (neither pairing resolves an advantage). Also here: a strace teardown of Claude Code Web’s Firecracker microVM, a take-apart of the leaderboards this digest keeps quoting, the Agent Incident Registry’s 487 source-linked cases, and the data-center pollution report behind the EPA story.

Lead — The pacing fight goes global

Agent frameworks & tooling

  • Skill Issue: Lessons from Optimizing Repository SKILLs for Coding Agents (arXiv 2609.12742, submitted Sep 11) — the first honest measurement of the SKILL.md pattern this stack leans on: tasks are mined from merged PRs reverted at a single frozen base commit, and a candidate document is scored by whether the same agent does better with it than without. Over three Kotlin repositories, GEPA-found documents raise that score 4.9pp on average while SkillOpt-found ones land 0.1pp above the seed — and the authors state plainly that at one repository’s dataset size the GEPA gain cannot be separated from the agent’s own run-to-run variance. The best evidence for the pattern is qualitative: a maintainer of one repo found knowledge in the docs “one only gets by working in the project.” (arXiv)
  • Harness or Model? Isolating the Harness Effect in Agentic Coding (arXiv 2609.11987, submitted Sep 8, revised) — the paper behind the assumption that vendor-native harnesses win: 792 of 800 planned runs graded on a private 256-task suite, same-model contrasts. Neither pairing resolves an average advantage (Opus 4.8: claude-agent-sdk 48.8% vs deepagents 50.0%, −1.25pp, 95% CI [−10.0, +7.5]; GPT-5.5: codex SDK 55.6% vs deepagents 54.4%, +1.25pp, CI [−4.4, +6.9]). The Opus average hides opposite strata (−9.0pp on 61 repository tasks, +23.7pp on 19 contest tasks, p = 0.003) and the authors flag that partition as chosen after seeing the data and needing designed replication. Two details worth carrying into your own eval harness: 22 of 81 runs cancelled at the wall-clock ceiling had already produced a passing patch, and the cost comparison was re-priced after a telemetry defect in the authors’ own earlier manuscript — with 58 Anthropic runs missing usage records, so the billed ordering stays unresolved. Orchestrator, grading oracle and reanalysis code released; the tasks stay private. (arXiv)
  • Reverse-Engineering Claude Web’s MicroVM: Anthropic’s hidden “Antspace” (HN 96 · 21 comments) — a full teardown of how Claude Code Web actually isolates a session, obtained with nothing but strace, strings and go tool objdump inside the session (no exploit): Firecracker microVM, 4 vCPU/16GB/252GB, kernel 6.18.5, PID 1 is a custom process_api exposing a WebSocket process supervisor on :2024 and an HTTP container-control API on :2025, sessions restored from frozen VM snapshots with block devices hot-swapped at load (vda rootfs + squashfs overlays), plus an unstripped Go environment-manager binary naming an unreleased internal PaaS. Directly reusable if you build agent sandboxes — snapshot-resume hygiene (init_on_free=1, CRNG reseed, cache drop, CAP_SYS_RESOURCE drop) is the checklist. Honest caveat: the analysis was performed in March 2026 and resurfaced on HN today; it is new to the feed, not new work. (HN)
  • Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires (HN 31 · 12 comments) — Dan Luu taking apart the leaderboards this digest has been quoting: DeepSWE’s 113 tasks × 4 runs rank GPT-5.5 above Fable 5 on names that are mostly in languages he doesn’t use agents for (4 of 79 differing tasks in Rust; ~1 of 113 resembles his work), and Senior SWE-Bench adds discontinuous “tasteful solve” thresholds — on one task GLM-5.2 passes at 121 LOC against a 61-LOC reference, so one more line flips the result, from a single run per condition. Same post shows the popular napkin-math numbers are wrong (random-memory latency measured without data dependencies; SSD reads that touch page cache; 8 GiB/s seq read vs Google’s 5,000 MiB/s spec for the 8-disk instance). Reviewed with Aaron Levin, who ran an evals team at Anthropic. (HN)

Models & research

  • Reality Is the Final Verifier: On Two Key Gaps in Agentic Software Engineering (arXiv 2609.12039, submitted Sep 10) — Berkeley (Krentsel, Cemri, Mao, Zaharia, Stoica et al.) frames agentic SWE failure as a requirement gap (requirements only approximate stakeholder intent) plus a model gap (the environment model only approximates deployment), and shows the two failure modes fall out of it: reward hacking exploits the omissions, hallucination widens the gaps by fabricating assumptions. Because neither can be certified closed in a changing world, the move is an assurance-revision loop that uses deployment evidence to revise requirements, environment model or evaluator — and the bottleneck is resource allocation across human judgment (requirement gap) and faithful, costly evaluation (model gap). A position paper, explicitly not an experiment — no results table, so nothing here to adopt except the framing. (arXiv)
  • The Agent Incident Registry (arXiv 2609.11030, v2 Sep 11) — 487 source-linked agent-related records from 2022–2026, each with evidence, a stable identifier and missingness-aware labels for causal role, disclosure class, mechanism and outcome; all 487 were re-checked by a second human reviewer. Of the 336 generative-system records where the agent acted, 81 involved realized harm (24%) — and the authors say outright that this share reflects collection composition rather than deployment risk. The useful output for eval design: InjecAgent’s 1,054 cases occupy three of the registry’s twelve surfaces and are all attacker-triggered, while the registry holds 92 no-adversary safety failures — i.e. scope your agent-security evals beyond injection. Project page at enkryptai.com/air; explicitly not a failure-rate estimate. (arXiv)
  • Open-Source AI & Open Models Reading List (Nathan Lambert, HN 115 · 22 comments) — a curated, actively maintained entry point to the open-weights debate, updated Sep 13 and open to additions: what open models are and why labs ship them (Solaiman’s release-gradient paper), the data-commons collapse, the US-China competition thread, safety positions (Thinking Machines’ safe-path-to-open-weights, the societal-impact paper), and adoption data (ATOM report, adoption dashboard, artifacts hub). Useful as the reference shelf for the open-weights argument rather than as news. (HN)

Policy & provenance

  • Trump is giving data centers a pass to pollute (The Verge, Sep 12 · via Techmeme) — the externalities story with an actual document behind it: the Environmental Protection Network’s report identifies 30 federal actions since January 2025 that its former-EPA authors say worsen health risks from data-center pollution, 17 of which specifically mention AI or target data centers, and proposes a “Data Center Health Protection Pledge.” Cited inside: a UC Riverside/Caltech/RIT study (arXiv 2412.06288) projecting up to 1,300 premature deaths and $20B+ in public health costs by 2028. Quotes from Lynn Goldman and Marc Boom, with the EPA’s response (“returned regulations to the best reading of the Clean Air Act”) included. Article dated Sep 12 — surfaced by Techmeme today. (Techmeme · The Verge)
All gathered items - what was cut and why (37)