Friday’s digest is led by the biggest open-weight release since GLM-5.3-Flash: Tencent’s Hy4 preview, a 770B-parameter MoE with 49B active and a 1M-token context, shipped under Apache 2.0 alongside an API. Around it, an elevated arXiv feed with unusually practical agent-tooling work — evidence that the harness, not the model, drives coding-agent scores, token economics for when reasoning pays, and a full FP8 pretraining recipe that fits on consumer GPUs for under $7K. Industry-side: Anthropic previewed a hardware standard for agents that operate lab instruments, a US judge blocked the Pentagon’s blacklisting of the lab, and OpenAI is testing an always-on “Persistent mode” for Codex.

Agent frameworks & tooling

  • Same Model, Different Harness: Different Coding-Agent Results — same frozen weights, different harness, very different scores: context-compaction raised SWE-bench Verified completions from 43 to 72 and held across four models; the clearest evidence yet that model+harness is the tested unit, so eval claims must name both. (arXiv)
  • The load-bearing vocabulary of Claude — Show HN scraping 461K GitHub PRs finds a 2026 cluster where “load-bearing”, “quietly”, “refuses”, “asserted” run ~39× over baseline — a measurable fingerprint of agent-written code, useful for evals and detection. (HN 542 · Show HN)

Models & research

  • Tencent open-sources Hy4 preview — 770B total / 49B active MoE, 1M-token context, Apache 2.0 weights plus API at ~$0.83/M input; Bloomberg reports it “outperforms Z.AI and Moonshot” in internal tests (vendor claim, flagged). Biggest open-weight release since GLM-5.3-Flash. (Techmeme · Bloomberg)
  • Redwood: a frontier AI accelerator designed by AI in 2 weeks — Architect Labs reports its system generated performance models, RTL, UVM, formal proofs, firmware, and kernels end-to-end for an inference accelerator (1.75× throughput, 3.4× perf/watt vs measured Jetson baseline). Real 7-page paper, but every headline number is self-reported with no independent verification yet. (arXiv)
  • Terminal-Bench-Science 0.1 — Stanford/Terminal-Bench team benchmarks agents on 70 expert-curated research workflows (920 proposals → 70 tasks); best model resolves 30%, GLM 5.3 is top open-weight at 8.1%. Open GitHub benchmark with cost/token Pareto frontiers. (HN 94)
  • Puro-2B: pretraining from scratch on RTX 5090s for under $6.9K — Tsinghua PACMan’s complete FP8 pretraining recipe (data, code, weights, Apache 2.0) on consumer GPUs approaches Qwen2.5-1.5B; fitted cost law suggests ~$4.4K matches Qwen2-1.5B. (arXiv)
  • The Reasoning Tax: token economics of LLM reasoning — TES metric over 151 model-benchmark runs: task structure predicts when reasoning tokens pay (AIME/LiveCodeBench high, MMLU-Pro low despite difficulty), and more thinking sometimes reduces accuracy — a benchmark-driven toggle rule for routing. (arXiv)

Industry

  • OpenAI is testing “Persistent mode” in Codex — agents that “continue working until put to sleep” and proactively generate follow-up tasks; the always-on agent direction for the tool most of your agents already ride on, sources-say. (Techmeme · Wired)
  • US judge blocks the Pentagon’s blacklisting of Anthropic — supply-chain-risk designation ruled “illegal and baseless”; a notable legal check on how federal agencies treat frontier labs. (Techmeme · Reuters)
  • Anthropic previews the Model Hardware Standard — MCP-style driver spec for agents operating microscopes, liquid handlers, and robot arms; research preview with Genentech, CMU, Janelia, and QuEra, open-source planned. Platform move worth knowing even if it’s not your stack. (HN 119 · Techmeme/Wired)
All gathered items - what was cut and why (8)