Friday’s digest is led by the biggest open-weight release since GLM-5.3-Flash: Tencent’s Hy4 preview, a 770B-parameter MoE with 49B active and a 1M-token context, shipped under Apache 2.0 alongside an API. Around it, an elevated arXiv feed with unusually practical agent-tooling work — evidence that the harness, not the model, drives coding-agent scores, token economics for when reasoning pays, and a full FP8 pretraining recipe that fits on consumer GPUs for under $7K. Industry-side: Anthropic previewed a hardware standard for agents that operate lab instruments, a US judge blocked the Pentagon’s blacklisting of the lab, and OpenAI is testing an always-on “Persistent mode” for Codex.
Agent frameworks & tooling
- Same Model, Different Harness: Different Coding-Agent Results — same frozen weights, different harness, very different scores: context-compaction raised SWE-bench Verified completions from 43 to 72 and held across four models; the clearest evidence yet that model+harness is the tested unit, so eval claims must name both. (arXiv)
- The load-bearing vocabulary of Claude — Show HN scraping 461K GitHub PRs finds a 2026 cluster where “load-bearing”, “quietly”, “refuses”, “asserted” run ~39× over baseline — a measurable fingerprint of agent-written code, useful for evals and detection. (HN 542 · Show HN)
Models & research
- Tencent open-sources Hy4 preview — 770B total / 49B active MoE, 1M-token context, Apache 2.0 weights plus API at ~$0.83/M input; Bloomberg reports it “outperforms Z.AI and Moonshot” in internal tests (vendor claim, flagged). Biggest open-weight release since GLM-5.3-Flash. (Techmeme · Bloomberg)
- Redwood: a frontier AI accelerator designed by AI in 2 weeks — Architect Labs reports its system generated performance models, RTL, UVM, formal proofs, firmware, and kernels end-to-end for an inference accelerator (1.75× throughput, 3.4× perf/watt vs measured Jetson baseline). Real 7-page paper, but every headline number is self-reported with no independent verification yet. (arXiv)
- Terminal-Bench-Science 0.1 — Stanford/Terminal-Bench team benchmarks agents on 70 expert-curated research workflows (920 proposals → 70 tasks); best model resolves 30%, GLM 5.3 is top open-weight at 8.1%. Open GitHub benchmark with cost/token Pareto frontiers. (HN 94)
- Puro-2B: pretraining from scratch on RTX 5090s for under $6.9K — Tsinghua PACMan’s complete FP8 pretraining recipe (data, code, weights, Apache 2.0) on consumer GPUs approaches Qwen2.5-1.5B; fitted cost law suggests ~$4.4K matches Qwen2-1.5B. (arXiv)
- The Reasoning Tax: token economics of LLM reasoning — TES metric over 151 model-benchmark runs: task structure predicts when reasoning tokens pay (AIME/LiveCodeBench high, MMLU-Pro low despite difficulty), and more thinking sometimes reduces accuracy — a benchmark-driven toggle rule for routing. (arXiv)
Industry
- OpenAI is testing “Persistent mode” in Codex — agents that “continue working until put to sleep” and proactively generate follow-up tasks; the always-on agent direction for the tool most of your agents already ride on, sources-say. (Techmeme · Wired)
- US judge blocks the Pentagon’s blacklisting of Anthropic — supply-chain-risk designation ruled “illegal and baseless”; a notable legal check on how federal agencies treat frontier labs. (Techmeme · Reuters)
- Anthropic previews the Model Hardware Standard — MCP-style driver spec for agents operating microscopes, liquid handlers, and robot arms; research preview with Genentech, CMU, Janelia, and QuEra, open-source planned. Platform move worth knowing even if it’s not your stack. (HN 119 · Techmeme/Wired)
All gathered items - what was cut and why (8)
- Nvidia agrees to acquire Hugging Face for $13B - DEDUP: same BI article as yesterday’s lead; the HN title now says “agrees” but the linked article still says “in talks” — no new reporting since yesterday (HN 1,899)
- Small Models Have Arrived - STANDALONE-DEDUP: already a site post (small-models-have-arrived); essay, no new artifact (HN 662)
- The turbulent AI era is here - STANDALONE-DEDUP: already a site post (turbulent-ai-era-bill-gates); essay (HN 311)
- Xiaomi AI Cube announced with 1.2TB/s memory bandwidth - DEDUP: Aug 25 keep, same announcement (r/LocalLLaMA)
- [Megathread] Qwen3.8-Flash-Next - Release Day - DEDUP: yesterday’s keep re-surfaced, zero new facts (r/LocalLLaMA)
- We found a division by zero bug in FFmpeg with a vibecoded fuzzer - CUT-capacity: real bug, genuinely fun, but single-anecdote evidence (HN 253)
- OpenAI, Anthropic, AWS, Microsoft, and 100+ companies warn there is “a limited window” - LOW_UTILITY: joint statement, no concrete deliverable (Techmeme/Axios)
- Gemini Omni 1.1 Flash - OFFSTACK: video-generation release, not the LLM/agent stack (HN 260 · X)