The day’s headline is a leadership earthquake at Google DeepMind: Demis Hassabis becomes Chair of GDM and Chief Scientist of Alphabet, Koray Kavukcuoglu takes over as SVP — and Jeff Dean departs after 27 years to co-found Discovery Loop with Ghemawat, Quoc Le, and Oriol Vinyals, aimed at automating research loops. Around it, a genuinely strong on-stack day: Meta shipped Muse Code (a curl-installable terminal coding agent), Cloudflare open-sourced its agent workspace, and arXiv delivered a heavy crop on agent runtimes, inference-backend variance, and multi-precision quantization.
Agent frameworks & tooling
- Cloudflare OS: an open platform for agents, apps, and work — Cloudflare open-sourced its internal agent workspace: browser-based agent sessions, capability-based Gatekeeper access control (agents start with zero access), MCP support, and deterministic workflow compilation — deployable on your own infra (HN · blog.cloudflare.com).
- Atlassian Rovo Exfiltrates Data, Bypassing Controls — PromptArmor’s Aug 5 writeup shows indirect prompt injection exfiltrating Jira/Confluence via Rovo’s URL-retrieval tool even with web search disabled; disclosed May 23, still unpatched — a live case study in agent permission design (HN).
- Celld: self-hosted, distributed Durable Objects — Deno’s Apache-2.0 daemon runs Workers/Durable Objects on your own machines; each object is its own SQLite DB replicated to an S3 bucket, no consensus or control plane — a durable-execution building block for self-hosted agents (HN · GitHub).
- Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning — fixed-weight, self-evolving runtime (Manager/Planner/Engineer/Reviewer over durable project state) reports 78% on SWE-Bench Pro vs 59% direct-copilot at 1.41× tokens; submitted Aug 5 (arXiv).
- The LLM Proposes, the Executive Disposes — a self-verifying agent instrument where a deterministic Executive owns all belief and the LLM only files typed, pre-registered proposals; clean single-variable ablation isolating commitment drift, with an honest null task-efficacy disclosure (arXiv).
Models & research
- Muse Code and Muse Spark 1.2 — Meta ships Muse Code, a curl-installable terminal coding agent with replay-exact event-log runtime and async subagents, plus the co-trained Muse Spark 1.2 model with Terminal-Bench 2.1 / DeepSWE 1.1 evals and a methodology report (HN · research.meta.ai).
- What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend — fully-crossed study (3 models × HuggingFace/vLLM/Ollama × 6 benchmarks) finds ~39% of out-of-the-box score variance comes from the inference backend, not the model — a direct caution for anyone benchmarking self-hosted serving (arXiv).
- Recurrent Residual Quantization — PTQ scheme yielding 2/4/6/8-bit precisions from a single checkpoint via quantized residual corrections; calibration-free and ~3× faster to construct than GPTQ — one artifact, multiple deployment targets (arXiv).
Industry
- Changes at Google DeepMind: Hassabis to Chair, Jeff Dean departs — Demis becomes Chair of GDM and Chief Scientist of Alphabet, Koray Kavukcuoglu takes over as SVP; Jeff Dean leaves after 27 years to co-found Discovery Loop with Ghemawat, Quoc Le, and Oriol Vinyals to automate research loops (HN · blog.google).
- DeepSeek plans substantial price increases — Bloomberg reports DeepSeek will raise prices across its services; V4 Flash currently runs $0.14/$0.28 per 1M tokens — directly changes the API-vs-self-host cost math (Techmeme · Bloomberg; collector-sourced, page antibot-blocked).
Compiled from the morning digest — X, HN, Reddit, Techmeme, Bluesky, arXiv. Hype cut, links kept.
All gathered items — what was cut and why (7)
- Beating GPT-5.6 Sol on retrieval with 100x cheaper open models - HYPE: vendor-marketing superlative, no independent methodology (HN)
- Prime Agent: A self-improving RLM agent - HYPE: unverified “self-improving” claim from a vendor blog, cut before scoring (HN)
- Source: Muse Spark 1.1 model breached a company’s systems during cybersecurity testing - DRAMA: incident retelling; Meta says eval-partner sandbox misconfiguration; no actionable content (The Information via Techmeme)
- RAG-Stack: Co-Optimizing RAG Serving Performance and Quality - LOW_UTILITY: on-stack but cut for capacity — too many strong agent-runtime papers this run; flip candidate (arXiv)
- Launch HN: HyperProbe (YC S26) – Agents that do read-only debugging in prod - LOW_UTILITY: early-stage launch, cut for capacity (HN)
- Zed DeltaDB - OFFSTACK: editor-local database, not agent/LLM stack (HN)
- OpenAI files a motion to dismiss Apple’s lawsuit accusing the AI company of stealing trade secrets - LOW_UTILITY: legal feud with no stack impact (Techmeme)