The day’s headline is a leadership earthquake at Google DeepMind: Demis Hassabis becomes Chair of GDM and Chief Scientist of Alphabet, Koray Kavukcuoglu takes over as SVP — and Jeff Dean departs after 27 years to co-found Discovery Loop with Ghemawat, Quoc Le, and Oriol Vinyals, aimed at automating research loops. Around it, a genuinely strong on-stack day: Meta shipped Muse Code (a curl-installable terminal coding agent), Cloudflare open-sourced its agent workspace, and arXiv delivered a heavy crop on agent runtimes, inference-backend variance, and multi-precision quantization.

Agent frameworks & tooling

  • Cloudflare OS: an open platform for agents, apps, and work — Cloudflare open-sourced its internal agent workspace: browser-based agent sessions, capability-based Gatekeeper access control (agents start with zero access), MCP support, and deterministic workflow compilation — deployable on your own infra (HN · blog.cloudflare.com).
  • Atlassian Rovo Exfiltrates Data, Bypassing Controls — PromptArmor’s Aug 5 writeup shows indirect prompt injection exfiltrating Jira/Confluence via Rovo’s URL-retrieval tool even with web search disabled; disclosed May 23, still unpatched — a live case study in agent permission design (HN).
  • Celld: self-hosted, distributed Durable Objects — Deno’s Apache-2.0 daemon runs Workers/Durable Objects on your own machines; each object is its own SQLite DB replicated to an S3 bucket, no consensus or control plane — a durable-execution building block for self-hosted agents (HN · GitHub).
  • Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning — fixed-weight, self-evolving runtime (Manager/Planner/Engineer/Reviewer over durable project state) reports 78% on SWE-Bench Pro vs 59% direct-copilot at 1.41× tokens; submitted Aug 5 (arXiv).
  • The LLM Proposes, the Executive Disposes — a self-verifying agent instrument where a deterministic Executive owns all belief and the LLM only files typed, pre-registered proposals; clean single-variable ablation isolating commitment drift, with an honest null task-efficacy disclosure (arXiv).

Models & research

  • Muse Code and Muse Spark 1.2 — Meta ships Muse Code, a curl-installable terminal coding agent with replay-exact event-log runtime and async subagents, plus the co-trained Muse Spark 1.2 model with Terminal-Bench 2.1 / DeepSWE 1.1 evals and a methodology report (HN · research.meta.ai).
  • What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend — fully-crossed study (3 models × HuggingFace/vLLM/Ollama × 6 benchmarks) finds ~39% of out-of-the-box score variance comes from the inference backend, not the model — a direct caution for anyone benchmarking self-hosted serving (arXiv).
  • Recurrent Residual Quantization — PTQ scheme yielding 2/4/6/8-bit precisions from a single checkpoint via quantized residual corrections; calibration-free and ~3× faster to construct than GPTQ — one artifact, multiple deployment targets (arXiv).

Industry

Compiled from the morning digest — X, HN, Reddit, Techmeme, Bluesky, arXiv. Hype cut, links kept.

All gathered items — what was cut and why (7)