Saturday’s top story: Qwen3.8-27B — Alibaba released Qwen3.8-27B weights under Apache 2.0: a 27B dense model with a vision encoder (images + hour-scale video), 262K native context extensible to 1M, and thinking-mode controls (reasoning_effort, preserve_thinking). Vendor-reported evals (Claude Code harness, 256K context): Terminal Bench 2.1 73.0, SWE-bench Pro 61.7, DeepSWE 1.1 42.2, OSWorld-Verified 84.3, WebArena-Verified 64.8. FP8 weights plus a ~17GB Q4_K_M GGUF are already up, vLLM/SGLang recipes are live, and llama.cpp runs it on DGX Spark. The numbers are the vendor’s own re-evals, but the artifact is real and downloadable now — this is the first dense Qwen3.8 size that fits a single GPU.

Agent frameworks & tooling

  • Introducing Toast 1 — Mixedbread’s specialized search agent: takes over the full retrieval loop (decompose → gather → inspect → curate) as a subagent, backend-agnostic over your existing indexes, ~$0.016–0.023/query at ~8s median latency. Vendor benchmarks claim frontier parity with Claude Opus 5 / GPT-5.6 Sol at up to 10× lower cost and 3.5× fewer tokens at identical task score on Harvey’s legal benchmark — mixedbread’s own numbers, but the API, golden harness repo, and live demo are real.

  • Maximizing the value of your Claude Code sessions — First-party Anthropic guidance on token economics in Claude Code: /clear between tasks, why switching /model or /effort mid-session busts the prompt cache (0.1× cache reads vs up to 2× writes, ~1h expiry), and /compact while the conversation is still cached. If you run Claude Code daily, this is the cheapest practical read of the week.

Models & research

  • Continued: GLM-5.3 — day 2 of coverage (base specs in yesterday’s digest). Reuters reports Z.ai’s 84.5% CyberGym now benchmarked against Anthropic’s Mythos 5 (83.8%) — the first Mythos comparison — and the most sensitive cybersecurity functions will be gated behind verified-user access.

  • Governed Persistent Memory: Source-Bound State Semantics and Fail-Closed Release for Long-Horizon Agents — Fresh arXiv (Aug 12): treats agent memory as an auditable bitemporal state machine with source-bound admission, retraction/deletion semantics, and fail-closed structured release — so stale or superseded records can’t support an outgoing claim. Sealed evals: governed lane 2,400/2,400 vs 600/2,400 ungoverned on a local 7B.

  • Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence — Fresh arXiv (Aug 13): tests the independence assumption behind multiplied reliability bounds — two instances of the same model co-fail on 90% of missions where either fails (18,000 preregistered missions, deterministic scoring). Redundancy is over-credited exactly when components share a model; the paper gives a finite-sample certificate LP instead. Code + preregistration released.

  • vToken: Token-Level Virtualization for Reclaimable KV Caches — Fresh arXiv (Aug 13): decouples logical token liveness from physical KV blocks so evicted tokens’ memory is actually reclaimable; implemented in vLLM, preserves PagedAttention/CUDA-Graph compatibility. Cuts retained KV blocks 27–72% and raises SLA-constrained throughput up to 1.37×.

  • Practice Makes Unsafe: Skill Misevolution in Self-Improving LLM Agents — Fresh arXiv (Aug 13): agents that distill successful trajectories into reusable skills can bake in unsafe ones — 3 malicious tasks raise carryover attack success 16.0%→35.3%. Ships SkillMisevo-Gym/Bench for lifecycle attribution plus SafeEvolve (cuts unsafe retrieval 26.7pp, fresh-session harm 17.3pp). Code released.

Industry

  • Anthropic revenue surged 14× YoY to $11.5B+ in Q2 — Investor docs show $11.5B+ Q2 revenue with positive adjusted operating income (vs $787M in Q2 2025); Reuters adds bankers pricing the IPO off a ~$190–200B 2028 revenue projection. The demand data point behind every Anthropic API price you pay.

Policy & provenance

  • Anthropic raises misalignment risk to “low,” won’t release internal “Model 2” — Latest risk report (Aug 14): risk estimate up from “very low” citing recent cybersecurity incidents; confirms a stronger internal model (“Model 2”) with “noticeable improvement” on internal tasks will NOT be released — and admits task-based evals “no longer capture increases in models’ capabilities.” The capability-measurement gap is the quietly important part.
All gathered items - what was cut and why (8)