The headline today: Claude Code is switching to auto mode by default for Pro, Max, and Team plans starting August 14. Anthropic’s evals claim 89% harmful-action blocking (vs 13.6% for human reviewers) and 0/720 prompt-injection attacks succeeded against Fable 5 / Opus 5 / Sonnet 5 in a third-party eval — though Simon Willison’s analysis keeps a healthy dose of skepticism about the attack surface, including malicious packages in test suites.
Agent frameworks & tooling
- Auto mode is now the default in Claude Code for Pro, Max, and Team plans — Starting Aug 14, Claude Code sessions will default to auto mode. Anthropic’s evals claim 89% harmful-action blocking vs 13.6% for human reviewers, and 0/720 prompt-injection attacks succeeded against Fable 5 / Opus 5 / Sonnet 5 in a third-party eval. Simon Willison’s analysis includes both the data and healthy skepticism (the “malicious package in test suite” attack surface remains). (Techmeme · simonwillison.net)
Industry
- Software Giant SAP Stops Most Travel and Hiring Because of AI’s Soaring Cost — Internal SAP email obtained by 404 Media shows the freeze is still in effect since July, with exceptions only for AI-related hires/travel. An employee notes the company is rolling out a new internal AI tool “which massively increases costs.” Real data point on the tokenpocalypse hitting enterprise budgets. (HN 32pts · 404 Media)
All gathered items - what was cut and why (8)
- Google’s AI shakeup suggests it may be prioritizing AI diffusion over frontier-model leadership - LOW_UTILITY: well-argued opinion/analysis of DeepMind shakeup, not a verifiable development (Asimov’s Addendum)
- Amazon is backing a 7.65 GW gas plant for an off-grid TX AI data center - OFFSTACK: infrastructure story with no actionable angle for agent/LLM engineers (NYT via Techmeme)
- Vulnerabilities Found in Major AI Agent Frameworks - UNVERIFIABLE: Bluesky post unloadable, no linked paper or report (@record-lab)
- Three commands turn every LLM into an OpenAI-compatible endpoint - DEDUP: simonw’s llm plugin already covered in Aug 4-5 LLM release (@jeremymorgan)
- Tencent’s Agent Memory now supports team memory - DEDUP: same release kept in Aug 8 digest (@sungkim)
- TheoremDB - A public workspace for machine mathematics - OFFSTACK: ML-for-math, not agent/LLM stack (HN 43pts)
- Kill My SaaS evals release - LOW_UTILITY: competition evals for a coding-agent contest, no standalone artifact (@swyx)
- Could old GPUs become a tiny HPC cluster? - UNVERIFIABLE: speculative idea, no working code (@pitillo)