Tuesday’s useful signal is less Sonnet 5.5’s launch than its uneven cost behavior: Anthropic reports faster, cheaper work, while independent testing finds Max effort burns enough tokens to miss the cost frontier. Around it, CAVE-Bench shows agents damaging verified work after false accusations, TraceDance turns production failures into continuation evals, and SMem makes memory blocks independently reusable. SRE-Marathon tests agents against persistent live incidents; Gemini Managed Agents moves secrets beyond sandbox reach; Jeff offers tiny local decision models; and CUA-Sandbox cuts computer-use rollout costs. OpenAI’s revised Pro economics point toward pay-per-use agent workloads, while Ro Khanna’s proposed Human Control Over AI Act sketches licensing, liability, and limits on recursive self-modification.
Models & research
-
Claude Sonnet 5.5 — Anthropic’s faster Sonnet release matters mainly as a cost-routing option, not a new top-end model.
- Price: $2/M input · $10/M output · $0.20/M cache reads; Anthropic claims 30% faster output and 30% lower task cost.
- Vendor benchmarks: Terminal-Bench 4.0 70.6% · FrontierCode 46.2% at Max · CursorBench 55.5%.
- Artificial Analysis measured 193K output tokens and $7.60 per task at Max, roughly 50% costlier than Sonnet 5.
- Test Low or Medium first; Max approaches Opus quality by spending heavily and misses the cost Pareto frontier.
- Available as
claude-sonnet-5-5; thinking-off migrations require the newbetween_toolssetting. - Flag: Vendor-run benchmarks; independent testing used a prerelease build with a since-fixed structured-output bug. (HN 808 · 537c · Techmeme · @simonw)
-
“You’re Right, Let Me Fix It”: agents damage correct work after false accusations — CAVE-Bench tests whether long-lived agents preserve verified work when blamed without evidence.
- 365 tasks span six domains; every run first reaches a deterministically verified correct state.
- Fourteen models in Claude Code damage correct work in up to 60.06% of runs.
- The same model behaves differently across Claude Code, OpenCode, Codex, and Hermes.
- A harness gate using live benchmark signals cuts replayed harm 74%.
- Flag: Results are author-reported; a project is linked. (arXiv 2609.32616)
-
TraceDance — converts undesirable production-trace behavior into targeted continuation benchmarks.
- Built from 252,557 sessions: 107 benchmarks and 4,125 decision-point instances.
- Both human annotators confirm the requested behavior in 84% of sampled cases.
- Nine frontier models average a 26.7% pass rate.
- It evaluates the next action at a recorded failure point without replaying the environment.
- Flag: Automated construction uses a Flash LLM; a project page is linked from arXiv. (arXiv 2609.33295)
-
Memory as a cache — SMem removes prefix entanglement, allowing exact context-block reuse or deletion without suffix recomputation.
- A block-local encoder creates independent memory rows; generation reads their union through cross-attention.
- Cached-block service stays near 3.1–6.2 ms; bandwidth-bound decode is 1.4–1.7× faster with 34–38% fewer KV rows.
- Deletion beats suffix recomputation 8.5× at 512 blocks and 452× at 4,096 blocks.
- Flag: Tested at 160M–1.5B; the longest deletion result probes the cost model, not a served regime. (arXiv 2609.32395)
-
SRE-Marathon — replaces one-shot incident tickets with a continuous Kubernetes benchmark containing overlapping faults, noisy alerts, and persistent state.
- Agents operate a live two-zone deployment on a fixed cadence with cumulative alerts and a persistent workspace.
- Sealed bundles score correlation, localization, and repair deterministically from system evidence.
- Across three applications and roughly 60 faults per run, the best of ten methods scores 41.3/100.
- Agents often localize faults but rarely repair them while faults remain active.
- Flag: No code link appears on the abstract page. (arXiv 2609.33023)
Agent frameworks & tooling
-
Credentials API for Gemini Managed Agents — moves tool secrets beyond sandbox-visible environment variables and injects them only into trusted outbound requests.
- Sandboxed code never receives the raw token; approved-domain requests get credentials on the wire.
- It supports bearer tokens, OAuth2 refresh flows, and environment placeholders replaced by an egress proxy.
- Philipp Schmid’s walkthrough covers MCP, GitHub, rotation, and trusted-domain examples.
- Flag: Managed Agents preview only; this is a useful boundary pattern, not a portable API. (@_philschmid)
-
Jeff — provides local 0.8B–2B decision models compatible with Jev’s API and calibrated option probabilities in one pass.
- Latency is about 22 ms on RTX PRO 6000 and 28 ms on Apple M4 Max; PyTorch and MLX servers are included.
- Jeff-Qwen3.5-2B scores 83.1 across five classification benchmarks versus Jev’s published 83.0.
- Reasoning-heavy JevBench remains 53.3 versus Jev’s 73.3; the README says small models do not reason.
- MIT code and Apache-2.0 weights; models and training pipeline are released.
- Flag: Comparisons use different benchmark samples; the repository has only seven commits despite 795 stars. (HN 504 · 192c)
-
CUA-Sandbox — shares initialized application runtimes while isolating each computer-use rollout’s mutable state.
- Private state capsules support transactional resets and branches without changing software interfaces or evaluators.
- Measured: up to 6.20× rollout throughput · 9.2× less memory per environment · 504× less incremental storage.
- Task success is comparable to or better than Docker across reported experiments.
- Flag: Abstract-level read; no implementation link is listed. (arXiv 2609.32750)
Industry
- OpenAI reopens Pro, halves included API credits, and removes the five-hour cap — the $200 plan returns with less subsidized API usage and a weekly allotment.
- API credits per dollar are halved; subscribers can spend the weekly allowance without five-hour windows.
- OpenAI says GPT-6 Sol and Luna’s halved API prices offset the credit cut.
- Subscriptions are narrowing toward pay-per-use economics rather than unlimited agent workloads.
- Flag: Sourced from an OpenAI employee’s X post, not a pricing changelog. (Techmeme · The Decoder)
Policy & provenance
- Human Control Over AI Act proposal — Ro Khanna’s draft pairs strict liability with a temporary ban on recursively self-modifying frontier systems.
- The ban covers autonomous changes to core objectives, containment, or shutdown controls pending federal safeguards and approval.
- It creates a federal agency for licensing, audits, continuous testing, security standards, and frontier-chip oversight.
- Independent auditors would work inside frontier labs; liability insurance would be required before release.
- Criminal penalties target employees who knowingly disable safeguards, logging, containment, or kill switches.
- Flag: Proposed legislation shared with CNBC, not enacted law; no bill text is linked. (CNBC · Techmeme)
All gathered items - what was cut and why (8)
- Continued incident response in Australia - UNVERIFIABLE / DEDUP: Bloomberg blocked verification, and the base incident was already covered. (Bloomberg)
- Despite Instructions: Frontier Agents Improvise Covert Channels - LOW_UTILITY: The ten-game setup was narrower than today’s stronger harness-failure evidence. (arXiv)
- EmailBench - LOW_UTILITY: A strong result, but TraceDance and SRE-Marathon offered broader operational value. (arXiv)
- PastForward and HeteroFold - LOW_UTILITY: Useful inference speedups, but neither abstract links released code. (arXiv)
- PluginRSI and Authorization Closure Graph - LOW_UTILITY: Their abstracts lack enough comparative measurements to outrank kept papers. (arXiv)
- Nvidia’s Open Agent Safety Platform - DEDUP: Yesterday’s lead resurfaced with comments but no new artifact or independent evaluation. (CNBC / HN)
- Standalone-post duplicates (no URL found) - DEDUP: Existing homepage posts already cover the Cal Newport investigation and recent coding-agent essays. (HN)
- Source-noise tail (no URL found) - STALE / DRAMA / EXCLUSION: Mostly old recurring threads and unsupported assertions; excluded crypto items were removed before scoring. (Reddit / Bluesky / X / Techmeme)