Thursday’s AI news is about the cost of narrow agent work and the tests that tell you when a cheaper route is safe. Claude Haiku 5.5 leads because its pricing could change how often hosted agents delegate small tasks, though its coding benchmark leaves room for larger models. A property-testing playbook offers runnable ways to catch agent-written bugs, while a released home-automation harness shows where local models hold up and where unsafe device actuation breaks the case for substitution. Docker Agent packages declarative workflows, two papers examine skill selection and memory ownership, and a looped-model study tackles cache waste. Ecosia’s Mistral migration adds a customer-side data point, not a verdict on the new model.

Lead: Haiku 5.5 changes the price of narrow agent tasks

  • Claude Haiku 5.5 launches — Anthropic’s hosted small model targets cheaper compaction, classification, and subagent calls.
    • Pricing: $0.10/$0.50 per million input/output tokens up to 100K prompt tokens; $0.50/$2.50 above.
    • Cache reads cost $0.01/M below 100K; adjustable effort trades quality against expense.
    • Sonnet 5.5 cache reads fall to $0.10/M; Python/TypeScript browser and computer use enter beta.
    • Anthropic reports 39.2% on Terminal-Bench 4.0 against Sonnet 5.5’s 70.6%; route harder coding accordingly.
    • Flag: vendor-run benchmarks; measure task cost and refusal rates yourself. No local weights.
    • (HN 898 · 431c)

Agent frameworks & tooling

  • Docker Agent exposes YAML-defined agents as a Docker CLI plugin — a runnable Apache-2.0 option for declarative multi-agent workflows.

    • Run docker agent run agent.yaml; supports MCP toolsets, local Docker Model Runner, and hosted providers.
    • Repo documents hybrid retrieval, reranking, OCI distribution, and telemetry; check defaults before using private data.
    • Flag: HN resurfaced an existing repository, not a confirmed new release today.
    • (HN 245 · 114c)
  • Property testing with agent swarms — an open skill guides coding agents through reference models, adversarial sequences, and shrinking tests.

    • Compare simple reference behavior against optimized code; seed interacting operations instead of unrelated random inputs.
    • Author links upstream reports and merged uv, Prometheus, and KittyCAD fixes; land minimal regression tests separately.
    • Flag: larger bug counts include proposed fixes, not all merged or independently validated.
    • (lobste.rs)
  • Agent Skills’ downstream utility varies by harness — test skills against no-skill baselines in the model and harness you actually use.

    • Across 87 SkillsBench tasks and nine configurations, the same skills help some and hurt others on 36.78% of tasks.
    • Operation-support reranking beats relevance rankings by 4.35–5.80 pass-rate points in three configurations.
    • Flag: limited tasks and candidates; results may not transfer to your skills.
    • (arXiv cs.AI)
  • Whose Memory Is It? proposes scope-aware commits — CASK prevents hypothetical plans and others’ claims from becoming persistent agent facts.

    • Commits retain discourse ownership; provisional content stays branch- or speaker-scoped.
    • Controlled conversations and tool traces show cleaner memory admission and later answers.
    • Flag: research method, not an OpenViking integration; the abstract does not quantify improvement.
    • (arXiv cs.AI)
  • Agent Plasticity measures whether experience transfers — assess inherited agent artifacts on held-out tasks while accounting for learning cost.

    • Checkpoints separate initial capability from improvement rate; training gains only partly generalize out of distribution.
    • Traces distinguish missed retrieval from failures after relevant artifacts were used.
    • Flag: controlled environments, not proof of reliable self-improving production agents.
    • (arXiv cs.AI)

Models & research

  • Small models can handle four of five home-automation agent call sites — a released harness and records make per-call-site routing testable.

    • Measured: nine models · 280 cases · 2,520 scored calls · two Home Assistant installations.
    • Best local model matched hosted models statistically at four sites; code generation still differed.
    • Local routing reached 91.8% versus 95.4% hosted; one small model actuated 87.2% of unauthorized-device requests.
    • Flag: test permissions before substituting local models for device actuation.
    • (arXiv cs.AI)
  • Looped LMs get a practical dynamic-KV-cache strategy — separate cache state per iteration can erase compute savings from early exits.

    • Best-available KV caching reports up to 30% less FLOPs and KV memory on Ouro at full-depth performance.
    • A revised exit-prior objective makes iteration depth depend on token difficulty.
    • Flag: looped-model research, not a packaged runtime or generic llama.cpp speedup.
    • (arXiv cs.AI)

Industry

  • GPT-6 reaches more ChatGPT users with Intelligent UI — a chat rollout generates interactive charts, forms, and controls, not a Codex upgrade.

    • Starts with Plus, Pro, Business, and Enterprise; Free and Go follow the next day.
    • Chat tiers use GPT-6 Sol and Luna; OpenAI says Work and Codex models do not change.
    • Flag: UI availability is not a developer API for embedding generated interfaces.
    • (HN 649 · 355c · Techmeme)
  • Continued: Ecosia moves away from Mistral to hosted open-weight models — day 2 of coverage (base specs in yesterday’s digest).

    • Ecosia’s CEO cites recurring server overload and model quality; Melious now hosts Qwen, GLM, and Kimi for it.
    • He claims roughly half the cost and better quality/performance, without a reproducible task-level comparison.
    • Flag: complaints concern the prior Mistral service, not a measured failure of the new preview.
    • (Techmeme · Politico)
All gathered items - what was cut and why (8)
  • The Mathocalypse - DEDUP: recent standalone site post already covers it. (Scott Aaronson)
  • OpenAI’s 372 math results - UNVERIFIABLE: headline-scale claim lacks result-by-result verification; yesterday’s cut stands. (OpenAI)
  • RippleCP - LOW_UTILITY: 12-task checkpoint pilot is less actionable than runnable testing and routing work. (arXiv)
  • Nano Banana 2.1 - OFFSTACK: image generation showcase, not agent or inference infrastructure. (X @_philschmid)
  • Bluesky’s old EDIT-tool post - STALE: May result resurfaced in October. (Bluesky @antirez.bsky.social)
  • Agent swarm joke - HYPE: satire without a runnable artifact. (Bluesky @uhactually.bsky.social)
  • PewDiePie ban thread - DRAMA: personality story without a stack change. (r/LocalLLaMA)
  • Crypto/blockchain and prediction-market hits (no URL found) - EXCLUSION: aggregate excluded before ranking; individual URLs not recorded in the note. (source not recorded)