Friday’s AI news is tooling-heavy, with durable execution, stateful policy, and better evaluation design linking the strongest releases and research. Pi 1.0 leads because its stable, MIT-licensed coding harness stays small while adding practical extension, routing, caching, and cost-inspection machinery. Pi Durable adds recoverable long-running agents, while DeepSeek ships a plugin-first harness preview; ContractRL repairs broken tool calls instead of regenerating them, and new work separates memory retention from retrieval, tests incident repair in Kubernetes, and exposes unreliable model rankings beneath stable agent scores. Sapien constrains actions using execution history, Cloudflare releases local decision models, and arXiv’s new submission cap shows how AI-assisted volume is straining research infrastructure.

Agent frameworks & tooling

  • Pi 1.0 — Pi reaches a stable MIT release while keeping its coding harness small and extension-oriented.

    • Adds Codemode/MCP, deferred tool loading, virtual models, Anthropic cache warming, and mid-conversation system messages.
    • Virtual models can route planning and implementation across providers, with a decision model choosing the handoff.
    • /session exposes per-model cost and cache use; current major-provider models are supported.
    • Installers cover Unix and Windows; code and documentation are public.
    • Flag: This is the project’s stability declaration, not an independent reliability evaluation.
    • (HN 1,330 · 423c)
  • Pi Durable — an experimental TypeScript substrate adds crash recovery and shared control to long-running Pi-based agents.

    • Memory, SQLite, and JSONL backends checkpoint every model and tool task before execution advances.
    • Interrupted tools rerun only when declared safe; requestId provides exactly-once submission handling.
    • Concurrent conversations can fork without copying history, use separate models and tools, and accept live steering.
    • Flag: APIs may change; one process owns a storage backend at a time.
    • (HN 394 · 50c)
  • DeepSeek Harness — an open-source desktop and web harness enters worldwide preview with a plugin-first architecture.

    • Runs everyday, coding, research, background, and scheduled tasks; detailed execution traces are inspectable.
    • Creator mode builds and installs plugins through chat; experimental plugins include subagents and approval review.
    • Launches with npx @deepseek-ai/dsh web; source, developer docs, and community plugins are linked.
    • Flag: Core plugin APIs are explicitly evolving during preview.
    • (HN 250 · 120c)

Models & research

  • ContractRL — repairs invalid tool-call JSON with bounded RFC 6902 patches instead of regenerating the entire object.

    • A contract-derived action mask blocks malformed or prohibited edits before deterministic validation.
    • Measured: 0.9362 semantic success · 34.4 generated tokens · regeneration baseline 0.9148 and 137.2 tokens.
    • Three-seed results beat Patch-SFT by 3.96 points; the 95% confidence interval is +1.37 to +6.62.
    • Flag: No implementation link appears on the abstract page.
    • (arXiv 2610.00328)
  • What Should an Agent Remember? — separates memory-retention failures from retrieval-ranking failures rather than mixing both into one recall score.

    • The benchmark holds access fixed across 300 seeded streaming episodes.
    • Query-aware selection adds 15.5 recall points; a confounded comparison misleadingly reports 68.7 points.
    • All 319 bounded-recency failures come from eviction; sufficiently old targets fall to zero recall.
    • Flag: Single-author study; code is linked, but the abstract reports no end-to-end agent-task result.
    • (arXiv 2610.00366)
  • Incident-Arena — evaluates coding agents on production-style incident repair inside ephemeral Kubernetes clusters.

    • Twenty human-built tasks inject faults across configuration and image layers under sustained load.
    • Functional verifiers check safe repair while holding system metrics stable, rather than relying on static tests.
    • Measured: 2.81M tokens per trial · 41 turns · frontier models below 64.3%.
    • Flag: No code or benchmark repository is linked on the abstract page.
    • (arXiv 2610.00648)
  • Agent Evaluation Reliability — reliable system rankings can still conceal unreliable rankings of the underlying models.

    • Bayesian variance decomposition covers 22 benchmarks from Holistic Agent Leaderboard and Harbor Index.
    • Measured: model-scaffold reliability 0.935–0.994 · underlying-model reliability 0.148–0.841 · similar-task gain at most 0.097.
    • Pooling diverse benchmarks raises projected reliability from 0.44 to 0.75 at equal task budget.
    • (arXiv 2610.00651)
  • Sapien — enforces stateful tool-call policies whose allowed actions depend on an agent’s prior steps and observations.

    • Policies combine regular-expression sequences, stateful predicates, deferred generation, and scoped semantic checks.
    • Utility stays within a few points of an unconstrained agent in the reported evaluation.
    • Under full hijacking, policies exclude 93–95% of AgentDojo attacks and 62–85% on Toolathlon.
    • Flag: The abstract links no implementation; results are author-reported.
    • (arXiv 2610.00797)

Industry

  • Clef and Clef-flash — Cloudflare releases Apache-2.0, Jev-compatible decision models for local use or Workers AI.
    • Clef adds image input and 64K context; backbones are Qwen3.8-27B and Qwen3.5-9B.
    • Measured median latency: Clef 209.3 ms · Flash 38.8 ms · Jev 524.1 ms.
    • Clef variants lead three of four workflow evaluations; Jev leads agent-trace observability.
    • Hosted fine-tuning begins as a managed service; self-serve training is promised later.
    • Flag: Benchmark and latency results are Cloudflare-run.
    • (HN 523 · 181c · Techmeme)

Policy & provenance

  • arXiv limits submitters to two papers per month — a repository-wide stopgap responds to AI-assisted submission volume and moderator overload.
    • September reached 40,363 submissions versus 20,569 in September 2024, generating almost 9,000 support tickets.
    • The limit spans categories; rejected submissions count, while co-authors are unaffected unless they submit.
    • The existing cap of three simultaneously active submissions remains.
    • cs.AI submissions grew more than sixfold in two years; arXiv cites thin, fragmented, AI-written papers.
    • (HN 122 · 57c · Techmeme)
All gathered items - what was cut and why (8)