Friday’s useful AI news is about making agents safer and more predictable, rather than another frontier-model launch. Claude’s usage-policy update leads because it changes deployment requirements for high-risk recommendations and physical control, including human review, stop mechanisms, and independently enforced limits. Around it, a skill-conflict study shows why successful task completion can hide broken constraints; NOMOS turns written policies into deterministic tool gates; reversible context archival tests token savings against reliability; and runtime-budget and trajectory-evaluation papers separate plausible progress from actual compliance. Smaller-model work rounds out the day: CPU speech recognition, compiled visual-model harnesses, and cross-model cache repair, alongside an opt-in vulnerability-reporting service for open-source maintainers.

Policy & provenance

  • Claude’s usage-policy update changes agent deployment requirements — operational changes concern surveillance, high-risk decisions, and physical control.
    • The current policy covers API integrations, cloud access, Claude Code, and downstream users.
    • Prohibited uses include nonconsensual tracking, surveillance tooling, weapons-control software, and deceptive political or commercial campaigns.
    • High-risk recommendations require meaningful qualified-person review and AI disclosure; specified wholly favorable decisions have review exceptions.
    • Dangerous physical actions require an observable stop mechanism, safe-state behavior on disconnection, and independently enforced operating limits.
    • Ordinary frustration and model testing are outside the extreme-abuse rule; conversation termination remains its primary enforcement mechanism.
    • Flag: governmental contracts can tailor restrictions; this is provider policy, not a new law or proof of model consciousness. (Techmeme)

Agent frameworks & tooling

  • OSS Scanner: free, opt-in security audits — eligible maintainers receive periodic Claude-generated vulnerability reports.

    • Enroll by PR to the registry; infrastructure and user-security impact determine eligibility.
    • Anthropic accepted 85/97 sampled high/critical findings; eleven were duplicates, one invalid.
    • Reports include reproducers, explanations, and patches or introduction bisections where available.
    • Flag: no human triage on delivery; selected-sample validation is not universal precision. (Techmeme)
  • Co-installed skills silently override constraints — task completion can stay unchanged while the intended skill’s rules disappear.

    • Measured: 312 conflicting pairs · three models · 6,368 runs · 169,294 tool calls.
    • Similar skills took over one in five runs; substitution was disclosed in only 0.9% of final replies.
    • A first-read pre-tool hook restored exclusive-function fidelity; installation location mattered more than listing order.
    • Flag: research, not a Hermes fix; score prohibitions separately from task success. (arXiv cs.AI)
  • NOMOS: deterministic tool-call gates — compile written policies instead of trusting prompts or per-call model judgments.

    • Four-pass compilation checks tool schemas, repairing or rejecting incompatible rule bindings.
    • Reported airline state-changing-call violations fall from 66.3% to 2.6%; retail falls from 30.8% to 6.9%.
    • Microsecond decisions need no model calls; compilation works with an on-premise 26B model.
    • Flag: no-tool-call failures escape gating; benign-utility costs are domain-dependent. (arXiv cs.AI)
  • Reversible forgetting archives tool results — a released harness replaces observations with notes while preserving originals.

    • User instructions and assistant messages are protected from archival.
    • One debugging case used 50% fewer cumulative input tokens but took 17% longer and made more requests.
    • Another workload saved nothing; neither debugging arm fully passed follow-up evaluation.
    • Flag: exploratory cases; an earlier continuation lost quality despite reduced context. (arXiv cs.AI)
  • On the Clock: agent runtime budgets — prompting a deadline neither controls runtime nor guarantees productive extra work.

    • Tests: Qwen3.6-27B, five MLE-Bench Lite competitions; Qwen3-4B, Zork I.
    • Timing feedback improves adherence without measurable performance loss; enforcement hooks tighten deadlines.
    • Budget-aware RL learns stopping, but extra time often produces repeated actions.
    • Flag: measure deadline compliance and productive time allocation separately. (arXiv cs.AI)
  • AgentHorizon: almost-correct computer-use trajectories — instruction swaps expose constraint violations that plausible completion summaries miss.

    • Dataset: 1,373 instruction–trajectory pairs · 166 human-recorded hours · three operating systems.
    • Evaluation: eleven judges · five harnesses · trajectories reaching 300 screenshots/actions.
    • Best agentic judge: 80.9% balanced accuracy; tool use worsens tested open-weight models.
    • Flag: verify final state and unwanted side effects independently. (arXiv cs.AI)

Models & research

  • Whistle: 16.9 MB CPU speech recognition — released weights share Needle’s runtime for small local voice-input deployments.

    • Seven languages, word timestamps, keyword biasing; maximum 30-second clips, 16 kHz mono.
    • Python: pip install cactus-needle; transcribe with needle.transcribe("clip.wav").
    • Vendor: 11.1 ms first token for ten-second audio on Apple M4 Pro.
    • Flag: October 2 release resurfaced on HN; speed uses different runtimes/precisions, accuracy uses published baselines. (HN 785 · 156c)
  • Harness Compilation for small VLMs — use traces to allocate decisions between the model and its harness.

    • A larger teacher revises reusable content/control offline; validation selects the deployed harness.
    • Seven visual-QA settings report 9.9–23.9-point gains with students no larger than 9B parameters.
    • No deployment teacher calls or weight updates; bounded choices and supplied-text reading remain useful student work.
    • Flag: uneven policy transfer across students; not a drop-in package. (arXiv cs.AI)
  • RaReCache repairs cross-model KV caches — targeted recomputation reduces prefill when escalating within a model family.

    • Qwen3-0.6B→14B retains 95–99% of target accuracy after recomputing 30% of positions across tested benchmarks.
    • Llama3-8B→70B retains 96.5% after recomputing 40%; reported maximum prefill speedup is 3.04×.
    • Rank disagreement identifies transfer-sensitive tokens rather than recomputing the entire shared context.
    • Flag: requires calibration and serving changes; not cross-provider sharing or a llama.cpp update. (arXiv cs.AI)
All gathered items - what was cut and why (8)
  • DeepSeek 4.1 Flash workflow essay - DEDUP: already covered in yesterday’s standalone site post. (HN)
  • Yes, and - DEDUP: recent standalone site coverage; no fresh development. (HN)
  • TypedBench - LOW_UTILITY: probability/cost evaluation has less immediate operational impact than today’s tool-gate and skill-conflict research. (arXiv)
  • TokenBank - OFFSTACK: inference-financing contracts, not model-serving tooling. (arXiv)
  • Braintrust logs alongside traces - UNVERIFIABLE: direct page blocked; no checked implementation artifact. (X @ankrgyl)
  • SynthID Detector announcement - UNVERIFIABLE: blocked post and unresolved shortened artifact links prevent verifying availability. (X @_philschmid)
  • GPT-6.1 “crushing Opus” thread - HYPE: vendor comparisons framed as settled superiority; September thread. (Reddit r/ClaudeAI)
  • 2019 Oil language story - STALE: October 2019 article, not current AI news. (lobste.rs)