Monday’s useful releases cluster around local inference and the question of whether agents make reliable decisions once deployed. Strata leads with a runnable way to split a large MoE across consumer GPU, RAM, and CPU, though its speed claims remain project-reported. Cloudflare’s Web Search API offers hosted live-search plumbing; VERSE tests changes to agent harnesses against held-out tasks; Red Hat’s guardrail comparison and two decision-model papers expose accuracy, cost, and label-sensitivity limits. GHOST finds older safety instructions can slip during long interactions, while terminal-agent training, email-agent robustness, and CITA each put a different part of the agent evaluation and tool-choice stack under test.

Lead — Local inference

  • Strata runs Qwen3.8-Flash-Next across GPU, RAM and CPU — an MIT inference engine makes a 125B-parameter MoE accessible on consumer NVIDIA/AMD machines.
    • Hot experts sit on GPU; others use RAM/CPU, with an SSD lookup table.
    • Requirements: 12 GB+ VRAM · 32 GB+ RAM · roughly 80 GB storage; 64 GB RAM supports larger quantizations.
    • Serves OpenAI-compatible /v1, Responses and Anthropic Messages endpoints for local coding agents.
    • Reported speed varies by quantization, card and prompt; the 4090 headline is not a universal benchmark.
    • Flag: Throughput and quality are project-reported; model weights retain their own licenses.
    • (HN 844 · 371c)

Agent frameworks & tooling

  • Cloudflare launches a Web Search API beta — AI Gateway exposes search for agents without making them guess live URLs.

    • Choose Ceramic.ai, Exa or Linkup through REST or Workers AI bindings.
    • Requests appear in gateway logs; provider list pricing applies without Cloudflare markup, or bring your own key.
    • Flag: Beta and hosted-provider dependency; not a replacement for private SearXNG.
    • (HN 96 · 52c)
  • VERSE tests agent-harness changes before shipping them — a verified optimizer edits executor harnesses and its own workflow without changing model weights.

    • Replays failures and checks regressions across held-out software-engineering tasks; code is published.
    • Validation-selected harness: 42.3% in-distribution, 37.7% out-of-distribution; baselines: 39.2% and 29.3%.
    • Flag: Author-reported results do not establish that unattended self-modification is safe in production.
    • (arXiv cs.AI)

Models & research

  • Red Hat compares decision-model guardrails with classifiers and LLM judges — prompt-injection and content-safety tests challenge blanket claims that Jev-style models outperform alternatives.

    • Injection accuracy: Qwen3.6-35B 89.31% · DeBERTa 89.01% · Jev 86.35%; DeBERTa median latency: 54.1 ms.
    • Jev scored 86.20% on content safety; policy tuning raised Laya from 57.87% to 75.20%.
    • Flag: Hardware and US API versus UK client complicate latency comparisons; only two risks were tested.
    • (HN 98 · 36c)
  • Fast Models, Slow Evidence audits System-1 decisions for agents — paired Jev/Laya tests and self-corrections show why gate accuracy alone does not justify deployment.

    • 7,283 base cases and 6,640 variants cover 11 decisions; raw outputs and code are available.
    • Jev beats Laya on nine decisions; neither beats chance on zero-shot model routing.
    • Authors revise a claimed 23.9% cost saving to 4.3% after counting prescreening.
    • Flag: Single self-audited study; evaluate on your task distribution.
    • (arXiv cs.AI)
  • Labels can override definitions in typed decision models — controlled prompt-rendering changes isolate a classification failure in open Jev-style implementations.

    • Dropping definitions left Laya accuracy essentially unchanged; neutral A/B labels improved accuracy by 15.11 percentage points.
    • One formatting change removed bias in three Laya checkpoints and introduced it in von, without weight changes.
    • Swap labels while holding definitions fixed before trusting typed routing or policy decisions.
    • Flag: Open implementations were tested, not proprietary Jev weights.
    • (arXiv cs.AI)
  • GHOST tests whether agents forget old safety constraints — benign long-running interactions can end in unsafe tool actions despite earlier restrictions.

    • Authors report 11.5% violations in their GPT-5.5 setup.
    • STAR-Guard restores historical constraints and audits before execution; no violations occurred in tested runs.
    • Flag: Zero observed failures in one experiment is not a general safety guarantee.
    • (arXiv cs.AI)
  • Terminal-agent training can stall on generated task difficulty — executable Docker tasks and tests alone do not ensure valid rewards or appropriate RL training difficulty.

    • Prompt and context changes raised baseline solvability 5.6×; a 9B model saturated at 81.3% mean pass@2.
    • Harder tasks dropped mean pass@2 to 20.6% without changing training settings.
    • Audit verifiers, infrastructure errors and model-specific solvability before attributing gains to training.
    • (arXiv cs.AI)
  • Email-agent evaluations miss request-style robustness — equivalent requests in different styles yield different retrieval and action results.

    • Tests vary five style axes and four dialect conditions across a RAG pipeline and two tool-using agents.
    • Indirect requests hurt all three; agents more often omit required actions than invent unsupported ones.
    • Evaluate completed actions separately from answer quality across realistic phrasings.
    • (arXiv cs.AI)
  • CITA estimates tool value before agents act — comparative supervision ranks alternative next tool calls instead of relying solely on final-outcome rewards.

    • Training combines logged actions, a Bayesian tool-graph simulator and LLM comparisons.
    • Authors report better Tool F1 and task success on three benchmarks and multiple backbones.
    • Flag: Abstract gives no numerical effect sizes or implementation link; inspect the paper before adopting.
    • (arXiv cs.AI)
All gathered items - what was cut and why (8)