Monday’s useful releases cluster around local inference and the question of whether agents make reliable decisions once deployed. Strata leads with a runnable way to split a large MoE across consumer GPU, RAM, and CPU, though its speed claims remain project-reported. Cloudflare’s Web Search API offers hosted live-search plumbing; VERSE tests changes to agent harnesses against held-out tasks; Red Hat’s guardrail comparison and two decision-model papers expose accuracy, cost, and label-sensitivity limits. GHOST finds older safety instructions can slip during long interactions, while terminal-agent training, email-agent robustness, and CITA each put a different part of the agent evaluation and tool-choice stack under test.
Lead — Local inference
- Strata runs Qwen3.8-Flash-Next across GPU, RAM and CPU — an MIT inference engine makes a 125B-parameter MoE accessible on consumer NVIDIA/AMD machines.
- Hot experts sit on GPU; others use RAM/CPU, with an SSD lookup table.
- Requirements: 12 GB+ VRAM · 32 GB+ RAM · roughly 80 GB storage; 64 GB RAM supports larger quantizations.
- Serves OpenAI-compatible
/v1, Responses and Anthropic Messages endpoints for local coding agents. - Reported speed varies by quantization, card and prompt; the 4090 headline is not a universal benchmark.
- Flag: Throughput and quality are project-reported; model weights retain their own licenses.
- (HN 844 · 371c)
Agent frameworks & tooling
-
Cloudflare launches a Web Search API beta — AI Gateway exposes search for agents without making them guess live URLs.
- Choose Ceramic.ai, Exa or Linkup through REST or Workers AI bindings.
- Requests appear in gateway logs; provider list pricing applies without Cloudflare markup, or bring your own key.
- Flag: Beta and hosted-provider dependency; not a replacement for private SearXNG.
- (HN 96 · 52c)
-
VERSE tests agent-harness changes before shipping them — a verified optimizer edits executor harnesses and its own workflow without changing model weights.
- Replays failures and checks regressions across held-out software-engineering tasks; code is published.
- Validation-selected harness: 42.3% in-distribution, 37.7% out-of-distribution; baselines: 39.2% and 29.3%.
- Flag: Author-reported results do not establish that unattended self-modification is safe in production.
- (arXiv cs.AI)
Models & research
-
Red Hat compares decision-model guardrails with classifiers and LLM judges — prompt-injection and content-safety tests challenge blanket claims that Jev-style models outperform alternatives.
- Injection accuracy: Qwen3.6-35B 89.31% · DeBERTa 89.01% · Jev 86.35%; DeBERTa median latency: 54.1 ms.
- Jev scored 86.20% on content safety; policy tuning raised Laya from 57.87% to 75.20%.
- Flag: Hardware and US API versus UK client complicate latency comparisons; only two risks were tested.
- (HN 98 · 36c)
-
Fast Models, Slow Evidence audits System-1 decisions for agents — paired Jev/Laya tests and self-corrections show why gate accuracy alone does not justify deployment.
- 7,283 base cases and 6,640 variants cover 11 decisions; raw outputs and code are available.
- Jev beats Laya on nine decisions; neither beats chance on zero-shot model routing.
- Authors revise a claimed 23.9% cost saving to 4.3% after counting prescreening.
- Flag: Single self-audited study; evaluate on your task distribution.
- (arXiv cs.AI)
-
Labels can override definitions in typed decision models — controlled prompt-rendering changes isolate a classification failure in open Jev-style implementations.
- Dropping definitions left Laya accuracy essentially unchanged; neutral A/B labels improved accuracy by 15.11 percentage points.
- One formatting change removed bias in three Laya checkpoints and introduced it in von, without weight changes.
- Swap labels while holding definitions fixed before trusting typed routing or policy decisions.
- Flag: Open implementations were tested, not proprietary Jev weights.
- (arXiv cs.AI)
-
GHOST tests whether agents forget old safety constraints — benign long-running interactions can end in unsafe tool actions despite earlier restrictions.
- Authors report 11.5% violations in their GPT-5.5 setup.
- STAR-Guard restores historical constraints and audits before execution; no violations occurred in tested runs.
- Flag: Zero observed failures in one experiment is not a general safety guarantee.
- (arXiv cs.AI)
-
Terminal-agent training can stall on generated task difficulty — executable Docker tasks and tests alone do not ensure valid rewards or appropriate RL training difficulty.
- Prompt and context changes raised baseline solvability 5.6×; a 9B model saturated at 81.3% mean pass@2.
- Harder tasks dropped mean pass@2 to 20.6% without changing training settings.
- Audit verifiers, infrastructure errors and model-specific solvability before attributing gains to training.
- (arXiv cs.AI)
-
Email-agent evaluations miss request-style robustness — equivalent requests in different styles yield different retrieval and action results.
- Tests vary five style axes and four dialect conditions across a RAG pipeline and two tool-using agents.
- Indirect requests hurt all three; agents more often omit required actions than invent unsupported ones.
- Evaluate completed actions separately from answer quality across realistic phrasings.
- (arXiv cs.AI)
-
CITA estimates tool value before agents act — comparative supervision ranks alternative next tool calls instead of relying solely on final-outcome rewards.
- Training combines logged actions, a Bayesian tool-graph simulator and LLM comparisons.
- Authors report better Tool F1 and task success on three benchmarks and multiple backbones.
- Flag: Abstract gives no numerical effect sizes or implementation link; inspect the paper before adopting.
- (arXiv cs.AI)
All gathered items - what was cut and why (8)
- Kolibri 1 - DEDUP: Yesterday’s model coverage; no new integration or independent benchmark. (Aleph Alpha)
- C2PA timestamp proof-of-concept - DEDUP: Yesterday’s provenance coverage; no correction or new verifier result. (David Buchanan)
- Trump’s Super Intelligence Force - UNVERIFIABLE: Direct page blocked; no policy text or concrete regulatory change confirmed. (Techmeme · Politico)
- Western open-weight models allegedly coming this month - UNVERIFIABLE: Unnamed sources and unreleased weights; revisit when an artifact ships. (Techmeme · Axios)
- China’s Claude token-reseller market - EXCLUSION: Excluded under topic policy. (Techmeme · The Information)
- Qwen on a 4090 at 100 tokens/s - HYPE / DEDUP: Repo covered above; one configuration’s speed is not a universal claim. (HN)
- Muse profiles friends and family - DEDUP: Recent standalone post covers the consumer-agent permission story. (Techmeme · Wired)
- AI spending forecast survey - LOW_UTILITY: No directly testable price comparison or method in the collected headline. (Techmeme · WSJ)