Wednesday’s AI news is about what you can actually run or measure, and what still needs testing before you trust it. Mistral Large 4 leads because its API preview has independent performance and cost measurements, but promised open weights are not available for local deployment. EmbeddingGemma 2 offers multimodal retrieval on your own hardware; OpenAI’s Decisions API and the open Strands Decider present different ways to gate agent actions. Pi’s Codemode sketches a constrained tool-orchestration pattern, while new papers test stale context, permission-seeking, wasted retrieval, serving coordination, and privacy violations inside agent traces.
Industry
- Mistral Large 4 opens an API preview — testable via API now, but not yet a self-hosted alternative.
- Measured: 1T total / 49B active parameters · multimodal input · 512K context.
- Artificial Analysis scores it 38 on its Intelligence Index, versus 39 for DeepSeek V4.1 Flash.
- Its measured cost per index task is $1.13, versus $0.27 for DeepSeek V4.1 Flash at standard prices.
- API pricing: $1.36/$4.18 per million input/output tokens; a two-week introductory discount halves those prices.
- Flag: weights and architecture details are promised for late October; API results are not local-deployment benchmarks.
- (HN 1834 · Techmeme · Artificial Analysis)
Agent frameworks & tooling
-
OpenAI Decisions API enters public beta — a hosted classifier endpoint for typed routing and guardrails without text generation.
POST /v1/decisionsacceptsgpt-6-lunaonly; questions return probabilities, choices, or rubric scores.- Text and images are supported; independent questions can share one input.
- Input costs $0.10/M tokens before regional or long-context premiums; calibrate thresholds on labeled examples.
- Flag: hosted beta, not private inference or structured-output extraction.
- (HN 310 · 153c)
-
Strands Decider 2B ships code, weights, and training data — a local model for scoring proposed agent actions before tool calls.
- Qwen3.5-2B torso with a pointer head; repository includes training scripts and an intervention example.
- Authors report ~115 ms median on RTX 3090 and 153 ms for small tasks on M3 MacBook.
- The example checks whether weather-lookup arguments came from the user before allowing the call.
- Flag: hand-picked thresholds and public-set accuracy do not establish safety on your own tasks.
- (HN 186 · 46c)
-
Pi 1.0’s Codemode puts orchestration inside a constrained harness sandbox — scripts compose model-native tools and MCP calls without exposing unrestricted execution.
- QuickJS in WASM has no filesystem or network; calls go through harness-exposed tools.
- Scripts can parallelize calls and retain session state without loading every intermediate result into model context.
- Flag: MCP output is often inconsistent; durability and binary handling remain unfinished.
- (HN 92 · 43c)
-
HEAR pairs agent-harness intent with inference-engine state — a proposed interface coordinates workflow dependencies with queue and KV-cache pressure.
- Authors test cache-aware coordination and agent-role-specific execution under concurrent, memory-constrained serving.
- Reported on SCBench: 1.61× batch speedup; 2.23× lower median first-token latency.
- Flag: research prototype, not a standard supported by llama.cpp or vLLM.
- (arXiv cs.AI)
-
Concord detects stale observations in an agent’s context — file changes after a tool read can invalidate facts still in the transcript.
- It links observations to mutable sources, then updates, annotates, or suppresses stale context before reuse.
- On constructed file-change cases across three models, authors report oracle-matching recovery and 46.4% fewer tokens than their strongest non-oracle baseline.
- Flag: constructed workspace benchmark; not evidence of general cross-agent coherence.
- (arXiv cs.AI)
Models & research
-
EmbeddingGemma 2 releases local multimodal embeddings — an Apache-2.0 model indexes code, text, images, audio, and video for private retrieval.
- Weights and model card: 740M total / 270M text-only, with optional vision and audio encoders.
- 8K context; 768-dimensional vectors truncatable to 512, 256, or 128. Benchmark recall before shrinking an existing index.
- Google lists llama.cpp, MLX, vLLM, transformers.js, and Ollama integrations; code MTEB rises from 68.76 to 78.68 versus its predecessor.
- Flag: quality and memory figures are vendor-reported; changing embedding models requires re-indexing.
- (HN 337 · 35c · X @ggerganov)
-
DelegationBench tests whether agents ask before acting — a released 156-scenario benchmark tests approval behavior across wording changes and actual tool use.
- Matched pairs vary request status, stakes, reversibility, or audience; outcomes include action, permission, clarification, and refusal.
- Across ten models, equivalent prompts changed act rates by up to 52.5 percentage points.
- Every model asked less often during tool use than when judging hypothetical actions; test execution, not just prompts.
- (arXiv cs.AI)
-
Agents label retrieval useless but keep querying — failing-source experiments show prompt-only stopping advice misses the actual decision point.
- Seven agents judged failing results useless 97–100% of the time, yet most seldom stopped for that reason.
- A harness stop after five consecutive useless results improved failing-source success; authors replicated on 300 fresh questions.
- Code is available; compare the five-call rule with your own tool budget.
- (arXiv cs.AI)
-
AgentPrivArena audits unnecessary data access during tool use — final-answer leak checks miss privacy violations inside agent trajectories.
- Its reproducible environment uses MCP tools and self-hosted services; metrics separate unnecessary access from output leakage.
- Authors propose runtime auditing, but the abstract provides no numerical effect or linked implementation.
- Flag: inspection method, not a verified deployable privacy guardrail.
- (arXiv cs.AI)
All gathered items - what was cut and why (8)
- llama.cpp DFlash tip - DEDUP: depends on yesterday’s covered release; no independently checked benchmark. (X @ggerganov)
- OpenAI’s 372 mathematical results - UNVERIFIABLE: breakthrough count needs result-by-result mathematical review. (HN · Techmeme)
- OpenTPU - OFFSTACK: accelerator design, not a deployment change for the existing inference stack. (HN)
- Nano Banana 2.1 - LOW_UTILITY: image-generator pricing is less relevant than runnable retrieval and agent tooling. (Techmeme · The Decoder)
- Mistral preview top-open-model claims - HYPE: broad superlatives omitted; measured performance, price, and availability are covered above. (VentureBeat)
- Hard-excluded finance story (no URL found) - EXCLUSION: filtered before ranking; no public source identified in the digest. (source not recorded)
- PewDiePie ban thread - DRAMA: high engagement, no stack change. (r/LocalLLaMA)
- September DeepSeek V4.1-Flash thread - STALE: old thread, not today’s model release. (r/LocalLLaMA)