Wednesday’s AI news is about what you can actually run or measure, and what still needs testing before you trust it. Mistral Large 4 leads because its API preview has independent performance and cost measurements, but promised open weights are not available for local deployment. EmbeddingGemma 2 offers multimodal retrieval on your own hardware; OpenAI’s Decisions API and the open Strands Decider present different ways to gate agent actions. Pi’s Codemode sketches a constrained tool-orchestration pattern, while new papers test stale context, permission-seeking, wasted retrieval, serving coordination, and privacy violations inside agent traces.

Industry

  • Mistral Large 4 opens an API preview — testable via API now, but not yet a self-hosted alternative.
    • Measured: 1T total / 49B active parameters · multimodal input · 512K context.
    • Artificial Analysis scores it 38 on its Intelligence Index, versus 39 for DeepSeek V4.1 Flash.
    • Its measured cost per index task is $1.13, versus $0.27 for DeepSeek V4.1 Flash at standard prices.
    • API pricing: $1.36/$4.18 per million input/output tokens; a two-week introductory discount halves those prices.
    • Flag: weights and architecture details are promised for late October; API results are not local-deployment benchmarks.
    • (HN 1834 · Techmeme · Artificial Analysis)

Agent frameworks & tooling

  • OpenAI Decisions API enters public beta — a hosted classifier endpoint for typed routing and guardrails without text generation.

    • POST /v1/decisions accepts gpt-6-luna only; questions return probabilities, choices, or rubric scores.
    • Text and images are supported; independent questions can share one input.
    • Input costs $0.10/M tokens before regional or long-context premiums; calibrate thresholds on labeled examples.
    • Flag: hosted beta, not private inference or structured-output extraction.
    • (HN 310 · 153c)
  • Strands Decider 2B ships code, weights, and training data — a local model for scoring proposed agent actions before tool calls.

    • Qwen3.5-2B torso with a pointer head; repository includes training scripts and an intervention example.
    • Authors report ~115 ms median on RTX 3090 and 153 ms for small tasks on M3 MacBook.
    • The example checks whether weather-lookup arguments came from the user before allowing the call.
    • Flag: hand-picked thresholds and public-set accuracy do not establish safety on your own tasks.
    • (HN 186 · 46c)
  • Pi 1.0’s Codemode puts orchestration inside a constrained harness sandbox — scripts compose model-native tools and MCP calls without exposing unrestricted execution.

    • QuickJS in WASM has no filesystem or network; calls go through harness-exposed tools.
    • Scripts can parallelize calls and retain session state without loading every intermediate result into model context.
    • Flag: MCP output is often inconsistent; durability and binary handling remain unfinished.
    • (HN 92 · 43c)
  • HEAR pairs agent-harness intent with inference-engine state — a proposed interface coordinates workflow dependencies with queue and KV-cache pressure.

    • Authors test cache-aware coordination and agent-role-specific execution under concurrent, memory-constrained serving.
    • Reported on SCBench: 1.61× batch speedup; 2.23× lower median first-token latency.
    • Flag: research prototype, not a standard supported by llama.cpp or vLLM.
    • (arXiv cs.AI)
  • Concord detects stale observations in an agent’s context — file changes after a tool read can invalidate facts still in the transcript.

    • It links observations to mutable sources, then updates, annotates, or suppresses stale context before reuse.
    • On constructed file-change cases across three models, authors report oracle-matching recovery and 46.4% fewer tokens than their strongest non-oracle baseline.
    • Flag: constructed workspace benchmark; not evidence of general cross-agent coherence.
    • (arXiv cs.AI)

Models & research

  • EmbeddingGemma 2 releases local multimodal embeddings — an Apache-2.0 model indexes code, text, images, audio, and video for private retrieval.

    • Weights and model card: 740M total / 270M text-only, with optional vision and audio encoders.
    • 8K context; 768-dimensional vectors truncatable to 512, 256, or 128. Benchmark recall before shrinking an existing index.
    • Google lists llama.cpp, MLX, vLLM, transformers.js, and Ollama integrations; code MTEB rises from 68.76 to 78.68 versus its predecessor.
    • Flag: quality and memory figures are vendor-reported; changing embedding models requires re-indexing.
    • (HN 337 · 35c · X @ggerganov)
  • DelegationBench tests whether agents ask before acting — a released 156-scenario benchmark tests approval behavior across wording changes and actual tool use.

    • Matched pairs vary request status, stakes, reversibility, or audience; outcomes include action, permission, clarification, and refusal.
    • Across ten models, equivalent prompts changed act rates by up to 52.5 percentage points.
    • Every model asked less often during tool use than when judging hypothetical actions; test execution, not just prompts.
    • (arXiv cs.AI)
  • Agents label retrieval useless but keep querying — failing-source experiments show prompt-only stopping advice misses the actual decision point.

    • Seven agents judged failing results useless 97–100% of the time, yet most seldom stopped for that reason.
    • A harness stop after five consecutive useless results improved failing-source success; authors replicated on 300 fresh questions.
    • Code is available; compare the five-call rule with your own tool budget.
    • (arXiv cs.AI)
  • AgentPrivArena audits unnecessary data access during tool use — final-answer leak checks miss privacy violations inside agent trajectories.

    • Its reproducible environment uses MCP tools and self-hosted services; metrics separate unnecessary access from output leakage.
    • Authors propose runtime auditing, but the abstract provides no numerical effect or linked implementation.
    • Flag: inspection method, not a verified deployable privacy guardrail.
    • (arXiv cs.AI)
All gathered items - what was cut and why (8)
  • llama.cpp DFlash tip - DEDUP: depends on yesterday’s covered release; no independently checked benchmark. (X @ggerganov)
  • OpenAI’s 372 mathematical results - UNVERIFIABLE: breakthrough count needs result-by-result mathematical review. (HN · Techmeme)
  • OpenTPU - OFFSTACK: accelerator design, not a deployment change for the existing inference stack. (HN)
  • Nano Banana 2.1 - LOW_UTILITY: image-generator pricing is less relevant than runnable retrieval and agent tooling. (Techmeme · The Decoder)
  • Mistral preview top-open-model claims - HYPE: broad superlatives omitted; measured performance, price, and availability are covered above. (VentureBeat)
  • Hard-excluded finance story (no URL found) - EXCLUSION: filtered before ranking; no public source identified in the digest. (source not recorded)
  • PewDiePie ban thread - DRAMA: high engagement, no stack change. (r/LocalLLaMA)
  • September DeepSeek V4.1-Flash thread - STALE: old thread, not today’s model release. (r/LocalLLaMA)