Thursday’s AI news is unusually concrete: measured model economics, task-specific agent harnesses, and deployment evidence outweigh launch-day claims. Gemini 4 Argon leads because independent testing turns its million-token output window into an operational question about cost and unusually long generations, while access remains limited. Around it, Magnitude brings self-tuning local inference to common coding agents; STITCH and a self-evolving harness argue for task-specific control loops; risk-aware evaluation and CESS make scarce tests and research traces more auditable. Long-generation reliability, adaptive vLLM prefill, and AgBench add hard limits and deployment trade-offs, while California’s No Robo Bosses Act establishes a human-corroboration floor for workplace discipline.

Industry

  • Gemini 4 Argon — Google previews a long-output frontier model; independent testing makes price, token use, and availability the useful story.
    • Introductory pricing is $2/M input, $10/M output, $0.10/M cached input; standard pricing doubles later.
    • Context and output limits are both 1M tokens; Long Decode Continuation resumes long responses across API calls.
    • Artificial Analysis scores high reasoning at 53, tied with GPT-6 Astra and one point above GPT-6.1 Sol.
    • Measured cost is $1.99/task now, rising to $3.98; average output is 62K tokens versus Astra’s 27K.
    • Flag: Selected cyber defenders have access now; developer, enterprise, and consumer access is promised, not generally available.
    • (HN 1,416 · 931c · Techmeme · @_philschmid)

Agent frameworks & tooling

  • Magnitude — an Apache-2.0 local inference engine compiles and tunes kernels for the machine running each agent.

    • Supports Apple Silicon, Nvidia, AMD, CPU, Linux, macOS, and Windows through an OpenAI-compatible API.
    • Integrates with Hermes, Pi, OpenCode, Codex, Claude Code, and Cline; concurrent sessions share prefix caches.
    • Project benchmarks report 9% faster prefill and 92% faster decode on Metal versus llama.cpp.
    • Flag: Speed measurements are project-run; optimized model-family coverage is narrower than a generalist engine’s.
    • (HN 164 · 84c)
  • STITCH — selects reusable harness primitives per task instead of forcing one global agent loop onto every workload.

    • Primitives are mined from failed trajectories and carry explicit scope and composition contracts.
    • Test-time selection avoids generating or debugging new mechanism code during a run.
    • Reported success improves by up to 12 points over fixed harnesses with 2.7% composition overhead.
    • Flag: No implementation link appears on the abstract page.
    • (arXiv 2609.38912)
  • Self-Evolving Harness on Multiple Tasks — one frozen model solves tasks, reads complete traces, and edits the 49-line harness running it.

    • Evolution uses five benchmarks; evaluation separates held-out tasks and adds five unseen benchmarks.
    • The evolved harness gains 4.48 points in-distribution and 12.64 points out-of-distribution.
    • Emergent mechanisms include output truncation, history compaction, and independent review.
    • Flag: Single-author study; no code link appears on the abstract page.
    • (arXiv 2609.38372)
  • Risk-Aware Adaptive Evaluation — allocates scarce agent-evaluation trials toward scenarios combining failure likelihood with impact.

    • Offline replay covers 70 τ-bench airline scenarios and 824 recorded trials.
    • At 50 trials, it finds 86% of oracle impact-weighted failures versus 25% under uniform allocation.
    • Discovery per dollar rises 5×; budget spent on never-failing scenarios falls from 34% to 2.8%.
    • Flag: Results are offline replay on one benchmark domain; no implementation link is listed.
    • (arXiv 2609.38914)
  • Search Shapes Conclusions — audits whether a research agent’s evidence reflects the candidate pool rather than its adaptive search path.

    • CESS logs document-selection and round-reach probabilities, then corrects the estimated evidence direction.
    • On MS2 questions, mean absolute error falls 9.2%; ranking sensitivity falls 39.4%.
    • Public Open Deep Research traces show corresponding reductions of 60.1% and 87.2%.
    • Flag: Requires a known candidate pool and logged selection probabilities; no code link appears.
    • (arXiv 2609.39026)

Models & research

  • Staying on Task — Long-Transduction isolates whether long generations lose state as context, formatting, or local task complexity changes.

    • Seven open-weight models perform repeated arithmetic, sorting, lookup, and table transformations.
    • Scaling context from 4K to 128K lowers performance 62.8%.
    • Input-format changes cost 36.5%; greater local complexity costs 39.9%.
    • Flag: Controlled transductions diagnose reliability foundations, not end-to-end tool-using agents; no code link appears.
    • (arXiv 2609.38712)
  • Decode-Latency Feedback Prefill — a vLLM controller resizes overlapping prefill chunks using observed decode latency instead of a fixed hardware-specific setting.

    • Three A100/Qwen3-0.6B trials reduce P99 inter-token latency by 24.8–30.1%, with exact output agreement.
    • Mean P99 time-to-first-token rises 34.8%, remaining inside the declared SLO.
    • The method fails to generalize to Qwen3-8B, 32B, or two-GPU tensor parallelism.
    • Flag: Reproducible proof-of-concept, not a general scheduler improvement; implementation is not linked.
    • (arXiv 2609.38386)
  • AgBench — open artifacts compare local, hybrid, and cloud agent execution across success, latency, cost, goodput, and data exposure.

    • Results span 162.07M data points across personal-device workloads and deployment architectures.
    • Local-only removes cloud API cost and exposure but usually lowers success and slows completion as concurrency rises.
    • Hybrid improves some tasks, but its cost and exposure depend on work partitioning and shared information.
    • Flag: The release artifact is anonymously hosted during review.
    • (arXiv 2609.38652)

Policy & provenance

  • California’s No Robo Bosses Act — SB 947 bars automated systems from being the sole or principal basis for workplace discipline and firing.
    • Primarily AI-based decisions require human corroboration using additional evidence such as reviews and personnel files.
    • Workers receive written notice, the data categories used, and a human contact who can explain the decision.
    • The enacted revision drops advance-notice requirements and gig-worker coverage from the vetoed 2025 bill.
    • Flag: “Primarily relies” is undefined, leaving implementation and litigation risk.
    • (CNBC · Techmeme)
All gathered items - what was cut and why (9)