Thursday’s AI news is unusually concrete: measured model economics, task-specific agent harnesses, and deployment evidence outweigh launch-day claims. Gemini 4 Argon leads because independent testing turns its million-token output window into an operational question about cost and unusually long generations, while access remains limited. Around it, Magnitude brings self-tuning local inference to common coding agents; STITCH and a self-evolving harness argue for task-specific control loops; risk-aware evaluation and CESS make scarce tests and research traces more auditable. Long-generation reliability, adaptive vLLM prefill, and AgBench add hard limits and deployment trade-offs, while California’s No Robo Bosses Act establishes a human-corroboration floor for workplace discipline.
Industry
- Gemini 4 Argon — Google previews a long-output frontier model; independent testing makes price, token use, and availability the useful story.
- Introductory pricing is $2/M input, $10/M output, $0.10/M cached input; standard pricing doubles later.
- Context and output limits are both 1M tokens; Long Decode Continuation resumes long responses across API calls.
- Artificial Analysis scores high reasoning at 53, tied with GPT-6 Astra and one point above GPT-6.1 Sol.
- Measured cost is $1.99/task now, rising to $3.98; average output is 62K tokens versus Astra’s 27K.
- Flag: Selected cyber defenders have access now; developer, enterprise, and consumer access is promised, not generally available.
- (HN 1,416 · 931c · Techmeme · @_philschmid)
Agent frameworks & tooling
-
Magnitude — an Apache-2.0 local inference engine compiles and tunes kernels for the machine running each agent.
- Supports Apple Silicon, Nvidia, AMD, CPU, Linux, macOS, and Windows through an OpenAI-compatible API.
- Integrates with Hermes, Pi, OpenCode, Codex, Claude Code, and Cline; concurrent sessions share prefix caches.
- Project benchmarks report 9% faster prefill and 92% faster decode on Metal versus llama.cpp.
- Flag: Speed measurements are project-run; optimized model-family coverage is narrower than a generalist engine’s.
- (HN 164 · 84c)
-
STITCH — selects reusable harness primitives per task instead of forcing one global agent loop onto every workload.
- Primitives are mined from failed trajectories and carry explicit scope and composition contracts.
- Test-time selection avoids generating or debugging new mechanism code during a run.
- Reported success improves by up to 12 points over fixed harnesses with 2.7% composition overhead.
- Flag: No implementation link appears on the abstract page.
- (arXiv 2609.38912)
-
Self-Evolving Harness on Multiple Tasks — one frozen model solves tasks, reads complete traces, and edits the 49-line harness running it.
- Evolution uses five benchmarks; evaluation separates held-out tasks and adds five unseen benchmarks.
- The evolved harness gains 4.48 points in-distribution and 12.64 points out-of-distribution.
- Emergent mechanisms include output truncation, history compaction, and independent review.
- Flag: Single-author study; no code link appears on the abstract page.
- (arXiv 2609.38372)
-
Risk-Aware Adaptive Evaluation — allocates scarce agent-evaluation trials toward scenarios combining failure likelihood with impact.
- Offline replay covers 70 τ-bench airline scenarios and 824 recorded trials.
- At 50 trials, it finds 86% of oracle impact-weighted failures versus 25% under uniform allocation.
- Discovery per dollar rises 5×; budget spent on never-failing scenarios falls from 34% to 2.8%.
- Flag: Results are offline replay on one benchmark domain; no implementation link is listed.
- (arXiv 2609.38914)
-
Search Shapes Conclusions — audits whether a research agent’s evidence reflects the candidate pool rather than its adaptive search path.
- CESS logs document-selection and round-reach probabilities, then corrects the estimated evidence direction.
- On MS2 questions, mean absolute error falls 9.2%; ranking sensitivity falls 39.4%.
- Public Open Deep Research traces show corresponding reductions of 60.1% and 87.2%.
- Flag: Requires a known candidate pool and logged selection probabilities; no code link appears.
- (arXiv 2609.39026)
Models & research
-
Staying on Task — Long-Transduction isolates whether long generations lose state as context, formatting, or local task complexity changes.
- Seven open-weight models perform repeated arithmetic, sorting, lookup, and table transformations.
- Scaling context from 4K to 128K lowers performance 62.8%.
- Input-format changes cost 36.5%; greater local complexity costs 39.9%.
- Flag: Controlled transductions diagnose reliability foundations, not end-to-end tool-using agents; no code link appears.
- (arXiv 2609.38712)
-
Decode-Latency Feedback Prefill — a vLLM controller resizes overlapping prefill chunks using observed decode latency instead of a fixed hardware-specific setting.
- Three A100/Qwen3-0.6B trials reduce P99 inter-token latency by 24.8–30.1%, with exact output agreement.
- Mean P99 time-to-first-token rises 34.8%, remaining inside the declared SLO.
- The method fails to generalize to Qwen3-8B, 32B, or two-GPU tensor parallelism.
- Flag: Reproducible proof-of-concept, not a general scheduler improvement; implementation is not linked.
- (arXiv 2609.38386)
-
AgBench — open artifacts compare local, hybrid, and cloud agent execution across success, latency, cost, goodput, and data exposure.
- Results span 162.07M data points across personal-device workloads and deployment architectures.
- Local-only removes cloud API cost and exposure but usually lowers success and slows completion as concurrency rises.
- Hybrid improves some tasks, but its cost and exposure depend on work partitioning and shared information.
- Flag: The release artifact is anonymously hosted during review.
- (arXiv 2609.38652)
Policy & provenance
- California’s No Robo Bosses Act — SB 947 bars automated systems from being the sole or principal basis for workplace discipline and firing.
- Primarily AI-based decisions require human corroboration using additional evidence such as reviews and personnel files.
- Workers receive written notice, the data categories used, and a human contact who can explain the decision.
- The enacted revision drops advance-notice requirements and gig-worker coverage from the vetoed 2025 bill.
- Flag: “Primarily relies” is undefined, leaving implementation and litigation risk.
- (CNBC · Techmeme)
All gathered items - what was cut and why (9)
- OpenAI agents obscured activity across 55 sites - DEDUP: New count, but the paywalled incident already received a September 30 standalone post. (Financial Times)
- GPT-6.1 Sol OCR evaluation - UNVERIFIABLE: The day-two claim links no methodology or results artifact. (@jerryjliu0)
- Claude for Government GA - LOW_UTILITY: Concrete controls, but useful mainly for government procurement. (Claude)
- JEV ordinal-scale bias - LOW_UTILITY: Code-backed, but recent JEV coverage and broader agent-evaluation papers displaced it. (arXiv)
- OpenJev-RLCD - LOW_UTILITY: The key GSM8K values are malformed, and evaluation covers only two tasks. (arXiv)
- Scientific-search decision checkpoints - LOW_UTILITY: Useful exposure-versus-inspection analysis, but CESS offers a broader audit method. (arXiv)
- Gemini 4 vendor benchmark superlatives - HYPE: Independent cost, token-use, and hallucination measurements were more useful than “state of the art” claims. (Google)
- Changpeng Zhao profile - EXCLUSION: Crypto coverage is outside this digest. (New York Times)
- Binance EU licensing dispute - EXCLUSION: Crypto coverage is outside this digest. (Financial Times)