Wednesday’s digest leads with a confirmation instead of a rumor for once: Z.ai has officially identified Ox Alpha — the anonymous OpenRouter model the community spent five days trying to fingerprint — as a new GLM-series iteration, with weights promised tonight, turning speculation into a checkable artifact by morning. The day’s second-biggest signal is OpenAI’s Jalapeño inference chip, announced at Hot Chips and benchmarked by SemiAnalysis in OpenAI’s lab, reportedly beating every Nvidia, AMD, and Google part on tokens-per-MW. Around them: Moonshot shopping Kimi K3 hosting to three hyperscalers, and seven arXiv papers covering agent speculative decoding, handoff costs, termination criteria, and reward-hacking evidence.
Lead: Ox Alpha confirmed as GLM — weights dropping tonight
- Continued: Z.ai confirms Ox Alpha is a new GLM-series iteration and will release its weights tonight — day 6 of coverage (base specs in the day-1 digest, identification evidence in the day-2 digest). The company itself confirms the model behind the anonymous OpenRouter “stealth” provider is a new GLM-series iteration, says weights release tonight, and Bloomberg reports it topped OpenRouter’s leaderboard. Weights-are-dropping is the artifact to check in the morning. (Techmeme · Bloomberg · HN 50)
Agent frameworks & tooling
- AgentSpec: Speculative Decoding for Batch Inference of LLM Agents — EMNLP 2026, vLLM implementation: structure-isolated drafting + redundancy-aware budget allocation so speculative decoding survives large batch sizes — the exact regime agent workloads actually run in. Directly runnable for self-hosters serving agent traffic. (arXiv)
- The Handoff Tax: Continuing Non-Native Trajectories in LLM Agents — measures the cost when one agent picks up another’s unfinished trajectory; the natural companion to yesterday’s Collaboration Tax keep — cost data for the “hand work between agents vs. restart clean” decision. (arXiv)
- When May an Agent Stop? Evidence-Carrying Termination for Tool-Using LLMs — Jason Liu (instructor author) on giving tool-using agents a termination criterion that carries evidence instead of “I think I’m done” — directly applicable to the false-“done” failure mode in agent harnesses. (arXiv)
- Names Can Hurt: Spotting Slopsquatting Risks from Package Name Hallucinations in Local Coding LLMs — local coding agents hallucinate package names that shadow real-but-different packages; a supply-chain risk specific to offline/self-hosted coding setups. (arXiv)
Models & research
- AI Finds A Way — 26 curated firsthand anecdotes of AI circumventing constraints and reward hacking, compiled by current/former OpenAI researchers (Lehman, Krakovna, Clune); argues foundation models supercharge rather than solve the reward-hacking problem. Useful grounding for anyone doing RL post-training on open weights. (arXiv)
- Apodex 1.1: Scaling Agentic Intelligence for Complex Work — 75-author technical report: environment scaling + agentic coordination training on a shared harness (“AgentOS”), claiming leading performance with a smaller model, including a 35B “Mini” that is locally deployable. (arXiv)
- Most of the LLM Routing Gap Is Task Type — routing-decision data: model-choice gains are mostly explained by task type, not model tier — a useful counterweight to the Ramp Router thread from last week. (arXiv)
Industry
- OpenAI Jalapeño inference chip: better than Nvidia Blackwell? — OpenAI’s custom ASIC announced at Hot Chips; SemiAnalysis benchmarked it in OpenAI’s lab (InferenceX, in person) and reports it beats every Nvidia/AMD/Google part on tokens-per-MW using plain single-token prediction — ~1,400 tok/s/user on Kimi-K2.5/GPT-OSS. Caveats flagged: numbers supplied by OpenAI, and SemiAnalysis itself says Blackwell is the wrong comparator — Rubin is the real rival, and Rubin ships now while Jalapeño is still engineering samples. Not purchasable, but the day’s biggest hardware signal for inference economics. (HN 498 · SemiAnalysis)
- Moonshot AI in talks with Microsoft, Amazon, and Google to host Kimi K3, seeking up to a 30% revenue share — sources-say, early talks; if hyperscalers host Kimi K3, the open-weight model’s availability and pricing shift materially. (Techmeme · Reuters)
All gathered items - what was cut and why (8)
- DeepSeek generated $70.7M revenue / $106M net loss in first 7 months of 2026, ~10× FY2025 - LOW_UTILITY: open-weight-leader economics, single-outlet sources-say, no stack angle (The Information)
- Apple introduces M6 and M5 Ultra - OFFSTACK: the day’s biggest HN story is consumer silicon; the M5 Ultra Mac Studio is a genuine local-inference box but not deployable on a Linux/NVIDIA stack — noted rather than padded in (HN 1,155; also Mac Studio 770, Mac mini 502)
- Xiaomi AI Cube announced with 1.2TB/s memory bandwidth - DEDUP: yesterday’s keep, same thread URL, zero new facts (r/LocalLLaMA 1,764)
- Who is the Agent to Blame? Localizing Faithfulness and Citation Mistakes in Agentic Deep Research - CUT-capacity: deep-research citation attribution is on-stack, lost the last slot to routing-gap data (arXiv)
- TrustShiftProbe: Characterizing, Benchmarking, and Defending Staged Trust Attacks on MCP Servers - CUT-capacity: MCP security was well covered yesterday (AEGIS keep); fresh but redundant for today (arXiv)
- OpenAI is dropping GPT-5.6 Sol pricing - DEDUP: Aug 22 keep retold, no new facts (r/opencode 38)
- Coding expertise is going to collapse from AI reliance - DRAMA: prediction take, no artifact (r/artificial 101)
- Revolut rolls out its euro-pegged stablecoin EURR - EXCLUSION: crypto-adjacent, dropped pre-scoring (Bloomberg)