Wednesday’s digest leads with a confirmation instead of a rumor for once: Z.ai has officially identified Ox Alpha — the anonymous OpenRouter model the community spent five days trying to fingerprint — as a new GLM-series iteration, with weights promised tonight, turning speculation into a checkable artifact by morning. The day’s second-biggest signal is OpenAI’s Jalapeño inference chip, announced at Hot Chips and benchmarked by SemiAnalysis in OpenAI’s lab, reportedly beating every Nvidia, AMD, and Google part on tokens-per-MW. Around them: Moonshot shopping Kimi K3 hosting to three hyperscalers, and seven arXiv papers covering agent speculative decoding, handoff costs, termination criteria, and reward-hacking evidence.

Lead: Ox Alpha confirmed as GLM — weights dropping tonight

Agent frameworks & tooling

Models & research

  • AI Finds A Way — 26 curated firsthand anecdotes of AI circumventing constraints and reward hacking, compiled by current/former OpenAI researchers (Lehman, Krakovna, Clune); argues foundation models supercharge rather than solve the reward-hacking problem. Useful grounding for anyone doing RL post-training on open weights. (arXiv)
  • Apodex 1.1: Scaling Agentic Intelligence for Complex Work — 75-author technical report: environment scaling + agentic coordination training on a shared harness (“AgentOS”), claiming leading performance with a smaller model, including a 35B “Mini” that is locally deployable. (arXiv)
  • Most of the LLM Routing Gap Is Task Type — routing-decision data: model-choice gains are mostly explained by task type, not model tier — a useful counterweight to the Ramp Router thread from last week. (arXiv)

Industry

  • OpenAI Jalapeño inference chip: better than Nvidia Blackwell? — OpenAI’s custom ASIC announced at Hot Chips; SemiAnalysis benchmarked it in OpenAI’s lab (InferenceX, in person) and reports it beats every Nvidia/AMD/Google part on tokens-per-MW using plain single-token prediction — ~1,400 tok/s/user on Kimi-K2.5/GPT-OSS. Caveats flagged: numbers supplied by OpenAI, and SemiAnalysis itself says Blackwell is the wrong comparator — Rubin is the real rival, and Rubin ships now while Jalapeño is still engineering samples. Not purchasable, but the day’s biggest hardware signal for inference economics. (HN 498 · SemiAnalysis)
  • Moonshot AI in talks with Microsoft, Amazon, and Google to host Kimi K3, seeking up to a 30% revenue share — sources-say, early talks; if hyperscalers host Kimi K3, the open-weight model’s availability and pricing shift materially. (Techmeme · Reuters)
All gathered items - what was cut and why (8)