Wednesday brought a busy OpenAI DevDay, but the durable theme is operational: controlling persistent agents, recovering from bad tool guidance, forecasting token costs, and testing claims reproducibly. GPT-6.1 Sol leads because its sharply cheaper cached input could materially change the economics of long-running agents, though its benchmark and cost comparisons remain first-party. Dots adds persistent cloud agents with explicit approval rules and protected credentials. Elsewhere, Pi adopts MCP through a sandbox, two studies probe misleading tool outputs and unusable recovery instructions, SAGE gates self-editing skills against regressions, and TokenCast forecasts changing run costs. A preregistered Claude drift benchmark replaces “nerf” anecdotes with daily evidence, while new prompt-injection work tests attacks assembled from fragments. China’s model-hub competition closes the day with a reminder that open weights do not eliminate infrastructure dependencies.

OpenAI DevDay: GPT-6.1 Sol and Dots

  • GPT-6.1 Sol — OpenAI targets Astra-class agent work at Sol pricing, with cached input as the operationally important change.

    • Measured: $2/M input · $10/M output · $0.10/M cached input; model ID: gpt-6.1-sol.
    • OpenAI reports DeepSWE matching Astra and OSWorld within 2.1 points at lower per-task cost.
    • Terminal-Bench Science costs $5.47/task at maximum effort versus $23.21 for Opus 5.5 and $23.80 for Astra.
    • Available in ChatGPT Work, Codex, and API; an up-to-8× faster Ultrafast tier is promised.
    • Flag: Capability and cost comparisons are OpenAI-run; no independent evaluation appeared in this pull.
    • (HN 958 · 842c)
  • Dots — OpenAI launches persistent Astra agents with cloud computers, cross-channel context, app access, and background work.

    • Each Dot has a browser and cloud computer; OpenAI claims integrations with 4,000+ apps.
    • Background proactive research is read-only; Custom Rules can allow, require approval, or block actions.
    • Saved credentials are injected without model access; sensitive actions such as password changes remain human-only.
    • Rollout starts with Pro and Business Premium; Enterprise beta requires admin enablement.
    • Flag: Launch examples and safety claims are first-party; usage allowances are not quantified.
    • (HN 635 · 496c · Techmeme)

Agent frameworks & tooling

  • Pi adds MCP through Codemode — Pi reverses its anti-MCP stance and exposes deferred MCP tools inside a stateful JavaScript sandbox.

    • Codemode composes calls without dumping every tool or intermediate result into model context.
    • It runs beside the trusted harness loop, not in the lower-trust shell environment.
    • Pi’s critique remains: many MCP servers return prose, lack composable schemas, and assume developer-side recovery.
    • The interface can call classifier models such as Jev while retaining results in session state.
    • (HN 67 · 25c)
  • MCP error messages can sabotage capable agents — Developer-oriented recovery text often tells tool-only agents to perform actions unavailable to them.

    • Measured: 949 of 3,001 messages · 150 MCP servers · half of prescribed steps depend on hidden caller state.
    • A terminal instruction recovered 45% of expired-credential tasks; GPT-6 Astra lost up to 69 points.
    • Naming the correct tool raised credential recovery to 84% and rate-limit recovery to 88%.
    • Removing unusable steps recovered 82%; code and data are released.
    • (arXiv 2609.35381)
  • When Tools Silently Lie — ToxicBench separates whether agents check poisoned tool evidence from whether they ultimately adopt wrong answers.

    • Its 118 tasks inject numerical, label, schema, and retrieval errors while holding source data fixed.
    • Poisoning cuts success by 26–39 percentage points across three GPT adapters.
    • Retries help one-shot faults; repeated poisoning still causes wrong-answer adoption after checking.
    • Human review of 200 trajectories agrees 96% with the frozen scorer; code and trajectories are released.
    • (arXiv 2609.37153)
  • SAGE — A paired statistical gate rejects self-evolving skill edits whose average gains hide regressions on solved items.

    • Candidate and current skills run on identical validation items; losses receive asymmetric weight.
    • A one-sided paired test commits only when wins are reliable, otherwise abstaining.
    • Across five benchmarks and four models, regressions fall in 19 of 20 settings.
    • DeepSeek-V4: LiveMath regressions 36.5% → 0% · OfficeQA 42.8% → 0%.
    • Flag: No implementation link appears on the abstract page.
    • (arXiv 2609.36043)
  • TokenCast — TokenCast forecasts an agent run’s bill as tool feedback and repeated context reads alter remaining cost.

    • Segment records compose direct token use with context growth imposed on later calls.
    • Measured: 14.5% average MAE reduction · 21.3% fewer replay tokens · 32.8 ms cumulative forecasting time.
    • Results cover four suites, six agents, and 96 combinations; replay matches fixed-budget trace completion.
    • Forecast updates add no LLM calls; code is released.
    • (arXiv 2609.35760)
  • livenerf — A preregistered, append-only benchmark tests whether Opus 5.5 changes after launch instead of trusting anecdotes.

    • The frozen 78-question panel uses a pinned Claude Code CLI and harness hash.
    • Raw logs are retained for 30 daily runs.
    • Low effort produced 62% fewer output tokens and 8.3 ± 4.5 fewer accuracy points in validation.
    • The smaller validation run could not distinguish Opus 5 from 5.5.
    • Flag: Only six baseline days exist; the first drift decision is expected around October 24, and no license is chosen.
    • (HN 654 · 261c)

Models & research

  • Divide and Inject — Fragmented indirect prompt injections can be reconstructed after retrieval, bypassing tests built around complete malicious instructions.
    • AdaLCPI splits an objective across external content and supplies a reconstruction cue.
    • OpenEvolve adapts fragments using graded scores and execution feedback from the target agent.
    • Measured attack success: 61.4% macro · 32.8% Trojan Hippo-style · 30.0% AgentVigil.
    • Injection evals should distribute incomplete fragments across long-context tool results.
    • Flag: No code or benchmark link appears on the abstract page.
    • (arXiv 2609.36576)

Industry

  • China’s ModelScope and MoArk race to replace Hugging Face — Domestic hubs adapt open models to Chinese chips while developers still prefer the global ecosystem.
    • ModelScope reports 170,000+ models and 250M users; MoArk hosts roughly 20,000 models.
    • Hugging Face hosts more than 3 million models and is normally reached through VPNs in China.
    • MoArk maintains engineers to port mainstream models across domestic accelerator stacks.
    • Downloadable weights do not remove hub, toolchain, or hardware dependencies.
    • (Rest of World · Techmeme)
All gathered items - what was cut and why (7)
  • KV-streams - LOW_UTILITY: Credible 2.6–5× agentic-RL training speedup, but less immediately adoptable than the kept runtime artifacts. (arXiv)
  • SafeCoEvo - LOW_UTILITY: Reports safer, more successful co-evolution, but provides no implementation. (arXiv)
  • Voluntary White House AI accord - UNVERIFIABLE: Reuters blocked the article, and the accord text was absent from the pull. (Reuters)
  • DeepSeek–Huawei toolchain partnership - UNVERIFIABLE: TileLang may be useful, but no repository was supplied and Reuters blocked verification. (Reuters)
  • Language models for text classification - LOW_UTILITY: Useful Jev context, but it is an explainer and yesterday covered runnable compatible models. (Sebastian Raschka)
  • Recent repeats and site-post overlap (no URL found) - DEDUP: Credentials API, Sonnet 5.5, and topics already covered on the site were not repackaged. (multiple sources)
  • Source-noise tail (no URL found) - STALE / DRAMA / EXCLUSION: Recurring old social posts and a wagering/perpetual-futures item failed freshness or editorial rules. (Reddit · Bluesky)