Wednesday brought a busy OpenAI DevDay, but the durable theme is operational: controlling persistent agents, recovering from bad tool guidance, forecasting token costs, and testing claims reproducibly. GPT-6.1 Sol leads because its sharply cheaper cached input could materially change the economics of long-running agents, though its benchmark and cost comparisons remain first-party. Dots adds persistent cloud agents with explicit approval rules and protected credentials. Elsewhere, Pi adopts MCP through a sandbox, two studies probe misleading tool outputs and unusable recovery instructions, SAGE gates self-editing skills against regressions, and TokenCast forecasts changing run costs. A preregistered Claude drift benchmark replaces “nerf” anecdotes with daily evidence, while new prompt-injection work tests attacks assembled from fragments. China’s model-hub competition closes the day with a reminder that open weights do not eliminate infrastructure dependencies.
OpenAI DevDay: GPT-6.1 Sol and Dots
-
GPT-6.1 Sol — OpenAI targets Astra-class agent work at Sol pricing, with cached input as the operationally important change.
- Measured: $2/M input · $10/M output · $0.10/M cached input; model ID:
gpt-6.1-sol. - OpenAI reports DeepSWE matching Astra and OSWorld within 2.1 points at lower per-task cost.
- Terminal-Bench Science costs $5.47/task at maximum effort versus $23.21 for Opus 5.5 and $23.80 for Astra.
- Available in ChatGPT Work, Codex, and API; an up-to-8× faster Ultrafast tier is promised.
- Flag: Capability and cost comparisons are OpenAI-run; no independent evaluation appeared in this pull.
- (HN 958 · 842c)
- Measured: $2/M input · $10/M output · $0.10/M cached input; model ID:
-
Dots — OpenAI launches persistent Astra agents with cloud computers, cross-channel context, app access, and background work.
- Each Dot has a browser and cloud computer; OpenAI claims integrations with 4,000+ apps.
- Background proactive research is read-only; Custom Rules can allow, require approval, or block actions.
- Saved credentials are injected without model access; sensitive actions such as password changes remain human-only.
- Rollout starts with Pro and Business Premium; Enterprise beta requires admin enablement.
- Flag: Launch examples and safety claims are first-party; usage allowances are not quantified.
- (HN 635 · 496c · Techmeme)
Agent frameworks & tooling
-
Pi adds MCP through Codemode — Pi reverses its anti-MCP stance and exposes deferred MCP tools inside a stateful JavaScript sandbox.
- Codemode composes calls without dumping every tool or intermediate result into model context.
- It runs beside the trusted harness loop, not in the lower-trust shell environment.
- Pi’s critique remains: many MCP servers return prose, lack composable schemas, and assume developer-side recovery.
- The interface can call classifier models such as Jev while retaining results in session state.
- (HN 67 · 25c)
-
MCP error messages can sabotage capable agents — Developer-oriented recovery text often tells tool-only agents to perform actions unavailable to them.
- Measured: 949 of 3,001 messages · 150 MCP servers · half of prescribed steps depend on hidden caller state.
- A terminal instruction recovered 45% of expired-credential tasks; GPT-6 Astra lost up to 69 points.
- Naming the correct tool raised credential recovery to 84% and rate-limit recovery to 88%.
- Removing unusable steps recovered 82%; code and data are released.
- (arXiv 2609.35381)
-
When Tools Silently Lie — ToxicBench separates whether agents check poisoned tool evidence from whether they ultimately adopt wrong answers.
- Its 118 tasks inject numerical, label, schema, and retrieval errors while holding source data fixed.
- Poisoning cuts success by 26–39 percentage points across three GPT adapters.
- Retries help one-shot faults; repeated poisoning still causes wrong-answer adoption after checking.
- Human review of 200 trajectories agrees 96% with the frozen scorer; code and trajectories are released.
- (arXiv 2609.37153)
-
SAGE — A paired statistical gate rejects self-evolving skill edits whose average gains hide regressions on solved items.
- Candidate and current skills run on identical validation items; losses receive asymmetric weight.
- A one-sided paired test commits only when wins are reliable, otherwise abstaining.
- Across five benchmarks and four models, regressions fall in 19 of 20 settings.
- DeepSeek-V4: LiveMath regressions 36.5% → 0% · OfficeQA 42.8% → 0%.
- Flag: No implementation link appears on the abstract page.
- (arXiv 2609.36043)
-
TokenCast — TokenCast forecasts an agent run’s bill as tool feedback and repeated context reads alter remaining cost.
- Segment records compose direct token use with context growth imposed on later calls.
- Measured: 14.5% average MAE reduction · 21.3% fewer replay tokens · 32.8 ms cumulative forecasting time.
- Results cover four suites, six agents, and 96 combinations; replay matches fixed-budget trace completion.
- Forecast updates add no LLM calls; code is released.
- (arXiv 2609.35760)
-
livenerf — A preregistered, append-only benchmark tests whether Opus 5.5 changes after launch instead of trusting anecdotes.
- The frozen 78-question panel uses a pinned Claude Code CLI and harness hash.
- Raw logs are retained for 30 daily runs.
- Low effort produced 62% fewer output tokens and 8.3 ± 4.5 fewer accuracy points in validation.
- The smaller validation run could not distinguish Opus 5 from 5.5.
- Flag: Only six baseline days exist; the first drift decision is expected around October 24, and no license is chosen.
- (HN 654 · 261c)
Models & research
- Divide and Inject — Fragmented indirect prompt injections can be reconstructed after retrieval, bypassing tests built around complete malicious instructions.
- AdaLCPI splits an objective across external content and supplies a reconstruction cue.
- OpenEvolve adapts fragments using graded scores and execution feedback from the target agent.
- Measured attack success: 61.4% macro · 32.8% Trojan Hippo-style · 30.0% AgentVigil.
- Injection evals should distribute incomplete fragments across long-context tool results.
- Flag: No code or benchmark link appears on the abstract page.
- (arXiv 2609.36576)
Industry
- China’s ModelScope and MoArk race to replace Hugging Face — Domestic hubs adapt open models to Chinese chips while developers still prefer the global ecosystem.
- ModelScope reports 170,000+ models and 250M users; MoArk hosts roughly 20,000 models.
- Hugging Face hosts more than 3 million models and is normally reached through VPNs in China.
- MoArk maintains engineers to port mainstream models across domestic accelerator stacks.
- Downloadable weights do not remove hub, toolchain, or hardware dependencies.
- (Rest of World · Techmeme)
All gathered items - what was cut and why (7)
- KV-streams - LOW_UTILITY: Credible 2.6–5× agentic-RL training speedup, but less immediately adoptable than the kept runtime artifacts. (arXiv)
- SafeCoEvo - LOW_UTILITY: Reports safer, more successful co-evolution, but provides no implementation. (arXiv)
- Voluntary White House AI accord - UNVERIFIABLE: Reuters blocked the article, and the accord text was absent from the pull. (Reuters)
- DeepSeek–Huawei toolchain partnership - UNVERIFIABLE: TileLang may be useful, but no repository was supplied and Reuters blocked verification. (Reuters)
- Language models for text classification - LOW_UTILITY: Useful Jev context, but it is an explainer and yesterday covered runnable compatible models. (Sebastian Raschka)
- Recent repeats and site-post overlap (no URL found) - DEDUP: Credentials API, Sonnet 5.5, and topics already covered on the site were not repackaged. (multiple sources)
- Source-noise tail (no URL found) - STALE / DRAMA / EXCLUSION: Recurring old social posts and a wagering/perpetual-futures item failed freshness or editorial rules. (Reddit · Bluesky)