Quiet Sunday — three items at the floor, with the GPT-6 Astra launch still the story into day three. The lead: Fortune’s archive snapshots document OpenAI quietly editing Astra’s published evaluation numbers after launch — hallucination rates changed and then reverted, and Sol’s ExploitBench score jumped to a level OpenAI says it may revert because it “reflects a reasoning level that is not commercially available” — a documented-edits story that is exactly why the day’s other Astra item matters: Robocurve’s independent robot-arm eval publishes every run, transcript and video. Around it, Seattle Times and Newsday sued OpenAI and Microsoft over training on their journalism.
Policy & provenance
- Continued: OpenAI quietly edited GPT-6 Astra’s published evaluation numbers after launch — day 3 of coverage (base specs in yesterday’s digest). What’s new, per Fortune’s archive-snapshot timeline of OpenAI’s Sep 3 launch post: the page was published then pulled for ~2 hours (OpenAI gave three different reasons, none benchmark-related), and across the snapshots numbers changed — Astra’s hallucination rate went 4.2% → 2% → back to 4.2%; Sol’s ExploitBench score jumped 5.5% → 11.5% (OpenAI now says it’s investigating reverting it because 11.5% “reflects a reasoning level that is not commercially available”); and Anthropic’s Fable 5.1 briefly lost ~10 points on FrontierMath (87.8% → 78% → 83% today) before settling. OpenAI’s statement: most evals have “noise within a few percentage points” and it “made fixes to ensure the numbers represent our best estimate.” Reported as documented edits — the archive.org evidence is the artifact. (Techmeme · Fortune)
Models & research
- Continued: GPT-6 Astra on robot arms — an independent eval that publishes every run — day 3 of coverage (base specs in yesterday’s digest). What’s new: Robocurve gave Astra control of the same YAM arms and Inspect Robots agent policy they used for their Fable comparison — 20 trials per model per task, human-graded on a 5-stage rubric, with every transcript, video, and raw run downloadable and the limitations stated plainly (bowl task ran on a different rig, trials not interleaved, grader knew which model it was). Results: Astra placed the block in the bowl 19/20 vs Fable 5.1’s 8/20, using ~⅙ the output tokens at roughly half the per-run cost; on the precision puzzle-piece insertion it tied Fable 5.1 at 2/20 — “it reaches the groove and stalls at the same final step Fable does.” Worth reading as an example of how to run an honest physical-agent eval more than as a benchmark claim. (HN 193)
Industry
- Seattle Times and Newsday sue Microsoft and OpenAI over training on their journalism — New publisher copyright suit (filed Sep 4, the same Friday the NYT-side case moved for summary judgment). The complaint alleges hundreds of thousands of articles were scraped past paywalls and terms of service, and cites ChatGPT reproducing an 88-word verbatim stretch of the Times’ Pulitzer-winning Boeing 737 MAX coverage from a headline-plus-URL prompt. The wrinkle: the Seattle Times is suing two of its own funders (Microsoft Philanthropies underwrites projects; both companies jointly funded the Lenfest AI fellowship the paper joined). Another concrete datapoint in the training-data litigation thread. (Techmeme · GeekWire)
All gathered items - what was cut and why (18)
- Formalizing Fermat’s Last Theorem (lobsters re-list of the Anthropic page) - DEDUP: day-2 re-list with nothing new since yesterday’s lead; full coverage already published as a standalone post (lobsters)
- Discovery of a new OpenAI agent message board (collusion.wiki) - DEDUP: same Sep 4 HN story whose day-2 beats were kept yesterday — score growth isn’t news (HN 2,190)
- LLMs as a Cognitive Virus - LOW_UTILITY: cultural/epidemiological framing paper about LLM diffusion — interesting read, no methods or stack angle (arXiv · HN 280)
- Cloud in a Bottle: making self-hosting accessible to everyone - OFFSTACK: genuinely popular self-hosting launch, but general-purpose hosting, not AI (HN 426)
- Warning: There is a new form of ai generated image - HYPE: anti-AI panic post with no verifiable artifact (Reddit 6,314)
- jerryjliu0: GPT-6 Astra on document parsing — “97.2%, new SOTA on our benchmark” - LOW_UTILITY: LlamaIndex benchmarking its own suite — real numbers, vendor-voice claim (X 157)
- DeepSeek to order 160,000 Huawei AI chips over Nvidia - STALE: re-list of the Bloomberg sources-say infra story already cut Sep 4 (Reddit 381)
- Terence Tao on “prematurely solving a maths problem by purely AI-powered methods” - UNVERIFIABLE: mathstodon thread blocked from extraction, and it’s commentary, not a development (lobsters)
- simonw: pelican-SVG reply - LOW_UTILITY: reply-level, no artifact (X 22)
- swyx: “we have crossed over into a new age of AI Engineering” - HYPE: Sep 3 launch take, still no published artifact (X 1,050)
- ggerganov: “Hugging Face has been acquired by NVIDIA” - DEDUP: restates the Nvidia–HF deal with no new facts (X 570)
- Bluesky 24/24 items stale (30th straight run) - STALE: source produced nothing fresh that clears the bar (no URL found) (Bluesky)
- Trump officials say Judeo-Christian principles inform AI policy - LOW_UTILITY: culture feature on religion and AI, no stack angle (Techmeme)
- Anthropomorphic portrayals of AI models as rogue agents obscure responsibility - DEDUP: framing already covered by the models-dont-go-rogue standalone post (Techmeme)
- Businesses in China package AI tokens as consumer rewards - LOW_UTILITY: consumer-marketing story, no stack angle (Techmeme)
- Sources: Travis Kalanick’s Atoms is developing robotaxi tech; Uber invested $100M - OFFSTACK: robotaxi business news, off the AI stack (Techmeme)
- The data center backlash is challenging Texas’ pro-business approach - OFFSTACK: data-center/energy news, off the AI stack (Techmeme)
- AI’s next ethical dilemma: should humans talk to animals? - OFFSTACK: bioethics/animal-communication research, off the AI stack (Techmeme)