Quiet Sunday — three items at the floor, with the GPT-6 Astra launch still the story into day three. The lead: Fortune’s archive snapshots document OpenAI quietly editing Astra’s published evaluation numbers after launch — hallucination rates changed and then reverted, and Sol’s ExploitBench score jumped to a level OpenAI says it may revert because it “reflects a reasoning level that is not commercially available” — a documented-edits story that is exactly why the day’s other Astra item matters: Robocurve’s independent robot-arm eval publishes every run, transcript and video. Around it, Seattle Times and Newsday sued OpenAI and Microsoft over training on their journalism.

Policy & provenance

  • Continued: OpenAI quietly edited GPT-6 Astra’s published evaluation numbers after launch — day 3 of coverage (base specs in yesterday’s digest). What’s new, per Fortune’s archive-snapshot timeline of OpenAI’s Sep 3 launch post: the page was published then pulled for ~2 hours (OpenAI gave three different reasons, none benchmark-related), and across the snapshots numbers changed — Astra’s hallucination rate went 4.2% → 2% → back to 4.2%; Sol’s ExploitBench score jumped 5.5% → 11.5% (OpenAI now says it’s investigating reverting it because 11.5% “reflects a reasoning level that is not commercially available”); and Anthropic’s Fable 5.1 briefly lost ~10 points on FrontierMath (87.8% → 78% → 83% today) before settling. OpenAI’s statement: most evals have “noise within a few percentage points” and it “made fixes to ensure the numbers represent our best estimate.” Reported as documented edits — the archive.org evidence is the artifact. (Techmeme · Fortune)

Models & research

  • Continued: GPT-6 Astra on robot arms — an independent eval that publishes every run — day 3 of coverage (base specs in yesterday’s digest). What’s new: Robocurve gave Astra control of the same YAM arms and Inspect Robots agent policy they used for their Fable comparison — 20 trials per model per task, human-graded on a 5-stage rubric, with every transcript, video, and raw run downloadable and the limitations stated plainly (bowl task ran on a different rig, trials not interleaved, grader knew which model it was). Results: Astra placed the block in the bowl 19/20 vs Fable 5.1’s 8/20, using ~⅙ the output tokens at roughly half the per-run cost; on the precision puzzle-piece insertion it tied Fable 5.1 at 2/20 — “it reaches the groove and stalls at the same final step Fable does.” Worth reading as an example of how to run an honest physical-agent eval more than as a benchmark claim. (HN 193)

Industry

  • Seattle Times and Newsday sue Microsoft and OpenAI over training on their journalism — New publisher copyright suit (filed Sep 4, the same Friday the NYT-side case moved for summary judgment). The complaint alleges hundreds of thousands of articles were scraped past paywalls and terms of service, and cites ChatGPT reproducing an 88-word verbatim stretch of the Times’ Pulitzer-winning Boeing 737 MAX coverage from a headline-plus-URL prompt. The wrinkle: the Seattle Times is suing two of its own funders (Microsoft Philanthropies underwrites projects; both companies jointly funded the Lenfest AI fellowship the paper joined). Another concrete datapoint in the training-data litigation thread. (Techmeme · GeekWire)
All gathered items - what was cut and why (18)