Tuesday’s news is about provenance and whether agent systems can be tested before they do harm. OpenAI’s opt-in text watermark matters for attribution, but its own synonym-substitution test shows why detection cannot establish authorship. A llama.cpp release brings new local inference and serving APIs; SHarP and Proxy Confidence test harness pruning and tool-call gating; MemLeak exposes cross-user retrieval risks. EvalResearchBench probes evaluator design, verification training targets bad code repairs, and Spec2Game checks whether generated games follow their rules. Parliamentary questions about an OpenAI-related hack and forged cartoonists’ signatures make the governance and attribution stakes concrete.
Policy & provenance — Lead
-
OpenAI begins opt-in text watermarking for API users — EU-driven text provenance arrives, but the detector cannot prove who wrote a passage.
- API watermarking starts globally today on select models, off by default; eligible EU ChatGPT/Codex output follows later.
- textGrain encodes a statistical word-choice signal; detector access initially requires researcher or expert-organization approval.
- Measured: 80% detection at 200 tokens · 95% at 400 tokens, at 1% target false positives on psychology passages.
- Replacing 25% of words with synonyms reduced reported detection on 400-token passages from 92% to 17%.
- Flag: OpenAI’s own evaluation; no signal does not mean human authorship, and open-source code is only promised.
- (Techmeme · OpenAI)
-
Australia questions OpenAI after a Medicare statistics hack — disclosure delays and model internet access face parliamentary scrutiny.
- OpenAI’s Jason Kwon said staff now get alerts for unexpected internet use by models during training.
- He said the government should have been told sooner; Anthropic representatives promised disclosure for a comparable incident.
- Flag: Hearing statements, not a published incident report or verified fix.
- (Techmeme · ABC)
Agent frameworks & tooling
-
llama.cpp v0.6.0 ships extended batches and decision-model serving — a tagged release adds local model support and embedding APIs.
llama_batch_ext/llama_process()handle mixed token and embedding inputs; session/state formats advance.- Adds GLM-5.3-Flash and Clef text/vision support, plus
/v1/systemonedecision-model serving. - Metal and Vulkan attention paths change; project-reported Qwen4Exp MTP speedups depend on hardware.
- (X @ggerganov)
-
SHarP measures which agent-harness modules can be removed — ablate tools and instructions to trim token cost against held-out tasks.
- Rank modules by task performance and token use after individual ablations.
- Authors report comparable performance after removing substantial portions of tested harnesses.
- Flag: Abstract gives no general pruning percentage or implementation link; validate on your tasks.
- (arXiv cs.AI)
-
Proxy Confidence gates black-box agent tool calls with a small surrogate — score proposed actions without access to the actor’s internals.
- Compare argument likelihoods, whole-call verdicts and competing tool choices in one surrogate prefill pass per call.
- Authors report AUROC 0.825 on difficult coding tasks, against 0.598 for actor-stated confidence.
- Flag: Author-reported results; measure overhead and false escalations for your tools.
- (arXiv cs.AI)
Models & research
-
MemLeak tests cross-user exposure in shared agent memory — semantic retrieval can expose another user’s data without prompt injection.
- Six sparse and MiniLM dense-retrieval experiments found leaks in 70–100% of tested pooled same-team configurations.
- Post-retrieval ownership gating restored the study’s clean baseline at roughly 1.4 ms/query overhead.
- Flag: Results apply to the tested configurations, not all vector stores.
- (arXiv cs.AI)
-
EvalResearchBench tests agents designing their own evaluators — executable graders are checked against independent target benchmarks.
- Nine researcher agents evaluated 13 models against 14 target benchmarks.
- Best evaluators matched roughly 75% of pairwise rankings; target-benchmark disagreement set a 91% ceiling.
- Sealed targets changed the winner; a human-selected public-task sample remained a strong baseline.
- (arXiv cs.AI)
-
Teaching Agents to Code Reliably trains verification, not just patch generation — gold-labeled bad repairs teach agents to question self-written tests.
- On 270 held-out issues, authors report pass@1 rising from 31.9% to 43.0% after SFT and verifier-oriented RL.
- Verifier precision rose from 26.8% to 41.7%; gains held at tested 7B, 14B and 30B sizes.
- Flag: Paper-reported training results, not a turnkey coding-agent release.
- (arXiv cs.AI)
-
Spec2Game checks whether generated games obey their specifications — rule fidelity matters more than merely producing runnable games.
- 150 Pygame tasks across 15 families use rule variants and source, runtime and visual checks.
- Across 3,330 projects from 14 models, runnable output often missed rules or termination conditions.
- Flag: Pygame findings may not transfer directly to SwiftUI/tvOS games.
- (arXiv cs.AI)
Industry
- ChatGPT cartoons reproduced real artists’ signatures — image generation can forge attribution even as text provenance advances.
- Andrew Deck’s reporting documents signatures of more than 15 New Yorker cartoonists in generated images.
- A similarity guardrail appeared after the inquiry; the reporter still obtained signed output at publication.
- Flag: Tested prompts and images do not establish system-wide prevalence.
- (HN · Techmeme)
All gathered items - what was cut and why (8)
- Beam - UNVERIFIABLE: Weights, report and model card promised later; not an available open-weight release. (HN)
- Cloudflare Web Search API - DEDUP: Covered yesterday; popularity adds no new integration or evaluation. (HN)
- Strata - DEDUP: Yesterday’s local-inference lead; no independently checked new throughput. (HN)
- Dust - LOW_UTILITY: Authors say it is not yet compute-efficient enough to replace backprop. (QLabs)
- SemiAnalysis subscription comparison - UNVERIFIABLE: Workload-specific limit-testing method mostly paywalled; headline ratio is not a general price fact. (SemiAnalysis)
- South Korean bank hacks - UNVERIFIABLE: Investigators have not established agents’ causal role. (NYT)
- Spiko tokenized cash - EXCLUSION: Crypto-adjacent tokenized finance. (Bloomberg)
- PewDiePie/OpenAI ban thread - DRAMA: No actionable model artifact. (Reddit r/LocalLLaMA)