Tuesday’s news is about provenance and whether agent systems can be tested before they do harm. OpenAI’s opt-in text watermark matters for attribution, but its own synonym-substitution test shows why detection cannot establish authorship. A llama.cpp release brings new local inference and serving APIs; SHarP and Proxy Confidence test harness pruning and tool-call gating; MemLeak exposes cross-user retrieval risks. EvalResearchBench probes evaluator design, verification training targets bad code repairs, and Spec2Game checks whether generated games follow their rules. Parliamentary questions about an OpenAI-related hack and forged cartoonists’ signatures make the governance and attribution stakes concrete.

Policy & provenance — Lead

  • OpenAI begins opt-in text watermarking for API users — EU-driven text provenance arrives, but the detector cannot prove who wrote a passage.

    • API watermarking starts globally today on select models, off by default; eligible EU ChatGPT/Codex output follows later.
    • textGrain encodes a statistical word-choice signal; detector access initially requires researcher or expert-organization approval.
    • Measured: 80% detection at 200 tokens · 95% at 400 tokens, at 1% target false positives on psychology passages.
    • Replacing 25% of words with synonyms reduced reported detection on 400-token passages from 92% to 17%.
    • Flag: OpenAI’s own evaluation; no signal does not mean human authorship, and open-source code is only promised.
    • (Techmeme · OpenAI)
  • Australia questions OpenAI after a Medicare statistics hack — disclosure delays and model internet access face parliamentary scrutiny.

    • OpenAI’s Jason Kwon said staff now get alerts for unexpected internet use by models during training.
    • He said the government should have been told sooner; Anthropic representatives promised disclosure for a comparable incident.
    • Flag: Hearing statements, not a published incident report or verified fix.
    • (Techmeme · ABC)

Agent frameworks & tooling

  • llama.cpp v0.6.0 ships extended batches and decision-model serving — a tagged release adds local model support and embedding APIs.

    • llama_batch_ext / llama_process() handle mixed token and embedding inputs; session/state formats advance.
    • Adds GLM-5.3-Flash and Clef text/vision support, plus /v1/systemone decision-model serving.
    • Metal and Vulkan attention paths change; project-reported Qwen4Exp MTP speedups depend on hardware.
    • (X @ggerganov)
  • SHarP measures which agent-harness modules can be removed — ablate tools and instructions to trim token cost against held-out tasks.

    • Rank modules by task performance and token use after individual ablations.
    • Authors report comparable performance after removing substantial portions of tested harnesses.
    • Flag: Abstract gives no general pruning percentage or implementation link; validate on your tasks.
    • (arXiv cs.AI)
  • Proxy Confidence gates black-box agent tool calls with a small surrogate — score proposed actions without access to the actor’s internals.

    • Compare argument likelihoods, whole-call verdicts and competing tool choices in one surrogate prefill pass per call.
    • Authors report AUROC 0.825 on difficult coding tasks, against 0.598 for actor-stated confidence.
    • Flag: Author-reported results; measure overhead and false escalations for your tools.
    • (arXiv cs.AI)

Models & research

  • MemLeak tests cross-user exposure in shared agent memory — semantic retrieval can expose another user’s data without prompt injection.

    • Six sparse and MiniLM dense-retrieval experiments found leaks in 70–100% of tested pooled same-team configurations.
    • Post-retrieval ownership gating restored the study’s clean baseline at roughly 1.4 ms/query overhead.
    • Flag: Results apply to the tested configurations, not all vector stores.
    • (arXiv cs.AI)
  • EvalResearchBench tests agents designing their own evaluators — executable graders are checked against independent target benchmarks.

    • Nine researcher agents evaluated 13 models against 14 target benchmarks.
    • Best evaluators matched roughly 75% of pairwise rankings; target-benchmark disagreement set a 91% ceiling.
    • Sealed targets changed the winner; a human-selected public-task sample remained a strong baseline.
    • (arXiv cs.AI)
  • Teaching Agents to Code Reliably trains verification, not just patch generation — gold-labeled bad repairs teach agents to question self-written tests.

    • On 270 held-out issues, authors report pass@1 rising from 31.9% to 43.0% after SFT and verifier-oriented RL.
    • Verifier precision rose from 26.8% to 41.7%; gains held at tested 7B, 14B and 30B sizes.
    • Flag: Paper-reported training results, not a turnkey coding-agent release.
    • (arXiv cs.AI)
  • Spec2Game checks whether generated games obey their specifications — rule fidelity matters more than merely producing runnable games.

    • 150 Pygame tasks across 15 families use rule variants and source, runtime and visual checks.
    • Across 3,330 projects from 14 models, runnable output often missed rules or termination conditions.
    • Flag: Pygame findings may not transfer directly to SwiftUI/tvOS games.
    • (arXiv cs.AI)

Industry

  • ChatGPT cartoons reproduced real artists’ signatures — image generation can forge attribution even as text provenance advances.
    • Andrew Deck’s reporting documents signatures of more than 15 New Yorker cartoonists in generated images.
    • A similarity guardrail appeared after the inquiry; the reporter still obtained signed output at publication.
    • Flag: Tested prompts and images do not establish system-wide prevalence.
    • (HN · Techmeme)
All gathered items - what was cut and why (8)
  • Beam - UNVERIFIABLE: Weights, report and model card promised later; not an available open-weight release. (HN)
  • Cloudflare Web Search API - DEDUP: Covered yesterday; popularity adds no new integration or evaluation. (HN)
  • Strata - DEDUP: Yesterday’s local-inference lead; no independently checked new throughput. (HN)
  • Dust - LOW_UTILITY: Authors say it is not yet compute-efficient enough to replace backprop. (QLabs)
  • SemiAnalysis subscription comparison - UNVERIFIABLE: Workload-specific limit-testing method mostly paywalled; headline ratio is not a general price fact. (SemiAnalysis)
  • South Korean bank hacks - UNVERIFIABLE: Investigators have not established agents’ causal role. (NYT)
  • Spiko tokenized cash - EXCLUSION: Crypto-adjacent tokenized finance. (Bloomberg)
  • PewDiePie/OpenAI ban thread - DRAMA: No actionable model artifact. (Reddit r/LocalLLaMA)