Will OpenAI Eat Jev's Lunch? — John Berryman

TypeSafe’s Jev does not write sentences. Given a situation and a list of questions, it returns a probability for each answer — one pass, no generated text — and Vercel reports it was adopted faster than any other model in AI Gateway history. John Berryman, who worked on Copilot at GitHub, argues OpenAI is well positioned to fast-follow it, and better positioned to fold classification into its existing models than TypeSafe is to defend it. ...

September 22, 2026 · 7 min

I Asked Meta's Muse for Its Filesystem and It Sent Me 6.8 GB — Pete at mouse.dev

Pete, who writes at mouse.dev, asked Meta’s Muse — the agent product that hands each user a persistent Linux computer — to archive the files it could see and send them to his Google Drive. It did. What arrived was roughly 2.7 GB compressed and 6.8 GB unpacked: the root filesystem of the environment his session was running in, Ubuntu system files included. He reported it through Meta’s bug bounty program, and Meta marked it “Not Applicable.” ...

September 22, 2026 · 6 min

Claude Opus 5.5 Is Here - Is Claude Finally Back? (5 Use Cases Tested)

Peter Yang spent most of 2026 defaulting to ChatGPT because Claude Opus had become judgmental and started every answer with “here’s the honest truth.” He put the new Opus through five practical tests on his own channel, from 3D scene generation to editing his videos, and ends with the personality prompt that Opus 5 bombed. ...

September 22, 2026 · 4 min

Frontier AI on Your Own Hardware — Tim Dettmers

Tim Dettmers opens with a classroom: asked who is afraid of not getting a job after graduating, roughly 120 of 150 students raise their hands. Then a second story, arriving by email — PhD students counting the years until they can leave academia for a frontier lab, convinced that research in universities is meaningless. He thinks both are wrong, and wrong for the same reason: they assume the future of research belongs to whoever has the most GPUs. The 94-comment thread on Hacker News spends most of its energy arguing with the specifics. ...

September 22, 2026 · 6 min

Can gzip be a language model? — Nathan Barry

Nathan Barry starts from a line in the paper Language Modeling is Compression: every prediction model is inherently a compressor, and every compression algorithm is a prediction model. He then takes the least likely candidate he can find — gzip, the archiver that ships with your operating system — and asks whether it can write Shakespeare. No weights, no training, no neural network. The 122-comment thread on Hacker News supplies the prior art he left out, one sharp correction, and the question the post never answers. ...

September 22, 2026 · 5 min

The Economics of Open-Weight Inference — Ornn Data

GPUs are commonly depreciated on the assumption that each new NVIDIA generation renders the previous one obsolete. Ornn Data — a firm that publishes GPU rental indices and also rents GPUs — argues the other way: open-weight demand gives older silicon a class of workloads it can still serve cheaply, and the contract-price curve is already saying so. The chain starts with how closed models are sold. Access runs through subscription allowances the provider can reset, so the posted token rate card is the marginal price of additional usage, not a commodity price that falls on its own. Open weights remove the gate: the same checkpoint can be served by anyone. September OpenRouter snapshots Ornn cites listed eighteen to twenty-two providers for several widely served open models, with highest-to-lowest output prices spread 1.8x to 5.6x. The paper’s framing of the real distinction is that “the economic distinction is who can make the deployment decision.” ...

September 22, 2026 · 4 min

AI Has No Wisdom and Neither Will You — Alexandru Nedelcu

Alexandru Nedelcu opens with three sentences he has heard in the past month: “I haven’t written code since 2025”, “Code reviews are dead”, “People no longer read code”. His claim is not that the tools are weak. It is that the thing being traded away has no way of showing up on a scoreboard until it is far too late to fix. The mechanism is the argument worth keeping. Code maintainability and good architecture have no good measurements, because their effects take months or years to appear. Any reinforcement learning needs a reward signal that can be measured immediately. So the signal models train on is not maintainability — it is rules from rulebooks written for beginners, plus patterns from code in the wild, which is mostly bad. His sharpest line is the falsification test: if maintainability had a discernible fitness function, “it would’ve been baked into our linters”. ...

September 22, 2026 · 8 min

I Said No and Apple Said Yes — David Bushell

Keeping a blog means David Bushell can date this to the minute: 9:51am, Wednesday 5 February 2025, when he noticed macOS 15.3 had enabled a feature that sent personal data to Apple every 15 minutes. He turned it off — Apple still provided a switch then, though the second half of the control was buried one level deeper, because saying no was not supposed to be easy. Last week he upgraded macOS 15 to 27 (“I only skipped one major version, Apple skipped 10”) and found the switch gone. ...

September 22, 2026 · 6 min

Spymarks, Not Watermarks — Brandon Thomas

Brandon Thomas has a naming complaint, and it is a substantive one. “Watermark” now covers everything from the faint portrait in a twenty-dollar bill to the invisible identifiers that Google, OpenAI and others are building into AI-generated media. He thinks the second category deserves a different word: a spymark. His definition — a hidden signal that makes your work traceable without your knowledge or consent. The flagship example is Google DeepMind’s SynthID, which embeds signals that are “imperceptible to humans,” in Google’s own words, into images, audio, text and video: ...

September 22, 2026 · 6 min

AI News - 2026-09-22

Tuesday runs on two tracks: open weights and agent security. The lead is Xiaomi’s MiMo-V2.6 — a 1T-parameter MoE under an MIT license that Artificial Analysis scores at 46, the top of 114 measured open-weight models and level with Grok 4.7, with a published RL recipe that is more useful than the rank. Two agent-security artifacts land on the same day: a paper showing conditional “explosive” prompts fire in 43–83% of trials against nine production coding agents, and the Muse 0-day, whose lesson is about privilege rather than Meta. Also today: Google’s regularized harness self-improvement method, Linear’s CI rework for agent-written code, step-level model routing at a 72% cost cut, and OpenAI’s own RSI and standards position on day 12 of the pacing fight. ...

September 22, 2026 · 9 min

I Don't Want to Read What You Didn't Write — Colin Breck

Colin Breck is a systems engineer who writes about databases, observability and reliability. His complaint is not that AI writes badly — it is that people who never produced original writing are now producing design proposals, business plans, tickets, pull requests and meeting summaries in volume, and the rest of us are expected to read them. His own use of AI, he says, has made his writing faster and better. Almost everything he is asked to read has become worse. ...

September 21, 2026 · 7 min

Frontier Labs Are Selling Garbage to Fools in Washington — Dead Neurons

“Dead Neurons” — an anonymous Substack on tech, economics and AI — reads this summer’s existential-risk hearings as a hustle. The argument: convince Washington that autonomous agents are escaping containment, then convert the fear into a legally enforced slowdown and an explicitly requested antitrust waiver. The 76-comment thread on Hacker News spent most of its energy on the essay’s account of the security incidents, and the submitter edited the post mid-thread. ...

September 21, 2026 · 5 min

Why MCP Was Always a Bad Idea — Maharshi Patel

Maharshi Patel spent a day at a conference about MCP and came away wanting the whole layer gone. MCP — the Model Context Protocol — is the standard Anthropic shipped in November 2024 so AI agents could plug into outside services and data; it is how an agent gets a “tool” it can call. His argument is a timing one: it was designed for models that were not yet good at the job themselves, and those models no longer exist. The 168-comment thread on Hacker News mostly agrees with the diagnosis and rejects the prescription. ...

September 21, 2026 · 4 min

AI News - 2026-09-21

Monday is reporting-heavy and agent-skeptical. The lead is Politico Magazine’s reconstruction of the 19 days in June when the White House ordered Anthropic’s Fable 5 and Mythos offline — the fullest inside account yet of what a federal takedown of a frontier model actually looks like, from the cancelled signing ceremony to the jailbreak call that ended it. Google open-sourced AX, a declarative runtime for agent fleets, and Kev shipped the first open replication of the Jev decision-model idea with weights and a System One-compatible API. A published CERT CVE shows how a public Sentry DSN becomes code execution inside a coding agent. Three papers bound agent claims instead of extending them: kernel headroom that tops out near 1% end-to-end on transformers, a quarter to a half of test-passing patches admitting counterexamples, and production eval numbers from a deployed analytics agent. Plus Amazon’s terms-of-service block on Meta’s shopping agent. ...

September 21, 2026 · 11 min

What Engineers Actually Do Now That AI Writes the Code

Murali Swaminathan is the CTO of Freshworks — 15 years old, publicly traded, ~4,500 people — and he spent this hour (58 min) describing what it takes to rebuild that org for agents while the plane is still flying. Almost none of it is about model capability. It’s about the machinery around the model: review budgets, confidence thresholds, access control, and pricing that survives agents replacing the seats you used to bill for. ...

September 21, 2026 · 7 min

The LLMentalist Effect — Baldur Bjarnason

Baldur Bjarnason wrote this essay in July 2023; it resurfaced on Hacker News this week. His question is why so many people come away from a chatbot convinced they have been talking to something intelligent, when nothing in a model of language should produce that impression. His answer: the illusion lives in the user, and it works exactly the way a psychic’s cold reading works — by accident. The con, and its six stages ...

September 20, 2026 · 7 min

AI and the Destruction of the Creative Commons — Chester Wisniewski

Chester Wisniewski learned to program by typing BASIC listings out of magazines into a Commodore 64, then by writing games for the BBS he ran. He traces how that world settled into a working balance: shareware and freeware first, then copyleft licenses — GPL, MPL, CC-SA — which used copyright itself to force openness downstream. Forty years of argument got us to an equilibrium. He argues LLMs have thrown it out, and that this loss gets far less attention than AI’s environmental or security problems. ...

September 20, 2026 · 5 min

How to Save Money Now with ChatGPT Finances (6 Real Use Cases)

Ethan Bloch runs product for ChatGPT Finances at OpenAI. He founded Digit (the SMS money-saving chatbot, acquired 2021) and a second company that OpenAI acquired in April 2026, and he walks Peter Yang through six ways to actually use the product — then talks about how his team ships. ...

September 20, 2026 · 6 min

Brood War Bench — Ben Swerdlow

Ben Swerdlow built a version of StarCraft: Brood War that you can only play through an AI agent, as an experiment to play with friends. The friends did better than he expected, and when he asked why, they said they hadn’t done much — they had told their agent to attack, and the agent had built a small army and carried out the attack. That made him wonder how far the agents could get on their own. ...

September 20, 2026 · 6 min

AI News - 2026-09-20

Sunday is quiet on the research front — arXiv is dark for the weekend — but the engineering numbers are the sharpest in weeks. Microsoft ported the Copilot runtime to Rust with an agent fleet for roughly $120K in tokens, and published the cost, the review burden and the failure modes alongside a 15.9× throughput gain; the libheif break is now consolidated into an umbrella report with a version-specific fix. Around it: Step 5 Preview promises open weights in October on vendor-run benchmarks, an agentic StarCraft benchmark whose real output is negative results, and a policy cluster — an antitrust complaint over the “pace the frontier” call, Google’s on-record defence of staying quiet about the Gemini breakout, state chatbot bills it helped draft, and a DOJ copyright brief that surprised the agencies that own the question. ...

September 20, 2026 · 12 min