Piotr Sarnacki examines AI-generated rewrites of Campfire, a chat application, and challenges the idea that experienced programmers can safely stop reading generated code. His point is not that one programming language wins: agents make unstated design decisions, and performance tests can reward the wrong behavior.

  • A rewrite can quietly become a different product. Sarnacki reports differences in compatibility and notification queues between versions. A vague request to translate the code leaves reliability requirements up to the agent.
  • Request speed is only one measurement. His tests examine whether chat updates actually reach users, how long they take, and whether requests fail—not just how many requests finish.
  • More delivered updates can mean a worse experience. In his Rust test, enlarging the notification buffer raised delivery from roughly 14% to 90%, but pushed the worst delivery delay from 11 seconds to over 130 seconds. He favors disconnecting a lagging client so it can reconnect and fetch current messages instead.

The useful lesson for AI-assisted engineering: specify acceptable failure behavior before asking for a rewrite. Code that runs, and a speed comparison that looks impressive, do not establish that the agent preserved the system you meant to build.

The 119-comment thread on Hacker News adds review tactics and a maintenance question the essay leaves open.

What the thread adds

  • Fr0styMatt88 — a concrete review exercise: “Tell your AI agent you’re doing a code review workshop, ask it to select ten files at random and then go through them with you one at a time.” They explicitly retain human review: “You’ll still need to manually review the code to keep it in check”.
  • Schiendelman — a warning about using the original application as the specification: “you can often start with documenting current behavior with tests and then carefully going through to figure out what bugs you’ve just enshrined in test cases…”
  • adventure331 — dissent on how much scrutiny low-stakes work deserves: “I’m not opposed to shipping something that I understand only 60 or 70% and save energy.” This is a proposed tradeoff, not evidence that these rewrites are reliable.
  • rootusrootus — pushing back on another commenter’s analogy between trusting agents and trusting compilers: “A compiler is far more deterministic than any LLM. I am skeptical that the two situations are really comparable.”

Who maintains the port?

stephbook asks: “Who implements new features, and in which code base?” bocklund reports the practical version of that uncertainty: “I have a port that passes all the tests, but there’s enough code that it’s hard to review that the LLM didn’t change implicit design decisions or implement things in an unmaintainble way.” Passing the port’s tests does not settle their question about confidence to switch over.

Handles are pseudonymous; HN publishes no per-comment scores. Ordering is HN’s own ranking, while these excerpts are selected for usefulness. This is a slice of the thread, not a consensus.