The post opens on a television gag: a character can’t open a door, twice, because something is jammed behind it, and mutters “stupid thing sucks.” The author, patrickxia, offers that as a bad model of doors — and a good description of how software failures increasingly get received. From there it turns into an argument about evals, confidence scores, and what happens when nobody measures the failure rate before shipping. The 88-comment thread on Hacker News supplies both a validated counterexample and a lot of disagreement about whether AI development is the thing making this worse.
The case against shipping first
- The target is Jev, a model from TypeSafe AI that returns typed values with probability estimates. The pitch is that it’s fast, cheap, and quick to build on.
- The structural objection: to know whether Jev works, you need an eval suite and a ground-truth pipeline — and if you have built those, you have done most of the work of fine-tuning your own solution anyway.
- His prediction about buyers isn’t flattering: “Nobody buying this is running evals. They’re just handing opaque questions to Jev and getting opaque responses,” checking the “AI-powered” box, and shipping before Friday.
- Failure modes, error budgets, and test sets get deferred: “The user can discover the failure rate! You’ve already shipped!”
- On confidence scores, two things are needed and rarely present — an understanding of how well the scores are calibrated, and a model of what the uncertainty costs in that specific decision. Jev’s marketing leans on benchmark scores rather than calibration; its own docs suggest a 0.5 threshold for “do nothing” and 0.9 for “high-risk actions” while noting the right thresholds depend on your domain.
- His summary of current practice: “At best, people use confidence scores in a cargo cult manner. At worst, people use them as an excuse for why the API call failed. The model was only 73% confident! That means my error budget is 27%!”
- The asymmetry he cares about: a broken button normally has a traceable cause and an owner whose job is to find it. Probabilistic systems make “it just does that sometimes” an acceptable answer.
- The fear, stated precisely: not more failures — LLM-accelerated development will produce those — but that “sometimes it just sucks” becomes the accepted endpoint of investigations. His own remedy is that the QA workflows and the eval that would justify or retire a service like Jev are now a few prompts away.
What the thread adds
- benjaminsky2 — the counterexample: they validated Jev’s confidence scores, found accuracy scaling linearly with confidence across three use cases, and >0.9 matching a human labeler; they also found an unexpected user behavior for about $3. avianlyric’s reply is the sharpest line in the thread: doing that validation means you built the evals — which is the article’s argument, not a refutation of it.
- pmarreck — practitioner counterpoint from someone loud about determinism and reproducibility: they run agent-assisted development too, and it works because they require “pretty much every check in the book.” Their advice is to raise personal standards, and they note unreliable software was already untenable before agents made it worse.
- oli5679 — writes AI systems for ops automation and doesn’t share the pessimism. Evals are scarce because defining “good” is laborious and politically fraught; given good evals, they say frontier models outperform human ops teams. Jev’s claim is about expanding what’s achievable at a given cost.
- adamddev1 — where the tolerance argument breaks: “good enough” may be survivable in a user-facing app, but normalizing it in libraries, infrastructure, and compilers slows everyone down. grumbel disagrees with the frame entirely — the actual argument for LLM coding is “it will get better,” and the practice is barely a year old.
- theamk — the phenomenon predates AI, and it isn’t only users: GitHub returning 5xx, an AWS service failing, mail not being delivered — “Nothing we (developers) can do, ‘stupid thing sucks.’” ryandrake adds that bugs are a choice companies make, not a natural property of software.
- IshKebab — Jev would be used where a person used to make the call (reviewing app updates, screening CVs, labelling bugs), and those processes already failed inexplicably; old-school deterministic code isn’t always the alternative.
- voidhorse — the narrow-tool argument: the narrower the function, the clearer the contract and the explanation of what went wrong. Probabilistic computation as the foundation trades predictability away for flexibility.
- WorldMaker — “confidence” is an anthropocentric word applied to a statistic; business readers read it as a letter grade, which is a large part of why ML produces dumb outcomes.
- sixdimensional — points to Perrow’s Normal Accidents: in tightly coupled complex systems, interacting failures are inevitable, start small, and cascade.
- layer8 — the normalization of inexplicability is coupled to a normalization of unaccountability.
- easterncalculus — dismissive of the post itself: an “appeal to adult animated show” framing reads as blogspam.
The comment count is from the thread itself. HN handles are pseudonymous and the site publishes no per-comment scores, so the ordering above is HN’s own ranking, not a vote — a slice of the thread, not a consensus. Vendor-performance claims from commenters are their own reported results.