Hamel Husain and Isaac Flath test a claim about Opus 5.5’s writing in a 20-minute blind, side-by-side review against Opus 5. Their verdict is deliberately narrower than “slop is solved.”
How they compare
- A small app hides the model names until review; the speakers use the same writing tasks and inspect outputs aloud.
- Tasks include an AI-evals FAQ, a skeptical-reader summary of Isaac’s Jev article, replies to technical posts, an intentionally cringeworthy influencer post, and a conference abstract.
- They judge flow and reader context, not just grammar or whether a response includes all the facts; they notice when one model makes fewer tool calls but cannot tell whether it read the same context.
What improves—and what doesn’t
- In the Jev summary, 5.5 opens more naturally, then swerves into numbers and task details before explaining why they matter; 5 is awkward and needlessly defensive.
- Replies such as “Great write-up” still sound like engagement bots. Prompting for an influencer voice elicits the expected hooks and clichés from both models.
- A request for a short, thoughtful reply gets outputs they reject from both models; the conference abstract has slightly clearer framing from 5.5 but remains generic and needs heavy editing.
- Across this informal handful, 5.5 tends to fare better, but both use strikingly similar structure. This is not a rigorous benchmark or evidence of why the model changed.
A more useful writing eval
- Ask whether an unfamiliar reader knows what a number or example refers to before it appears.
- Prefer specific, grounded examples to an inventory of details dumped into an introduction.
- Add a both failed option rather than forcing a winner between two drafts you would rewrite.
- Blind the labels and inspect the entire draft: an opening can sound good while later sentences fall apart.
“Slop is not dead. 5.5 might be a little bit better in some cases, but it’s not like a landslide.” — Hamel Husain