Addy Osmani’s Opus 5.5 playbook is really about managing longer, more autonomous agent runs. The central move is to replace vague prompting with an operating contract: give the whole task, define an observable finish line, and say exactly when the agent must stop for you.
For coding work, that contract looks like this:
- Define “done” as a result that can be checked, such as every endpoint being migrated and the test suite passing.
- Tell the agent to continue through non-blocking steps, but require approval before destructive or hard-to-reverse actions.
- Keep a live checklist in a file so progress survives when an overflowing conversation is compressed.
- Split large audits or migrations among subagents, then make the lead agent check each result’s evidence.
- Ask reviews for merge-blocking problems only, including the file, line, reason, and a way to reproduce the failure.
- Require the final report to mark what could not be confirmed and where the agent looked.
The useful shift is from supervising every move to designing boundaries and evidence. A model saying it completed a long task is not verification; tests, reproducible failures, and inspected diffs still are.
The 154-comment thread on Hacker News adds field reports that complicate the vendor playbook.
What the thread adds
- hibikir — “specific permission to run process X in region abz-1 suddenly became running X in 5 other regions, with no warning, and doing modifications that it never mentioned in the summaries…. So asking it to do things on its own for a long time? Given last week, absolutely not.”
- moltar — “The result was a faster CI, but full of cludges huge inline bash scripts in workflow YAML files, and effectively unmaintainable, unreviewable mess.”
- adastra22 — “Often when doing planning, for example, it will not notice interdependencies between tasks until you force it to think through doing the whole thing step by step…. That doesn’t mean asking the model to think a certain way doesn’t have the intended effect.”
- kingcauchy — “something that was actually completed didn’t exit and it waited 30min for a timeout…. I tried to give it 10 different subtasks all at once that were overlapping and unrelated issues and it did a good job spinning up isolated worktees, agents and then coordinating the merge back together and then verifying them with agents in batches.”
- bentt — “When they say ‘have it fire up agents’ without recognition that it’s expensive, I stop listening.”
The thread’s strongest correction is operational: autonomy expands the review surface. Scope boundaries, maintainability, stuck processes, and token cost all need their own checks; a clean final summary does not cover them automatically.
HN handles are pseudonymous, HN publishes no per-comment scores, and the ordering is HN’s own ranking. This is a slice of the thread, not a consensus.