Matt Pocock walks through the v1.3 release of his open-source skills. The theme is not getting an agent to write more code; it is making the whole path from specification to review more reliable, then learning from what went wrong.

Implement-spec: one spec, many focused sessions

  • Start with a spec and tickets that identify what blocks what. Treat the tickets as a dependency graph, not a flat checklist: independent tickets can run in parallel.
  • An orchestrator reads the work, explores the codebase, creates an integration branch, and assigns ready tickets to fresh implementer subagents in separate Git worktrees. Each ticket uses test-driven development; completed work is merged back, then the next unblocked tickets begin.
  • After the tickets land, the whole integration branch gets a code review and fixes. The target is a coherent branch, not automatically a PR; a draft PR is conditional on the team’s workflow or a request.
  • Pocock calls this a useful middle ground for unattended work. He still prefers a deterministic script to drive the ticket loop when one is available: it is more predictable and cheaper than asking an agent to babysit the queue. Worktrees isolate edits, but integration conflicts can still appear at merge time.

PR: make the evidence easy to inspect

  • The PR skill supplies a body with three parts: a visual summary (small diagram, pseudocode, call tree, or diff sketch), before/after evidence from actual tests or runtime behavior, and merge danger.
  • Merge danger distinguishes a two-way door (easy to revert) from a one-way door (a change whose external effects cannot simply be undone), then names the blast radius if it fails.
  • Asking for evidence is itself a useful intervention: it pushes the agent to run another test or collect a screenshot instead of declaring a change correct because the code looks plausible. The skill changes how a PR is explained; it does not replace review.

Retro: improve the environment, not just the output

  • Retro reads real coding-agent sessions, including ones where the feature shipped despite friction the agent never reported. It looks for navigation problems, missing automated checks, weak review standards, bloated instructions, wasted tool calls, and missing information.
  • Its strongest prescription: put mechanical rules in executable checks (lint, hooks, CI), and reserve written coding standards for judgment calls. Remove instructions that say nothing actionable.
  • It is deliberately human-in-the-loop: rank findings, let a person choose which ones matter, and only then change the environment. Pocock warns against automatically running retro and applying every suggested fix; false positives can compound into a worse repo.
  • In the demo, it spots an unguarded check script, context lost during long sessions, and a token-hungry CLI. It also flags a release decision whose severity Pocock himself judges differently—a good reminder that findings need human interpretation.

One migration to check

The release renames CONTEXT.md to GLOSSARY.md and CONTEXT-MAP.md to GLOSSARY-MAP.md. Skills updated to the new convention look for the new filenames, so projects using Pocock’s domain-modeling workflow should move their existing files rather than creating empty replacements. His v1.3 changelog has the update details.

The useful distinction: implement-spec coordinates the work, PR makes the result legible and verifiable, and retro turns observed friction into better checks and workflows. None removes the need for a human to judge what should ship.