Cockroach Labs built MOLT Sinai, a coding-agent pipeline modeled on a teaching hospital. Issues become patients; separate agents investigate, propose treatment plans, implement fixes, review them, and check that the review was properly completed before merging. GitHub Actions and issue labels run the workflow.
The useful part is less the metaphor than the safeguards:
- Review the plan before writing code. Another agent challenges the diagnosis, proposed changes, tests, and risks. If scope changes, the implementation stops for a revised plan review.
- Prove the tests exercise the fix. Reviewers disable the implementation and check that the new tests fail. Changing a supposedly incorrect test requires separate justification and review.
- Hand off rather than flail. A stuck agent records the problem, attempted work, next actions, and risks; the receiving agent writes back its understanding before proceeding.
- Audit the review itself. The final gate checks approval, unresolved discussion, passing automated checks, and clean commit history—not just whether an agent said “looks good.”
The authors report adding IBM Db2 migration support in less than two days for $4,172 in tokens, compared with nine months and roughly $160,000 for earlier Oracle support. The experiment also produced more than a million lines of code over roughly five months.
Those numbers need a qualification. The integrations are different projects, not a controlled comparison. After the first week, the team required human approval for every merge. And the authors say they are still confirming the correctness of the initial Db2 work, which humans had not reviewed. This is evidence for substantial agent implementation inside a supervised process—not proof of safe, fully autonomous database development.
The operational failures are just as instructive:
- One urgent, one-line fix went through eleven rework rounds over two days. They added escalation thresholds to interrupt non-converging reviews.
- Agents generated nearly half the issues. Research proposals needed both throttling and a cap on the waiting queue.
- Their instructions grew beyond 100,000 words; an audit flagged 23% as removable redundancy.
- Tokens cost just over $135,000 across the experiment, separate from the human engineering and review involved.
The lesson is to borrow the controls before borrowing the whole institution: reviewed plans, meaningful tests, independent review, bounded retries, and explicit escalation. Generating more code—and more work—is easy. Deciding what deserves to ship remains the scarce skill.