Philippe Xanthopoulos ran offshore delivery at a scale of ~240,000 development hours per year. His signature sat at the bottom of software budgets. So when he says what software costs, he says it as a witness — not a theorist.

Over weeks this year, working alone with two AI collaborators under written contracts, he shipped a production decision engine whose quality metrics beat anything he ever saw come out of a team. The engine makes technology strategy auditable using mixed-integer optimization — complete search with optimality certificates, no heuristics, every shadow price printed as a Lagrange multiplier wearing its business name.

  • Nine enterprise seats (PM, BA, architect, QA lead, testers, release manager, steering) collapsed into one person running a task queue. WIP limit of one. State in records, not in the author’s head.
  • 63 pull requests, 531 commits, one git author. Traditional estimate: months with a full bench.
  • “Vibe coding with a constitution”: laws written down, binding on everything including the author. Data speaks. Verdicts stay human. Nothing ships untested.
  • The validator battery’s first act was filing 990 findings against its own author’s corpus — and most indicted the battery itself. Reviewers reviewed everyone.

“I graded everything. I typed almost nothing. I could not have done this alone. Everything collapsed except the place where judgment sits.”

The heaviest insight is also the most uncomfortable for the industry: these tools do not fail beginners loudly. They flatter them into disaster. Without the knowledge to grade, succeeding falsely is worse than failing — you find out in production, in front of the people who trusted the report.