Matt Pocock talks with Lauren Tan—known as poteto, creator of pstack, and an engineer at SpaceX—about the system behind her reported 2,500 production PRs in one month. The number is the attention-grabber, but her real argument is about trust: autonomy comes after agents can observe their own work, verify it, and operate inside an environment engineered to make the right path the easy path.

Extract the process, not just the code

  • After leaving Meta’s React team, Tan began a side project while burned out and found herself micromanaging a single coding agent for hours
  • The early pstack work grew from trying to extract her own engineering process into reusable instructions; its history remains visible in the brain directory of her open-source Noodle repository
  • At Cursor, she initially debugged the Agents window manually with flame graphs and heap snapshots, acting as what she calls the “meat proxy” between the agent and Chrome DevTools
  • Better models do not remove the need for expertise. They shift the bottleneck toward expressing intent, recognizing good work, and giving the model precise concepts
  • Terms such as test-driven development or tautological test compress a larger method into language that reorganizes how the agent approaches a problem

The Michelin-kitchen model

  • Tan dislikes software factory as a metaphor because it implies volume without craft. She prefers a Michelin kitchen: high throughput can still be deliberate and high quality
  • A solo developer is the home cook who shops, prepares, cooks, and cleans. Adding agents without reorganizing the environment just overcrowds the kitchen
  • The human gradually becomes the executive chef or technical lead: divide the work, prepare the environment, supply tools and context, coordinate dependencies, and remain accountable for the result
  • Skills, CLIs, code structure, lint rules, and runtime access are the kitchen. Agent capability alone is not the system

Verification starts the trust ladder

  • Tan calls verification the most important capability in an agent toolkit because it gives the agent hands and eyes
  • The agent must be able to run the application, interact with it like a user, inspect traces and snapshots, compare the result with a rubric, and revise its work
  • That closes the loop: change → run → observe → measure → debug → repeat
  • For performance work, the loop supports hill climbing: make an attempted improvement, measure it, keep it only when the score improves
  • She says every AI application at Cursor or SpaceX now has an automatically maintained verification skill shared by the team
  • The trust ladder cannot be skipped. If the agent cannot verify its own output, the human remains the verification layer and autonomy stalls

Deterministic work belongs in tools

  • Tan’s verification CLI wraps Playwright, the Chrome DevTools Protocol, and application APIs so every agent does not rebuild the same machinery from scratch
  • Her division of labor is simple: leave interpretation, synthesis, and tradeoffs to the model; put repeatable mechanical operations in scripts and CLIs
  • This saves context and preserves solutions that earlier agents already made reliable
  • The same rule applies to migrations: use codemods and syntax-tree transformations for deterministic edits instead of asking a model to rediscover the transformation in every file
  • A skill should expose the workflow and the tool; it should not spend thousands of tokens reteaching the agent how to implement its own instrumentation

Make bad code difficult to write

  • Agent engineering increasingly means improving the environment in which agents work, not merely improving prompts
  • Tan uses TypeScript’s type narrowing as the model: progressively reduce the space of possible outputs until invalid forms are excluded
  • Her team built Dune, an internal framework she describes as an internal Next.js for Electron applications
  • Dune enforces convention-driven feature directories, restrictive lint rules, registries, and codebase discovery so there is one obvious way to add functionality
  • The framework followed painful early Grokbot versions with roughly eight “god files,” each at least 10,000 lines long
  • When several agents repeat a mistake, the fix should move into the system: can a lint rule catch it, can the type system reject it, or can the framework make it impossible?
  • Constraints reduce the instruction burden because guidance appears exactly where the agent reaches a boundary

Connect coding agents to the outer loop

  • Tan separates the inner loop—coding agents working from a snapshot of intent—from the outer loop, where new information arrives through Slack, Linear, X, email, and other systems
  • Without a connection between them, the human must gather new context and relay it into every coding session
  • Her goal is to remove that relay work without letting agents guess: subscribe to a bug-report channel, reproduce the issue, check whether it still exists on main, then distinguish a product bug from bad setup, missing data, or a dependency problem
  • “Company brain” and “context graph” are less mystical in this framing: teach the agent how to retrieve the information a human would otherwise have to carry into the chat
  • Her recurring diagnostic is: Where am I still the bottleneck, and why does the agent need me to answer this?

Coordinator agents manage the work

  • Tan does not manually open thousands of agent chats. Cursor Projects provide cloud-based coordinator agents with their own computers
  • A coordinator maintains tasks, delegates to subagents, monitors progress, transfers context, and chooses an agent topology instead of implementing every change itself
  • Grokbot supplies outer-loop connectors to Slack, email, calendars, Linear, X, and other systems, then routes relevant reports into a project
  • If 30 related issues arrive, the coordinator should not launch 30 unrelated fixes. It can group them, preserve shared context, and look for a common root cause
  • She describes running the equivalent of more than ten chief-of-staff agents, each responsible for a domain such as performance, user-reported bugs, or an experimental native rewrite

Gardening produces much of the volume

  • Tan is explicit that 2,500 PRs did not mean 2,500 product features. Much of the output was gardening: maintenance, code quality, architecture, and improvements to the working environment
  • One recurring routine scans React code for known footguns and banned patterns, but does not immediately fix every result
  • Findings accumulate in a document that she reviews every few days, allowing related symptoms to be grouped before agents are dispatched
  • The buffer prevents duplicate work and helps the coordinator see a systemic problem instead of treating every occurrence as an isolated task
  • A constrained, well-maintained environment also benefits human teammates and new hires; the agent infrastructure is still software infrastructure

Review by sampling—and repair the system

  • At this scale, Tan does not review every PR before it lands. She samples output as quality control and inspects landed commit history afterward
  • She looks for repeated shortcuts, workarounds, and weak patterns rather than treating each defect as an independent failure
  • A one-off mistake may only need a correction. The same mistake across agents should produce a new skill, lint rule, type constraint, framework convention, or verification check
  • The feedback loop is therefore not just agent → code. It is agent failures → better environment → better future work
  • Tan warns that installing pstack does not create this state. The autonomy rests on substantial prior investment in tools, constraints, and verification

Full autopilot still has controls

  • In pstack’s full-autopilot mode, an agent can open a PR, spawn verifier agents, run the application, click through it like a user, fuzz for regressions, fix what it finds, repeat the verification, and merge
  • Tan describes configurations with roughly ten verifier agents per PR, while noting that this is expensive and can be reduced to one verifier or self-verification
  • Her agents now work and merge around the clock; she samples the landed work, then reverts, adjusts, or adds guardrails when needed
  • This is not “dark” software development in the sense of abandoning quality. It is unattended execution built on unusually heavy observation and feedback
  • Pocock distinguishes reversible two-way-door changes from irreversible one-way-door risks. Tan does not claim every domain can safely reach this level of autonomy

Autonomy is bounded by verifiability

  • Software is unusually suitable for agent autonomy because many properties can be compiled, tested, fuzzed, traced, or checked programmatically
  • Formal methods can push the boundary further: Tan points to Lean, TLA+, and languages such as Bend as examples of bringing programs and machine-checkable proofs closer together
  • But some work cannot yet be verified with enough rigor to justify autonomous merging
  • The real ceiling is therefore not how many agents can be launched. It is how confidently their work can be evaluated without a human becoming the hidden test harness

Skills are portable workflows

  • Tan and Pocock reject the idea that skill libraries are competing products. Skills are processes expressed in language, usually Markdown, and can be combined or rewritten
  • Past agent transcripts are especially valuable because they record the process as it actually happened: repeated corrections, interventions, mistakes, and unstated assumptions
  • Tan’s Recall skill mines those old conversations, compresses useful context, and transfers it into a new task
  • As models improve, she expects skills to contain fewer low-level command details and focus more narrowly on durable workflow steps
  • The practical starting point is to find where you repeatedly intervene, then decide whether the correction belongs in a skill, deterministic tool, lint rule, type constraint, or verification loop

“You can do anything you want. It’s just language.” — Lauren Tan