Skip to content

How to be like @poteto

A practical guide to setting up engineering bots and agents the way Lauren Tan (@poteto) does it. It's built from her X posts, articles, talks and the pstack plugin, Aug 3 to Oct 3, 2026. She leads Grok Bot engineering, previously worked at Cursor, and says she shipped 2,500 PRs to production in one month (talk post).

Every point links to its source. Paraphrases are marked as paraphrases; anything in quotation marks is her exact wording.


The one-paragraph version

Make it cheap for an agent to prove its work, then make it impossible for the same mistake to happen twice. You build a verification skill and CLI for each app. Every correction you make gets pushed down into the code: architecture first, then lint or tests, then skills, and human review last. Work runs on cloud agents, each with its own computer, driven by a coordinator (a Cursor Project or a bot) that never writes code itself. Bots and routines feed the outer loop (bug reports, user complaints, ideas), and Projects run the inner loop (code, verify, ship). As the verification gets better, you let agents merge their own PRs and review what landed on main afterwards. (1, 2, 3, 4)


1. Principles

  1. Verification is everything. In her words, a high-quality verification skill is "the most critical skill to have in your toolbox." Treat it like critical infrastructure, not "just" a skill. She even suggests putting an on-call rotation on it. (pstack guide Pt. 1)
  2. Remove the correction, don't repeat it. Every time you correct an agent, ask how to get rid of that correction for good. In order of value: (1) remove the problem with better architecture or data structures; (2) make it a lint rule or test so CI catches it; (3) make it a skill or rule; (4) rely on human review, which she labels "ngmi." (post)
  3. Hard rules beat soft rules. CI checks, lints and compiler diagnostics turn CI red; rules, skills and Bugbot are "soft" and agents won't apply them consistently. Layer both kinds, but never rely on soft rules alone. (Maven workshop, ~47:45–48:45) The best review is the one your codebase and lint rules enforce. (reply)
  4. The codebase is the agent's memory. Agents copy the patterns they see, so clean up tech debt and record rules inside the code itself. (Gene Gi's notes on the 2,500 PRs talk; talk)
  5. Be the gardener. Someone has to watch the PR stream for smells, like the third isRecord helper this week or lint suppressions creeping in. "Every weed you pull becomes a rule, so it can't grow back." (post)
  6. Work backwards from "the agent merges its own code." She calls this her big unlock for reaching 2,000+ PRs. (Matt Pocock interview, ~39:40)
  7. The three virtues of Grok Bot use: Laziness, Impatience and Hubris. Hubris means owning the outcome even when your own hands didn't make it. (post)

2. Setup steps, in order

  1. Install pstack in Cursor (marketplace, source) and in Grok Bot (plugin). In Cursor and Grok Bot it updates automatically (post). Run /setup-pstack to pick models per role. Her advice is a big model as coordinator and an efficient one (Auto, or Grok non-fast) for the coding workhorses. (reply)
  2. Start with zero skills, then add them based on evidence. Watch where agents fail and pull in only the skills that measurably help. (reply, reply)
  3. Run /create-verification-skill on each app. It writes .cursor/skills/verify-<app>/ with Launch, Doctor, Drive, Evidence and Cleanup sections, plus a feature map under features/. Before handing it over, it proves itself once end to end. (post, guide 06, example repo)
  4. Build the lever: a small, agent-friendly CLI that scripts interaction with your app. It should have composable subcommands, --dry-run on anything destructive, descriptive errors, rich --help, and machine-readable JSON output. (Pt. 1)
  5. Add a feature map. It's a searchable markdown index of every feature: how to reach it and what result proves it works. She calls it "materialized memory." (Pt. 1)
  6. Schedule /maintain-verification-skill daily as a Grok Bot routine or a Cursor Automation. It ends as clean, changed (one PR confined to the skill's own directory) or blocked, and it never edits product code. (post, Pt. 1, guide 06)
  7. Make cloud agents the default. "worktrees are dead. cloud agents are the future." Each agent has its own computer, so it can run the app, take videos and screenshots, and drive the UI. That's why she can trust the output. (post)
  8. Group work into Cursor Projects. A coordinator supervises every agent in the Project, and you can drag existing chats in. She runs 10+ Projects in parallel. (post, post, Cursor blog)

3. Rules, AGENTS.md and skills

  • Put rules in code first. Use strong types, a framework that constrains structure, and lint rules for observed bad patterns. Her internal framework, "Dune", works like an internal Next.js for Cursor's Electron apps and limits the shapes code can take; she credits it with letting her agents merge their own PRs. (Matt Pocock interview, ~30:00, post)
  • Use opinionated anti-slop lint rules, for example dmmulroy/anti-slop for TypeScript. (post)
  • Ban code comments. Agents use comments to justify workarounds, and you end up with "hacks built on top of hacks." pstack has a /no-comments skill. (reply, post)
  • Prefer skills over a giant AGENTS.md. Skills compose and aren't always loaded. (reply) She set disable-model-invocation: true on all pstack skills so they only run when invoked. (reply)
  • Eval skill changes before trusting them. Use /poteto-mode update and eval the skill to <change>, and you can hill-climb on it. (reply) She runs evals for her own skill changes on cloud agents. (Peter Yang episode, ~17:00)
  • Make your own "-mode" skill. Use /automate-me and /reflect, and mine your past transcripts for the corrections you keep making. Internally she calls hers /lauren-mode; in pstack it's /poteto-mode. (post, pstack guide 09)
  • Don't write SKILL.md freehand. Route it through the authoring playbook so it gets validated. (guide 10)

4. Verification and eval loops

  • State the finish condition up front: concrete checks the agent can run, not "make it better." A confident reply with no evidence is a red flag; "inconclusive" is an honest answer. (guide 06)
  • Match the check to the change. A CLI change runs the real command. A UI change walks the flow in the running app. A migration replays saved input. A perf change compares profiles. A storage change reads the value back. (guide 06)
  • Use a verifier swarm. Spawn a /swarm of agents to "fuzz" every stack of PRs, so agents can merge their own work while you sleep. (post) The agent that judges a change is never the one that wrote it. (guide 06, Shipping)
  • Use the overnight contract: goal, finish condition, permissions, escape hatch, /loop, and a decision log. A plateau means pivot, and the finish condition never quietly relaxes. (guide 07)
  • Test behavior, not implementation, and attack the premise. Both are pstack principles. (pstack README)

5. Review flow and the trust ladder

  • Agents open small, ordered PRs with evidence in the description. Five narrow PRs beat one fat one. /poteto-mode babysit this pr takes blockers in order (conflicts, then review threads, then CI) and stops at merge-ready. (guide 06)
  • Review comments get triaged with skepticism, not accepted wholesale. /interrogate sorts real catches from noise. (guide 10)
  • Bugbot plus agentic review. Her agents fix the real bugs those reviews find and then merge on their own. (reply)
  • Code owners' PRs get auto-approved; the codebase and lint rules do the enforcing. (reply)
  • Review main after the fact. Her agents merge their own PRs, and she reviews what landed while she was away or asleep. (post) Some projects merge their own PRs; she reviews after they land. (reply)
  • Remove the human step by step as trust grows. Make each step an agent. (Gene Gi's talk notes)

6. Planning (through code)

  • Her most-used prompt: "restate in your own words what you think my goals are and what the problem i'm trying to solve is." (post)
  • Understanding skills: /teach calls /how (runtime mechanics) and /why (history and intent); /recall pulls context from past chats. (Pt. 2)
  • Plan with prototypes and an architecture arena. Parallel candidate designs, often from different model families, are scored by a cross-judge that uses a different model, and the whole design gets scrapped if it turns out wrong. (Pt. 2)
  • Don't adversarially review abstract plans; agents invent theoretical risks. If you commit plans for a big project, delete them when you're done. (Pt. 2)

7. Parallel agents and the bot hierarchy

  • Bots are managers, not coders. She asks a bot to spawn cloud agents rather than do the work itself, so there's one coordinator and many agents. (reply) Her favorite setup is a team engineer bot in Slack that you @ for coding work and that can create Projects. (post)
  • Her hierarchy: she talks to her chief of staff, which talks to the eng lead bot. Dr Eggbot set the eng lead up to break tasks down, delegate and supervise rather than do work itself. (Peter Yang episode, ~25:20–26:40)
  • Outer loop and inner loop. Bug reports land in Slack, Linear or X, which aren't connected to your inner loop. Cursor Projects are her inner loop and Grok Bot is her outer loop: the bot gathers external context and messages it to the Projects. (Matt Pocock interview, ~35:15–44:15) Grok Bot routines gather context for her from Slack bug reports, X complaints and feature ideas. (post)
  • Local coordinator, cloud subagents. Also /in-cloud, and a /swarm of 10+ agents that fuzz PRs on autopilot. (post)

8. Grok Bot practice

  • Keep routines infrequent. A 15-minute routine runs almost 100 times a day; hourly or a few times a day is usually enough. Long chats make routines more expensive, so give recurring routines to a fresh bot. (post)
  • Use Dr Eggbot to design bots. Its routine health check flags unused or expensive routines (weekly by default), and a daily friction routine skims your chats for corrections and suggests skills, bots or routines. (post, Peter Yang episode, ~20:30)
  • Ask Dr Eggbot for an engineer bot that runs /create-verification-skill and sets up a daily /maintain-verification-skill routine. (Pt. 1)
  • Use tinkabot to wrap an API as a Grok Bot plugin. (post)

9. Tooling she uses

  • Cursor cloud agents, Projects and Automations (post, Pt. 1)
  • Grok Bot with routines and connectors (post)
  • pstack, including /poteto-mode, /goal, /loop, /swarm and the Full Autopilot playbook (post, autopilot-full)
  • Bugbot (reply)
  • The Dune framework (interview)
  • The benny automation pack for Slack triage, repro and fix (source)

10. Anti-patterns

  • Naming skills step by step in the prompt instead of stating the goal (guide 10)
  • A vague finish condition; a duration is not a finish condition (guide 10, guide 07)
  • Running parallel agents in one worktree (guide 10)
  • Accepting every review comment (guide 10)
  • Reporting success off a green build (guide 10)
  • Relying only on soft rules: AGENTS.md, skills and Bugbot without CI enforcement (Maven, ~48:20)
  • Frequent routines and long-chat bots running recurring jobs (post)
  • Installing a pile of skills up front (reply)
  • Leaving plan files in the repo after the work is done (Pt. 2)
  • Comments as band-aids (reply)

Sources

X posts and articles by @poteto (Aug 3 – Oct 3, 2026)

Talks, videos and podcasts

Code and docs