A 65 minute conversation, rebuilt as a reading guide

Lauren Tan's agent kitchen: how she scaled to 2,500 PRs

Lauren Tan, known online as poteto, explains how she moved from supervising one coding agent to running many coordinated projects. Her method starts with verification, then changes the codebase and tools until agents face fewer bad choices. Only after that foundation does parallel work become useful.

Read the headline carefully. The 2,500 pull requests in a month are Lauren's account of her own output. The figure was not independently audited in the interview, and she says many of those PRs were maintenance and codebase gardening. They were not 2,500 product features.
2,500 PRsLauren's reported monthly total, including maintenance
>10 coordinatorsher reported chief of staff agents across separate work areas
8 huge filesher description of early Grok Bot code, each at least 10,000 lines
65:36full interview duration, covered through the closing skills discussion

Matt Pocock interviews Lauren on 3 October 2026. This guide pairs a small number of real interview frames with explanatory diagrams. Every frame comes from the stated interval. Watch the source video.

01 · 01:35 to 06:20

Trust began with removing herself as the human adapter

Matt's opening questionHow did Lauren climb from one supervised agent to large-scale parallel work?

Lauren traces pstack back to a side project after leaving Meta's React team. In January or February, she found herself spending hours micromanaging one agent. She started writing skills and a brain directory to transfer parts of her own working method into the agent. She could iterate quickly, but she had little evidence that a skill improved the result.

She joined Cursor in March and soon worked on performance problems in its agent window. The early work was manual. She read flame graphs, inspected heap snapshots, and passed information between the agent and Chrome DevTools. Her own labor connected the model to the evidence it needed.

That bottleneck shaped the rest of her method. A capable model still reaches for shortcuts. The useful intervention is to redesign the process so that the quick path also produces the desired result.

"how do I make the easy thing the right thing?"

Lauren Tan, 06:16
The trust ladder is earned through evidence.Lauren does not describe trust as confidence in a model brand. Each higher level follows from giving the agent better tools, constraints, context, and feedback.
Transcript summary · 00:00 to 06:20
Matt frames the interview as a follow-up to Lauren's earlier talk. Lauren covers her Meta background, the side project that became the basis for pstack, her March start at Cursor, and the manual performance debugging that exposed her role as the bridge between agent and browser evidence.
02 · 06:20 to 10:45

Better models make precise intent more valuable

Lauren argues that domain expertise gains value when the model can carry out more of the implementation. The limiting step moves toward the human's ability to define the goal, recognize a good result, and explain the constraints. A doctor, lawyer, or engineer with enough technical fluency can use an agent to build around knowledge the model does not own.

Matt connects this to language. A term such as test-driven development changes which evidence the agent seeks. His grilling method asks questions until hidden requirements become explicit. Lauren adds tautological tests as an example of a precise label. Such a test restates the implementation and can pass without increasing confidence in behavior.

Vocabulary compresses a working concept into a phrase the agent can reuse. That compression helps communication, but it does not prove the agent understood the requirement or implemented it correctly. Verification still has to supply that evidence.

In plain words

When execution gets cheaper, deciding what good means becomes more of the job.

Transcript summary · 06:20 to 10:45
Lauren says domain experts have an advantage because clear intent becomes the bottleneck. Matt discusses how terms such as TDD guide agent attention. They use tautological tests to show how a compact technical term can carry a detailed criticism.
03 · 10:45 to 16:06

The kitchen metaphor keeps quality in the throughput story

Lauren accepts that software factory describes high-volume production, but she prefers a Michelin kitchen. A factory can suggest interchangeable work and indifferent quality. A serious kitchen still produces at scale, yet its organization exists to protect craft, user experience, and the finished dish.

A solo cook prepares ingredients, cooks, and cleans. Adding more people without stations, tools, or shared methods creates congestion. The chef's work changes as the kitchen grows. She chooses ingredients, prepares stations, defines methods, and owns the result even when other people do most of the cooking.

Lauren maps that role to engineering. The codebase, tools, skills, and verification system become the kitchen. Agents do more implementation, while the engineer designs the conditions in which implementation happens. The engineer's name and reputation still attach to the outcome.

Solo cookone agent, constant supervisionhuman carries every order tools Organized kitchenstations, constraints, verificationshared methods and evidence trust Several projectscoordinators route contextworkers execute in parallel Explainer's reconstruction of the progression, not formal pstack levels
The scale comes after the middle box. More agents can add coordination problems when the environment is unprepared.
The kitchen analogyLauren compares adding agents to adding cooks who need stations, tools, and a division of work. 12:40
Transcript summary · 10:45 to 16:06
Lauren explains why she prefers the kitchen metaphor, walks from a home cook to a chef who organizes many people, and assigns the human engineer final responsibility for the product. Matt turns the discussion toward the environment in which agents work.
04 · 16:06 to 18:58

Verification is the part that turns a sequence into a loop

Lauren's first high-value skill at Cursor let the agent operate the application and observe the result. For interface work, that can mean running the app, interacting with it as a user, inspecting traces, reading browser state, and taking snapshots. Without those abilities, the agent has to ask a person what happened.

Planning and explanation skills can improve the proposed change, but they do not close the feedback loop. The agent must compare observed behavior with an explicit goal and revise. Lauren calls this foundation more important than any particular pstack command.

She connects the loop to hill climbing. If a performance goal has a reliable score or rubric, the agent can try a change, measure it, and keep improvements. She relates the idea to Andrej Karpathy's autoresearch. The interview gives no measured performance result, so this is a method, not a benchmark claim.

Goalrubric or behavior Implementmake one change Runuse the real app Observetraces and output Judgecompare with goal revise with evidence
Verification supplies the observation and comparison steps. A test command alone is useful only when it checks the behavior that matters.
Give the agent a way to see the resultThe discussion covers application interaction, debugging, traces, and snapshots. 16:50
Transcript summary · 16:06 to 18:58
Lauren defines verification as giving an agent ways to run and inspect its work. She explains why her earlier analysis skills still left her in the middle and how a measurable rubric makes iterative improvement possible.
05 · 18:58 to 24:55

Move repeatable mechanics out of the agent's judgment path

Lauren reports that each application in her organization has a shared, automatically maintained verification skill. A custom command-line tool handles browser interaction through Playwright and the Chrome DevTools Protocol. She is careful about its novelty. It is glue around existing APIs, not a new verification engine.

Concern in the agent community that compaction weakened results first motivated the tool, because a small command preserved context space. The longer-lasting benefit is consistency. Without the tool, every agent rebuilt a browser driver, debugged it, and discarded it. The next agent repeated the same work with a different result.

Her general rule is to separate judgment from mechanics. An agent can decide what evidence matters or how to respond to a failure. A script should handle steps whose correct execution is known in advance. She extends the same idea to migrations, where codemods and abstract syntax tree transforms can perform mechanical changes before an agent handles ambiguous cases.

Instructions onlyEach agent invents its own driver, spends tokens on setup, and fails in different ways.
Reusable deterministic toolOne tested interface performs known steps. The agent spends judgment on what to inspect and what to change.
A skill can be a thin guide around a tool.The prose explains when to use the tool and how to interpret evidence. The script preserves the repeatable part.
Transcript summary · 18:58 to 24:55
Lauren describes shared verification skills, a CLI over Playwright and browser APIs, the original context-window concern, and the speed gained by reusing deterministic code. She applies the same split to codemods and migrations.
06 · 24:55 to 28:37

Parallelism should wait until the environment can absorb it

A low-trust setup traps the engineer in supervision. Each agent needs correction, so deadlines consume all available attention. The team then has no time to build better tools, which keeps trust low. Lauren compares the cycle to working with a dull knife because sharpening feels too slow during dinner service.

Her remedy is an explicit investment in the environment. That includes fast tests, useful types, focused modules, lint rules, stable tool interfaces, and application-level verification. The interview does not prescribe a percentage of engineering time or a fixed setup period. The work depends on observed failure modes.

A good environment helps human engineers too. It reduces the number of local conventions someone must remember and makes a safe change easier to recognize. Agent readiness and maintainability point in the same direction here.

Paraphrase of Lauren's explanation

If every deadline forces more supervision, the missing task is improving the place where the work happens.

Transcript summary · 24:55 to 28:37
Matt frames quality as ease of safe change. Lauren explains the low-trust micromanagement cycle and argues for time spent on tools and constraints before trying to multiply agents.
07 · 28:37 to 33:32

Dune narrows the set of possible mistakes

Lauren uses TypeScript narrowing as the conceptual model. A broad string type permits many invalid states. Runtime checks and type guards can narrow it to a specific valid form. She applies the same move to repository design by reducing the number of places and patterns an agent can choose.

Dune is her internal framework for Electron applications. She describes it as similar to Next.js in the way it supplies conventions. It is not open source. Features live in dedicated directories, a registry discovers them, and restrictive lint rules reject patterns the team has decided against.

The framework grew from concrete failure. Lauren says early Grok Bot versions accumulated about eight huge files, each at least 10,000 lines. Feature directories prevent the easy move of appending more code to a central file. Every recurring mistake becomes a design question: can a type, directory convention, linter, or API make that mistake impossible?

Dune is an internal example. The interview gives its principles and a few conventions, not a public API. This guide does not infer implementation details beyond those statements.
Transcript summary · 28:37 to 33:32
The discussion moves from TypeScript narrowing to Dune. Lauren describes feature directories, registry discovery, lint rules, early Grok Bot files, and the practice of turning observed agent failures into codebase constraints.
08 · 33:32 to 38:50

The inner loop needs fresh context from the outer loop

Once the implementation environment works, the next human bottleneck is context transfer. Lauren calls engineering agents working against a captured goal the inner loop. That goal is only a snapshot. New bug reports, product decisions, infrastructure limits, and user feedback arrive elsewhere and can make the snapshot stale.

The outer loop is the set of systems where that information appears. She names Slack, Linear, X, and email. Grok Bot routines can subscribe to relevant sources and pass reports into an engineering project. The key value is not the number of connectors. It is removing the person who repeatedly copies information between systems.

A report still needs investigation. The engineering agent should reproduce the problem on the current main branch and distinguish a product bug from account data, local setup, or a missing dependency. Context ingestion should supply real evidence, not permission to guess.

Outer loopSlack reportsLinear issuesemail and Xnew context Ingestfilter and group reports Coordinatepass intent and context Inner loopreproduce on mainimplement changerun verificationreturn evidence verified outcome updates the next decision
The outer loop refreshes intent. The inner loop still has to reproduce and verify the report before changing code.
Transcript summary · 33:32 to 38:50
Lauren says the prepared environment is a prerequisite for volume. She defines inner projects and outer information sources, then explains how subscriptions can deliver real reports to agents that reproduce and triage them.
09 · 38:50 to 40:10

The unit of scale is a project, not a pile of chats

Lauren is not manually opening 2,500 conversations. She describes Cursor Projects as persistent coordinators. A coordinator supervises a task list, creates worker agents, carries context between them, and tracks the work. Its main job is management rather than implementation.

She worked backward from a concrete target: what conditions would let an agent merge its own code? That question forces attention onto verification, repository constraints, context freshness, and failure recovery. The PR count is an effect of those systems, not the first control to turn.

In plain words

Do not scale prompts. Scale a prepared unit of work that can gather context, delegate, verify, and report.

Transcript summary · 38:50 to 40:10
Lauren connects high PR volume to persistent projects and coordinator agents. She says the turning point was designing toward safe agent merges rather than manually reviewing an ever-growing number of chats.
10 · 40:10 to 45:31

Coordination preserves shared context across related reports

Lauren's concrete setup combines Grok Bot routines with Cursor Projects. Grok Bot watches selected external sources and sends relevant work to a project. The project runs in the cloud with a coordinator and its own computer. The coordinator can choose worker arrangements, delegate tasks, and pass findings between them.

She connects the design to a management lesson from Netflix: give context that enables self-sufficiency. A burst of 30 reports is Matt's hypothetical example, not Lauren's measured workload. The useful point is that a coordinator can group related reports before assigning work.

If several users report similar performance failures, one worker per report may duplicate the same investigation or produce competing fixes. A coordinator can compare the reports, identify the common layer, and send workers after distinct parts of one root problem.

One report, one workerFast dispatch, but repeated investigation and local fixes can hide the shared cause.
Group, then delegateThe coordinator retains the thread between reports and splits work after forming a broader hypothesis.
A coordinator manages workersLauren compares the project coordinator to an executive chef or chief of staff. 43:00
Transcript summary · 40:10 to 45:31
Matt asks what Lauren sees when managing many agents. She separates Grok Bot's external context role from Cursor Projects, describes coordinators and worker topologies, and explains why grouping related bug reports can reveal a higher-level cause.
11 · 45:31 to 47:33

The PR total contains a large amount of gardening

Lauren explicitly corrects an easy misreading of the headline. The monthly total is not 2,500 features. Much of the work is codebase gardening: small maintenance changes, cleanup, refactoring, and improvements to the environment.

That work benefits every engineer. A new hire inherits clearer structure and better defaults. Existing engineers spend less time navigating old shortcuts. Agents also encounter fewer ambiguous choices. The same PR can therefore raise future throughput without adding visible product behavior.

PR count remains a volume measure. It does not reveal size, risk, user value, defect rate, or review quality. Lauren's interview explains the operating system behind the count, but it does not supply an independent outcome audit.

2,500 PRs does not mean 2,500 features. Treat the number as Lauren's reported work volume. Evaluate the method through the evidence and safeguards she describes, not through the count alone.
Transcript summary · 45:31 to 47:33
Matt asks what else feeds the work queue. Lauren says code reading and gardening produce many PRs and explains how a prepared environment helps agents, the existing team, and new hires.
12 · 47:33 to 49:19

A queue can reveal a pattern that immediate fixes would hide

One recurring agent scans React code for problematic patterns. It does not open a fix for every finding. It appends observations to a document. Every few days, Lauren reviews the accumulated items and asks whether they are instances of the same underlying problem.

The delay is deliberate. Immediate execution optimizes each item in isolation. A buffer creates material for comparison and gives a human or coordinator time to choose one systemic change, such as a lint rule, shared abstraction, codemod, or repository constraint.

finding Afinding Bfinding C Buffered documentcompare every few days One common patternrule, structure, or shared fix
The queue is an analysis tool. It prevents a swarm of local fixes from erasing the evidence of a shared cause.
Transcript summary · 47:33 to 49:19
Lauren gives the React anti-pattern scan as an example. Findings accumulate in a document so repeated cases can be grouped before the system chooses a broader intervention.
13 · 49:19 to 51:55

Review samples rigorously, then repair the production process

Lauren does not claim to read every PR. She samples code and pull requests, reportedly every day, and inspects them closely. The object of review is both the change and the system that produced it.

A one-off error may need a local correction. Several agents taking the same shortcut point to an environmental defect. She then changes a skill, lint rule, type, constraint, or shared abstraction so the next worker does not repeat it.

This resembles statistical quality control, but the interview gives no sampling rate or defect threshold. The practical requirement is enough scrutiny to notice recurring failure. Lauren also warns that reaching this state takes substantial effort. Installing pstack does not transfer her trust or her codebase conditions to another team.

Fix the sampled PRNecessary when the change is wrong, but it improves only one output.
Fix the repeated causeChange the rule or environment when the same failure appears across workers.
Transcript summary · 49:19 to 51:55
Lauren says scale requires sampling rather than reading every change. She explains how recurring shortcuts trigger changes to the environment and stresses the time needed to build enough confidence for this workflow.
14 · 51:55 to 55:40

Some changes merge before human review

Lauren says her agents run around the clock and that she has more than ten coordinator agents. They cover areas such as desktop performance, user-reported bugs, and an exploratory native rewrite that she calls a toy. These are separate work areas with separate management context.

Full autopilot can start several verifier agents for one PR. They run the application, interact with it, look for regressions, report problems, and trigger another repair cycle. Lauren loosely calls this fuzzing. In this context, it includes exploratory application interaction, so readers should not assume it means only a formal coverage-guided fuzzer.

The process consumes many tokens and can be tuned. Lauren gives ten verifiers, one verifier, or self-verification as examples. Some PRs merge before her review. In the morning she samples commit history and can modify, revert, or add rules after seeing a problem.

This is post-merge human review backed by pre-merge machine checks. Lauren grounds her confidence in the constrained codebase and verification system she described earlier.

Explainer clarification. Post-merge review carries different risk from mandatory human approval. UI interaction cannot prove every system property, and a larger verifier count does not guarantee coverage.
Transcript summary · 51:55 to 55:40
Lauren describes more than ten coordinators, overnight merges, multiple verifier agents, repeated repair, token cost, and morning history review. Matt distinguishes this setup from treating code as irrelevant.
15 · 55:40 to 59:37

Autonomy depends on what can be verified and reversed

Matt's challengeWhat happens when a change can lose data or cause harm in medicine, law, or finance?

Lauren says the answer depends on the quality of available verification. She regards much software work as verifiable and gives mathematical proofs as another example, while acknowledging that neither field is entirely verifiable. Other domains are much harder to evaluate programmatically. She says she does not have a universal answer.

She suggests that better formal methods and agent-oriented languages may expand what can be checked. She mentions Bend, Lean, and TLA+ while discussing code and proofs. Her direction is toward stronger machine-checkable evidence, not a claim that current tools make all changes safe.

Explainer clarification. Verification and reversibility are distinct. A test may lower the probability of a harmful change, but it does not restore deleted data or reverse an external action. A proof establishes the properties encoded in the proof under its assumptions. It does not cover a requirement that nobody wrote down.

Explainer clarification. Lauren's method scales best where work can be checked. High-stakes, irreversible work needs controls beyond the workflow shown here.
No universal answerThe conversation turns to irreversible changes and domains that resist programmatic verification. 57:00
Transcript summary · 55:40 to 59:37
Matt raises one-way changes and regulated domains. Lauren ties safe autonomy to verifiability, acknowledges the gap for hard-to-check work, and discusses formal proof systems and future languages as one possible direction.
16 · 59:37 to 62:52

A useful skill library should become your own

Matt asks how people should combine his skills with pstack. Lauren rejects the idea that the libraries compete as complete systems. A skill is a process written down. Users can combine Matt's questioning or Wayfinder work with pstack execution, or take only the pieces that fit their environment.

Trust here means understanding the tool well enough to predict its behavior and correct it. A chef carries familiar knives between jobs. In the same way, an engineer should adapt a personal set of prompts, scripts, rules, and verification methods rather than copy a library unchanged.

Past agent chats are evidence for that library. Repeated corrections reveal missing instructions. Repeated workarounds may belong in a lint rule or tool. A good skill begins with a real intervention that the person no longer wants to repeat manually.

"everyone should have their own set of knives"

Lauren Tan, 61:27
Paraphrase of Lauren's explanation

Borrow processes freely, then keep only the ones you understand and can verify in your own work.

Transcript summary · 59:37 to 62:52
Matt asks whether users should choose one skill library. Lauren says workflows are composable, encourages people to understand their tools, and recommends mining past chats for corrections that should become reusable processes or rules.
17 · 62:52 to 65:10

Recall turns repeated context recovery into a workflow

Lauren's recall skill came from debugging virtualization in Cursor. Each new chat needed context from earlier attempts, and manually explaining where to look became repetitive. Recall compresses that instruction into a reusable workflow that finds and carries forward relevant history.

She expects stronger models to need fewer low-level command recipes and more concise process descriptions. That does not cancel her earlier argument for deterministic tools. Commands still belong in scripts when exact execution matters. The skill can become shorter because it points to stable tools and states the decision process around them.

The conversation ends with a permissive model of skill design. Combine Wayfinder and poteto-mode, rewrite them, or extract a smaller method. The durable asset is the tested process behind the words.

The complete path is circular.Observe your own repeated intervention, encode it as a process or constraint, verify the result, and use the next failure to improve the environment again.
Transcript summary · 62:52 to 65:36
Lauren explains how repeated virtualization debugging led to recall, predicts more compact workflow-oriented skills, and encourages users to combine existing processes. The remaining seconds are thanks and sign-off.
Reference

Terms used in the conversation

TermMeaning in this guideFirst useful point
pstackLauren's skill and workflow collection. Automatic captions often render the name as PAC.03:40
VerificationGiving an agent tools to run the product, observe behavior, and compare evidence with an explicit goal.16:06
Hill climbingRepeatedly changing a system against a score or rubric and retaining improvements.17:40
DuneAn internal, closed framework Lauren describes as a convention-heavy base for Electron applications.29:40
Outer loopExternal context sources and routines that keep engineering projects current.35:10
Inner loopEngineering agents implementing and verifying work against a captured intent.35:35
GardeningMaintenance, cleanup, refactoring, and improvements to the working environment.46:00
Full autopilotLauren's pstack mode that runs intensive machine verification and can merge before human review.52:40
Method and provenance

Source notes

This guide is grounded in the complete 65 minute 36 second video and its English automatic captions. The source was reviewed from start to finish. The public page contains concise transcript summaries rather than a reproduced transcript. The untouched SRT and raw JSON remain in the local research record outside the public explainer directory.

Automatic captions repeatedly misheard names. The guide normalizes PAC to pstack, potato to poteto, several Grok Bot variants to the official spelling, and the language reference to the current Bend2 project. The source manifest records all section intervals and frame times.

Numbers about PRs, file sizes, coordinator count, daily sampling, and internal practices are attributed to Lauren because the interview supplies them as her account. All diagrams are original explainer reconstructions of the processes she describes. Editorial clarifications distinguish verification from reversibility and formal proof from unencoded requirements.