返回目录

LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX

Matt Pocock 请 poteto:每月上千条 PR 不是多开聊天,而是先把验证和环境做硬,再让外环把活送进内环

这一期在说什么

上千条 PR 是她先从回路里走开的结果,不是聊天窗口开得更多

本文综合。poteto 把两件事先做完,月度 PR 数字才出现。一件是厨房:验证 skill、确定性 CLI、lint 和内部框架 Dune,把“顺手的写法”收成“正确的写法”,agent 能自己看见结果。另一件是传菜:Slack、X、Linear、邮件不再由她当人肉代理,而是送进 Cursor projects 里的 coordinator。两件事接上以后,审查从每盘都尝变成抽样,autopilot 才能在她睡觉时合并。她同时说这些 PR 里大量是园艺,不是同等数量的功能。数字是这个结构的后果,不是结构本身。

一句话

poteto 向 Matt Pocock 解释她在 SpaceXAI 每月合并上千条 PR 的前提:先把验证和环境做硬,再让外环自己把活送进内环,人改成抽样而不是逐条拦住。

pstack · 信任阶梯 · 验证 · 米其林厨房 · 内环外环 · Grok Bot

Insight

阶梯的第一级不是多开 agent,是承认自己在当代理。

  1. 她说那是 2026 年 1 月或 2 月初,Cursor agents window 还没流行,大家还在终端里做自己的编排器。她做了 skill,但测不出影响,等于盲飞(约 00:02–00:04)。
  2. 她说 3 月加入 Cursor,4 月初做 agents window 的性能,看火焰图和 heap snapshot。烦的是自己夹在 agent 和 Chrome DevTools 中间(约 00:04–00:06)。

So, hello folks. I've got another treat for you today. Last time on this kind of podcasty thing, I suppose, we had Uncle Bob and we talked about software quality. We talked about agents. We talked about lots of cool stuff. Now, we have uh an incredible guest, someone who I'm delighted to welcome on, who's been exploding on Twitter recently about software factories, um about increasing the quality of your work, about

主持人 increasing your velocity and climbing the trust ladder with agents so that you can ship more and more and more. And it is potato. Welcome. Thank you so much for joining.

嘉宾 Thanks for having me. Yeah, very excited to be here. Yeah, big fan of yours

主持人 and a huge fan of yours. I think people have been talking about this like it's like the meeting of the skill minds, the skill Mount Olympus or something because both of us have very popular skill libraries. Um I've not, as I was saying before we started, I've not used a ton of yours and like I want to get all of the juice out of your brain so that I can go and use it properly and use it better. And I think where I want to start with this is you gave a talk um pretty recently

like um about 10 days ago and posted on X which went absolutely nuts as about how I shipped 2,500 PRs last month to production got about 3 million views or something on X and I watched it and I loved it and I recommended it and I kind of want to run this as almost like a Q&A of that talk basically of giving you because it just I just had tons of questions about it and I wanted to dive into it. And I think where I want to start is you talk about a trust ladder with agents wher

e you as you trust agents more, you can get them to do better and better things and or scale them to up to use more and more agents. So what is your story of how you climbed the trust ladder and how did that work when like you got SpaceX and

嘉宾 started climbing more and more?So I think this the the journey sort of began even before I joined cursor uh which is now SpaceX AI. Uh so the story is um after Meta so I I used to work at Meta on the React team. Uh I took a month off uh because I was feeling kind of burnt out and of course when what do you do when you're burnt out?

You go and start a new side project. Um and so I started a side project. you know, I was uh of course using AI to to write code. Uh but then I started to realize uh you know, I was spending like so many hours just micromanaging one agent, right?

And you know, at the time, this was back in February, maybe February, early February or January, you know, people were really obsessed with this idea of like orchestration. This was like, you know, before, you know, things like cursor, you know, like the agents window was had become popular. So people were still in like like 2 land you know in their terminal and they were all talking about okay here you know I built a custom orchestrator right and so of course I had I was a bit nerd sniped by that and you know as I was building my toy project uh I got nerd sniped by oh how do I make my AI coding setup more efficient and so you know I I kind of started the journey there where I just you know took a step back and realized you know I was spending all this time micromanaging a single agent you know I was creating skills and I was like finding it quite difficult to measure the output or the the result the impact of the skill as well so I was kind of flying blind but I was you know iterating really fast um and um so that project eventually sort of became the basis of PAC even though I didn't know it at the time um and a lot of some tricks I had learned like building that early set of skills. Actually, it's still open source if you want to if anybody wants to take a look. It's on my GitHub like potato noodle n o d l e. Um, and in there you will see some skills and a brain directory. And so I was really interested in this idea of how do I, you know, extract my own ability, if that makes sense, and give it to the agent, right?

cuz I was I I realized that you know all I was trying to do was trying to teach the agent to write code more like me you know do do you do do workflows more like me. So you know the skills were like an entry point to doing that. Um and then you know after I joined cursor uh I was starting to work on the agents window and uh it had a lot of performance issues. Uh it was it was it was pretty laggy. Uh and so since I had experience working in React, I was asked like, "Hey, do you want to come and help out uh with the agents window?

" Um and so the the this beginning of the cursor journey was very manual. Uh I was deep in like looking at like flame graphs and heap snapshots and trying to see like why exactly is the app so slow. Uh but then coming back to the same realization like you know I was sort of the bottleneck. I was doing everything manually. I was sort of the meat proxy in a way, right?I was the meat proxy between my agent and Chrome DevTools. Uh and I was like really annoyed by that.

主持人 And what month of the year is that? Let's say where are we in the timeline?

嘉宾 Uh so I joined Cursor in March. So this was like early early April probably early April is when you know uh I joined and I didn't have any skills, right?

I had I I sort of abandoned my personal skills because I didn't think they'd be relevant anymore. Uh but then working on the agents window uh and now working on grockbot uh I sort of realized that a lot of the lessons I had learned from those skill time building the the initial set of skills were very relevant especially around things like verification uh you know being very rigorous in your work um because I think from my experience even the the frontier ones tend to take shortcuts. Uh they tend to do the easy thing. Uh so uh a lot of the skills that I've built have been around how do I make the easy thing the right thing?

Insight

专长没有被模型替代;专长要先变成 agent 能接住的词。

  1. Matt 用 TDD 和 grilling 说明:词会被 agent 接住,写进思考,于是优先级变了。她补充 tautology 这个词能把“别写同义反复的废测试”压成一个短指令(约 00:08–00:10)。
  2. 她把这件事说成自然语言和编程语言碰到一起,底下都是沟通。Matt 用自己的戏剧学位接话(约 00:10–00:11)。

嘉宾 You know, how do I make that the best thing?

主持人 The idea of sort of distilling your expertise and turning what you do every day into processes, that's something that feels super familiar to me. That's exactly what I've been doing with the skills. And I suppose there's something in that which is a lot of people think domain expertise is getting less useful now as people uh start to rely more on AI where what do you think about that just as a sort of vibe check before we start talking

嘉宾 I actually feel like domain expertise is more important than ever you know uh I think I wrote this on my ex at some point but you know at times I sometimes think of you know AI as is like especially as the models get smarter and more capable and the frontier models are just getting so good like I love Opus 5.5 by the way um uh you know as the models get really really good it almost becomes like the bottleneck is no longer the agent right it becomes your ability to express your intent and your goals in a clear way that the agent can understand and actually carry out and That's why I think you know like people with a lot of domain expertise are extremely have a have a huge advantage in my opinion especially if you're a little bit like you know tech technoc curious you know so I I think of people like you know like uh like a doctor or a lawyer or you know someone who who has a deep expertise in a particular non-engineering domain and if they're actually just a little bit techsavvy and they can figure out how to use agents they can actually build really really great products, right?

If they if they have a clear enough vision in their head and they can articulate it in a way that the agent can build it, you know, I think that that is really the the bottleneck these days is is like the transfer of your intent, right, and your vision to the agent.

主持人 Yeah. I've been obsessed with language basically since agents um dropped. are just obsessed 100% and thinking constantly about the the composition of words, how I can make things sharper, what um what might be hidden in the phrases that I'm using. And it's and finding what I love is when you find a word that the agent then hooks on to and then goes, "Okay, I'm going to reinforce that word. I'm going to reuse that in my thinking traces." You know, I found that with um TDD was

an early example of that. a lot of chat about TDD recently of like, you know, people say, should you use TDD with agents?

Doesn't matter. What you're doing is you're getting the agent to think about TDD, getting it to write tests, getting it to prioritize things in a different way than it did before. And that's why sort of grilling, I think, works effectively. Grilling is a

嘉宾 Yeah. Yeah. It it draws those words out of you, right? or or at least it helps the agent understand your thinking so that they can propose those words to you and you can pick up and say yes exactly that.

主持人 Uh I've actually copied some of the the tips that you've shared as well where you know one of my favorite ones that you've shared recently or or not or like maybe in the past couple weeks is about uh reducing or eliminating tautological tests. Like one of my pet peeves of agents is like all of the useless tests that they write. And so, you know, that was one thing where, you know, the word tutology, right, is is is I guess, you know, not many people necessarily know that if if especially if English isn't your first language, but there's a lot of meaning to that word. And it's like it's almost like compressed, right?

Like you compress a lot of intent and meaning into words. And so I I I totally agree with you. I think language I've always been interested in language actually uh like programming languages natural human languages and how they came to be and it's so interesting that now with agents it's sort of like this meeting of natural language with programming language but it's all it's all language out of the hood it's all communication

嘉宾 totally I did a drama degree right so you know I've been thinking about language and Shakespeare and stuff for a long time and so this all feels very familiar

主持人 um so okay there's sort before we get into like because I think the thing I want from you is like software factory stuff, right? Software factory is the big buzzword. Software factory is the thing that I'm thinking about too. I'm sort of releasing a course in that direction too.

嘉宾 And it's this sort of scaling yourself up to un unrealistic numbers of PRs basically or PR numbers that sound ridiculous to people who don't understand how this works. So where I want to get to is sort of from people who are doing kind of like one to five agents today up to, you know, hundreds of agents running at once and how that sort of functions. And so I'd love to hear about your metaphor of the Michelin Kitchen instead of the software factory because I think that says a

Insight

规模来自厨房的布置,不是来自把更多人推进同一口锅。

  1. 她把主厨比成厨房的 tech lead,甚至 CEO。对应到工程上,人不再写每一行,但名字和声誉仍在最终结果上(约 00:13–00:15)。
  2. Matt 把这收成:与其改模型或 harness 内部,不如改 agent 所在的环境和它用来自证的工具(约 00:15–00:16)。

嘉宾 bit about the way you think about this stuff.

主持人 Yes. Yeah. I I I've I've never really liked the term software factory. Not because you know it's not accurate but I think I think the a lot of people when they think factory right they don't necessarily equate that with quality or craft right things which are very important to me and a lot of people and technologists who work you know building products we care about the user experience we care about the things we're building. So while so while I do think software factory is an apt term, it also I guess maybe conjures up negative, you know, maybe sometimes negative connotations. So Michelin Kitchen is the thing that I've sort of landed on where it's much more I feel like it's much more aspirational and uh I like the metaphor a lot cuz you know I like food. I'm called potato of course and I like cooking and I see a lot of parallels right like with food right when you're cooking a meal for yourself for example it's both utilitarian like you're trying to just feed yourself right and and survive uh but it can actually be transformed into art right and that's what what a Michelin starred chef or even just a chef or a cook can do with food is take something very ordinary and turn it into a delicious meal that you know takes you back to your childhood days or something like that. Um and so it almost like mirrors that trust letter that I talk about where uh you can sort of imagine your own journey as a home cook, right?

Uh as a home cook, you are doing all of the food, the cooking yourself. You cut all the vegetables, you do all the prep work, you do all the cleanup, you know, you are the one man or one woman show really. Um, and it's an interesting thought experiment like, okay, if you were to cook a meal and then you add people, right, your your your partner trying to your brother, your sister, and now suddenly you have your whole family in the kitchen. I think most people would get very stressed by that, right?

The thought of, oh, so many people are just mocking around in my kitchen. They have no no idea where all the utensils are.

嘉宾 I have a max capacity of one person in the kitchen. Yeah, absolutely. So, I feel like that that's really apt because when you ask yourself that question of how do I go from being a solo cook, right, to having an army or even not not even an army but a few sue chefs, right, that that are helping me in the kitchen. How do I think about dividing the work in a way that makes sense?

You know, I'm not dividing work just for the sake of it, but in a way that actually makes the sum the to the better than, you know, the total of its parts. And so the Michelin kitchen metaphor to me like works really well in that regard because you know as a chef you're you know if you become a chef you're in a position where you're not necessarily cooking all the food yourself anymore but you are thinking you're almost like the tech lead right for the kitchen where uh you know chefs have to think about you know not just cooking but they have to basically organize the whole kitchen and they're like the CEO of the kitchen they have to think about when do you order ingredients,how do you store them, how do you prepare them, when do they have to be prepared, you know, it's a whole job, right?

That's not just cooking. Um, and I think that again it mirrors so much of how engineers write code today where you are not writing the code yourself anymore. You have agents, right?But you as the human are still responsible for the final outcome, right?

your name still is associated with the work that you do, your reputation and you know so how you set up your kitchen right and how you set up your skills your environment your codebase I think are ultimately the new ingredients that go into um building product

主持人 yeah I think what I love about your approach is the amount of focus that you put into the environment that the agent operates in right because I think a lot of people they think, right, the agent is good. I'm probably not going to be able to make it better. Let's just trust what these magic model people have put into the harness and the model combination. Uh, there's nothing I can really do, like I can't mess about with claw codes internals or something or whatever you're using. Um, but what I love about your approach, and it's something I advocate for too, is that you can change the environment the agent operates in, right?

you can make changes in the codebase and also give it tools for verification as well and allow it to verify its own work. So the thing I I loved about watching that talk is the amount of focus you put in verification and like that is the lever that you can start to generate trust. Can you talk about that and what that concretely looks like?Let's start like looking at practical ways that people can improve their own processes, their own kitchens.

Insight

循环能闭上,是因为验证在;CLI 是为了不把机械步骤交给临场发挥。

  1. 第一个让她往上爬的用途,是 Cursor agents window 的性能。有了评分和循环,才能做 hill climbing。她把 Andrej Karpathy 的 autoresearch 算进同一类想法(约 00:17–00:19)。
  2. 她说 Cursor / SpaceXAI 的每个应用现在都有自动维护的验证 skill,已经是团队的关键设施(约 00:19)。
  3. 验证 CLI 起初是为了省上下文:当时压缩一次,后半段会话就会变笨。现在她更强调把确定性编码成脚本,让 agent 不必每次重建世界。迁移则用 codemod 走语法树,而不是让模型临场发明(约 00:19–00:25)。

主持人 Yeah, I've I've I've said this a lot actually that you know even if you don't use PAC or you know your skills I think that the single most important skill that should be in your toolkit is verification because without verification and for for by the way for those watching who don't know what that means it's this idea that you can give you can sort of give your agent uh hands and eyes in a way that's the the analogy I where the agent is able to run the code, right? And actually

嘉宾 uh interact with it like a normal human user would and also do things like you know debug it, you know, take traces and snapshots. Um and uh funnily enough like that was actually the first skill I built when I joined Cursor. uh that gave me a lot of that was that was the thing that actually started to let me ascend the trust ladder a little bit in a way that some of the other skills I had looked at or built had not really let me do because no matter how good you know some of

the other skills were like the how skill, the why skill, the unsop skill were, I was still relying on me right as the proxy between my agent and the output. So that you know if the agent can't actually see the result of its work there's no way it can actually iterate right and so this is where people start to talk about loops this idea of a loop and really I think the term loop you know seems kind of uh almost abstract like people like what what is a loop what is an agent loo

p but really to me like the most important part of a loop that allows it to be a loop is the verification part because the agent is able to to verify by its own work and uh you know that takes you out of the equation where now I can actually do something like so the very one of the very first use cases I had for verification was you know like the performance work that I was doing on cursors agent window and I want I wanted to get to a point where I could do something called h

ill climbing uh which is a term that I I think the labs uh talk about a lot which is this idea that you know you have some kind of rubric or a way to judge or score something And now because you have a loop, you can have an agent continually try to make improvements to that. Uh I think Carpathy, Andre Carpathy also famously released uh something called auto research that has a lot of these ideas. Um but yeah, verification I would say is probably the most important skill in PA

C uh and many other you know tool sets. Uh and I think it's the most important thing to focus on. So a lot of the a lot of I spent a lot of time actually you know tuning the verification the creative verification skill um and also internally the the we have so many verification skills now like every app that cursor has or spaceexai has has a uh verification skill that is automaintained as well

主持人 uh and it's become critical infrastructure for our team because everybody uses it

嘉宾 and you went pretty far with that too right like you had a um in your talk I saw that you actually built a custom CLI for that too. So what does that CLI do? Like how does it execute things and why did you I mean that's proof of how deep you're going right of how much you're pushing that.

主持人 Yeah. So this is actually a tip I learned early on where um I guess you know back in January or or late last year the thing that people were concerned about was context window, right?

That was the big the big topic at the time was how do I you know manage the context window because you know compaction summarization wasn't really that good yet and people were always people had this there was almost this meme in the community that you know once your agent summarized or compacted once it would become sort of stupid right for the rest of your session. So there was a lot of thinking around like you know being very efficient with your context usage and so that was actually the inspiration for some of the uh the CLI work inside of the verification skills. I guess now it's less so about context because uh you know agents are much better or harnesses have gotten a lot better with summarization.

Um, I still think there's some benefits to, you know, uh, having a clean context window. Uh, so the CLI is really just more of a way for me to take the deterministic parts of what the skill does and encode that into a script or CLI to reduce to kind of take away the judgment that would otherwise unnecessarily be used because with judgment so I also think of you know agents and skills in sort of like it's like a gradient you have some parts of the work that are entirely judge measurement based right you know something that requires thought you know putting together multiple pieces of context thinking um and then you have the more deterministic parts like I don't know if you wanted to uh refactor some code right from one pattern to another that's very mechanical right you don't you don't need an agent to think about it and come up with it in a novel way each time right and so that that was really the inspiration for the CLI and you'll see this in a lot of the other skills that I built is like I try to extract out the deterministic parts and turn that into code and just leave only the parts that actually require judgment to the agent. So in a way I think of the seal as kind of like your wrapper, right?

It's a wrapper with some light instructions around how to use these custom tools that are inside of the skill. Um but yeah, I don't think the CL is really that interesting in its own really. It's not like a novel piece of software. It's just something that interacts with like Playright and the Chrome DevTools protocol and calls a bunch of APIs. It's like it's just a bunch of glue.

嘉宾 No, it's fascinating because it's a way of hiding information from the skill, right? It's a way of conserving the skill, keeping the skill quite small, I imagine, and then you're able to delegate more of the complicated deterministic stuff into a script within the skill. So it's almost you're compressing information and making the agent do more consistent things more consistently.

主持人 Yeah,

嘉宾 that's fascinating. And it it also helps I guess if you care about context window it it does help because now the agent doesn't need to uh you know re reinvent uh things cuz uh one thing I had noticed early on when we didn't have a CLI was that uh well the the agent would try to verify it work but it would basically rebuild the world each time and then every agent did it differently and I was starting to notice like that's very inefficient right I was wasting it it was actual

ly not just about context usage but also speed, right?

Like because now an agent had to actually go off and write the scripts or the CLI and test it and you know and it doesn't work and the last agent did it and it worked but it discarded it. So it was just very obvious at that point like I should just turn this into a CLI and put that inside of the skill uh so that every agent that uses it now benefits from that same piece. Um but I I also think like you know it's a good push for people to think about is how much of your skills

and rules could actually be deterministic.

Um that's like another core thing or or one of my core principles that I like to think about is yeah how do I uh make very efficient use of determinism and non-determinism and you know let Asians shine at the non-deterministic parts right because that's what they're trained to do. Um and the other parts which are much more mechanical or you know straightforward can be just pure determinism. Um, and you'll see this as well for things like doing migrations. Um, which is another big thing that I've I've talked about is, you know, going from one technology to another, especially one that is better for agents, right?

And a lot of how you can do that migration is, I think, through things like scripts and CLIs, like the deterministic parts like code mods, you know, like crawling the abstract syntax tree and transforming code literally mechanically, right?like a script does it for you instead of the agent.

主持人 Totally makes sense. I I mean I think what there's another thing there which is you're taking stuff away from the agent and you're kind of putting it in the environment too a little bit which is let's say you have a a thing that you notice the agent always gets wrong. You want to make that um just impossible within the environment. And that sort of comes down to code quality as well. I mean, I talk about a lot like having a what a good codebase means, right?

What is a good codebase?And there's a definition I like which is a a good codebase is a codebase that's easy to make changes in, right?Easy to um change stuff without things screwing up. And that means that you have a lot of guard rails that you have a lot of um the agent or the human is constrained to very narrow paths. And that's again something you talk about in your talk.

嘉宾 And you talk about this not only on the kind of sort of automated checks side of things. So linting and type checking blah blah blah but also in the way you design abstractions. And you guys even I think built a framework uh for your agent to work in too.

新工作是把环境收窄

00:26:00–00:33:00

Insight

规则放进环境,agent 是撞上它,而不是被要求记住它。

  1. Dune 未开源。她把它描述成 Electron 应用的内部 Next.js:feature 有固定目录,注册表去扫代码库,lint 很紧,很难写出坏代码,也不再往 god file 上追加(约 00:29–00:31)。
  2. 她说 Grok Bot 最初几版大约是八个 god file,每个至少一万行。每次看见 agent 失败,就问能不能变成 lint,让代码库使这种错误无法出现(约 00:31–00:32)。

主持人 I think what I'd love to hear is you obviously think of that as very important, right?

And that's how important is that compared to other things you could be doing like building features or shipping work. Yeah, I think that's a um I almost feel like the new job of the engineer is really to to spend time on the environment. Um I almost actually wrote a tweet about this yesterday, but I but I didn't. But I think that you you know if you if you haven't really spent time, you know, building trust in your agents and building skills and tools, you can get stuck in this mode where you're very low on that trust ladder, right?

you don't have a lot of trust in your agents work. And so the only way to cope in that when you're in that situation is just to kind of lock in and micromanage your agents. And that's very time consuming. And when you're stuck in that mode, you don't really have the luxury to think about, you know, uh higher level things like like making yourself more productive. In the same way that uh I guess analogy would be like if you've never taken the time to learn like your tools right as a developer when you were writing code yourself and you know you've never heard of VS Code, you've never heard of Vim, you only knew about Notepad uh and you had hadn't even heard about Git. That's sort of the analogy. It's like you you haven't spent the time sharpening your own knives, right?

And so, of course, if you have a dull knife, then everything's going to take a long time. Um, and you're going to be you're just going to be and and especially if you know deadlines are looming, then you don't have the now you're stuck in this rut, right?Where where you you you don't have sharp knives, you don't have good tools, but you're under all this pressure to ship, right?

And so, all you can do is just focus on that. But I do think that, you know, if you can find yourself the time to actually spend time thinking about your setup, it's again going back to the cooking, you know, analogy, it's like uh, you know, if you, for example, if if cutting cutting cutting the garlic is like super slow, right?There are garlic mashers, right?You can buy and you put it in the thing and you like squeeze it out, right?

It's super fast. Uh, machines and tools were invented for a reason, right?

And so if you're operating a Michelin kitchen and your your your cooks had no tools, then of course everything's going to be extremely inefficient, very very, you know, every every every cook is going to make something up of their own. So I think the tools and the determinism to me are you know taking that part away and and just like you said about constraints as well. It's the constraints are are to me as well like uh actually a slight tangent on that is uh I think we should talk about TypeScript cuz like we we actually both share like a background in Typescript where you know you obviously have done a lot of work with TypeScript and total TypeScript and you know you're a leader in that space and I uh had adopted TypeScript pretty early and I had given like a talk or two at Typescript conf uh many years ago and so one of the the talk that I did actually was about type systems and constraining the constraining types. Like one of my most favorite things about Typescript is actually type narrowing, right?

This idea that you go from a very broad type,right?That could be anything and then you through type guards and you know type narrowing and you know runtime checks you can actually narrow the space and say like oh this isn't just a string this is a very special type of string. It's a constant, right?

like I but I I determine that through the type system and in a way it's like uh there's a lot of parallels I think to that with constraints in your codebase where is it you're you're constraining the space right if you if you think about category theory as well you know you're constraining the the number of possible types right that can can exist and you're saying there's only one type right and for for us like that framework that I'm called Dune. Uh it's not an open source framework. It's the the way I describe it to people. It's it's kind of like a internal Nex.js for our Electron apps. Uh but it comes with a lot of really really restrictive lit rules and the codebase is designed in a way that there's really only one way to do something. So we make use a we make use of a lot of conventional patterns. So like features all go into a specific directory. Well, every feature has its own directory. As an example, you know, there's like a a thing that discovers features like through a registry and like crawling the codebase and stuff like that. But this conventional pattern and the lint rules make for an environment where it's actually very hard to write bad code. And that sort of frees up the it both frees up your own mental uh you know capacity as well as the agent sort of doesn't have to think about that anymore where it's just like oh there's only there's I should just if I want to add a new feature it just goes in the feature the new feature directory and all the code goes in there and I'm not going to append to a god file right that was really actually the inspiration for those feature directories is the very first couple of versions of Grockbot were composed of like eight god files which were like at least 10,000 lines long if not longer and so I kind of had to break it up into smaller pieces. Uh but it was just observing you know actually that's another important part is observing how agents fail and then every time you see a mistake every time you see something that could be done better you think you step back and think how do I turn this into a lint rule?

How do I make it so that the code base makes this impossible?Right?And it comes back to me for my you know my background learning Typescript and types uh type systems is how do I constrain the space so that you know I know precisely what I'm working with and I think yeah there's a lot of parallels there.

嘉宾 Totally makes sense. And don't I mean it's funny that you mentioned TypeScript and Goth files in the same sentence because Typescript famously has a 25,000line type uh file.

主持人 Although I don't know if they've rewritten that and go as they probably have, haven't they?

Um, okay. So, environment is important. You should watch your agent like a hawk to make sure that any mistakes it makes. You turn them into things in the environment. And the benefit of the environment is you're not overloading your agent, right, in terms of rules, in terms of things it has to remember. It's just in the environment. And so it stumbles into the rules and exactly um you know bounces off them and hits them at the right moment.

Insight

先问这一步为什么还要我,再决定是教 agent 自己取数据,还是继续当传话人。

  1. 她用 Grok Bot 做外环:订阅 Slack、X、邮件、Linear,发现 Grok Bot 桌面端的缺陷就送进 Cursor project。project 是云端 coordinator,有自己的电脑,自己不干活,只分派和传递上下文(约 00:40–00:44)。
  2. 三十个问题一起到,不应该一条一个 agent。coordinator 要能换拓扑,把相关的性能问题放在同一条线上,避免重复劳动,也避免看错层级(约 00:43–00:46)。
  3. 她把 Netflix 时期听到的 context, not control 直接挪过来:用控制可以逼出结果,但目标是把上下文交给对方,让它自己够到(约 00:41–00:42)。

嘉宾 So okay, we still haven't talked about the 2,500 PRs. Where do those come from?Like how do you you've built your trust ladder, you've worked on your environment, and you understand, okay, um I now want to scale up. So what are the mechanics of that scaling?Are you um initiating 2500 like chats per month?That can't be right. So there must be are there any kind of automated triggers that trigger stuff in your repo?

Like how do you get the software factory kind of triggering work by itself?

主持人 Right. Um I'll definitely say that the prerequisite to you know something like a very high volume of of pull requests um is the environment. you know, the the stuff we just talked about where I definitely would not have been able to do this if I had not spent the time, you know, thinking about the kitchen, right, and the knives and the tools for my Asians. And so, in a way, I I think of this as I've spent the time building one kitchen and one restaurant. And now I I'm in a position where I don't actually have to be there anymore because the environment, you know, that the same analogy, right?

It works really well. Yeah. You open chain of restaurants, right?That's

嘉宾 Yeah. Exactly. Yeah. Exactly. It's like you're Gordon Ramsay and you know, you you've taught your executive chef like all the tricks of the of coming up with great menu. Uh and like the kitchen is set up really well. Everything's just perfect and you're now in a position where you can open your second your third restaurant. And I guess I I sort of see each project that I work on, like each big chat is sort of like a restaurant, right?

and I'm I'm I have multiple of them operating at the same time and I'm sort of like helicoptering between them sometimes some more than others depending on how in the loop I am but yeah definitely I think there's there are external triggers and context that those projects don't have that for a long time I was the proxy for that so uh the best example I have is like you know you have a project that's working on a feature uh or you're trying to fix a bug and you're getting bug reports, but the bug reports are going to things like Slack or linear or X, right?

And these are external systems that aren't connected to your inner loop. So, I like to talk about this outer loop and the inner loop. Uh I don't know if I'm using the definition correctly but to me my inner loop is like basically my engineers my agent engineers working on the code to an building towards an intent or snapshot of my intent right and the thing about that is that the snapshot can go stale right new information comes to light that I then have to be the proxy of and you know transfer that context to my agent so you know if you don't have these triggers pulling information back into the interloop, then you sort of have to play that role where you're you're off, you know, in Slack or X or or whatever and you're gathering context, right?

You're getting context about bug reports, about feature requests, about, you know, something someone said about, you know, our backend infrastructure has some limitation, you know, all that information, you have to f that across to your agent. So that's where I think like tools like Grogbot are really good because they help you automate the outer loop as well. And when you connect those two loops, it's very very powerful because now all of a sudden your agents have the ability to get context for this for themselves, right?

If for example uh you know either through just as a simple example like maybe you have the Slack MCP, right?Or you have uh your own harness, right, that you've built a Slack subscription into for a particular Slack channel. Now all of a sudden you can tell your agents, okay, subscribe to the Slack channel. Every time there's a uh, you know, bug report about something, go off and triage that thing, right?Go reproduce the issue, right?

Using the verification skills that we've already spent time building and all of those other skills that we've set up so that I have a lot of trust, right?

I have a lot of trust that these agents can actually go off and understand the bug, you know, uh verify that the bug actually still exists on main and it wasn't something about you maybe the users setup or their data or maybe I don't know they didn't install a dependency or something like that like uh basically I think uh creating that yeah creating those two loops and connecting them is really a very important part of the job these days. Um, especially if you are thinking about how to scale yourself. So, a big theme here is really just like always thinking about like what where am I the bottleneck in this process?

Why do my agents need me, you know, to answer this question?I I always like to think about that. And so I try to think about how do I actually get the agent to answer its own question, right?

But not by hallucinating, not by guessing, but actually real data. And you know, a lot of people talk about this idea of a company brain, right, or a context graph. I feel like those terms are unnecessarily complex uh or even abstract. To me, it's just about um how do I take information that my agent needs that I would otherwise have to go and pass it myself and just teach it how to do it, right?

And that removes me from the equation. And so how I arrive at 2,000 or however many PRs is the fact that I have all these loops set up, right?And so uh it allows me to open chain restaurants, right?I can I can really parallels myself. So yeah, I'm not sitting there creating 2,500 chats, right?

Of course, it's really like these projects are um actually cursor has a new feature called projects which are these like coordinator agents. Um and so the coordination co coordinator agents are really good at sort of delegating and not doing work of their own but they manage and supervise like almost a list of tasks and they spawn sub agents to go and do them. And so I'm just constantly feeding context or teaching the agents how to get their own context and then they're going off and doing the work for me. Uh and really the big the last thing I'll say to this is like the big unlock for me for getting to 2,000 PRs is starting from the question and working backwards of how do I get to the point where my agent can merge its own code?

Because the obvious thing people ask me when they when I tell them, "Oh, I shipped 2,000 and 2,500 pull requests last month." They'll be like, "How did you review that?" Right?That that's a lot of PRs to review. Like your team must hate you.

主持人 Do do you mind if we go there in a second? Because a good question about that.

嘉宾 Yeah. Yeah. Yeah.

主持人 I want to like this analogy is great. I want to like deepen it a bit which is before if you're like manually initiating all those chats it's like you're bringing the orders to your chefs manually right whereas if you've got an agent sort of like doing the expo then you're able to sort of run it yourself itself what is what does that concretely look like then you've got these sort of grock bots that are um subscribing to channels pulling in Slack messages and you it sounds lik

e have a couple of coordinator agents or like chief of staff agents that like monitor that or something like when you look at your computer to manage your agents, what does it look like?

嘉宾 Yeah. So, so uh this is I guess somewhat confusing but we're working on you know simplifying and unifying but so uh there's graphbot uh which or you know you can use other tools of course as well but I I largely think of these tools as like your outer loop. These are tools like you know Grabbot that have connectors right these are connectors I guess they a lot of people call them personal agents um but they're connectors to things like your email your calendar slack uh plaid I don't know like all these different services and they are a great source of pulling context in to your work so the same way that a human like you know if I were if I was a manager and I was leading a team of engineers years. Um, you know, like when I used to work in Netflix, one of the biggest things that managers would talk about was this idea of context not control, which funnily enough, you know, has so much uh has so much uh carry over to the agents world. Uh, of you know, you you know, you you of course can drive to an outcome you want by control, right?

Like by micromanaging, but what you want is to provide context instead, right?like teach the agent, teach your engineers how to be self-sufficient and then you don't have to micromanage them.

主持人 Um, and so I see a lot of parallels there. Uh, but yeah, graphbot. So, concretely, I have some graph bots that look at my Slack channels, look at my X, uh, or my emails, uh, or linear, and they're just constantly they have routines that subscribe. So they're constantly watching and I have I I'll tell them things like you know uh I'll watch for issues with uh bugs in the graphbot desktop app as an example. Uh and whenever you find that send it to my cursor project. So one of the really cool things about grabbot is it connects to cursor. So cursor has uh like I I just mentioned this new feature called projects. And a project is really a uh again like a you get a coordinator agent that's in the cloud. It has its own computer and all it really does is like it's a manager of agents. It's like your executive chef, right?

Your your chief of staff. It doesn't do the work itself. It delegates and orchestrates and manages the work of other sub agents to you know that report to your chief your chief uh of staff. And it basically is responsible for driving the work forward and managing things and uh passing context to them.

嘉宾 So if you get a sudden burst of issues, let's say you get 30 issues at once in one payload or something or very quickly the coordinator agent can figure it out and delegate.

主持人 Yeah, exactly. It gets like uh you know 30 the 30 or so payloads and spawns a sub agent or a single coordinator agent. It can actually do a bunch of different topologies of agents and it will sort of figure out the best way to uh you know efficiently distribute the tasks to your team of agents. Um so I use uh cursor projects a lot um and I also use grapot a lot and cursor projects are my inner loop and grabbot is my outer loop. Grabbot takes all the context, external context, gives it to the projects because it can actually just send messages to those projects, right?

You don't even have to open cursor. You can just tell your grabbot, okay, create a project, right, for these series of tasks. They're all related, right?Maybe as an example, you know, you've had a uh a big burst of issues that are all about performance, right?Your app is slow uh and they're all connected, right?Maybe some of them even have a similar fix, right?

But and you can certainly go off and just spawn one agent per task, but then you've lost that sort of thread between them, right?

And and you may duplicate work or you may not really think about the higher level problem. You know, sometimes when you you you solve bugs, you know, it helps to have multiple bug reports that are are slightly different because it helps you really, you know, zoom out and see actually, you know, the problem when I looked at this one report, I thought the bug was here, but actually when when I see the other multitude of bugs is actually up here,

嘉宾 right?

主持人 Yeah. Got you. So that that's why you have so many agents in that loop then,right? Because it's not just you have um like you have a bug report comes in, you spawn a single agent to look at that bug report. that a that single agent will be duplicating work with other um other agents, right? Because if there are multiple bug reports coming in through the same thing, that can be duplicated work.

嘉宾 Yeah.

主持人 Fascinating.

嘉宾 That's really fascinating. Okay. And so this just this endless series of triggers um coming from real users reporting real reports um builds up this sort of and accelerates the factory sort of adds more orders in. Other than bug reports, are there any other sources that you use for like um accelerating for pushing these PRs?

主持人 Uh well, funnily enough, it's some of it comes from uh reading the code, too. So, I guess I have sort of uh well, so to clarify that, you know, the 2,500 PRs, they're not obviously like 2,500 features, right?

Insight

重复两次的错误属于厨房,不属于那一个 agent。

  1. 缓冲是为了不被执行速度带走。幕僚长式的 agent 要能看见整片,而不只是清掉队列(约 00:47–00:49)。
  2. 她不会每条都尝。每天抽样看 PR 和代码,同一类抄近路、同一种权宜之计扩散开,就改 skill、lint 和类型,而不是只纠正那一个 agent(约 00:49–00:51)。
  3. 她明确说这不是装上 pstack 就会发生,要花时间看 agent 在哪里失败(约 00:51)。

主持人 they are a lot of the work actually is spent on gardening like another term that I really love. Uh where so I guess this is more important when you have a big team of engineers human engineers that you work with where and also this goes back a little bit to what I was talking about with the environment. You know, setting up a really good environment that doesn't just help you and your agents, but everybody on your team, right?

Think of a new hire who doesn't have a lot of context on all of your engineering practices joining your your team. And if you have a really good environment, they can be productive from day one, right?They can they don't have to like, you know, make open a bunch of lowquality PRs. They can start, you know, they can start just turning out really good code. Um, and uh, yeah, I think I I sort of lost my train of thought.

嘉宾 I've got a I've got a followup, which is what's what are the mechanics of like triggering

主持人 like how when when do you trigger a a agent to go and look at the code, right? Because some people might say, "Oh, let's just do that every hour or something or like on a chron job or what's

嘉宾 Oh, yeah. Yeah. Yeah. Yeah. Uh I saw some of your recent tweets as well about like you know the some of the tweets you've been doing which are great for setting up your routines. Uh I have some routines like that as well. Um so uh one of them is uh like looking through just another simple example is you know React has a lot of foot guns. Um, so, uh, as as I'm sure you're aware. And so I have an agent that's just constantly looking for band patterns. And the interesting thing

about that one is that I don't actually tell it to fix the issue first. I tell it to append it to a document. And then every couple of days I look at it and I see actually these are all the same thing, you know, and so that gives me, you know, you almost want like a buffer, a queue. Sometimes that's actually more effective than just spawning off a couple of like a lot of sub agents to fix every single thing because when you are in kind of pure execution mode and just trying t

o like you know f uh yo u know execute on the orders that are coming in very fast you sometimes miss the big picture. So sometimes having a buffer forces you to think about the big picture because you you you have these artifacts and things that you can look at as a human um and sort of use your own human judgment to or I guess you can use an agent to do that as well. But you give the agent and yourself a way to identify patterns,right, that you might otherwise miss if you're

just only solving each bug at a time. And that's also really the benefit of having something like a chief of staff agent is uh it can see the forest right uh in addition to actually doing the execution.

主持人 Fascinating. That's I mean my brain is exploding a bit there with the sort of chief of staff at the software factory. I might have to change some of the course that I'm filming next week.

That's

嘉宾 uh all right. Let's talk about let's talk about review, right? because this is the reply that you get, you know, is

主持人 did you read did you taste all 2500 of those dishes as they swept past you?

嘉宾 And I assume the answer is a variety is a version of no.

主持人 Yeah, I think you you you don't want to be in a position where you're not tasting your food ever again. Uh but you also, you know, for scale, you cannot be tasting every single dish that comes out of your kitchen, especially if you have multiple restaurants. So it becomes more about sampling right and thinking about the processes in the same way that you know if I guess maybe this is where the the factory analogy is a bit more apt is you know as a quality supervisor on a factory you you can't look at every single item you sample right you take you you you go in there every day and you look at the quality of the pull requests you look at the code that the agents are writing and you scrutinize it very rig rigorously and you think about all the inefficiencies, the bad patterns that the agents are doing and then you think about how to course correct the environment, right?

Not not that single agent. Uh because if maybe if it if it was a one-off incident, it's fine. You know that maybe there's nothing to fix there. But if you actually notice that multiple agents are are having the same issue, right?They're taking the same shortcut. they're they're propagating the same workaround everywhere. Uh that's a sign that you should go off and think about how to uh amend your kitchen or your factory, right?

Like thinking about your skills, your constraints,your lints, your type systems um and setting or adjusting it so that that problem doesn't happen again. And when you do that enough times, then you get to a place where the codebase is again like the environment is so constrained and so it guides you so well that you can just you can just step away, right?

That's the dream. And I'll I'll definitely say it um it's very hard to get to this point. I don't want to sell this as like, you know, something that you can just do easily by using PAC. Like it takes a lot of time and effort to think about your code and where you see your agents failing and thinking very thoughtfully, intentionally and setting up guard rails and constraints so that they do the right thing by default.

嘉宾 And you're not like if to go back to the software factory analogy, this isn't a dark factory, right? This the lights are on, right?

主持人 It kind of is. Yeah, actually.

嘉宾 Is it?

主持人 Yeah. Well, it's dark in the sense that so um it's dark in the sense that well I think if my agents are merging their own pull requests it's sort of become dark where I go to sleep my agents now work 24/7 uh I have I have the equivalent of like more than 10 chiefs of staff right each working on a different area like for example I have one that's working on performance of the Grockbot desktop app I have one that's working on uh fixing bugs that users report I have one that's exploring rewriting it in a different language just for fun, you know, like what if what if, you know, just reimagining what what it would be if it was like a native app. It's just a toy. Um, but the idea is like yeah, I uh I when you spend the time setting up your environment, I've gotten to a point where I review the pull request after it's landed, right?

I I tell my agents full autopilot is is something that you can do in in PAC and that will trigger off this very intense rigorous verification loop where it will spawn a bunch of verifier agents for every pull request and it will fuzz right fuzzing meaning that it will actually run the application. It's going to click around and try to use it like a real human.

look for regressions, look for bugs in your implementation and um it will try to find issues with the thing and then it will fix it itself. It'll do that again and eventually get the PR to a state where it can land. Uh so it does it does it is quite token intensive. You can tune this of course. Uh so you know instead of like 10 verifier agents you might do like one, right?

Or you just tell the agent to verify it's done work. But yeah, the key thing is the verification part is really the key piece that gives me a lot of confidence that I guess verification plus the environment, right?It's these the combination of these two things that allow me to step away and say agents go off and merge your thing. I'll review it in the morning by looking at my commit history

嘉宾 and if I see problems, I go and course correct,

主持人 right? And I'll go and revert or modify, add new link rules and whatever. Um, so it does it does take time to get to that point, but once you get it, oh, it's so it feels so magical. Uh, I I I tell people like I'm sleeping so much better now because, you know, it took it the very first day I turned on the sort of dark factory was very scary because I was like, "Ooh, what if I call the SE, right? What if I break something overnight?"

嘉宾 Uh, and it took a lot of it took a lot of uh bravery, I think, to do that, but

主持人 somehow I did it. And yeah, now I'm in a place where my Asians are are merging their own code while I sleep.

嘉宾 It sounds like

主持人 I think it's dark in that sense.

嘉宾 Yes, it's dark sometimes, right? You do sometimes.

主持人 That's true. That's true.

嘉宾 Because I think of a dark factory is like almost like if you take the original definition of Kapathy's vibe coding, right, which is the code almost doesn't exist. You forget that code might be a thing. I think your approach is totally different from that, which is that code and the environment is essential. And if the code in the environment are bad, then you will get bad outputs. Garbage in, garbage out. So I I I think this is a this is a different thing. It's like, you know, the I don't know, maybe there's a dimmer switch or something, right?

Insight

敢走开,是因为验证加环境已经替她挡了一层;挡不住的领域,这期没有给出替代流程。

  1. 第一晚她说很害怕,怕半夜弄坏东西。她也说 verifier 数量可调,十个换成一个,或者只让 agent 验证自己做完,但很费 token(约 00:53–00:55)。
  2. Matt 认为这和 Karpathy 那种“代码几乎不存在”的 vibe coding 不是一回事:环境和代码差,产出就差。她把厨房类比留到这里:老板仍会回店里尝一口(约 00:54–00:56)。
  3. 单向门如果不能程序化验证,她认为到不了这一步。她的预测是会出现更面向 agent、把程序和证明写在一起的语言;她举了 Bend,并回忆以前要另用 Lean 或 TLA+。证明能过,她会问为什么还不合并。她同时说不是所有领域都可验证,这个问题她没有答案(约 00:56–01:00)。
  4. 收尾她把 skill 定义成过程,不是实现细节。模型更强以后,精确的脚本命令可以从 skill 里删掉,只留步骤。pstack 的 recall 来自 Cursor 虚拟化那段:每次新对话都想把上一份 transcript 带过来。她建议让 agent 扫自己的历史干预,提 lint 或新 skill。Wayfinder 和 poteto mode 可以自己拼(约 01:00–01:05)。

嘉宾 Like, you know, some parts of dark, some parts were light. This is why the maybe the kitchen is a better analogy

主持人 the restaurant because you know even as a as a as a restaurant you still might go to your restaurants every now and then to take a peek in taste the food right

嘉宾 uh I think that's

主持人 the idea of sampling instead of blocking I think is really important

嘉宾 I think what would you say to people who are in I guess you're obviously in a pretty security conscious environment where you're working very security conscious

主持人 mh Um maybe there are folks working in like um medical applications or law or finance or something. I think of the like some PRs are kind of like two-way doors which is you can merge it and then revert it, right? It's cheap back through. But there are some PRs that are one-way doors, right? That will cause data loss of some kind that will

嘉宾 do something that can't be easily walked back. How do you deal with situations where most of your PRs, let's say, are one-way doors? Like, is this something you just wouldn't recommend or like what do you think?

主持人 Yeah, I think that's a really good question. I think that it all comes back to me to the quality of the verification that you're able to um get out of your agent. And I think for domains where the work is verifiable, this is easier, right?

and the the oneway doors become two-way doors in a sense. But I guess I don't know if you're working on something that is like extremely is very hard to verify programmatically then I think yeah you're definitely in a position where it's very hard to get to that point. Um so I do think like yeah verifiability of the domain is an important aspect to be able to do this. Um and software engineering is just one of those things where it's quite verifiable in in a lot of cases maybe not totally um you know like other domains like mathematics I think are another example of not all of it of course but some aspects of mathematics can be verifiable if you write a proof for example um and so yeah I think it's a great question that I don't really have the answer to and I think that this is something the industry and us as engineers will have to figure out is you know my sort of uh hope and prediction for the future is that we'll see more and more interesting new agentoriented programming languages and one of the most fascinating ones that I've seen so far is this one called bend bend d bend um and that language is one where it kind of marries programming with proofs right there used to be a time, you know, where you actually had to write your proofs in a different language. And proofs, by the way, for those uh who who aren't familiar is this idea of uh that you can sort of formally verify that some code is correct mathematically, right?

Especially if you've written your code in a very functional programming way. uh but for the longest time you had to do that in a separate language like lean or tla+ or uh I'm blanking on some of the other other examples but uh like languages like that where you would construct the mathematical proof and then use a solver essentially to det that that you've covered all the cases you don't have like a race condition or whatever. So yeah, I think trying to sum up the question, I think yeah, if you are in a position where you can figure out how your agents can truly verify the work in a way that gives you confidence, you can actually, you know,uh have the PRs merge cuz if it compiles, right, if it if the proofs show you that it's correct, then why wouldn't you just merge it?

Um but of course, yeah, not all the means are verifiable. Yeah, it's a tough one. Um, okay. I think we've got to think about wrapping up because we are nearly on the hour. Have you Have you got something after this?I mean, I've got something before I give my son dinner, but

嘉宾 I I can go a bit longer after you.

主持人 Okay, let's let's go five minutes longer then. Um I think I just want to have one more question which is I think I want to ask how you see Pstack and how you see skills in general like in terms of we talked about this before we went on air which is like people think of as like my skills versus your skills and how do you combine frameworks together?How do you use Pstack with my stuff?

like what should you take from each one and because I think I see skills as sort of just derived from process basically like they're just processes turned into words and I would love to know how you recommend people take Pstack and take my stuff as well and turn it into their own processes. I think you shared a tip actually today that I thought was actually very relevant, which is this idea that you go off and look at your previous transcripts, right?

And you sort of mine for information of, you know, your own look through your own your own prompts, right, to the agents where you correct them where you have to constantly intervene and uh you know take that higher level learning and turn that into a a reusable skill, right?

So that agents stop repeating that mistake. I think that uh PAC and your skills are very complimementaryary. I I totally agree with you that they're like a skill is really much just process. I mean it's just at the end of the day a skill is just English or or language. It's just markdown.

嘉宾 Um and I think you can you can definitely weave them, combine them in a way that makes sense to you. But I do think that uh everyone should have their own set of knives, right?I I keep going back to the the the cooking analogy, but it's so apt because like, you know, every chef when they go to a different job, right, when they go to a different restaurant, they carry they bring their knives with them. The tools go with them, right?

And so trust to me is really about trust in your own tools. And when you spend the time sharpening them and understanding them really, really well, you can do great things. And everybody's skills and tool set is going to look different. you know, someone might find a lot of success combining, you know, like your grill me with docs, uh, or wayfinder skill with some of the execution skills in Pstack as an example. Some people might use more of your skills, some people might use

more of my skills. I think at the end end of the day, it really just comes back to how much do you trust, you know, me and Matt, right?

Like if you if you trust us both, of course, use our skills, but I also encourage you to, you know, look at your own transcripts. Um um tell the agent to look through, you know, some of the all of the patterns that you've used, the times you've had to intervene, you know, suggest turning them into lint rules or new skills, right?The the past chats I I often say is like a a treasure trove of context because that, you know,it's it's like the process materialized, right?

Like it's the real process. It's not an abstract idea in your head. it's the actual thing right and you can actually see how it happened in practice and extract so much information from that and there's so much so that I actually turn I have a skill in pac called recall which is exactly that um where this was a pattern where you know I was working in I was working on a similar problem so specifically I was working on virtualization for the cursor application and there were a lot of bugs and so you know every time I started a new chat I was like ah this there's so much good context from the last one So, you know, I want to bring it over to the new chat. How do I do that?

And that's where the transcript came, uh, you know, looking at the past transcript came about. And then recall was just a way for me to collapse and compress that workflow into a skill. So that I didn't have to just say I didn't have to write a long essay every time. Go look at all these chats, right?

And, you know, blah blah blah. So, I I largely think of skills, especially as agents get more capable as really encoding workflows. you know skills from last year were really more about like almost like implementation details like here are the exact script commands you know you should use right I think with the latest models you can just delete those parts and just really focus on the workflow right until it it's more the skill becomes more like a series of steps a series of your process uh and I think over time we'll see that skills get smaller and smaller you know more compact Um, and yeah, they're very compatible. Or you can, you know, if you want, why not read our skills, right, and com and combine them in your of your own, right?

Combine Wfinder with potato mode and make your own custom mode, right?Like like skills are the the thing I love about skills that is are that they're so malleable. You can do anything you want. It's just language.

主持人 Absolutely. There's nothing magical in them, right? They're just words. And

嘉宾 Exactly. If if there is any magic in them, it's just the words chosen and the phrases used and the thinking that's been done to turn those like take abstract process and turn them into language. And once that thinking has been done, then it's just there. It's available. It's on the surface and you just nick it. Um Lauren, thank you so much. This has been glorious.

主持人 Yeah, this has been super fun. I really enjoyed talking to you. Hope we can do it again. I'd love to do it again. I'd love to do it again. Absolutely. Um yeah, we'll check in in uh

嘉宾 I don't know. Yeah, Monday. Let's do it.

主持人 Yeah, let's do it. Part two.

嘉宾 Well, thank you so much. I'm going to close the stream here. Laura and I will uh chat a little bit and stay here. But thank you guys so much for watching. The glorious.

回到顶部