# LIVE: Poteto (creator of pstack) on shipping 1,000's of PR's a month at SpaceX · 中英对照逐字稿

- 原节目：Matt Pocock
- 英文原始来源：https://www.youtube.com/watch?v=MN9dGgmLyso
- 中文译制版入口：https://www.youtube.com/watch?v=MN9dGgmLyso
- 时长：01:05:36
- 方法与限制：英文来自已验证的原始节目 transcript/caption；中文由 Codex 逐段翻译，未做逐字人工校对，公开引用前请回到英文原文与音频复核。

## 中英对照逐字稿

### [00:00:00–00:00:29]

**EN**  So, hello folks. I've got another treat for you today. Last time on this kind of podcasty thing, I suppose, we had Uncle Bob and we talked about software quality. We talked about agents. We talked about lots of cool stuff. Now, we have uh an incredible guest, someone who I'm delighted to welcome on, who's been exploding on Twitter recently about software factories, um about increasing the quality of your work, about

**中文**  所以，大家好。今天我又给你们准备了一份好内容。上次在这种我姑且算作有点像播客的节目里，我们请到了 Uncle Bob，聊了软件质量。我们聊了 agent。我们聊了很多很酷的东西。现在，我们有呃一位了不起的嘉宾，一位我很高兴请上台的人，她最近在 Twitter 上关于软件工厂的讨论爆了，嗯，关于提高你工作的质量，关于

### [00:00:26–00:00:54]

**EN**  increasing your velocity and climbing the trust ladder with agents so that you can ship more and more and more. And it is potato. Welcome. Thank you so much for joining. >> Thanks for having me. Yeah, very excited to be here. Yeah, big fan of yours >> and a huge fan of yours. I [laughter] think people have been talking about this like it's like the meeting of the skill minds, the skill Mount Olympus or something because both of us have very popular skill libraries. Um I've not, as

**中文**  提高你的速度，以及和 agent 一起攀登信任阶梯（trust ladder），这样你就能交付越来越多、越来越多。这位就是 poteto。欢迎。非常感谢你来。poteto：谢谢邀请我。是的，非常高兴来到这里。是的，我是你的大粉丝。主持人：我也是你的超级粉丝。我[笑声]觉得大家一直在说这件事，就像 skill 头脑的聚会，skill 的奥林匹斯山之类的，因为我们俩都有非常受欢迎的 skill 库。嗯，我还没有，正如

### [00:00:53–00:01:22]

**EN**  I was saying before we started, I've not used a ton of yours and like I want to get all of the juice out of your brain so that I can go and use it properly and use it better. And I think where I want to start with this is you gave a talk um pretty recently like um about 10 days ago and posted on X which went absolutely nuts as about how I shipped 2,500 PRs last month to production got about 3 million views or something on X

**中文**  我在我们开始之前说的，我还没大量用过你的那些，而且我想把你脑子里的精华都榨出来，这样我就能去正确地用它，并且用得更好。我想从这里开始的是，你不久前做了一场分享，嗯，大概十天前，发在了 X 上，彻底爆了，讲的是我上个月怎样把 2,500 个 PR 交付到了生产环境，在 X 上大概有 300 万次浏览之类的

### [00:01:19–00:01:47]

**EN**  and I watched it and I loved it and I recommended it and I kind of want to run this as almost like a Q&A of that talk basically of giving you because it just I just had tons of questions about it and I wanted to dive into it. And I think where I want to start is you talk about a trust ladder with agents where you as you trust agents more, you can get them to do better and better things

**中文**  而且我看了，我很喜欢，我也推荐了它，我有点想把这次几乎做成那场分享的一场问答，基本上就是把时间给你，因为它就是，我就是有一大堆关于它的问题，我想深入进去。我想从这里开始：你谈到了和 agent 一起的信任阶梯（trust ladder），随着你越来越信任 agent，你就能让它们做越来越好的事情

### [00:01:45–00:02:15]

**EN**  and or scale them to up to use more and more agents. So what is your story of how you climbed the trust ladder and how did that work when like you got SpaceX and >> started climbing more and more? So I think this the the journey sort of began even before I joined cursor uh which is now SpaceX AI. Uh [snorts] so the story is um after Meta so I I used to work at Meta

**中文**  以及，或者把它们扩大到去使用越来越多的 agent。所以，你攀登信任阶梯（trust ladder）的故事是什么，以及当你到了 SpaceX 之后那是怎么运作的，poteto：然后越爬越高？所以我觉得这段旅程差不多甚至在我加入 Cursor 之前就开始了，呃，Cursor 现在是 SpaceXAI。呃[哼鼻]所以故事是，嗯，在 Meta 之后，我曾经在 Meta 工作

### [00:02:13–00:02:41]

**EN**  on the React team. [snorts] Uh I took a month off uh because I was feeling kind of burnt out and of course when what do you do when you're burnt out? You go and start a new side project. Um and so I started a side project. you know, I was uh of course using AI to to write code. Uh [snorts] but then I started to realize uh you know, I was spending like so many hours just micromanaging one agent, right? And you know, at the time, this was back in

**中文**  在 React 团队。[哼鼻]呃，我休了一个月，呃，因为我感觉有点倦怠了，当然，倦怠的时候你会做什么？你就会去开一个新的业余项目。嗯，所以我开了一个业余项目。你知道，我当时呃当然在用 AI 来写代码。呃[哼鼻]但后来我开始意识到，呃，你知道，我花了好像好多好多个小时，只是在微观管理一个 agent，对吧？而且你知道，在那个时候，那是回到

### [00:02:39–00:03:08]

**EN**  February, maybe February, early February or January, you know, people were really obsessed with this idea of like orchestration. This was like, you know, before, you know, things like cursor, you know, like the agents window was had become popular. So people were still in like like 2 land you know in their terminal and they were all talking about okay here you know I built a custom orchestrator right and so of course I had I was a bit nerd sniped by that [snorts] and you know as I was building

**中文**  二月，也许是二月，二月初或者一月，你知道，人们当时真的痴迷于编排这个想法。这大概是，你知道，在 Cursor 这类东西之前，你知道，在 agents window 变得流行之前。所以人们还待在 two land 里，你知道，在他们的终端里，大家都在说，好，你看，我做了一个自定义编排器，对吧，所以我当然有点被这件事勾住了[哼鼻]，你知道，就在我做着

### [00:03:05–00:03:35]

**EN**  my toy project uh I got nerd sniped by oh how do I make my AI coding setup more efficient and so you know I I kind of started the journey there where I just you know took a step back and realized you know I was spending all this time micromanaging a single agent you know I was [snorts] creating skills and I was like finding it quite difficult to measure the output or the the result the impact of the skill as well so I was

**中文**  我的玩具项目的时候，呃，我被「哦，我怎样才能让我的 AI 写代码配置更高效」勾住了，所以你知道，我差不多是从那里开始这段旅程的，我只是你知道退了一步，意识到你知道我把所有这些时间都花在微观管理一个单独的 agent 上，你知道，我[哼鼻]在创建 skill，而且我发现很难衡量产出，或者结果，skill 的影响，所以我当时

### [00:03:32–00:03:59]

**EN**  kind of flying blind but I was you know iterating really fast um and um so that project eventually sort of became the basis of PAC even though I didn't know it at the time um and a lot of some tricks I had learned like building that early set of skills. Actually, it's still open source if you want to if anybody wants to take a look. It's on my

**中文**  有点是在盲飞，但我你知道迭代得非常快，嗯，所以那个项目最终差不多成了 pstack 的基础，尽管我当时并不知道，嗯，还有我学到的不少技巧，比如构建那一套早期的 skill。实际上，它现在仍然是开源的，如果你想看，如果有人想看一看。它在我的

### [00:03:56–00:04:25]

**EN**  GitHub like potato noodle n o d l e. Um, and in there you will see some skills and a brain directory. [snorts] And so I was really interested in this idea of how do I, you know, extract my own ability, if that makes sense, and give it to the agent, right? cuz I was I I realized that you know all I was trying to do was trying to teach the agent to write code more like me you know do do you do do

**中文**  GitHub 上，像是 poteto noodle，n o d l e。嗯，在那里你会看到一些 skill 和一个 brain 目录。[哼鼻]所以我当时真的对这个想法很感兴趣：我怎样，你知道，提取我自己的能力，如果这说得通的话，然后把它交给 agent，对吧？因为我意识到，你知道，我一直试图做的，就是教 agent 写代码更像我，你知道，把事情做得更像我，做做你做做

### [00:04:22–00:04:48]

**EN**  workflows more like me. So you know the skills were like an entry point to doing that. Um and then you know after I joined cursor uh I was starting to work on the agents window and uh it had a lot of performance issues. Uh it was it was it was pretty laggy. Uh and so since I had experience working in React, I was asked like, "Hey, do you want to come and help

**中文**  工作流更像我。所以你知道，skill 就像是做这件事的一个入口。嗯，然后你知道，在我加入 Cursor 之后，呃，我开始做 agents window，呃，它有很多性能问题。呃，它相当卡。呃，所以既然我有在 React 里工作的经验，就有人来问我，比如，「嘿，你想过来帮忙

### [00:04:45–00:05:13]

**EN**  out uh with the agents window?" Um and so the the this beginning of the cursor journey was very manual. [snorts] Uh I was deep in like looking at like flame graphs and heap snapshots and trying to see like why exactly is the app so slow. Uh but then coming back to the same realization like you know I was sort of the bottleneck. I was doing everything manually. I was sort of the meat proxy

**中文**  一下，呃，做 agents window 吗？」嗯，所以这段 Cursor 旅程的开头非常手工。[哼鼻]呃，我埋头在看火焰图和堆快照，试图看清楚这个应用到底为什么这么慢。呃，但后来又回到同一个认识，就是你知道，我差不多就是那个瓶颈。我每件事都在手工做。我差不多就是那个肉身代理

### [00:05:12–00:05:41]

**EN**  in a way, right? I was the meat proxy between my agent and Chrome DevTools. Uh and I was like really annoyed by that. >> And what month of the year is that? Let's say where are we in the timeline? >> Uh so I joined Cursor in March. So this was like early early April probably early April is when [snorts] you know uh I joined and I didn't have any skills, right? I had I I sort of abandoned my personal skills because I didn't think they'd be relevant anymore. Uh but then

**中文**  从某种意义上说，对吧？我是我的 agent 和 Chrome DevTools 之间的肉身代理。呃，我当时真的很烦这件事。主持人：那是一年里的哪个月？我们说一下，时间线上我们在哪？poteto：呃，所以我是三月加入 Cursor 的。所以这大概是四月初，很可能是四月初，就是[哼鼻]你知道，呃，我加入的时候，我没有任何 skill，对吧？我差不多放下了我个人的那些 skill，因为我觉得它们不再相关了。呃，但后来

### [00:05:39–00:06:08]

**EN**  working on the agents window uh and [snorts] now working on grockbot uh I sort of realized that a lot of the lessons I had learned from those skill time building the the initial set of skills were very relevant especially around things like verification uh you know being very rigorous in your work um [clears throat] because I think from my experience even the the frontier

**中文**  在做 agents window 的时候，呃，[哼鼻]现在在做 Grok Bot 的时候，呃，我差不多意识到，我在那些做 skill 的时间里、构建最初那一套 skill 时学到的很多教训都非常相关，尤其是围绕验证这类事情，呃，你知道，在工作里非常严谨，嗯[清嗓子]，因为我觉得以我的经验，即使是前沿的

### [00:06:05–00:06:34]

**EN**  ones tend to take shortcuts. Uh they tend to do the easy thing. Uh so uh a lot of the skills that I've built have been around how do I make the easy thing the right thing? You know, how do I make that the best thing? >> The idea of sort of distilling your expertise and turning what you do every day into processes, that's something that feels super familiar to me. That's

**中文**  那些也倾向于走捷径。呃，它们倾向于做容易的事。呃，所以呃，我构建的很多 skill 都围绕着：我怎样让容易的事变成正确的事？你知道，我怎样让那件事变成最好的事？主持人：把你的专长提炼出来，把你每天做的事变成流程，这个想法对我来说感觉特别熟悉。那就是

### [00:06:32–00:07:02]

**EN**  exactly what I've been doing with the skills. And I suppose there's something in that which is a lot of people think domain expertise is getting less useful now as people uh start to rely more on AI where what do you think about that just as a sort of vibe check before we start talking [laughter] >> I actually feel like domain expertise is more important than ever you know uh I think I wrote this on my ex at some

**中文**  正是我一直在用 skill 做的事。而且我觉得这里面有一点，就是很多人认为领域专长现在越来越没用了，随着人们呃开始更依赖 AI，你对这件事怎么看，就在我们开始聊之前，先做一种氛围核对[笑声]poteto：我其实觉得领域专长比以往任何时候都更重要，你知道，呃，我想我在某个时候把这个写在了我的 X 上

### [00:06:59–00:07:27]

**EN**  point but you know at times I sometimes think of you know AI as is like especially as the models get smarter and more capable and the frontier models are just getting so good like I [snorts] love Opus 5.5 by the way um uh you know as the models get really really good it almost becomes like the bottleneck is no longer the agent right it becomes your ability to

**中文**  某个时候，但你知道，有时候我有时会把，你知道，AI 想成这样，尤其是随着模型变得更聪明、更有能力，前沿模型正在变得非常好，顺便说我[哼鼻]很喜欢 Opus 5.5，嗯，呃，你知道，随着模型变得非常非常好，几乎就变成了瓶颈不再是 agent，对吧，它变成了你的能力，去

### [00:07:25–00:07:54]

**EN**  express your intent and your goals in a clear way that [snorts] the agent can understand and actually carry out and That's why I think you know like people with a lot of domain expertise are extremely have a have a huge advantage in my opinion especially if you're a little bit like you know tech technoc curious you know so I I think of people like you know like uh like a doctor or a lawyer or you know someone who who has a

**中文**  用一种清晰的方式表达你的意图和你的目标，让[哼鼻] agent 能够理解并真正执行，这就是为什么我觉得，你知道，有很多领域专长的人，在我看来有巨大的优势，尤其是如果你有那么一点，你知道，对技术感到好奇，你知道，所以我会想到像，你知道，像呃一位医生或者一位律师，或者你知道，某个有着

### [00:07:52–00:08:20]

**EN**  deep expertise in a particular non-engineering domain and [snorts] if they're actually just a little bit techsavvy and they can figure out how to use agents they can actually build really really great products, right? If they if they have a clear enough vision in their head and they can articulate it in a way that the agent can build it, you know, I think that that is really the the bottleneck these days is is like

**中文**  在某个特定的非工程领域有深厚专长的人，而且[哼鼻]如果他们其实只是有一点技术上手的能力，能搞清楚怎么使用 agent，他们其实就能做出非常非常好的产品，对吧？如果他们脑子里有足够清晰的愿景，并且能用一种 agent 可以把它做出来的方式把愿景说清楚，你知道，我觉得这才真正是如今的瓶颈，就是

### [00:08:18–00:08:47]

**EN**  the transfer of your intent, right, and your vision to the agent. >> Yeah. I've been obsessed with language basically since agents um dropped. are just obsessed 100% and thinking constantly about the the composition of words, how I can make things sharper, what um what might be hidden in the phrases that I'm using. And it's and finding what I love is when you find a word that the agent then hooks on to and

**中文**  把你的意图，对吧，以及你的愿景，传递给 agent。主持人：是的。自从 agent 嗯出现以来，我基本上就对语言着迷了。就是百分之百地着迷，并且不断地想词语的组合，我怎样能让事情更锐利，嗯，我正在用的短语里可能藏着什么。而且我喜欢发现的是，当你找到一个词，agent 随后就钩住它，并且

### [00:08:46–00:09:14]

**EN**  then goes, "Okay, I'm going to reinforce that word. I'm going to reuse that in my thinking traces." You know, I found that with um TDD was an early example of that. a lot of chat about TDD recently of like, you know, people say, should you use TDD with agents? Doesn't matter. What you're doing is you're getting the agent to think about TDD, getting it to write tests, getting it to prioritize things in a different way than it did before. And that's why sort of grilling, I think, works effectively. Grilling is a

**中文**  然后会说，「好，我要强化这个词。我要在我的思考轨迹里复用它。」你知道，我发现嗯 TDD 就是一个早期的例子。最近有很多关于 TDD 的讨论，比如，你知道，人们说，你应该和 agent 一起用 TDD 吗？这不重要。你在做的是让 agent 去想 TDD，让它写测试，让它用一种和以前不同的方式给事情排优先级。这就是为什么某种 grilling，我觉得，会有效。Grilling 是一个

### [00:09:11–00:09:39]

**EN**  >> Yeah. Yeah. It it draws those words out of you, right? or or at least it helps the agent understand your thinking so that they can propose those words to you and you can pick up and say yes exactly that. >> Uh I've actually copied some of the the tips that you've shared as well where you know one of my favorite ones that you've shared recently or or not or like maybe in the past couple weeks is [snorts] about uh reducing or eliminating tautological tests. Like one

**中文**  poteto：是的。是的。它把那些词从你身上引出来，对吧？或者至少它帮助 agent 理解你的思考，这样它们就能把那些词提议给你，你可以接住并说，对，就是那个。poteto：呃，我其实也抄了一些你分享过的技巧，你知道，你最近分享的、或者不是、或者大概过去几周里我最喜欢的一个是[哼鼻]关于呃减少或消除同义反复的测试。比如一个

### [00:09:38–00:10:06]

**EN**  of my pet peeves of agents is like all of the useless tests that they write. And so, you know, that was one thing where, you know, the word tutology, right, is is is I guess, you know, not many people necessarily know that if if especially if English isn't your first language, but there's a lot of meaning to that word. And it's like it's almost like compressed, right? Like you compress a lot of intent and meaning into words. And so I I I totally agree

**中文**  我对 agent 特别烦的一点，就是它们写的那些全都没用的测试。所以，你知道，那就是一件事，你知道，tautology（同义反复）这个词，对吧，是，我想，你知道，不一定有很多人知道，如果，尤其如果英语不是你的第一语言，但这个词有很多含义。而且它几乎像是被压缩过的，对吧？就像你把很多意图和含义压缩进词语里。所以我完全同意

### [00:10:04–00:10:32]

**EN**  with you. I think language I've always been interested in language actually uh like programming languages natural human languages and how they came to be and it's so interesting that now with agents it's sort of like this meeting of natural language with programming language but it's all it's all language out of the hood it's all communication >> totally I did a drama degree right so you know I've been thinking about language and Shakespeare and stuff for a long time [laughter] and so this all

**中文**  你。我觉得语言，我其实一直对语言感兴趣，呃，像编程语言、自然的人类语言，以及它们是怎么形成的，而且很有意思的是，现在有了 agent，这就有点像自然语言和编程语言的这场相遇，但在底层它全是语言，全是沟通。主持人：太对了，我读的是戏剧学位，对吧，所以你知道我已经想了很久语言、莎士比亚这些[笑声]，所以这一切

### [00:10:30–00:11:00]

**EN**  feels very familiar >> um so okay there's sort before we get into like because I think the thing I want from you is like software factory stuff, right? Software factory is the big buzzword. Software factory is the thing that I'm thinking about too. I'm sort of releasing a course in that direction too. >> And it's this sort of scaling yourself up to un unrealistic numbers of PRs basically or PR numbers that sound ridiculous to people who don't

**中文**  感觉非常熟悉，嗯，所以好，在我们进去之前有点，因为我觉得我想从你这里得到的是软件工厂那套东西，对吧？软件工厂是那个大热词。软件工厂也是我在想的那件事。我也有点要往那个方向出一门课。而这有点是把自己扩展到不、不现实的 PR 数量，基本上，或者对那些不

### [00:10:58–00:11:25]

**EN**  understand how this works. So where I want to get to is sort of from people who are doing kind of like one to five agents today up to, you know, hundreds of agents running at once and how that sort of functions. And so I'd love to hear about your metaphor of the Michelin Kitchen instead of the software factory because I think that says a bit about the way you think about this stuff. >> Yes. Yeah. I I I've I've never really liked

**中文**  明白这是怎么运作的人来说听起来荒唐的 PR 数量。所以我想到达的地方，是从今天大概在用一到五个 agent 的人，到，你知道，一次同时跑几百个 agent，以及那大致是怎么运转的。所以我很想听听你用米其林厨房而不是软件工厂的那个比喻，因为我觉得这说出了一点你是怎么想这些事的。poteto：是的。是的。我从来都不太喜欢

### [00:11:23–00:11:51]

**EN**  the term software factory. Not because you know it's not accurate but I think I think the a lot of people when they think factory right they don't necessarily equate that with quality or craft right things which are very important to me and a lot of people and technologists who work you know building products we care about the user experience we care about the things we're building. So while so while I do

**中文**  软件工厂这个词。不是因为，你知道，它不准确，而是我觉得，很多人想到工厂的时候，对吧，他们不一定会把它和质量或手艺等同起来，对吧，而这些对我、对很多人、对你知道做着产品的技术人来说非常重要，我们在乎用户体验，我们在乎我们正在做的东西。所以虽然，所以虽然我确实

### [00:11:48–00:12:18]

**EN**  think software factory is an apt term, it also I guess maybe conjures up negative, you know, maybe sometimes negative connotations. So Michelin Kitchen is the thing that I've sort of landed on where it's much more I feel like it's much more aspirational and uh I like the metaphor a lot cuz you know I like food. I'm called potato of course [snorts] and I like cooking [laughter] and I see a lot of parallels right like with food right when you're cooking a

**中文**  觉得软件工厂是一个贴切的词，它也，我想，也许会唤起负面的，你知道，也许有时候是负面的联想。所以米其林厨房是我差不多落定的那个说法，我觉得它更让人向往，呃，我很喜欢这个比喻，因为你知道我喜欢食物。我当然叫 poteto[哼鼻]，而且我喜欢做饭[笑声]，我看到很多相似之处，对吧，就像食物，对吧，当你做一

### [00:12:16–00:12:46]

**EN**  meal for yourself for example [clears throat] it's both utilitarian like you're trying to just feed yourself right and and survive uh but it can actually be transformed into art right and that's what what a Michelin starred chef or even just a chef or a cook can do with food is take something very ordinary and turn it into a delicious meal that you know takes you back to your childhood days or something like that. Um and so it almost like mirrors

**中文**  顿给自己吃的饭，比如[清嗓子]，它既是实用的，就像你只是想喂饱自己，对吧，并且活下去，呃，但它其实可以被变成艺术，对吧，而这就是一位米其林星级主厨，或者哪怕只是一位主厨或一位做饭的人，能用食物做的事：拿一件很普通的东西，把它变成一顿美味的饭，你知道，把你带回你的童年之类的。嗯，所以它几乎就像映照着

### [00:12:44–00:13:12]

**EN**  that trust letter that I talk about where uh you can sort of imagine your own journey as a home cook, right? Uh as a home cook, you are doing all of the food, the cooking yourself. You cut all the vegetables, you do all the prep work, you do all the cleanup, you know, you are the one man or one woman show really. Um, and it's an interesting thought experiment like, okay, if you

**中文**  我谈的那个信任阶梯（trust ladder），在那里，呃，你可以大致把你自己的旅程想象成一位家庭厨师，对吧？呃，作为家庭厨师，所有的食物、所有的烹饪都是你自己在做。你切所有的蔬菜，你做所有的备菜，你做所有的收拾，你知道，你真的就是一个人的独角戏。嗯，这是一个有趣的思想实验，比如，好，如果你

### [00:13:09–00:13:38]

**EN**  were to cook a meal and then you add people, right, your your your partner trying to your brother, your sister, and now suddenly you have your whole family in the kitchen. I think most people would get very stressed by that, right? The thought of, oh, so many people are just mocking around in my kitchen. They have no no idea where all the utensils are. >> I have a max capacity of one person in the kitchen. Yeah, absolutely. So, I feel like that that's really apt because when you ask yourself that question of

**中文**  要做一顿饭，然后你加人，对吧，你的、你的、你的伴侣，试着你的兄弟、你的姐妹，现在突然你的全家都在厨房里。我觉得大多数人会因此非常紧张，对吧？一想到，哦，这么多人就在我的厨房里瞎折腾。他们完全不知道所有餐具在哪里。主持人：我的厨房最大容量是一个人。poteto：是的，绝对是。所以，我觉得这真的很贴切，因为当你问自己那个问题，

### [00:13:35–00:14:05]

**EN**  how do I go from being a solo cook, right, to having an army or even not not even an army but a few sue chefs, right, that that are helping me in the kitchen. How do I think about dividing the work in a way that makes sense? You know, I'm not dividing work just for the sake of it, but in a way that actually makes the sum the to the better than, you know, the total of its parts. And so the Michelin kitchen metaphor to me like

**中文**  我怎样从一位独自做饭的人，对吧，变成拥有一支军队，或者甚至不是、甚至不是一支军队，而是几个 sous-chef，对吧，在厨房里帮我。我怎样思考用一种说得通的方式来分工？你知道，我不是为了分工而分工，而是用一种确实让总和变得比，你知道，各部分的总量更好的方式。所以米其林厨房这个比喻对我来说

### [00:14:03–00:14:33]

**EN**  works really well in that regard because you know as a chef you're you know if you become a chef you're in a position where you're not necessarily cooking all the food yourself anymore but you are thinking you're almost like the tech lead right for the kitchen where uh you know chefs have to think about you know not just cooking but they have to basically organize the whole kitchen and they're like the CEO of the kitchen they have to think about when do you order ingredients, how do you store them, how

**中文**  在这方面非常好用，因为你知道，作为一位主厨，你知道，如果你成为主厨，你所处的位置是你不一定再亲自做所有的食物，但你在思考，你几乎就像厨房的 tech lead，对吧，在那里，呃，你知道，主厨必须想的，你知道，不只是烹饪，他们基本上必须把整个厨房组织起来，他们就像厨房的 CEO，他们必须想你什么时候订食材，你怎么储存它们，怎么

### [00:14:31–00:14:59]

**EN**  do you prepare them, when do they have to be prepared, you know, it's a whole job, right? That's not just cooking. Um, and I think that again it mirrors so much of how engineers write code today where you are not writing the code yourself anymore. You have agents, right? But you as the human are still responsible for the final outcome, right? your name still is associated with the work that you do, your

**中文**  你怎么准备它们，它们必须在什么时候准备好，你知道，这是一整份工作，对吧？那不只是烹饪。嗯，而且我觉得这又一次映照了今天工程师写代码的方式有那么多相似：你不再自己写代码了。你有 agent，对吧？但作为人的你仍然对最终结果负责，对吧？你的名字仍然和你做的工作连在一起，你的

### [00:14:55–00:15:25]

**EN**  reputation and you know so how you set up your kitchen right and how you set up your skills your environment your codebase I think are ultimately the new ingredients that go into um building product >> yeah I think what I love about your approach is the amount of focus that you put into the environment that the agent operates in right because I think a lot of people they think, right, the agent

**中文**  声誉，而且你知道，所以你怎样布置你的厨房，对吧，以及你怎样设置你的 skill、你的环境、你的代码库，我觉得最终就是进入嗯、做产品的那些新食材。主持人：是的，我觉得我喜欢你这个做法的地方，是你放进 agent 所在的那个环境里的关注有多少，对吧，因为我觉得很多人会想，对吧，这个 agent

### [00:15:23–00:15:51]

**EN**  is good. I'm probably not going to be able to make it better. Let's just trust what these magic model people have put into the harness and the model combination. Uh, there's nothing I can really do, like I can't mess about with claw codes internals or something or whatever you're using. Um, but what I love about your approach, and it's something I advocate for too, is that you can change the environment the agent operates in, right? you can make changes

**中文**  是好的。我大概没办法让它变得更好。我们就信任这些搞魔法模型的人放进 harness 和模型组合里的东西吧。呃，我其实没什么真能做的，比如我不能去乱动 claw code 的内部之类的，或者不管你在用的是什么。嗯，但我喜欢你这个做法的地方，而且这也是我同样主张的，是你可以改变 agent 所在的环境，对吧？你可以做出改变

### [00:15:48–00:16:16]

**EN**  in the codebase and also give it tools for verification as well and allow it to verify its own work. So the thing I I loved about watching that talk is the amount of focus you put in verification and like that is the lever that you can start to generate trust. Can you talk about that and what that concretely looks like? Let's start like looking at practical ways that people can improve their own processes, their own kitchens.

**中文**  在代码库里，并且也给它验证用的工具，让它验证它自己的工作。所以我看那场分享时喜欢的一点，是你放在验证上的关注有多少，而那就是你可以开始用来产生信任的那个杠杆。你能谈谈这个，以及它具体看起来是什么样吗？我们开始看看人们可以用来改进自己的流程、自己的厨房的实际办法。

### [00:16:14–00:16:44]

**EN**  Yeah, I've I've I've said this a lot actually that you know even if you don't use PAC or you know your skills I think that the single most important skill that should be in your toolkit is verification because without verification and for for by the way for those watching who don't know what that means it's this idea that you can give you can sort of give your agent uh hands and eyes in a way that's the the analogy

**中文**  是的，我其实说过很多次，你知道，即使你不用 pstack，或者你知道，你的 skill，我觉得你的工具箱里最该有的、最重要的那一项 skill 就是验证，因为没有验证，而且，顺便，对那些在看、还不知道这是什么意思的人来说，这个想法是你可以给，你可以有点给你的 agent 呃一双手和一双眼睛，从某种意义上，这就是那个类比

### [00:16:42–00:17:11]

**EN**  I where the agent is able to run the code, right? And actually >> [snorts] >> uh interact with it like a normal human user would and also do things like you know debug it, you know, take traces and snapshots. Um and uh funnily enough like that was actually the first skill I built when I joined Cursor. uh that gave me a lot of that was that was the thing that actually started to let me ascend the trust ladder a little bit in a way

**中文**  我，在那里 agent 能够运行代码，对吧？而且实际上[哼鼻]呃像一个正常的人类用户那样和它交互，并且也做一些事，比如你知道调试它，你知道，采集 trace 和快照。嗯，而且呃有意思的是，那其实是我加入 Cursor 时构建的第一个 skill。呃，那给了我很多，那就是那件实际上开始让我以某种方式往信任阶梯（trust ladder）上稍微爬一点的事

### [00:17:09–00:17:38]

**EN**  that some of the other skills I had looked at or built had not really let me do because no matter how good you know some of the other skills were like the how skill, the why skill, the unsop skill were, I was still relying on me right as the proxy between my agent and the output. So that you know if the agent can't actually see the result of its work there's no way it can actually iterate right and so this is where

**中文**  而我看过或构建过的其他一些 skill 并没有真正让我做到，因为无论你知道其他一些 skill 有多好，比如 how skill、why skill、unslop skill，我仍然依靠我自己，对吧，作为我的 agent 和产出之间的代理。所以你知道，如果 agent 不能真正看到它工作的结果，它就没有办法真正迭代，对吧，所以这就是

### [00:17:36–00:18:04]

**EN**  people start to talk about loops this idea of a loop and really I think the term loop you know seems kind of uh almost abstract like people like what what is a loop what is an agent loop but really to me like the most important part of a loop that allows it to be a loop is the verification part because the agent is able to to verify by its own work and uh you know that

**中文**  人们开始谈 loop 的地方，loop 这个想法，而我真的觉得 loop 这个词，你知道，看起来有点呃几乎是抽象的，就像人们会问什么是 loop，什么是 agent loop，但对我来说，一个 loop 里让它能够成为一个 loop 的最重要部分是验证这一部分，因为 agent 能够验证它自己的工作，而且呃你知道，那

### [00:18:02–00:18:30]

**EN**  takes you out of the equation where now I can actually do something like so the very one of the very first use cases I had for verification was you know like the performance work that I was doing on cursors agent window and I want I wanted to get to a point where I could do something called hill climbing uh which is a term that I I think the labs uh talk about a lot which is this idea that you know you have some kind of rubric or a way to judge or score something And

**中文**  把你从这道式子里拿了出去，于是现在我其实可以做一些事，比如，我对验证最早的用例之一，你知道，就是我在 Cursor 的 agents window 上做的性能工作，我想到达一个点，在那里我可以做一种叫 hill climbing 的事，呃，这是一个我觉得那些实验室呃经常谈的词，这个想法是，你知道，你有某种评分标准，或者一种评判或给某件事打分的方式，而且

### [00:18:29–00:18:58]

**EN**  now because you have a loop, you can have an agent continually try to make improvements to that. Uh I think Carpathy, Andre Carpathy also famously released uh something called auto research that has a lot of these ideas. Um but yeah, verification I would say is probably the most important skill in PAC uh and many [clears throat] other you know tool sets. Uh and I think it's the

**中文**  现在因为你有一个 loop，你可以让一个 agent 持续尝试对它做出改进。呃，我觉得 Andrej Karpathy，Andrej Karpathy 也很有名地发布了呃一个叫 autoresearch 的东西，里面有很多这些想法。嗯，但是的，验证我会说大概是 pstack 里最重要的 skill，呃，以及很多[清嗓子]其他你知道的工具集。呃，而且我觉得它是

### [00:18:55–00:19:25]

**EN**  most important thing to focus on. So a lot of the a lot of I spent a lot of time actually you know tuning the verification the creative verification skill um and also internally the the we have so many verification skills now like every app that cursor has or spaceexai has has a uh verification skill that is automaintained as well [clears throat] >> uh and it's become critical infrastructure for our team because

**中文**  最该关注的事。所以很多，很多，我其实花了很多时间，你知道，在调校验证、那个 creative verification skill 上，嗯，而且在内部，我们现在有这么多验证 skill，比如 Cursor 有的、或者 SpaceXAI 有的每一个应用都有一个呃验证 skill，而且也是自动维护的[清嗓子]，呃，而且它已经成了我们团队的关键基础设施，因为

### [00:19:23–00:19:52]

**EN**  everybody uses it >> and you went pretty far with that too right like you had a um in your talk I saw that you actually built a custom CLI for that too. So what does that CLI do? Like how does it [clears throat] execute things and why did you I mean that's proof of how deep you're going right of how much you're pushing that. >> Yeah. So this is actually a tip I learned early on where um I guess you know back in January or

**中文**  每个人都在用它。主持人：而且你在这件事上也走得相当远，对吧，比如你有一个嗯，在你的分享里我看到你其实还为此做了一个自定义 CLI。那么这个 CLI 做什么？比如它怎样[清嗓子]执行事情，以及你为什么，我的意思是这证明了你走得有多深，对吧，你把这件事推得有多用力。poteto：是的。所以这其实是我很早学到的一个技巧，在那里，嗯，我想你知道，回到一月，或者

### [00:19:49–00:20:18]

**EN**  or late last year the thing that people were concerned about was context window, right? That was the big the big topic at the time was how do I you know manage the context window because you know compaction summarization wasn't really that good yet [snorts] and people were always people had this there was almost this meme in the community that you know once your agent summarized or compacted once it would become sort of stupid

**中文**  或者去年晚些时候，人们担心的事情是上下文窗口，对吧？那是当时的大话题，就是我怎样，你知道，管理上下文窗口，因为你知道，compaction、summarization 当时还没那么好[哼鼻]，而且人们总是，人们有这个，社区里几乎有这么一个梗，你知道，一旦你的 agent 做过一次摘要或一次 compaction，它就会变得有点笨

### [00:20:15–00:20:45]

**EN**  right for the rest of your session. So there was a lot of thinking around like you know being very efficient with your context usage and so that was actually the inspiration for some of the uh the CLI work inside of the verification skills. I guess now it's less so about context because uh you know agents are much better or harnesses have gotten a lot better with summarization.

**中文**  对吧，在你这次会话剩下的时间里。所以有很多思考，围绕着你知道非常高效地使用你的上下文，所以那其实就是验证 skill 里面一些呃 CLI 工作的灵感来源。我想现在就不那么关乎上下文了，因为呃你知道 agent 好得多了，或者 harness 在 summarization 上已经好了很多。

### [00:20:42–00:21:11]

**EN**  Um, I still think there's some benefits to, you know, uh, having a clean context window. Uh, so the CLI is really just more of a way for me to take the deterministic parts of what the skill does and encode that into a script or CLI to reduce to kind of take away the judgment that would otherwise unnecessarily be used because with

**中文**  嗯，我仍然觉得，你知道，呃，有一个干净的上下文窗口是有一些好处的。呃，所以这个 CLI 其实更多只是一种方式，让我把 skill 所做的事情里那些确定性的部分拿出来，编码进一个脚本或 CLI，从而减少，有点是拿掉那些否则会被不必要地用上的判断，因为有了

### [00:21:08–00:21:38]

**EN**  judgment so I also think of you know agents and skills in sort of like it's like a gradient you have some parts of the work that are entirely judge measurement based right you know something that requires thought you know putting together multiple pieces of context thinking um and then you have the more deterministic parts like I don't know if you wanted to uh refactor some code right from one pattern to another that's very

**中文**  判断，所以我也会把，你知道，agent 和 skill 想成有点像，它像一个梯度，你有一些工作的部分是完全基于判断的，对吧，你知道，需要思考的东西，你知道，把多块上下文放在一起思考，嗯，然后你有更确定的那些部分，比如我不知道，如果你想呃把一些代码从一种模式重构到另一种，那是非常

### [00:21:36–00:22:05]

**EN**  mechanical right you don't you don't need an agent to think about it and come up with it in a novel way each time right and so that that was really the inspiration for the CLI and you'll see this in a lot of the other skills that I built is like I try to extract out the deterministic parts and turn that into code and just leave only the parts that actually require judgment to the agent. So in a way I think of the seal as kind of like your wrapper, right? It's a

**中文**  机械的，对吧，你不需要、你不需要一个 agent 去想它，并且每次都以一种新颖的方式把它想出来，对吧，所以那真的就是这个 CLI 的灵感来源，而且你会在我构建的很多其他 skill 里看到这一点，就是我试着把确定性的部分提取出来，把它变成代码，只把真正需要判断的部分留给 agent。所以从某种意义上，我觉得 skill 有点像你的包装层，对吧？它是一个

### [00:22:03–00:22:33]

**EN**  wrapper with some light instructions around how to use these custom tools that are inside of the skill. Um but yeah, I don't think the CL is really that interesting in its own really. It's not like a novel piece of software. It's just something that interacts with like Playright and the Chrome DevTools protocol and calls a bunch of APIs. It's like it's just a bunch of glue. >> No, it's fascinating because it's a way of hiding information from the skill, right? It's a way of conserving the skill, keeping the skill quite small, I

**中文**  wrapper，外面包着一些轻量说明，讲怎么使用这个 skill 里面的这些自定义工具。嗯，不过是的，我并不觉得这个 CLI 就其本身而言真的有那么有意思。它并不像一套新颖的软件。它只是去和 Playwright 以及 Chrome DevTools 协议交互，并且调用一堆 API。它就像，它只是一堆胶水。主持人：不，它很吸引人，因为它是一种向 skill 隐藏信息的方式，对吧？它是一种节省这份 skill 的方式，让 skill 保持得相当小，我

### [00:22:31–00:22:59]

**EN**  imagine, and then you're able to delegate more of the complicated deterministic stuff into a script within the skill. So it's almost you're compressing information and making the agent do more consistent things more consistently. >> Yeah, >> that's fascinating. And it it [clears throat] also helps I guess if you care about context window it it does help because now the agent doesn't need to uh you know re reinvent

**中文**  设想一下，然后你就能把更多复杂的确定性工作，委托给 skill 里面的一段脚本。所以这几乎就是你在压缩信息，并且让 agent 去做更一致的事情，而且更稳定地去做。poteto：是啊，这很吸引人。而且它［清嗓子］我觉得也有帮助，如果你在意上下文窗口的话，它确实有帮助，因为现在这个 agent 不需要，你知道，去重新发明

### [00:22:56–00:23:26]

**EN**  uh things cuz uh one thing I had noticed early on when we didn't have a CLI was that uh well the the agent would try to verify it work but it would basically rebuild the world each time and then every agent did it differently and I was starting to notice like that's very inefficient right I was wasting it it was actually not just about context usage but also speed, right? Like because now an agent had to actually go off and write the scripts or the CLI and test it and you know and it doesn't work

**中文**  一些东西，因为有一件事我很早的时候就注意到了，那时候我们还没有 CLI，就是，嗯，agent 会试着验证它的工作，但它基本上每次都要把整个世界重建一遍，然后每个 agent 的做法都不一样，我开始注意到，这样非常低效，对吧。我是在浪费它，其实不只是上下文用量的问题，还有速度，对吧？因为现在一个 agent 必须真的跑去写那些脚本或者那个 CLI，并且测试它，你知道，然后它并不能用

### [00:23:24–00:23:54]

**EN**  and the last agent did it and it worked but it discarded it. So it was just very obvious at that point like I should just turn this into a CLI and put that inside of the skill uh so that every agent that uses it now benefits from that same piece. Um but I I also think like you know it's a good push for people to think about is how much of your skills and rules could actually be deterministic.

**中文**  而上一个 agent 做过这件事，而且做成了，但它把结果丢掉了。所以到那个时候就非常明显了，我应该直接把这个做成一个 CLI，放进这个 skill 里面，这样现在每一个使用它的 agent 都能从同一份东西里受益。嗯，不过我也觉得，你知道，这很好地推动人们去想：你的 skill 和规则里，到底有多少其实可以是确定性的。

### [00:23:50–00:24:19]

**EN**  Um that's like another core thing or or one of my core principles that I like to think about is yeah how do I uh make very efficient use of determinism and non-determinism and you know let Asians shine at the non-deterministic parts right because that's what they're trained to do. Um and the other parts which are much more mechanical or you know straightforward

**中文**  嗯，这像是另一件核心的事，或者说我喜欢去想的核心原则之一，就是，是的，我怎样非常高效地使用确定性和非确定性，并且，你知道，让 agent 在非确定性的部分发挥出来，对吧，因为那正是它们被训练来做的。嗯，而其他那些机械得多，或者，你知道，直截了当的部分

### [00:24:17–00:24:44]

**EN**  can be just pure determinism. Um, and you'll see this as well for things like doing migrations. Um, which is another big thing that I've I've talked about is, you know, going from one technology to another, especially one that is better for agents, right? And a lot of how you can do that migration is, I think, through things like scripts and CLIs, like the deterministic parts like

**中文**  就可以只是纯粹的确定性。嗯，你在做迁移这类事情上也会看到这一点。嗯，这是我谈过的另一件大事，就是，你知道，从一种技术换到另一种技术，尤其是换到一种对 agent 更好的技术，对吧？而我觉得，做这种迁移的很大一部分办法，是通过脚本和 CLI 这类东西，也就是那些确定性的部分，比如

### [00:24:40–00:25:08]

**EN**  code mods, you know, like crawling the abstract syntax tree and transforming code literally mechanically, right? like a script does it for you instead of the agent. >> Totally makes sense. I I mean I think what there's another thing there which is you're taking stuff away from the agent and you're kind of putting it in the environment too a little bit which

**中文**  codemod，你知道，比如爬抽象语法树，然后实实在在地、机械地转换代码，对吧？像是让一段脚本来替你做，而不是让 agent 来做。主持人：完全说得通。我，我的意思是，我觉得这里还有另一件事，就是你在把一些东西从 agent 那里拿走，而且你也有一点点把它放进环境里，这一点

### [00:25:06–00:25:34]

**EN**  is let's say you have a a thing that you notice the agent always gets wrong. You want to make that um just impossible within the environment. And that sort of comes down to code quality as well. I mean, I talk about a lot like having a what a good codebase means, right? What is a good codebase? And there's a definition I like which is a a good codebase is a codebase that's easy to

**中文**  就是，比方说你注意到有一件事，agent 总是做错。你想让这件事，嗯，在环境里干脆变成不可能。而这多少也落到代码质量上。我的意思是，我经常谈的是，一个好的代码库意味着什么，对吧？什么是好的代码库？我喜欢的一个定义是，一个好的代码库，是一个很容易去

### [00:25:30–00:26:00]

**EN**  make changes in, right? Easy to um change stuff without things screwing up. And that means that you have a lot of guard rails that you have a lot of um the agent or the human is constrained to very narrow paths. And that's again something you talk about in your talk. >> And you talk about this not only on the kind of sort of automated checks side of things. So linting and type checking blah blah blah but also in the way you design abstractions. And you guys even I

**中文**  在里面做改动的代码库，对吧？很容易，嗯，改东西，而不会把事情搞砸。这意味着你有很多护栏，你有很多，嗯，agent 或者人被约束在非常窄的路径上。而这又是你在你的演讲里谈的东西。而且你谈这个，不只是在那种自动化检查那一侧。所以 lint、类型检查，巴拉巴拉，也包括你设计抽象的方式。而且你们甚至，我

### [00:25:58–00:26:25]

**EN**  think built a framework uh for your agent to work in too. >> I think what I'd love to hear is you obviously think of that as very important, right? And that's how important is that compared to other things you could be doing like building features or shipping work. Yeah, I think that's a um I almost feel like the new job of the engineer is really to to spend time on the environment. Um I almost actually wrote

**中文**  觉得，也搭了一个框架，嗯，让你们的 agent 也在里面工作。主持人：我想，我很想听的是，你显然觉得那件事非常重要，对吧？那么这件事有多重要，比起你能做的其他事情，比如做功能，或者交付工作？poteto：是啊，我觉得这是一个，嗯，我几乎觉得工程师的新工作，真的是把时间花在环境上。嗯，我其实差不多写了

### [00:26:24–00:26:52]

**EN**  a tweet about this yesterday, but I but I didn't. But I think that you you know if you if you haven't really spent time, you know, building trust in your agents and building skills and tools, you can get stuck in this mode where you're very low on that trust ladder, right? you don't have a lot of trust in your agents work. And so the only way to cope in that when you're in that situation is just to kind

**中文**  一条关于这件事的推文，就在昨天，但是我，但是我没发。不过我觉得，你，你知道，如果你，如果你还没有真正花时间，你知道，去建立你对 agent 的信任，去打造 skill 和工具，你可能会卡在这样一种模式里：你在那条信任阶梯（trust ladder）上非常靠下，对吧？你对你的 agent 的工作并没有多少信任。所以，在那个时候，当你处在那种情况里，唯一的应对办法就只是去

### [00:26:49–00:27:17]

**EN**  of lock in and micromanage your agents. And that's very time consuming. And when you're stuck in that mode, you don't really have the luxury to think about, you know, uh higher level things like like making yourself more productive. In the same way that uh I guess analogy would be like if you've never taken the time to learn like your tools right as a developer when you were writing code

**中文**  一头扎进去，并且微管理你的 agent。而这非常耗时间。而当你卡在那种模式里，你其实没有余力去想，你知道，嗯，更高一层的事情，比如比如让你自己更有生产力。同样地，嗯，我想那个类比会是，如果你从来没有花时间去学，比如你的工具，对吧，作为一个开发者，在你自己写代码的

### [00:27:15–00:27:43]

**EN**  yourself and you know you've never heard of VS Code, you've never heard of Vim, you only knew about Notepad [laughter] uh and you had hadn't even heard about Git. That's sort of the analogy. It's like you you haven't spent the time sharpening your own knives, right? And so, of course, if you have a dull knife, then everything's going to take a long time. Um, and you're going to be you're just going to be and and especially if you know deadlines are looming, then you

**中文**  时候，而且，你知道，你从来没听说过 VS Code，你从来没听说过 Vim，你只知道 Notepad［笑声］，嗯，而且你甚至没听说过 Git。这大概就是那个类比。就像你，你没有花时间去磨你自己的刀，对吧？所以当然，如果你的刀是钝的，那每件事都会花很长时间。嗯，而且你会变成，你只会变成，而且而且尤其是，如果你知道截止日期正在逼近，那么你

### [00:27:41–00:28:10]

**EN**  don't have the now you're stuck in this rut, right? Where where you you you don't have sharp knives, you don't have good tools, but you're under all this pressure to ship, right? And so, all you can do is just focus on that. But I do think that, you know, if you can find yourself the time to actually spend time thinking about your setup, it's again going back to the cooking, you know, analogy, it's like uh, you know, if you, for

**中文**  现在就没有那个了，你陷在这个窠臼里，对吧？在这里，你，你，你没有锋利的刀，你没有好工具，可你承受着所有这些要交付的压力，对吧？所以你能做的，就只是盯着那件事。不过我确实觉得，你知道，如果你能给自己找出时间，真正花时间去想你的这套配置，这又回到做饭，你知道，那个类比，就像，嗯，你知道，如果你，比方说

### [00:28:06–00:28:36]

**EN**  example, if if cutting cutting cutting the garlic is like super slow, right? There are garlic mashers, right? You can buy and you put it in the thing and you like squeeze it out, right? It's super fast. Uh, machines and tools were invented for a reason, right? And so if you're operating a Michelin kitchen and your your your cooks had no tools, then of course everything's going to be extremely inefficient, very very, you know, every every every cook is going to

**中文**  如果，如果切、切、切大蒜切得超级慢，对吧？是有压蒜器的，对吧？你可以买，把蒜放进那个东西里，然后你就这么一挤，把它挤出来，对吧？超级快。嗯，机器和工具被发明出来是有原因的，对吧？所以，如果你在经营一间米其林厨房，而你的，你的，你的厨师没有任何工具，那当然每件事都会极其低效，非常非常，你知道，每一个，每一个，每一个厨师都会

### [00:28:33–00:29:02]

**EN**  make something up of their own. So I think the tools and the determinism to me are you know taking that part away and and just like you said about constraints as well. It's the constraints are are to me as well like uh actually a slight tangent on that is uh I think we should talk about TypeScript cuz like we we actually both share like a background in Typescript where you know you obviously have done a lot of work with TypeScript and total

**中文**  自己现编一套。所以我觉得，工具和确定性对我来说，你知道，就是把那一部分拿走，而且就像你说到约束时也是这样。这些约束对我来说也是，就像，嗯，其实在这上面稍微岔开一下，嗯，我觉得我们应该谈谈 TypeScript，因为像我们，我们其实都有 TypeScript 的背景，你知道，你显然在 TypeScript 上做了很多工作，还有 total

### [00:29:00–00:29:28]

**EN**  TypeScript and you know you're a leader in that space and I uh had adopted TypeScript pretty early and I had given like a talk or two at Typescript conf uh many years ago and so one of the the talk that I did actually was about type systems and constraining the constraining types. Like one of my most favorite things about Typescript is actually type narrowing, right? This idea that you go from a very broad type,

**中文**  TypeScript，而且你知道，你是那个领域的领头人，而我，嗯，很早就采用了 TypeScript，而且很多年前我在 TypeScript conf 上做过大概一两场演讲，嗯，所以我实际做过的其中一场演讲，是关于类型系统，以及约束、约束类型。比如，我对 TypeScript 最喜欢的事情之一，其实是类型收窄，对吧？就是这个想法：你从一个非常宽的类型出发，

### [00:29:26–00:29:56]

**EN**  right? That could be anything and then you through type guards and you know type narrowing and you know runtime checks you can actually narrow the space and say like oh this isn't just a string this is a very special type of string. It's a constant, right? like I but I I determine that through the type system and in a way it's like uh there's a lot of parallels I think to that with constraints in your codebase where is it

**中文**  对吧？那个类型什么都可能是，然后你通过类型守卫，还有，你知道，类型收窄，还有，你知道，运行时检查，你其实可以把这个空间收窄，然后说，哦，这不只是一个 string，这是一种非常特殊的 string。它是一个常量，对吧？就像我，不过我，我是通过类型系统确定这一点的，在某种意义上，我觉得这和你代码库里的约束有很多平行之处，就是在这种地方

### [00:29:52–00:30:20]

**EN**  you're you're constraining the space right if you if you think about category theory as well you know you're constraining the the number of possible types right that can can exist and you're saying there's only one type right and for for us like that framework that I'm called Dune. Uh it's not an open source framework. It's the the way I describe it to people. It's it's kind

**中文**  你，你在约束这个空间，对吧，如果你，如果你也去想范畴论，你知道，你在约束能够存在的可能类型的数量，对吧，然后你说只有一种类型，对吧，而对，对我们来说，像那个我叫做 Dune 的框架。嗯，它不是一个开源框架。这是，这是我向别人描述它的方式。它，它有点

### [00:30:18–00:30:47]

**EN**  of like a internal Nex.js for our Electron apps. Uh but it comes with a lot of really really restrictive lit rules and the codebase is designed in a way that there's really only one way to do something. So we make use a we make use of a lot of conventional patterns. So like features all go into a specific directory. Well, every feature has its own directory. As an example, you know, there's like a a thing that

**中文**  像一个给我们的 Electron 应用用的内部 Next.js。嗯，不过它带有很多非常非常严格的 lint 规则，而且代码库被设计成，做一件事其实只有一种方式。所以我们大量使用，我们大量使用约定式的模式。所以，像功能全部放进一个特定的目录。嗯，每个功能有它自己的目录。举个例子，你知道，有一个东西会

### [00:30:45–00:31:14]

**EN**  discovers features like through a registry and like crawling the codebase and stuff like that. But this conventional pattern and the lint rules make for an environment where it's actually very hard to write bad code. And that sort of frees up the it both frees up your own mental uh you know capacity as well as the agent sort of

**中文**  通过注册表，以及爬代码库之类的方式，去发现这些功能。但这种约定式模式和这些 lint 规则，造就了一个环境，在里面其实很难写出坏代码。而这多少腾了出来，它既腾出你自己的心智，嗯，你知道，容量，也让 agent 多少

### [00:31:12–00:31:41]

**EN**  doesn't have to think about that anymore where it's just like oh there's only there's I should just if I want to add a new feature it just goes in the feature the new feature directory and all the code goes in there and I'm not going to append to a god file right that was really actually the inspiration for those feature directories is the very first couple of versions of Grockbot were composed of like eight god files which were like at least 10,000 lines

**中文**  不用再去想那个了，就是那种，哦，只有，只有，我应该就，如果我想加一个新功能，它就放进功能目录、那个新功能目录，所有代码都放在那里，我不会再往一个 god file 后面追加，对吧。那些功能目录的灵感，其实真的来自这一点：Grok Bot 最初的几个版本，是由大约八个 god file 组成的，它们大概至少有 10,000 行

### [00:31:38–00:32:06]

**EN**  long if not longer and so I kind of had to break it up into smaller pieces. Uh but it was just observing you know actually that's another important part is observing how agents fail and then every time you see a mistake every time you see something that could be done better you think you step back and think how do I turn this into a lint rule? How do I make it so that the code base makes this impossible? Right? And it comes

**中文**  长，如果不是更长的话，所以我多少得把它拆成更小的块。嗯，但这就是在观察，你知道，其实另一件重要的事，就是观察 agent 是怎么失败的，然后每当你看到一个错误，每当你看到有什么可以做得更好，你就会想，你退一步去想：我怎样把这个变成一条 lint 规则？我怎样让代码库使这件事变成不可能？对吧？而这又回到

### [00:32:03–00:32:32]

**EN**  back to me for my you know my background learning Typescript and types uh type systems is how do I constrain the space so that you know I know precisely what I'm working with and I think yeah there's a lot of parallels there. >> Totally makes sense. And don't I mean it's funny that you mentioned TypeScript and Goth files in the same sentence because Typescript famously has a 25,000line type uh file. >> [laughter] >> Although I don't know if they've

**中文**  我这里，因为我，你知道，我学习 TypeScript 和类型、嗯、类型系统的背景，就是我怎样约束这个空间，好让，你知道，我精确地知道我在跟什么打交道，而且我觉得，是的，那里有很多平行之处。主持人：完全说得通。而且别，我的意思是，有意思的是你在同一句话里提到了 TypeScript 和 god file，因为 TypeScript 出了名地有一个 25,000 行的类型，嗯，文件。［笑声］虽然我不知道他们有没有

### [00:32:31–00:32:59]

**EN**  rewritten that and go as they probably have, haven't they? Um, okay. So, environment is important. You should watch your agent like a hawk to make sure that any mistakes it makes. You turn them into things in the environment. And the benefit of the environment is you're not overloading your agent, right, in terms of rules, in terms of things it has to remember. It's just in the environment. And so it stumbles into the rules and exactly um you know bounces off them and hits them

**中文**  把那个用 Go 重写了，他们大概已经这么做了，是不是？嗯，好。所以，环境很重要。你应该像鹰一样盯着你的 agent，好确认它犯下的任何错误。你把这些错误变成环境里的东西。而环境的好处是，你没有让你的 agent 过载，对吧，在规则方面，在它必须记住的事情方面。那些就只是在环境里。所以它会撞上这些规则，而且正好，嗯，你知道，被它们弹开，并撞上它们

### [00:32:58–00:33:26]

**EN**  at the right moment. >> So okay, we still haven't talked about the 2,500 PRs. Where do those come from? Like how do you you've built your trust ladder, you've worked on your environment, and you understand, okay, um I now want to scale up. So what are the mechanics of that scaling? Are you um initiating 2500 like chats per month? That can't be right. So there must be are there any

**中文**  在恰当的时刻。主持人：所以，好，我们还没谈到那 2,500 个 PR。那些是从哪来的？就像，你是怎么做到的，你已经建好了你的信任阶梯（trust ladder），你已经在你的环境上做了工作，而且你明白了，好，嗯，我现在想扩大规模。那么这种扩大的机制是什么？你是在，嗯，每个月发起 2500 个那样的聊天吗？那不可能是对的。所以一定还有，有没有任何

### [00:33:24–00:33:53]

**EN**  kind of automated triggers that trigger stuff in your repo? Like how do you get the software factory kind of triggering work by itself? >> Right. Um I'll definitely say that the prerequisite to you know something like a very high volume of of pull requests um is the environment. you know, the the stuff we just talked about where I definitely would not have been able to do this if I had not spent the time, you

**中文**  会在你的仓库里触发事情的那种自动化触发器？就像，你怎样让软件工厂自己去触发工作？poteto：对。嗯，我肯定会说，你知道，像数量非常大的 pull request 这种事情，它的前提，嗯，就是环境。你知道，就是我们刚刚谈过的那些东西，如果我没有花那些时间，我肯定做不到这件事，你

### [00:33:52–00:34:21]

**EN**  know, thinking about the kitchen, right, and the knives and the tools for my Asians. And so, in a way, I I think of this as I've spent the time building one kitchen and one restaurant. And now I I'm in a position where I don't actually have to be there anymore because the environment, you know, that the same analogy, right? It works really well. [laughter] Yeah. You open chain of restaurants, right? That's >> Yeah. Exactly. Yeah. Exactly. It's like

**中文**  知道，去想那间厨房，对吧，还有给我的 agent 用的那些刀和工具。所以在某种意义上，我，我把这看成，我已经花时间建好了一间厨房和一家餐馆。而现在我，我处在这样一个位置：我其实不必再亲自待在那里了，因为那个环境，你知道，还是同一个类比，对吧？它运转得非常好。［笑声］是啊。你去开连锁餐馆，对吧？那就是 poteto：是啊。完全对。是啊。完全对。就像

### [00:34:19–00:34:48]

**EN**  you're Gordon Ramsay and you know, you you've taught your executive chef like all the tricks of the of coming up with great menu. Uh and like the kitchen is set up really well. Everything's just perfect and you're now in a position where you can open your second your third restaurant. And [snorts] I guess I I sort of see each project that I work on, like each big chat is sort of like a restaurant, right? and I'm I'm I have multiple of them

**中文**  你就是 Gordon Ramsay，而且你知道，你，你已经把想出好菜单的所有诀窍都教给了你的行政总厨。嗯，而且那间厨房布置得非常好。一切都刚刚好，而你现在处在可以开你的第二家、你的第三家餐馆的位置。而且［哼了一声］我想，我，我多少把我在做的每一个项目，比如每一段大的聊天，都看成一家餐馆，对吧？而且我，我，我同时有好几家

### [00:34:45–00:35:13]

**EN**  operating at the same time and I'm sort of like helicoptering between them sometimes some more than others depending on how in the loop I am but yeah definitely I think there's there are external triggers and context that those projects don't have that for a long time I was the proxy for that so uh the best example I have is like you know you have a project that's

**中文**  在同时运转，而我有点像在它们之间坐直升机来回，有时候在其中一些上面比另一些更多，这取决于我自己介入得有多深，不过是的，我肯定觉得，存在一些外部的触发器和上下文，是那些项目自己没有的，而在很长一段时间里，是我在为那个做代理，所以，嗯，我最好的例子是，比如，你知道，你有一个项目，它正在

### [00:35:12–00:35:38]

**EN**  working on a feature uh or you're trying to fix a bug and you're getting bug reports, but the bug reports are going to things like Slack or linear or X, right? And these are external systems that aren't connected to your inner loop. So, I like to talk about this outer loop and the inner loop. Uh I don't know if I'm using the definition correctly but to me my inner

**中文**  做一个功能，嗯，或者你在试着修一个 bug，而且你在收到 bug 报告，但这些 bug 报告去往 Slack，或者 Linear，或者 X 这类地方，对吧？而这些是没有连到你的内环上的外部系统。所以，我喜欢谈这个外环和这个内环。嗯，我不知道我是不是把定义用对了，但对我来说，我的内

### [00:35:35–00:36:04]

**EN**  loop is like basically my engineers my agent engineers working on the code to an building towards an intent or snapshot of my intent right and the thing about that is that the snapshot can go stale right new information comes to light that I then have to be the proxy of and you know transfer that context to my agent so you know if you don't have these triggers pulling

**中文**  环，基本上就是我的工程师，我的 agent 工程师，在代码上工作，朝着一个意图，或者我的意图的一份快照去构建，对吧，而关于这件事的一点是，那份快照会过时，对吧，会有新的信息出现，然后我必须去做它的代理，并且，你知道，把那个上下文传给我的 agent，所以，你知道，如果你没有这些把

### [00:36:02–00:36:30]

**EN**  information back into the interloop, then you sort of have to play that role where you're you're off, you know, in Slack or X or or whatever and you're gathering context, right? You're getting context about bug reports, about feature requests, about, you know, something someone said about, you know, our backend infrastructure has some limitation, you know, all that information, you have to f that across

**中文**  信息拉回内环的触发器，那你就多少得扮演那个角色，你，你人在别处，你知道，在 Slack 里，或者在 X 里，或者在随便什么地方，你在收集上下文，对吧？你在拿到关于 bug 报告的上下文，关于功能请求的上下文，关于，你知道，某个人说到的，你知道，我们的后端基础设施有某种限制，你知道，所有这些信息，你都得把它们传

### [00:36:27–00:36:56]

**EN**  to your agent. So that's where I think like tools like Grogbot are really good because they help you automate the outer loop as well. And when you connect those two loops, it's very very powerful because now all of a sudden your agents have the ability to get context for this for themselves, right? If for example uh you know either through just as a simple example like maybe you have the

**中文**  给你的 agent。所以这就是我觉得像 Grok Bot 这样的工具真正好的地方，因为它们也帮你把外环自动化。而当你把这两个环连起来，就非常非常强，因为现在突然之间，你的 agent 有能力自己去拿到这个上下文，对吧？如果比方说，嗯，你知道，哪怕只作为一个简单的例子，比如也许你有那个

### [00:36:53–00:37:23]

**EN**  Slack MCP, right? Or you have uh your own harness, right, that you've built a Slack subscription into for a particular Slack channel. Now all of a sudden you can tell your agents, okay, subscribe to the Slack channel. Every time there's a uh, you know, bug report about something, go off and triage that thing, right? Go reproduce the issue, right? Using the verification skills that we've already spent time building and all of

**中文**  Slack MCP，对吧？或者你有，嗯，你自己的 harness，对吧，你已经为一个特定的 Slack 频道建好了 Slack 订阅。现在突然之间，你可以告诉你的 agent，好，去订阅这个 Slack 频道。每当出现一条，嗯，你知道，关于某件事的 bug 报告，就去把那件事分诊掉，对吧？去复现这个问题，对吧？用我们已经花时间建好的那些验证 skill，以及所有

### [00:37:21–00:37:50]

**EN**  those other skills that we've set up so that I have a lot of trust, right? I have a lot of trust that these agents can actually go off and understand the bug, you know, uh verify that the bug actually still exists on main and it wasn't something about you maybe the users setup or their data or maybe I don't know they didn't install a dependency or something like that like uh basically I

**中文**  我们搭好的那些其他 skill，这样我就有很多信任，对吧？我非常信任这些 agent 真的能自己跑去理解这个 bug，你知道，嗯，去核实这个 bug 在 main 上确实还在，而不是关于你的某件事，也许是用户的环境，或者他们的数据，或者也许，我不知道，他们没有安装某个依赖，或者类似的事情，像，嗯，基本上我

### [00:37:45–00:38:15]

**EN**  think uh creating that yeah creating those two loops and connecting them is really a very important part of the job these days. Um, especially if you are thinking about how to scale yourself. So, a big theme here is really just like always thinking about like what where am I the bottleneck in this process? Why do my agents need me, you know, to answer this question? I I always like to think about that. And so I try to think about

**中文**  觉得，嗯，把那个建起来，是的，把这两个环建起来，并且把它们连上，真的是如今这份工作里非常重要的一部分。嗯，尤其是如果你在想怎样放大你自己。所以，这里的一个大主题其实就是，一直在想，比如，在这个过程里，我这个瓶颈到底在哪里？为什么我的 agent 需要我，你知道，来回答这个问题？我，我总喜欢想这个。所以我试着去想

### [00:38:13–00:38:42]

**EN**  how do I actually get the agent to answer its own question, right? But not by hallucinating, not by guessing, but actually real data. And you know, a lot of people talk about this idea of a company brain, right, or a context graph. I feel like those terms are unnecessarily complex uh or even abstract. To me, it's just about um how do I take information that my agent needs that I would

**中文**  我怎样真正让 agent 去回答它自己的问题，对吧？但不是靠产生幻觉，不是靠猜，而是靠真正的数据。而且，你知道，很多人谈公司大脑这个想法，对吧，或者 context graph。我觉得那些术语不必要地复杂，嗯，甚至抽象。对我来说，这只不过是，嗯，我怎样把我的 agent 需要、而我本来会

### [00:38:39–00:39:07]

**EN**  otherwise have to go and pass it myself and just teach it how to do it, right? And that removes me from the equation. And so how I arrive at 2,000 or however many PRs is the fact that I have all these loops set up, right? And so uh it allows me to open chain restaurants, right? I can I can really parallels myself. So yeah, I'm not sitting there

**中文**  自己去拿、再亲手传给它的那些信息拿过来，并且直接教它怎么做，对吧？这样就把我从这道式子里去掉了。所以，我能做到 2,000 个，或者不管多少个 PR，是因为我把这些环都搭好了，对吧？所以，嗯，这让我可以去开连锁餐馆，对吧？我可以，我可以真正把我自己并行起来。所以，是的，我并不是坐在那里

### [00:39:04–00:39:34]

**EN**  creating 2,500 chats, right? Of course, it's really like these projects are um actually cursor has a new feature called projects which are these like coordinator agents. Um and so the coordination co coordinator agents are really good at sort of delegating and not doing work of their own but they manage and supervise like almost a list of tasks and they spawn sub agents to go and do them. And so I'm just constantly feeding context or teaching the agents

**中文**  创建 2,500 个聊天，对吧？当然，这其实是，这些项目，嗯，实际上 Cursor 有一个新功能，叫做 projects，就是这些协调者 agent。嗯，所以这些协调、这些协调者 agent，非常擅长去做委派，而不是做它们自己的活，但它们管理和监督几乎一份任务清单，并且它们会派生 sub agent 去做这些任务。所以我只是在不断地喂上下文，或者教这些 agent

### [00:39:32–00:40:01]

**EN**  how to get their own context and then they're going off and doing the work for me. Uh and really the big the last thing I'll say to this is like the big unlock for me for getting to 2,000 PRs is starting from the question and working backwards of how do I get to the point where my agent can merge its own code? Because the obvious thing people ask me when they when I tell them, "Oh, I shipped 2,000 and 2,500 pull requests

**中文**  怎样自己去获取上下文，然后它们就跑去替我把活干了。嗯，而关于这件事，我真正要说的最后一点，大的一点是，对我来说，能到 2,000 个 PR 的那个大的解锁，是从这个问题出发往回推：我怎样走到我的 agent 可以合并它自己的代码的那一步？因为，很明显，人们会问我的是，当他们，当我告诉他们，「哦，我交付了 2,000 和 2,500 个 pull request

### [00:39:59–00:40:27]

**EN**  last month." They'll be like, "How did you review that?" Right? That that's a lot of PRs to review. Like your team must hate you. >> Do do you mind if we go there in a second? Because a good question about that. >> Yeah. Yeah. Yeah. >> I want to like this analogy is great. I want to like deepen it a bit which is before if you're like manually initiating all those chats it's like you're bringing the orders to your chefs manually right whereas if you've got an agent sort of like doing the expo then

**中文**  是在上个月。」他们会说，「那你是怎么审查的？」对吧？那，那是很多要审查的 PR。就像，你的团队一定很讨厌你。主持人：你，你介意我们等一下再去谈那个吗？因为关于那个，有一个好问题。poteto：对。对。对。主持人：我想，这个类比很好。我想把它再挖深一点，也就是，以前，如果你是在手动发起所有这些聊天，就像你在手动把订单送到你的厨师那里，对吧，而如果你有一个 agent，有点像在做出餐协调（expo），那么

### [00:40:26–00:40:55]

**EN**  you're able to sort of run it yourself itself what is what does that concretely look like then you've got these sort of grock bots that are um subscribing to channels pulling in Slack messages and you it sounds like have a couple of coordinator agents or like chief of staff agents that like monitor that or something like when you look at your computer to manage your agents, what does it look like? >> Yeah. So, [clears throat] so uh this is I guess somewhat confusing but we're

**中文**  你就能让它自己把这件事跑起来，那么这具体看起来是什么样？然后你有这些 Grok Bot，嗯，在订阅频道，把 Slack 消息拉进来，而且听起来你有几个协调者 agent，或者像幕僚长（chief of staff）那样的 agent，在监视那个，或者类似的情况，当你看着你的电脑来管理你的 agent 时，它看起来是什么样？poteto：对。所以，［清嗓子］所以，嗯，我想这多少有点让人糊涂，不过我们

### [00:40:54–00:41:22]

**EN**  working on you know simplifying and unifying but so uh there's graphbot uh which or you know you can use other tools of course as well but I I largely think of these tools as like your outer loop. These are tools like you know Grabbot that have connectors right these are connectors I guess they a lot of people call them personal agents um but they're connectors to things like your email your calendar slack uh plaid I

**中文**  正在，你知道，做简化和统一，不过，所以，嗯，有 Grok Bot，嗯，或者，你知道，你当然也可以用别的工具，不过我，我很大程度上把这些工具看成你的外环。这些是，你知道，像 Grok Bot 这样带有连接器的工具，对吧，这些是连接器，我想很多人把它们叫做个人 agent，嗯，但它们是接到你的邮件、你的日历、Slack、嗯、Plaid 上的连接器，我

### [00:41:21–00:41:49]

**EN**  don't know like all these different services and they are a great source of pulling context in to your work so the same way that a human like you know if I were if I was a manager and I was leading a team of engineers years. Um, you know, like when I used to work in Netflix, one of the biggest things that managers would talk about was this idea of context not control, which funnily

**中文**  不知道，像所有这些不同的服务，而且它们是把上下文拉进你的工作的一个很好的来源。所以，同样地，一个人，比如，你知道，如果我是，如果我当年是一个经理，我带着一个工程师团队的那些年。嗯，你知道，像我以前在 Netflix 工作的时候，经理们会谈的最大的事情之一，就是给上下文而不是给控制（context not control）这个想法，说来有意思

### [00:41:45–00:42:14]

**EN**  enough, you know, has so much uh has so much uh carry over to the agents world. Uh, of you know, you you know, you you of course can drive to an outcome you want by control, right? Like by micromanaging, but what you want is to provide context instead, right? like teach the agent, teach your engineers how to be self-sufficient and then you don't have to micromanage them. >> Um, and so I see a lot of parallels

**中文**  的是，你知道，有那么多，嗯，有那么多，嗯，可以照搬到 agent 的世界里。嗯，就是，你知道，你，你知道，你，你当然可以靠控制来推向你想要的那个结果，对吧？比如靠微管理，但你想要的是改为提供上下文，对吧？比如教这个 agent，教你的工程师怎样自给自足，然后你就不必微管理他们。嗯，所以我看到很多平行之处

### [00:42:12–00:42:40]

**EN**  there. Uh, but yeah, graphbot. So, concretely, I have some graph bots that look at my Slack channels, look at my X, uh, or my emails, uh, or linear, and they're just constantly they have routines that subscribe. So they're constantly watching and I have I I'll tell them things like you know uh I'll watch for issues with uh bugs in the graphbot desktop app as an example. Uh

**中文**  在那里。嗯，不过是的，Grok Bot。所以，具体来说，我有一些 Grok Bot，它们看我的 Slack 频道，看我的 X，嗯，或者我的邮件，嗯，或者 Linear，而且它们就是一直在看，它们有去订阅的 routine。所以它们一直在看着，而且我有，我会告诉它们一些话，比如，你知道，嗯，以 Grok Bot 桌面应用里的 bug 为例，我会盯着那些问题。嗯

### [00:42:37–00:43:06]

**EN**  and whenever you find that send it to my cursor project. So one of the really cool things about grabbot is it connects to cursor. So cursor has uh like I I just mentioned this new feature called projects. And a project is really a uh again like a you get a coordinator agent that's in the cloud. It has its own computer and all it really does is like it's a manager of agents. It's like your executive chef, right? Your your chief

**中文**  而且每当你发现那种情况，就把它发到我的 Cursor project。所以 Grok Bot 真正很酷的一点是，它能连上 Cursor。所以 Cursor 有，嗯，像我，我刚才提到的这个叫做 projects 的新功能。而一个 project 其实是一个，嗯，再说一次，像是你得到一个在云端的协调者 agent。它有它自己的一台电脑，而它真正做的全部事情就是，它是这些 agent 的经理。它就像你的行政总厨，对吧？你的，你的幕僚

### [00:43:03–00:43:32]

**EN**  of staff. It doesn't do the work itself. It delegates and orchestrates and manages the work of other sub agents to you know that report to your chief your chief uh of staff. And it basically is responsible for driving the work forward and managing things and uh passing context to them. >> So if you get a sudden burst of issues, let's say you get 30 issues at once in

**中文**  长（chief of staff）。它自己并不干活。它委派、编排，并管理其他 sub agent 的工作，你知道，那些向你的幕僚长、你的幕僚，嗯，长（chief of staff）汇报的那些。而它基本上负责把工作往前推进，把事情管起来，嗯，并且把上下文传给它们。主持人：所以，如果你突然涌来一批 issue，比方说你一次来了 30 个 issue，在

### [00:43:31–00:43:58]

**EN**  one payload or something or very quickly the coordinator agent can figure it out and delegate. >> Yeah, exactly. It gets like uh you know 30 the 30 or so payloads and spawns a sub agent or a single coordinator agent. It can actually do a bunch of different topologies of agents and it will sort of figure out the best way to uh you know efficiently distribute the tasks to your

**中文**  一个 payload 里，或者来得非常快，协调者 agent 就能弄清楚，并且把事情委派出去。poteto：对，正是这样。它会拿到像，嗯，你知道，30 个，大约 30 个 payload，然后派生一个 sub agent，或者一个单独的协调者 agent。它其实可以摆出好几种不同的 agent 拓扑，而且它会多少想出最好的办法，嗯，你知道，把这些任务高效地分发到你的

### [00:43:53–00:44:21]

**EN**  team of agents. Um so I use uh cursor projects a lot um and I also use grapot a lot and cursor projects are my inner loop and grabbot is my outer loop. Grabbot takes all the context, external context, gives it to the projects because it can actually just send messages to those projects, right? You don't even have to open cursor. You can just tell your grabbot, okay, create a

**中文**  agent 团队。嗯，所以我很常用 Cursor projects，嗯，我也很常用 Grok Bot，而 Cursor projects 是我的内环，Grok Bot 是我的外环。Grok Bot 把全部的上下文，外部的上下文，交给这些 project，因为它其实可以直接给那些 project 发消息，对吧？你甚至都不用打开 Cursor。你只要告诉你的 Grok Bot，好，创建一个

### [00:44:18–00:44:47]

**EN**  project, right, for these series of tasks. They're all related, right? Maybe as an example, you know, you've had a uh [clears throat] a big burst of issues that are all about performance, right? Your app is slow uh and they're all connected, right? Maybe some of them even have a similar fix, right? But and you can certainly go off and just spawn one agent per task, but then you've lost that sort of thread between them, right?

**中文**  project，对吧，来做这一串任务。它们都是相关的，对吧？也许举个例子，你知道，你遇到过，呃，[清嗓子] 一下子涌来一大批全都关于性能的 issue，对吧？你的应用很慢，呃，而且它们全都连在一起，对吧？也许其中有一些甚至修法都差不多，对吧？不过，你当然也可以直接去，每个任务只拉起一个 agent，可那样你就失去了它们之间的那条线索，对吧？

### [00:44:45–00:45:13]

**EN**  And and you may duplicate work or you may not really think about the higher level problem. You know, sometimes when you you you solve bugs, you know, it helps to have multiple bug reports that are are slightly different because it helps you really, you know, zoom out and see actually, you know, the problem when I looked at this one report, I thought the bug was here, but actually when when I see the other multitude of bugs is actually up here, >> right? >> Yeah. Got you. So that that's why you have so many agents in that loop then,

**中文**  而且，而且你可能会把工作做重复，或者你可能并没有真正去想更高一层的问题。你知道，有时候当你，你，你在修 bug 的时候，你知道，手头有多份略有不同的 bug 报告是有帮助的，因为它帮你真正地，你知道，把视角拉开，然后看见，其实，你知道，这个问题，当我看这一份报告的时候，我以为 bug 在这里，但实际上当，当我看到另外那许多 bug 的时候，它其实在上面这里，主持人：对吧？ poteto：是的。主持人：明白了。所以这，这就是为什么你在那个环里有这么多 agent，

### [00:45:11–00:45:41]

**EN**  right? Because it's not just you have um like you have a bug report comes in, you spawn a single agent to look at that bug report. that a that single agent will be duplicating work with other um other agents, right? Because if there are multiple bug reports coming in through the same thing, that can be duplicated work. >> Yeah. >> Fascinating. [snorts] >> That's really fascinating. Okay. And so this just this endless series of triggers um coming from real users reporting real reports um builds up this

**中文**  对吧？因为这并不只是你有，嗯，就像一份 bug 报告进来，你拉起一个单独的 agent 去看那份 bug 报告。那个，那个单独的 agent 会和其他，嗯，其他 agent 做重复的工作，对吧？因为如果有多份 bug 报告从同一件事进来，那就会是重复的工作。poteto：是的。主持人：很有意思。[哼了一声] 这真的很有意思。好。于是这一连串没完没了的触发，嗯，来自真实用户在报告真实的报告，嗯，就积累起这个

### [00:45:39–00:46:09]

**EN**  sort of and accelerates the factory sort of adds more orders in. Other than bug reports, are there any other sources that you use for like um accelerating for pushing these PRs? >> Uh well, funnily enough, it's some of it comes from uh reading the code, too. So, I guess I have sort of uh well, so to clarify that, you know, the 2,500 PRs, they're not obviously like 2,500

**中文**  东西，并且让这个工厂加快，算是把更多订单加进去。除了 bug 报告以外，你还有没有别的来源，用来，比如，嗯，加快、用来推动这些 PR？ poteto：呃，说来有点好笑，其中有一部分也来自，呃，读代码。所以，我想我算是有，呃，好吧，先把这一点说清楚，你知道，那 2500 个 PR，它们显然并不像是 2500 个

### [00:46:06–00:46:35]

**EN**  features, right? they are a lot of the work actually is spent on gardening like another term that I really love. Uh where so I guess this is more important when you have a big team of engineers human engineers that you work with where and also this goes back a little bit to what I was talking about with the environment. You know, setting up a really good environment that doesn't just help you and your agents, but everybody on your team, right? Think of

**中文**  功能，对吧？它们，有大量的工作其实花在打理上，gardening，这是另一个我真的很喜欢的说法。呃，所以我想，当你有一支很大的工程师团队，那些和你一起工作的人类工程师，在那种情况下这一点更重要，而且这也稍微回到我之前讲环境时说的那些。你知道，搭一个真正好的环境，它不只帮你和你的 agent，也帮你团队里的每一个人，对吧？想象一下

### [00:46:33–00:47:03]

**EN**  a new hire who doesn't have a lot of context on all of your engineering practices joining your your team. And if you have a really good environment, they can be productive from day one, right? They can they don't have to like, you know, make open a bunch of lowquality PRs. They can start, you know, they can start just turning out really good code. Um, and uh, yeah, I think I I sort of lost my train of thought. >> I've got a I've got a followup, which is

**中文**  一个新人，他对你们所有的工程实践并没有很多上下文，加入你们的、你们的团队。而如果你有一个真正好的环境，他们从第一天起就能有产出，对吧？他们可以，他们不必去，你知道，弄出、打开一大堆低质量的 PR。他们可以开始，你知道，他们可以开始直接写出真正好的代码。嗯，而且，呃，是的，我想我，我有点把思路丢了。主持人：我有一个，我有一个追问，是

### [00:47:00–00:47:28]

**EN**  what's what are the mechanics of like triggering >> like how when when do you trigger a a agent to go and look at the code, right? Because some people might say, "Oh, let's just do that every hour or something or like on a chron job or what's >> Oh, yeah. Yeah. Yeah. Yeah. Uh I saw some of your recent tweets as well about like you know the some of the tweets you've been doing which are great for

**中文**  这种触发的机制到底是什么，比如你在什么时候、在什么时候触发一个 agent 去看代码，对吧？因为有些人可能会说，「哦，我们就每小时做一次之类的，或者像挂在一个 cron job 上，或者怎样。」 poteto：哦，对。对。对。对。呃，我也看到你最近的一些推文了，关于，你知道，你一直在发的其中一些推文，它们很适合

### [00:47:24–00:47:54]

**EN**  setting up your routines. Uh I have some routines like that as well. Um so uh one of them is uh like looking through just another simple example is you know [clears throat] React has a lot of foot guns. Um, so, uh, as as I'm sure you're aware. And so I have an agent that's just constantly looking for band patterns. And the interesting thing about that one is that

**中文**  用来把你的这些例行做法搭起来。呃，我也有一些那样的例行做法。嗯，所以，呃，其中一件是，呃，就像去翻看，再举一个简单的例子就是，你知道，[清嗓子] React 有很多 footgun。嗯，所以，呃，这一点我肯定你也清楚。于是我有一个 agent，它就在不断地寻找 bad patterns。而这一件有意思的地方是

### [00:47:52–00:48:22]

**EN**  I don't actually tell it to fix the issue first. I tell it to append it to a document. And then every couple of days I look at it and I see actually these are all the same thing, you know, and so that gives me, you know, you almost want like a buffer, a queue. Sometimes that's actually more effective than just spawning off a couple of like a lot of sub agents to fix every single thing because when you are [clears throat] in kind of pure execution mode and just trying to like you know f uh you know

**中文**  我其实并不先叫它去修这个问题。我叫它把这个问题追加到一份文档里。然后每隔几天我看一看，我就看见，其实这些全都是同一件事，你知道，于是这给了我，你知道，你几乎会想要一个缓冲区，一个队列。有时候这其实比直接派出几个、像是很多 sub agent 去把每一件单独的事都修好更有效，因为当你 [清嗓子] 处在一种纯粹的执行模式里，只是想要，你知道，去，呃，你知道

### [00:48:20–00:48:48]

**EN**  execute on the orders that are coming in very fast you sometimes miss the big picture. So sometimes having a buffer forces you to think about the big picture because you you you have these artifacts and things that you can look at as a human um and sort of use your own human judgment to or I guess you can use an agent to do that as well. But you give the [clears throat] agent and yourself a way to identify patterns,

**中文**  执行那些来得非常快的订单的时候，你有时候会错过大局。所以有时候有一个缓冲区，会逼着你去想大局，因为你，你，你有这些产物，有这些东西，你可以当作一个人去看，嗯，并且用你自己的人的判断去看，或者我想你也可以用一个 agent 来做这件事。但你给这个 [清嗓子] agent，也给你自己，一种把模式认出来的办法，

### [00:48:46–00:49:16]

**EN**  right, that you might otherwise miss if you're just only solving each bug at a time. And that's also really the benefit of having something like a chief of staff agent is uh it can see the forest right uh in addition to actually doing the execution. >> Fascinating. That's I mean my brain is exploding a bit there with the sort of chief of staff at the software factory. I might have to change some of the course that I'm filming next week.

**中文**  对吧，如果你只是一次只解决每一个 bug，这些模式你本来可能会错过。而有一个像幕僚长（chief of staff）agent 这样的东西，好处也真的在这里，呃，它看得见森林，对吧，呃，而不只是真的去把执行做了。主持人：很有意思。这，我是说，我的脑子在这儿有点炸开了，软件工厂里的这种幕僚长（chief of staff）。我可能得改一改我下周要拍的课程里的一些内容。

### [00:49:14–00:49:42]

**EN**  That's [laughter] >> uh all right. Let's talk about let's talk about review, right? because this is the reply that you get, you know, is >> did you read did you taste all 2500 of those dishes as they swept past you? >> And I assume the answer is a variety is a version of no. >> Yeah, I think you you you don't want to be in a position where you're not tasting your food ever again. Uh but you

**中文**  就是这样。[笑] 主持人：呃，好。我们来谈谈，我们来谈谈审查，对吧？因为你得到的那种回复，你知道，就是，你读过没有，那些菜从你面前掠过去的时候，那 2500 道你都尝过了吗？而我假定那个答案是「没有」的一个版本。 poteto：是的，我觉得你，你，你并不想处在一个你再也不品尝自己的食物的位置上。呃，但是你

### [00:49:41–00:50:10]

**EN**  also, you know, for scale, you cannot be tasting every single dish that comes out of your kitchen, especially if you have multiple restaurants. So it becomes more about sampling right and thinking about the processes in the same way that you know if I guess maybe this is where the the factory analogy is a bit more apt is you know as a quality supervisor on a factory you you can't look at every single item you sample right you take

**中文**  也，你知道，要上规模的话，你不可能把从你厨房里端出来的每一道菜都尝到，尤其是如果你有好几家餐厅。所以这就更多地变成抽样，对吧，并且去想那些流程，同样地，你知道，如果我猜的话，也许这正是工厂这个类比更贴切的地方，你知道，作为工厂里的一名质量主管，你，你不可能去看每一件东西，你是抽样，对吧，你去

### [00:50:07–00:50:35]

**EN**  you you you go in there every day and you look at the quality of the pull requests you look at the code that the agents are writing and you scrutinize it very rig rigorously and you think about all the inefficiencies, the bad patterns that the agents are doing and then you think about how to course correct the environment, right? Not not that single agent. Uh because if maybe if it if it

**中文**  你，你，你每天走进去，看这些 pull request 的质量，看 agent 正在写的代码，并且非常、非常严格地审它，然后你去想所有那些低效的地方、agent 正在做的那些 bad patterns，然后再去想怎样把环境的方向纠正过来，对吧？不是，不是那一个单独的 agent。呃，因为如果也许，如果它，如果它

### [00:50:33–00:51:02]

**EN**  was a one-off incident, it's fine. You know that maybe there's nothing to fix there. But if you actually notice that multiple agents are are having the same issue, right? They're taking the same shortcut. they're they're propagating the same workaround everywhere. Uh that's a sign that you should go off and think about how to uh amend your kitchen or your factory, right? Like thinking about your skills, your constraints,

**中文**  只是一次偶发的事，那没关系。你知道，也许那里并没有什么要修的。但如果你真的注意到，多个 agent 都，都有同一个问题，对吧？它们在走同一条近路。它们，它们在把同一个变通办法传播到每一个地方。呃，这就是一个迹象，说明你该转去想一想，怎样，呃，修整你的厨房，或者你的工厂，对吧？比如去想你的 skill、你的约束、

### [00:50:58–00:51:26]

**EN**  your lints, your type systems um and setting or adjusting it so that that problem doesn't happen again. And when you do that enough times, then you get to a place where the codebase is again like the environment is so constrained and so it guides you so well that you can just you can just step away, right? That's the dream. And I'll I'll definitely say it um it's very hard to

**中文**  你的 lint、你的类型系统，嗯，并且去设置它，或者调整它，让那个问题不再发生。而当你把这件事做上足够多次，你就到了这样一个地方，代码库又一次，就像这个环境被约束得那么紧，而且它把你引导得那么好，以至于你可以，你可以就走开，对吧？那就是那个梦想。而我，我肯定会这么说，嗯，要

### [00:51:24–00:51:54]

**EN**  get to this point. I don't want to sell this as like, you know, something that you can just do easily by using PAC. Like it takes a lot of time and effort to think about your code and where you see your agents failing and thinking very thoughtfully, intentionally and setting up guard rails and constraints so that they do the right thing by default. >> And you're not [clears throat] like if to go back to the software factory

**中文**  走到这一步是非常难的。我不想把这件事说成，你知道，一件你只要用上 pstack 就能轻松做到的事。它要花很多时间和力气，去想你的代码，以及你在哪些地方看见你的 agent 在失败，并且非常用心地、有意地去想，然后设上护栏和约束，好让它们在默认情况下就做对的事。主持人：而你并不是，[清嗓子] 就像，如果回到软件工厂

### [00:51:51–00:52:21]

**EN**  analogy, this isn't a dark factory, right? This the lights are on, right? >> It kind of is. Yeah, actually. >> Is it? >> Yeah. Well, it's dark in the sense that so um it's dark in the sense that well I think if my agents are merging their own pull requests it's sort of become dark where I go to sleep my agents now work 24/7 uh I have I have the equivalent of like more than 10 chiefs of staff right each working on a different area like for

**中文**  这个类比，这并不是一座 dark factory，对吧？这里的灯是开着的，对吧？ poteto：它有点是。是的，其实是。主持人：是吗？ poteto：是的。嗯，说它暗，是在这个意义上，所以，嗯，说它暗，是在这个意义上，嗯，我想，如果我的 agent 在合并它们自己的 pull request，它就算变得有些暗了，我去睡觉，我的 agent 现在 7 天 24 小时都在工作，呃，我有，我有相当于超过 10 个幕僚长（chief of staff），对吧，每一个在一块不同的领域上工作，比如

### [00:52:19–00:52:48]

**EN**  example I have one that's working on performance of the Grockbot desktop app I have one that's working on uh fixing bugs that users report I have one that's exploring rewriting it in a different language just for fun, you know, like what if what if, you know, just reimagining what what it would be if it was like a native app. It's just a toy. Um, but the idea is like yeah, I uh I when you spend the time setting up your environment, I've gotten to a point

**中文**  举个例子，我有一个在做 Grok Bot 桌面应用的性能，我有一个在，呃，修用户报上来的 bug，我有一个在探索用另一种语言把它重写一遍，只是为了好玩，你知道，就像假如，假如，你知道，只是重新想象，如果它像一个原生应用，它会是什么样子。它只是一个玩具。嗯，但这个意思是，是的，我，呃，我，当你把时间花在把环境搭好上，我已经到了这样一个点

### [00:52:46–00:53:14]

**EN**  where I review the pull request after it's landed, right? I I tell my agents full autopilot is is something that you can do in in PAC and that will trigger off this very intense rigorous verification loop where it will spawn a bunch of verifier agents for every pull request and it will fuzz right fuzzing meaning that it will actually run the application. It's going to click around and try to use it like a real human.

**中文**  在这个点上，我是在 pull request 落地之后才去审查它，对吧？我，我告诉我的 agent，full autopilot 是，是一件你可以在 pstack 里做的事，而那会触发这一套非常猛、非常严格的验证循环，它会给每一个 pull request 拉起一大批做验证的 agent，而且它会做 fuzz，对吧，fuzzing 的意思是它会真的把应用跑起来。它会到处点击，试图像一个真人那样去使用它。

### [00:53:11–00:53:41]

**EN**  look for regressions, look for bugs in your implementation and um it will try to find issues with the thing and then it will fix it itself. It'll do that again and eventually get the PR to a state where it can land. Uh so it does it does it is quite token intensive. You can tune this of course. Uh so you know instead of like 10 verifier agents you might do like one, right? Or you just tell the agent to

**中文**  去找回归，去找你这份实现里的 bug，而且，嗯，它会试着找出这东西的问题，然后它自己把问题修掉。它会再做一遍，最后把这个 PR 弄到一个可以落地的状态。呃，所以它确实，它确实，它相当消耗 token。你当然可以把这个调一调。呃，所以你知道，你可以不放像 10 个验证 agent，而做成像 1 个，对吧？或者你只是叫那个 agent 去

### [00:53:39–00:54:07]

**EN**  verify it's done work. But yeah, the key thing is the verification part is really the key piece that gives me a lot of confidence that I guess verification plus the environment, right? It's these the combination of these two things that allow me to step away and say agents go off and merge your thing. I'll review it in the morning by looking at my commit history >> and if I see problems, I go and course correct,

**中文**  验证它已经做完的工作。不过，是的，关键的一点是，验证这一部分真的是给我很多信心的那一块，我想，是验证再加上环境，对吧？正是这两样东西合在一起，才让我可以走开，然后说，agent 们去把你们的东西合并吧。我早上会看着我的 commit 历史来审查它，然后如果我看见问题，我就去把方向纠正过来，

### [00:54:03–00:54:31]

**EN**  >> right? And I'll go and revert or modify, add new link rules and whatever. Um, so it does it does take time to get to that point, but once you get it, oh, it's so it feels so magical. Uh, I I I tell people like I'm sleeping so much better now because, you know, it took it the very first day I turned on the sort of dark factory was very scary because I was like, "Ooh, what if I call the SE, right? What if I break something

**中文**  对吧？然后我会去回退，或者修改，加上新的 lint 规则，以及别的。嗯，所以它确实，它确实要花时间才能到那个地步，可一旦你到了，哦，它是那么，它让人觉得那么奇妙。呃，我，我，我跟人们说，比如我现在睡得好多了，因为，你知道，它，我把这种 dark factory 打开的第一天非常吓人，因为我当时在想，「噢，万一我打电话给 SE，对吧？万一我弄坏什么东西

### [00:54:30–00:54:59]

**EN**  overnight?" >> Uh, and it took a lot of it took a lot of uh bravery, I think, to do that, but >> somehow I did it. And yeah, now I'm in a place where my Asians are are merging their own code while I sleep. >> It sounds like >> I think it's dark in that sense. >> Yes, it's dark sometimes, right? You do sometimes. >> That's true. That's true. >> Because I think of a dark factory is like almost like if you take the original definition of Kapathy's vibe coding, right, which is the code almost

**中文**  过夜？」 呃，而且我觉得，那样做花了很多，花了很多，呃，勇气，但是不知怎么的我还是做了。而且，是的，现在我处在这样一个地方，我的 agent 在我睡觉的时候合并它们自己的代码。主持人：听起来像是 poteto：我觉得在那个意义上它是暗的。主持人：是的，它有时候是暗的，对吧？你有时候确实会这样。poteto：这话没错。这话没错。主持人：因为我心里的 dark factory，几乎就像，如果你取 Andrej Karpathy 对 vibe coding 的那个原始定义，对吧，也就是代码几乎

### [00:54:58–00:55:25]

**EN**  doesn't exist. You forget that code might be a thing. I think your approach is totally different from that, which is that code and the environment is essential. And if the code in the environment are bad, then you will get bad outputs. Garbage in, garbage out. So I I I think this is a this is a different thing. It's like, you know, the I don't know, maybe there's a dimmer switch or something, right? Like, you know, some parts of dark, some parts

**中文**  并不存在。你会忘掉代码可能还是一件东西。我觉得你的做法和那个完全不同，你的做法是，代码和环境是必不可少的。而且如果环境里的代码是差的，那你就会得到差的产出。垃圾进，垃圾出。所以我，我，我觉得这是，这是另一回事。它就像，你知道，这个，我不知道，也许有一个调光开关之类的，对吧？就像，你知道，有些部分是暗的，有些部分

### [00:55:24–00:55:53]

**EN**  were light. This is why the maybe the kitchen is a better analogy >> the restaurant because you know even as a as a as a [clears throat] restaurant you still might go to your restaurants every now and then to take a peek in taste the food right >> uh I think that's >> the idea of sampling instead of blocking I think is really important >> I think what would you say to people who are in I guess you're obviously in a pretty security conscious environment where you're working very security

**中文**  是亮的。这就是为什么，也许厨房是一个更好的类比。poteto：是餐厅，因为你知道，即便作为一家，作为一家，作为一家 [清嗓子] 餐厅，你仍然可能会时不时去你的那些餐厅看一眼，尝尝食物，对吧。主持人：呃，我觉得那就是 poteto：抽样，而不是拦住，这个想法我觉得真的很重要。主持人：我觉得，你会怎么跟那些人说，我想你显然是在一个相当注重安全的环境里，你做的工作非常注重安全

### [00:55:52–00:56:21]

**EN**  conscious >> mh [clears throat] Um maybe there are folks working in like um medical applications or law or finance or something. I think of the like some PRs are kind of like two-way doors which is you can merge it and then revert it, right? It's cheap back through. But there are some PRs that are one-way doors, right? That will cause data loss of some kind that will >> do something that can't be easily walked

**中文**  注重安全。poteto：嗯。[清嗓子] 主持人：嗯，也许有些人做的是，比如，嗯，医疗应用，或者法律，或者金融之类的。我是这么看的，有些 PR 有点像双向门，也就是你可以把它合并，然后再把它回退，对吧？退回去代价很低。但有一些 PR 是单向门，对吧？它们会造成某种数据丢失，它们会去做一件不能轻易走回来的事

### [00:56:19–00:56:46]

**EN**  back. How do you deal with situations where most of your PRs, let's say, are one-way doors? Like, is this something you just wouldn't recommend or like what do you think? >> Yeah, I think that's a really good question. I think that it all comes back to me to the quality of the verification that you're able to um get out of your agent. And I think for domains where the work is

**中文**  回来。遇到这样的情况你怎么处理，比如说，你的大多数 PR 都是单向门？比如，这是一件你就是不会推荐的事，还是说，你怎么想？ poteto：是的，我觉得这是一个非常好的问题。我觉得，这一切对我来说，都回到你能够，嗯，从你的 agent 那里得到的验证质量上。而且我觉得，对于那些工作是

### [00:56:44–00:57:13]

**EN**  verifiable, this is easier, right? and the the oneway doors become two-way doors in a sense. But I guess I don't know if you're working on something that is like extremely is very hard to verify programmatically then I think yeah you're definitely in a position where it's very hard to get to that point. Um so I do think like yeah verifiability of the domain is an

**中文**  可以验证的领域，这件事就更容易，对吧？而且那些单向门在某种意义上变成了双向门。不过我想，我并不知道，如果你在做的东西是那种极端地、非常难以用程序来验证的，那我觉得，是的，你肯定处在一个很难走到那一步的位置。嗯，所以我确实认为，是的，这个领域可不可验证，是一个

### [00:57:10–00:57:40]

**EN**  important aspect to be able to do this. Um and software engineering is just one of those things where it's quite verifiable in in a lot of cases maybe not totally um you know like other domains like mathematics I think are another example of not all of it of course but some aspects of mathematics can be verifiable if you write a proof for example um and so yeah I think it's a great question

**中文**  能够做成这件事的重要方面。嗯，而软件工程恰好就是这类事情之一，它在，在很多情况下相当可以验证，也许不是完全可以验证，嗯，你知道，像别的领域，比如数学，我觉得是另一个例子，当然不是它的全部，但数学的某些方面是可以验证的，比如说如果你写下一个证明，嗯，所以，是的，我觉得这是一个很好的问题

### [00:57:38–00:58:05]

**EN**  that I don't really have the answer to and I think that this is something the industry and us as engineers will have to figure out is you know my sort of uh hope and prediction for the future is that we'll see more and more interesting new agentoriented programming languages and one of the most fascinating ones that I've seen so far is this one called

**中文**  而我并没有真正的答案，而且我觉得，这是这个行业、以及我们这些工程师将不得不去弄清楚的事，你知道，我对未来的那种，呃，希望和预测是，我们会看到越来越多有意思的、新的、面向 agent 的编程语言，而到目前为止我见过的最让人着迷的其中一种，是这个叫做

### [00:57:59–00:58:27]

**EN**  bend bend d bend um and that language is one where it kind of marries programming with proofs right there used to be a time, you know, where you actually had to write your proofs in a different language. And proofs, by the way, for those uh who who aren't familiar is this idea of uh that you can sort of formally

**中文**  Bend 的，嗯，而那种语言是一种把编程和证明结合在一起的语言，对吧，以前有过一段时间，你知道，你实际上得用另一种语言来写你的证明。而证明，顺便说一下，对那些，呃，并不熟悉的人来说，是这样一个想法，呃，就是你可以比较形式化地

### [00:58:24–00:58:52]

**EN**  verify that some code is correct mathematically, right? Especially if you've written your code in a very functional programming way. uh but for the longest time you had to do that in a separate language like lean or tla+ or uh I'm blanking on some of the other other examples but uh like languages like that where you would construct the mathematical proof and

**中文**  从数学上验证某段代码是正确的，对吧？尤其是如果你用一种非常函数式的编程方式把代码写了出来。呃，可是在很长很长一段时间里，你得在一种分开的语言里做这件事，比如 Lean，或者 TLA+，或者，呃，另外几个例子我一时想不起来了，但是，呃，就是那样的语言，你会在里面把数学证明构造出来，并且

### [00:58:48–00:59:17]

**EN**  then use a solver essentially to det that that you've covered all the cases you don't have like a race condition or whatever. So yeah, I think trying to sum up the question, I think yeah, if you are in a position where you can figure out how your agents can truly verify the work in a way that gives you confidence, you can actually, you know,

**中文**  然后基本上用一个 solver，来确定你已经把所有情况都覆盖到了，你并没有像竞态条件之类的问题。所以，是的，我想试着把这个问题总结一下，我觉得，是的，如果你处在这样一个位置，你能弄清楚你的 agent 怎样才能真正验证这项工作，而且以一种让你有信心的方式，那你实际上可以，你知道，

### [00:59:14–00:59:44]

**EN**  uh have the PRs merge cuz if it compiles, right, if it if the proofs show you that it's correct, then why wouldn't you just merge it? Um but of course, yeah, not all the means are verifiable. Yeah, it's a tough one. Um, okay. I think we've got to think about wrapping up because we are nearly on the hour. Have you Have you got something after this? I mean, I've got something before I give my son dinner, but >> I I can go a bit longer after you.

**中文**  呃，让这些 PR 合并，因为如果它编译通过了，对吧，如果它，如果这些证明向你表明它是正确的，那你为什么不直接把它合并了呢？嗯，不过当然，是的，并不是所有的手段都是可以验证的。是的，这是个难题。嗯，好。我觉得我们得考虑收尾了，因为我们差不多到整点了。你，你这之后有事吗？我是说，给我儿子开饭之前我有一件事，不过 poteto：我，我可以再多留一会儿，随你。

### [00:59:42–01:00:11]

**EN**  >> Okay, let's let's go five minutes longer then. Um I think I just want to have one more question which is I think I want to ask how you see Pstack and how you see skills in general like in terms of we talked about this before we went on air which is like people think of as like my skills versus your skills and how do you combine frameworks

**中文**  主持人：好，那我们就再多五分钟。嗯，我觉得我只想再要一个问题，就是，我想问你怎么看待 pstack，以及你总体上怎么看待 skill，就像我们开播之前谈过的那样，也就是人们会把它想成我的 skill 对上你的 skill，以及你怎样把这些框架组合

### [01:00:09–01:00:38]

**EN**  together? How do you use Pstack with my stuff? like what should you take from each one and because I think I see skills as sort of just derived from process basically like they're just processes turned into words and I would love to know how you recommend people take Pstack and take my stuff as well and turn it into their own processes. I think you shared a tip actually today

**中文**  到一起？你怎样把 pstack 和我的东西一起用？比如每一样你该从里面取走什么，而且因为我觉得，我基本上把 skill 看成只是从流程里来的，它们就只是被写成了文字的流程，我很想知道，你建议人们怎样拿 pstack，也拿我的东西，再把它们变成他们自己的流程。我觉得你今天其实分享过一个窍门

### [01:00:37–01:01:04]

**EN**  that I thought was actually very relevant, which is this idea that you go off and look at your previous transcripts, right? And you sort of mine for information of, you know, your own look through your own your own prompts, right, to the agents where you correct them where you have to constantly intervene and uh you know take that higher level learning and turn that into a a reusable skill, right? So that

**中文**  而我觉得那个窍门其实非常相关，就是这个想法：你转去看你以前的那些转录，对吧？然后你算是从里面挖信息，你知道，你自己的，翻看你自己的、你自己给 agent 的那些 prompt，对吧，在那些地方你纠正它们，在那些地方你不得不不断地插手，而且，呃，你知道，把那个更高一层的心得变成一个、一个可以再用的 skill，对吧？好让

### [01:01:01–01:01:30]

**EN**  agents stop repeating that mistake. I think that uh PAC and your skills are very complimementaryary. I I totally agree with you that they're like a skill is really much just process. I mean it's just at the end of the day a skill is just English or or language. It's just markdown. >> Um and I think you can you can definitely weave them, combine them in a way that makes sense to you. But I do

**中文**  agent 不再重复那个错误。我觉得，呃，pstack 和你的 skill 是非常互补的。我，我完全同意你，一个 skill 其实在很大程度上就只是流程。我是说，到头来，一个 skill 就只是英语，或者，或者就是语言。它就只是 markdown。 poteto：嗯，而且我觉得你可以，你完全可以把它们织到一起，以一种对你说得通的方式组合起来。不过我确实

### [01:01:27–01:01:57]

**EN**  think that uh everyone should have their own set of knives, right? I I keep going back to the the the cooking analogy, but it's so apt because like, you know, every chef when they go to a different job, right, when they go to a different restaurant, they carry they bring their knives with them. The tools go with them, right? And so trust to me is really about trust in your own tools. And when you spend the time sharpening them and understanding them really, really well, you can do great

**中文**  认为，呃，每个人都应该有自己的一套刀，对吧？我，我总是回到这个，这个，这个烹饪的类比，但它实在太贴切了，因为，你知道，每一位厨师去到一份不同的工作的时候，对吧，当他们去到一家不同的餐厅的时候，他们带着，他们把自己的刀带在身上。工具是跟着他们走的，对吧？所以对我来说，信任其实就是对你自己这些工具的信任。而当你花时间把它们磨快，并且把它们理解得真的、真的很透，你就能做出很棒的

### [01:01:54–01:02:24]

**EN**  things. And everybody's skills and tool set is going to look different. you know, someone might find a lot of success combining, you know, like your grill me with docs, uh, or wayfinder skill with some of the execution skills in Pstack as an example. Some people might use more of your skills, some people might use more of my skills. I think at the end end of the day, it really just comes back to how much do you trust, you know, me and Matt, right? Like if you if you trust us both, of

**中文**  事情。而每个人的 skill 和工具集都会看起来不一样。你知道，有的人可能会在组合上很成功，你知道，比如把你的 grill-me 和 docs 合在一起，呃，或者举个例子，把 Wayfinder 这个 skill 和 pstack 里一些负责执行的 skill 合在一起。有的人可能会更多地用你的 skill，有的人可能会更多地用我的 skill。我觉得，到头来，到头来，这真的只是回到你有多信任，你知道，我和 Matt，对吧？比如如果你们，如果你们信任我们两个，那

### [01:02:22–01:02:52]

**EN**  course, use our skills, but I also encourage you to, you know, look at your own transcripts. Um um tell the agent to look through, you know, some of the all of the patterns that you've used, the times you've had to intervene, you know, suggest turning them into lint rules or new skills, right? The the past chats I I often say is like a a treasure trove of context because that, you know,

**中文**  当然就用我们的 skill，但我也鼓励你，你知道，去看你自己的转录。嗯，嗯，叫这个 agent 去翻看，你知道，你用过的其中一些、全部那些模式，那些你不得不插手的时候，你知道，建议把它们变成 lint 规则，或者新的 skill，对吧？那些，那些过去的聊天，我，我常常说，就像一个，一个上下文的宝库，因为那个，你知道，

### [01:02:47–01:03:15]

**EN**  it's it's like the process materialized, right? Like it's the real process. It's not an abstract idea in your head. it's the actual thing right and you can actually see how it happened in practice and extract so much information from that and there's so much so that I actually turn I have a skill in pac called recall which is exactly that um where this was a pattern where you know I was working in I was working

**中文**  它，它就像流程已经变成了具体的东西，对吧？就像它就是真实的流程。它不是你脑子里的一个抽象想法。它就是实际的那件事，对吧，而且你真的可以看见它在实践当中是怎么发生的，并且从里面抽出那么多信息，而且有那么多，以至于我实际上做成了，我在 pstack 里有一个叫做 recall 的 skill，它做的正是这件事，嗯，这当初是一种反复出现的情况，你知道，我当时正在做，我当时正在做

### [01:03:13–01:03:42]

**EN**  on a similar problem so specifically I was working on virtualization for the cursor application and there were a lot of bugs and so you know every time I started a new chat I was like ah this there's so much good context from the last one So, you know, I want to bring it over to the new chat. How do I do that? And that's where the transcript came, uh, you know, looking at the past transcript came about. And then recall was just a way for me to collapse and compress that workflow into a skill. So that I didn't have to just

**中文**  一个相似的问题，所以具体地说，我当时在给 Cursor 这个应用做虚拟化，而且有很多 bug，所以你知道，每一次我新开一个聊天，我都觉得，啊，这个，上一个里面有那么多好的上下文。所以，你知道，我想把它带到新的聊天里。我要怎么做？而这就是转录出现的地方，呃，你知道，去看过去的转录，就是这么来的。然后 recall 就只是一个办法，让我把那个工作流折叠并且压缩成一个 skill。这样我就不必再只是

### [01:03:40–01:04:09]

**EN**  say I didn't have to write a long essay every time. Go look at all these chats, right? And, you know, blah blah blah. So, I I largely think of skills, especially as agents get more capable as really encoding workflows. you know skills from last year were really more about like almost like implementation details like here are the exact script commands you know you should use right I think with the latest models you can just delete those parts and just really

**中文**  去说，我就不必每一次都写一篇很长的东西。去把这些聊天全都看一遍，对吧？而且，你知道，巴拉巴拉。所以，我，我大体上把 skill 看成，尤其是当 agent 变得更能干的时候，真正是在把工作流写下来。你知道，去年的 skill 其实更多地像是，几乎像实现上的细节，比如这是你该用的那些确切的脚本命令，你知道，对吧。我觉得有了最新的这些模型，你可以把那些部分直接删掉，然后真正地

### [01:04:07–01:04:36]

**EN**  focus on the workflow right until it it's more the skill becomes more like a series of steps a series of your process uh and I think over time we'll see that skills get smaller and smaller you know more compact Um, and yeah, they're very compatible. Or you can, you know, if you want, why not read our skills, right, and com and combine them in your of your own, right? Combine Wfinder with potato

**中文**  把注意力放在工作流上，对吧，直到它，它更，这个 skill 变得更像一连串步骤，一连串你的流程，呃，而且我觉得随着时间过去，我们会看到 skill 变得越来越小，你知道，更紧凑。嗯，而且，是的，它们非常兼容。或者你可以，你知道，如果你愿意，为什么不去读一读我们的 skill，对吧，然后把，把它们组合成你自己的，对吧？把 Wayfinder 和 poteto

### [01:04:34–01:05:04]

**EN**  mode and make your own custom mode, right? Like like skills are the the thing I love about skills that is are that they're so malleable. You can do anything you want. It's just language. >> Absolutely. There's nothing magical in them, right? They're just words. And >> Exactly. If if there is any magic in them, it's just the words chosen and the phrases used and the thinking that's been done to turn those like take abstract process and turn them into

**中文**  mode 合在一起，做出你自己的自定义 mode，对吧？就像，就像 skill 是，我喜欢 skill 的那一点就在于，它们是那么可以任你去改。你想做什么都可以。它就只是语言。主持人：完全是这样。它们里面并没有什么魔法，对吧？它们就只是一些词。而且 poteto：正是。如果，如果它们里面有任何魔法，那也只是选用了哪些词、用了哪些说法，以及已经做过的那些思考，去把那些，就像拿起抽象的流程，把它们变成

### [01:05:02–01:05:31]

**EN**  language. And once that thinking has been done, then it's just there. It's available. It's on the surface and you just nick it. Um Lauren, thank you so much. This has been glorious. >> Yeah, this has been super fun. I really enjoyed talking to you. Hope we can do it again. I'd love to do it again. I'd love to do it again. Absolutely. Um yeah, we'll check in in uh >> I don't know. Yeah, Monday. Let's do it. >> Yeah, [laughter] let's do it. Part two. >> Well, thank you so much. I'm going to

**中文**  语言。而一旦那个思考已经做完，它就在那儿了。它是拿得到的。它就在表面上，你拿走就行。嗯，Lauren，太感谢你了。这实在很精彩。 poteto：是的，这超级好玩。我真的很喜欢和你聊。希望我们还能再来一次。我很想再来一次。主持人：我很想再来一次。当然。嗯，是的，我们再联系，在，呃 poteto：我不知道。对，星期一。就这么定。主持人：对，[笑] 就这么定。第二部分。嗯，太感谢了。我要

### [01:05:29–01:05:36]

**EN**  close the stream here. Laura and I will uh chat a little bit and stay here. But thank you guys so much for watching. The glorious.

**中文**  在这里把这场直播结束掉。Lauren 和我会，呃，再聊一会儿，并且留在这里。不过非常感谢各位收看。这场精彩的。
