# Why The Harness Matters More Than The Model | YC Paper Club · 中英对照逐字稿

- 原节目：Y Combinator
- 英文原始来源：https://www.youtube.com/watch?v=n9xKblqyQ28
- 中文译制版入口：https://www.xiaoyuzhoufm.com/episode/6a9f788ea0210c197dcf3b8c
- 时长：01:00:11
- 方法与限制：英文来自已验证的原始节目 transcript/caption；中文由 Codex 逐段翻译，未做逐字人工校对，公开引用前请回到英文原文与音频复核。

## 中英对照逐字稿

### [00:00:01–00:00:03]

**EN**  [music]

**中文**  [音乐]

### [00:00:08–00:00:38]

**EN**  Welcome to YC Harness Club. I worked really hard. It took me like 10 its with Gemini to have Harsha riding a lobster. Gemini didn't want to do it, but we figured it out. Okay, welcome to harness night. Um, first order of business. Uh, how does everyone like the new YC paper club look? Well, this is like an idea that I kind of blurted out at EV, our head of design

**中文**  欢迎来到 YC Harness Club。我花了很大力气。大概用 Gemini 迭代了十次，才做出 Harsha 骑龙虾的图。Gemini 一开始不肯做，但我们还是搞定了。好，欢迎来到 harness 之夜。嗯，第一件事。大家觉得新的 YC Paper Club 外观怎么样？这个想法其实是我随口跟我们的设计负责人 EV 说的

### [00:00:35–00:01:03]

**EN**  at YC, and two days later, she came back with this, and I'm like, "This is amazing." So, Ev in the back there, please take a bow. [applause] Very exciting. All right, why harnesses? Um, I mean, it's just a rapper. This is just scaffolding. This is just like prompt engineering. Um, why would this be at all worthy of a night? Uh, here's a great little Reddit that was only a

**中文**  她在 YC，两天后她就拿出了这个，我当时就说：「这也太棒了。」所以后面那位 Ev，请鞠一躬。[掌声] 非常激动。好，为什么要讲 harness？嗯，我的意思是，它不就是一层 wrapper 吗。这不就是 scaffolding。这不就是 prompt engineering。为什么这值得专门办一晚？这里有一条很典型的 Reddit，就在

### [00:01:01–00:01:28]

**EN**  month ago, which is actually uh the most aggressive. I'm not sure this kind of prompt engineering belongs at a top tier machine learning conference. Um, here's another great one. Uh, context engineering is not a research problem. Um, so I think harnesses have long been belittled uh as um subpar research. yet it is uh literally gives us an 18% bump

**中文**  一个月前，而且语气相当冲。有人说：我不确定这种 prompt engineering 配不配出现在顶级机器学习会议上。还有一条也很典型：context engineering 不是研究问题。所以我觉得 harness 长期被贬低，被当成不够格的研究。可它实打实地带来了 18% 的提升

### [00:01:26–00:01:55]

**EN**  in the difference between harness one and harness two. Um, and as Seth will tell us is the difference between getting ArcGI to work and not. And so this is a uh obviously worthy of some amount of research. And I think if you look at this the classic meter plots release date and how long uh an agent can be running. Um, literally we go a lot of this progress has been because of harnesses. And so I call this the static

**中文**  这是 harness 一和 harness 二之间的差距。Seth 等会儿也会讲，这就是 ARC-AGI 能不能跑通的差别。所以这显然值得做一些研究。再看这张经典的 METR 图：发布日期，以及一个 agent 能连续跑多久。很多这样的进展，实际上就是因为 harness。所以我把这一阶段叫做静态

### [00:01:53–00:02:22]

**EN**  harness era where there isn't self-improvement on the harness. And then this later latest latest um maybe the last 6 months has all been on the self improving harnesses and think and we'll get through all these and it's I actually saw this plot in a presentation by the CEO of trajectory. I actually really liked it. was talking about how the we we keep measuring perplexity and you know this somewhat correlates to IQ and and how intelligent the model is and

**中文**  harness 时代，harness 本身还没有自我改进。再往后、最近、大概最近六个月，全部都在做会自我改进的 harness。我们会把这些都过一遍。这张图我是在 Trajectory 的 CEO 的分享里看到的，我很喜欢。他在讲：我们一直在测 perplexity，它和 IQ、和模型有多聪明有一定相关

### [00:02:20–00:02:48]

**EN**  so we keep pushing more and more up this intelligence uh IQ dimension but we're not leveraging test time experience very much so we're generating all this test time experience we have a new domain uh and it does and we don't really quickly adapt and if you remember one of the first YC paper clubs we did um I've been doing this experiment where um the as I increase the number of samples online how do you actually learn from just

**中文**  所以我们一直在这个智力、IQ 维度上往上推，但几乎没有用上 test-time experience。我们会产生大量 test-time experience，遇到新的 domain，却没法很快适应。如果你们还记得最早几期 YC Paper Club，我做过一个实验：当我在线增加样本数量时，你到底怎么从

### [00:02:46–00:03:13]

**EN**  batch size one we don't really don't have that structure we have ICL and then once ICL gets saturated after just loaded like 40 or 50 it doesn't actually improve um on the valet at all then you have to go to Laura small rank then you go to Laura big rank then you go to SFT and it's kind of weird that we have this like different training procedures and so I think that's really where harnesses are shining right

**中文**  batch size 为 1 里真正学到东西。我们其实没有这套结构。我们有 ICL，ICL 在大概塞进 40 或 50 条之后就饱和了，validation 上完全不再涨。然后你得上小 rank 的 LoRA，再上大 rank 的 LoRA，再上 SFT。我们有这么多套不同的训练流程，这其实挺怪的。我觉得 harness 真正发光的地方就在这儿

### [00:03:10–00:03:39]

**EN**  And what Arc AGI exposes, that's kind of the main uh um uh point of ARGI is how quickly it adapts to uh a new problem, a new a new distribution and does well in it. And so, um, ArcGI actually went through the batch with me, winter 26, and we helped Greg, you know, look at this and like the amount of thought and attention, as I mentioned last time, that goes into these these these games to make sure they're all orthogonal

**中文**  ARC-AGI 暴露出来的，也就是它的核心，是：面对一个新问题、一个新的分布，它能多快适应，并且表现得好。ARC-AGI 其实跟我一起过了 Winter 26 这期 batch，我们帮 Greg 看过这些。就像我上次说的，这些游戏花了大量心思，就是为了保证从游戏一到游戏二，技能是正交的

### [00:03:37–00:04:06]

**EN**  skills from game one to game two, so that it's isolating this fluid intelligence measure. Um really uh Claude uh Opus was one of the the first that was actually verified um on the the hold out on the private that no one else has access to other than Greg and Chalet. Um and the best that they got was 30%. And just with um some harness uh this thing that doesn't deserve any research just some wrapper and some

**中文**  这样才能把流体智力这个指标单独测出来。Claude Opus 是最早在 hold-out、在那份除了 Greg 和 Chollet 谁都拿不到的私有集上被验证过的模型之一。他们最好也就到 30%。而只靠某种 harness，就是这种据说不配做研究的东西，只是一层 wrapper 和一些

### [00:04:03–00:04:32]

**EN**  scaffolding we can get to 95 and AVO from Nvidia got to 100%. So, Prime Agent and Nvidia, which both recently just came out. And the other thing I want to add, so when uh Carpathy launched his auto researcher thing in March, I want to say it was um I forked it and I was playing around with it. And all I wanted to do was make like a little user interface to kind of see what's happening and track it and and I ended up building a harness

**中文**  scaffolding，我们就能到 95；NVIDIA 的 AVO 到了 100%。所以是 Prime Agent 和 NVIDIA，两边都是最近刚出来的。另外我想补一句：大概三月，Karpathy 发布了他的 auto researcher，我 fork 了一份自己玩。我本来只想做个小界面，看看它在干什么、跟踪进度，结果一不小心做出了一个 harness

### [00:04:31–00:05:00]

**EN**  by accident. I didn't mean to, but it was just like I wanted to see it. And basically what it is is you specify a purpose. And in this example, which is actually a true one I gave, is like diffusion LM don't beat ARLM. But maybe if I ensemble, if I shard the diffusion LM into a bunch of different ones because there's such high arithmetic intensity per GPU on a diffusion model versus AR that I can actually um in aggregate by sharding them, I can get

**中文**  我不是故意的，我只是想看见它在干什么。它本质上就是：你指定一个 purpose。这个例子是我真给过的：diffusion LM 打不过 AR LM。但也许我做 ensemble，把 diffusion LM 切成很多份，因为 diffusion 模型每块 GPU 上的算术强度比 AR 高得多，所以把它们分片之后，在总体上看，我也许能得到

### [00:04:59–00:05:27]

**EN**  better better results. And so I just give this to a per as the purpose to the to my uh uh to my suite of agents um my swarm of agents. I give it some seed ideas. I want to vary the ensemble size shard at different amounts 100 times, 10 times, five times, whatever. Um and then I specify a valometric. Um and maybe I want to do GSMAK and like GPT2 setup or something

**中文**  更好的结果。于是我就把这句话当作 purpose，交给我那一套 agent、我的 swarm。我再给一些种子想法。我想改 ensemble 的规模，按不同分片量来切：100 倍、10 倍、5 倍，随便。然后指定一个 validation metric。也许我想跑 GSM8K，配一个 GPT-2 那种设置，之类的

### [00:05:25–00:05:53]

**EN**  like that. And then I have a scoping agent that will kind of look up papers that have that are are similar, look up GitHub repos. Um, we'll give it to a PI agent named Chris Ray. Um, who will then uh give it to a research agent um named John South Khan. Uh, there he is. Um, and then he'll work on it and then he'll start doing some stuff, give it to a council for some help. That's where I come in and me and Yaso will give you

**中文**  然后我有一个 scoping agent，会去查相近的论文、查 GitHub 仓库。再交给一个叫 Chris Ré 的 PI agent。他再交给一个叫 John 的 research agent。他就在那儿。然后他开始干活，再拿去给一个 council 求助。这就是我出场的地方，我和 Yaso 会给你

### [00:05:51–00:06:21]

**EN**  some feedback and then you work on it a bit more and then whenever you're ready and Chris Ray will kind of keep tabs on you, keep nagging you every every hour or so then it goes to this author agent to say okay now start stop freeze the idea start writing ablations and um and then start writing the paper and so then what this has turned into so this is the scoping agent uh and then you have this cockpit that kind of you can view from anywhere we do the uh tail scale up so you can actually view this URL from

**中文**  一些反馈，你再接着做。等你准备好了——Chris Ré 会一直盯着你，大概每小时催你一次——就会交给这个 author agent，说：好了，先停，把这个 idea 冻住，开始写 ablation，然后开始写论文。于是这就变成了：这是 scoping agent，然后你有一个 cockpit，在哪儿都能看。我们开了 Tailscale，所以你可以从任何地方打开这个 URL

### [00:06:19–00:06:49]

**EN**  anywhere and see how it's progressing and and talk to it right there. Um, it just sends you email updates if you build the whole 1B thing in here. And then it actually starts publishing some papers and like I've read the papers. They're actually like good. They started out in March, April like not so good like you know May or whatever. But and now I just basically give eight ideas to eight H100 nodes and each one has eight H100s. and just keep going. And then I check in and they give give me

**中文**  看它进展到哪了，还可以直接跟它说话。如果你在这里把整个 1B 的东西跑起来，它还会给你发邮件更新。然后它真的开始发论文。我读过那些论文，其实写得不错。三月、四月刚起步时不太行，大概五月左右还一般。但现在我基本上就是把八个 idea 丢给八台 H100 节点，每台有八张 H100，让它们一直跑。我再去看一眼，它们就把论文交

### [00:06:47–00:07:17]

**EN**  back papers. And it's kind of wild what we're what we're dealing with there. And what's changed really is just the scaffolding that you put on top of it to allow it. And I I didn't spend an enormous amount of time on this, but this is largely how I I do a lot of my research, at least the the initial idea. Um, and so yeah, so I'll just give this six different ideas of things I want to try and just leave it alone and let it rip. And so the things that we can now do just because of harnesses on the same exact weight file is just wild. Um, so I

**中文**  回来。我们现在面对的这件事，其实挺离谱的。真正变了的，只是你叠在上面、让它能这么干的 scaffolding。我并没有在这上面花巨量时间，但这基本上就是我做很多研究的方式，至少是最初那个 idea 阶段。所以对，我会丢给它六个我想试的 idea，然后放手让它狂跑。同一份 weight 文件，只因为有了 harness，我们现在能做的事情已经很夸张了。所以我

### [00:07:15–00:07:43]

**EN**  looked on I spent the the weekend just in prep for this reading um a whole bunch of the classic literature from self-refined to reflection to Voyager to all the tool former and I just wanted to just do like a my my shot at a five minute like history how did we get here? Um I think it was it's kind of important and I don't think a lot of people that are entering AI now kind of like have the context. So this is not going to be in chronological order. It actually

**中文**  为了准备这场，整个周末都在读经典文献，从 Self-Refine 到 Reflection，到 Voyager，到 Toolformer。我就想用五分钟，按我的理解讲一遍：我们是怎么走到这一步的？我觉得这很重要，很多现在才进 AI 的人其实没有这段上下文。所以我不会按时间顺序讲。按时间顺序来讲或来展示

### [00:07:41–00:08:10]

**EN**  doesn't make sense to to teach it that way or show it that way. I'm not I'm going to skip a lot of papers and I may course grain the papers absurdly. So, apologies. So, the initial harness the GPT2 February 2019 is a while not end of sequence loop. That's and then we basically have top P sampling and then we have the environment and that's the harness. And there's really not that much. There's no tool calling, there's no skills, there's nothing like that.

**中文**  其实并不合适。我会跳过很多论文，也可能把论文讲得非常粗。提前道歉。最初的 harness，是 2019 年 2 月的 GPT-2：一个 while not end-of-sequence 的循环。然后基本上是 top-p sampling，再加上环境，这就是 harness。真的没多少东西。没有 tool calling，没有 skills，什么都没有。

### [00:08:07–00:08:35]

**EN**  And so that's the V0ero harness. And so basically if you have this GSMAK example, you're a math, remember the system prompt. You are a math teacher. This is like the persona stuff we used to have. Um there's a context. Susie has five bucks, she spends three. How much does she have now? And then there's just no chain of thought. It was just like four hashes and two end of state and end of sequence that is uh the GS MK format. We measure uh what happens after the

**中文**  这就是 V0 的 harness。假如你有这样一道 GSM8K 例题，你是数学老师，还记得当时的 system prompt。你是一名数学老师。那就是我们以前搞的 persona。然后有一段 context：Susie 有五块钱，花了三块，她现在还剩多少？当时没有思维链。就是四个井号，再加 2，然后 end of sequence，这就是 GSM8K 的格式。我们量的是四个井号后面

### [00:08:33–00:09:03]

**EN**  four hashes. We get accuracy and you get a plus one or minus one if it's wrong. And then the entire last six years has been giving more functionality into the harness into a static harness. And so we said okay well what if we give that back into the context a bunch of examples like that like well here's an example and let's say now we'll switch it to four and one and hopefully we'll have um from that previous example we'll say oh okay this makes sense I can will help to

**中文**  出现了什么。对了加一，错了减一，得到准确率。过去整整六年，就是往 harness 里加更多功能，加进一个静态 harness。于是我们说：那如果把一堆这样的例子塞回 context 呢？比如这是一个例子，现在改成四和一，希望它能从前一个例子里反应过来：哦，原来是这样，这能帮我

### [00:09:01–00:09:29]

**EN**  learn and actually this is with a fshot learners paper um back in July in 2020 and then we had chain of thought and says well directly predicting hash hash 2 is maybe difficult what we'll do is we'll smear the the compute the logic over many more tokens and we'll train it to get to that two now rather than just output the two and that was the chain of thought idea uh very cool and that was all context innovations uh and output

**中文**  学会。这其实就是 2020 年 7 月那篇 few-shot learners 论文。然后有了思维链：直接预测井号后面的 2 可能太难，那我们把计算、把逻辑摊到更多 token 上，训练它走到那个 2，而不是直接输出 2。这就是思维链的想法，非常酷。这些都是 context 上的创新，以及 output

### [00:09:28–00:09:57]

**EN**  space innovations action space innovations um and then then we came up with these tool former and webgbt webgbt actually came out first then tool former which is the idea of giving it tools and it can call a tool that's just really just a JSON object with a bunch of things specified but in this example just let's say subtraction instead of me calculating in the weight file what is 5 - 2 I can just call Python and call sub 53 and it'll tell me two and so that's

**中文**  space 的创新、action space 的创新。然后有了 Toolformer 和 WebGPT。WebGPT 其实先出来，然后才是 Toolformer。想法就是给它工具，让它能调工具。工具其实就是一个写明了一堆字段的 JSON 对象。这个例子里就说减法：与其在 weight 文件里自己算 5 减 3，我可以直接调 Python，调用 sub(5, 3)，它告诉我是 2。这

### [00:09:54–00:10:23]

**EN**  pretty cool and you can expose all the tools in the system prompt and that's where tools come from and then we figured out megpt which is big the one of the coolest tools is being able to read and write to your own context and so before then all we could do is just append append append append to the context now we said what if I actually give you create read, update, delete on the context itself and we'll se separate just this one little chunk called a

**中文**  就很酷。你还可以把所有工具都暴露在 system prompt 里，工具就是这么来的。然后我们搞出了 MemGPT，这很重要。最酷的工具之一，是能对自己的 context 做读写。在那之前，我们只能往 context 上 append、append、append。现在我们说：如果我给你对 context 本身做增删改查呢？我们单独划出一小块，叫做

### [00:10:19–00:10:46]

**EN**  memory that you'll be able to to update. And then Voyager said, okay, well, we have these tools, but like what if I want to chain together these tools to achieve a task and then I learn it and how do I distill it back into the the system prompt to make it learn forever? And this is the skill this where skills kind of came about. And uh they did this on Minecraft. Um and this the Voyager paper and it's very cool paper. This is

**中文**  memory，你可以去更新它。然后 Voyager 说：好，我们有这些工具了，但如果我想把这些工具串起来完成一个任务，再把它学会，怎么蒸馏回 system prompt，让它永远记住？这就是 skill，skills 大概就是从这儿来的。他们是在 Minecraft 上做的。Voyager 这篇论文非常酷。这基本上就是

### [00:10:44–00:11:11]

**EN**  like largely now what a skill is. And so I have this skills.md and I have the name of it that I can go and search and here's the procedure. And then uh intercode this idea of like again in action space innovation if I can actually output code uh then I can basically now I have on the-ly uh tools or on the fly skills. I actually don't know if it's a considered a tool because a tool is technically an API. Am I

**中文**  现在 skill 的样子。所以我有一个 skills.md，有名字，我可以去搜，下面是步骤。然后是 InterCode，又是 action space 上的创新：如果我能直接输出代码，那我就相当于有了即时工具，或者说即时 skills。我其实不太确定这算不算 tool，因为 tool 从技术上讲是一个 API。我输出的函数

### [00:11:09–00:11:39]

**EN**  outputting a function as a tool or a skill? I actually am still unsure. Um and then the react came first then self-refine then reflection. But this idea if I have multiple um agents that have different roles and they can help self-improve self-improve on the context uh then I can um get smarter and smarter. So, I'll have take some action. In this example, I I I flipped the three

**中文**  算 tool 还是 skill？我到现在都没想清楚。然后是 ReAct 先出来，再是 Self-Refine，再是 Reflection。但这个想法是：如果我有多个不同角色的 agent，它们能在 context 上互相帮助、自我改进，那我就能越来越聪明。于是我会采取某个动作。这个例子里，我把 3

### [00:11:35–00:12:05]

**EN**  and the five. Whoops. Um, I sent it to an internal evaluator to say, is this right? No, it doesn't look right. I can either, you know, keep looping here or I can go to the actual environment to get back a reward signal. Um, I go here uh to the reflection. They'll say, "Hey, you actually flip these two, go back, and I can improve the result." Um and this is the one of the first like uh uh ideas of like being multi- aent uh being

**中文**  和 5 写反了。哎呀。我把它送给内部评估器：这对不对？不对，看起来不对。我可以继续在这里循环，也可以走到真实环境里拿 reward 信号。我走到 reflection 这儿。它会说：「嘿，你把这两个写反了，回去改。」我就能把结果改好。这是最早的多智能体想法之一，也就是

### [00:12:03–00:12:32]

**EN**  let the letting the agent reflect on its own um output and improve it. And then this idea of multi- aent goes even further where I can actually spawn one of the tools could be I can spawn a a set of sub aents um and they they persist in these persistent ripples and they'll be able to be running and I can interact with them and I'll have the sub aent list. I can invoke those ones and I can keep adding to that the the launch

**中文**  让 agent 反思自己的输出并改进。多智能体这个想法还可以再往前走：其中一个工具可以是，我能拉起一组 sub-agent。它们活在持久的 REPL 里，能一直跑，我可以跟它们交互。我会有一份 sub-agent 列表，可以调用它们，也可以继续往上加，去 launch

### [00:12:29–00:12:58]

**EN**  sub aents and then RLM went even crazier to allow this in a recursive fashion so that um and I exposes this RLM query so I keep uh recursively calling the RLM query to uh solve a larger class of problems and that the leafs whenever I want I can call the LM query to spawn that LM agent as well and then I have this main orchestrator agent that's running all that and that's what I call like harness v1 one, this whole thing is

**中文**  这些 sub-agent。然后 RLM 更猛，允许用递归的方式来做。它暴露了 RLM query，我不断递归地调用 RLM query，去解更大一类问题。在叶子节点，我随时可以调 LM query，再拉起那个 LM agent。然后有一个主 orchestrator agent 在跑这一切。这就是我所说的 harness v1。这整套东西

### [00:12:56–00:13:24]

**EN**  like static harnesses. I'm not improving on the system prompt. I'm not um updating the harness itself. And this is this you, of course, we're going to have GStack on the top. He's the one that lets us do all this. So, we thank you, Gary. Love GStack. Um but there's some other ones that deserve a call out as well. And basically what that means to summarize all this, you have some agent spec, you have some system prompt here. You say, "How many turns am I allowed? How many tool calls am I allowed?" You don't want to allow

**中文**  都还是静态 harness。我没有在改进 system prompt，也没有在更新 harness 本身。当然，最上面我们会有 GStack。能让我们做这一切的就是他。所以谢谢你，Gary。爱 GStack。不过还有一些也值得点名。把这些总结一下，意思就是：你有一份 agent spec，这里有一份 system prompt。你规定：我允许多少轮？允许多少次 tool call？你不会想允许

### [00:13:22–00:13:51]

**EN**  infinite. You specify a tool list. you specify a skills list, the sub agent list, and that's largely what the V1 um is, and then you put it in a loop. And this can be spawned on a prompt if I'm asking it to do something in my Slack channel, which we'll hear about QM, which is like I use it every day. It's the team that made it is here. It's super exciting. Um and or it's on a crunch and just wakes up every hour and it decides to do work just like you

**中文**  无限次。你指定 tool 列表、skills 列表、sub-agent 列表，这大体就是 V1，然后放进一个循环。它可以由一句 prompt 拉起来，比如我在 Slack 频道里让它做事。等会儿会讲到 QM，我每天都在用。做它的团队就在现场，非常兴奋。或者它挂在 cron 上，每小时醒一次，自己决定要不要干活，就像你

### [00:13:49–00:14:16]

**EN**  know, anyone else. Um, and so there's some session management, there's a loop, there's this context compilation. We're actually creating the context. I'm putting uh uh all that into an LM call. I get back uh the action. And then I I may or may not have some tools that I need to invoke and append back into the context. And that's basically harness v1. And then the cool part, this is the most exciting part where we're seeing a

**中文**  其他人一样。所以这里有 session 管理，有循环，有 context compilation。我们其实是在构造 context。我把这些塞进一次 LM 调用，拿回 action。然后我可能要调一些工具，再把结果 append 回 context。这基本上就是 harness v1。然后酷的部分来了，也是我们现在看到最多进展的地方

### [00:14:14–00:14:42]

**EN**  lot of advancements, um, where you're letting the harness itself learn. either we're learning the system prompt or we're learning the harness itself, which is very trippy. And so, one of the famous ones that I actually wanted them to talk, but they're actually running a 150 person DSPY meetup tonight and in in this in San Francisco. They couldn't make it. Um, but I've done some podcast with them before and a super great community. Is this DSPY? So,

**中文**  就是让 harness 自己去学。要么学 system prompt，要么学 harness 本身，这很玄。其中一个很有名的，我本来想请他们来讲，但他们今晚在旧金山办一场 150 人的 DSPy meetup，来不了。我以前跟他们做过播客，社区非常棒。这就是 DSPy。所以

### [00:14:39–00:15:09]

**EN**  demonstrate, search, uh, predict. uh they have this idea of of basically uh taking a train set a small set of examples and then learning the optimal system prompt. So I keep iterating, iterating. I can't back prop through that process, but I can do uh something called genetic programming where I'm finding candidates, I'm merging candidates, and with some merge rule, I'm evaluating, seeing what happens, and I keep working and working. And I basically gives me CRUD over the

**中文**  Demonstrate、Search、Predict。他们的想法基本上是：拿一个训练集，一小批例子，去学最优的 system prompt。我不断迭代、再迭代。这个过程没法 backprop，但我可以做所谓的遗传编程：找候选、合并候选，按某种合并规则评估，看结果，然后再接着干。这基本上就是给了我对

### [00:15:07–00:15:36]

**EN**  system prompt itself, and allows me to choose any system prompt. And then Darwin machines actually go a step further. Not only are you allowed to change the system prompt, but you're allowed to change the harness itself, the harness code that is actually running. And so you can imagine basically what happens is you have this archive of um many different uh agents which is harness and system prompt. Uh and you sample from them. You push them

**中文**  system prompt 本身的 CRUD，让我可以选用任意 system prompt。Darwin machines 又往前走了一步。你不但能改 system prompt，还能改 harness 本身，也就是正在跑的 harness 代码。可以想象过程是这样：你有一份档案，里面是很多不同的 agent，每个 agent 就是 harness 加 system prompt。你从里面采样，把它们推

### [00:15:34–00:16:03]

**EN**  uh through you'll actually evaluate how they did on some fit fitness function. Uh you'll uh add it back into this archive state. And I skipped over the the self modify. You can actually you have a a meta harness that actually allows the the agent to modify its own uh harness so that it can become a different harness and then that basically loops around loops around and eventually you get better and better agents over time. And then the meta harness which is the main harness is to

**中文**  过去，用某个适应度函数评估它们的表现，再加回这份档案状态。我刚才跳过了自我修改。你其实有一个 meta harness，允许 agent 修改自己的 harness，让它变成另一个 harness。然后这就转起来、再转起来，时间一长，agent 会越来越好。而这个 meta harness，也就是主 harness，要做的是

### [00:16:01–00:16:30]

**EN**  produce harnesses, right? Which is a really meta concept. Um and so this is the output space where you're configuring multi- aent context compilation. Um, and you're you're doing this and you have allow CRUD on all of it. And so it just keeps adding more and more uh CRUD into the harness code, all the green that you see here, the uh meta prompt, the system prompts of all the agents, how many agents there are. Um, and you kind of grow this uh this meta

**中文**  生产 harness，对吧？这是一个非常 meta 的概念。所以这是 output space：你在配置多智能体的 context compilation。你在做这些，并且对这一切都开放 CRUD。于是它不断往 harness 代码里加更多 CRUD，就是你们在这儿看到的那些绿色：meta prompt、各个 agent 的 system prompt、有多少个 agent。你就这样把这个 meta

### [00:16:28–00:16:57]

**EN**  harness over time. And then one of the authors of this paper is the lead author is actually here tonight. Super exciting. You got to drive down and talk to him a lot of it about um continual harness, which is I love it. And they go even a step further where one they add they they add some extra color on the classes of memory and so they add this history thing. You know he'll he'll go into a bunch bunch more detail on that memory breakdown. But the coolest part I

**中文**  harness 慢慢养大。这篇论文的第一作者今晚就在现场。非常激动。你得开车过来，跟他好好聊聊 continual harness，我超喜欢。他们又往前走了一步：给 memory 的类别加了更多划分，加了这个 history。他等会儿会把 memory 的拆分讲得很细。但我觉得最酷的部分

### [00:16:55–00:17:23]

**EN**  think is for the classic RL people. I see Robert back there. He definitely would would enjoy this Dagger style online learning where you can actually update the weight file itself. So you're actually doing a test time training on the LLM based on a small amount of examples that you just learned which I think is actually a huge huge important research direction that we should get working. Anyway, that's it. How you do all

**中文**  是给经典 RL 的人看的。我看见后面的 Robert 了，他肯定会喜欢这种 DAgger 风格的在线学习：你真的可以更新 weight 文件本身。也就是根据刚学到的少量例子，对 LLM 做 test-time training。我觉得这是一个极其重要、我们应该把它做起来的研究方向。好，就这些。大家还好吗

### [00:17:23–00:17:50]

**EN**  right so um tonight we have uh three authors. Uh Ben got food poisoning this morning so he couldn't make it. It was very sad but we have three tremendous authors. uh Seth uh student under shein is I say it right uh at Princeton a researcher at Prime Intellect and the author of Prime Agent um John Sadvalone who's a the first time we've had a a

**中文**  好，今晚我们有三位作者。Ben 今早食物中毒，来不了，挺遗憾的，但我们仍有三位非常厉害的作者。Seth，Princeton 的学生，导师我念得对吧，也是 Prime Intellect 的研究员，Prime Agent 的作者。还有 John，这是我们第一次有

### [00:17:47–00:18:15]

**EN**  call back um pres presenter very excited about that PhD under Chris an Aelia um and he's going to be he's the author of open Jarvis with Ivonica uh from my lab uh hazy which is a personal uh open Jarvis and which I think is really cool. And then for Josh and Rean who Josh just got promoted to be head of YC labs which is really exciting which I think deserves a little round of applause as well.

**中文**  返场的演讲者，这点我特别兴奋。他在 Chris Ré 手下读博士。他是 open Jarvis 的作者，和 Ivonica 一起，来自我的实验室 Hazy。这是一个个人向的 open Jarvis，我觉得非常酷。然后是 Josh 和 Rean。Josh 刚升任 YC Labs 负责人，这很值得再鼓一次掌。

### [00:18:13–00:18:43]

**EN**  [applause] Um be talking about QM and so QM there at YC we had this general agent that we use and they came out with QM a month ago and it is like meaningfully uh better and is so functional and I use it every day. All right, that's all I got. Thank you so much. [applause] >> Hi everyone. Uh, great to be here tonight. Uh, my name is Seth and I'm going to be presenting Prime Agent,

**中文**  [掌声] 他们会讲 QM。在 YC，我们本来有一个大家在用的 general agent，他们一个月前推出了 QM，明显更好，也好用得多，我每天都在用。好，我就讲这些。非常感谢。[掌声] >> 大家好。今晚很高兴能在这儿。我是 Seth，接下来讲 Prime Agent，

### [00:18:40–00:19:09]

**EN**  which is a self-improving ROM harness. Um, my job here tonight, I think, is to try to convince all of you to take a very first principles style approach to thinking about how you build your harness. So, we're going to be very, very basic here to begin with. If you think about and I love the uh introduction right from of all the background literature um which I thought was so cool I think is very complimentary the how I think about this as well um but if you think about just

**中文**  一个会自我改进的 RLM harness。我今晚的任务，是想说服大家用非常第一性原理的方式来想：你该怎么建自己的 harness。所以一开始我们会非常、非常基础。刚才那段背景文献的介绍我特别喜欢，我觉得它和我的思路非常互补。但如果你只看

### [00:19:06–00:19:35]

**EN**  the raw LLM itself it's actually just this sequential processor that you know has some fixed weights and has some visible context and it's it's taking some tokens in and it's putting some tokens out uh to make the next decision. Um we don't really think of an LLM that way these days. We have a set of files that it has access to. We give endless programs and tools for it to use. Um, you can even message other sessions of

**中文**  裸的 LLM 本身，它其实只是一个序列处理器：有固定的 weights，有一段可见的 context，吃进一些 token，吐出一些 token，用来做下一个决定。我们现在不太会这样看 LLM。我们会给它一组能访问的文件，给它无穷无尽的程序和工具。你甚至可以给正在跑的其他

### [00:19:34–00:20:04]

**EN**  LLMs that are going on and create sub agents in order to do all of these very cool things. But, but in its basics, it's just tokens in, tokens out. It's a neural network making a prediction. The harness itself is the layer between the LLM and the world that adds things like this persistent state tools and compute. Uh, here we have a diagram of how we think about prime agent. um from the human's perspective. So you open up your prime agent uh on your computer just like you would cloud code, codec, pi,

**中文**  LLM session 发消息，创建 sub-agent，去做所有这些很酷的事。但在最根本上，它就是 tokens in、tokens out。就是一个在做预测的神经网络。harness 本身是 LLM 和世界之间的那一层，加上持久状态、工具和算力这类东西。这是一张我们怎么看待 Prime Agent 的图，从人的视角。你在电脑上打开 Prime Agent，就像打开 Claude Code、Codex、Pi 一样

### [00:20:02–00:20:30]

**EN**  etc. Um and it gets you into this agents view and this agents view is an overview of all the agents that you have going on for all your parallel sessions with like a very tight like tight summary of what you're using and then you can hop into one of those and check it out. there you're going to be at this root session and this [clears throat] root session is your basically the project orchestrator over all of these different sub aents that it's controlling and you don't have to ask it to start sub agents it will leverage sub agents when they're useful

**中文**  等等。然后你会进入这个 agents 视图。这个视图是你所有并行 session 里正在跑的 agent 总览，带一份很紧的摘要，说明你在用什么。然后你可以点进去看某一个。你会落在这个 root session 上。[清嗓] 这个 root session 基本上就是项目的 orchestrator，管着它控制的那些不同的 sub-agent。你不必让它去启动 sub-agent，有用的时候它自己会用。

### [00:20:28–00:20:57]

**EN**  um and these are all programmatically called because this is all based on uh the recursive language model principle where everything is in this IPython shell um all your tools all your memories all your sub aents uh and and we expose uh for further coordination these uh these messaging paradigms. So you can manage all of these um and then these agents can directly interact with your environment which could just be like the programs and files on your computer or running an H200 node cluster

**中文**  这些都是程序化调用的，因为整套都建立在 Recursive Language Model 原则上：所有东西都在这个 IPython shell 里，你的工具、memory、sub-agent。为了进一步协调，我们还暴露了这些消息范式。所以你可以管理所有这些。然后这些 agent 能直接跟你的环境交互，环境可以只是你电脑上的程序和文件，也可以是跑着的 H200 节点集群

### [00:20:56–00:21:24]

**EN**  maybe for your auto research. Each of these agents are then backed by this persistent dam on your computer. Uh this is so that you know when you close uh your laptop or you you control C out of the session, it's still running in the background. You have to actually stop the session so that make sure that you're continuing to running. And then we also have these other features we exposed from continual harness where it's able to provide like live CRUD operations on all of the components that

**中文**  也许是给你做 auto research 的。每个 agent 背后都有你电脑上的一个持久 daemon。这样你合上笔记本，或者对 session 按了 Ctrl+C，它仍在后台跑。你必须真正停掉 session，否则它会继续跑。我们还从 continual harness 里暴露了另一些功能：它能对所有组件做实时 CRUD

### [00:21:20–00:21:49]

**EN**  we mentioned um in order to manage its uh memory skills sub agents um persistent and and prompt its own system prompt persistently. The way I like to think about all of this context that we're building up is that we have this sort of almost like like here I have like L1, L2, L3. is like a cache, right? It's like what is the most accessible information that we're working with and at the very like fastest like readily available information, you know, the

**中文**  来管理它的 memory、skills、sub-agent，并且持久地改它自己的 system prompt。我喜欢这样看我们堆起来的这些 context：差不多像这里的 L1、L2、L3，像一层缓存。也就是：我们正在用的信息里，什么是最好拿的。在最快、最现成的那一层

### [00:21:48–00:22:16]

**EN**  models to be able to retrieve that really quickly. It's the model weights. So, everyone always wants to get all the information in the model weights. Um, but then we said, okay, well, maybe we don't have all the information because we don't want to have to fine-tune every single time to update because that's very expensive. So, we have this uh active input context. So, we're using lots and lots of tokens on the input. we might have some in context examples like we've seen previously in order to add to these different capabilities but at a

**中文**  模型要能很快取到，那就是 model weights。所以大家都想把信息全塞进 weights。但后来我们说：也许不必全放进去，因为我们不想每次更新都微调，那太贵了。于是我们有了 active input context。我们在输入上用掉大量 token。我们可能放一些 in-context 例子，就像前面看到的，用来补上这些不同能力。但到了

### [00:22:13–00:22:42]

**EN**  certain point we run out of context. Um and so the very like earliest form of harnesses that we've seen that are still used to this day even by those who say we want the most minimal harness possible is compaction because compaction is a very generalized tool for the agent to be able to uh summarize its own context history in order to work past its context length working window. You can think about um once we go beyond like what are directly like inputs and

**中文**  某一刻，context 就用完了。所以我们见过的最早、直到今天还在用的 harness 形态——就算那些声称只要最极简 harness 的人也会用——就是 compaction。因为 compaction 是一种很通用的工具，让 agent 能总结自己的 context 历史，从而越过 context 长度这个工作窗口。可以这样想：一旦我们超出模型直接的输入和

### [00:22:41–00:23:10]

**EN**  outputs from the the model here into this L2L3. You might be familiar with the L3 which is more of the dispatch state. So if you're working with a file system, you can read and write from uh main memory. Uh if we're at the L2, which is I could think uh at a means in between uh what the active context is and working with your file system, you might have a live uh ripple, which could just be running things directly in Python uh an IPython shell like you're in a Jupyter notebook. And all of those

**中文**  输出，进入这个 L2、L3。你们可能熟悉 L3，它更像 dispatch 状态。如果你在跟文件系统打交道，可以从主存里读写。如果在 L2，我会把它想成 active context 和文件系统之间的一层，你可能有一个活的 REPL，可以直接在 Python 里跑东西，一个 IPython shell，就像在 Jupyter notebook 里。而那些

### [00:23:08–00:23:38]

**EN**  variables are saved directly in your RAM. your agent can then programmatically manipulate them and run all sorts of programs directly on the information there saving tons of tokens rather than putting it directly into context. Uh you can also create sub agents and it's the same thing you're basically saving context here because you can task the agent with a specific set of information in order to perform some operations and the report back at the end. What is interesting here is what we talked about compaction for the active context, right? You have your

**中文**  变量直接存在 RAM 里。你的 agent 可以程序化地操作它们，直接在那些信息上跑各种程序，省下大量 token，而不必把它们塞进 context。你也可以创建 sub-agent，道理一样：你其实是在省 context，因为你可以给 agent 指定一小份信息去执行操作，最后再汇报回来。有意思的是，我们刚才说的 active context 上的 compaction，对吧？你有你的

### [00:23:36–00:24:04]

**EN**  context history. This is helping to update it over time. So you can continue to leverage this. But once we go beyond this, we need to be thinking about how are we doing these update. We talked about CRUD. How are we do beyond just creating reading? How are we updating and deleting our context over time beyond our uh or the state over time beyond the context line. Uh I like to think of this at the ripple is this aentic garbage collection where we're just cleaning up the variables in our

**中文**  context 历史。这能帮它随时间更新，让你能继续用这一层。但一旦超出这一层，我们就要想：这些更新怎么做。我们讲过 CRUD。不只是创建和读取，我们怎么随时间更新和删除 context，或者说删除 context 窗口以外的状态。我喜欢把 REPL 这一层想成 agent 式的垃圾回收：我们在清理自己状态里的变量

### [00:24:03–00:24:32]

**EN**  state as well as like what sub agents could be used. And then so to make sure that our RAM doesn't crash my computer laptop every day. And then on top of that, we have this uh notion of refinement where we're updating and deleting the skills and memories and prompts that are stored on your system. You can think of that so that way you don't crash your actual uh out of space on your hard drive as well. And so this very much is a here's how we express this thing and here's how we revise it over time. The other perspective that I

**中文**  以及哪些 sub-agent 还能用，免得 RAM 每天把我的笔记本搞崩。再往上，我们有 refinement 这个概念：更新和删除存在系统上的 skills、memory 和 prompt。可以把它想成：这样你的硬盘也不会被撑爆。所以这很大程度上就是：我们怎么表达这件事，以及我们怎么随时间改它。我很喜欢的另一个视角

### [00:24:31–00:24:59]

**EN**  really like to think about and I really trying to push because harnesses are almost going towards this like agentic operating system that we're creating uh is when I think of it um more metaphorically here is that when you look at the raw LLM it kind of looks more like a touring machine where you have this ticker tape uh and you have all these instructions that are going in and then it's performing some set of operations and going out. But when you look at a harness it's looking a lot more vono like a vonoyman computer.

**中文**  我也一直在推，因为 harness 几乎在走向我们正在做的这种 agent 操作系统。更比喻一点说：看裸的 LLM，它更像一台图灵机，有一条纸带，指令进去，做一组操作，再出来。但看一个 harness，它更像冯·诺依曼计算机。

### [00:24:57–00:25:26]

**EN**  you're able to do these read and write operations on external memory and that makes it much more powerful and another class of problems than just what a touring machine is able to express on its own. And so yeah, the idea of like how do you build a good hardness? You want it to be the most expressable thing you can imagine. So some some like early harnesses before it gets into the data flywheel where the models can do themselves are very specific. Plan, act,

**中文**  你能对外部 memory 做读写，这让它强大得多，能表达的问题类别也和图灵机自己能表达的不一样。所以，怎么建一个好的 harness？你希望它是你能想象到的、表达力最强的东西。早期一些 harness，在进入模型能自己转起来的数据飞轮之前，是非常具体的。规划、行动、

### [00:25:22–00:25:52]

**EN**  critique, do these exact specific um steps. Well, now har uh the models are able to do that themselves. You can imagine like we we don't have like a react loop that we necessarily need to explicitly impose. The models kind of have natively uh figured this out. But what they haven't figured out is how to um you know they have to be able to have the expressibility to call compact. They have to be able to have a Python ripple so they can run programs. Um they have to have the ability to programmatically

**中文**  批评，按这些非常具体的步骤来。而现在，模型自己就能做这些了。可以想象，我们不必再显式地套一个 ReAct 循环。模型差不多已经原生学会了。但它们还没学会的是：它们必须有调用 compact 的表达能力；必须有一个 Python REPL，才能跑程序；必须能程序化地

### [00:25:50–00:26:18]

**EN**  create sub agents and access state and have different feedback mechanisms. Th those are model controlled expressibility features and if you removed one of those you're actually removing a capability that it won't be able to do otherwise. The way we manage um and I'm sure you're all familiar with the RLM paper uh from my co-author Alex um fantastic bit work. What we do beyond what was in the RLM paper is we think about the age sub aents as these persistent subsessions.

**中文**  创建 sub-agent、访问状态，并拥有不同的反馈机制。这些都是由模型控制的表达力特性。你拿掉其中任何一个，就是在拿掉一种它否则做不到的能力。我们的管理方式——我相信大家都熟悉我的合作者 Alex 的 RLM 论文，那篇工作非常出色。我们在 RLM 论文之上多做的是：把这些 sub-agent 想成持久的 subsession。

### [00:26:15–00:26:42]

**EN**  So each the parent station can create uh spin up a new RLM sub aent and each of these are then emitted. They run some task and then they finish and report back to some end state to the parent session. These are then idle. They're still working in your RAM. At any point, the parent session can then send a message to one of the sub aents to continue working and it has all that good context that you built up over time so that you're not missing information

**中文**  所以父 session 可以创建、拉起一个新的 RLM sub-agent，它们被发出去，跑某个任务，完成后把某个结束状态汇报给父 session。然后它们闲着，但还在你的 RAM 里。父 session 随时可以给其中一个 sub-agent 发消息，让它继续干。它还保有你一路攒下来的那些好 context，这样你不会丢信息

### [00:26:40–00:27:10]

**EN**  or have to reuse information that was already developed in a prior context. And then of course you know we don't want to use a lot of RAM. So we can move them offloaded uh in an inactive state which then can be called back at any time by messaging them in this persistent sub agent setup. I talked a little bit about uh continual harness uh which we have in a a prior paper of mine which talks about cuding the entire uh harness state. Um this is another feature that we want in our coding agents leverage all of our prior

**中文**  也不必重用先前 context 里已经长出来的信息。当然我们也不想占用太多 RAM，所以可以把它们卸载成非活动状态，之后随时发消息再召回来，这就是这套持久 sub-agent 的设定。我刚才提了一点 continual harness，那是我之前一篇论文，讲的是对整个 harness 状态做 CRUD。这也是我们希望编码 agent 具备的能力：用上我们先前所有的

### [00:27:08–00:27:37]

**EN**  history. So you can imagine like some set of trajectories where they have some actions and outcomes or something happened um at a at each turn. And so we just kind of want to expose the ability for the agent to leverage all that information in order to update what the future harness is going to look like. Are we do we need to change our system prompt? Do we need to create some skills? And skills I think of as a set of instructions or a program in order to achieve some some specific goal. uh

**中文**  历史。可以想象有一组轨迹，每一轮都有一些动作和结果，或者发生了什么。我们只是想把这种能力暴露给 agent，让它用上全部这些信息，去更新未来的 harness 会长成什么样。要不要改 system prompt？要不要创建一些 skills？我把 skills 理解成一套指令或一段程序，用来达成某个具体目标。

### [00:27:35–00:28:03]

**EN**  memory which could just be long-term storage about things that are important as well as the sub aent specifications that we talked about in this very persistent manner. Were there uh certain sub aents that we want to reuse at a later time because the context is useful and just having the ability to do this kind of reflection or refinement um over time. It's very powerful for the models to have. They're not perfect at this right now, but this is one of the the capabilities that you want to you want

**中文**  还有 memory，可以只是关于重要事情的长期存储，以及我们刚才说的、以非常持久的方式保存的 sub-agent 规格。有没有某些 sub-agent 以后还想复用，因为那份 context 有用？再加上随时间做这种 reflection 或 refinement 的能力。这对模型来说非常强。它们现在做得还不完美，但这正是你希望

### [00:28:00–00:28:30]

**EN**  to build your harness such that it is a bit better than what the current models are able to do. So then you can get those reasoning traces and use that to leverage your next iteration of model and they'll be able to handle the harness and be able to bootstrap themselves into a higher and higher performance. One of the coolest features that we have in uh Prime Agent um that we we've had since the beginning of when I was working on this, this is one of the first things I added um is the ability to message between any any two

**中文**  把 harness 建成比当前模型能力稍高一点的东西。这样你就能拿到那些推理轨迹，用来推动下一轮模型，它们就能驾驭这个 harness，并把自身 bootstrap 到越来越高的性能。Prime Agent 里最酷的功能之一，从我一开始做这事就有，也是我最先加的之一：任意两个

### [00:28:27–00:28:57]

**EN**  agents um within like some nuclear family setup, parents, children, uh siblings. Um and the reason why I did this is because I was I was constantly trying to figure out what's the best way to like myself to manage all of the agents I have doing everything for me in five different directions, five billion different directions every day. Um, and it would be so much better if they could just like share their contacts directly with each other and coordinate. And turns out that's fantastic for like typical software engineering and long horizon jobs as well. Uh, the last thing

**中文**  agent 之间能发消息，在一种核心家庭结构里：父、子、兄弟。我之所以这样做，是因为我自己一直在想：我每天让这些 agent 朝五个方向、五十亿个方向替我干活，最好的管理方式是什么。如果它们能直接共享各自的 context、互相协调，会好得多。结果发现，这对典型的软件工程和长程任务也同样非常好。最后一件

### [00:28:55–00:29:23]

**EN**  that we look at when it comes to how did we want to design our harness is we were really thinking about long horizon performance. I want to go run some jobs and I don't want to have to babysit my agents the entire time and when when I'm ready to come back and check in, I can check in with them and see what's going on. And this is a perspective that I also really lack seeing in a lot of the evaluations that we're looking at. Uh a lot of times if you run a model for not enough time or say, oh well the model

**中文**  我们在设计 harness 时看的，是长程表现。我想去跑一些任务，不想全程盯着我的 agent。等我准备好回来查看时，我可以去跟它们对一下，看看进展。这个视角，我在很多现有评测里几乎看不到。很多时候，如果你跑模型的时间不够，或者说：哦，这个模型

### [00:29:21–00:29:49]

**EN**  stopped working after this amount of budgets, but then this other model kept working with using more budgets. Well, first of all, you're not even using the same fixed expenditure to compare the models. But second of all, that could also be hiding performance that you're missing. Uh the way that I look at long horizon performance eval is that I want to see what's the practical plateau. At what point will we only get incremental gains in performance as I throw more test time tokens at it? I have a couple experiments that I'm going to show after

**中文**  在这么多预算之后就不干了，另一个模型用了更多预算还在继续。首先，你比较模型时甚至没用同一笔固定开销。其次，这也可能把你错过的性能藏起来了。我看长程表现评测的方式是：我想看到实际的平台在哪。从哪一点开始，我再砸更多 test-time token，性能也只是增量上涨？等会儿讲完我们怎么做成

### [00:29:48–00:30:15]

**EN**  we've shared design philosophy here about how we created uh prime agent. Um we're going to talk a little bit about test time scaling and uh our our TI results as well as looking at um you know does is it actually helpful and why is it actually helpful for our information management for the ripple that we're working on these long contexts and then um when we have these really really long like almost ultra horizon long horizon uh tasks uh how do we sustain these like multi-day work and

**中文**  Prime Agent 的设计理念之后，我会展示几个实验。我们会讲一点 test-time scaling，以及我们的 TI 结果；再看它对信息管理、对我们在这些长 context 上用的 REPL，到底有没有用、为什么有用。然后，当我们面对这些非常非常长、几乎是超长程的任务时，我们怎么把这种持续多日的工作撑住

### [00:30:13–00:30:42]

**EN**  like what actually goes on when we have these refinements um over like these settings that can last like a week at a time or more. So, this is a result that you probably all seen. We actually have a one additional data point that we added here that we didn't include in our original result uh just to compare across harnesses. We solved this uh we went out, we're trying to figure out what is the the best uh eval that people care about these days when we're running our harnesses and we're like, "Oh, we should do RKGI." I like, "Oh, yeah. Yeah, I remember. I I ran some results

**中文**  以及在这些一次能持续一周甚至更久的设定里，refinement 实际在发生什么。这是一个你们大概都见过的结果。我们这里多加了一个原来结果里没有的数据点，只是为了跨 harness 比较。我们当时在想：现在跑 harness，大家最关心的评测是什么？然后说：哦，应该做 ARC-AGI。我心想：对，我记得，我跑过一些结果

### [00:30:40–00:31:09]

**EN**  with continual harness and we got 20% with um Gemini Flash uh or sorry, Gemini Pearl." Um so, I I think we can get at least 20%. people who think that's really cool that our like general harness that didn't even like wasn't even structured for ARHI did really well. So I went online I was like okay I need to find a good system prompt because I don't want to make sure that we're losing information. So I found a another community leaderboard called prolong and I just grabbed their system prompt and I was like okay I'm going to grab their system prompt forget the rest and I'm just going to throw this

**中文**  是用 continual harness，Gemini Flash——抱歉，Gemini Pro——拿到了 20%。所以我想至少能到 20%。有人觉得这很酷：我们这个通用 harness 甚至不是为 ARC-AGI 专门设计的，却表现得很好。于是我上网，心想得找一份好的 system prompt，别把信息丢掉。我找到另一个社区榜，叫 Prolong，直接把他们的 system prompt 拿过来。我说：好，只要他们的 system prompt，别的不管，直接扔进

### [00:31:07–00:31:36]

**EN**  directly into prime agent. Uh and then I ran this and I was like oh my god the first run that I got it hit 99.9% and then I looked at the logs and I was cheating. Okay. So, I was like, "Okay, I got to do proper sandboxing here. Like, let's set this up properly." Uh, and then so I spent another day on this. And then, and then I went back and I was like, "Oh my god, I got 78% with GPT soul. Like, this is going to be a great result." Um, and then we're back. It's like, oh, let's compare a couple other ones. And

**中文**  Prime Agent。然后我一跑，第一轮就到了 99.9%。我去看日志，发现自己在作弊。好吧。我说：得做正确的沙箱，把实验搭干净。于是我又花了一天。然后再跑，我心想：天哪，GPT Soul 拿到了 78%，这会是个很好的结果。然后我们回来，说：再拿几个别的比一比。

### [00:31:34–00:32:02]

**EN**  so, it's again, we just took the prompt, uh, general prompt that basically says, uh, use a world model to solve ARC AGI 3. Uh, here are the actions that you can take. um you have uh and then the general system prompt for prime agent which is like you have a ripple you can call sub agents uh you can use the it programmatically um and and so we went through I went through the traces and it's basically doing a bunch of different um like calls of the coding in order to like check out these different scenarios and analyzing the images and

**中文**  所以还是一样，我们只用了那份 prompt，一份通用 prompt，大意是：用一个世界模型来解 ARC-AGI 3。这些是你可以采取的动作。再加上 Prime Agent 的通用 system prompt：你有一个 REPL，可以调 sub-agent，可以程序化地用它。我把轨迹过了一遍，它基本上就是在做各种编码调用，去查看不同场景、分析图像

### [00:32:00–00:32:27]

**EN**  doing like image processing and it's a lot of really cool um stuff that uh it seems like it was doing reasonable reasoning while leveraging the the ripple that we had um as like one of the main things that was able to enable build this. Uh so I went through and I ran a couple other ones. We did GPT tero 25.7% which is really cool. You can see that compared to like what were the um like the week before we did this uh open AAI was like the guys the harness matters a lot when you're doing

**中文**  以及做图像处理，很多很酷的操作。看起来它在做合理的推理，同时把我们的 REPL 当作能把这套东西做起来的主要能力之一。于是我又跑了几个。GPT-5 拿到了 25.7%，这很酷。你可以对比一下：我们做这事的前一周，OpenAI 还在说，做评测的时候 harness 非常重要

### [00:32:25–00:32:55]

**EN**  evaluations. We use the responses API. This is the result that we got. Um and we we ran Terra and and got almost like we we didn't run to completion this one but we got really good results in comparison. And then we go um that that we're already achieving higher than some of like the GBT soul extra high which was crazy. And then uh we went and we did Opus and hit 95.5%. We're like that's insane. We also compared to a lot

**中文**  我们用的是 responses API，这是我们拿到的结果。我们跑了 Terra，虽然这一次没有跑完，但对比下来结果已经很好。然后我们发现，我们已经高于某些 GPT Soul extra high 的成绩了，这很疯狂。接着我们跑了 Opus，到了 95.5%。我们当时就觉得这太离谱了。我们还跟很多

### [00:32:53–00:33:22]

**EN**  of the other harnesses. So some people ask me like did you run this with cloud code? Uh I did. Um unfortunately the results weren't very good. Um, and so rather than having bad results, I just deferred to the the original cloud code results and some other people have run it uh with similar configurations to prime agent and gotten much better results since then. Um, but what's interesting is that a lot of the really popular harnesses don't necessarily do well when prime agent does well. So like for air agent, um, we spent a lot of money very quickly and uh, we had to cut

**中文**  其他 harness 比过。有人问我：你用 Claude Code 跑过吗？跑过。可惜结果不太好。所以与其摆出差结果，我直接沿用了 Claude Code 原来的成绩。后来也有人用和 Prime Agent 类似的配置跑，结果好了很多。但有意思的是：很多很火的 harness，在 Prime Agent 表现好的地方，它们未必好。比如 Air Agent，我们很快花掉很多钱，不得不砍掉

### [00:33:21–00:33:50]

**EN**  it off because I spent like $5,000 without making much performance. Um, not saying this is the best they could do, but it cost a lot of money to do so. Uh so I think that the cost to performance uh ratio is very important and one of the things that does save money is being able to programmatically work with your context. Uh we ran a bunch of long horizon um evals as well like oolong and some coding uh emulator bench which is going to come out soon which is a program bench alternative and we found

**中文**  因为我花了大概五千美元，性能却没涨多少。不是说他们不可能做得更好，但这样做很贵。所以我觉得性价比非常重要。能真正省钱的一件事，就是能程序化地处理你的 context。我们还跑了一批长程评测，比如 Oolong，还有一些编码的 EmulatorBench，马上会发布，是 ProgramBench 的替代。我们发现

### [00:33:48–00:34:15]

**EN**  that it was mainly parody or slightly better than these other harnesses like you across different models versus doing like pimono cloud codecs with glm 5.2 to Opus 5 and 5.6 as our setting. Another one I thought was really cool is we have this like program bench alternative called emulator bench where we're trying to reproduce entire emulators of computer systems or in this case creating like a Game Boy Color and check that out. And we found that what's

**中文**  它大体持平，或略好于这些其他 harness。我们在不同模型上比过，比如 Prime、Claude Code、Codex，设定是 GLM 5.2 到 Opus 5 和 5.6。另一个我觉得很酷的，是我们这个叫 EmulatorBench 的 ProgramBench 替代：我们试图复现整套计算机系统的模拟器，这个例子里是做一台 Game Boy Color，再去验收。我们发现有意思的是

### [00:34:13–00:34:42]

**EN**  really interesting is because it has this um ripple access in the RLM, it's able to use these programs in order to kind of do these like out of experiment uh loop designs in order to um try things out in a lot more expressable and free way before submitting the final solution to the greater. Uh we also tried this with uh GPU kernels um and we got about par results uh across different um both soul and kimico. One

**中文**  因为它在 RLM 里有 REPL 访问权，它能用这些程序做实验循环之外的设计，用表达力更强、更自由的方式先试，再把最终方案提交给评测。我们也用 GPU kernel 试过，结果大致持平，Soul 和 Kimi 上都是。一个

### [00:34:40–00:35:08]

**EN**  is better, one is worse. about par um which so we we're not overfit to like any one particular um evaluation here. Um what's interesting for the long horizon stuff is we had some auto research uh experiments that we did with the nano GPT speedrun but we scaled it up. We said let's give it uh 8 by H200 for um a week and see what happens. And you might be like, okay, prime age is going to do so much better, right? Because it's able to do all this

**中文**  更好，一个更差，总体持平。所以我们并没有过拟合到某一个评测上。长程这边有意思的是，我们做了一些 auto research 实验，基于 nanoGPT speedrun，但规模放大了。我们说：给它 8 张 H200，跑一周，看看会怎样。你可能会想：Prime Agent 会好得多，对吧？因为它能做所有这些

### [00:35:06–00:35:35]

**EN**  programming. Uh, it's a little high variance. We can't attribute um any of the benefits to with the harness versus the model there because it's a very hard task. But what we can do is inspect a lot of the behavior that we've seen. And what's really interesting is that we're seeing models like deep 6v4, GLM 5.3, and Kim K3. Um, you can tell these were done a little more recently than our first results. Uh, and we took these and they were doing like what we call out of loop experiments. So we were trying to say how can I run experiments on like the CPU and like look at the

**中文**  编程。方差有点高。这个任务太难，我们没法把收益归因到 harness 还是模型。但我们可以去看大量观察到的行为。很有意思的是，我们看到 DeepSeek V4、GLM 5.3、Kimi K3 这类模型——可以看出这些比我们第一批结果更新一点——它们在做我们所谓的循环外实验。也就是在说：我怎么在 CPU 上跑实验，去看

### [00:35:33–00:36:02]

**EN**  parameterization and do hyperparameter search and analyze the data so that I don't have to spend like all my time running expensive H200 experiments uh because that takes the majority of the time. So it's it's running experiments that are not the main experiment in order to optimize them. I think that's really cool behavior that we're seeing uh as we we shape what would be what kind of things we need to for the expressability for prime agent. So you can use like really good auto research because you can imagine if it's good at

**中文**  参数化、做超参搜索、分析数据，这样就不必把所有时间花在昂贵的 H200 实验上，因为那才占了大部分时间。所以它在跑的不是主实验，而是为了优化主实验的那些实验。我觉得这是我们看到的很酷的行为，也在塑造 Prime Agent 需要什么样的表达力。所以你可以拿它做很好的 auto research。可以想象，如果它擅长

### [00:36:01–00:36:29]

**EN**  auto research, it'll be good with you. It be even better with a human in the loop to bootstrap your experiments. Uh and finally, we also streamed a 7-day factorial run which used a total of 633 agents um across uh 23 million tok output tokens in order to make like steady uh tech technological advancement across the tech tree to continue to progress over time. And here it uh one of the main benefits is I can use like

**中文**  auto research，它跟你配合也会好；有人在环里帮你 bootstrap 实验，会更好。最后，我们还直播了一场 7 天的 Factorio 运行，一共用了 633 个 agent、2300 万输出 token，为的是在科技树上稳步推进，持续往前走。这里一个主要好处是，我可以用

### [00:36:27–00:36:56]

**EN**  these sub aents that can divvy up into different tasks in the factory in order to research and build and gather resources and build the next items to design the factory. Um as well as it can use the refinement to leverage what happened in the past in order to help in the future um over these very long context so it doesn't get stuck. And one of the most interesting things here is that it does not get stuck and it continues to make technology progression even at the end of our uh stage. Um this

**中文**  这些 sub-agent，把工厂里的不同任务拆开：研究、建造、收集资源、做出下一件物品，去设计工厂。它也能用 refinement，把过去发生的事用到未来，撑过这些非常长的 context，这样它不会卡住。最有意思的一点是：它没有卡住，即使到了我们这个阶段的末尾，仍在推进科技。这更像

### [00:36:54–00:37:22]

**EN**  is more like a Gemini plays Pokemon kind of uh conclusion here. If there's one thing that uh I find interesting today u but like what takeaways you should actually add to your own harness. Um I think that you should think about agentic context management. Uh you should think about swarms and looking into further depth RLMs and trying to run standardized eval. All of the results that we can they showed today can be run with our uh verifiers uh package that we have at Prime Inslect. Um and shout out to my collaborators who

**中文**  Gemini 玩宝可梦那种结论。如果今天只带一样东西走，也就是你真正该加进自己 harness 的 takeaway：我认为你该想想 agent 式的 context 管理；该想想 swarm，更深入地看 RLM，并试着跑标准化评测。我们今天展示的所有结果，都可以用我们 Prime Intellect 的 verifiers 包来跑。也感谢我的合作者们

### [00:37:20–00:37:50]

**EN**  are fantastic and I love working with. Thanks. [applause] >> All right, next up we have John. >> Hey everybody. I'm super excited to talk about a project um that we've been working on at Stanford. Um, I've been working on this with Ivanka Orion, my my co-lead author, as well as our adviserss Hazeni and Christopher Ray. So, personal AI is everywhere, but it's mostly

**中文**  他们非常出色，我很喜欢和他们共事。谢谢。[掌声] >> 好，下一位是 John。 >> 大家好。我非常兴奋能讲我们在 Stanford 做的一个项目。我和共同第一作者 Ivanka Orion 一起做的，还有我们的导师 Hazy 这边，以及 Christopher Ré。所以，个人 AI 到处都是，但今天大多还是

### [00:37:48–00:38:18]

**EN**  cloudbound today. Uh, we see lots of different harnesses and projects focused on making daily writing, research, coding, and scheduling. But projects like OpenClaw and Hermes agent typically rely on cloud LMS um for most of the intelligence and for most of the most of the queries. Um, what does this mean? It means that it's pretty costly. You're getting thousands and thousands of dollars in API costs if you aggregate it over a year. Um, it's not private.

**中文**  绑在云上的。我们看到很多不同的 harness 和项目，在做日常写作、研究、编码和日程。但像 OpenClaw 和 Hermes agent 这类项目，智能的大部分、查询的大部分，通常都依赖云端 LLM。这意味着什么？意味着很贵。按一年加起来，API 费用能到成千上万刀。而且不私密。

### [00:38:15–00:38:43]

**EN**  You're often sending your most personal um, data to LMS up in the cloud and you don't necessarily know where all that data is going. Um, it also requires you to rent your intelligence as opposed to just simply owning it out of the box. And finally, it tends to consume orders of magnitude more energy than just running these LMS on your laptop. And so the local LMS are finally good enough to actually run a lot of these queries that people care about. And so we see that um

**中文**  你常常把最私人的数据送到云上的 LLM，而且不一定知道这些数据会去哪。这还意味着你得租用智能，而不是开箱即拥有。最后，它消耗的能量，往往比直接在笔记本上跑这些 LLM 高几个数量级。而本地 LLM 终于够好了，真能跑很多人在乎的那些查询。所以我们看到

### [00:38:41–00:39:09]

**EN**  the the current LMS of today are only 6 to 12 months uh behind whatever is the state-of-the-art frontier models um of before. So you see um LMS today such as Quen 3.8 27B um that achieve roughly the same performance as like Claude 4.6 Opus um back in the day. So that was kind of the state-of-the-art model back in August 2025. Um, and that gap seems to be closing uh more and more as the hardware accelerators that we have um

**中文**  今天的 LLM，只落后于此前最前沿模型大约 6 到 12 个月。比如今天的 Qwen 3.8 27B，大致能达到当年 Claude 4.6 Opus 的水平。那大概是 2025 年 8 月的前沿模型。而且这个差距似乎还在不断缩小，因为我们拥有的硬件加速器

### [00:39:07–00:39:37]

**EN**  for our laptops and for our workstations get better and better. Uh, just this week we saw a new release from Apple um with the new Mac Mini. And so we're seeing this renewed focus from Apple as well as Nvidia to build accelerators specifically for personal use cases. And so with this project, we wanted to explore the the main question of can we build the core of a personal AI stack, namely the model inference, the agent execution, the memory, the learning, basically the parts that are mostly

**中文**  笔记本和工作站上的，都越来越好。就在这周，Apple 发布了新的 Mac Mini。所以我们看到 Apple 和 NVIDIA 都重新把重点放在为个人场景做加速器。于是这个项目想探索的核心问题是：我们能不能把个人 AI 技术栈的核心，也就是模型推理、agent 执行、memory、学习，基本上就是今天大多还

### [00:39:34–00:40:03]

**EN**  reliant on the cloud today entirely on device while staying competitive with these cloudonly stacks. And so we decided to propose open Jarvis. Um name needs no needs no explanation. Um but we wanted to explore just how much of this we could run on device completely for free uh while preserving uh security, privacy and quality. And so to construct open Jarvis we wanted to create the simplest set of primitives for which you define any sort of harness or or personal AI stack. Um the first

**中文**  依赖云的那些部分，完全放到端上，同时仍能跟这些纯云技术栈竞争。于是我们提出了 open Jarvis。名字不用解释。我们想探索：这里面有多少可以完全免费地在端上跑，同时保住安全、隐私和质量。为了构建 open Jarvis，我们想给出最简单的一组原语，用来定义任何一种 harness 或个人 AI 技术栈。第一

### [00:40:02–00:40:31]

**EN**  one is whatever user interfaces you need to use. Um the second one is the actual agentic logic around composable reasoning and using different kinds of intelligence and tools. Uh for the intelligence, it's whatever LM you're using as your engine for keeping everything going. Um so this could be Quen, GBDO, OSS, Gemma 3N. Um and then whatever actual inference engine you need to run it. So this could be O Lama, um Llama CBP, VLM, SG Lang, um including whatever hardware you're running it on.

**中文**  个是你需要用的那些用户界面。第二个是真正的 agent 逻辑：可组合的推理，以及使用不同种类的智能和工具。智能这一层，就是你用来驱动一切的那台 LM 引擎。可以是 Qwen、GPT-OSS、Gemma 3N。然后是实际跑它所需的推理引擎，可以是 Ollama、llama.cpp、vLLM、SGLang，再加上你跑它所在的硬件。

### [00:40:29–00:40:58]

**EN**  So this could be Apple Silicon, Nvidia, whatever you need. um for actually making all of these agents and intelligence useful you need some set of tools in memory that can be run through a standard MCP protocol um and you need some sort of uh set of primitives for actually doing learning whether it's prompt based techniques like Japa or DSPI um whether it's weight based techniques like gpo and sftt and Laura um you need some way to actually get this agent to improve over time and actually be able to make it more

**中文**  可以是 Apple Silicon、NVIDIA，你需要什么就是什么。要让这些 agent 和智能真正有用，你需要一组能通过标准 MCP 协议跑的工具和 memory；还需要一组做学习的原语，不管是基于 prompt 的技术，比如 GEPA 或 DSPy，还是基于权重的技术，比如 GRPO、SFT 和 LoRA。你需要某种方式让这个 agent 随时间改进，真正能让它更

### [00:40:56–00:41:25]

**EN**  personal and more effective and so to kind of walk through like what opens looks like um we tried to go with all of the standard um interfaces that people are already accustomed to. Um so we wanted to give people the ability to interact with it through a desktop and actually just run it as they would normally expect, but then see all of the savings that they're getting in terms of dollars and energy. Um we also wanted to give people the ability um to run different kinds of continuous agents. So different kinds of agents that are

**中文**  个人化、更有效。顺着讲一下 open Jarvis 长什么样：我们尽量走人们已经习惯的那些标准界面。我们希望大家能通过桌面跟它交互，按他们平时预期的方式去跑，同时看到自己在钱和能耗上省了多少。我们也希望大家能跑不同类型的持续 agent。也就是不同类型的、

### [00:41:23–00:41:53]

**EN**  persistent in terms of cron jobs and being able to run standard protocols um day after day. Um, basically we just wanted to to make this like plug-and-play with all of the workflows that people are already accustomed to running. Um, and we wanted to make this something that can get people to have their first experience with LMS on device the same way people had their first experience with ChatGpt or Claude back in the day. Um, so yeah, and so yeah, to step through a little quicker, but yeah, here's like a nice way to like

**中文**  以 cron job 的方式持久存在、能日复一日跑标准协议的 agent。基本上我们只是想让它跟人们已经习惯的那些工作流即插即用。我们还希望它能成为人们第一次在端上使用 LLM 的入口，就像当年人们第一次用 ChatGPT 或 Claude 那样。好，我们稍微讲快一点。这里有一个不错的方式，可以

### [00:41:50–00:42:19]

**EN**  set up new persistent jobs. Um, yeah, we have all of these different components. We need some way to actually optimize it. And so we wanted to get out of the way of the LM as much as possible by just creating a simple spec of these five primitives by which they could go through the optimization. And what we found is that by going through this whole optimization loop, we were not only able to get significant dollar costs uh dollar cost reductions, but also significant latency reduction and significant um improvements to overall

**中文**  设置新的持久任务。对，我们有这么多不同组件。我们需要某种方式真正去优化它。所以我们想尽量别挡 LM 的路，只给出这五个原语的一份简单 spec，让它们按这个去做优化。结果是：走完这整套优化循环，我们不但显著降低了美元成本，还显著降低了延迟，整体

### [00:42:17–00:42:46]

**EN**  quality on these tests. And so this this configuration is meant to simplify down to just the five main things that people care about when they're building these LM uh harnesses. So the intelligence, the engine, the actual agentic logic around it, the tools or learning systems required for running it um and the whole optimization um for the whole spec as a whole. And so something that we thought could be interesting to help bridge this gap between local and cloud LMS is to actually have the cloud LM go through and manual and uh and

**中文**  质量在这些测试上也有显著提升。所以这套配置是想简化成人们在建这些 LM harness 时真正关心的五件事：智能、引擎、外围的 agent 逻辑、跑它所需要的工具或学习系统，以及对整份 spec 的整体优化。我们觉得有一件事可能有助于弥合本地 LLM 和云端 LLM 的差距：让云端 LM 去走一遍，手动地，以及

### [00:42:44–00:43:12]

**EN**  automatically optimize the whole LM the whole local stack. And so this is a nice way of taking advantages of the capabilities of cloud LM to diagnose proposed changes and gate um to create improved solutions for these local LMS while not incurring the cost of those cloud LMS when you actually deploy these um local stacks at inference. And so what we found is that these uh these open Jarvis jobs that were these open Jarvis um configurations that were

**中文**  自动地优化整个本地技术栈。这是一种很好的办法：借云端 LM 的能力来诊断、提出改动、做门控，给这些本地 LLM 做出更好的方案，而在真正部署这些本地技术栈做推理时，又不必承担那些云端 LLM 的成本。我们发现，这些由云端 LLM 优化过的 open Jarvis 任务、这些 open Jarvis 配置

### [00:43:10–00:43:38]

**EN**  actually optimized by cloud LMS like cloud or chat GPT um were much more effective than uh local stacks that were just deployed out of the box because you could actually cater to the specific LMS the specific harness uh that was needed uh for for different kinds of workloads. And what we found is that even with the ondevice LMS of today, we can rival cloud LMS on different workflows around personal AI um personal use cases,

**中文**  比如由 Claude 或 ChatGPT 优化的，比开箱即用部署的本地技术栈有效得多，因为你能针对不同工作负载，去适配具体的 LLM、具体的 harness。我们还发现，即便用今天的端侧 LLM，在个人 AI、个人场景的不同工作流上，我们已经能和云端 LLM 抗衡

### [00:43:35–00:44:04]

**EN**  coding, agentic tasks. While there remains like many tasks for which um like local local size LMS are not enough um the gap is surprisingly closing um month after month um as these LMS become better distilled, more effective and also we get better accelerators um for running them. And so even with LM of today, we can get 800x lower lower costum in terms of actually running them as well as a significant reduction in

**中文**  编码、agent 任务上也能比。虽然仍有很多任务是本地体量的 LLM 不够用的，但这个差距在出人意料地逐月缩小，因为这些 LLM 蒸馏得更好、更有效，我们也有了更好的加速器去跑它们。所以即便用今天的 LM，实际跑起来的成本也能低 800 倍，延迟也有显著

### [00:44:02–00:44:30]

**EN**  latency. Um what we also found is that no matter which cloud LM that we cloud LM we picked um it was useful um in terms of optimizing the whole aentic loop for these uh local local open Jarvis configurations. Um we found that um the Opus series, Opus 5 as well as GBD 5.6 Soul were were naturally the best. Um but was interesting to see is that you could pick Gemini, you could pick um other other uh larger um LM

**中文**  下降。我们还发现，不管选哪家云端 LM，对优化这些本地 open Jarvis 配置的整条 agent 循环都有用。我们发现 Opus 系列，Opus 5 以及 GPT 5.6 Soul，自然是最好的。但有意思的是：你可以选 Gemini，也可以选其他更大的 LM

### [00:44:28–00:44:57]

**EN**  families like Kimmy and GLM and use them to optimize these local configurations so that you could capture those efficiency gains, capture those performance gains um for local inference later. What we also found is that the whole open Jarvis harness was cheaper to optimize than alternatives which might require more data or more LM calls. Um we found that like this set of specs um and this set of primitives um was most effective for local LM settings because it got um the whole optimization

**中文**  家族，比如 Kimi 和 GLM，用它们来优化这些本地配置，这样之后做本地推理时就能吃到那些效率收益、性能收益。我们还发现，整套 open Jarvis harness 的优化成本比那些可能需要更多数据或更多 LM 调用的替代方案更低。我们发现这组 spec、这组原语，对本地 LM 设定最有效，因为它把整套优化

### [00:44:55–00:45:24]

**EN**  loop and the whole um set of LM abstractions out of the way of the cloud LM to just optimize the whole system and just make it make it really fast and really effective. Uh looking forward, we're excited to keep building out this project. Uh we think in the very near future you're going to see um a huge maj a huge proportion maybe even a majority of uh people's daily inference calls going to local devices and on-prem uh laptops or on-prem workstations as opposed to the kind of standard of today where everything's being pushed out to

**中文**  循环和整套 LM 抽象都从云端 LM 面前挪开，让它直接优化整个系统，把它做得非常快、非常有效。往后看，我们很兴奋能继续把这个项目做大。我们认为在很近的将来，你会看到人们日常推理调用中很大一部分，甚至可能是大多数，会走到本地设备、本地笔记本或本地工作站上，而不再是今天这种标准：什么都往

### [00:45:22–00:45:48]

**EN**  the cloud. Um we think these trends are only going to continue because the accelerators keep getting better and the LMS keep getting better. Um and so if you're excited about um anything in the stack, whether it's uh better local LMS, better accelerators, better inference engines uh for deploying uh beyond data centers, um please reach out. Uh we'd be excited to chat. Include a QR code of the project. Um if if folks are around here afterwards, we'd love to chat. Thanks. [applause]

**中文**  云上推。我们认为这些趋势只会继续，因为加速器越来越好，LLM 也越来越好。所以如果你对栈里的任何一层感兴趣，无论是更好的本地 LLM、更好的加速器，还是能部署到数据中心之外的更好的推理引擎，请联系我们。我们很乐意聊。这里有项目的二维码。如果之后还有人在现场，我们很愿意交流。谢谢。[掌声]

### [00:45:49–00:46:17]

**EN**  All right. Now, we have our own YC. We have Josh and Rean. Hi, I'm Josh and uh this is Rean and we're working on QM, which is YC's open-source uh agent harness for work. QM is one system that uh gives every employee at YC an open claw-like assistant uh that's

**中文**  好，现在是我们 YC 自己的。有 Josh 和 Rean。大家好，我是 Josh，这位是 Rean。我们在做 QM，这是 YC 开源的、面向工作的 agent harness。QM 是一套系统，给 YC 的每一位员工一个类似 OpenClaw 的助手，它

### [00:46:14–00:46:44]

**EN**  like fully customizable and available in Slack or via web UI, which is uh what we're looking at here. Um each person works within QM in their own personal context that has its own sandbox files uh and crons and it can they can also work with QM in a multiplayer setting like a slack channel. People use QM for

**中文**  可以完全定制，在 Slack 里或通过 Web UI 使用，也就是我们现在看到的这个。每个人在 QM 里都有自己的个人 context，有自己的沙箱文件和 cron。他们也可以在多人场景里用 QM，比如一个 Slack 频道。人们用 QM 做的事情范围很广

### [00:46:40–00:47:10]

**EN**  a pretty broad range of things uh like a lot of automations like email triage like legal and finance workflows. It's really good at editing documents and pulling data out of our internal database. Uh it can also spin up live internal web apps and help with stuff like planning events. Uh but it's designed to be broadly helpful for the range of tasks that someone might encounter at YC uh on a day-to-day

**中文**  比如大量自动化，邮件分拣，法务和财务工作流。它很擅长改文档，以及从我们的内部数据库里抽数据。它还能拉起活的内部 Web 应用，帮忙策划活动之类。但它的设计目标，是对一个人在 YC 日常可能碰到的那一类任务，都能广泛地有帮助

### [00:47:08–00:47:37]

**EN**  basis. So you might be wondering uh why we built this and it's really the result of a string of internal agent projects that have kind of unwound throughout the years and all of which were really riding this tailwind of increasingly capable models. Um the first one we built was in like January of 2025. We internally refer to

**中文**  你可能想问我们为什么要做这个。它其实是这些年一串内部 agent 项目展开下来的结果，而所有这些都搭上了模型能力越来越强的这股东风。我们做的第一个大概是 2025 年 1 月。我们内部叫它

### [00:47:35–00:48:02]

**EN**  it as the quote unquote like general agent. Uh, but it was pretty straightforward. Just kind of a system prompt with tools uh in a loop. It was one sizefits-all. Um, sort of like everyone was talking to the same thing. Uh, and it was pretty straightforward architecturally, but like still uh very or surprisingly good at answering data questions. Uh,

**中文**  所谓的 general agent。但它相当直接：就是一份带工具的 system prompt，放在一个循环里。它是一刀切的，差不多所有人都在跟同一个东西说话。架构上很简单，但在回答数据问题上，仍然出奇地好。

### [00:48:01–00:48:30]

**EN**  interestingly like the scope of what the general agent was good at just increased I guess unsurprisingly as the under underlying models got better. Uh, and we eventually hooked it up to Slack. We added crons uh and gave it a a few more tools so that it could be more be capable across more domains. Um, in June of 2025, uh, by then like a lot of our engineers started using cloud code and codecs. Uh,

**中文**  有意思的是，general agent 擅长的范围，随着底层模型变强，也就不足为奇地变大了。我们后来把它接到了 Slack，加了 cron，又多给了几个工具，让它能覆盖更多领域。到 2025 年 6 月，那时我们很多工程师已经开始用 Claude Code 和 Codex。

### [00:48:28–00:48:55]

**EN**  and we realized that you could pretty easily run these in a VM. And then, uh, we hooked that up to a Slack tag, which was a pretty powerful medium for people who just wanted to like run a one-off code change. Uh, we also configure configured it to run our CI pipelines and then spin up dev environments for testing. Uh, and so people could come in like describe a bug or something they wanted to see happen and the bot would

**中文**  我们意识到这些可以很轻松地在虚拟机里跑。然后我们把它接到一个 Slack 标签上，这对只想跑一次临时代码改动的人来说，是个很强的入口。我们还把它配成能跑我们的 CI 流水线，再拉起测试用的开发环境。所以人们可以进来说一个 bug，或他们想看到的某种效果，这个 bot 就会

### [00:48:54–00:49:23]

**EN**  go off and actually solve it, which is like a pretty powerful um thing for someone who like maybe hadn't uh made a code change before in their life even. Uh but we also on top of that had a small loop going where we would observe sort of how the bot failed uh where it went wrong and then update the agents.mmd which was uh present in the codebase at the time um to make sure

**中文**  真的去把它解决掉。对一个也许这辈子还没改过代码的人来说，这相当有力量。除此之外，我们还有一个小循环：观察 bot 是怎么失败的、错在哪，然后更新当时代码库里的 AGENTS.md，确保

### [00:49:20–00:49:49]

**EN**  that the thing got better as as uh we like observe the usage and so in January of this year a lot of the partners started using openclaw and one thing to know about YC partners is that they're incredibly busy uh between like office hours uh they get tons of inbound email they're always reading applications so like any tools that can give them additional leverage are incredibly valuable to YC so open cloud particular was useful because it

**中文**  这东西会随着我们观察使用情况而变好。于是今年 1 月，很多 partner 开始用 OpenClaw。关于 YC 的 partner，有一点要知道：他们极忙。办公室时间之间，大量涌入的邮件，还一直在看申请。所以任何能给他们额外杠杆的工具，对 YC 都极其值钱。OpenClaw 尤其有用，因为它

### [00:49:48–00:50:18]

**EN**  was the first agent that a lot of them had used uh that had their own computer that had its own computer and so this made it like very customizable in a way that the previous paradigm of agents was not uh and it functioned almost like a personal assistant And so in April uh the question became like could we provide this to every employee at YC uh without uh like buying everyone a Mac Mini effectively. And so we ended up provisioning a fleet of like

**中文**  是他们很多人用过的第一个拥有自己电脑的 agent。这让它非常可定制，这是此前那套 agent 范式做不到的，它几乎就像私人助理。于是到了 4 月，问题变成：我们能不能在实质上不必给每个人买一台 Mac Mini 的前提下，把这东西提供给 YC 的每一位员工。于是我们最终配了一整队

### [00:50:15–00:50:44]

**EN**  50 plus uh Hermes agents that were running in VMs. And these were definitely pretty helpful but they required a lot of configuring for people to get value out of them. And it was just like inherently kind of difficult to manage this fleet. Um, it was sort of like a whack-a-ole situation where I would have to sort of like SSH into these individual instances and fix them. And so the follow-up question became

**中文**  五十多个跑在虚拟机里的 Hermes agent。它们确实挺有帮助，但人们要做很多配置才能用出价值。管理这支舰队本身就很难。有点像打地鼠：我得 SSH 进一台台实例去修。于是接下来的问题变成

### [00:50:41–00:51:10]

**EN**  like we've got we've gotten a lot of value out of these agentic systems. Uh, like let's build something that tries to address some of the downsides of running this big fleet of Hermes agents. Uh, while still maintaining the personalizability and some of the like the stuff that people were really getting value out of. And so, >> um, yeah. Okay. So, basically, uh, there's a pretty clear trend from

**中文**  我们已经从这些 agent 系统里拿到了很多价值。那我们就做点东西，试着解决跑这么大一队 Hermes agent 的一些缺点，同时保住可个性化，以及人们真正用出价值的那些部分。所以，>> 嗯，好。所以基本上，从 Josh 刚才给你们看的内容里，有一条相当清楚的趋势

### [00:51:08–00:51:37]

**EN**  whether you could see there from what Josh was showing you. Basically, um, the models are getting better exponentially. Um, and we were starting to see just increasingly impressive returns from giving them more and more capabilities. So, OpenCloud gives the agent its own computer and we start to see really impressive returns from that. So around May this year, we started thinking just like how far can we push this if we just keep pulling on this threat. Um we really like the the lens of sort of unhobling. I don't know if you guys uh

**中文**  模型在指数级变好。我们开始看到：给它们越来越多的能力，回报越来越惊人。OpenClaw 给了 agent 自己的电脑，我们从中看到了非常惊人的回报。所以大约今年 5 月，我们开始想：如果就沿着这根线一直拉，能推到多远。我们很喜欢所谓 unhobbling 这个视角。不知道你们有没有

### [00:51:35–00:52:03]

**EN**  read situational awareness when it came out in like 2024. Um but that was kind of an era when like test time compute was just starting to become a thing and like you know we're starting to give agents tools for the first time and there's sort of this intuition that what agents can do like there's a little bit there's kind of more intelligence in the models than we're than we're using in a lot of cases. Um, and it's like if we really push uh the frontier in terms of just like what we're offering up the agents as as capabilities that they can

**中文**  读过 2024 年那篇 Situational Awareness。那大概是 test-time compute 刚开始成为一件事的年代，我们也第一次开始给 agent 工具。当时有一种直觉：agent 能做的事里，模型里的智能在很多情况下比我们用上的要多。就好像如果我们真的把边界往前推，把我们提供给 agent、让它们能

### [00:52:00–00:52:30]

**EN**  make use of uh like magic can start happening. Um so the first way that we do that with QM um is by essentially pulling the brain of the system up out of the sandbox. So with uh you know with Hermes and with OpenClaw you effectively have um the agent has its own computer which is super powerful but it's also trapped inside that computer. So that causes a few issues just from you know uh us trying to administer that system

**中文**  用上的能力做足，奇迹就会开始发生。所以我们在 QM 里做的第一件事，本质上就是把系统的大脑从沙箱里抽出来。在 Hermes 和 OpenClaw 里，agent 实际上有自己的电脑，这非常强，但它也被困在那台电脑里。这会带来一些问题，单从我们去管理那套系统来说

### [00:52:28–00:52:58]

**EN**  when there were you know uh even a few dozen of these things it starts to become unwieldy almost immediately. Um but the other issue with that is that you um all all of the like all the sessions that you would have uh they're trapped inside that computer. And so what we did instead is we uh we just offload everything into Postgress. So everything is centralized um from all the agent conversations that people are having. Um and then we expose those to the agent itself. So it can look at all the context that's sort of aggregating

**中文**  哪怕只有几十个，也几乎立刻就会变得难搞。另一个问题是：你所有的 session 都被困在那台电脑里。所以我们改成把所有东西卸到 Postgres。一切都集中起来，包括人们和 agent 的全部对话。然后我们再把这些暴露给 agent 自己，让它能看到系统里正在汇总的全部 context

### [00:52:56–00:53:25]

**EN**  from across the system. And then the other thing we do is we start thinking about sandboxes less as this home where the uh where the agent lives and where it's kind of stuck in a lot of ways. And sandboxes become more of this thing uh more of a resource that the agent can dip into and use as needed. Um but it's um it's a lot less limiting. Uh the other thing that this starts to open up um is this idea of you're accumulating this large eval set of all the traces that you have from the conversations that people are having with the agent.

**中文**  然后我们做的另一件事，是不再把沙箱想成 agent 住的家，那个它在很多方面被卡住的地方。沙箱更变成一种资源，agent 需要时可以伸进去用。这样就限制少得多。这件事还打开了另一个想法：你在积累一个很大的评测集，里面是人们和 agent 对话留下的全部轨迹。

### [00:53:24–00:53:52]

**EN**  And in principle, you can think about going and and hill climbing on that um and sort of having this automated improvement loop. Um we've had sort of mixed results with that. I would say I think typically if you're just dispatching this like torrent of agents that are supposed to um fix all of the bugs that they're encountering when you have the LLM as a judge, you start to get this kind of uh like main character syndrome where the agents are making fixes that are, you know, only seeing their their um their piece of the

**中文**  原则上，你可以想着在这上面做 hill climbing，弄一套自动改进循环。我们在这上面的结果比较参差。我会说，通常如果你只是派出一大群 agent，让它们去修自己碰到的所有 bug，再用 LLM 当裁判，你就会开始看到一种主角综合征：这些 agent 做的修复，只看见了自己那一块

### [00:53:51–00:54:20]

**EN**  elephant effectively. like they're really um they can be sort of yeah not not seeing the whole um whole system and so having the human in the loop there has continued to be really important although we're really looking forward to this uh working uh all the way around. Um so the other major thing we do that's pretty obvious is just wire the agents to all of the resources across the company that we can. Uh we already happened to have a CLI at YC that worked really well um that wired a

**中文**  大象。它们往往看不见整个系统。所以人还在环里，这一点一直很重要，尽管我们非常期待这件事能完全转起来。我们做的另一件比较明显的大事，就是把 agent 接到公司里我们能接到的所有资源上。YC 碰巧已经有一个很好用的 CLI，把很多系统接在了

### [00:54:17–00:54:46]

**EN**  lot of systems together. Um uh but anything that wasn't in there uh we basically allow um adding just arbitrary API keys um that sort of thing. And then we also want to ensure that we have par with uh just an employee working on their laptop. So you know device code OOTH um we go ahead and ingest that into a keychain and then refresh it for you. So it's it's uh ideally supposed to imitate the experience of a person on their

**中文**  一起。不在里面的，我们基本上允许加入任意 API key 之类的东西。我们还想确保它能跟一个员工在自己笔记本上干活打平。所以像 device code OAuth，我们会收进 keychain，再帮你刷新。理想情况下，它应该模仿一个人在自己电脑上的体验。

### [00:54:43–00:55:13]

**EN**  computer. We mostly keep this to to be read only uh in the database but we do allow for rights um via uh human reviewed bulk upserts. So the way that works is the agent will put forward a plan to um edit the database um that a person can kind of give a once over and ensure it's not doing anything crazy before the write actually happens. Um one thing we've observed with this is that we've started just kind of rubber stamping these. It's a little bit like um I think if you guys use cloud code in

**中文**  数据库这边我们大多保持只读，但我们允许通过人工审过的批量 upsert 来写。做法是：agent 会提出一份改数据库的计划，人可以扫一眼，确认它没在干离谱的事，写操作才会真正发生。我们观察到的一件事是：我们已经开始对这些盖橡皮图章了。有点像你们早期用 Claude Code 的时候

### [00:55:12–00:55:41]

**EN**  the early days like you might have been reviewing the tool uses very closely and eventually um you sort of build up more trust in the agent. So this is something that we're uh looking at very closely over the over the next few months. Yeah. So sort of like I was saying um the sandboxes in this system we like to think of as a resource for the agent. So and and less where the agent actually lives. Um so in QM the agent can um basically uh by default it's going to be using a particular sandbox. So that's

**中文**  可能会把每一次工具调用看得很紧，后来你会对这个 agent 建立起更多信任。这是我们接下来几个月会盯得很紧的一件事。对，就像我刚才说的，这套系统里的沙箱，我们更愿意把它想成 agent 的一种资源，而不是 agent 真正住的地方。所以在 QM 里，agent 默认会用某一个沙箱，也就是

### [00:55:40–00:56:10]

**EN**  going to be one that's been allocated to the user that it's talking to. But in general uh there are there are environments that the agent can kind of converge on and uh and can collaborate with. Um, and then the other uh key thing is that um if the agent is working on like a um like a heavier dev workload, it can go and reach for a machine that has more resources. Um if it's working on something that's simpler, it'll just go for a sandbox that's that's less powerful. Um and so pushing that decision into the agent

**中文**  分配给正在跟它说话的那个用户的那一个。但一般来说，还有一些环境是 agent 可以汇聚上去、可以协作的。另一件关键的事是：如果 agent 在做更重的开发负载，它可以去够一台资源更多的机器；如果在做更简单的事，它就会选一个不那么强的沙箱。所以把这个决定推给 agent 自己

### [00:56:08–00:56:36]

**EN**  itself rather than the harness has been a really um really powerful thing. Similarly, um allowing the agent to uh tap into its own runtime. So basically, if it can choose the provider um that it's working with, you start to get out of situations like um I'm sure you guys have run into this with Fable. If you try to do AI research, if you try to do cyber security, anything, you'll get a bunch of refusals. Um so what we can do when the agent can control it on runtime, it can just pop out into

**中文**  而不是交给 harness，这件事非常强。同样，允许 agent 触达自己的运行时。基本上，如果它能选择自己在用的 provider，你就能躲开这类情况——我相信你们用 Fable 时都碰到过：你想做 AI 研究，想做网络安全，什么都会被一堆拒绝。所以当 agent 能控制自己的运行时，它就可以跳到

### [00:56:34–00:57:03]

**EN**  another model uh when it needs to avoid a situation like that. Um similarly, uh like in the earlier situation, um it's often useful to pop between different sandbox providers. Um and so that's something we can do quite easily. And as a general rule, what we've tried to do is keep the harness extremely thin. Uh we think of the kind of core of the system as being these three tools um where you have execution in a remote sandbox, reading and writing from object

**中文**  另一个模型，来避开那种情况。类似地，像前面那种情况，在不同沙箱 provider 之间跳转也常常有用。这件事我们做起来相当容易。作为一条总原则，我们一直想把 harness 保持得极薄。我们把系统的核心想成这三个工具：在远程沙箱里执行，对对象存储做读写

### [00:57:00–00:57:29]

**EN**  storage and then uh publishing internal apps, a pretty simple sort of getbacked uh system. And then we have other tools for interacting with memory and crons and that sort of thing. But we really think of these as sort of temporary um papering over rough edges in the system. And really the core of it is is these three up here. Um it's we we really try to keep it as small as we can. Yeah, we're we're sort of trying to be this like uh AGI anticipating harness.

**中文**  以及发布内部应用，一套相当简单的、由 git 支撑的系统。我们还有其他工具来交互 memory、cron 之类。但我们其实把这些当成暂时糊住系统毛边的东西。真正的核心就是上面这三个。我们真的尽量把它保持得小。对，我们有点想做一种在预期 AGI 的 harness。

### [00:57:27–00:57:56]

**EN**  Although uh since we aren't there yet, um there are a few things that we've run into. Um, one of these is that the agents have been we've tried to put them in this really capable environment where they have all these uh tools available to them, but um they often give up way too early. Uh so one thing we've experimented with especially over the past month or so has been setting uh we call it like a grind tool or um basically we set budgets on goals. So

**中文**  不过既然我们还没到那一步，就碰到了一些问题。其中之一是：我们试着把 agent 放进一个能力很强、工具很多的环境里，但它们常常过早放弃。所以我们最近一个月左右在试的一件事是设一种我们叫 grind 的工具，基本上就是给目标设预算。所以

### [00:57:54–00:58:23]

**EN**  the agent is not allowed to give up on its task before a certain amount of like walk clock time. just like a couple hours uh or a certain amount of token spend. And so um what that can accomplish is just like uh really a lot better, you know, uh research outputs, better reports, that sort of thing. Um and it's been really fun actually to see uh like OpenAI and anthropic um you know, crack some open problems in math with like a very similar technique. Uh but it also works for just you know,

**中文**  agent 在一定的墙上时钟时间之前不许放弃任务，大概就是几个小时，或者一定的 token 花费。这样能带来的，就是好得多的研究产出、更好的报告之类。其实挺有意思的是，OpenAI 和 Anthropic 用非常相似的手法，攻克了一些数学上的开放问题。但这套对普通的

### [00:58:20–00:58:49]

**EN**  normal office work stuff too. Um the other thing that we've seen a lot of is so this harness is supposed to work uh it works in multiplayer. It works in Slack. Um, but because of the the artifacts of its of its training, effectively uh what we see is that uh the agent gets can get very confused about the situation that it's in even if we specify this pretty clearly in the system prompt. So having like local affordances for this has been something that's been uh that's been really important.

**中文**  办公室工作也同样有效。我们还经常看到另一件事：这个 harness 本应能在多人场景工作，它确实在 Slack 里工作。但因为它训练留下的痕迹，我们看到的是：即便我们在 system prompt 里写得很清楚，agent 仍可能对它所处的情境非常困惑。所以为这件事提供本地的可操作接口，就变得非常重要。

### [00:58:51–00:59:20]

**EN**  So yeah, um the other thing that has been a problem is that agents really don't uh understand social contexts. Um to make this a little more concrete, like if I tell Regan a piece of information, uh he intuitively sort of knows uh or at least has like a good mental framework of where it is okay to share that information. Uh but it takes some actual work to recreate this with an agent. uh like privileged information

**中文**  对，另一个问题是：agent 真的不懂社交情境。说具体一点，如果我告诉 Rean 一条信息，他凭直觉就大概知道，或者至少有一套不错的心智框架，知道这段信息适合分享到哪里。但要用 agent 复现这一点，得真下功夫。像特权信息

### [00:59:18–00:59:48]

**EN**  can very easily just leak into these contexts where it should not be and so uh the information that you can put in the brain is effectively like bounded by how good your permission system is uh and so YC luckily has an existing software system with like fine grained permissioning uh that has been built over the years but a lot of people just don't have that uh and so it takes uh work to allow for knowledge sharing in

**中文**  很容易就会漏进不该出现的 context。所以你能放进大脑的信息，实际上受限于你的权限系统有多好。YC 幸运的是，已经有一套做了多年的、带细粒度权限的软件系统，但很多人并没有这个。所以要在细腻的方式下允许知识共享，得花功夫

### [00:59:45–01:00:10]

**EN**  an in nuanced Okay, so thanks everybody for listening. Um, you can try out QM. It's open source. Uh, and coding agents are pretty good at standing it up. If you run into any problems, feel free to put up an issue and we'll look at it. Uh, we're also hiring. So if any of this resonated with you or would be exciting, then feel free to send us an email. [applause]

**中文**  。好，感谢大家来听。你们可以试试 QM，它是开源的。编码 agent 把它架起来已经相当拿手。如果碰到问题，欢迎提 issue，我们会看。我们也在招人。如果这些内容让你有共鸣，或让你觉得兴奋，欢迎给我们发邮件。[掌声]
