# Russ Salakhutdinov - Kimi K3 CEO’s PhD Advisor Predicts the Future of AI Agents · 中英对照逐字稿

- 原节目：Basil Chatha
- 英文原始来源：https://www.youtube.com/watch?v=upzrwzPpTf8
- 中文译制版入口：https://www.xiaoyuzhoufm.com/episode/6a6eb8a6ab3a91c24a0e5d40
- 时长：01:44:14
- 方法与限制：英文来自已验证的原始节目 transcript/caption；中文由 Codex 逐段翻译，未做逐字人工校对，公开引用前请回到英文原文与音频复核。

## 中英对照逐字稿

### [00:00:00–00:00:30]

**EN**  McKenzie, why does Mckenzie exist? Boston Consulting. These are big organizations. All of them should be replaced by people say AGI is two years. And people who actually work in these lab and frontier labs, I know a lot of them they basically saying just like we're not going to have [music] AGI in 2 years. A lot of people were basically saying, my god, OpenAI has figured out a new architecture. They figured out like AGI that's absolutely not true. The architectures are the same. The CEO of this company, Zillian Yan, he's my former student at CU. He is one of the smartest people I know. China has strong

**中文**  McKinsey 为什么会存在？Boston Consulting 也是。这些都是大机构，而那些说 AGI 两年就会到来的人认为，它们全都应该被取代。但真正任职于这些实验室和前沿实验室的人——我认识很多——基本都会说：我们不可能在两年内实现 AGI。很多人当时都在说：天啊，OpenAI 找到了一种新架构，他们已经攻克了 AGI。事实绝非如此，大家使用的架构都一样。这家公司的 CEO 杨植麟（Zhilin Yang）是我以前在 CU 的学生，也是我认识的最聪明的人之一。中国拥有强大的

### [00:00:28–00:00:57]

**EN**  talent. They now have the resources. So I think they'll be able to compete with the US. So if I am a CEO of a company and I have a better view of the future than my competitor, I'm going to win. >> So you think that all the LLMs are basically commodities? >> I think they're going to be commodities. There's only one example which is Roomba. There's no other example right now where a robot is useful. In 2005, neural networks were a joke. Third choice behind support vector machines. At the time, Russ Celakudinoff was working in finance, but by a chance encounter, he bumps into Jeff Hinton on the street in Toronto. Hinton drags him into his office and he walks out doing a PhD in deep learning before deep

**中文**  人才，现在也有了资源，所以我认为中国有能力与美国竞争。假如我是一家公司的 CEO，而且比竞争对手更清楚未来会怎样，我就会赢。主持人：所以你认为所有 LLM 最终基本都会商品化？嘉宾：我认为它们会成为商品。目前真正有用的机器人只有 Roomba 这一个例子，没有第二个。2005 年，神经网络还是个笑话，只是排在支持向量机之后的第三选择。当时 Russ Salakhutdinov 在金融业工作，却偶然在多伦多街头碰到了 Geoffrey Hinton。Hinton 把他拉进办公室；等他出来时，他已经决定去攻读“深度

### [00:00:56–00:01:25]

**EN**  learning was even a term. And if you're not familiar with Jeff Hinton, he's widely considered the godfather of AI. So he was in the middle of it all. 20 years later, Russ has done basically everything. He sold his startup to Apple. He worked on their secret self-driving car project, Titan. He became a professor at Carnegie Melon. And he just left Meta Super Intelligence Lab to start soothe labs, which just raised $50 million to build AI that forecasts the future. This episode, Russ tells me that the frontier labs are all building commodities. How the people actually building AGI don't think that it's 2 years away. Explains why Curser and half the startups you know are actually running on Chinese open source models. How Yang Zillian, the CEO of Kimmy, was one of his smartest PhD students. Why the only useful robot in

**中文**  学习”的博士，而当时甚至还没有 deep learning 这个说法。如果你不熟悉 Geoffrey Hinton，他被广泛视为 AI 教父，所以 Russ 从一开始就在这场变革的中心。20 年后，Russ 几乎什么都做过：把创业公司卖给 Apple，参与其秘密自动驾驶汽车项目 Titan，成为 Carnegie Mellon 教授，之后离开 Meta Superintelligence Labs，创办了刚融资 5,000 万美元、致力于构建预测未来之 AI 的 soothe labs（英文自动字幕拼写）。本期节目里，Russ 会谈到为什么前沿实验室构建的东西最终都会商品化，为什么真正从事 AGI 的人并不认为它会在两年内到来，为什么 Cursor 和你熟悉的一半创业公司其实在运行中国开源模型，Kimi 的 CEO 杨植麟为什么是他最聪明的博士生之一，以及为什么唯一有用的机器人

### [00:01:24–00:01:51]

**EN**  the world is still the Roomba. and why AI should replace Mackenzie Bane and BCG. Let's get into it. So you've been at the center of modern AI since you know basically the beginning. So I want to walk through that journey a little bit uh before we get into some of the stuff that you're working on now. So first of all like why did you decide to do a PhD in machine learning in the mid200s >> like where did you see things headed? >> You know it's a very good question. So my story is kind of a little bit unusual because I was very interested in AI. So

**中文**  仍然是 Roomba，以及为什么 AI 应当取代 McKinsey、Bain 和 BCG。我们开始吧。你从现代 AI 基本诞生之初就处在它的中心，所以在谈你现在做的事情之前，我想先梳理一下你的经历。首先，你为什么会在 2000 年代中期决定攻读 machine learning 博士？你当时觉得这个领域会走向哪里？嘉宾：这是个很好的问题。我的经历有些不寻常，因为我一直对 AI 很感兴趣。所以

### [00:01:50–00:02:19]

**EN**  I started my master's degree at the University of Toronto in early 2000s. Um I did my masters and then I kind of left and and went into banking. So I was in financial sector for a year and it's almost like by luck I started doing my PhD. Um at a time I was working with Jeff Hinton. He was you know working with me when I was doing my masters and I didn't know whether I wanted to do PhD or not. But then at one point I was walking down the street and I bumped

**中文**  我在 2000 年代初进入 University of Toronto 读硕士。硕士毕业后，我离开学校进了银行业，在金融行业待了一年。后来开始读博士几乎是机缘巧合。我读硕士时曾和 Geoffrey Hinton 合作，由他指导，但我并不知道自己是否想读博。后来有一天，我走在街上，碰巧遇见了 Geoffrey Hinton。他基本上对我说：“我有一个特别了不起的想法，你知道，这些模型、这些 deep learning 模型可以学习多层表征，等等。”于是他几乎是把我直接拉进了办公室，开始给我看他一直在研究的东西。我觉得那件事极具启发性，于是我基本就是问他：“

### [00:02:16–00:02:45]

**EN**  into Jeff Hinton and he basically said like I have this amazing idea you know these these models these deep learning models you can learn multiple layers of representation and everything and so he essentially dragged me to his office and started showing to me like what he's been working on and I think that was so inspirational to me that I was basically saying can I can I come and do a PhD with you on this topic and he sort of you know he made it happen he's like okay why don't you apply for your PhD program and everything come and work

**中文**  我能不能来跟你读博士、研究这个课题？”他就设法促成了这件事。他说：“好啊，你申请博士项目吧，然后过来跟我一起做。”于是我在 2005、2006 年开始读博。当时 deep learning 还处于非常早期，前景并不清楚；这些神经网络很受轻视，因为那时大家更关注 statistical machine learning 的理论以及数学上更严谨的东西。神经网络一直被看成这样一种非线性系统：大家只是在优化，却并不真正理解

### [00:02:43–00:03:13]

**EN**  with me and so you know in I started my PhD in 2005 2006 and at that time I think deep learning has been sort of uh you know it just was very early wasn't very clear what it's going to work these neural networks they were very sort of looked down upon uh because at the time people were looking at statistical machine learning sort of theory and you know uh something that's more rigorous and neural networks have always been looked at as like well these are nonlinear systems that people just optimizing without really understanding

**中文**  其中究竟发生了什么，对吧？不过，也正是在那时，一些早期 deep learning 模型被训练了出来。我跟 Geoffrey 完成博士，之后又去了 MIT 做 postdoc，诸如此类。后来的事情基本就是历史了。

### [00:03:11–00:03:40]

**EN**  what's what's going on right and So, but that's when some of the early deep learning models were were were trained and I, you know, did my PhD and then with with Jeff and then moved to MIT to do my pawn dog and and such. So, and then the rest is the history basically. >> Yeah. Yeah. But why why were neural nets like so exciting to you? Like what was he articulating to you that made it seem so interesting? >> Yeah. Yeah. Yeah. Jeff has like a very

**中文**  主持人：是的。但神经网络究竟为什么让你那么兴奋？他当时描述了什么，让你觉得这件事如此有意思？嘉宾：是啊。Geoffrey 很擅长……我觉得就是他的兴奋感染了我，让我相信这会是一个很有意思的博士课题。

### [00:03:36–00:04:04]

**EN**  good I think he was just his excitement and it convinced me that this is going to be an interesting PhD to do like interesting topic to do my PhD in. >> Yeah. I think it was part of it because like I think that at a time you know people were again neural networks have all like at that time neural networks been looked down upon as like yeah these sort of like systems they've been around for 20 years they never worked. They're like number third choice of any sort of

**中文**  我觉得这确实是一部分原因。因为当时人们一直看不起神经网络，觉得这种系统已经存在 20 年、却从来没奏效过；在各种方法里，它最多只是第三选择。

### [00:04:02–00:04:31]

**EN**  method at the time. you know, methods like support vector machines and more rigorously like more mathematical rigor behind those models. But you know, Jeff was one of the first people who showed that learning these stacked systems, learning these deep systems, right, can actually give you so you had he had like a mathematical justification for us, which is a beautiful, you know, beautiful math behind showing that, you know, these models can improve something. It's called variation

**中文**  当时像支持向量机这样的方法背后有更严谨的数学基础。但 Geoffrey 是最早证明学习这种堆叠系统、学习这种深层系统确实可以带来收益的人之一。他给出了一套数学论证；其中有非常漂亮的数学，表明这些模型能够改进某个指标，叫作 variational

### [00:04:28–00:04:58]

**EN**  lowbound. And and I think it it's it was exciting that I could see that this could have a strong potential in the future. >> And what what were the use cases that he was excited about or that you were excited about or was it more just oh the math was beautiful? >> No, no, no. Yeah. So, so there was the math was beautiful. That's number one because for the first time you could actually say okay there is you know there is like mathematical rigor behind these models. It's not like before that people were sort of you know training neural nets. was always like a little

**中文**  lower bound。我觉得这很令人兴奋，因为我能看到它未来潜力巨大。主持人：当时让他或让你兴奋的 use case 是什么？还是说纯粹因为数学很漂亮？嘉宾：不，不。数学很漂亮，这是第一点，因为这是第一次你可以说这些模型背后确实有数学严谨性。此前人们也会训练 neural net，但总有点像是

### [00:04:56–00:05:23]

**EN**  bit like you know there is some optimization you do optimization but there's no like understanding there is no sort of connection to statistical machine learning I mean there were some but not not as rigorous right and it was always like this field have always been looked down upon I mean you'd probably ask Yan Lun Yan Lakun was like working on these models for a long time but he could never get recognition in early days for for any of his work because it was just kind of like didn't didn't work right because the scale was small and

**中文**  做了一些优化；你确实在做 optimization，却没有真正的理解，也没有与 statistical machine learning 建立联系。其实也不是完全没有，只是没那么严谨。这个领域一直都被轻视。你大概可以问 Yann LeCun，他研究这些模型很久了，但早期始终很难让工作获得认可，因为它们就是没有真正奏效：规模太小，而且

### [00:05:21–00:05:51]

**EN**  and you know one of the interesting things that showed up early on is a glimpse into generative models. So at the time we were generating images of handwritten digits uh these M >> this was in the mid 2000s. >> It's the mid200s. Yes. And it was one of the early days that you could show that these deep models something is called deep belief network and then we were working on debols machines were initially the models that can sort of like generate you know uh these nice

**中文**  早期出现的一件有意思的事，是大家第一次窥见了 generative model。当时我们在生成手写数字图像，也就是 MNIST 数字。主持人：这是 2000 年代中期？嘉宾：对，2000 年代中期。那是很早的时候，我们可以展示这些 deep model——有一种叫 deep belief network，后来我们还研究了 Boltzmann machine——最初这些模型能够生成

### [00:05:49–00:06:19]

**EN**  looking handwritten digits like amnes digits right and it showed that hey you know these models are capable of going beyond multiple level they can learn multiple levels of of representation and that basically excited a lot of people both the math behind Plus that they showed the promise and they were sufficiently different from anything else that people were working in the community and sort of like at a time in 2005

**中文**  看起来相当不错的 MNIST 手写数字。它说明这些模型能够超越单层，学习多层表征。背后的数学，加上它们展现出的潜力，以及它们与当时社区中其他研究足够不同，令很多人感到兴奋。当时，也就是大约 2005

### [00:06:16–00:06:45]

**EN**  2010 it was like a small group of people who were essentially working on these models making progress on these on these models. >> And when did you finish your PhD? >> Finished my PhD in 2009. >> Okay. And then you went to MIT to do >> post to do my postdoc. Yes. I was doing my post dog u with Josh Tenbomb. fantastic professor and fantastic colleague of mine. Josh was sort of looking at deep learning models but he was also looking on these sort of

**中文**  到 2010 年间，实际上只有一小群人在研究并推进这些模型。主持人：你什么时候博士毕业？嘉宾：2009 年。主持人：然后你去了 MIT 做……嘉宾：做 postdoc，对。我的合作导师是 Josh Tenenbaum，他是一位非常出色的教授和同事。Josh 会研究 deep learning 模型，但也会研究

### [00:06:43–00:07:11]

**EN**  non-parametric basin models. These are sort of classes of models. I mean Josh is a cognitive neuroscient cognitive scientist and so um it was great hanging out and at MIT MIT also had a lot of strong computer vision people and such. So was a good experience and then I came back to Toronto as a faculty. >> Yeah. So when so it seems like yeah you were very excited about like the generating digits like that was very interesting but then it seems like when

**中文**  non-parametric Bayesian model 这一类模型。Josh 本身是 cognitive scientist，所以在 MIT 和他相处、共事很棒。MIT 还有很多很强的 computer vision 研究者，那段经历很好。之后我回 Toronto 当了 faculty。主持人：听起来你当时对生成数字很兴奋，但 AlexNet 出现后，所有人的重心似乎基本都转向了图像分类，而不是生成。

### [00:07:08–00:07:38]

**EN**  Alex net came out that everyone's focus basically shifted to classifying images rather than generating >> so what's what's so so what's happened early on is that if you take like at a time we didn't have GPUs if you sort of train these multi-layer neural networks on something like Amnest there's like there's a class of data sets like you know these old toy data sets C far you could never beat sort of competing methods like support factor machines and

**中文**  嘉宾：事情是这样的：早期我们没有 GPU。如果你在 MNIST、CIFAR 这类旧的小型数据集上训练多层神经网络，始终无法击败支持向量机之类的竞争方法。

### [00:07:36–00:08:05]

**EN**  these generative models showed the promise that if you actually learn these generative models and then convert them into discriminative models like classification systems they would actually they would actually work much better. So this this notion is that you know learning how to generate images basically gives us the representations that we need to actually do classification and then after imageet competition after alexnet people sort of realized well convolutional neural networks actually do work they work you

**中文**  而那些 generative model 展示出一种可能性：如果先学习生成模型，再把它们转换成 discriminative model，也就是分类系统，效果会好得多。核心想法是，学会如何生成图像，会给我们真正做分类所需的表征。ImageNet 竞赛和 AlexNet 出现后，人们意识到 convolutional neural network 的确有效，而且效果极其突出。于是大家转向这些 discriminative model，因为

### [00:08:02–00:08:29]

**EN**  know they work remarkably well and after alexnet people actually shifted to these discriminative models because that was you know they they've solved I think Alex net solved one of the key vision problems right >> which was the you know it was imageet competition and it's was like it was a million images very diverse images and people were looking at it over the over

**中文**  它们解决了问题。我认为 AlexNet 解决了 vision 的一个关键难题，对吧？主持人：也就是……嘉宾：也就是 ImageNet 竞赛。那是一百万张非常多样的图像，过去

### [00:08:28–00:08:56]

**EN**  the past you know two to three years like gigantic labs at Berkeley Oxford you know MIT they had like you know lots of you know PhDs and computer vision people were participating and I think Alexet was the first model that showed that you can improve performance by like 50% So you can imagine Alex Keski, Izzkerver and Jeff Hinton, three people they've built, you know, fairly simple connet,

**中文**  两三年间，Berkeley、Oxford、MIT 的大型实验室投入很多博士生和 computer vision 研究者参与。我认为 AlexNet 是第一个把性能提高大约 50% 的模型。你可以想象，Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton 三个人构建了一个相当简单的 ConvNet，

### [00:08:54–00:09:24]

**EN**  but they scaled it and they were the first ones to scale it on GPUs and they've shown that this model was able to improve over whatever people have done by 50%. And they've able to do it to the point where there was a lot of skepticism initially like people couldn't believe it, you know, like it's it's remarkable, right? When you see the result, you say that cannot be true. It's too good to be true. So it took people like couple of months to actually verify and say, "Okay, this is real." >> Yeah. >> So you were So you were there when Alex

**中文**  但他们把规模做大了，而且率先把它扩展到 GPU 上，最终比此前所有人的成果提升了 50%。他们做出的效果太好，以至于一开始有很多质疑，人们不敢相信。看到结果时你会说：“这不可能，简直好得不真实。”于是大家花了几个月去验证，最后才确认：“好，这是真的。”主持人：是啊。所以 Alex…

### [00:09:23–00:09:51]

**EN**  came out. >> I was there when Alexand came out. That's right. I I came from my pawns dog back to Toronto and Yeah. It was >> Were you skeptical? >> No, I wasn't skeptical of of of the results. I knew like Jeff was kind of like on it. So it's it's like it's you know one of the things that Alex Keski has done very well is he optimized some of the computations on early versions of GPUs which which required like a very strong engineering talent to do it because scale that was one of the

**中文**  Net 出现时，你也在那里。嘉宾：AlexNet 出现时我确实在那里。那时我结束 postdoc，回到了 Toronto。主持人：你怀疑过吗？嘉宾：不，我没有怀疑结果。我知道 Geoffrey 一直在推动它。Alex Krizhevsky 做得特别好的一件事，是针对早期 GPU 优化了一些计算。这需要极强的工程能力，因为 scale 是其中一个

### [00:09:49–00:10:18]

**EN**  earlier kind of like points where the scale actually showed big big success. >> Interesting. Um so that yeah that was the turning point for computer vision that was the turning point for computer vision researchers like from that point on Google all the other major labs and academics started basically converting to uh to deep learning. So what were you focused on before like what was your research focused on before Alex net came

**中文**  最早显示出巨大成功的关键点。主持人：有意思。所以它成了 computer vision 的转折点？嘉宾：对，是 computer vision 研究者的转折点。从那以后，Google、其他主要实验室以及学术界基本都开始转向 deep learning。主持人：AlexNet 之前你的研究重点是什么？

### [00:10:16–00:10:44]

**EN**  and then did your research shift when it came out and you were got really excited by it? So my research was kind of like I never looked at I mean I was looking at vision and generative models. I think my research you know I wasn't part of the Alex net because it was like a fairly strong engineering effort. My research kind of like started shifting after Alex net because one of the things that my lab has been looking at at Toronto also is earlier versions of caption generation. So once we started

**中文**  它出现、让你十分兴奋以后，你的研究方向有变化吗？嘉宾：我并不是只研究某一个方向，我当时在研究 vision 和 generative model。我没有参与 AlexNet，因为那是一项很强的工程工作。AlexNet 后，我的研究方向开始变化，因为我在 Toronto 的实验室还研究早期的 caption generation。所以，当我们开始

### [00:10:42–00:11:12]

**EN**  seeing that these models actually pretty good then we've started looking at models that given images generate descriptions for those images. This was one of the early I remember uh talking to some of the computer vision folks and then at a time classification detection was the thing to do but a lot of computer vision people were saying that well I'll actually I want a system where I show you the image and I just don't want to say well there's a car in this image over here and then there is like you know a person standing here. I

**中文**  看到这些模型已经相当不错后，就开始研究：给定图像，让模型为图像生成描述。我记得当时和一些 computer vision 研究者交流，那时 classification 和 detection 才是主流，但很多人说，我真正想要的是这样一个系统：给它看一张图，我不只想让它说这里有一辆车、那里站着一个人。我

### [00:11:10–00:11:39]

**EN**  actually want the system to describe me the scene. I want the system to look in and just like human tell me like oh that's you know it's a sunset you have a beautiful beach here you have people walking on the beach and sort of like just tell me what's like that's the you know a a way of much better understanding than if me running a classifier say is there a person there or not yes or no is there a car there yes yes or no right which is what the field was and so like early on these recurrent neural networks coupled with

**中文**  真正希望系统能描述场景。我希望系统看过之后像人一样告诉我：“这是日落，这里有一片漂亮的海滩，有人在沙滩上散步。”也就是直接告诉我里面有什么。相比运行一个 classifier，分别问“有没有人”“有没有车”并得到 yes/no，这体现了更深入的理解，而后者就是当时这个领域的做法。因此早期的 recurrent neural network 与

### [00:11:36–00:12:05]

**EN**  vision systems started showing progress and this is what you know we were looking at >> I had never even thought about that as a thing until like LLM came out. >> No, at a time people were kind of like and it's this was you know this was one of the early sort of models that showed promise and then early on people started looking at attention mechanisms. So we had a paper called um generating image with visual attention. So the model would attend to different objects when

**中文**  vision system 结合后开始显示进展，这就是我们当时在研究的东西。主持人：直到 LLM 出现之前，我甚至从没想过这会成为一个问题。嘉宾：对，当时人们已经开始做这个方向，它是早期显示出潜力的模型之一。之后大家开始研究早期的 attention mechanism。我们有一篇论文叫《Generating Images with Visual Attention》。模型会在

### [00:12:03–00:12:33]

**EN**  it describes different parts of the scene. So this was like earlier versions. I mean they never worked very well because it all was like very small scale right and then in parallel my lab was also looking at image generation early version of image generation and we had one of the papers one of the students Mitish Sastama was working with another student on sort of these models that could generate small images small patches of image I give you the description so it's like reverse give an image you generate the caption but then we started asking well if I give you the

**中文**  描述场景不同部分时关注不同对象。这就是早期版本。我是说，它们从来没有工作得很好，因为规模都非常小。与此同时，我的实验室也在做早期 image generation。我们有一篇论文，一位学生 Nitish Srivastava 和另一位学生共同研究一种模型，能生成小图像、小图像块：我给你一段描述，而你为我生成图像。这相当于反过来——给图像生成 caption；但我们开始问，如果我给你

### [00:12:31–00:13:00]

**EN**  captioning I tell you what it is can you generate me image um and there was a fun story we had uh we had a paper That was one of the early papers that would generate those im those images and wasn't very good wasn't very good model again like the scale was pretty small we didn't use GPUs but we started seeing early versions of you know like for example I can tell you a plane is flying in rainy skies and we try to sort of generate something that looks like a

**中文**  caption，告诉你里面是什么，你能否为我生成图像？这里还有个有趣的故事。我们有一篇很早的论文能够生成那些图像，模型效果并不好。同样，规模很小，我们也没用 GPU。但我们开始看到一些早期能力。例如，我可以告诉你“一架飞机在阴雨的天空中飞行”，然后我们会试着生成某个看起来像

### [00:12:58–00:13:28]

**EN**  plane in kind of like darkish background and then we would say a plane is flying in in in uh not in a rainy sky but in blue sky and the image would change completely like same plane but the background would be different. So it was one of the early systems that like if you look at distribution of pixels in one image versus another image it's very different. So the model would be able to generate something very different even though the caption was exactly the same except for rainy you would change rainy skies to blue skies. Right? So one word would change would impact the entire

**中文**  飞机的东西，放在偏暗的背景里。接着我们会说“一架飞机在……不是阴雨的天空，而是蓝天中飞行”，图像就会完全改变：同一架飞机，背景却不一样。所以，如果观察一张图像与另一张图像的 pixel distribution，它们会非常不同。即使 caption 完全相同，只把 rainy sky 改成 blue sky，模型也能够生成极其不同的东西。一个词的改变会影响它生成结果的整个

### [00:13:26–00:13:55]

**EN**  distribution of what the model was able to generate. And and then at that time we started thinking kind of like well can we actually generate images that don't exist in the real world because we wanted to test the ability of these models to generalize. So you can kind of like say well a school bus is parked in a street or we can say a school bus is flying in the sky. So can the model generate the school bus? Can the model generate the sky and kind of like make it look like it's flying in the sky.

**中文**  分布。那时我们开始想，能否生成现实世界中不存在的图像，以检验模型的 generalization 能力。比如可以说“一辆校车停在街上”，也可以说“一辆校车在天空中飞”。模型能否生成校车、生成天空，并且让它看起来确实在天上飞？

### [00:13:52–00:14:20]

**EN**  Right? So it was early versions of of these systems and you know there's arguments is does the system understand what it's doing or does it not understand what it's doing and then you know Jeff Hintington was earlier kind of like promoter by basically saying look if I if I generate you know a cat sitting on the beach drinking beer and the model can actually do it does it mean that it understands what it's generating or not right so these are arguments because those things don't exist in the real world right but if the

**中文**  对吧？这是这些系统的早期版本。那时有很多争论：系统究竟是否理解自己在做什么？Geoffrey Hinton 是早期的推动者之一。他基本会说：“假如我让模型生成‘一只猫坐在沙滩上喝啤酒’，而它真的做到了，这是否意味着它理解自己生成的东西？”这些争论的原因是，那种场景在现实中并不存在；但如果

### [00:14:19–00:14:48]

**EN**  model can generate and kind of hallucinate and such it means that it has has an understanding and we had this one uh the student uh Elman Mansimov he's I think he then you know was a student PhD student at NYU but he had like a cool idea of basically he generated the image of um a toilet sit sits open in the grass field right so something a toilet sit in a grass field like in the open sort of open grass

**中文**  模型能够生成，能够“hallucinate”，诸如此类，就意味着它拥有某种理解。我们有一位学生 Elman Mansimov，我想他后来成了 NYU 的博士生。他有个很有意思的想法，生成了一张“马桶座圈打开、放在草地上”的图像。也就是一只马桶座放在开阔的草

### [00:14:46–00:15:15]

**EN**  field and it would sort of generate like a grass like you can see like green and again these models were very primitive And we generate like a white kind of like looks like a toilet. We just did it for fun, right? And I was giving a talk and then somebody noticed that like actually if you ask Google images, you know, give me a toilet seat in the in the green grass. There was actually somebody took a picture and you know if you do Google search like one of the top

**中文**  地上。它会生成看起来像草的绿色背景；当然那些模型非常原始，又生成一个白色、看上去像马桶的东西。我们只是觉得好玩。有一次我做演讲，有人发现，如果让 Google Images 搜索“绿草地里的马桶座”，其实真有人拍过一张照片。如果你做 Google 搜索，排名靠前的

### [00:15:13–00:15:43]

**EN**  pictures is the picture of like somebody put like a toilet seat in the in the grass field, right? And people were saying well but is your model really can generate because it's already exists there right I mean Google can do better than than your model can but what was funny is that in couple of weeks if you ask the same question for Google images our picture would come up before any other picture. Oh wow. Right. >> Like it was indexed by Google. That's

**中文**  图片里就有一张别人把马桶座放在草地上的照片。于是有人说：“你的模型真的算是在生成吗？现实里本来就已经有了，对吧？Google 的结果都比你的模型好。”但有意思的是，几周后再用同样的问题搜索 Google Images，我们生成的图反而会排在所有其他图片之前。主持人：哇，因为它被 Google 收录了。嘉宾：

### [00:15:41–00:16:11]

**EN**  >> it was indexed by Google. And the reason why is because many more people would be clicking on our image because there's an interest in generation than on the real image. And what Google's ranking system would basically say, well, if I ask you like, you know, for for if I ask you for uh this prompt and then more more people click on this image than on this image, then this image gets upgraded. And so our image was at the very top. So we're very proud that we kind of you know so

**中文**  对，被 Google 收录了。原因是更多人会点击我们的图，因为大家对 generation 感兴趣，而不会点那张真实照片。Google 的 ranking system 会判断：如果查询这个 prompt 时，人们点击这张图的次数比另一张更多，这张图的排名就会上升。所以我们的图最终排到了最前面。我们还挺自豪的，因为我们算是……

### [00:16:09–00:16:37]

**EN**  it basically was showing that you know the ranking systems they don't I mean there's no reasoning right it's just indexing right and so >> yeah what did the data labeling look like at that point like where was a lot of your time spent basically having somebody look at images and just describing them in natural language >> no I think that we we there there were a few data sets there is something that's called Microsoft Coco like big labs would would label the images and we

**中文**  这其实说明 ranking system 并没有 reasoning，它只是在 indexing，对吧？主持人：是啊。当时 data labeling 是什么样的？你们是不是会花很多时间，让人看图片再用自然语言描述？嘉宾：不是。我想当时已经有一些数据集，比如 Microsoft COCO，大型实验室会负责给图像做标注，而我们

### [00:16:33–00:17:00]

**EN**  would just basically try to either you know, so we I I never spend time labeling the data. I think because like I was in machine learning in in in machine learning kind of like and so students in ML they don't want to I mean students in computer vision they kind of like tend to label more. We're just relying on data sets and images that are coming from from from other labs. So we never labeled but we we were trying to basically build algorithms and try to

**中文**  基本只会使用这些数据。我从来没有亲自花时间标数据，因为我做的是 machine learning；做 ML 的学生不想标，computer vision 的学生往往会多标一些。我们依靠其他实验室产出的数据集和图像，从未自己标注。我们努力构建算法，并尝试

### [00:16:58–00:17:26]

**EN**  identify uh because it's kind of like you know the fascinating field the fascinating piece that we have right now is that people don't like hallucinations right like people don't like when models hallucinate because you know it's generally bad. On the other hand, when we've looked at computer vision, we actually do want model to hallucinate in the sense that we want models to produce something new, something that's not part of the training data. If

**中文**  找到……因为这里有个很有意思的点：现在人们不喜欢 hallucination，因为模型产生幻觉通常很糟。但我们研究 computer vision 时，反而希望模型能够以某种意义“产生幻觉”：希望它创造新的、训练数据里没有的东西。如果

### [00:17:25–00:17:54]

**EN**  they're able to produce something that's not part of the training data, then it's then we believe that the model is actually has some form of generalization. I remember there was this argument where we were looking at models like genative adversarial networks, GANs and try to evaluate these models and you look at GANs these beautiful models they would generate like very nice looking images and then the question has always been well this particular image that it generated is

**中文**  它们能产生训练数据中不存在的内容，我们就会认为模型具有某种 generalization 能力。我记得当时有一场争论：我们在研究 generative adversarial network，也就是 GAN，并尝试评估这些模型。GAN 是很漂亮的模型，能生成非常好看的图像；接下来的问题总是：它生成的这张特定图像是否

### [00:17:51–00:18:19]

**EN**  this is this like truly generated image or is it something that you kind of picked from the training set and just modified a little bit and you know because when we were looking at images like a One one sort of great generative model would be to just pick a training example and show it to you. Right? So, so I ask you to generate me images of horses. I go to my training data, I pick images of horse and I show it to you, right? And if you as a user, you look at like this is great, right? These are

**中文**  这究竟是真正生成的图像，还是模型从训练集里挑了一张、稍作修改？因为研究图像时，一个“很棒”的 generative model 完全可以只挑一条训练样本展示给你，对吧？比如我让它生成马的图像，它就去训练数据里挑几张马的图片给我。如果用户看了说：“真棒，这些都是

### [00:18:18–00:18:46]

**EN**  true images of horses, but it's meaningless. All right? So, you so so we were sort of trying to test, you know, can it generate different variations of the horses? Can it generate different viewpoints, different backgrounds? So does the model really sort of understand pieces and then you started getting into these problems where even for image recognition tasks you know if you try to recognize cows early versions you know the model would do really good at

**中文**  真实的马匹图像。”但这毫无意义。所以我们会测试：它能不能生成马的不同变体、不同视角和不同背景？模型是否真的理解了各个组成部分？于是你也会碰到 image recognition 的问题：早期模型识别奶牛时表现很好，

### [00:18:44–00:19:13]

**EN**  recognizing a cow right but then if you put the cow on the beach and you tried ask system what is this the model would start confusing these cows with like boats right >> where it's kind of like when people always use the example of like a chihuahua with a bunch of blueberry muffins. Yeah. Yeah. That's that that's one example. That's right. That's right. And and so like it's a sort of like and then part of it is because you know when you look at cows, you know, they're usually like in a grass field and you

**中文**  但如果把奶牛放到海滩上，再问系统“这是什么”，模型就会开始把奶牛和船混淆。主持人：就像人们总举的例子，把 Chihuahua 和一堆 blueberry muffin 放在一起。嘉宾：对，那就是一个例子。部分原因是，我们平时看到的奶牛通常都在草地上，而你

### [00:19:11–00:19:38]

**EN**  know like you'd ever see cows on the beach, all right? So if I artificially put cow on the beach and I ask them all what is this like it will start confusing and so this was like one of the early failure points and you know right now people sort of can handle this. So like now vision systems are getting to the point where like really well you know they work remarkably well com even compared to like you know to to human accuracies.

**中文**  几乎不会在海滩上看到奶牛。所以，如果人为把奶牛放到海滩上再问模型它是什么，模型就会混淆。这是早期的 failure point 之一。现在大家基本能处理它了；如今 vision system 已经做得非常好，即便与人类准确率相比也表现出色。

### [00:19:37–00:20:07]

**EN**  >> Yeah. >> So let's go to like 2015. So you started a company called perceptual machines. >> Yes. So like what like what was the genesis of that business? So we basically it's it's interesting because you know at the time I couldn't talk about this but myself and couple of my students Mitastav and Charlie Tang we started a company looking at detection systems. So early on at a time people were looking how well can you detect you

**中文**  主持人：我们谈谈 2015 年。你创办了一家叫 Perceptual Machines 的公司。这个业务是怎么来的？嘉宾：很有意思，当时我不能谈这些。我和两位学生 Nitish Srivastava、Charlie Tang 一起创办公司，研究 detection system。当时人们关心图像中的对象检测能做到多好，而系统开始

### [00:20:05–00:20:33]

**EN**  know how well can you detect objects in images and at a time it started working well enough that the systems could actually do well right before it was you know after addict paper classification started working people started looking at at at at sort of detection and detection was one of these sort of you know challenges where I want to be able to look at the image and just say like oh there's like five people here this is where are this is the way they uh this

**中文**  好到真正可用。AlexNet 论文之后，classification 开始奏效，人们继而研究 detection。这类挑战要求系统看一张图后说：这里有五个人，分别在这些位置；这是一辆车，它在这里，等等。

### [00:20:31–00:20:59]

**EN**  is a car, this is where it is and such. And we started working with Apple. At the time, Apple was also looking at autonomy and sort of like trying to build a strong computer vision systems. And we started working with them. And I think that eventually we sold the company to Apple. But uh it was interesting because you know at a time a lot of work was done using traditional computer vision techniques and it was very clear that the new wave of deep

**中文**  我们开始和 Apple 合作。Apple 当时也在研究 autonomy，尝试构建强大的 computer vision system。后来我们把公司卖给了 Apple。那时很多工作仍使用传统 computer vision 技术，但很明显，新一波 deep

### [00:20:57–00:21:26]

**EN**  learning models was just much better than whatever traditional computer vision techniques could do right and >> and that new what was that new technique >> this deep deploying based so these >> like CNN's or >> like there was CNN's uh these are convenient layers learning multiple layers of representation specific laws functions, detection loss functions. Whereas traditional computer vision techniques were based on like well let's find bunch of features and design those features,

**中文**  learning 模型远胜于任何传统 computer vision 技术。主持人：那种新技术具体是什么？嘉宾：以 deep learning 为基础，比如 CNN、convolutional layer、学习多层表征、特定的 loss function 和 detection loss function。传统技术则会人为设计一批 feature，

### [00:21:24–00:21:54]

**EN**  build some kind of like linear models, build like support vector machines and you know and then like these are kind of like very kind of arcane sort of systems. Whereas these deep learning system combat systems started working well enough that when you learn them at scale they showed way stronger performance than whatever was happening, right? And so then we started working with Apple and showed that this technology was way better than whatever they were doing. And so it was making sense for us to go and and work. And I

**中文**  再构建 linear model、support vector machine 等相当晦涩的系统。deep learning 与 ConvNet 系统已经足够成熟，一旦进行大规模学习，性能就远超以往。我们和 Apple 合作，证明这项技术远胜于他们原来的方案，于是加入他们就很合理。

### [00:21:52–00:22:20]

**EN**  spent three years at Apple uh with a team. We were working on autonomy at a time. >> Autonomous cars. >> Autonomous cars with a project called Project Titan. And so this was you know we essentially were building the perception stack like early version of perception stack like you know recognize objects and images and so forth, right? So this was uh this was pretty cool very cool application and of course you know it goes far beyond that there is other

**中文**  我和团队在 Apple 工作了三年，当时研究 autonomy。主持人：自动驾驶汽车。嘉宾：对，是 Project Titan。我们基本上在构建 perception stack 的早期版本，例如识别图像中的对象。这是非常酷的应用。当然整套系统远不止这些，还有

### [00:22:18–00:22:48]

**EN**  other parts of the stack and actually but the perception system the initial perception system was based on on on deep learning and it's sort of yeah >> sorry no go ahead. Yeah. Yeah, I was going to say so early on so with self-driving cars basically the unlock was hey now we can actually recognize what's going on in front of the car and so we can use that to actually maneuver the car. Is that basically >> that's exactly right. So the way >> go ahead. >> No, go ahead. Go ahead.

**中文**  其他部分，但最初的 perception system 是基于 deep learning 的。主持人：所以自动驾驶汽车早期的关键突破是，我们终于能识别车前发生的事情，再用它操纵车辆，对吗？嘉宾：完全正确。主持人：请继续。嘉宾：不，你先说。

### [00:22:47–00:23:15]

**EN**  >> I was going to say I was going to say like Yeah. So yeah, it seems like around that time there was a lot of interest in self-driving cars. It seems like then for a few years it kind of went nowhere and then basically in the last couple years things have picked back up and now cars can actually full self-drive themselves. So like what what changed? I think that what was happening so the whole sort of self-driving business started in 2015

**中文**  主持人：那段时间自动驾驶汽车很受关注，随后几年似乎没有进展，而最近几年又重新加速，现在汽车已经能真正全自动驾驶。发生了什么变化？嘉宾：整个自动驾驶行业大约始于 2015、

### [00:23:13–00:23:41]

**EN**  2016 right and what happened with self-driving cars and with technology like perception started working way better and what happened is that we went from zero to like 80% in a span of like a year or two years. >> What did 80% look like? It's sort of like you know like you let's say you sit in a car right you can detect like most of the objects you can detect cars you can detect pedestrians you can detect the road and sort of like

**中文**  2016 年。当时 perception 技术明显进步，我们在一两年内从 0 做到了大约 80%。主持人：80% 是什么样？嘉宾：比如你坐进车里，它能检测绝大多数对象，包括汽车、行人和道路，

### [00:23:39–00:24:06]

**EN**  but then you started hitting corner cases right one famous example was by Google because they were the first in this business is that there was like a bag flying on the highway and the model just recognized it as like some object didn't know what it was and start slamming on the brakes right in the middle of the highway way, right? Like or there was another example where towards like end of spring vegetation

**中文**  但接着就遇到 corner case。Google 是最早进入这个行业的公司，一个著名例子是高速公路上有只袋子飘过，模型只认出那是某个对象，却不知道是什么，于是在高速公路中央猛踩刹车。另一个例子是春末时树木的枝叶

### [00:24:05–00:24:33]

**EN**  started like you know like from the trees and the systems that would look at perception looked at the lighter and everything would basically say there's like some objects are secluding kind of like my my path so I can't go through that even though it was just like bunch of bushes kind of like you know over kind of and so what I mean by from 0 to 80 is like basic stuff was there and then you started hitting these corner cases construction like people start putting

**中文**  开始长出来，perception 系统结合 LiDAR 看到它们后，会认为有物体遮住行驶路径、车辆不能通过，尽管那只是一簇伸过来的灌木。所以所谓从 0 到 80%，就是基本能力已经有了，接下来却撞上施工等 corner case，例如有人开始在路上摆

### [00:24:30–00:24:58]

**EN**  cones or like something you know on the road right and that's what basically stopped a lot so it was a lot of excitement because like even in the early days I think people were basically saying okay in two years we're just going to have full self-driving cars people were now studying reports what's going to happen to truck drivers what's going to happen to taxi drivers within the next two to three years like the whole economy is going to change right so people were like very sort of feels a

**中文**  锥桶或其他东西。这基本阻断了进展。早期大家太兴奋，甚至说两年内就会有完全自动驾驶汽车；人们开始研究未来两三年卡车司机、出租车司机会怎样，认为整个经济都会改变。这种感觉

### [00:24:56–00:25:24]

**EN**  little bit like AGI we Right now people say AGI is 2 years and people who actually work in these lab and frontier labs I know a lot of them they're basically saying just like we're not going to have AGI in 2 years. So I think there was excitement in self-driving cars is like look in two years we're going to have full self-driving cars everywhere. The other corner cases that started picking up is weather snow like all these other things right and perception is one part of it like then

**中文**  很像今天的 AGI。现在有人说 AGI 两年可达，但真正任职于这些实验室和前沿实验室的人——我认识很多——基本都会说，两年内不会有 AGI。当时自动驾驶也一样：大家说两年后会遍地都是完全自动驾驶汽车。之后还有天气、降雪等 corner case。perception 也只是一部分，接下来还有

### [00:25:22–00:25:51]

**EN**  planning. How do you plan things right? Like for example, now Whimo has a way of actually being a little bit aggressive. Like earlier versions, like earlier versions of these cars, like you couldn't you couldn't merge into the highway because there was a lot of safety things. So you you you know, you try to merge on the highway and just people just don't let you merge, right? Because that's how we drive. You know, if you're slow and you don't, you know, you're not aggressive enough, just nobody lets you in. And so they they just couldn't merge, you know, if they

**中文**  planning：怎样规划动作？例如现在 Waymo 已经会稍微强势一点。早期版本的车辆因为安全限制，连汇入高速都做不到：你想并线，但别人根本不让你进，因为人类就是这样开车的。你慢吞吞、不够果断，谁都不会让你。所以有其他车辆时，它们就无法并线。

### [00:25:50–00:26:19]

**EN**  were cars. And and the way that humans do it is they basically like well you know I'm going to go right and and whoever is behind me we'll have to stop because otherwise you know and and so we kind of adapt to this right and early version systems could not and so they were just not you know and so the whole field was stuck for about four or five years primarily because of these corner cases and like all the nuances right and even today you have way more even today

**中文**  人类的做法其实是：“我要并进去了，后车必须减速，否则就会撞上。”我们会彼此适应，但早期系统做不到。于是整个领域主要因为这些 corner case 和种种细节停滞了四五年。即便今天，Waymo 仍会碰到 corner case；

### [00:26:16–00:26:42]

**EN**  you know you see like corner cases happening and even they you know they expanding to cities but this expansion is very slow part of it because you have to have high definition maps you have to sort of know exactly what the road structure is and everything so it's not like you have a system you just turn it on and it just works anywhere >> but it seems like Tesla FSD can basically run everywhere because like I've rented it and I've gone to

**中文**  它虽然在扩展到不同城市，但速度很慢。部分原因是需要 high-definition map，必须精确掌握道路结构等信息；并不是有套系统，打开就能在任何地方运行。主持人：但 Tesla FSD 似乎基本可以在任何地方运行。我租过它，去过

### [00:26:41–00:27:10]

**EN**  different cities and it it just works >> it works it works yes but with Tesla FSD it it's a great sort of system But at the same time, you know, with a self-driving cars, being 99 correct is not good enough. I think that's that's one of the things that, you know, it's kind of it's interesting because I think that these the technology could save a lot of lives. Right now, when I'm in San Francisco and I drive on Whimo, I

**中文**  不同城市，它都能工作。嘉宾：确实能工作，Tesla FSD 是套很棒的系统。但对自动驾驶汽车来说，99% 正确仍不够。我觉得有意思的是，这项技术本可挽救很多生命。现在我在 San Francisco 坐 Waymo 时，

### [00:27:08–00:27:36]

**EN**  feel safer than when I have an Uber driver. Like, you know, I've d I've driven multiple times. >> I agree. But at the same time there is this sort of argument that you know 99.9% is just not good enough you know like because you get into the accident that's you know Uber had one accident and they've shut down their entire kind of like operations right

**中文**  感觉比坐 Uber 司机开的车更安全。我坐过很多次。主持人：我同意。嘉宾：但另一方面，也有人认为 99.9% 仍不够，因为一旦出事故，后果严重。Uber 发生过一次事故，就关闭了整个自动驾驶业务。

### [00:27:32–00:28:01]

**EN**  >> yeah what changed since like you know 2017 2018 to now where okay you know we were at 80% now we're at 99%. Sure, we still have that, you know,.99% to go maybe, but like what allowed that >> I think dramatic increase. >> I I think it's just the uh data the data collection and continuous improving of the models. >> Okay. >> Obviously, Whimo has LAR. So they, you

**中文**  主持人：从 2017、2018 年到现在发生了什么，让我们从 80% 达到 99%？当然可能还剩最后 0.99%，但进步靠什么？嘉宾：我认为是巨大的提升……核心就是数据、数据收集，以及持续改进模型。主持人：明白。嘉宾：Waymo 显然有 LiDAR，所以它们

### [00:28:00–00:28:30]

**EN**  know, they look at the cameras, they're running deep learning models on the cameras. There some of the architectures might have changed. Maybe we went through CNN's convolutional neural net to transformer based architectures. There's a few sort of nuances. These are hybrid systems and you know Tesla FST is still relying on I think CNN's and transformers. Low level processing happening with CNN's the high level processing happens with transformer based architectures. I mean LLMs are based on that. I think it's just you know to train these models you need a

**中文**  既看 camera，也在 camera 上运行 deep learning 模型。一些架构也许变了，比如从 CNN、convolutional neural net 转向 transformer-based architecture，还有一些细节。这些都是 hybrid system。Tesla FSD 我认为仍结合 CNN 与 transformer：low-level processing 用 CNN，high-level processing 用 transformer-based architecture，LLM 也是基于后者。训练这些模型需要

### [00:28:28–00:28:58]

**EN**  lot of data and you need a lot of corner cases right. >> Yeah. Which I guess Tesla has a lot of because they just have cars driving all over the place. >> Yes. Exactly. Cars driving all over the I think they're they're very well positioned to just because of the amount of data that they have. On the other hand, what you know, like right now, you know, with with taxi caps, like what Tesla is doing, like just having cars without wheels, that's going to be the final frontier, right? So, like >> right now, you sit in your Tesla and you still kind of like, you know, you have

**中文**  大量数据和大量 corner case。主持人：Tesla 应该有很多数据，毕竟车辆遍布各地。嘉宾：没错，单凭数据量，它们的位置就很有利。另一方面，Tesla 现在想做的出租车——造出没有方向盘的车——会是最后一道关卡。现在你坐在 Tesla 里仍然

### [00:28:56–00:29:26]

**EN**  to control like you still have to be in the car like the, you know, >> Yeah. like at the at the very end like when you get to your parking your parking lot then it kind of messes up sometimes. Yeah. >> Yeah. Yeah. So it's like it's not like and then no truck industry is the same sort of you know but at the same time you know like you still have to be at the wheel. The question is when are we going to get to the point where there's no wheel in the car. You just sit the back of the car and you just don't even need to worry about you know.

**中文**  需要控制，仍必须人在车里。主持人：对，最后驶入停车场时它有时还是会出错。嘉宾：对，卡车行业也一样，你仍必须握着方向盘。问题是何时能做到车里根本没有方向盘：你只坐在后排，什么都不用操心。

### [00:29:24–00:29:53]

**EN**  So that's you know that's going to take some time but you know >> Yeah. Do you think do you think there's going to be something like an iOS versus Android thing where every other company that wants a full self-driving system will basically build will work together to do it because they have way more cars compared to Tesla? >> Yeah, it's a good question. I don't know. I think every every car manufacturer is trying to look into this, but I suspect, you know, Google's business is probably going to be they'll

**中文**  这还需要一段时间。主持人：未来会不会像 iOS 与 Android？其他所有想要完全自动驾驶的公司会联合开发，因为它们合起来的车比 Tesla 更多。嘉宾：这是个好问题，我不知道。每家汽车厂商都在研究，但我猜 Google 的业务模式可能是由他们

### [00:29:51–00:30:20]

**EN**  develop the tech. Uh the only the only problem with Google to be to be honest is that right now the economics is not working out. just because they have so many sensors, lighter and everything. When something breaks, you have to maintain it. And it's just when something goes wrong, you have to maintain it. So, you kind of like right now it's more expensive to have self-driving cars than just having, you know, a person driving a car, right? But some point it will kind of break even and then you know it's but but I do

**中文**  开发技术。坦率说，Google 唯一的问题是现在经济账还算不过来，因为车上有太多 sensor、LiDAR 等设备；设备会坏，也要维护。出了问题就得维修。所以眼下自动驾驶汽车反而比雇人开车更贵。不过总有一天成本会达到平衡。我

### [00:30:17–00:30:45]

**EN**  think that essentially all cars in the future are going to be self-driving cars and there's probably like somebody like Google or Tesla what whoever whoever is building that tech will build in such a way that can be integrated into any other car. >> I see. So >> I'm actually just curious about the architecture of the the the neuronet that's like running FSD. I don't know if you know much or maybe you you can try to guess. >> Yeah, I can try to guess. I don't know. I don't know what FSD is doing, but I

**中文**  认为未来所有汽车最终都会自动驾驶；Google、Tesla 或其他技术提供者，大概会把技术做成可集成到任何汽车里的方案。主持人：明白。我很好奇运行 FSD 的 neural net 架构。你也许了解，或者可以猜测。嘉宾：我可以猜，但不知道 FSD 的具体做法。

### [00:30:43–00:31:12]

**EN**  suspect they're doing some form of CNN plus transformer based architecture like CNN for the low-level processing plus some form of um you know transform architecture that kind of like synthesizes bunch of different sensors. And I suspect a lot of you know Whimo for example what they doing is that you know it's a hybrid system. So it's not people typically think well deep learning system you know like images come in and then actions come out like how how much you should steer the wheel and such. There are some companies like

**中文**  我猜是某种 CNN 加 transformer 的架构：CNN 做 low-level processing，某种 transformer architecture 汇总不同 sensor。我想 Waymo 的方案是 hybrid system。人们通常以为 deep learning system 就是图像输入、驾驶动作输出，比如该把方向盘打多少。确有一些公司，例如

### [00:31:10–00:31:39]

**EN**  Wave, I think, that are trying to do that. But what way more and more serious companies are doing is they basically taking all the input and then creating a 3D map of the environment. They know exactly where the cars are. They know exactly what the roads, they know exactly how far these cars are moving. And then there is something that's called a planning module. A planner basically plans, replans like I think at eight times a second. And it sort of tries to say like okay should I stay if I if I'm trying to change the lane what

**中文**  Wayve，试图这样做。但 Waymo 和更严肃的公司会接收所有输入，再建立环境的 3D map。它们准确知道汽车在哪里、道路是什么样、车辆移动多快。随后有一个 planning module；planner 大概每秒规划、重规划八次，判断例如我要变道时

### [00:31:37–00:32:07]

**EN**  path should I take what sequence of actions should I take so it's almost like you know sensors give you a 3D representation of what's happening around you and then you have a planet that's actually saying like okay changing the changing lane so if I need to go faster if I see the red light I need to stop all of that is happening on top of that representation. I see. >> Interesting. And then like how how like are is it basically they have like multiple decoders or something that are

**中文**  该走什么路径、采取什么动作序列。也就是说，sensor 给出周围环境的 3D representation，planner 再基于它判断：变道时是否要加速，看到红灯是否要停车。主持人：明白。有意思。那它是不是有多个 decoder 一起

### [00:32:04–00:32:34]

**EN**  working together because like how how do you sync up the steering wheel along with the gas pedal along with the blinkers? >> Yeah. Yeah. So they they basically have some form of uh motion like something called what is it MPC? This is kind of like predictive control systems where all the sensors sort of come in and then you have a policy that tells you exactly what the wheel configuration should be, what the gas configuration should be and like all the other sort of you know so

**中文**  工作？怎么让方向盘、油门和转向灯同步？嘉宾：它们会使用某种 motion……可能叫 MPC 的 predictive control system。所有 sensor 输入后，policy 会精确给出方向盘和油门应处的状态，以及其他配置。

### [00:32:32–00:33:01]

**EN**  it's essentially a policy that spits out actions and these actions would be you know composition of steering wheel pedal on the gas or brakes and such. So I think that the critical piece is perception and that's why you have LARS and you know Tesla doesn't because Tesla believes we just have cameras but Tesla sort of was doing it. I I you know to me I believe the more senses you have the better because you can kind of like observe the world around you better but

**中文**  本质上是由一个 policy 输出动作，而动作由转向、踩油门或刹车等组合而成。关键部分是 perception，所以才有 LiDAR。Tesla 没用 LiDAR，因为他们认为 camera 就够了。我个人相信 sensor 越多越好，因为能更好地观察周围世界，但

### [00:32:59–00:33:28]

**EN**  lighter just looks ugly. You know you cannot sell cars with lighters and so like Tesla could not afford like you know it just looked so ugly right? So they >> is lighter really expensive as well? I I don't know. >> Lighter is expensive. The prices are going down but it is expensive but it's this thing that sort of like sits off and like you know it spins and it's just like like nobody would buy a car like I would never buy a car with like bunch of lighters like it just looks ugly, right? It's [laughter] it just so Tesla had to go with a vision only because that's the

**中文**  LiDAR 太难看了。车顶装着 LiDAR 的车卖不出去，Tesla 承受不起这种外观。主持人：LiDAR 也很贵吧？嘉宾：很贵，虽然价格正在下降。它突出在车外，不停旋转；没人愿意买装着一堆 LiDAR 的车，我就绝不会买，实在太丑了。所以 Tesla 只能走纯 vision 路线，因为那是

### [00:33:27–00:33:56]

**EN**  only way. Whimo took a different path where basically said we're going to put all the sensors because that allows us to be more precise around what objects around you, right? I can tell you there's a car here. >> I don't need computer vision techniques to detect and understand. I know there's a car because you have these laser bouncing around and sort of tell me where things are. So, and of course, if your goal is to design a cab that just something that comes to you, picks you

**中文**  唯一可行的办法。Waymo 选择了不同路径，装上所有 sensor，以便更精确地了解周围对象。我可以知道这里有辆车，不必依靠 computer vision 去检测和理解，因为 laser 反射会告诉我物体在哪里。当然，如果目标只是设计一辆接你、把你从 A 点送到 B 点的出租车，你大概不会在乎它有没有 LiDAR 或其他设备。

### [00:33:53–00:34:22]

**EN**  up and drives you from point A to point B, you probably don't care if it has ladders and whatnot, right? >> Yeah. >> But if I want to own the car, I don't want it's like expensive. It breaks. It's uh you know so it's interesting to see like you know two different viewpoints. One is just vision based the other one is like more sensor based. >> Yeah interesting. So after perceptual labs you went back to be a professor >> at CMU. >> Yeah.

**中文**  主持人：对。嘉宾：但如果我要拥有这辆车，就不希望它昂贵、容易坏。所以两种观点很有意思：一条路线只靠 vision，另一条则更依赖 sensor。主持人：Perceptual Machines 之后，你回 CMU 当教授。

### [00:34:20–00:34:49]

**EN**  >> And then you joined meta super intelligence lab. Why did you join initially? >> Yeah. So this is kind of like interesting because like one of the projects that we're working at CMU on and we continue working on right now is uh what we call computer agents. >> What year was this by the way? >> This was 2022. >> When you joined Meta? >> Uh no I joined the meta in 2024. >> Okay. Okay. So you were at CMU2. Yes. I

**中文**  后来你加入 Meta Superintelligence Labs，最初为什么加入？嘉宾：这很有意思，因为我们在 CMU 做、如今仍在继续的一个项目，是所谓 computer agent。主持人：这是哪一年？嘉宾：2022 年。主持人：你加入 Meta 的时候？嘉宾：不，我 2024 年加入 Meta。主持人：好，所以当时你在 CMU？嘉宾：对。我

### [00:34:47–00:35:16]

**EN**  was I was at CMU since 2020. Uh well actually was at CMU since 9 since uh but I was at Apple during that time. So I came back to CMU and in 2022 we started 2022 2023 we started looking at these agentic systems when Chad GPT came out it was kind of like amazing what it can do what it could do at a time so we started looking with a student his name is JY K um he was at Google before he's doing a PhD with me and another

**中文**  从 2020 年起在 CMU——其实更早就在 CMU，只是中间去了 Apple。回 CMU 后，我们在 2022、2023 年开始研究 agentic system。ChatGPT 出现时，它当时的能力令人惊叹。于是我和学生 Jiayuan Mao 一起研究；他之前在 Google，现在跟我读博士，还有另一位

### [00:35:13–00:35:42]

**EN**  professor Daniel Fred and so we started looking at building agentic systems because the thinking was that well early on if you have an agent that can sort of do things on your behalf you know that would be interesting and so we we there was a research uh there's a lot of research happening at CMU one of the things that we were looking at is we were looking at these multimodel systems because we always were looking at combining images and text and actions so

**中文**  教授 Daniel Fried。我们开始构建 agentic system，因为早期的想法是：如果有个 agent 能代你做事，那会很有意思。CMU 有很多相关研究；我们研究的方向之一是 multimodal system，因为我们一直关注怎样结合图像、文本与动作。

### [00:35:40–00:36:09]

**EN**  we've built something that's called visual web arena um a bunch of students at CMU and this was one of the early systems that you know you give it a task like hey find me find me the cheapest iPhone on Craigslist and you know make a posting offering $20 less than the asking price. And so the system would have to sort of navigate to the appropriate web pages, find the right information, know try to compare what's happening, what's you know what's the

**中文**  我和 CMU 的一群学生构建了 VisualWebArena，这是早期系统之一。你可以交给它一项任务，例如：“去 Craigslist 找到最便宜的 iPhone，并发帖出价，比卖方要价低 20 美元。”系统必须导航到合适的网页、找到正确信息、比较各种情况，弄清

### [00:36:07–00:36:36]

**EN**  cheapest price, what's the right thing and then you know post the message and uh so you can kind of like autonomously trying to solve these tasks. We built this simulation environment that simulates Amazon simulates Craigslist simulates Reddit with some of the tasks I think we had in about thousand tasks 950 tasks that we've designed ourselves and we sort of started building this environment and you know building these agentic systems and at a time we got an

**中文**  最低价格和正确选项，然后发布消息，从而自主解决任务。我们建立了一个模拟环境，模拟 Amazon、Craigslist 和 Reddit，并自行设计了约 950 项、接近 1,000 项任务。我们开始构建这个环境和 agentic system，当时引起了

### [00:36:33–00:37:02]

**EN**  interest from Microsoft and Meta so we went to Microsoft we've talked to the Microsoft folks they were very excited and it's you know at the same time I started talking to Meta and Meta at a time they were open source ing their models. So that was a big appeal to us, you know, like early versions of Llama models at the time they were coming up with Llama 2, Llama 3, uh, and they were like fully open sourcing for everybody to use. And, you know, Llama 3 was a great model. And so it was making sense for us to say, okay, well, if they open

**中文**  Microsoft 和 Meta 的兴趣。我们与 Microsoft 团队交流，他们非常兴奋；同时我也开始和 Meta 谈。Meta 当时在 open source 他们的模型，这对我们很有吸引力。Llama 的早期版本、Llama 2 和 Llama 3 陆续推出，而且完全开源给所有人使用。Llama 3 是很棒的模型。所以我们会想，如果他们要把

### [00:37:01–00:37:31]

**EN**  sourcing everything from from our perspective was it was pretty cool environment to go and try to build some of these agentic systems inside a bigger lab because they had more compute, more data, more infrastructure for training these models. Why was why was them building open source models really interesting to you? >> I think it's because being in academia, most of my students were using Llama 3. >> Oh, I see. >> And Llama 2, right? And this was like, okay, well, if we go and they build some

**中文**  所有东西开源，那么在更大的实验室内部尝试构建 agentic system 会很理想，因为那里有更多 compute、data 和训练 infrastructure。主持人：他们构建开源模型为什么对你如此重要？嘉宾：因为在学术界，我的大多数学生都在使用 Llama 3。主持人：明白。嘉宾：也用 Llama 2。所以，如果我们过去构建一些

### [00:37:29–00:37:58]

**EN**  of the agentic systems like if they open sourcing it, then it's kind of like helps us as academics, it helps kind of like also for us, we're not constrained. We would publish the findings and you know, and I think that was appealing to me. was appealing to students and so you know we started working with we we actually joined Meta and spent a couple of years at Meta. >> Yeah. So like so you guys were building like early computer use models basically. >> That's right. We were building early

**中文**  agentic system，而他们又把它们开源，那既帮助学术界，也让我们不受限制，可以发表研究成果。这对我和学生都有吸引力。于是我们加入 Meta，在那里工作了几年。主持人：你们基本在构建早期的 computer-use model。嘉宾：对。

### [00:37:57–00:38:26]

**EN**  computer use models. >> So like even >> Yeah. >> Sorry. Go ahead. >> There was a team that's also was looking at that. So we were you know we're joining we were joining forces and sort of like trying to source the right sort of data because like for computer use web navigation you need you need a lot of grounding data you need a lot of kind of datas to to teach the model how to click on things how to you know and so yeah >> and now that's like a big business it seems like everyone wants to build RL environments

**中文**  我们在构建早期 computer-use model。另一个团队也在研究它，于是我们合并力量，并尝试获取合适的数据。computer use 和 web navigation 需要大量 grounding data，用来教模型怎样点击等。主持人：如今这似乎成了大生意，每个人都想构建 RL environment。

### [00:38:24–00:38:52]

**EN**  >> yes that's actually actually yes but right now with code building these RL environments is becomes a little bit easier I mean right now in my lab like after we left CM you in my lab, we continue building some of some of these models. We we've built an environment called Odysseies, which is sort of an environment. I mean, it's not an environment, but it's more like a task where I give you very complex tasks that would require, you know, a human a few hours to solve. And sort of now we're

**中文**  嘉宾：确实如此。现在借助 code，构建 RL environment 变得更容易。离开 CMU 后，我的实验室仍在继续开发这些模型。我们构建了一个叫 Odysseies（英文自动字幕拼写）的任务集；与其说是环境，不如说是一组任务。我会给模型非常复杂、需要人类数小时才能完成的任务。现在我们

### [00:38:51–00:39:20]

**EN**  progressing from sort of simpler tasks that we were doing before to more complex tasks that would again require you many hours. And there something that we call long horizon tasks where the the system actually has to go do research, find information and sort of do the right sort of planning and such. >> Yeah. So like what were some of the early challenges that you guys had to solve when you were building those computer models because even now I feel like computer models like the the >> It's still early. It's still early.

**中文**  正从过去较简单的任务转向更复杂、同样需要花数小时的任务。这些叫 long-horizon task，系统必须开展研究、查找信息、正确规划等。主持人：构建那些 computer model 时，你们早期必须解决哪些挑战？因为即便现在 computer model 仍然……嘉宾：仍然很早期。

### [00:39:18–00:39:47]

**EN**  >> Yeah. Yeah. It's it's not Yeah. They're not amazing compared to the LLMs for example. Yeah. >> No. And I think that's what makes it exciting because you know obviously these models like the thing with these models with computer use models there's a terminal aspect of it where like through MCP and such where you can communicate with a particular website with a particular program through the specific protocol using LLM and they were kind of like working and kind of don't and there is this argument that if

**中文**  主持人：对，与 LLM 相比它们并不出色。嘉宾：这正是令人兴奋之处。computer-use model 一方面可以通过 MCP 等终端协议，让 LLM 与特定网站或程序通信；这种方式有时有效、有时不行。另一种观点认为，如果

### [00:39:45–00:40:14]

**EN**  you're able to build a system that can control visual interface would it be your computer website or whatever visual interface you have because people are inherently visual and we sort of you know then these models would be more general. One of the problems right now is that the models are not good enough at this point in terms of grounding like clicking on the right things because there's lots of things to do and sometimes they fail on planning. So even

**中文**  系统能控制 computer、website 或其他任何 visual interface，它就会更通用，因为人类天生依赖视觉。目前的问题是模型 grounding 还不够好，无法总是点击正确的东西；可操作项很多，planning 也有时失败。因此即便

### [00:40:11–00:40:40]

**EN**  frontier models today, you know, on the other hand, they're getting to the point where you're giving you giving them very like for example, one one example that my students always sort of ask is that I'm graduating with my PhD like I want to do post dog or I want to apply for faculty positions. Okay, which schools are looking for faculty positions? What are the deadlines? Are they looking in my area? What what uh information I should be providing? And

**中文**  今天的 frontier model 仍有问题。不过它们正在接近可用。例如，我的学生常问：“我博士毕业了，想做 postdoc 或申请 faculty position。哪些学校在招聘？截止日期是什么？招聘方向与我相符吗？我要提供什么材料？”

### [00:40:37–00:41:05]

**EN**  so, you know, it's not a really hard task, but PhD students spend like a few days, right, trying to search like, okay, Stanford looking for people. Oh, look at this. This this website says yes, they're going to be hiring people. What areas? Well, they hiring in robotics. No, they're not hiring in machine learning, so maybe it's not for me. And sort of, you know, imagine you sort of and so we we we've built these tasks where we ask the model to, you know, say, okay, I'm trying to look for

**中文**  这任务并不特别难，但博士生会花几天搜索：Stanford 是否招聘？这个网页说会招，具体领域是什么？他们招 robotics，却不招 machine learning，那也许不适合我。于是我们构建这类任务，要求模型：“我要找

### [00:41:02–00:41:30]

**EN**  a faculty position, right? give me the list like go to the top 25 schools in the US try to find whom are they hiring what are the deadlines make me excel spreadsheet with all this information and you give it precise kind of like what it needs to collect and also so that it doesn't hallucinate actually go on the website verify that it's correct and keep it open as a tab for me so that if I'm interested I can go and actually look at it myself right so it doesn't

**中文**  faculty position。请查看美国排名前 25 的学校，找出它们招什么人、截止日期，并把信息做成 Excel spreadsheet。”我们会精确说明要收集什么。为防止 hallucination，还要求它访问网站验证准确性，并保留打开的 tab，让我感兴趣时能亲自查看。这样它就不会

### [00:41:28–00:41:58]

**EN**  cuz sometimes the models would just say like yeah you know I did it >> I did it and they're Where is this? Like I don't know where this is. >> Sorry I didn't actually do it. [laughter] >> Yes. Exactly. Exactly. So So these are computer use because it actually requires you to go on the web and like open the thing and you know like so that I can actually see not just like hey Stanford says that they're looking for people and this this like but you actually go open the the link yourself look at the link and keep the tab for me

**中文**  只是说：“对，我做完了。”主持人：做完了？东西在哪里？模型：我不知道，对不起，其实没做。嘉宾：没错。这就是 computer use，因为它必须真正上网、打开内容，让我看见，而不只是说“Stanford 在招人”。它必须自己打开链接、查看，再为我保留 tab，

### [00:41:57–00:42:26]

**EN**  so that I can you know verify it myself. So it requires to actually engage with your computer and the way that we've done it actually we've we we've asked bunch of people to provide us with their his history. We paid the uh people to to provide like what what are they doing on computers? What are they using them for? And so this was kind of like inspiration for us to build some of some of these environments, some of these data sets. It's still early on because we're still trying to understand what

**中文**  让我可以亲自验证，也就是必须真正操作 computer。我们的做法是付费请许多人提供使用历史，告诉我们他们在电脑上做什么、用电脑干什么。这启发我们构建一些 environment 和 dataset。现在仍很早，因为我们还在理解

### [00:42:24–00:42:52]

**EN**  people are using. One sort of success story is the open claw where you know people sort of like figure it out and giving access to people how people want to use it and people want to do a lot of different things. People want to hook up your LLM into your sound system and do stuff and then so it's one of those things that people are using computers for a lot of different things but it's very hard to know exactly what they want to do it and and how much they want to automate. I can give you one example.

**中文**  人们如何使用电脑。一个成功案例是 OpenClaw：它给用户使用权限，让大家自己摸索用法，于是人们会做各种事情，比如把 LLM 接入音响系统。人们用电脑做的事情太多了，很难确切知道他们想做什么、想自动化到什么程度。举个例子。

### [00:42:48–00:43:15]

**EN**  There's wonderful company called Ytori um that does web navigation like agentic systems that kind of like go the web and I think when somebody was telling me that some of the early use cases was when people buy something they want to look for coupons. So it's like I'm buying these shoes. Hey agent go find me coupons. Find me 20% discount. Who's selling these shoes at 20% discount? It takes me an hour and it comes back and

**中文**  有家很棒的公司叫 Yutori，开发 web-navigation agentic system。有人告诉我，一个早期 use case 是购物时找优惠券：我要买鞋，就让 agent 找 coupon、找 20% 折扣、找哪家店打八折。它花一个小时回来

### [00:43:14–00:43:42]

**EN**  says, "Hey, you can buy these shoes over here and with this coupon you get 20% discount. Hey, I'm going to use this product." >> Yeah. >> So, it's like it's not like safety critical thing because maybe there is a discount but if I didn't find it, I'm like, "Okay, fine." You know, I But if it does, >> I was going to buy it anyway. Yeah. >> Yeah. I was going to buy it anyway. Like, so these things you kind of like you start seeing where people But it's still early. It's still early, not robust enough where you can sort of just delegate the task and say like, "Go do

**中文**  说：“在这里买，使用这个 coupon 可省 20%。”那我就会使用这个产品。这不是 safety-critical 的任务；也许有折扣，但没找到也没关系，因为我本来就要买。你开始看见人们会怎样使用它，但现在仍很早，还不够 robust，无法把任务交出去就说：“去

### [00:43:41–00:44:11]

**EN**  it." Then >> like one use case that I was working on with computer use agents was we were working with this company where they have like an online marketplace of courses and what they wanted was they wanted an agent that would they have like 8,000 courses and they wanted an agent where you could give it a list of the course links and then the agent would literally like go and take the course. So it would watch the watch the tutorial videos, it would answer the quiz questions, it would download the homework assignments and then the reason why they wanted it was number one they

**中文**  做吧。”主持人：我曾参与一个 computer-use agent 项目。客户经营在线课程 marketplace，有 8,000 门课程，希望 agent 接收一组链接后，真的进去“上课”：看教学视频、回答 quiz、下载 homework。第一个原因是他们想下载 asset，

### [00:44:08–00:44:37]

**EN**  wanted to basically download the assets. So then they could standardize the course outlines for all 8,000 courses because they were just made by random people. So it wasn't standardized. So that was number one. And then number two was they wanted to take screenshots of the course so that they could generate an accessibility report that hey like for people who are for people who are deaf like you need to have captions on the videos. So you know does it pass that rule? Does it not pass this rule? >> Yeah. Noet etc.

**中文**  从而把随机创作者制作、格式不统一的 8,000 门课，统一整理课程大纲。第二，他们想给课程截图，生成 accessibility report：例如面向 deaf 用户，视频必须有 caption，那么课程是否符合这条规则？等等。

### [00:44:35–00:45:03]

**EN**  >> Excellent. And we and we tried um we tried using like the Google computer use model but maybe we just didn't spend enough time on it but it just didn't seem like it was like good enough and so >> yeah I think >> basically the >> Go ahead. >> Yeah I think that existing models are not at the point where you can sort of like reliably rely on them to go and do these kinds of tasks. Yeah, I think that these are kind of like difficult tasks, right? Because that it it it it requires

**中文**  嘉宾：很好。主持人：我们试过 Google 的 computer-use model，也许只是投入时间不够，但它似乎不够好。嘉宾：我认为现有模型还无法可靠完成这类任务。它们是困难任务，因为既要求

### [00:45:01–00:45:28]

**EN**  you to understand both the textual input like HTML or whatever the behind, but it also requires you to have very good grounding, right? Like understanding what's on the website, what are being displayed and you know, so it's sort of these multimodal models are still sort of >> I think they'll get better. They'll get better, but there's a bunch of startups that are also trying to build better multimodel models. But yeah, >> like what what do you like what makes them better? Is it not just more data or you have to have like a different

**中文**  理解 HTML 等文字输入和底层信息，也要求非常好的 grounding：理解网页显示了什么。这些 multimodal model 仍然……我认为会变好，也有很多创业公司在构建更好的 multimodal model。主持人：怎样才能变好？只是增加数据，还是需要不同

### [00:45:27–00:45:56]

**EN**  architecture or like what is it? >> It's a good question. I think since 2002 when we we're building these models to where we are right now it's night and day, right? So Frontier models now actually you know GPTs and and Gemini and and Claude uh the latest models are actually pretty decent. Decent in terms of like the performance has improved dramatically. It's still not at the point where you can trust them completely to just go do the task. like you know I can I can ask you like you

**中文**  架构？嘉宾：从 2022 年我们构建这些模型到今天，已经天差地别。现在的 frontier model，如最新 GPT、Gemini 和 Claude，已经相当不错，性能大幅提高。但仍未达到可以完全信任、放手完成任务的程度。比如我把任务交给一名

### [00:45:55–00:46:24]

**EN**  know like here's the task go do it and I know you know a capable person will go and do it like a good sort of engineer will go and do it you can't unfortunately do it with with existing models >> right so you kind of like have to you know and so people are building hardness I think it's going to take a little bit of time but I I think that's an interesting interesting research from research perspective we're trying to understand so for example we also did this analysis of the frontier models where they fail a lot of times for example

**中文**  有能力的人或优秀工程师，我知道他能完成；现有模型不行，仍要有人把控。所以人们在构建 harness，这还需要一些时间。但从研究角度很有趣，我们在试图理解原因。例如我们分析 frontier model 的失败方式，很多时候

### [00:46:22–00:46:50]

**EN**  You know, GBT does very good planning of what needs to do, how it needs to do, and then it just doesn't do it. Or like, okay, it's like done, I can't do it. OP, for example, does a lot of research and is doing, but then sort of forgets to replan and such, gets lost and then like kind of especially for these long kind of like task where you have to do a lot of different things, it sort of just forgets or doesn't doesn't plan. So like

**中文**  GPT 很会规划需要做什么、怎样做，之后却根本不执行，或者直接说“做完了，我做不到”。Opus 会做大量研究，也会实际执行，但之后忘了重新规划、迷失方向。尤其在需要完成许多不同事项的 long task 中，它会忘记或不再规划。因此

### [00:46:48–00:47:16]

**EN**  again we have these benchmark odices and the best models hit like 60% or maybe you know 65%. And you need to have like you know above 99.9% for them to be like for me to be reliable to say like hey go do it and knowing that they'll go and do it. So yeah there's still some also another interesting work on like computer use agents for controlling your phone >> you know like if if I give certain commands and cross the apps and like

**中文**  在我们的 Odices benchmark（英文自动字幕拼写）上，最佳模型大约只有 60% 或 65%。要让我可靠地说“去做吧”并确信它会完成，必须超过 99.9%。此外还有 computer-use agent 控制手机的研究，例如我下达跨 app 指令。

### [00:47:14–00:47:44]

**EN**  there's also notion of personalization. So if the agent knows a lot about myself, you know, my apps are cross-connected. Again, this is something that we're trying to build to test how well existing models are doing me. We have something called iOS benchmark, uh, which is sort of mimics iPhone and apps on the iPhone. And it's this notion that if it's my phone, you know, when I order Uber, right, I can go on Uber map and see when my food's going to arrive. Then I can go to my banking app and see like,

**中文**  还涉及 personalization：如果 agent 很了解我、知道各 app 相互连接。我们正在构建测试来衡量现有模型，有一个叫 iOS benchmark 的项目，模拟 iPhone 及其 app。比如在我的手机上，我用 Uber 下单后，可以在 Uber map 看外卖何时到，再去 banking app 查看

### [00:47:43–00:48:12]

**EN**  hey, did they charge me the right amount or not, right? What's different about building comput agents for your phone versus just a normal desktop? >> I think that they're very similar. The set of actions is different like in one case you swiping and such. The visually they look different but they pretty much the same systems. I think that sometimes also like a lot of existing benchmarks and computer use are kind of like these generic things like go to the cat system, look at the video game, you know, control this and such for the

**中文**  收费是否正确。主持人：为手机和普通 desktop 构建 computer agent 有什么不同？嘉宾：其实非常相似，只是动作集合不同，比如手机要 swipe，视觉外观也不同，但本质是同一类系统。很多现有 computer-use benchmark 是通用任务，例如进入某系统、看 video game、控制某物；而在

### [00:48:11–00:48:39]

**EN**  phone and maybe for your personal computer. There is this notion of again I think the notion of personalization, right? My phone, you know, I have Slack, I have my email, I have my banking, I have my flights, they all kind of like interconnected, right? >> Yeah. Is this this good example I think I really like this example that Apple gave I think a year and a half ago when they were talking about super int talking about intelligence like Apple intelligence which really resonated with me which is they basically said hey

**中文**  手机或 personal computer 上，还有 personalization：我的手机有 Slack、email、banking、flight，它们彼此相连。主持人：Apple 大约一年半前介绍 Apple Intelligence 时举过一个我很喜欢、也很有共鸣的例子。他们说：

### [00:48:36–00:49:04]

**EN**  simple question I'm going to ask my I'm going to ask Siri you know when should I go and pick up my mom from from the airport what you know San Francisco airport today a smart system would basically go through your emails and find the flight that your mother is on let's say it's whatever the flight is look at this flight Go to the airport, see if the flight is on time or delayed, right? If it's delayed by 1 hour, estimate when the flight's going to

**中文**  我问 Siri：“今天该几点去 San Francisco 机场接妈妈？”聪明的系统会搜索 email，找到她乘坐的航班；查询机场，看航班是否准点或延误；如果晚一小时，就估算降落时间；

### [00:49:02–00:49:31]

**EN**  land. Look at where you are. Open the Google or the Apple uh Google Apple Maps and look at like how long is it going to take you from the airport to to from your current place to the airport. Usually, if there's a lot of traffic, estimate what the time should be and then basically give a simple answer. Okay, you should leave your home by 2:15 and then sort of say, well, the flight isn't thing, but it's delayed by 1 hour. Also, the kind of traffic conditions are like this. We expect this to be to be worse in an hour because that's the traffic patterns and you kind of like

**中文**  再查看你的位置，打开 Google Maps 或 Apple Maps，计算从当前位置到机场要多久，结合拥堵和未来一小时交通模式，最后给出简单回答：“你应该在 2:15 离家。航班晚点一小时，而且交通状况如此，预计一小时后会更糟，因为你走的是这条路。”

### [00:49:29–00:49:59]

**EN**  see because that's the path that you're traveling. And so now an intelligent system, you know, a dumb system would basically say like, oh yeah, flight lands at like 12:00. That's what's in the email. So you figure out when you should get there, right? and intelligent and and you kind of like don't even think about this but that sort of requires you to you know like open your email figure out when the flight once you know what the flight you have to go to the airport's website see is the flight delayed right because they kind of like listing or is it on time is it

**中文**  所以一个智能系统……笨系统只会说：“邮件里写着航班 12 点落地，你自己算几点到。”你平时不会意识到，但完成任务需要打开 email、找到航班，再访问机场网站，确认航班究竟是否延误或准点——航班经常晚点——还要

### [00:49:58–00:50:27]

**EN**  really on time because a lot of times flights get delayed go open the map see what the traffic conditions are like so this is sort of gives it like a very interesting sort of you know basic intelligence right like >> yeah Interesting. What what were some surprising things that you came across when you were at Meta about training models at that level of scale? One interesting thing that I came out of,

**中文**  打开地图查看交通。这体现了很有意思的基础智能。主持人：有意思。你在 Meta 时，大规模训练模型有哪些出乎意料的发现？嘉宾：我得到的一个有趣认识是，

### [00:50:26–00:50:55]

**EN**  you know, and it's true of all the Frontier Labs is that the architectural choices are very similar like similar to the original transform architecture. There's a few nuances and everything, but the main sort of effort is engineering and data. infrastructure, engineering, and data. It's kind of like, you know, when GPT5, when GPT3 or GPT4 came out, a lot of people were basically saying, my god, OpenAI has has like figured out a new architecture.

**中文**  所有 Frontier Lab 的架构选择都很相似，仍类似原始 transformer architecture。细节会有差别，但主要投入其实是 engineering 和 data，也就是 infrastructure、engineering、data。GPT-3 或 GPT-4 出现时，很多人基本在说：天啊，OpenAI 找到了一种新架构。

### [00:50:53–00:51:22]

**EN**  They figured out like AGI. They figured out like a new sort of thing that's sort of that nobody else knows about. That's absolutely not true. The architectures are the same. These are transformer based architectures. Everything comes down to the data, the quality of the data, engineering, infrastructure. you know how do you do reinforcement one I think that you know the breakthrough like in reasoning systems was was a good one but when you look at these architecture and such it's like it's fairly standard so you know and obviously like there is a lot of work on

**中文**  他们攻克了 AGI，找到了别人不知道的新东西。事实绝非如此，架构都一样，都是 transformer-based architecture。一切归结为数据、数据质量、engineering 和 infrastructure，以及怎样做 reinforcement learning。我认为 reasoning system 的突破很好，但观察这些架构时，它们相当标准。当然还有大量

### [00:51:20–00:51:48]

**EN**  self-improving systems these all fascinating system but a lot of them are sort of I guess what I'm trying to say is that it's fascinating what's been happening but the real drive has been more on the infrastructure the data >> the way that code is actually working it's fascinating because it's almost like a game in the sense that I give you the task, I have my code base, I have my unit test and if the model is able to

**中文**  self-improving system 的研究，都很迷人。我想说的是：正在发生的事情确实精彩，但真正驱动力更多来自 infrastructure 和 data。code 的工作方式也很有意思，它几乎像游戏：我给你任务、codebase 和 unit test，如果模型能够

### [00:51:47–00:52:16]

**EN**  pass all the unit tests, I give it the reward of one and I use reinforcement learning to reinforce whatever I did is the correct thing to do. And then what you can do is you can basically say, well, what I'm going to do is I'm going to introduce subtle bugs manually myself through combination and such. And then that's another training example. Now my agent has to go try to find fix it and if it fixes I have unit tests and that's like it's almost like a game where I I run the thing I get the score and I

**中文**  通过所有 unit test，就获得 1 的 reward；再用 reinforcement learning 强化刚才的正确做法。之后我可以通过组合等方式，人工引入细微 bug，作为另一条训练样本。agent 必须找到并修复它；如果修好，unit test 就会通过。这像一场游戏：运行后得到分数，而且我

### [00:52:14–00:52:43]

**EN**  always know whether I fixed it or not >> and so it's not like with web agents and sort of like computer use sometimes I can solve the same problem many different ways and sometimes it's like partial solutions are also okay but in the systems like math and and coding is that's why we saw we seeing a lot of progress is that these are kind of what we call verifiable domains where I quit a task there is one like I know where I got the answer correctly. So >> um

**中文**  始终知道是否修好了。web agent 和 computer use 不同：同一问题可能有许多解法，partial solution 有时也可以；而 math 与 coding 是所谓 verifiable domain，进步快正因为我提出任务后，能明确知道答案是否正确。

### [00:52:41–00:53:10]

**EN**  >> do you think do you think we're at a point where that is plateauing where just the amount of data is not actually improving the outcomes as much as they were two years ago? So but what's interesting right now is that there is sort of you know the continuation of this piece where you sort of try to generate synthetic data right so in coding you try to introduce bugs artificial bugs and get the system to fix those bugs and you kind of like continuously self-improve uh the system I think that the self-improving aspect

**中文**  主持人：我们是否到了 plateau？增加数据带来的 outcome 改善，是否已不如两年前？嘉宾：现在有个有趣的延续方向，就是尝试生成 synthetic data。例如在 coding 中人为引入 bug，让系统修复，从而让系统持续 self-improve。我认为 self-improving

### [00:53:09–00:53:39]

**EN**  of it is is an interesting one because you can kind of introduce like disturbances in the system and then try to fix it. There's a lot of work also on synthetic data generation. A lot of times you know you have multiple models they all generate data and you kind of like clean it up. You pass it through the filters and say now use that data for training. So there's kind of like many many nuances. I think I don't know whether we've reached completely the the scaling law kind of like saturation where we basically say we've done because I've been wrong before you know

**中文**  很有意思，因为可以给系统引入 disturbance，再让它修复。synthetic data generation 也有大量工作：让多个模型生成数据，清洗并通过 filter，再用于训练，有很多细节。我不知道 scaling law 是否已彻底饱和，能否说已经到头，因为我以前也判断错过。

### [00:53:37–00:54:06]

**EN**  like where people would say well scale is not going to get you far but scale is getting you very far right like all the innovations that we had so far is due to scale of the data right and these sort of like emerging behaviors and of course you need to sort of make sure that and then you have if you have strong models you can sort of generate synthetic data from those models filter them and that now becomes a cleaner training data for your model. So I I don't know. I don't know. I do

**中文**  比如有人会说 scale 走不了多远，但实际上 scale 已经把我们带得非常远。迄今为止的所有创新都来自 data 的 scale，以及由此出现的 emergent behavior。当然还要保证其他条件；如果有强大的模型，就可以用它们生成 synthetic data，再进行过滤，得到更干净的训练数据。我不知道最终答案，但

### [00:54:03–00:54:32]

**EN**  see the appeal of multimodel data where you go beyond language. You go into the physical world. You have this is where robotics applications are now happening embodied AI where in addition to language you have videos. You have world models. You have like these other rich sources of data. I have um one of my former students who's a professor at MIT right now Paul Leang was working on multi-ensory foundational models. So for example, he's now working on building models for smell, right?

**中文**  我确实看到了 multimodal data 的吸引力：超越语言，进入 physical world。这正是 robotics application 和 embodied AI 正在发生的地方。除了语言，还可以使用 video、world model 和其他丰富的数据源。我的一位 former student、现在任 MIT 教授的 Paul Liang，研究 multisensory foundation model。例如，他现在正在构建嗅觉模型。

### [00:54:31–00:55:00]

**EN**  >> Interesting. >> Can I like if if I have my senses smelling, can I understand what what am I smelling? >> Right? And you know like touch senses if I'm touching something, can I identify like without vision like am I touching specific type of so you you you can imagine there's like many many more different areas. It's just the industry right now is focused on images and text videos as well because for companies

**中文**  主持人：有意思。嘉宾：如果我拥有嗅觉感知，能不能理解自己闻到了什么？touch sense 也一样：触摸某个东西时，不借助 vision 能否辨别自己摸的是哪一类物体？可以想象，还有许许多多不同领域。只不过 industry 现在集中于 image、text 和 video，因为对公司来说

### [00:54:57–00:55:27]

**EN**  like uh you know Facebook generating images like one of the success stories of Gemini is not a banana apparently like it's >> that killed our business when we first started. >> Yeah. That's like >> we were doing like virtual tryons for like for like retailers and then Nano Banana comes out and you can basically do that with one API call. >> Yes. Exactly. And it's like it just became very popular because like you know I have a friend and he's like my

**中文**  ——例如 Facebook——生成图像很重要。Gemini 的成功案例之一显然是 Nano Banana。主持人：它刚出现时把我们的业务毁了。嘉宾：是啊。主持人：我们原本为 retailer 做 virtual try-on，Nano Banana 出来后，一个 API call 就能完成。嘉宾：没错，它特别流行。我有个朋友，他说自己的

### [00:55:24–00:55:54]

**EN**  dad is in his 70s and he's like using nada banana for like put himself into some weird kind of like you know situations and then sends it to to his family to his friends and it's like oh everybody's like using it right but there's obviously like much more beyond just images right there is vision videos there is sort of understanding the physical world and this is where robotics will have to face this right I mean self-driving cars is kind you know one one aspect of it but of course for

**中文**  父亲已经七十多岁，还会用 Nano Banana 把自己放进各种奇怪场景，再发给家人朋友，大家都在用。但显然，图像之外还有更多领域，包括 vision、video 和理解 physical world，robotics 必须面对这些问题。self-driving car 是其中一方面，不过对

### [00:55:52–00:56:22]

**EN**  for businesses I mean code cloud is amazing product like I mean I'm using it all my students are using it anytime I have a bug somewhere just give it to code claude and tells me like oh you have you know you don't have a the latest version of Python like do you want me to install it for you I'm like sure go ahead and it's like pipe install whatever the commands that I would basically spend lots of time doing it just does it for you right so it's >> what are your thoughts on what are your thoughts on the future of software engineering because Like most of the

**中文**  企业而言，Claude Code 是很棒的产品，我和所有学生都在用。任何地方出了 bug，交给 Claude Code，它会说：“你没有最新 Python 版本，需要我安装吗？”我说好，它就执行 pip install 等命令，把原本要花我很多时间的事情做完。主持人：你怎么看 software engineering 的未来？因为我认识的多数

### [00:56:20–00:56:50]

**EN**  engineers that I know, they're basically saying what you're saying where I haven't written a line of code in months is what I keep hearing. >> That's true. But at the same time, I think that you still need to have good understanding of software engineering and how things work because if you actually have to go and debug it, you have to be good at it. But of course, like you need to understand how to use AI tools as well. Uh so I think in the future it's going to be some kind of hybrid. I would not recommend people to

**中文**  工程师都说自己几个月没写过一行 code。嘉宾：这是真的。但仍需要很好地理解 software engineering 及其工作原理，因为一旦真要 debug，你必须有能力；当然，也要懂得使用 AI tool。未来会是某种 hybrid。我不建议人们

### [00:56:47–00:57:17]

**EN**  sort of like say well don't do coding or completely abandoned coding. I mean there's there's an argument for that. Like when I was doing my undergrad we were studying assembly. Hey, this instruction inject this beat over here and then this beat over here and like I don't know like nobody's doing assembly, right? Like it's fully kind of, you know, but at the same time if I had to I could go and like figure stuff out if I needed to, right? I I think that pure software engineering is probably not going to be the thing like where you

**中文**  完全不学或放弃 coding。当然有人这么主张。我读本科时学过 assembly，研究一条 instruction 怎样注入这个 bit 和那个 bit；现在没人写 assembly，几乎完全被抽象掉了。但如有必要，我仍能重新弄懂它。我认为单纯的 software engineering——只做一个 coder——可能不会继续成为主流。

### [00:57:14–00:57:44]

**EN**  just like just a coder. I think people with strong machine learning statistics kind of expertise plus strong coding expertise is is what's going to be in demand. >> So basically like research roles is basically what's left. research or research engineers like you know like engineering pure engineering like I mean yes obviously you have amazing people who can optimize and everything but people who understand how these models

**中文**  兼具 machine learning、statistics 专业能力和强 coding 能力的人，才会是市场所需。主持人：所以剩下的基本是 research role？嘉宾：researcher 或 research engineer。纯 engineering 当然也有能够做优化的优秀人才，但既理解模型

### [00:57:42–00:58:10]

**EN**  work with strong coding skills I think are going to are still going to continue uh will continue being in in in in in demand >> interesting right >> so you think like >> would you say like most software engineering roles are not that right now and so do you do you believe that those roles are going to go away >> that's That's a tough question. That's a tough question. I think that you know some of these roles will stay. But I think that what what I see happening right now with a lot of businesses is

**中文**  工作原理，又有强 coding skill 的人仍会持续有需求。主持人：有意思。你是否认为现在多数 software engineering role 并非如此，所以会消失？嘉宾：这是个难题。我认为部分岗位会保留，但现在许多企业正在发生的情况是，

### [00:58:08–00:58:38]

**EN**  that instead of having like 10, you know, engine 10, you know, computer programmers, I only need five or four or three right now. Of course, like I mean it's it's uh it sort of makes the whole thing more efficient for sure. It also makes kind of like some of the research that I'm seeing like faster because I can give you an example. One of my students is you know like they running experiments and they found like coding agents to be extremely useful like if something happens overnight and you know

**中文**  过去需要十名 computer programmer，现在只要五个、四个或三个。它确实让一切更高效，也让一些 research 更快。举例说，我的一名学生在运行 experiment，发现 coding agent 极有用：如果实验夜间失败，agent 可以分析 error，判断 memory 配置不当，自动增加 memory 并重新运行。

### [00:58:36–00:59:04]

**EN**  it fails the agent can go analyze the error and say you know the memory wasn't set up properly let me increase the memory and start running it again. Yeah, >> is extremely useful because you know you you've just wasted you know six seven hours of GPU use if if it just sits there right and so >> I actually just did that like last night >> that's literally that exactly and so it now becomes extremely and and so I think that yeah I it's a good question I don't know there's a lot of software engineers

**中文**  这极其有用，否则任务停在那里，就浪费六七小时 GPU 时间。主持人：我昨晚刚做了这件事。嘉宾：正是如此。因此它非常有用。至于大量 software engineer 会怎样，这是个好问题，我不知道。市场仍在到处寻找优秀 software engineer。

### [00:59:03–00:59:31]

**EN**  people are still looking for strong software engineers like like everywhere even in in my company right now there's always demand for like you know we've hired a person who's like amazing engineer but also very strong machine learning person. So those people are always going to be in demand because you know it's not just like we we say like hey go fix this thing and the person goes fixes this. The person actually fixes and then comes back and says I think you're doing this wrong like we should do it differently right those people are going to be in always

**中文**  即便在我的公司，我们也总需要人才。我们聘请过一位很出色的 engineer，同时也很懂 machine learning。这样的人永远有需求，因为你不会只是说“去修这个东西”，他修完就回来；他会修好后告诉你：“我认为你的做法不对，应该换一种方式。”这类人永远

### [00:59:30–00:59:58]

**EN**  high demand. >> Yeah. Yeah. People with the judgment essentially >> people with judgments people with understanding of of machine learning AI but also strong coders. >> Yeah. >> Right. That's Yeah. So maybe that leads us to like you starting Soothe Lab. So could you explain like >> Yeah. What was the genesis of Sooth? >> Yeah. So >> So you left Meta? >> Yes. I left Metam and I started the company with a couple of co-founders.

**中文**  会有很高需求。主持人：本质上是有 judgment 的人。嘉宾：有 judgment、理解 machine learning 和 AI，同时也是强 coder 的人。主持人：这也许正好引到你创办 Soothe Labs。它是怎么来的？你离开了 Meta？嘉宾：对。我离开 Meta，与几位 co-founder 一起创办公司。

### [00:59:56–01:00:24]

**EN**  One of them is Yaser Shik. He was a former professor at CMU. I've known him for a long time. He also built he was at Meta for 10 years building some of the Oculus and some of the predictive models like how people behave and such. And so the premise for SU Labs is we, you know, we want to be able to build a foresight system, a system that can forecast future events or at least

**中文**  其中一位是 Yaser Sheikh，他曾任 CMU 教授，我认识他很久了。他也曾在 Meta 工作十年，开发 Oculus 的部分技术以及预测人类行为的模型。Soothe Labs 的前提是，我们想构建 foresight system，也就是能 forecast 未来事件，或至少

### [01:00:21–01:00:51]

**EN**  estimates, you know, create what we call calibrated probabilities for future events. Um, so you can imagine that the system that basically takes all the available information today. You know this would include some structured data which is like time series like you know new reports earnings reports or maybe like unemployment rates right like things of that and then combining it with with unstructured data news things that are happening in the environment

**中文**  为未来事件给出所谓 calibrated probability 的系统。可以想象，系统接收今天所有可用信息，其中包括 structured data，例如 time series、新闻报告、earnings report、unemployment rate，也把它与 unstructured data、新闻及环境中发生的事情结合起来。

### [01:00:49–01:01:18]

**EN**  right and whatever that information is you can fuse it together to try to forecast certain probabilities of certain events that might be coming up I can give you an example perhaps a toy example >> yeah so I have a student hopefully he's going to be okay you know me giving this example he's a great student he's like one of my best student his his name is Martin Mh he's working on you know multimodel aspects and I'm co-dising him with another professor but he's working on multimodal aspects and such and so he's graduating in December so I was

**中文**  无论有哪些信息，都可以融合起来，forecast 某些未来事件的 probability。我可以举个偏 toy 的例子。我有一位学生，希望他不介意我讲。他很优秀，是我最好的学生之一，叫 Martin Ma，研究 multimodal 方向，我和另一位教授共同指导他。他将在 12 月毕业，所以我

### [01:01:16–01:01:45]

**EN**  talking to him what he wants to do next and I said like hey look we're building this sort of agendic systems for forecast let's ask the following question will Martin Mah s you know su student graduate by December 2016 so the system goes finds him and says like will Martin Mah PhD PhD CMU student a PhD student at CMU who's been there for 4 years graduate by December 2026. So it formulates the question creates a contract much like what you would do in poly market or kali like this is a thing

**中文**  问他之后想做什么，并提议用我们为 forecast 构建的 agentic system 问：“Martin Ma 是否会在 2026 年 12 月前毕业？”系统找到他后，把问题表述为：“在 CMU 已读四年的博士生 Martin Ma，会否于 2026 年 12 月前毕业？”它像 Polymarket 或 Kalshi 一样建立 contract，明确

### [01:01:43–01:02:12]

**EN**  this is how we're going to resolve it that's the question so the model goes through reasoning it thinks about you know what's the probability of this event occurring comes up with the answer of 13% 13%. Basically saying no way it's going to graduate by December. So it was like hm that's surprising why is it so low. So you go to the system. The system provides reasoning like what's what it found and what's what's the reasoning behind it. And it basically says well Martin hasn't really hasn't proposed yet

**中文**  问题和结算方式。模型开始 reasoning，估算事件发生 probability，得出 13%，基本就是说他绝不可能在 12 月前毕业。我们很意外，追问为什么这么低。系统给出它找到的事实和 reasoning：Martin 还没有做 thesis proposal，

### [01:02:10–01:02:40]

**EN**  which is true. He hasn't done his thesis proposal yet. He was going to do his thesis proposal a couple of weeks after and the system man said okay that's fine. You know a lot of students you know propose later and then it said well and there is a new uh requirement that was introduced in the department. It found a document through sort of research that there is a rule right now at CMU that the time between your proposed date and your defense date should be at least 12 months.

**中文**  这确实是真的。他原本准备几周后做 proposal。系统认为晚一点 proposal 本来也没关系，但发现系里新增了一项要求：通过研究找到一份文件，说明 CMU 现在规定 proposal date 与 defense date 至少相隔 12 个月。

### [01:02:39–01:03:07]

**EN**  >> Starts doing reasoning and basically saying well you haven't proposed it's May you try to propose you try to defend in June. There's a rule that says like the time should be at least 12 months. That means that your probability of defending goes down significantly. So I'm asking Martin, did you know about this rule? that's something new that they've done. And the reason why CMU's done it is that a lot of times students would like propose and in three months defend that wasn't the intention. The intention for students to propose like

**中文**  它于是 reasoning：你还没 proposal，现在是 5 月，却想在 6 月 defense；规则要求相隔至少 12 个月，所以如期 defense 的 probability 大幅下降。我问 Martin 是否知道这项新规。CMU 之所以这么做，是因为很多学生 proposal 后三个月就 defense，但原本意图并非如此，而是希望学生 proposal 时说明

### [01:03:06–01:03:35]

**EN**  you know what they've done, what they going to do and the committee like looks at them and says like okay this is what you're going to do in the next year and for your defense and such. was designed specifically to avoid situations where people just sort of you know try to you know and he said yes I know about this rule and then the model would basically generate what we what what we call like trigger events or what we call like decision points and one of the decision points was that Martin gets a waiver so

**中文**  已经完成什么、接下来准备做什么，committee 再评估他未来一年到 defense 期间要完成的工作。新规专门用来避免匆忙走流程。Martin 说自己知道。随后模型会生成我们所谓的 trigger event 或 decision point，其中一个 decision point 就是 Martin 获得 waiver。

### [01:03:33–01:04:01]

**EN**  in the next month there's some probability that you will get a waiver condition on this event that you get a waiver your probability of defense goes up from 13% to something like 80% >> mm So I'm telling Martin like, "Have you thought about getting your waiver like for this thing?" He's like, "Yeah, I should probably get a waiver." And I'm like, "Yeah, you probably should. If you want to graduate, you probably should go get a waiver because otherwise you're not graduating, right?" So this is a

**中文**  下个月获得 waiver 有一定 probability；以获得 waiver 为条件，按时 defense 的 probability 会从 13% 上升到约 80%。我问 Martin：“你有没有考虑申请 waiver？”他说应该申请。我说，如果想毕业，你大概必须拿到 waiver，否则就毕不了业。这个例子

### [01:03:59–01:04:28]

**EN**  little bit of a toy example, but you can imagine that a system that forecast any future from will will anthropic IPO this year, will SpaceX go to Mars by 2030, right? Any sort of question. uh we believe that right now I mean there are forecasters there's something called super forecasters which people have a deep knowledge of certain areas and then they try to estimate the probability of future events of course nobody can be precise of saying whether

**中文**  稍显 toy，但可以想象一个能够 forecast 任何未来问题的系统：“Anthropic 今年会 IPO 吗？”“SpaceX 会在 2030 年前到达 Mars 吗？”目前已有 forecaster，还有所谓 superforecaster；他们对特定领域理解很深，再估算未来事件的 probability。当然没人能准确断言事件

### [01:04:26–01:04:54]

**EN**  events happen or not but you can estimate the probability and calibrated estimation basically means that you know if the system uh says 70% then out of 100 events with 70% we'd expect 70 of them to happen and 30 of them not to happen right that's that's the meaning of the calibrated probabilities and so we believe that you know in this sort of vertical AI should be able to do much better than people um and then I think

**中文**  是否发生，但可以估算 probability。所谓 calibrated estimation，是系统对一百个事件都给出 70% 时，我们预期其中约七十个发生、三十个不发生；这就是 calibrated probability。我们相信，在这个垂直领域，AI 应能远胜于人类。

### [01:04:52–01:05:20]

**EN**  it will open up a lot of these decision points so it's not just the outcome but it's like a lot of decision points like you know certain things maybe certain things you can influence certain things you can't like in Martin's case like he actually went and and and got his waiver so he basically said okay if I'm trying to optimize me graduating like these are things that I have to do right and the system can sort of are find and suggest and it has a lot of implications because I think that a lot of times our world is

**中文**  它还会揭示许多 decision point，不只给 outcome，也指出哪些事能影响、哪些不能。Martin 的案例里，他真的去拿到了 waiver。他会想：“如果我要优化自己如期毕业的结果，这些就是我必须做的事。”系统能发现并提出建议。它影响很大，因为我们的世界

### [01:05:17–01:05:47]

**EN**  like very interconnected. So things you know geopolitics impacts you know culture and culture impacts how we operate and sort of like all these things are very much intertwined and right now with the power of AI how much data can process in real time this I think is going to be the next frontier. Mhm. >> So I went from we went from aentic systems that like super useful to systems that can kind of like look into the future a little bit better given all

**中文**  高度互联：geopolitics 影响 culture，culture 又影响我们的运作方式，一切彼此交织。凭借 AI 实时处理大量 data 的能力，我认为这会成为下一个 frontier。我们从非常实用的 agentic system，走向一种能依据

### [01:05:45–01:06:13]

**EN**  the information we have today and also try to kind of like make it part of the decision- making. So if I am a CEO of a company and I have a better view of the future than my competitor, I'm going to win. >> Yeah. So why can't I just kick off a deep research assignment in Chad GBT to try to predict something? >> You can't. You can. >> Okay. The the caveat is that these models like CH GPT first of all they don't really take structured data into

**中文**  今天所有信息更好地观察未来，并把预测纳入 decision-making 的系统。作为 CEO，如果我对未来的判断优于竞争对手，我就会赢。主持人：为什么我不能直接在 ChatGPT 发起一个 deep research 任务，让它预测某件事？嘉宾：可以。但 caveat 是，ChatGPT 这类模型首先不会真正纳入 structured data，

### [01:06:11–01:06:40]

**EN**  account. They're very bad at processing numeric data. It's number one. >> And these models have been trained to predict the next token, right? They sort of like they're not designed to sort of predict they're not designed for forecasting. >> They are forming like fantastic base models but you still need to postrain such that you can adjust and such. And then there is another thing which is accumulation which makes this problem

**中文**  处理 numeric data 的能力很差，这是第一点。其次，这些模型训练目标是预测 next token，并不是为 forecasting 而设计。它们是非常优秀的 base model，但仍要 post-train，才能把行为调整到这个目标。还有另一个不同之处，就是 accumulation，

### [01:06:38–01:07:06]

**EN**  actually a little different is the context accumulation. Which sources I'm relying on, which sources I should use, which sources I shouldn't use, which ones I trust, which ones I don't trust. And so it it it requires you know obviously Chad GPT can like I can ask it like you know will Martin M graduate in 2027 you know we'll come up with something >> but I don't think it it will do a good job at like giving me the reasons and giving me these sort of like decision

**中文**  也就是 context accumulation：依赖哪些 source、该用哪些、不该用哪些、信任哪些、不信任哪些。你当然可以问 ChatGPT：“Martin Ma 会在 2027 年毕业吗？”它会给出某种回答，但我不认为它能很好地给出理由，以及沿途这些 decision

### [01:07:04–01:07:33]

**EN**  points along the way. >> So as much as you're able to talk about it I guess is this more of a model training problem or is this like a harness problem? >> Both. >> Okay. You need to train your own model because training becomes a critical aspect aspect for actually steering these models not just towards predicting the next token but actually trying to predict future future outcomes right so it's it's it's the training uh but

**中文**  point。主持人：在你能透露的范围内，这是 model-training problem，还是 harness problem？嘉宾：两者都是。需要训练自己的模型，因为 training 对 steer 模型至关重要：不只让它预测 next token，而是让它预测未来 outcome。所以 training 很重要；但

### [01:07:31–01:08:00]

**EN**  harness is also kind of you know obviously you can rely on existing models as well and one sort of interesting aspect of it is that for example Grock has access to Twitter no other models have access to Twitter sometimes is a very good source of the news, right? So, if you're relying on chat GPT alone, I think you're always going to be inferior >> just in terms of the ability to find and retrieve the relevant content. When you're also building these systems, I think for businesses and such, a lot of

**中文**  harness 也重要，当然也可依赖现有模型。有个有趣例子：Grok 能访问 Twitter，其他模型做不到，而 Twitter 有时是很好的新闻来源。因此只依靠 ChatGPT，寻找和检索相关内容的能力始终会更弱。为企业构建这类系统时，

### [01:07:58–01:08:27]

**EN**  times corporations will have their own private data that they accumulating and such. And so, how do you build systems that can also leverage internal data and such? So, >> yeah, >> this is >> so why why did you pick this problem to work on? I think that from high level perspective, I think the next two years is going to be the year of agents, you know, self-improving agents, agents that can go do research, agents that can go do like, you know, scientific

**中文**  企业往往还有自己积累的 private data，系统怎样利用 internal data 也很重要。主持人：你为什么选择这个问题？嘉宾：从宏观上看，未来两年会是 agent 的时代：self-improving agent、能做 research 的 agent、能做 scientific

### [01:08:25–01:08:54]

**EN**  experiments and whatnot. After two years, I think it's going to shift towards decision-m systems. I want to optimize for my future, you know, like what should I do so that in in in two years, I'm going to have this job, right? I want what's I I want you know so it's it's almost like there's this notion of personalization for corporations this notion of trying to anticipate the future. So it's it's more like if you have a system that can predict future like I mean it you can

**中文**  experiment 的 agent 等。两年之后，重心会转向 decision-making system：为了实现我想要的未来，应该做什么？怎样才能在两年后得到这份工作？这里既有个人的 personalization，也有企业对未来的 anticipatation。拥有能预测未来的系统——当然永远无法完美预测——

### [01:08:53–01:09:20]

**EN**  never predict the future like perfectly right but if you can kind of like it's almost like it's like blurry but if I can make it a little bit more precise you know then and take I can make decisions so it's no longer about like hey agent go like find stuff for me or like do stuff it's actually okay what decision I should take so that I'm optimizing for my own future outcomes. >> Yeah. Are there are there certain are

**中文**  就像未来原本很模糊，而它能让画面稍微更清晰，我就能据此做决定。问题不再是“agent，替我查东西或做事情”，而是“我应该做什么决定，才能优化自己的未来 outcome”。主持人：你是否针对某些特定事件类型做 training optimization？

### [01:09:18–01:09:47]

**EN**  there certain types of certain types of events that you are optimizing your training for? Um because I'm assuming what you're doing is you're doing some kind of back testing on events that happened in the past and then you're using >> that's a fantastic point and that's the point that's basically is so hard. This probably makes it so hard because back testing I can go back test and publish bunch of papers and say look my model is better. Nobody's going to trust me. >> Why is that? Because back testing

**中文**  我猜会对过去事件做某种 backtesting，然后用于……嘉宾：这是非常关键的一点，也正是问题最难之处。我可以 backtest，再发表很多论文，说模型表现更好，但没人会相信。主持人：为什么？嘉宾：因为 backtesting

### [01:09:44–01:10:14]

**EN**  there's always leakage of information. I see like like I can give you an example like you know we had a way of like trying to predict some outcomes of the sports game right and we we were like searching through the sources we were way constraining to only look at the sources before July 2025 right we picked up this Wikipedia article it says like you know was standby July 2025 like you know and was one of the sources and we

**中文**  总会发生 information leakage。例如我们曾尝试预测体育比赛 outcome，严格限制系统只能查看 2025 年 7 月以前的 source。它找到一篇 Wikipedia article，上面带着 2025 年 7 月的时间戳，于是被用作 source；然后我们

### [01:10:11–01:10:39]

**EN**  could perfectly predict yeah we could perfectly predict the the outcome of the And we're like this fantastic our model works. But what happened is that somebody went modified this Wikipedia article with the outcome of the game and a stamp was like still so so the the the model that was the data that was stand by July 2025 contain the data about 2026 >> right so then it's like okay yeah I can

**中文**  完美预测了比赛结果，以为模型太厉害了。但实际情况是，有人在赛后修改了 Wikipedia article，网页时间戳仍保留原样。于是标成 2025 年 7 月的数据里，其实包含了 2026 年的结果。所以，我当然可以

### [01:10:38–01:11:06]

**EN**  I can tell you in the back test I'm doing like fantastic it means nothing like the only way and this what makes kind of like these prediction systems interesting because you cannot you cannot not there's no way to contaminate your data there's no way to overfeit the data the only way if I make a prediction I have to wait and see like did it come out or not and this is where it becomes an interesting problem because you know like at meta I remember there was an example where the model would basically generate an answer to a particular

**中文**  声称 backtest 表现绝佳，但这毫无意义。prediction system 有意思之处就在于，它无法污染未来数据，也无法对未来 overfit。做出 prediction 后，唯一办法就是等待，看它是否成真。这让问题很有意思。我记得在 Meta 有个例子：模型会生成某个

### [01:11:05–01:11:34]

**EN**  coding question and it would generate the comment section as well and supposed to be the test point like where we test the model and the model would generate comments which were too precise to be true like the comments would be like hey John like You know, I think we should fix this part, right? Okay. You look at the test example. Test example has the same comment. Hey, John, I think we should fix this part. Right? So, what happened is that somehow this example actually ended up somewhere on the web.

**中文**  coding question 的答案，连 comment section 都一并生成。那本应是一条 test point，但生成的 comment 精确得不真实，例如：“John，我认为应该修复这一部分。”查看 test example，里面竟有完全相同的 comment。原来这个 example 不知怎么出现在网上。

### [01:11:32–01:12:02]

**EN**  You know, the you know, people scraped it, trained on it, and now we're testing. So, we're testing on the training basically, right? So, >> this is known as the contamination problem. It's a big problem. It's not just it's all different from TLS because you know there was this example of TIG unicorns where it's like people can say hey look like if I prompt my model it generates me these TIG unicorns like and it's like it's truly generalizable until somebody found that there is like entire web page dedicated to this problem. it

**中文**  人们抓取它、用它训练，现在又拿它测试，所以本质上是在 training data 上测试。这就是 contamination problem，是个大问题。类似例子不只这一个，比如所谓 TIG unicorn：有人说“看，我给模型 prompt，它能生成这些 TIG unicorn，真的具备 generalization”，直到有人发现网上有一整页专门收集这个问题，

### [01:12:00–01:12:29]

**EN**  has like all the examples and so you kind of you know it's almost like you're testing on the training right which is not but if you're trying to test on the future events that haven't happened there's no way for you to cheat right >> it's just and this is like true generalization right and that's what makes this problem very hard because there is a time component to it any sort of back testing you can do okay that's fine for model development and everything it's okay but to really test it you have to do forward testing and

**中文**  里面有所有 example。于是你几乎又是在 training data 上测试，这当然不对。但如果测试尚未发生的 future event，就没办法作弊，这才是真正的 generalization。它也让问题很难，因为包含 time component。backtesting 可以用于 model development，但真正测试必须做 forward testing。

### [01:12:26–01:12:55]

**EN**  you know like like poly market and ki is the way to sort of you So um >> so so I guess does that mean that you will be building your own hedge fund then using using this? >> I think that hedge fund is sort of fairly narrow application because you know if it works it would have a lot of potential in many many different areas you know from insurance intelligence

**中文**  Polymarket 和 Kalshi 就提供了这样的方式。主持人：所以你会用它来建立自己的 hedge fund 吗？嘉宾：hedge fund 是相当狭窄的 application。如果系统有效，它会在很多领域拥有潜力，从 insurance、intelligence

### [01:12:52–01:13:22]

**EN**  communities to kind of like financial predictions and such. So >> do you have a specific industry in mind as you think this would be best used for? that that's I think that I think that you know there there are multiple industries we haven't sort of looked at exactly you know I think that obviously financial kind of like places is is is a good example I think that intelligence community is a good example and there like bunch of like anytime I think that you know it's it's the problem that if you if you're accurate enough and you

**中文**  community 到 financial prediction。主持人：你认为最适合哪个行业？嘉宾：可能有多个行业，我们尚未精确选择。financial sector 显然是很好的例子，intelligence community 也是。任何情况下，只要模型足够准确，胜过现有模型，application 就

### [01:13:20–01:13:48]

**EN**  can do better than any of the existing models then the applications are >> it's it's pretty obvious what the applications would be right like from financials to intelligence to insurance it's like >> I mean somebody was telling me pricing right Now insurance pricing for like fires in LA there is a poly market for it a culture for it which is much better prediction what insurance companies are doing. >> Yeah >> right. >> Yeah that's the other question I wanted to ask you. So like what do you think

**中文**  显而易见，从 finance 到 intelligence、insurance。有人告诉我，LA 火灾的 insurance pricing 已有 Polymarket 或 Kalshi 市场，它给出的 prediction 比保险公司更好。主持人：这正是我想问的：为什么要用你的 decision system，而不是直接用 Kalshi 或 Polymarket 的 probability？

### [01:13:45–01:14:15]

**EN**  about using this decision system versus just using the outcome of um or like the probabilities on kaly or poly market. >> That's a very good question. What we found is that the crowd prediction is actually pretty good. like Cali Poly Market people are you know one person one one person prediction is probably not good >> but when you have thousand people >> and potentially insiders >> so if you take insiders out so let's assume that there's no insider

**中文**  嘉宾：很好的问题。我们发现 crowd prediction 相当好。在 Kalshi、Polymarket 上，单个人的 prediction 也许不好，但一千人聚在一起就不同了。主持人：而且可能有 insider。嘉宾：先排除 insider information，假设没有内幕。

### [01:14:14–01:14:44]

**EN**  information obviously if somebody has an inside information there's nothing you can do right like if I know certain events going to happen you know then you know I I can't do much about >> but when you have the power of the crowd like many people and many people are betting money then people become efficient right you know it's because when I bet money I actually think I know what the outcome will be or at least I can estimate the so and so the crowd prediction are actually very accurate and it's you know it's one of the best

**中文**  如果有人确实掌握 inside information，那什么系统都没办法。但有 crowd 的力量，许多人真正下注时，市场会变得 efficient。因为押上钱时，我是真的相信自己知道 outcome，或至少能估算 probability。因此 crowd prediction 非常准确，是最好的

### [01:14:43–01:15:12]

**EN**  predictors like if you look at the prediction of like elections or sports events or such prediction markets are actually fairly accurate baseline like very good baseline right because otherwise if it wasn't people would kind of make money out of it they would be arbit chart opportunities. I think that we can do better with AI. We can do better than prediction markets than what people are doing. That's one hypothesis because the access to information and the breadth of the information that you

**中文**  predictor 之一。观察 election 或 sports event，prediction market 是相当准确、很强的 baseline；否则人们会利用错误定价赚钱，出现 arbitrage opportunity。我们的 hypothesis 是 AI 能做得更好，胜过 prediction market，因为它可访问和处理的信息广度

### [01:15:10–01:15:39]

**EN**  can process will be always bigger than any person can do. >> Yeah. >> But at the same time, imagine a system that can be just as good as the crowd or even better for any question. Right. >> Yeah. like and I'm assuming that you can also use the use the poly market information as an input into your system as well any moment in time. >> Yeah, >> that's that's exactly right. But at the same time, imagine you can create sort

**中文**  永远超过个人。但想象一下，如果系统能对任何问题都做到与 crowd 一样好，甚至更好。主持人：也可以把 Polymarket 信息作为系统在任意时刻的一项 input。嘉宾：完全正确。同时，想象你能为任何问题建立某种

### [01:15:37–01:16:07]

**EN**  of, you know, it's almost like you're creating prediction market for any question. Right? Right now, >> prediction markets are sort of centered around what people want to and it's like the most popular categories are I think it's like Bitcoin and sports. >> Yeah. And maybe politics. Yeah. >> Politics less so. Uh but like Bitcoin, people want to know what's happening to the Bitcoin and kind of and people want to know what's happening to the sports. Like those are two things. And then you know these are good people want to

**中文**  prediction market。目前的 prediction market 只围绕人们愿意下注的事情，最热门类别好像是 Bitcoin 和 sports。主持人：也许还有 politics。嘉宾：politics 少一些。人们想知道 Bitcoin 和体育比赛会怎样，

### [01:16:05–01:16:34]

**EN**  but but these are very limited you know in terms of trying to forecast where things might be where the economy is going to go what what the and and the prediction markets don't tell you the reasons for why right so you actually want to have a very strong reasoning system that can actually tell you why it's you know 70% this team is going to win over this team why is this the case like who and it's sort of you know it's aggregation of the experts It's a

**中文**  这些当然有用，但对预测整体趋势、经济走向等而言非常有限。而且 prediction market 不告诉你原因。你需要很强的 reasoning system，解释为什么某队有 70% 胜率：哪些人、哪些因素？它要 aggregation expert opinion，聚合

### [01:16:32–01:17:02]

**EN**  segregation of sports statistics like what they've done and such. And so I think that ultimately AI is going to be, you know, I I like to think about this is like, and again, maybe it's a little bit of a kind of my just skewed view, but McKenzie, why does McKenzie exist? Boston Consulting Group, these are big organizations, you know, all of them should be replaced by AI because AI is going to give you a much better information. Like McKenzie does these things where they say, well, company

**中文**  sports statistic 等。我认为 AI 最终会……也许这是我偏颇的看法，但 McKinsey 为什么存在？Boston Consulting Group 为什么存在？这些大机构都应该被 AI 取代，因为 AI 会给出更好的信息。McKinsey 的工作是：公司

### [01:17:00–01:17:28]

**EN**  wants to invest in this thing. McKenzie goes does the report you have analysts doing research and everything they come to you and say like hey yeah the probability of you succeeding is like 80% I think this is the right decision this is what you should do and this this right and this will all be replaced by AI >> big big statement wow >> I mean all of >> I'm just saying like because those are huge huge businesses right so >> those are huge businesses but a lot of it is going to be replaced by I think a lot of it I mean maybe you're going to

**中文**  想投资某件事，他们让 analyst 研究并出报告，然后说：“成功 probability 是 80%，这是正确决定，应该这么做。”这些都会被 AI 取代。主持人：好大的论断，这些可是规模庞大的业务。嘉宾：确实，但我认为其中很多会被取代。也许

### [01:17:26–01:17:55]

**EN**  have humans in the loop kind of like but there's no reason why I mean McKenzie and those guys. I mean, they have one of the caveats, which is the insurance caveat, which is, you know, if I'm a CEO and I'm making this big investment, like a multi-billion investment. I'm hiring McKenzie, they go do the research to me and they tell me like that's the right thing to do. And if I do it and I fail, right, I go to my stakeholders and say, "Hey, McKenzie, you also said this is a good thing to do, right?"

**中文**  仍有 human in the loop。不过 McKinsey 等公司有一个 insurance caveat：作为 CEO，我要做数十亿美元的大投资，于是聘请 McKinsey 研究；他们告诉我这是正确选择。投资失败时，我可以对 stakeholder 说：“McKinsey 也说这是好决定。”

### [01:17:52–01:18:21]

**EN**  >> So, that's the role that those play, right? It's like an insurance, but I don't think their predictions are going to be better. like AI is going to be able to make much better prediction because again the amount of information it can process and everything is just going to be looking >> way more than a human. Yeah, >> they way more than like 10 analyst or maybe you need one analyst working with AI. >> Yeah. >> You know I mean maybe that's that's the scenario because you still might need a human to kind of like but >> just coordinate things but

**中文**  这就是它们提供的 insurance 角色。但我不认为它们的 prediction 会更好；AI 能处理的信息量远超人类，因此会做出更好的 prediction。主持人：远超一个人。嘉宾：也远超十名 analyst。也许只需一名 analyst 与 AI 合作，因为仍可能需要人

### [01:18:20–01:18:49]

**EN**  >> Yeah. >> Exactly. But I I I think that all of And so this is like systems, decision-m systems that can predict the future is going to be that's kind of like what I think again I might be wrong. Maybe the whole thing is going to fall apart because a lot of you know I I I do have like a lot of opinions. Some of my best students are basically saying I think it's a hard problem. It's not going to work >> and it's great you know and it's we'll see. I I know maybe it's a hard problem and maybe it's it's such a hard problem that

**中文**  协调。总之，能预测未来的 decision-making system 会非常重要——这是我的判断，也可能错。也许整个设想会失败。我确实有很多观点，而我最优秀的一些学生就说：“这是个难题，不会奏效。”这很好，我们会看到结果。也许它确实太难。

### [01:18:48–01:19:17]

**EN**  >> but to me >> prediction markets was a confirmation that with the power of the crowd it's the best prediction right now of any sort of model. >> Yeah. >> Right. It's very hard to beat the crowd because if you could beat the crowd you could make money you know just off the cal poly mark. It's very hard to crowd. It's not like you have a system that's like, well, I can just dominate, right? But it also points out that humans and at the and the crowd itself is fairly

**中文**  对我而言，prediction market 是一种确认：依靠 crowd 的力量，它是眼下所有模型中最好的 prediction。很难胜过 crowd，否则你可以直接在 Kalshi、Polymarket 上赚钱。并没有哪个系统能够轻易碾压市场。这也说明人类形成的 crowd 本身相当

### [01:19:14–01:19:43]

**EN**  efficient and you know, but but I I still think AI should be able to do better. >> Yeah, >> it's the wealth of information I can consume and if I can build the right sort of systems to do it, >> um, we'll see. It's an interesting hypothesis to test. So, >> yeah, for sure. So number one, I wanted to get your thoughts on all these Chinese distilled models. I think a lot of people are talking about this. Um like what are your thoughts on the

**中文**  efficient，但我仍认为 AI 应该能做得更好，因为它能吸收海量信息——前提是我们构建出正确的系统。我们会看到，这是假说，可以测试。主持人：我还想问所有这些中国 distilled model。很多人都在谈，它们会怎样影响美国 LLM business，尤其 frontier model？

### [01:19:41–01:20:10]

**EN**  impact of these models on the future of the LLM business in the US maybe like you know the frontier models? >> So I think that the the models that are coming out of China are actually very good models. I don't know if they distilling or not distilling. I mean obviously Anthropic says they distilling but I have no idea. I do know. Do you know a company called Moonshot AI? Very recent models. >> Kimmy models. >> So the CEO of this company, Zillinga. He's my former student at CMU.

**中文**  嘉宾：中国推出的模型其实非常好。我不知道它们是否在 distill；Anthropic 说它们在做，但我不知道。你知道 Moonshot AI 吗？最近的 Kimi model。它的 CEO 杨植麟（Zhilin Yang）是我以前在 CMU 的学生。

### [01:20:09–01:20:34]

**EN**  >> Oh, cool. >> So he is one of the smartest people I know. >> Like he was he was a PhD student at CMU. He's like one of the best. And so China has strong talent, right? They now have the resources. So I think they'll be able to compete with the US. And >> and they're just making everything open source, which is crazy. which is I I think to the community like

**中文**  主持人：很酷。嘉宾：他是我认识的最聪明的人之一，是 CMU 最优秀的博士生之一。中国拥有强大人才，现在也有资源，所以我认为有能力和美国竞争。他们还把一切 open source，这很惊人，也让 community 获益。

### [01:20:32–01:21:00]

**EN**  I can tell you all of my students right now using quen models, sche models to do research with. >> Yeah. >> I mean >> and maybe GLM now >> and GLM. GLM is a little bit expensive but it's a great model for distilling for like doing research and such. So because these are open models you can see the reasoning traces and everything. So I think that the strategy of open-sourcing models is actually a very

**中文**  我的所有学生现在都用 Qwen 和 DeepSeek model 做研究。主持人：也许现在还会用 GLM。嘉宾：对，GLM 略贵，但很适合 distilling 和研究。因为这些是 open model，你能查看 reasoning trace 等。我认为 open source model 是非常

### [01:20:58–01:21:25]

**EN**  good strategy because it first of all it gives academics it also like you know gives startups as a way to you know iterate on these models fine-tune see what works see what doesn't work right and and so now with GLM the gap becomes fairly small and so it gives opportunities for for others to sort of you know test and and build upon right I'm actually very surprised that US doesn't have a good strategy for open

**中文**  好的 strategy：首先让 academic 能使用，也让 startup 迭代、fine-tune、测试哪些方法有效。如今 GLM 把差距缩得很小，让其他人有机会测试并在其上继续开发。我很惊讶美国没有好的 open

### [01:21:23–01:21:52]

**EN**  source ing the models. You know, Llama was the last llama 3. Llama Llama 4 didn't work so well, but Llama 3 was one of the last sort of good open source models that people were using and relying and people are still relying on. But now, I think it's a sad kind of state of affairs. I have a hypothesis about this. My my hypothesis is if if these Chinese companies are actually distilling the the like anthropic and open AAI models, well, they're not doing

**中文**  source strategy。Llama 3 是最后一批真正优秀、大家广泛使用和依赖的 open source model；Llama 4 表现不佳。现在的局面令人遗憾。主持人：我有个假说：如果中国公司确实在 distill Anthropic 与 OpenAI 的模型，美国公司没有这样做是因为法律限制，而中国公司也许没那么在乎美国法律。

### [01:21:49–01:22:18]

**EN**  it in the US because of laws and in China maybe they just don't care as much about following the US laws. So that's probably the main reason is my guess. I mean it could be it could be but again yes if they are kind of like again but if they're open sourcing then it kind of like also gives >> a whole bunch of startups and academics and such to actually have access to these models run and and develop continue developing the technology.

**中文**  嘉宾：可能如此。但如果它们 open source，也会让大量 startup 和 academic 获得模型，运行它们并继续开发技术。

### [01:22:16–01:22:46]

**EN**  >> Yeah. So what do you think? So what do you think these like open source models like the Chinese open source models mean for the like US industry then? Like what do you think like how do you think the industry is going to have to evolve? >> I don't know. I don't know because I think that at some point what's going to happen is that I mean there are a few startups that are kind of like very well funded. I think reflection that are supposed to kind of like produce strong open source models. Uh but we'll have to see. Yeah,

**中文**  主持人：这些中国 open source model 对美国 industry 意味着什么？行业要怎样演变？嘉宾：我不知道。现在有几家 funding 很充足的 startup，例如 Reflection，据说会推出强大的 open source model，还要继续观察。

### [01:22:44–01:23:13]

**EN**  Lava was again I I think Meta kind of like positioned itself differently now. I don't think they're going to be open sourcing the models and so that's a sad state of affairs. I think that it's true the laws in the US and what you can distill and such are stricter but yeah I don't know I don't know how it's going to evolve. I mean it's it's uh I I I think that we need more startups in the US like who will you know actually end the right the problem is that you need a lot of capital to do it right and it's

**中文**  Meta 如今似乎重新定位，我认为不会再 open source model，这是令人遗憾的局面。美国关于能 distill 什么的法律确实更严格，但我不知道未来会怎样。美国需要更多 startup 来做这件事；问题是需要大量 capital。

### [01:23:11–01:23:39]

**EN**  not it's we know exactly what needs to be done it requires infrastructure capital and GPUs to sort of you know build build an open source but yeah I don't know if if US needs to do something you know definitely US needs to do something about it. Yeah, >> because kind of like said like if you look at all the prominent startups and such they all post training like even cursor >> right

**中文**  需要做什么其实很明确，只是需要 infrastructure、capital 和 GPU 才能构建 open source model。美国肯定需要采取行动。观察知名 startup，它们都在做 post-training，就连 Cursor 也是。

### [01:23:37–01:24:07]

**EN**  they use Kimmy post train the models right because like I mean it's it's it's a good model and then you just start post training and then a lot of startups doing the same thing relying on Quen Deepc Kimmy GLM to post train because it's cheaper way way way cheaper than frontier models and if I can post train I can own the model like That's my model like nobody touches it like I it's not like you know people think that if it's if it's a model coming from China then somehow like the information leaks

**中文**  它们使用 Kimi 再 post-train。很多 startup 都依赖 Qwen、DeepSeek、Kimi、GLM 做 post-training，因为比 frontier model 便宜得多。如果我能 post-train，就能拥有自己的模型。有人以为模型来自中国就会泄露信息，但不是这样，它只是 weights。

### [01:24:04–01:24:33]

**EN**  no it's contained it's just the weights right and if I can post train it's my model >> so why oh sorry go ahead >> and if it's if base model is already good enough has base knowledge that I can post train them then it's it's great I mean I see startups from speech and such they all relying on these open source models post training them and creating products out of them. >> Yeah, I was going to ask why is it important for the US to have our own

**中文**  如果我能 post-train，它就是我的模型。如果 base model 已足够好、拥有基础知识，再做 post-training 就很理想。我看到 speech 等领域的 startup 都依赖这些 open source model，做 post-training 后构建产品。主持人：既然中国模型很好，美国为什么还需要自己的

### [01:24:32–01:25:00]

**EN**  open source models? If the Chinese models are really good, then why do we need so >> American open source? >> Yeah, that's a good question. I think one sort of caveat is we don't know. I mean for things like coding that's fine, but for general knowledge there is a little bit of bias that's being objected to any model, >> right? And so if you're relying if your business your entire business relies on open source models produced by a different country right and so it might

**中文**  open source model？嘉宾：好问题。有一点 caveat：coding 之类的任务没问题，但 general knowledge 会向任何模型注入一定 bias。如果整个 business 都依赖另一个国家生产的 open source model，

### [01:24:59–01:25:28]

**EN**  have some kind of like you know you can ask it some questions about like kamin square see what it does right like there's there's certain kind of you know nuances that some of these models and I've talked to some people and they're like yeah we know that these models like math coding that's all fine like coding there's no sort of you know political bias or like cultural bias right it's like you know math is the same thing. It's like here's the problem. This is the answer like either you got it or not. uh but more kind of you know if you

**中文**  就可能有问题。可以问它某些关于 Tiananmen Square 的问题，看看怎样回答。我与一些人讨论过，他们知道 math 和 coding 没问题，因为 coding 没有 political 或 cultural bias；math 也有确定答案。但如果用它

### [01:25:25–01:25:54]

**EN**  kind of I don't know use it for creative writing or if you use it for like other things or in the future if you're going to be relying this as a ground truth source like sometimes right now I'm no longer going to like I do go to Google but like sometimes you go to like cloud or Gemini and just ask it like hey what about tell me about this event and it kind of like tells you right and I don't bother verifying I just assume that it's truth and so like these biases may come in right and so I yeah I don't know

**中文**  做 creative writing 等，或未来把它当作 ground-truth source，bias 就可能进入。现在我仍用 Google，但也常直接问 Claude 或 Gemini 某个事件，它会告诉我；我懒得验证，就假定为真。因此这些 bias 可能进入，我也不知道最终影响。

### [01:25:52–01:26:20]

**EN**  you know Yeah. I had a Yeah. So, another question was so it seems like there's a ton of these like RL environment businesses or these data labeling businesses that are sprouting up all like all over the place. Uh do you think that is a a sticky business like it will be around a few years from now? >> No, I don't think it's a sticky business. I mean people who are building RL environments and people I mean having a building a good environment is

**中文**  主持人：还有个问题。许多 RL environment business 或 data-labeling business 到处涌现，你认为这种业务 sticky 吗？几年后还会存在吗？嘉宾：我不认为它是 sticky business。构建优质 RL environment 当然重要，但

### [01:26:17–01:26:47]

**EN**  obviously important but what I've seen right now especially in the environment space now that the codings become so good creating these environments just becomes cheaper and yeah I I I don't know I I I suspect there's still going to be like couple of players big players who will provide the environments the data because you need the data and uh people who work very closely with the frontier labs and such they might they might survive but I don't see it as a sustainable business long term. >> Yeah. Especially because you were saying

**中文**  现在 coding 能力提升，创建 environment 变得更便宜。我猜仍会有少数大公司提供 environment 和 data，因为 data 仍然需要；与 frontier lab 紧密合作的公司也许能生存，但我不认为它长期可持续。主持人：尤其因为 synthetic data？

### [01:26:45–01:27:14]

**EN**  because of synthetic data, right? >> Synthetic data. And then the other thing is that like I think like some of these businesses are now trying to shift towards robotics and I suspect a lot of them will, right? Uh >> like OpenAI anthropic are shifting to robotics or >> no some of the data collection. >> Oh, I see. I see. >> Right. Like people who kind of like develop data like collecting data environments, they're trying to kind of like see okay like can we can we go into the physical world and start collecting data in the physical world for robotics applications. So these data kind of like

**中文**  嘉宾：对。另一点是，有些公司现在尝试转向 robotics，我猜许多都会这么做。主持人：OpenAI、Anthropic 转向 robotics？嘉宾：不，是 data collection 公司。它们在想，能否进入 physical world，为 robotics application 收集现实数据。所以这些 data

### [01:27:12–01:27:41]

**EN**  collection companies will adapt but obviously you know it's there's always going to be the need for you know frontier when they deploy the models they kind of like find some edge cases they might go to these data kind of like or or these companies say like create me an environment create me the data so it can kind of like fix those edge cases but I don't know yeah >> I don't know I mean data is definitely going to be important no matter what for physical AI and for but I I still think there's going to be maybe a few players

**中文**  collection company 会适应变化。frontier lab 部署模型后总会发现 edge case，可能要求这类公司创建 environment 或 data 来修复。所以 data 无论对 physical AI 还是其他领域都很重要，但我仍认为最终可能只有少数 player。

### [01:27:39–01:28:09]

**EN**  Yeah, but not like I mean but not like you know there's so many people building RL environments. I think it was two years ago it was like people like okay we need those because we can train our models. Now they just becoming cheaper and it's probably >> and now now like I don't know if you saw like do you know the company Handshake? So Handshake was this platform where like a lot of undergrads would use it to get jobs and now they have basically pivoted to selling data to the lab. So it's like this weird dystopian thing

**中文**  不像现在这么多人都在构建 RL environment。两年前大家说训练模型需要它们；现在成本越来越低。主持人：你知道 Handshake 吗？它原本帮助许多本科生找工作，现在转型向实验室出售 data。这有种 dystopian 色彩：原本帮本科生找工作，现在却让本科生标注 data，帮助实验室取代他们原本想得到的工作。

### [01:28:07–01:28:36]

**EN**  where it's oh you were trying to get undergrads jobs but now you're basically telling the undergrads to label data for the labs to basically take the jobs that the undergrads were trying to get in the first place. >> Fair enough. Yes. Yes. I think that's kind of like you know Yeah. Yeah. Yeah. I mean data is still like high quality good data is still going to be valuable. Uh you know like and and humans collecting the data high quality data is still going to be valuable and so that's why they're probably selling it because it's it's val but but I agree with you.

**中文**  嘉宾：确实。高质量的优质 data 仍会有价值，人类采集的高质量 data 也一样，所以他们会出售它，因为它有价值。但我同意，这种情况很奇怪。

### [01:28:33–01:29:03]

**EN**  kind of like it's it's weird kind of. >> Yeah. So my my next question was on SAS. So it seems like a lot of a lot of people are very scared that basically this is like the end of SAS because it's so easy to copy software now. What are your thoughts on that? It's a good question. I don't know. >> Okay. I don't know. I think that the SAS companies that the software companies will need to adapt and a lot of it is based on how much mode they have, the client relationships and the legacy

**中文**  主持人：下一个问题是 SaaS。很多人担心这意味着 SaaS 末日，因为复制 software 太容易了。你怎么看？嘉宾：好问题，我不知道。SaaS company 和 software company 必须适应；关键在于有多少 moat、client relationship 和 legacy

### [01:29:00–01:29:29]

**EN**  software they have. But I I I do think now with these agentic systems and the coding systems like you know SAS gonna take a hit. >> Yeah. You know like it depends how it's going to evolve. I mean we've seen so for example Salesforce been down for like a lot in the last two years right? Service Now same thing like these old SAS kind of like companies and snowflake and you know those kind of so I think that they'll have to innovate. I think

**中文**  software。但我确实认为 agentic system 与 coding system 会打击 SaaS。具体取决于发展。过去两年 Salesforce 下跌很多，ServiceNow 也一样；这些传统 SaaS company，以及 Snowflake 等，都必须创新。我认为

### [01:29:27–01:29:55]

**EN**  that software is going to be cheap. >> Yeah. It's it's >> do you agree with people when they say that the cost of intelligence is approaching zero now is going to zero? >> I don't know if it's going to go completely to zero but what's going to happen is that if you have models like 5.6 or methas right super expensive it now basically comes down to intelligence per token like cost per token like how intelligent your model is and how

**中文**  software 会变便宜。主持人：你同意“intelligence 的成本正在接近零、最终会归零”吗？嘉宾：我不知道是否会完全归零。像 5.6 或 Opus 这样的模型非常昂贵，最终会归结为 intelligence per token，也就是 cost per token：模型有多智能，以及

### [01:29:52–01:30:22]

**EN**  efficient it is and how how cheap it is >> compared to a human essentially >> compared to or compared to any other models. Okay. >> Like you know like I was talking to who was I talking like somebody at LinkedIn and then they're like well early on they were they were trying to use like 8B or 9B models because these are cheaper models and for what they do that's enough. So there is like there's this notion of intelligent. You obviously want to have intelligent models but then the cost is the other kind of like I want intelligence but I want to have it at a very low cost. And so obviously

**中文**  多 efficient、多便宜。主持人：本质上是与人相比？嘉宾：也可以与其他模型相比。我曾与 LinkedIn 的某个人谈过，他们早期试图用 8B 或 9B model，因为更便宜，而且足够完成任务。也就是说，一方面想要 intelligent model，另一方面又希望以很低成本获得 intelligence。

### [01:30:21–01:30:49]

**EN**  like this inference like all these people doing inference chips and and whatnot. They're all trying to basically say I want to be able to give you intelligent tokens but for cheaper. So I don't know if the cost is going to go down to completely be uh close to zero. And this is where open source models are kind of like coming online right because they're like hey they're intelligent but like tenth of the cost. >> I mean as a business of course I'm going to be using that like it's not at the same level as methods but it's good

**中文**  所以做 inference chip 的人都在试图提供更便宜的 intelligent token。我不知道成本会不会接近零。open source model 正是在这里发挥作用：它们会说，自己的模型同样智能，却只有十分之一成本。作为企业，我当然会用；即使不如最强模型，只要足够

### [01:30:48–01:31:17]

**EN**  enough for what I do. >> 90% of their 90% of the way there for 10% of the cost. Something like that. >> Yeah. Exactly. And then that's that's a good business argument for me to use it, right? Like I mean why would I use So I don't think it's going to go to zero. I think it's but it will it will be cheaper and cheaper because there's always ways of distilling these models into smaller models. >> Why why do the Frontier models keep getting more expensive? Is it just because they're using up more compute?

**中文**  完成任务，花 10% 成本得到 90% 能力，就是很好的 business argument。我不认为成本会归零，但会越来越便宜，因为总有办法把模型 distill 成更小模型。主持人：Frontier model 为什么越来越贵？只是用了更多 compute 吗？

### [01:31:14–01:31:43]

**EN**  >> No, because I think there is a demand. >> Okay. So it's just they want to increase the price. >> They want to increase the price. There's a demand. And then you know there there's been some wars before right like I mean like I've seen these examples where open the eye would say like here's our price and then like 2 months later they would like cut the price by like five. Okay why' you do that? I mean they doing it because it's competition like because there are cheaper models that might be just as good maybe a little bit worse but like way cheaper. So it's like

**中文**  嘉宾：不是，因为有 demand。主持人：所以只是想涨价。嘉宾：有需求，就会涨价。过去也发生过 price war：OpenAI 先报一个价格，两个月后又降到五分之一。为什么？因为有竞争，有些更便宜的模型也许一样好，或稍差一点，却便宜得多。

### [01:31:41–01:32:11]

**EN**  it's all about and this is where they kind of the efficiency comes into play. Intelligence is one thing efficiency is another thing. So like it's it's yeah it's brutal because you know if I can do my job at 90% but at a tenth of the cost maybe I'm okay kind of you know >> maybe it makes sense. >> Yeah it's like e economically I think it all is going to come down to economics right like because I can't whatever is cheaper I'm going to use.

**中文**  所以 efficiency 开始重要。intelligence 是一回事，efficiency 是另一回事。竞争很残酷：如果十分之一的成本就能把任务做到 90%，也许已经足够，在经济上更合理。最终都归结为 economics，我会使用更便宜的方案。

### [01:32:09–01:32:39]

**EN**  >> So you think that all the LLMs are basically commodities? >> I think they're going to be commodities. Yes. >> Wow. Yeah. I mean, it's interesting because like early on when the LLM first came out, that's what everyone was saying. But then, you know, then they all the labs just kept competing with each other. It kept like outdoing each other. But yeah, it seems like everyone's basically converging back on that point, which is they're basically going to converge on the same thing. Yeah, I think that it's like you know some models are going to be a little bit better and it's all now is going to come

**中文**  主持人：所以所有 LLM 基本都会成为 commodity？嘉宾：对，我认为会。主持人：很有意思。LLM 刚出现时大家都这么说，后来实验室不断竞争、相互超越；现在似乎又回到同一个结论，最终会趋同。嘉宾：是，有些模型会略好，但一切将

### [01:32:37–01:33:06]

**EN**  down to the product >> like you know codex and code clo is a fantastic product like because you can integrate itself with your code base and everything and like once you kind of like when it becomes sticky kind of like use it right even if somebody's a little bit better it's all going to come down to like products like what product you have and how easy it is to use and how cheap it is. >> Yeah. >> Right. It's >> what what are you really excited about that's happening in the industry maybe research- wise or otherwise that you

**中文**  取决于 product。Codex 和 Claude Code 都是很棒的 product，可以集成到 codebase 中；一旦产生 sticky usage，即使别人模型稍好，你仍可能继续使用。最后看产品是什么、是否易用和便宜。主持人：你对行业里正在发生的哪些 research 或其他变化最兴奋？哪件事会带来下一次 step-function change？

### [01:33:05–01:33:34]

**EN**  think is going to create like that next step function change and what's going on? >> It's a good question. It's a good question. I don't know. I mean there are some kind of there is a lot of work on self-improving systems self-approving agents self-improving which seems to be promising. So these systems again this is notion of you have existing models it generates some data you filter it you use this data for for training or you use like these sort of what we call parallel systems where you do rollouts you do you do multiple rollouts you figure out which one solve the problem which ones didn't and the ones that

**中文**  嘉宾：好问题。我不知道。self-improving system 和 self-improving agent 的研究很多，似乎很有希望。现有模型生成 data，经过 filter 后用于 training；或者使用所谓 parallel system，做多次 rollout，找出哪些解决了问题，再用成功 rollout 的

### [01:33:32–01:34:02]

**EN**  solve the problem uses the data for kind of improving improving the system so you know like AI that kind of like self-improves is get getting a little bit better is one direction that people are looking at which is an interesting one I think decision making forecasting is going to be another one and I think The third one is robotics. This is sort of like this notion that this is one big area that we're missing right now. And that's the area where there's probably going to be a lot more investment, but it's also the area that's going to take the longest to actually produce something meaningful. >> And for robotics, you're specifically

**中文**  data 改进系统。能够 self-improve、逐渐变好的 AI 是一个有趣方向。decision-making 和 forecasting 是另一个。第三是 robotics：这是目前缺失的一大领域，未来会获得更多 investment，但也会花最长时间才产出有意义的东西。主持人：robotics 特指 humanoid 吗？

### [01:34:01–01:34:30]

**EN**  talking humanoids. >> Doesn't have to be humanoids, but it's kind of like one of those things is that somebody was telling me, you know, which robots kind of like have shown to be useful like actually kind of, you know, people use and in in the real world. And there's only one example which is Roomba, right? Just cleans your house. There's no other example right now where a robot is useful where I would pay like whatever 5,000 for it to like do stuff for me. >> Yeah. For consumers, I guess, right?

**中文**  嘉宾：不一定。有人问我，哪些机器人已证明真正有用、在现实世界里有人使用？只有 Roomba 这一个例子，它会打扫房子。目前没有其他机器人值得我花 5,000 美元来替我做事。主持人：是指 consumer 市场？

### [01:34:28–01:34:56]

**EN**  >> For consumers. But even for enterprises, I mean, yes, there is like manufacturing and stuff like that, but even there hasn't been I mean, if you look at Amazon, but all the robots of like kind of like moving things around, these all like predefined. So, these are not really these not really robotics. It's more like like you know a system that like moves things around. It's all like very specific or like assembly robots like they all very specific, right? It's

**中文**  嘉宾：consumer 是这样。enterprise 也类似，当然 manufacturing 有机器人，但 Amazon 仓库里搬东西的机器人都按预定义方式行动，不算真正的通用 robotics，只是搬运系统；assembly robot 也都针对特定任务。

### [01:34:53–01:35:22]

**EN**  not but there aren't any examples where you would basically either in enterprise maybe I don't know maybe there's some enterprise but consumers in enterprise like it's not like hey Tesla Optimus that goes and like assembles the cars and and like replace like a 100 humans. Yeah. It's just it's just not there. Right. So, and it's going to take I mean the best case we have right now is self-driving cars as kind of like the actual robot that sort of does something

**中文**  目前没有 Tesla Optimus 真正走进工厂组装汽车、取代一百名工人的例子。它还不存在。现在最好的实际案例是 self-driving car，确实是一种能做有意义事情的 robot。其他方面，我们看过机器人 cooking、走动等 demo，但真正应用还没到来。

### [01:35:19–01:35:48]

**EN**  meaningful. Others have not you know we see these examples people like cooking like robots cooking or or like doing stuff they walking around and but it's all at this point it's still yet to to come you know and so so I think that's the area there's going to be a lot of research and development and investments as well. >> Why is that why why is that not progressing as fast as LLMs? I think one of the things in robotics, you know, we've seen examples of like robots doing

**中文**  因此 robotics 会有大量 R&D 与 investment。主持人：为什么它没有像 LLM 一样快速进步？嘉宾：我们看过机器人 backflip、跳舞、移动、上下楼梯，这些 locomotion 都是进步。但 robotics 最大的 bottleneck 之一是 manipulation，也就是实际完成有用事情的能力。

### [01:35:46–01:36:15]

**EN**  like back flips, dancing, moving around, like locomotion, going up and down the stairs. It's all great. That's the progress. One of the biggest bottlenecks right now in robotics is manipulation. Like the ability to actually do useful things, you know, >> with your hands specifically >> with your hands or like doing any sort of like can the robot clean up my kitchen? >> Can the robot like load my dishwasher, wash it for me, unload it for me? Can my robot like you know if it's a factory can it go and like do some kind of

**中文**  主持人：特指用手？嘉宾：用手，或者执行任何操作。机器人能清理厨房、装载 dishwasher、清洗再卸下餐具吗？在 factory 里能完成 wiring 或其他

### [01:36:13–01:36:43]

**EN**  wiring right or do like multiple tasks and do it reliably not just sort of like you know people show that it's like sorting can kind of like sort and things but it's like very very sort of >> and very carefully structured like little things of the whole thing will kind of there was this and so like manipulation like the ability of actually do meaningful uh tasks is still not there. manipulation is one of the big problems that like yeah >> it's not not there. >> What what about that is difficult

**中文**  多种任务，而且可靠执行吗？现在的 sorting demo 都经过极为精心的结构化设置，稍有变化整个系统就失效。因此完成有意义任务的 manipulation 能力仍不存在，这是个大问题。主持人：它究竟难在哪里？

### [01:36:42–01:37:10]

**EN**  because it seems like okay so for example for self-driving cars the big unlock was just throw throw tons more data at it and it'll figure it out. Why can't we just do manipulation? >> Yeah. So for this one it's it's difficult because like we don't have we don't have good hardware universal hardware like Tesla develops its own unit develops its own like everybody's developing its own. That's number one. And number two, building a million robots that will do data collection is very expensive. And so like and so you

**中文**  self-driving car 的突破似乎只是加入海量 data，系统就学会了。为什么 manipulation 不行？嘉宾：首先，我们没有优秀的 universal hardware。Tesla、Unitree 等都在各自开发自己的设备。其次，造一百万个机器人收集 data 极其昂贵，所以

### [01:37:09–01:37:36]

**EN**  need sort of like and so this this data collection and manipulation and everything is still like very early right now. >> I see. And so like this is data collection kind of like companies are trying to pivot right now to sort of say well we can we can take humans like take humans in the factories put you know >> goggles on and then like and then they collect data for robots right I mean maybe that's the way to do it but nobody has figured out like what what you know

**中文**  manipulation data collection 仍处于极早期。现在 data collection company 正试图转型：让工厂工人戴上 goggle，为机器人采集 data。也许这可行，但没人弄清怎样高效使用这些数据。

### [01:37:35–01:38:04]

**EN**  how to use that data efficiently right and LLMs for LLMs like the data is on the web like that's unlaw right there's no like internet scale data for robotics it's just it's just out there. >> Yeah, maybe that's why Apple is trying to put cameras in the ear earpods now. >> Could be. >> I don't know. Yeah. [laughter] But yeah, so my my last question was if you were in college right now, what would you study? >> Good question. Um my view are going to

**中文**  LLM 所需 data 本来就在 web 上，而 robotics 没有 internet-scale data。主持人：也许这就是 Apple 想把 camera 放进 AirPods 的原因。嘉宾：可能吧。主持人：最后一个问题：如果你现在上大学，会学什么？嘉宾：好问题。我的观点会

### [01:38:02–01:38:31]

**EN**  be very biased. I think studying STEM related areas, I still think that's the right thing to do. I think that studying computer science, statistics, AI, machine learning, that's what I would do. Obviously there is like engineering computer engineering and engineering efforts like anything sort of like in STEM I think you know is going to be relevant. >> So so you don't agree that we should just stop studying computer science because you've probably seen the graphs on on Twitter where it shows like

**中文**  非常有偏向性。我仍认为应该学 STEM。我会学 computer science、statistics、AI、machine learning；当然还有 computer engineering 等 engineering 领域，任何 STEM 方向都会有用。主持人：所以你不同意因为 AI 而停止学习 computer science？你可能看过 Twitter 上的图，

### [01:38:28–01:38:58]

**EN**  st the number of Stanford CS majors has gone down by like 40% in the last two years or something. >> You know my kid is going to be kind of like at some point few years from now applying. I do hope that the trend continues because it's going to be easy to get into computer science [laughter] you know like in the last three years since AI took over like CS like I mean you can see at CMU our PhD applications over the last 5 years just skyrocketed >> you know we have 1,200 applicants

**中文**  显示 Stanford CS major 人数两年下降约 40%。嘉宾：我的孩子几年后要申请，我希望趋势继续，这样 computer science 更好进。AI 兴起后的三年，CMU 的 PhD application 激增。过去五年更是如此，现在有 1,200 人

### [01:38:55–01:39:25]

**EN**  applying for PhDs we we we accepting like 50 and these are like strong people these are not like you know like you kind of like B and C type of students who are applying like these are >> best of the best yeah >> best of the best right And so you know for undergrads I I don't know I I actually think that in the future maintaining these systems improving these systems and everything is just going to require more and more uh computer scientists and look at the job

**中文**  申请 PhD，我们只录取约 50 人，而且申请者都很强，不是成绩普通的人，而是最优秀的人。我认为未来维护、改进这些系统会需要越来越多 computer scientist。看看 job opening 就知道，

### [01:39:23–01:39:53]

**EN**  applications like there's so many openings right now like if if the whole thing was automated like even anthropic is hiring bunch of software engineers or like ML research engineers right it's uh I don't know I know there's an argument for like you know if things completely shift certain way where you know these agentic systems going to be writing the code and everything and then you know like psychology and social sciences and such where like a little bit of a human interaction is needed might go up but I don't know I don't know

**中文**  现在仍有很多岗位。如果一切都自动化了，为什么连 Anthropic 还在招聘大量 software engineer 和 ML research engineer？当然，也有人认为如果 agentic system 真的接管全部 coding，需要 human interaction 的 psychology 和 social science 可能上升。我不知道，

### [01:39:51–01:40:19]

**EN**  >> what do you I think that technical people will always be in demand >> yeah I I agree but then there's people who say don't study anything that's verifiable >> oh like math >> verifiable like you were saying where we can easily spin a RL to train on it that to What's the answer? >> My my view is like I mean maybe like it's going to be a little controversial view, right? Which is math. If I take math, math is a fantastic subject to

**中文**  但 technical people 永远会有需求。主持人：我同意。也有人说，不要学任何 verifiable 的东西。嘉宾：比如 math？主持人：对，你刚说它很容易建立 RL 来训练，因为答案可验证。嘉宾：我的看法可能有些争议。math 是

### [01:40:17–01:40:46]

**EN**  study in like high school in college, but it's probably not kind of like useful like if you're studying like all these abstract kind of, you know, abstract algebra and like all these like crazy maybe it's not really applicable in the real world, right? So only few people kind of like very smart people kind of, you know, stay in this area. But studying math gives you the analytical thinking. >> So when you study computer science, when you study like data structures and everything, it just gives you very

**中文**  高中和大学里极好的学科；但 abstract algebra 等抽象领域也许在现实中并不实用，只有少数非常聪明的人留在其中。不过学习 math 会培养 analytical thinking。学习 computer science 和 data structure 也会给你一种高度

### [01:40:44–01:41:12]

**EN**  analytical way of thinking. And to me, that was like extremely useful for me to kind of like go and do research in that area. >> Yeah. Like the logical reasoning ability. >> Exactly. It's the reasoning. It's like math. It just it's just like, you know, I wouldn't go and kind of like go and do like pure math as my career, right? Because I don't find it appealing. I mean, some people do. I don't. But I did major in computer science and math. And the reason why I did math is because I thought that this

**中文**  analytical 的思考方式，这对我在该领域做 research 极有帮助。主持人：也就是 logical reasoning 能力。嘉宾：没错，就是 reasoning。我不会把 pure math 当成职业，因为自己不感兴趣，虽然有人喜欢。但我主修 computer science 和 math；选择 math，是因为我认为它

### [01:41:09–01:41:38]

**EN**  gives me strong analytical abilities to think and whatever the future is going to be, you always will have to adapt. People who are able to adapt will sort of be on top of things, right? And so, you know, >> Yeah. >> So, do you do so do you do you think that the skill sets that will matter in five years are the same skill sets that mattered 5 years ago? >> I think it will change. I think the change will come with the use of AI.

**中文**  赋予我强大的 analytical ability。无论未来怎样，人都必须适应；能适应的人会保持领先。主持人：五年后重要的 skill set，是否与五年前相同？嘉宾：会改变，变化来自对 AI 的使用。

### [01:41:37–01:42:06]

**EN**  >> Yeah. >> So, and I can see it. I can see it. I have a kid who is now in high school and then they just basically using VS Code, code and building websites, building startups and stuff like that, right? Because you can do it right now. And so people who can have these strong analytical skills plus able to you know use these tools and shift and kind of like make the use of both are going to be at the top.

**中文**  我已经看到这种变化。我的孩子正在读高中，直接使用 VS Code、Codex 构建网站和 startup，因为现在可以做到。兼具强 analytical skill，又能使用这些 tool、适应变化并结合二者的人，会处在最前面。

### [01:42:03–01:42:32]

**EN**  >> Yeah. So like hustle plus uh just that you know you can just move fast essentially. >> You can you can move fast. Yeah. I mean it's it's it's it's it's it completely changes the and people who don't use AI tools are going to be left you know and I've seen this like some of the smartest people I mean it it will be interesting to see especially people like in the 30s right now who kind of like 30s and the 40s who like built you know have in like this rapid shift right now you have to adapt so you have to

**中文**  主持人：也就是 hustle 加上快速行动。嘉宾：对，可以快速行动。它彻底改变了工作。不会使用 AI tool 的人会落后。我已经看到这种情况。尤其值得观察的是现在三四十岁、已经形成职业技能的人，面对这次快速转变必须适应，

### [01:42:30–01:43:00]

**EN**  know how to use these tools and make kind of like you know be more productive I think that's I'm seeing it with my PhD students they're like my workflow has completely changed last year. >> Yeah. >> Now we're interacting with AI tools and I can sort of like say let's do this, let's do this and kind like before it would take me 3 days to code something up. Now it takes me an hour and so if it takes me now I can actually progress on my thing. And so people who can take advantage of this they moving way faster. >> Yeah. And then I've I've I've seen a lot of people that are like losing their

**中文**  学会使用这些 tool，提高 productivity。我的博士生会说，过去一年 workflow 完全改变。现在我们与 AI tool 互动，可以不断让它做不同任务。过去写出某个东西要三天，现在只要一小时，于是能真正推进工作。善用工具的人会快得多。主持人：我也看到很多人正在失去

### [01:42:59–01:43:28]

**EN**  identities because they're like I just spent my entire career learning like being a good coder and now I can just do what I learned in 20 years in like 20 minutes basically. So yeah, there's also that side of it. Yeah, >> there is. Yes, that that's that's kind of like yeah, there is there is and it also happens in academia because there's a lot of people who worked on like dialogue systems and everything and it's just it's like completely changed but change is inevitable. kind of like you

**中文**  identity，因为他们整个职业生涯都在学习成为优秀 coder，如今二十分钟就能做完过去二十年所学的东西。这是另一面。嘉宾：确实。academia 也一样，许多人过去研究 dialogue system 等，整个领域完全改变了。但 change 不可避免，

### [01:43:25–01:43:53]

**EN**  have to you have to you know it's like you have to adapt like it's it's but it's an exciting I think it's exciting development like I mean it's definitely changed kind of and it empowers like I mean I do coding but now I can be so much more productive right like I can actually go and you know I'm doing this thing with my students and it's like some weird bizarre error on some of the packages would take me days to figure out because I would go to Google and say I have this error what do I do and it's

**中文**  必须适应。我认为这是令人兴奋的发展，也赋予人力量。我会 coding，但现在 productivity 高得多。我和学生做项目时，遇到 package 中莫名其妙的 error，过去要花几天 Google 搜索解决办法，而现在

### [01:43:51–01:44:16]

**EN**  like look at this look at it now I can just ask code like you know this is an error like just figure this out how to install these things like I don't even need to think about like that's not what I want to think about and just like yeah like this and this and >> gets the right to the right thing that I don't even have to think about right so >> yeah awesome well I mean those are all the questions that I had thank you for staying over >> awesome awesome well thanks for having

**中文**  只需把 error 交给 coding agent，让它弄清怎样安装依赖。我根本不想把精力花在这种事情上，而它会直接找到正确解决方案，我无需思考。主持人：太好了，这就是我的所有问题。感谢你延长时间。嘉宾：很棒，谢谢邀请。
