Deep learning 的早期反共识与 scale 转折
00:01:23–00:11:36Insight
- Geoffrey Hinton 对深层模型的数学论证和研究热情,促使 Russ 从金融业回到 University of Toronto 攻读博士。
- AlexNet 在 ImageNet 上的大幅提升最初好到令人怀疑,几个月后才被社区接受,并推动 Google、学术界和主要实验室全面转向 deep learning。
McKenzie, why does Mckenzie exist? Boston Consulting. These are big organizations. All of them should be replaced by people say AGI is two years. And people who actually work in these lab and frontier labs, I know a lot of them they basically saying just like we're not going to have AGI in 2 years. A lot of people were basically saying, my god, OpenAI has figured out a new architecture. They figured out like AGI that's absolutely not true. The architectures are the same. McKinsey 为什么会存在? Boston Consulting 也是。这些都是大机构,而那些说 AGI 两年就会到来的人认为,它们全都应该被取代。但真正任职于这些实验室和前沿实验室的人——我认识很多——基本都会说:我们不可能在两年内实现 AGI。很多人当时都在说:天啊,OpenAI 找到了一种新架构,他们已经攻克了 AGI。事实绝非如此,大家使用的架构都一样。
The CEO of this company, Zillian Yan, he's my former student at CU. He is one of the smartest people I know. China has strong 这家公司的 CEO 杨植麟(Zhilin Yang)是我以前在 CU 的学生,也是我认识的最聪明的人之一。中国拥有强大的
主持人 talent. They now have the resources. So I think they'll be able to compete with the US. So if I am a CEO of a company and I have a better view of the future than my competitor, I'm going to win. So you think that all the LLMs are basically commodities? I think they're going to be commodities. There's only one example which is Roomba. There's no other example right now where a robot is useful. In 2005, neural networks were a joke. Third choice behind support vector machines. 人才,现在也有了资源,所以我认为中国有能力与美国竞争。假如我是一家公司的 CEO,而且比竞争对手更清楚未来会怎样,我就会赢。主持人:所以你认为所有 LLM 最终基本都会商品化?嘉宾:我认为它们会成为商品。目前真正有用的机器人只有 Roomba 这一个例子,没有第二个。 2005 年,神经网络还是个笑话,只是排在支持向量机之后的第三选择。
At the time, Russ Celakudinoff was working in finance, but by a chance encounter, he bumps into Jeff Hinton on the street in Toronto. Hinton drags him into his office and he walks out doing a PhD in deep learning before deep learning was even a term. And if you're not familiar with Jeff Hinton, he's widely considered the godfather of AI. So he was in the middle of it all. 20 years later, Russ has done basically everything. He sold his startup to Apple. 当时 Russ Salakhutdinov 在金融业工作,却偶然在多伦多街头碰到了 Geoffrey Hinton。 Hinton 把他拉进办公室;等他出来时,他已经决定去攻读“深度学习”的博士,而当时甚至还没有 deep learning 这个说法。
He worked on their secret self-driving car project, Titan. He became a professor at Carnegie Melon. And he just left Meta Super Intelligence Lab to start soothe labs, which just raised $50 million to build AI that forecasts the future. This episode, Russ tells me that the frontier labs are all building commodities. How the people actually building AGI don't think that it's 2 years away. 如果你不熟悉 Geoffrey Hinton,他被广泛视为 AI 教父,所以 Russ 从一开始就在这场变革的中心。 20 年后,Russ 几乎什么都做过:把创业公司卖给 Apple,参与其秘密自动驾驶汽车项目 Titan,成为 Carnegie Mellon 教授,之后离开 Meta Superintelligence Labs,创办了刚融资 5,000 万美元、致力于构建预测未来之 AI 的 soothe labs(英文自动字幕拼写)。
Explains why Curser and half the startups you know are actually running on Chinese open source models. How Yang Zillian, the CEO of Kimmy, was one of his smartest PhD students. Why the only useful robot in the world is still the Roomba. and why AI should replace Mackenzie Bane and BCG. Let's get into it. So you've been at the center of modern AI since you know basically the beginning. 本期节目里,Russ 会谈到为什么前沿实验室构建的东西最终都会商品化,为什么真正从事 AGI 的人并不认为它会在两年内到来,为什么 Cursor 和你熟悉的一半创业公司。
So I want to walk through that journey a little bit uh before we get into some of the stuff that you're working on now. So first of all like why did you decide to do a PhD in machine learning in the mid200s like where did you see things headed? You know it's a very good question. So my story is kind of a little bit unusual because I was very interested in AI. 其实在运行中国开源模型,Kimi 的 CEO 杨植麟为什么是他最聪明的博士生之一,以及为什么唯一有用的机器人。仍然是 Roomba,以及为什么 AI 应当取代 McKinsey、Bain 和 BCG。
So I started my master's degree at the University of Toronto in early 2000s. Um I did my masters and then I kind of left and and went into banking. So I was in financial sector for a year and it's almost like by luck I started doing my PhD. Um at a time I was working with Jeff Hinton. He was you know working with me when I was doing my masters and I didn't know whether I wanted to do PhD or not. 我们开始吧。你从现代 AI 基本诞生之初就处在它的中心,所以在谈你现在做的事情之前,我想先梳理一下你的经历。首先,你为什么会在 2000 年代中期决定攻读 machine learning 博士?你当时觉得这个领域会走向哪里?
But then at one point I was walking down the street and I bumped into Jeff Hinton and he basically said like I have this amazing idea you know these these models these deep learning models you can learn multiple layers of representation and everything and so he essentially dragged me to his office and started showing to me like what he's been working on and I think that was so inspirational to me that I was basically saying can I can I come and do a PhD with you on this topic and he sort of you know he made it happen he's like okay why don't you apply for your PhD program and everything come and work with me and so you know in I started my PhD in 2005 2006 and at that time I think deep learning has been sort of uh you know it just was very early wasn't very clear what it's going to work these neural networks they were very sort of looked down upon uh because at the time people were looking at statistical machine learning sort of theory and you know uh something that's more rigorous and neural networks have always been looked at as like well these are nonlinear systems that people just optimizing without really understanding what's what's going on right and So, but that's when some of the early deep learning models were were were trained and I, you know, did my PhD and then with with Jeff and then moved to MIT to do my pawn dog and and such. 嘉宾:这是个很好的问题。我的经历有些不寻常,因为我一直对 AI 很感兴趣。所以我在 2000 年代初进入 University of Toronto 读硕士。硕士毕业后,我离开学校进了银行业,在金融行业待了一年。后来开始读博士几乎是机缘巧合。我读硕士时曾和 Geoffrey Hinton 合作,由他指导,但我并不知道自己是否想读博。后来有一天,我走在街上,碰巧遇见了 Geoffrey Hinton。他基本上对我说:“我有一个特别了不起的想法,你知道,这些模型、这些 deep learning 模型可以学习多层表征,等等。 ”于是他几乎是把我直接拉进了办公室,开始给我看他一直在研究的东西。我觉得那件事极具启发性,于是我基本就是问他:“我能不能来跟你读博士、研究这个课题? ”他就设法促成了这件事。他说:“好啊,你申请博士项目吧,然后过来跟我一起做。 ”于是我在 2005、2006 年开始读博。当时 deep learning 还处于非常早期,前景并不清楚;这些神经网络很受轻视,因为那时大家更关注 statistical machine learning 的理论以及数学上更严谨的东西。
So, and then the rest is the history basically. Yeah. Yeah. But why why were neural nets like so exciting to you? Like what was he articulating to you that made it seem so interesting? Yeah. Yeah. Yeah. Jeff has like a very good I think he was just his excitement and it convinced me that this is going to be an interesting PhD to do like interesting topic to do my PhD in. 神经网络一直被看成这样一种非线性系统:大家只是在优化,却并不真正理解其中究竟发生了什么,对吧?不过,也正是在那时,一些早期 deep learning 模型被训练了出来。
Yeah. I think it was part of it because like I think that at a time you know people were again neural networks have all like at that time neural networks been looked down upon as like yeah these sort of like systems they've been around for 20 years they never worked. They're like number third choice of any sort of method at the time. you know, methods like support vector machines and more rigorously like more mathematical rigor behind those models. 我跟 Geoffrey 完成博士,之后又去了 MIT 做 postdoc,诸如此类。后来的事情基本就是历史了。主持人:是的。但神经网络究竟为什么让你那么兴奋?他当时描述了什么,让你觉得这件事如此有意思?嘉宾:是啊。 Geoffrey 很擅长… …我觉得就是他的兴奋感染了我,让我相信这会是一个很有意思的博士课题。
But you know, Jeff was one of the first people who showed that learning these stacked systems, learning these deep systems, right, can actually give you so you had he had like a mathematical justification for us, which is a beautiful, you know, beautiful math behind showing that, you know, these models can improve something. It's called variation lowbound. And and I think it it's it was exciting that I could see that this could have a strong potential in the future. 我觉得这确实是一部分原因。因为当时人们一直看不起神经网络,觉得这种系统已经存在 20 年、却从来没奏效过;在各种方法里,它最多只是第三选择。当时像支持向量机这样的方法背后有更严谨的数学基础。但 Geoffrey 是最早证明学习这种堆叠系统、学习这种深层系统确实可以带来收益的人之一。
And what what were the use cases that he was excited about or that you were excited about or was it more just oh the math was beautiful? No, no, no. Yeah. So, so there was the math was beautiful. That's number one because for the first time you could actually say okay there is you know there is like mathematical rigor behind these models. 他给出了一套数学论证;其中有非常漂亮的数学,表明这些模型能够改进某个指标,叫作 variational lower bound。我觉得这很令人兴奋,因为我能看到它未来潜力巨大。主持人:当时让他或让你兴奋的 use case 是什么?还是说纯粹因为数学很漂亮?
It's not like before that people were sort of you know training neural nets. was always like a little bit like you know there is some optimization you do optimization but there's no like understanding there is no sort of connection to statistical machine learning I mean there were some but not not as rigorous right and it was always like this field have always been looked down upon I mean you'd probably ask Yan Lun Yan Lakun was like working on these models for a long time but he could never get recognition in early days for for any of his work because it was just kind of like didn't didn't work right because the scale was small and and you know one of the interesting things that showed up early on is a glimpse into generative models. 嘉宾:不,不。数学很漂亮,这是第一点,因为这是第一次你可以说这些模型背后确实有数学严谨性。此前人们也会训练 neural net,但总有点像是做了一些优化;你确实在做 optimization,却没有真正的理解,也没有与 statistical machine learning 建立联系。其实也不是完全没有,只是没那么严谨。这个领域一直都被轻视。你大概可以问 Yann LeCun,他研究这些模型很久了,但早期始终很难让工作获得认可,因为它们就是没有真正奏效:规模太小,而且早期出现的一件有意思的事,是大家第一次窥见了 generative model。
So at the time we were generating images of handwritten digits uh these M this was in the mid 2000s. It's the mid200s. Yes.
And it was one of the early days that you could show that these deep models something is called deep belief network and then we were working on debols machines were initially the models that can sort of like generate you know uh these nice looking handwritten digits like amnes digits right and it showed that hey you know these models are capable of going beyond multiple level they can learn multiple levels of of representation and that basically excited a lot of people both the math behind Plus that they showed the promise and they were sufficiently different from anything else that people were working in the community and sort of like at a time in 2005 2010 it was like a small group of people who were essentially working on these models making progress on these on these models. 当时我们在生成手写数字图像,也就是 MNIST 数字。主持人:这是 2000 年代中期?嘉宾:对,2000 年代中期。那是很早的时候,我们可以展示这些 deep model——有一种叫 deep belief network,后来我们还研究了 Boltzmann machine——最初这些模型能够生成看起来相当不错的 MNIST 手写数字。它说明这些模型能够超越单层,学习多层表征。背后的数学,加上它们展现出的潜力,以及它们与当时社区中其他研究足够不同,令很多人感到兴奋。当时,也就是大约 2005到 2010 年间,实际上只有一小群人在研究并推进这些模型。
And when did you finish your PhD? Finished my PhD in 2009. Okay. And then you went to MIT to do post to do my postdoc. Yes. I was doing my post dog u with Josh Tenbomb. fantastic professor and fantastic colleague of mine. Josh was sort of looking at deep learning models but he was also looking on these sort of non-parametric basin models. These are sort of classes of models. 主持人:你什么时候博士毕业?嘉宾:2009 年。主持人:然后你去了 MIT 做… …嘉宾:做 postdoc,对。我的合作导师是 Josh Tenenbaum,他是一位非常出色的教授和同事。
I mean Josh is a cognitive neuroscient cognitive scientist and so um it was great hanging out and at MIT MIT also had a lot of strong computer vision people and such. So was a good experience and then I came back to Toronto as a faculty. Josh 会研究 deep learning 模型,但也会研究non-parametric Bayesian model 这一类模型。 Josh 本身是 cognitive scientist,所以在 MIT 和他相处、共事很棒。
Yeah. So when so it seems like yeah you were very excited about like the generating digits like that was very interesting but then it seems like when Alex net came out that everyone's focus basically shifted to classifying images rather than generating so what's what's so so what's happened early on is that if you take like at a time we didn't have GPUs if you sort of train these multi-layer neural networks on something like Amnest there's like there's a class of data sets like you know these old toy data sets C far you could never beat sort of competing methods like support factor machines and these generative models showed the promise that if you actually learn these generative models and then convert them into discriminative models like classification systems they would actually they would actually work much better. MIT 还有很多很强的 computer vision 研究者,那段经历很好。之后我回 Toronto 当了 faculty。主持人:听起来你当时对生成数字很兴奋,但 AlexNet 出现后,所有人的重心似乎基本都转向了图像分类,而不是生成。嘉宾:事情是这样的:早期我们没有 GPU。如果你在 MNIST、CIFAR 这类旧的小型数据集上训练多层神经网络,始终无法击败支持向量机之类的竞争方法。而那些 generative model 展示出一种可能性:如果先学习生成模型,再把它们转换成 discriminative model,也就是分类系统,效果会好得多。
So this this notion is that you know learning how to generate images basically gives us the representations that we need to actually do classification and then after imageet competition after alexnet people sort of realized well convolutional neural networks actually do work they work you know they work remarkably well and after alexnet people actually shifted to these discriminative models because that was you know they they've solved I think Alex net solved one of the key vision problems right which was the you know it was imageet competition and it's was like it was a million images very diverse images and people were looking at it over the over the past you know two to three years like gigantic labs at Berkeley Oxford you know MIT they had like you know lots of you know PhDs and computer vision people were participating and I think Alexet was the first model that showed that you can improve performance by like 50% So you can imagine Alex Keski, Izzkerver and Jeff Hinton, three people they've built, you know, fairly simple connet,but they scaled it and they were the first ones to scale it on GPUs and they've shown that this model was able to improve over whatever people have done by 50%. And they've able to do it to the point where there was a lot of skepticism initially like people couldn't believe it, you know, like it's it's remarkable, right? 核心想法是,学会如何生成图像,会给我们真正做分类所需的表征。 ImageNet 竞赛和 AlexNet 出现后,人们意识到 convolutional neural network 的确有效,而且效果极其突出。于是大家转向这些 discriminative model,因为它们解决了问题。我认为 AlexNet 解决了 vision 的一个关键难题,对吧?主持人:也就是… …嘉宾:也就是 ImageNet 竞赛。那是一百万张非常多样的图像,过去两三年间,Berkeley、Oxford、MIT 的大型实验室投入很多博士生和 computer vision 研究者参与。我认为 AlexNet 是第一个把性能提高大约 50% 的模型。你可以想象,Alex Krizhevsky、Ilya Sutskever 和 Geoffrey Hinton 三个人构建了一个相当简单的 ConvNet,但他们把规模做大了,而且率先把它扩展到 GPU 上,最终比此前所有人的成果提升了 50%。
When you see the result, you say that cannot be true. It's too good to be true. So it took people like couple of months to actually verify and say, "Okay, this is real." Yeah. So you were So you were there when Alex came out. I was there when Alexand came out. That's right. I I came from my pawns dog back to Toronto and Yeah. It was Were you skeptical? No, I wasn't skeptical of of of the results. I knew like Jeff was kind of like on it. 他们做出的效果太好,以至于一开始有很多质疑,人们不敢相信。看到结果时你会说:“这不可能,简直好得不真实。 ”于是大家花了几个月去验证,最后才确认:“好,这是真的。 ”主持人:是啊。所以 Alex… Net 出现时,你也在那里。嘉宾:AlexNet 出现时我确实在那里。
So it's it's like it's you know one of the things that Alex Keski has done very well is he optimized some of the computations on early versions of GPUs which which required like a very strong engineering talent to do it because scale that was one of the earlier kind of like points where the scale actually showed big big success. 那时我结束 postdoc,回到了 Toronto。主持人:你怀疑过吗?嘉宾:不,我没有怀疑结果。我知道 Geoffrey 一直在推动它。 Alex Krizhevsky 做得特别好的一件事,是针对早期 GPU 优化了一些计算。
Interesting. Um so that yeah that was the turning point for computer vision that was the turning point for computer vision researchers like from that point on Google all the other major labs and academics started basically converting to uh to deep learning. So what were you focused on before like what was your research focused on before Alex net came and then did your research shift when it came out and you were got really excited by it? 这需要极强的工程能力,因为 scale 是其中一个最早显示出巨大成功的关键点。主持人:有意思。所以它成了 computer vision 的转折点?嘉宾:对,是 computer vision 研究者的转折点。从那以后,Google、其他主要实验室以及学术界基本都开始转向 deep learning。
So my research was kind of like I never looked at I mean I was looking at vision and generative models. I think my research you know I wasn't part of the Alex net because it was like a fairly strong engineering effort. My research kind of like started shifting after Alex net because one of the things that my lab has been looking at at Toronto also is earlier versions of caption generation. 主持人:AlexNet 之前你的研究重点是什么?它出现、让你十分兴奋以后,你的研究方向有变化吗?嘉宾:我并不是只研究某一个方向,我当时在研究 vision 和 generative model。我没有参与 AlexNet,因为那是一项很强的工程工作。
So once we started seeing that these models actually pretty good then we've started looking at models that given images generate descriptions for those images. AlexNet 后,我的研究方向开始变化,因为我在 Toronto 的实验室还研究早期的 caption generation。
This was one of the early I remember uh talking to some of the computer vision folks and then at a time classification detection was the thing to do but a lot of computer vision people were saying that well I'll actually I want a system where I show you the image and I just don't want to say well there's a car in this image over here and then there is like you know a person standing here. I actually want the system to describe me the scene. 所以,当我们开始看到这些模型已经相当不错后,就开始研究:给定图像,让模型为图像生成描述。我记得当时和一些 computer vision 研究者交流,那时 classification 和 detection 才是主流,但很多人说,我真正想要的是这样一个系统:给它看一张图,我不只想让它说这里有一辆车、那里站着一个人。
I want the system to look in and just like human tell me like oh that's you know it's a sunset you have a beautiful beach here you have people walking on the beach and sort of like just tell me what's like that's the you know a a way of much better understanding than if me running a classifier say is there a person there or not yes or no is there a car there yes yes or no right which is what the field was and so like early on these recurrent neural networks coupled with vision systems started showing progress and this is what you know we were looking at I had never even thought about that as a thing until like LLM came out. 我真正希望系统能描述场景。我希望系统看过之后像人一样告诉我:“这是日落,这里有一片漂亮的海滩,有人在沙滩上散步。 ”也就是直接告诉我里面有什么。相比运行一个 classifier,分别问“有没有人”“有没有车”并得到 yes/no,这体现了更深入的理解,而后者就是当时这个领域的做法。因此早期的 recurrent neural network 与vision system 结合后开始显示进展,这就是我们当时在研究的东西。