返回目录

Why Fei-Fei Li Is Betting on Spatial Intelligence

a16z 请李飞飞、Justin 和 Ben 讲 Atlas:同一模型里同时做生成和重建,原语是 new view prediction,不是再做一个会出片的视频模型

这一期在说什么

他们把空间智能收成一个可 scale 的原语:先预测新视点,再让生成去补重建永远拍不到的洞

本文综合:节目前半把 Atlas 从视频模型里拆出来——每帧带着相机位姿进入 spatial context,同一套权重既做精确重建,也做生成式补洞。后半把同一原语接到机器人:稠密重建卡住了 real-to-sim 的速度,稀疏输入加神经仿真才可能给策略足够的随机化。两条线连起来,Atlas 不是更好看的文生视频,而是把“看见、体验、交互”闭环的第一步;4D 和工业级可编辑性被明确标成还没做完的下一截。

一句话

World Labs 的 Atlas 把生成和重建做进同一个 world model,核心原语是 new view prediction:几张带位姿的图就能重建、飞穿和仿真,采集量比稠密重建少约 50 到 100 倍。

Atlas · spatial intelligence · new view prediction · World Labs · Gaussian splat · robotics simulation

Insight

新的不是“又能出视频了”,而是视点成为模型的原生输入。

  1. 最多约 100 帧真实视角可以重建空间,结果可以是飞穿视频,也可以是显式三维(约 00:01–00:02)。
  2. 李飞飞把像素生成和像素重建的统一,放到计算机视觉半个多世纪的分赛道背景里:会议里生成、识别、三维重建一直分开(约 00:07–00:08)。

主持人 On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. 在通往空间智能的路上,生成真正具有空间上下文、并且落地的像素,这正是 Atlas 已经迈出的那一步,也是最难的一步。

嘉宾 You know, LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. 你知道,LLM 建立在 next token prediction 之上。我们看到视频模型建立在 next frame prediction 之上。Atlas 真正做的是 new view prediction。

主持人 This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100 extra reduction. There's a famous shot in the first Matrix movie where Neo is like falling down. Exactly. 这才是 AI 真正能为人和他们的工作流程解锁大量价值的地方。我们说的是大约 50 倍、100 倍的额外缩减。第一部《黑客帝国》里有一个著名镜头,Neo 好像在往下掉。没错。

主持人 They had hundreds of cameras doing that angle on a green screen. With Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. 他们当时用了几百台相机,在绿幕前拍那个角度。有了 Atlas,我们只用三台相机就能做到。不需要棚拍,不需要绿幕,也不需要昂贵的标定。

嘉宾 No one has ever seen these results. 从来没有人见过这样的结果。

主持人 When you set out to do this, did you know it was going to work? 你们开始做这件事的时候,知道它会成功吗?

嘉宾 I was pretty sure. Each time we made the model bigger, and each time we trained it for longer, it got significantly better. 我相当有把握。每一次我们把模型做大,每一次我们训练得更久,它都会明显变好。

主持人 Does that mean we're going to get 4D video? Can I go walk around? 那是不是意味着我们会得到 4D 视频?我能走进去走一圈吗?

嘉宾 So, big day yesterday, you launched a new frontier model, which has got amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So, maybe Justin, do you want to talk about what was launched yesterday, why it's significant. 昨天是个大日子,你们发布了一个新的前沿模型,反响非常好,还在持续。进来。我觉得组织这次对话的一个好办法是,先准确讲讲昨天发布的是什么,然后再回到历史,一路往上讲。Justin,你要不要讲讲昨天发布了什么,以及它为什么重要。

主持人 Yeah, so Atlas is our new next generation world model. Um, it has three basic things. It can generate, reconstruct, and simulate the world. Um, so within that, there's a couple different major capabilities. It has really good camera condition generation. So, you can input an image together with a camera trajectory with a camera trajectory and steer the model and have it generate, you know, video frames along any any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to 100 frames, um, that are views of the real world, and use those to reconstruct the real world. And that reconstruction can take the case either of of novel of video flying through the space, or an explicit 3D reconstruction of the space. Um, then finally, it can be used for simulation. Um, and for this, we show off, um, you know, these awesome bullet time videos, which got a lot of attention online, and then also robotics simulation. 好。Atlas 是我们新一代的 world model。它有三件基本能力:生成世界、重建世界、模拟世界。在这之下,还有几项主要能力。它的相机条件生成非常强。你可以输入一张图,再配上一条相机轨迹,用这条相机轨迹来引导模型,让它生成视频帧沿着你想要的任意视角。它在稀疏三维重建上也很强。你可以输入一帧,或多帧,最多 100 帧真实世界的视角,用它们来重建真实世界。重建的结果可以是一段在空间里飞过的视频,也可以是这个空间的显式三维重建。最后,它还可以用于仿真。在这方面,我们展示了那些很酷的子弹时间视频,网上关注度很高,另外还有机器人仿真。

嘉宾 What's a What's a bullet time video? 什么是子弹时间视频?

主持人 A bullet time video, this comes from the Matrix. You know, there there's a famous shot in the first Matrix movie where Neo is like falling down. 子弹时间视频来自《黑客帝国》。第一部《黑客帝国》里有一个著名镜头,Neo 好像在往下掉。

嘉宾 Oh, yeah. 对,对。

主持人 Exactly. So, and then remember in that famous shot he's like falling down, it's in slow motion, and the camera flies all the way around. Um so, that's the way that they did that shot is they had a ring of like hundreds of cameras. So, then like he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that that famous shot in in Matrix. But now with Atlas we can do this with just as few as through three cameras. So, like no studio capture, no green screen, no expensive calibration. We can literally stick like three cameras on three iPhones on tripods, um use these to take sort of a video of something happening, um like someone shooting a basket, someone dropping a strawberry into a bowl of milk. And then from those three like iPhone videos, we can then reframe the shot, and imagine like a like freeze time, have the camera fly in like as the milk is splashing up, 没错。然后你记得那个著名镜头里,他往下掉,是慢动作,相机一路绕着他飞。他们拍那个镜头的办法是,摆了一圈大概几百台相机。他在棚里倒下,几百台相机在绿幕前拍那个角度,然后他们用那几百台相机做出了那个著名镜头,就是《黑客帝国》里的那个。但现在有了 Atlas,我们最少只用三台相机就能做到。不需要棚拍,不需要绿幕,也不需要昂贵的标定。我们真的可以把三台相机、三部 iPhone 架在三脚架上,用它们拍一段正在发生的事,比如有人投篮,有人把一颗草莓丢进一碗牛奶里。然后从这三段 iPhone 视频,我们可以重新构图,想象时间冻结,让相机飞进去,就在牛奶溅起来的时候,

主持人 and get these amazing frozen time views. Um and we can do this with just uh just a couple cameras. 得到这些非常惊人的冻结时间视角。而且我们只用几台相机就能做到。

嘉宾 Can Can you um just maybe What is the simplest description of what Atlas does? Like what goes in and what comes out? 你能不能用最简单的话说一下 Atlas 是做什么的?进去的是什么,出来的是什么?

主持人 Yeah, so one of the one of the really core principles of Atlas, like the most fundamental thing, is it does new view prediction. Um and this is a a really fundamental primitive that we think is super exciting, a super new primitive for for base models that no one's ever done before. Right? So, we know LLMs are built on next token prediction, we've seen video models as being built on next frame prediction. Atlas is really new view prediction. Right? That given some number of views of a scene or a description of a scene, um those go into what we call a spatial context that describes implicitly what is the world that we want to talk about, then you can point a virtual camera at at an arbitrary point in space and time, and Atlas will understand what that world is supposed to look like from that position in space and time. 好。Atlas 真正核心的原则,最根本的一件事,就是它做 new view prediction。我们认为这是一个非常根本的原语,对基础模型来说是一个全新的原语,以前没有人做过。对吧?我们知道 LLM 建立在 next token prediction 之上,我们也看到视频模型建立在 next frame prediction 之上。Atlas 真正做的是 new view prediction。对吧?给定一个场景的若干视角,或者对这个场景的一段描述,它们会进入我们所说的 spatial context,隐式地描述我们要讨论的那个世界是什么。然后你可以把一台虚拟相机指向时空中的任意一点,Atlas 就会理解,从这个时空位置看,那个世界应该是什么样子。

嘉宾 You know, you know, Ben, that you know, with with a with a bajillion video models out there all claiming to be world models and all claiming to have novel views, and can you maybe tease apart can more concretely how this is different from like the myriad models that have come before? Ben,你知道,外面有无数的视频。模型,都声称自己是 world model,都声称能做新视角,你能不能更具体地把这个和以前那一大堆模型的差别拆开讲清楚?

主持人 Mhm. Yeah, I think what Justin was saying about the spatial context aspect is super important here. So there's many video models, a lot of video models actually got their claim to fame from their single dimension input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omni referencing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of like specially grounded meaning to every frame you put into it. So it's not just an image that the model's going to interpret whatever way it wants or you can kind of try to argue with it in the the text prompting and get it to do something specific. With Atlas, every image actually has an associated three-dimensional camera pose and that means that you can perform this task of reconstruction with an extremely high degree of precision, right? So if we had 嗯。对,我觉得 Justin 刚才说的 spatial context 这一点在这里非常重要。有很多视频模型,不少视频模型成名靠的是单维输入,或者从首帧到末帧的插值。现在我们开始看到一些模型能够做这种全方位参考,用 20 张、30 张、50 张图。但 Atlas 的关键在于,你输入的每一帧实际上都有一种空间上落地的含义。所以它不只是一张图,让模型随便怎么理解,或者你靠文本提示去跟它争、让它做某件具体的事。在 Atlas 里,每张图都对应一个三维的相机位姿,这意味着你可以用极高的精度来做重建这件事。对吧?所以如果我们有

主持人 four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room and it's not going to guess what's in the other corner, like the relationship between things. It's just going to reproduce exactly what you give it. And you can also do that in a kind of creative or imaginative sense, too. If you take two photos from different, you know, AI generations or real-world locations, you can actually position and stage those to build these kind of intentionally directed fly-throughs that are really governed by exactly the precise place that you put the content you want and where the camera's going to look and travel, which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models. 这个房间的四个视角,每个角落一张,你可以把它们放进模型,然后得到你在这个房间里看到的一切的精确复制。它不会去猜另一个角落里有什么,比如物体之间的关系。它就是精确复现你给它的东西。你也可以用一种创造性、想象性的方式来做这件事。如果你从不同的 AI 生成结果,或者真实世界的地点,各拿两张照片,你实际上可以摆放和布置它们,来构建这种有意安排的定向飞穿镜头,真正由你把想要的内容放在精确位置、以及相机要看向哪里、沿哪里走来决定。这和视频模型那种更像老虎机的效果很不一样,那种只能靠更高层的文本控制,一遍遍重试生成。

嘉宾 Is this just kind of an obvious, you know, scaled-up version of a traditional video model or is it a new architecture? 这只是传统视频模型一个显而易见的放大版,还是一种新架构?

主持人 I think it's it's a pretty new thing for a couple of different reasons. One that we talk about is it does both generation and reconstruction jointly in the same model. Like Ben was saying, this thing can take a couple of views of this room and then reconstruct everything in this room exactly as you see it. And historically, reconstruction has been its own subfield in computer vision with its own specialized task, its own specialized models. And generation is what all the text all the text to video models are really good at like what all the all the big diffusion models we've seen the last couple years. And those are great for creative applications. I want to imagine 我觉得这是一件相当新的事,原因有。几个不同的方面。我们常说的一点是,它在同一个模型里同时做生成和重建。就像 Ben 说的,它可以拿这个房间的几个视角,然后把这个房间里的一切按你看到的样子精确重建出来。历史上,重建一直是计算机视觉里一个独立的子领域,有自己专门的任务,有自己专门的模型。而生成是所有文生视频模型真正擅长的事,也就是我们过去几年看到的那些大型 diffusion 模型。那些对创意应用很好。我想想象

主持人 something that's never been there before. Um but now with Atlas for the first time we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture. So to do that we had to make a couple changes. Um one is we had to make it multimodal from the start. So this thing natively works on text, it works on images, it works on videos. It also works on camera poses um as a native input to the model which I don't think anyone's ever done at the pre-training phase before. Um and it uses um it uses 3D as a native modality that it works on. So this thing from the beginning was designed to be natively multimodal in a way that no one else I think 以前从未存在过的东西。但现在有了 Atlas,我们第一次把视觉智能的这两个不同部分放进同一个模型。所以它可以在同一套架构里同时做三维重建和生成。要做到这一点,我们做了几处改动。一是我们必须从一开始就做成多模态。所以它原生处理文本,处理图像,处理视频。它也把相机位姿作为模型的原生输入,我认为在预训练阶段以前没有人这样做过。而且它把 3D 作为自己处理的一种原生模态。所以这个东西从一开始就被设计成原生多模态,我认为没有别人是这样

嘉宾 Sorry, I just um I don't know this space super well but 3D is this like depth or models and like what is that what is that mean? 抱歉,我对这个领域不是特别熟,3D 是指深度,还是模型,这到底是什么意思?

主持人 Yeah, so the formulation we used so far is depth maps. 对,我们目前用的表述是深度图。

嘉宾 Okay. 好。

主持人 Right? So um right now you can have a when you have a frame that has a a a virtual camera telling its position in 3D space, yeah, that camera position and camera parameters are a native input to the model. And then what attached to that camera position you can have both um like RGB telling you what is that position in space look like and you can have a depth map that tells you what is the spatial structure of that position in 3D space. So then you know text, image, video, 3D cameras are these modalities that this thing all does jointly in a multimodal way. 对。所以现在,当你有一帧,上面有一台虚拟相机告诉你它在三维空间里的位置,对,那个相机位置和相机参数是模型的原生输入。然后附着在这个相机位置上,你可以既有 RGB,告诉你那个空间位置看起来是什么样,也可以有一张深度图,告诉你那个位置在三维空间里的空间结构。所以文本、图像、视频、三维相机,这些模态这个模型都是以多模态的方式一起处理的。

嘉宾 I want to add something cuz I think what Justin just said is actually so important and also what Ben said that it's under appreciated. It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century. Um sitting here having been in this field for decades, I cannot tell you how many PhD thesis have been written on the 我想补充一点,因为我觉得。 Justin 刚才说的其实非常重要,Ben 说的那一点也被低估了。这是我们第一次把像素生成和像素重建统一起来。在计算机视觉这个领域,已经存在了半个多世纪。我坐在这里,在这个领域待了几十年,我说不清有多少篇博士论文写的就是

嘉宾 problem of reconstruction or novel view things synthesis and also our field traditionally uh have multiple tracks. You go to a computer vision conference, you have the pixel generation track, you have some recognition track and you have 3D reconstruction track. This is a elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and the the viewpoint estimation and that's just incredibly powerful. 重建问题,或者新视角合成。而且我们这个领域传统上有多条赛道。你去一个计算机视觉会议,会有像素生成赛道,会有识别赛道,还会有三维重建赛道。这是一个优雅的模型,它把重建和生成结合起来、统一起来,锚点是视点和视点估计,这非常强大。

Insight

产品迭代的真正决定,是输出模态不再绑死在一种三维表示上。

  1. 生成像素是早期一步,已经有无数模型做过;生成真正具有空间上下文、落地的像素,才是 Atlas 走的难步(约 00:10–00:11)。
  2. Marble 输出 Gaussian splat,方便移动端、VR 和游戏引擎,但也成了瓶颈。Atlas 需要时仍可出 splat,不必处处经过这一关(约 00:11–00:13)。

主持人 Can Can you Can you maybe Well, can we take a step back and then maybe you just fill something out? So, when when you started the company, I remember you saying uh you know, you want to tackle um spatial intelligence, right? And uh you know, now we have this new model. And so, it feel I mean, like as a layperson, it feels very general to me. You've got next product prediction and this is next new view prediction. 我们能不能先退一步,你再补充一点?你们创办公司的时候,我记得你说过,你们想攻克空间智能,对吧?现在我们有了这个新模型。所以,作为一个外行,我觉得它非常通用。你们有 next product prediction,而这是 next new view prediction。

嘉宾 New view prediction, right? 是 new view prediction,对。

主持人 So, like you can get one view out of a set of views and you have a new view. Can maybe you pencil out like like how this is a significant step to this general problem of spatial intelligence? And maybe by like starting to describe what spatial intelligence is? 所以你可以从一组视角里得到一个视角,你就有了一个新视角。你能不能大致画一下,这如何构成通往空间智能这个一般问题的重要一步?也许可以从描述空间智能是什么开始?

嘉宾 Well, spatial intelligence eventually must enable us to both generate what the space is, reason within it, and being able to edit and interact within it. 空间智能最终必须让我们既能生成这个空间是什么,能在其中推理,。也能在其中编辑和交互。

主持人 Yeah. 对。

嘉宾 Now, we can argue is it 3D or 4D? Ultimately, it's 4D with the time dimension, but even just 3D, these are the fundamental tasks that one has to do or spatial intelligence has to enable. And then we talk about with that, you can render, you can simulate, and you can plan actions. But to do that, a fundamental problem to solve is to understand the geometry and structure and the physics of the space. 现在我们可以争论这是 3D 还是 4D。最终它是带时间维度的 4D,但哪怕只是 3D,这些也是必须完成的基本任务,或者说空间智能必须使能的任务。然后我们说,有了这些,你可以渲染,可以仿真,可以规划动作。但要做到这些,必须解决的一个根本问题是理解这个空间的几何、结构和物理。

主持人 Yeah. 对。

嘉宾 And I do believe Atlas is a significant step forward because now with every single frame, you have a you can generate an estimate a important piece of information, which is the the view viewpoint, the camera pose. And that is the most critical information one needs about the geometry of the of the space. And that can lead to all the emergent behaviors we see in our in the downstream of the model, which we showed in the blog. So in the on the path to spatial intelligence, generating pixels is definitely a a early step, which we have seen with what you call it gazillions of models. But generating pixels that are truly 我确实认为 Atlas 是重要的一步,因为现在每一帧,你都可以生成、估计出一条重要信息,也就是视点、相机位姿。而这是关于这个空间的几何,人最需要的关键信息。而这可以带来我们在模型下游看到的那些涌现行为,我们在博客里展示过。所以在通往空间智能的路上,生成像素当然是早期的一步,我们已经在你所说的无数模型里见过。但生成真正

嘉宾 spatially contextualized and grounded is absolutely another major step. And that is the very hard step that Atlas has taken. I we definitely have, you know, we can just keep going here, right? Like there is the fourth dimension of time, which will bring in dynamics. And there is more higher fidelity simulation and delineation of the space. So this is part of the the road map of spatial intelligence. 具有空间上下文、并且落地的像素,绝对是另一大步。而这正是 Atlas 已经迈出的那一步,也是最难的一步。我们当然还可以继续往前走,对吧?还有时间这第四个维度,它会带入动态。还有更高保真的仿真,以及对空间更细致的刻画。所以这是空间智能路线图的一部分。很好。我当然想深入讲讲接下来往哪走。但先说说是怎么走到这一步的。World Labs 成立多久了?

主持人 Great, yeah. I mean, I definitely want to like dig into like where this is going. But first, maybe let's talk about getting here. How long has World Models been in existence? 两年半。两年半,对。

嘉宾 Two and a half Two and a half, yeah. 你们其实以前也发布过模型。那为什么不直接跳到 Atlas?

主持人 And so you you've actually released models before. So what why didn't you just jump right to Atlas? 好问题。

嘉宾 Good question. 这太神奇了,对吧。

主持人 It's so magic, right? Yeah. Justin 的团队需要很多芯片。

嘉宾 Justin's team needs a lot of chips. 对,我们需要很多 GPU 才能真正。

主持人 Yeah, we need a lot of GPUs to actually scale this thing up. So what one of the last year we released our Marble World model, and that was the first kind of big major World model that we put out. That that powers our current Marble product. And Marble and Marble is really cool. Marble can take images, it can take videos, it can take text prompts, and use these to generate 3D worlds. But one of the biggest differences between Marble and Atlas is exactly what is that output modality. So, marble was really focused on Gaussian splats as an output representation. So, whatever you're inputting, um it's going to output a 3D world represented as a as a Gaussian 把这个东西做大。所以去年我们发布了 Marble 这个 world model,那是我们推出的第一个大型、主要的 world model。它驱动着我们现在的 Marble 产品。Marble 非常酷。Marble 可以接收图像,可以接收视频,可以接收文本提示,用这些来生成三维世界。但 Marble 和 Atlas 最大的差别之一,恰恰是输出模态是什么。Marble 当时非常聚焦把 Gaussian splat 作为输出表示。所以无论你输入什么,它都会输出一个用 Gaussian splat 表示的三维世界。

主持人 splat. And Gaussian splats are really useful, right? They're really nice, they're easy to render, they are they can they can render efficiently on mobile devices, on VR devices, they can interoperate with other with a with game engines, with simulation engines. There's a lot of nice things about Gaussian splats. But, you know, I think that was kind of a bottleneck in the previous marble model. So, what we did with Atlas is redesign the thing bit. Um and we realized that we need to bifurcate these modalities earlier and actually have these things all these modalities working in a more unified way in the model. So, now with Atlas, um the the fundamental primitive is not like a good generate a Gaussian splat world. The fundamental primitive is is as we said new view prediction. Um and that can generate RGB frames, that can generate 3D, and we can use those to generate a beautiful Gaussian splat worlds when you need them. Um but, we don't need to bottleneck our our our outputs through the Gaussian splats when we don't need to. And that was a that actually took a lot of, you know, blood, sweat, and tears to understand like what are all the pros and cons of these Gaussian splat 确实很有用,对吧?它们很好,容易渲染,可以在移动设备上、VR 设备上高效渲染,可以和其他东西互通,和游戏引擎、仿真引擎互通。Gaussian splat 有很多优点。但我觉得那也是之前 Marble 模型的一种瓶颈。所以我们做 Atlas 时,把这件事重新设计了一遍。我们意识到需要更早地把这些模态分开,同时让所有这些模态以更统一的方式在模型里工作。所以现在有了 Atlas,根本原语不再是生成一个好的 Gaussian splat 世界。根本原语就是我们说的 new view prediction。它可以生成 RGB 帧,可以生成 3D,需要的时候我们也可以用这些去生成漂亮的 Gaussian splat 世界。但当我们不需要的时候,就不必把输出卡在 Gaussian splat 这一关上。而这实际上花了大量心血,才弄清楚这些

主持人 different representations. So, that's one part of it. Um the other part is you got to like climb the scaling ladder, right? You got to like work your way up and like do smaller experiments, do smaller models like to build your conviction on what's going to work and what's going to scale. Um and there's you know, if you could instantly know the right thing that's going to scale, you know, you should just do that. But, when we when we started the company, the world was a very different place. Like there's no scaling law of special intelligence. Right. So, like when we started the company, like the world was in a very different place, the tech was in a very different place. We had a lot of ambitions for where we wanted it to go, but it took a couple it took a a couple iterations for us to hit upon this formulation that we thought is actually is actually like this is the one. This is this is the one that can scale out. 不同表示各自的利弊。所以这是一部分。另一部分是,你必须爬 scaling 这把梯子,对吧?你必须一步步往上走,做更小的实验,做更小的模型,来建立你对什么能行、什么能 scale 的信心。如果你能立刻知道什么东西能 scale,那你就直接做那个就好。但我们创办公司的时候 ,世界是完全不同的。空间智能并没有一条 scaling law。对。所以我们创办公司的时候,世界是完全不同的,技术也是完全不同的。我们对它要去的方向有很多雄心,但我们花了几次迭代,才碰到这个我们觉得对了的表述:就是这个。就是这个能够 scale 出去。

把 300 张图降到 3 张

00:13:00–00:21:00

Insight

稀疏不是偷懒,是承认永远拍不全,必须让生成补洞。

  1. Stanford 四方院:输入全是地面照片,展示却是航拍。看到的是生成结果,但按重建法则生成(约 00:17–00:18)。
  2. Marble 塞不进超过几张图;Atlas 可以把 2000 张图的房子降到 30、40 张输入,飞穿看起来基本一样。重建被说成带超长 context 的生成(约 00:19–00:20)。

嘉宾 You know, Ben, you know, being the creator of Nerf and doing a lot of 3D and reconstruction, so it's not so obvious to me that like if you have multiple views that you actually end up with a 3D thing. But, like you've kind of like made a career of ending up with a 3D thing. So, maybe talk a little bit about like kind of that step. Ben,你是 NeRF 的创造者,做了大量三维和重建工作。对我来说,并不是那么显而易见:如果你有多个视角,最后就会得到一个三维的东西。但你几乎把职业生涯都花在最后得到一个三维的东西上。所以也许讲讲那一步。对。就像你说的,我职业生涯里很多很多年,实际上职业生涯的绝大部分,都在做从图像生成三维的东西。这其实也是我们公司早期经常讨论的:这会不会就是产出三维的那条路?我们是先合成多个视角,再从中构建三维?还是直接走向三维?这个领域里一直有很多不确定性,这些路径里哪一条会胜出,或者

主持人 Yeah. Yeah. Um yeah, I mean as you said, I've spent many, many years of my career as a mass majority of my career actually working on producing 3D things from images. Um and this is actually something we we talked about a lot early on in the company even of like is this going to be the approach that produces 3D, right? Are we going to synthesize multiple views and then build 3D out of that? Are we going to try to go direct to 3D? Like there's been a lot of um uncertainty in the field around like which of those approaches kind of will win out or will kind of like reap the best advantages earlier on. Um but I did have a lot of conviction just from seeing the kind of power of what I would almost call the brute force scaling scaling at a very, very, very small baby scale, not like real model scaling, but the scaling of dense reconstruction that we had seen happening over the past 3 years before. Um so, basically we put up 会更早拿到最大的好处。但我确实有很强的信心,只是因为看到了一种力量,我几乎想称之为蛮力 scaling,而且是非常、非常、非常小的婴儿级 scale,不是真正的模型 scaling,而是我们在过去三年里看到的稠密重建的 scaling。所以基本上我们提出

嘉宾 What Why is dense reconstruction Yeah. dense? 为什么叫稠密重建?稠密是什么意思?

主持人 Yeah, dense. 对,稠密。

嘉宾 Dense dense is 稠密就是

主持人 Because I know we're going to talk about sparse and I want to make sure that people understand what is dense and what is sparse. 因为我知道我们接下来会讲稀疏,我想确保大家理解什么是稠密,什么是。稀疏。

嘉宾 Yeah, so I think this is actually even on the kind of like business and commercial side I I think been one of the challenges of productizing uh 3D reconstruction technology like at a fundamental level, right? People kind of don't uh in in a in a casual sense like you think I took three photos of this object or I took six photos of this room. Like I I look at the photos, I can understand in my mind like how this piece together. I can kind of fill in the gaps and get it,but there's just never been really any kind of reconciliation between those like really data-driven priors and then the kind of brute force dense reconstruction, which it actually is much more akin to almost like scientific or medical imaging what we did in dense reconstruction, right? You basically have to say every single thing I want to appear in this reconstruction, I need at least three or four views of it. And if you think about that, like even just in this room, right? There's like under the microphone, under the table, between every different crack and crevice and 对,我觉得即使在商业和产品侧,这也一直是把三维重建技术产品化的挑战之一,而且是根本层面上的。对吧?人们在日常意义上会想:我给这个物体拍了三张照片,或者给这个房间拍了六张。我看着这些照片,脑子里能理解它们怎么拼在一起。我能把缺口补上,能看懂。但那些真正数据驱动的先验,和那种蛮力稠密重建之间,从来没有真正对上过。我们在稠密重建里做的事,其实更接近科学成像或医学成像。对吧?你基本上必须说:每一个我想出现在这次重建里的东西,我都至少需要三到四个视角。你想一下,哪怕只是这个房间,麦克风下面,桌子下面,每一条缝、每一个角落,还有

嘉宾 the plant leaves, right? To actually truly get a picture that covers every one of those spots, it's this like very tedious and exhaustive effort to walk around the room. I think you've all seen me running around various places like capturing them. It takes, you know, for someone who's well-trained per se, like it can take minutes, but if you hand a casual consumer or even some like kind of professional trying to do this for the first time, an average like cell phone camera or capture device, like it's going to take them probably an hour. I've seen someone for the first time trying to scan a multi-room environment spend like two hours walking through it and get enough coverage. And that's just this like very, very exhaustive and tedious loop. And so yeah, when we say dense, we really mean dense. It's like this room I want 植物的叶子。对吧?要真正拍到覆盖每一个点的画面,你得非常乏味、非常彻底地在房间里走一圈。我想你们都见过我在各种地方跑来跑去采集。对受过训练的人来说,可能只要几分钟;但如果你把一部普通手机或采集设备交给一个普通消费者,甚至一个第一次做这件事的专业人士,大概要花一个小时。我见过有人第一次扫一个多房间环境,走了两个小时才拿到足够覆盖。那就是一个非常、非常耗尽精力、非常乏味的循环。所以我们说稠密,是真的稠密。就像这个房间我想

主持人 You just need like lots of 你就是需要很多。

嘉宾 so many photos, right? I want like 100, 200, 300 photos of this room to to capture it. And what we're trying to do is bring that down to like three. 非常多的照片,对吧?我想要这个房间的 100 张、200 张、300 张照片才能采集下来。而我们想做的是把这个降到大概三张。

主持人 Three. 三张。

嘉宾 Right? Well, we're saying like 50, 100 x reduction. And then that's at that scale where it just completely like flips that calculus on its head of like what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and like kind of unearth a lot of footage in the past you would never have treated as reconstructible and go back and like bring it to life as 3D potentially. This is something we've been playing around with a lot with Atlas, right? Like taking old clips. Like I've taken a bunch of my own old captures that never worked before and then put them through the system and I kind of seen a 对。我们说的是大约 50 倍、100 倍的缩减。到了那个量级,它就会彻底把那套计算翻转过来:你可以重建什么样的采集。你可以回到手里已有的影像。你可以去网上找到的素材,甚至用那些来建场景。你可以去随手拍的视频,把过去大量你从未当成可重建的素材挖出来,再把它们以三维的形式重新激活。这是我们用 Atlas 玩得很多的事。拿旧片段。我拿过很多自己以前从未成功过的旧采集,丢进这个系统,然后第一次看到了

嘉宾 reconstruction for the first time or taken my old captures and thrown away 95% of the photos I took and you know, imagine angles that I never would have gotten from a traditional kind of like Nerfer Splat type reconstruction. 重建;或者拿我的旧采集,扔掉我拍的 95% 的照片,去想象那些传统 NeRF 或 splat 类重建里我永远拿不到的角度。

主持人 One thing that's under appreciated on the website of the demos is the Stanford demo where Ben showed uh anywhere between 3 to 25 images you can reconstruct that entire Stanford quad. 官网上的演示里有一点被低估了,就是 Stanford 那个演示。Ben 展示了用 3 到 25 张图,就可以重建整个 Stanford 四方院。

主持人 But the thing is we had to show it from aerial view. But every single input image is been standing on the ground taking a picture from the ground. So, everything you see are generated but according to the laws of reconstruction. And this is really magical. 但问题是我们必须从航拍视角来展示。而每一张输入图都是站在地面上、从地面拍的。所以你看到的一切都是生成的,但却是按照重建的法则生成的。这真的很神奇。

嘉宾 And this is where like generation and reconstruction need to interplay in a really fundamental way to solve this problem. Because under the classic kind of reconstruction stuff that Ben was talking about, like the reason you need so many views is because I need like multiple images and I need to triangulate this point in 3D space and see it from multiple viewpoints. So, that means like that's required in the traditional version. And on the flip side, anything that wasn't captured in these views, like any pixel that was not visible in one of the input views will be a hole in a 3D reconstruction. Because fundamentally like if a thing wasn't visible in the input views, you know, you need to imagine it to fill in the gaps. And that's fundamentally a 而这正是生成和重建必须以一种非常根本的方式相互配合,才能解决这个问题的地方。因为在 Ben 说的那种经典重建里,你需要那么多视角的原因是:我需要多张图,我需要在三维空间里把这个点三角化,从多个视点看到它。所以在传统做法里这是必需的。反过来,这些视角里没有采到的任何东西,任何在输入视角里不可见的像素,在三维重建里都会是一个洞。因为从根本上说,如果一样东西在输入视角里看不见,你就需要把它想象出来才能补上缺口。而这从根本上是一个

嘉宾 generative process. So, even in this room, even if we set Ben loose with a DSLR and like let him like capture like hundreds of views of this room, even the world expert on doing these dense captures is still going to miss some spots. Like he's not going to get like underneath all of the microphones or underneath all the tables or like in between all the chair legs, you're always going to miss something, no matter how many views you get. So, that that's where you need generation as another mechanism in the model. Because you're never going to get everything. So, you need to have some generative capacity for the model to imagine, oh, based on what I'm seeing, then like first triangulate what I what I can see, but then fill in the gaps of the stuff that inevitably inevitably was not captured. 生成过程。所以哪怕在这个房间里,哪怕我们把 Ben 放出去,给他一台单反,让他采集这个房间的几百个视角,就算是做这种稠密采集的世界级专家,还是会漏掉一些点。他拍不到所有麦克风下面,拍不到所有桌子下面,拍不到所有椅腿之间。无论你拿多少视角,你总会漏掉一些东西。所以这就是你需要把生成作为模型里另一种机制的原因。因为你永远拿不到全部。所以你需要有一定的生成能力,让模型去想象:根据我看到的,先把我能看到的三角化,然后再把那些不可避免没采到的缺口补上。

主持人 Yeah, and there's something like super cool about this that LLMs have really understood this for a for a long time, right? There was kind of one of these like context wars like the first couple years I was like, oh, we got to 128 to 256 to like 512. We got a million, right? And everyone kind of understands now at a pretty tangible level the value of, you know, you crank your context length to high when you're using your coding models. It's a hard problem. Like everyone has a feel for that. But like no one has pushed that at all on the image and video model side in the same kind of like principled way. Like no one's out there trying to like put a like an hour-long video through and do a needle in a haystack retrieval of like a frame at the 30 37-minute mark. Whereas with reconstruction and generation, you actually have the same exact thing of like reconstruction is just like generation with a really long context and you put a lot of stuff in it. Right? Like that's the way to actually build this continuum where you kind of bridge 对,这里有一件非常酷的事,LLM 很早就真正理解了这一点。对吧?曾经有一场 context 战争,头几年是:哦,我们到了 128,到了 256,到了 512。我们到了一百万。对吧?现在大家都在一个相当可感的层面上理解了这个价值:你把 context长度拉高,在用编程模型的时候。这是个难题。大家都有体感。但在图像和视频模型这边,没有人用同样有原则的方式去推这件事。没有人试图把一个小时长的视频塞进去,再去做大海捞针式的检索,去找第 30 分钟、第 37 分钟的某一帧。而在重建和生成这里,你其实面对的是同一件事:重建就是带超长 context 的生成,你往里面塞很多东西。对吧?这才是真正建立这条连续谱的方式,你在这两件事之间架起桥。

主持人 between those two things. And like Atlas, like being able to like this is something we could never do with Marvel. Marvel had this kind of fundamental blocker of like you couldn't really jam more than honestly like a couple images in. But Atlas, I can go and I can actually take like a 64-image capture and do like a fly-through of an entire house and everything is grounded by being, you know, seen or like almost seen or like slightly extrapolated from what's not there. But you're just getting these, you know, I'm taking captures I did with 2,000 images of a multi-room house and taking it down to like 30, 40 inputs and the fly-through looks like basically the same. And this is just like totally inconceivable before and it's all enabled by building this gracefully scaling kind of context window that you can dump stuff into. 而这是我们用 Marble 永远做不到的。Marble 有一种根本阻碍:老实说,你塞不进去超过几张图。但 Atlas 可以。我可以拿一次 64 张图的采集,做整栋房子的飞穿,而一切都是落地的,因为被看到过,或者几乎被看到过,或者从缺失的部分稍作外推。但你得到的就是这些。我把一次用 2000 张图采集的多房间房子,降到大概 30、40 张输入,飞穿看起来基本上一样。这在以前完全无法想象,而这一切都是因为我们建了一个能够优雅 scaling 的 context window,你可以往里面倒东西。

嘉宾 And so the way to think about it is like the the sparseness are the pictures that you physically took and then Atlas is a model creates the rest of the views and then you use classic reconstruction techniques. Is that roughly the way to think about it or 所以理解方式是:稀疏的部分是你实际拍下的照片,然后 Atlas 这个模型生成其余视角,再用经典重建技术。大致是这样理解,还是

主持人 In some sense, yeah, yeah. I mean,that's the beauty of Atlas is like you can take however many inputs you have down to like a single view and then you can always use Atlas as this this rendering engine to produce anything else you want. Right? You can you can navigate it like a virtual camera. Yeah, exactly. You can just you can say like, "Okay, I have a picture here. I want a picture there, there, there." You can make a couple of those then you can say to dense fly-through. You can do this in sequence because it's an auto regressive model. It's up to you, right, to kind of pick and choose what you add interactively into the context as you generate. 某种意义上,对。我的意思是,这就是 Atlas 的美妙之处:无论你有多少输入,哪怕少到单个视角,你都可以始终把 Atlas 当作这个渲染引擎,去产出你想要的任何其他东西。对吧?你可以像虚拟相机一样在里面导航。对,没错。你可以说:好,我这里有一张图。我想要那里、那里、那里的图。你可以先做几张,然后再说要稠密飞穿。你可以按顺序做,因为它是自回归模型。由你来选,在生成时交互地把什么加进 context。

嘉宾 I mean, the thing that I just blows my mind is Listen, I just have a very simple mental mental model. I I I I four pictures and then I've got to like have the model extrapolate between them and then it has to fit when you reconstruct. Like it's got to be 3D. Like and I always think of these diffusion models as like being visually great but not accurate. And so like And and I don't even know if there's a question here, but like how how how is I like the room fits? So like how is it that it's three 3D consistent? Is it just lots of data or 让我感到震撼的是,听着,我只有一个非常简单的心智模型。我有四张图,然后必须让模型在它们之间外推,重建的时候还得对得上。它必须是三维的。而我一直觉得这些 diffusion 模型视觉上很好,但不准确。所以,我甚至不知道这里有没有一个问题,但到底是怎么让房间对得上的?它怎么做到三维一致的?只是因为数据很多吗?

Insight

内部故事把科学判断收成一次看见:相机飞进桌下,才算活了。

  1. 初夏一个更小的模型,相机从 NeRF 论文那张花园桌下飞过,足球也在。三人早上对视,五秒内决定做这个(约 00:23–00:25)。
  2. 李飞飞看着团队从不知道要花多久,到有生命迹象,再到这会成功。没有人做成过,但他们对这两条假设有确信。

主持人 Yeah, I mean it's a part partially it's a belief in the scaling hypothesis, right? Like, you know, 对,部分是对 scaling hypothesis 的信念,对吧?

嘉宾 Did you By the way, but I have to ask, when you set out to do this, did you know it was going to work? 对了,我必须问,你们开始做这件事的时候,知道它会成功吗?

主持人 I was pretty sure. 我相当有把握。

嘉宾 Were you sure? 你有把握吗?

主持人 I think three of us have total conviction about the scaling law. That I think we do. I do think the exact architecture choices and data mixtures is where the the devils are in the details. I, you know, have watched Justin and his team going from we really don't know how long this is going to take to oh, maybe sign of life to wow, this is going to work. So it it 我觉得我们三个人都有完全的。确信,关于 scaling law。这一点我想我们是有的。我确实认为,精确的架构选择和数据配比,才是细节里的魔鬼所在。我看着 Justin 和他的团队,从我们真的不知道这要花多久,到哦,也许有生命迹象了,再到哇,这会成功。所以

主持人 no one what no one has done it, but I think the hypothesis, two hypothesis, one is scaling law hypothesis, the other one is next viewpoint prediction. We had conviction of these two things primarily. 没有人做成过,但我认为假设有两个:一个是 scaling law 假设,另一个是 next viewpoint prediction。我们主要对这两件事有确信。

嘉宾 So I think I was very convicted that it was going to work. I was not sure it was going to work at this well at this fast, right? Like I thought there's a chance that we do this. Maybe it's not clear that like the first cycle of pre-training a new model with a new architecture and a new paradigm. Like the first cycle of that working is insane. So I thought there was a chance in which we had to we might have had to do a couple more turns of that of that pre-training cycle before we got to the level of quality we we wanted. 所以我觉得我非常确信它会成功。我不确定的是它会这么好、这么快,对吧?我当时想,有可能我们做了这个。也许并不清楚:用新架构、新范式去预训练一个新模型的第一个周期,第一个周期就成功,这是疯狂的。所以我当时想,有可能我们还得再走几轮那样的预训练周期,才能到我们想要的质量水平。

主持人 Is there is it are we kind of like at the end of like the scaling for this architecture approach? We need another breakthrough or is there 这种架构路径的 scaling 是不是已经到头了?我们还需要下一次突破,还是

嘉宾 No, no, we're at the beginning. 不,不,我们才刚开始。

主持人 Really? 真的?

嘉宾 Yeah. 对。

主持人 Without without changing the architecture? 在不改架构的情况下?

嘉宾 Yeah, we're basically at the beginning. I think we're basically at the beginning and we're basically limited by compute at this point. All right, like data is important as Feifei likes to point out, but like everything has a bottleneck and I think the main bottleneck on continuing to scale this thing is actually training compute. Right? Like during development, we trained a sequence of models. We wrote about this in the blog post a little bit, but we trained a couple of models that like the first couple of runs of the scaling ladder. Um and each time we made the model bigger and each time we trained it for longer, each time we put it on more chips, like it got significantly better. And the model size that we like the model that we showed in the blog post is obviously the biggest and best one that we trained, but the thing that was limiting it was not the 对,我们基本上才刚开始。我觉得我们基本上才刚开始,而目前基本上受限于算力。好,数据很重要,Fei-Fei 喜欢强调这一点,但每件事都有瓶颈,而我认为主要瓶颈,要继续把这个东西 scale 上去,实际上是训练算力。对吧?开发期间我们训练了一系列模型。我们在博客里写过一点,我们训练了几个模型,相当于 scaling 梯子上的头几级。每一次我们把模型做大,每一次我们训练得更久,每一次我们放到更多芯片上,它都会明显变好。我们在博客里展示的那个模型,显然是我们训练过的最大、最好的一个,但限制它的并不是

嘉宾 scale or the data or anything like that. It was literally like we had a deadline of when we wanted to release this thing and therefore we backed up what we could afford to train in time for that deadline. 规模或数据之类的东西。字面上就是:我们有一个想发布这个东西的截止日期,所以我们倒推了在那个截止日期前赶得上、训练得起的规模。

主持人 But here's a little bit of a insider story, right? Like Justin and team are training from the smaller and slightly bigger, you know, are train having these roadmaps. And then there was one day in summer, early summer, that it's not even this the current Atlas model size, it's a smaller model. And then Ben, Justin, Ben feeded into, you know, the viewpoint generation. And remember that famous table, the garden table for Nerf paper and many papers, that overnight I got a slack. I mean, we all saw the slack from Ben that our camera flew through under the table. 但这里有一点内部故事,对吧?Justin 和团队从更小的、稍大一点的开始训练,有这些路线图。然后夏天有一天,初夏,那甚至还不是现在这个 Atlas 的模型规模,是一个更小的模型。然后 Ben、Justin,Ben 把它喂进视点生成。还记得那张著名的桌子吗,NeRF 论文和很多论文里的那张花园桌子。那天晚上我收到一条 Slack。我们都看到了 Ben 发的 Slack:我们的相机从桌子下面飞过去了。

嘉宾 With the soccer ball. 还有那个足球。

主持人 Uh yes, with the soccer ball. 对,还有那个足球。

嘉宾 ball emergent or is that in the original picture? 球是涌现出来的,还是原来。图里就有的?

主持人 It's It's It was real, right? 它是真的,对吧?

嘉宾 Okay. 好。

主持人 That morning the three of us looked at each other in the eyes and say, "That's it. This is We're going to build this." Like we we made a decision within 5 seconds. How This is just a no one has ever seen this result. 那天早上我们三个人对视,说:就是它。我们要做这个。我们在五秒内做了决定。从来没有人见过这样的结果。

嘉宾 Ben, um can you talk through maybe more specifically the use cases? So so uh World Labs has historically had a lot of users that were creatives and they use it for like consistency in, you know, like whatever 2D images and for movies and for 3D and for games etc. And so maybe can you talk about how this extends use cases or cater to the existing ones and then we'll I'd like to talk about robotics actually. Ben,你能不能更具体地讲讲用例?World Labs 历史上有很多用户是创作者,他们用它来做一致性,无论是二维图像、电影、三维还是游戏等等。所以你能不能讲讲这如何扩展用例,或如何服务现有用例,然后我其实还想谈谈机器人。

主持人 Yeah, sure. Um yeah, I mean it's kind of funny actually one of the kind of main ways we even saw people using Marvel plays exactly into this new view prediction case. Like a lot of our 好。其实有点有意思,我们看到人们使用 Marble 的主要方式之一,恰好就落在这个 new view prediction 的用例上。我们很多

Insight

一致性本身就是产品,不只是更好看的帧。

  1. 人们习惯舞台、道具和随时间发展的环境,不想生成完就扔掉、只留下提示词(约 00:26–00:28)。
  2. 会议展位、建筑施工都有一段辛苦的虚拟设计。AI 被认为能把“玩乐高、陶艺、铅笔”那种直观,接到现在几十年没改的三维软件上(约 00:28–00:29)。

嘉宾 Marvel being the previous sorry our Marvel our previous product. Marble 是之前的,抱歉,我们。之前的产品 Marble。

主持人 Um like people would take that product, put an image in, get a full 3D scene as a Gaussian splat, take a couple of screenshots of it from different points of view and leave. Right? And we're like We can just make these images and that's generative AI cool, right? So I think like and you know, there's a lot of degradation there. They're like, "Oh, this splat could look better." And it's like, "Okay, what if we just generatively model those viewpoints with that exact modality of control?" So I think like even that core capability of just like view synthesis, um it's sort of been this academic problem for a long time. But in the sense of "Oh, you're going to do this really dense capture." Like like generative view synthesis is a relatively quite a new problem. And we just see so many people who uh in this creative pipeline, right? People have a multi-stage workflow, right? I don't think there's a single person out there using one monolithic model, not even C dance or whatever for their entire task. Uh people will have this like, you know, kind of a bunch of storyboards and mood boards of images they pull out from like their favorite collection of image models. And then they'll go to different video tools 人们会拿那个产品,放进一张图,得到一个完整的三维场景,形式是 Gaussian splat,再从不同视点截几张图,然后就走了。对吧?然后我们就想,我们直接生成这些图就好了 ,那也挺 generative AI、挺酷的,对吧?所以我觉得,那里有很多质量下降。他们会说:哦,这个 splat 可以看起来更好。然后就是:好,如果我们直接用那种精确的控制模态,用生成的方式去建模那些视点呢?所以我觉得,即便只是视角合成这项核心能力,它也在很长一段时间里都是一个学术问题。但那种「哦,你要做这种非常稠密的采集」的意义上,生成式视角合成是一个相对相当新的问题。我们看到很多人在这条创意流水线里。人们有多阶段工作流。对吧?我不认为外面有任何一个人用一个单体模型完成全部任务,哪怕是 C dance 之类也不行。人们会有一堆分镜和情绪板,图是从他们最喜欢的那批图像模型里抽出来的。然后再去不同的视频工具

主持人 and like build those together as keyframes. Then they'll go and like clip and edit those later, right? So we were seeing this like sort of, you know, niche but very specific use case for Marvel as just providing that like sanity that you can ground your generations in some kind of 3D consistent world, right? People, you know, I don't I I fought with with various image models to ask them to like give me different viewpoints of a room. And every time you can just look and see, "Oh, things kind of moved around. Like it's not stable." And like even that one seed of a use case I think kind of signals that there's this value and there's hiding under the surface there. Like there's just decades of people being used to persistent 3D state like virtually modeling what they would be doing in the real world and having, you know, a stage and props and like elements there, whether it is for a movie or a show or a marketing shot or like building out game environments. Like this this statefulness and persistence is so key in how people think about spatial reasoning and like developing an environment over time. 把它们拼成关键帧。然后再去剪辑。对吧?所以我们看到 Marble 有一种小众但非常具体的用例,就是提供那种踏实感:你可以把生成结果落在某种三维一致的世界里。对吧?我跟各种图像模型较过劲,让它们给我一个房间的不同视点。每次你一看就会发现:哦,东西有点挪了,它不稳定。即便只是这一个用例的种子,我觉得也某种程度上说明那里有价值,就藏在表面下面。人们几十年来已经习惯持久的三维状态,像在虚拟里建模他们在真实世界会做的事,有一个舞台、道具和元素,无论是电影、节目、营销拍摄,还是搭建游戏环境。这种有状态、这种持久,对人如何思考空间推理、如何随时间发展一个环境,至关重要。

主持人 Like people don't think in this ephemeral like generate a thing, like just throw it away, keep my text prompts. Like people want to build this like collection of assets and like model a world in that way. So we're we're trying to provide like again with with the spatial context mechanism and other things like we're trying to provide that level of control and precision and the ability to adjust different modalities of input starting with the post images, but you know, we want to give people more control over the elements of the things in the scenes they're looking at and editing and interaction and all that as we go forward. And I think that that it unlocks like further use cases in those areas we're already seeing, but also expanding out into kind of any place people want to create a virtual replication or or like, you know, a pre-imagination of a real-world space they need to build, right? For architecture and construction. Like I talked to a guy at some point building booths for conferences, right? There's just so many things in the world you don't think about need to be fabricated and every single one of those basically goes through this like pretty 人们不会用那种转瞬即逝的方式思考:生成一个东西,扔掉,只留下我的文本提示。人们想建立一套资产集合,用那种方式去建模一个世界。所以我们在尝试提供,同样借助 spatial context 机制和其他东西,我们在尝试提供那种程度的控制和精度,以及调整不同输入模态的能力,从位姿图像开始。但我们想给人更多控制,去控制他们正在看、正在编辑的场景里那些元素,以及交互,以及我们接下来会做的那些。我觉得这会解锁我们已经在那些领域看到的更多用例,也会扩展到任何人们想创造虚拟复制、或者对他们需要建造的真实空间做预先想象的地方。对吧?建筑和施工。我曾经和一个给会议搭展位的人聊过。对吧?世界上有太多你想不到需要被制作出来的东西,而其中每一件基本上都要经过这种相当

主持人 painstaking virtual design phase. And of that process like the part where you go into 3D software is kind of one of the most like arduous and like labor-intensive parts right now. Like taking feedback on a 3D design from kind of like verbal commentary or sketch or really really quick stuff you got from like a creative director or like a design director or an architect or whatever. Like mapping that back into the 3D representation is like 95% of the work, right? You can have a meeting get feedback and then you go back and do a week of revisions. And that's just because like our software is kind of decades old at this point and it's it's just never became as intuitive as, you know, playing with Legos or like pottery or doing this stuff with your hands or sketching with a pencil. And this is the real place where AI can actually unlock a ton of value for people in their process, whether it's a creative application or something more industrial or design or whatever. Uh, and that like really motivates me to 辛苦的虚拟设计阶段。而在那个过程里,进入三维软件的那部分,目前是最艰难、最耗人力的部分之一。把对三维设计的反馈,来自口头评论、草图,或者创意总监、设计总监、建筑师随口给你的非常快的东西,映射回三维表示,这几乎是95% 的工作。对吧?你可以开一次会拿到反馈,然后回去做一周修改。那只是因为我们的软件到现在已经有几十年了,它从未变得像玩乐高、像陶艺、像用手做这些事、像用铅笔画草图那样直观。而这才是 AI 真正能为人和他们的工作流程解锁大量价值的地方,无论是创意应用,还是更工业、更设计的事情。而这非常驱动我去

主持人 kind of build different flavors of of our model to cater to those kind of people. 做我们模型的不同口味,去服务那类人。

Insight

会动的世界不是装饰。对机器人来说,仿真器就是训练场,也是未来的规划器。

  1. 李飞飞说机器人现在最大的问题是数据,有一天会是芯片。Atlas 被看成当前 real-to-sim 的下一代技术;下一步是吃动态数据,弥合动作规划(约 00:30–00:33)。
  2. Justin 提出神经仿真器,甚至问:仿真器自己为什么不变成规划器。理解世界如何响应动作,和想象该采取什么动作,是同一件事(约 00:34–00:35)。
  3. 专家反馈要更多动态。Marble 架构是静态的;Atlas 已支持动态,演示里有水波和小车。预训练见过大量动态,这次发布的后训练更偏静态空间运动(约 00:35–00:38)。

嘉宾 Yeah, I I I can understand how it helps with the creatives cuz like Marvel did that. And also how that extends to things like designer architecture. Uh, Baidu you acquired a robotics company. And so it's 对,我能理解它如何帮助创作者,因为 Marble 就做过这个。也能理解它如何延伸到设计师、建筑这类事。对了,你们收购了一家机器人公司。所以这就

主持人 it's less less 没那么

嘉宾 talked about it too. We just talked about 也谈过。我们刚刚谈过

主持人 know. But it's the last time we talked to me especially in the context of Atlas like how that maps to robotics. So if you wouldn't mind just penciling that out. 知道。但上次我们谈的时候,尤其是在 Atlas 这个语境下,它如何映射到机器人。所以如果你不介意,大致画一下。

嘉宾 Yeah, actually Atlas is a a key part of the puzzle. So, um, we acquired this company that was formerly known as Synnex. And what is their key technology? Right now their key technology is a system that goes from real to sim and and then sim to real. And what does that mean in robotic situation? You want to train a robotic arm to, you know, figure out how to, um, do cabling, let's say, in a in a 好,实际上 Atlas 是拼图里的关键一块。我们收购了这家公司,它以前叫 Synnex。他们的关键技术是什么?现在他们的关键技术是一套从真实到仿真、再从仿真到真实的系统。这在机器人场景里意味着什么?你想训练一条机械臂,让它搞清楚怎么做布线,比如说在一个

嘉宾 industrial setting. Well, you need a whole bunch of data to first train a robotic policy to do these cable cables cabling activity. And then you want to evaluate if the robotic policy is doing a good job. And then you deploy the robot into the cabling environment. In order to train, what you what this company, uh, Synnex and now our robotics team used to be doing is doing exactly what Baidu was saying, dense reconstruction. You take 工业环境里。那你首先需要大量数据,来训练一个机器人策略去做这些布线活动。然后你想评估这个机器人策略是否做得好。然后你把机器人部署到布线环境里。为了训练,这家公司 Synnex,也就是现在我们的机器人团队,以前做的是正好是 Ben 在说的稠密重建。你拍

主持人 All right, so Synnex 好,所以 Synnex

嘉宾 pictures of a situation and then and try to reconstruct that environment. It's excruciatingly painful. Takes a long time, laborious, and it really blocks the velocity of robotic simulation real to sim, right? So Atlas really is the next generation technology for that. And this is not just for robotics cabling or anything. We should zoom out and recognize the biggest problem right now in robotics is actually data. One day it'll be chips, but for now it's data. Because it's so hard to collect real-world data 一个场景的照片,然后试图重建那个环境。极其痛苦。花很长时间,很费力,而且它真正卡住了机器人仿真从真实到仿真的速度。对吧?所以 Atlas 真正是下一代技术。而这不只是为了机器人布线或任何一件具体的事。我们应该拉远来看,认识到机器人领域现在最大的问题实际上是数据。有一天会是芯片,但现在是数据。因为采集真实世界数据太难了

嘉宾 where robots, you know, are operating and and uh in order to uh not only you need to uh collect the data of, let's say, the cabling situation or dishwashing situation, whatever. There is also a very important step called randomization. Is that you have to take the same environment and then randomize the conditions. So, the cable doesn't literally only, you know, uh bend this way. It can bend a different way or the box can have different sizes, colors, different lids, and all that. And and or in different parts of the scene. So, you have to go through a real-to-sim situation in order to get enough of that data in addition to other data you can get from internet. So, this real-to-sim uh step will be um you know, really helped by by Atlas. 机器人在其中操作的地方。为了做到这一点,你不仅需要采集数据,比如说布线场景,或者洗碗场景,无论什么。还有非常重要的一步叫随机化。也就是你必须拿同一个环境,然后把条件随机化。这样电缆不会字面上只朝这一个方向弯它可以朝另一个方向弯,或者盒子可以有不同尺寸、颜色、不同的盖子,诸如此类。或者在场景的不同位置。所以你必须走一遍从真实到仿真的过程,才能拿到足够的那种数据,再加上你可以从网上得到的其他数据。所以这个从真实到仿真的步骤,会真正被 Atlas 帮到。

嘉宾 That's just the first part of this is meeting the the the robotics needs in the current technology cuz we don't yet have a a frontier foundation model that's robust enough for for uh robotics. But Atlas is a uh omni model. It's a multi uh multi-modal model. It takes on different kinds of input and generates different kind of output. You can totally imagine the next step is Atlas um taking in uh data that's in the uh in that's dynamical. And that can really start to bridge the gap between, you know, action planning and uh and um a robotics and and the Atlas output. So, that's all narrow map. What are you going to say something? 这只是第一部分,是用当前技术去满足机器人需求,因为我们还没有一个足够稳健、面向机器人的前沿基础模型。但 Atlas 是一个 omni 模型。它是一个多模态模型。它接收不同类型的输入,生成不同类型的输出。你完全可以想象下一步是 Atlas 接收动态的数据。而这可以真正开始弥合动作规划和机器人、以及 Atlas 输出之间的差距。所以我就先把这条映射画到这里。你要说什么?

主持人 Yeah, I I was going to say there's something fundamentally different about training a robotics policy compared to really any other application in AI we've seen before. Um and that's like if you're generating a piece of code, like you're generating an image, you're generating a video, I'm the model is fundamentally creating this this artifact. And that artifact like there's a lot of examples of artifacts that you can go out on the web or somewhere and collect. Right? You want to generate images, there's a lot of images out there. You want to generate videos, there's a lot of videos out there. You want to generate a code base, there's a lot of code bases out there you can learn from. 对,我正要说,训练一个机器人策略,和我们以前在 AI 里见过的几乎任何其他应用相比,有一些根本不同的地方。如果你在生成一段代码,生成一张图,生成一段视频,模型从根本上是在创造一件制品。而这件制品,网上或别处有大量制品样本可以去收集。对吧?你想生成图像,外面有很多图像。你想生成视频,外面有很多视频。你想生成一个代码库,外面有很多代码库可以学习。

嘉宾 Yeah. 对。

主持人 A robotics policy is something fundamentally different. It's not producing a static thing. It's instead a policy that's going to go out into the world, make actions, and like try to achieve a goal. And the world's not always going to respond the way you expect, right? Unexpected stuff is going to happen. So, a robotics policy is like fundamentally an agent that is out in the real world interacting with the real world, and stuff happens. So, you need like a critical part of that is the those those policies during training need to be exposed to every possible thing that could go wrong during a 机器人策略是根本不同的东西。它不是在产出一个静态的东西。它是一个策略,要走进世界,做出动作,试图达成一个目标。而世界并不总是按你预期的方式回应,对吧?会有意想不到的事发生。所以机器人策略从根本上是一个智能体,在真实世界里与真实世界交互,然后事情会发生。所以你需要的关键一部分是:这些策略在训练期间,必须被暴露到部署时可能出错的每一种情况。

主持人 deployment. Um and that's where simulation is really key for robotics, right? So, there's the then there's two angles on that. Like one is the kind of classical simulation. You can go out You can go and like go to your favorite physics engine and like try to imagine creatively as a human designer, what are all the scenarios that might happen in this in this in this when achieving this task, and then try to write explicit code that models them all. That That's one angle. And that's an interesting angle with coding agents. Like that actually gets supercharged, too. But there's another angle, which is try to more data-driven simulation. Right? Like maybe we can Can we have a learned model that can understand how the environment how the world is going to respond to actions? And maybe it might respond in unexpected ways sometimes. Then could we build these neural simulators that are trained on as much data as we can, then use these neural simulators, these learned neural simulators, you know, as a simulation bed to train robotic policies. Um and that's that's a really interesting future direction of Atlas. 而这正是仿真对机器人真正关键的地方。对吧?这里有两个角度。一个是经典仿真。你可以去你最喜欢的物理引擎,作为人类设计者去创造性地想象:完成这个任务时可能发生的所有情景,然后尝试写显式代码把它们全部建模出来。这是一个角度。而这个角度配上编程智能体也很有意思,它实际上也会被大幅加强。但是还有另一个角度,是尝试更数据驱动的仿真。对吧?也许我们能不能有一个学来的模型,理解环境、世界会如何响应动作?也许它有时会以意想不到的方式响应。那我们能不能构建这些神经仿真器,用我们能拿到的尽可能多的数据来训练,然后用这些神经仿真器、这些学来的神经仿真器,作为仿真床去训练机器人策略。而这是 Atlas 一个非常有意思的未来方向。

主持人 But but then it doesn't stop there, right? So 但然后并不止于此,对吧?所以

嘉宾 so um but once you have, you know, this learned simulator, like this learned simulator kind of already has in its like mental brain, like it understands the world, it understands how the world is going to respond to actions, and why doesn't the simulator itself become the planner? Right? Like the same That's kind of the core thesis that we've had around world models and their generality, that there's some core stuff that a model should understand around generating worlds, simulating them,understanding how they appear in different situations, and, you know, understanding how the world's going to respond to an action is highly related to imagining what kind of action I need to take to make the world respond in a particular way. 所以一旦你有了。这个学来的仿真器,这个学来的仿真器在它的脑子里已经有了:它理解世界,它理解世界会如何响应动作,那为什么仿真器自己不变成规划器呢?对吧?这就是我们围绕 world model 及其通用性的核心论点:有一些核心的东西,模型应该理解,关于生成世界、仿真世界,。理解它们在不同情况下如何呈现。而理解世界会如何响应一个动作,和想象我需要采取什么样的动作才能让世界以特定方式响应,是高度相关的。

主持人 Yeah. What what um one piece of feedback that I got I've been By the way, congrats on the launch. It was overwhelmingly positive. I think it was probably the most significant model launch this year. And yeah, everybody said glowing things. But one person who's an expert in the space who I texted was like, "What do you think?" And great person says, "It's fantastic, it's amazing, but there needs to be more dynamics." 对。我收到的一条反馈是,对了,恭喜发布。反响一边倒地正面。我觉得这可能是今年最重要的一次模型发布。大家都说了很好的话。但有一个这个领域的专家,我发了消息问:你怎么看?那个很棒的人说:太棒了,太惊人了,但需要有更多动态。

嘉宾 And so it would seem to be at least in the robotics case, but generally it was kind of ideal to actually have a world that moves. And so maybe talk a little bit about that and then any other future directions that A, you're comfortable sharing, but you think are worth talking through. 所以至少在机器人这个例子里,而且一般而言,真正有一个会动的世界似乎是理想的。所以也许讲讲这一点,然后再讲任何其他未来方向,一是你们方便分享的,二是你们觉得值得讲的。

主持人 Yeah, I mean, like dynamics is clearly going to happen. Like actually um 对,动态显然会发生。实际上

嘉宾 We have baby dynamics. 我们有婴儿级的动态。

主持人 We actually do have baby dynamics already. And this is something I think people didn't quite appreciate, we didn't really highlight in the blog post, but like the previous marble world model, it was like fundamentally static. Like the the model just like could not handle any dynamics at all. And that was just like baked into the model architecture, baked into the training, like the whole thing was fundamentally static. 我们其实已经有婴儿级的动态了。而这是我觉得人们没有完全意识到,我们在博客里也没有真正强调:之前的 Marble world model 从根本上是静态的。模型完全处理不了任何动态。而这是写进模型架构、写进训练里的,整件事从根本上就是静态的。

嘉宾 Yeah. 对。

主持人 Um we already knew that that was a big problem post marble, and we already fixed it in Atlas, right? Like the Atlas architecture is already fundamentally supports dynamics. Um the Atlas training data fundamentally has dynamics. And if you look carefully in some of the videos that we've even posted, 我们在 Marble 之后就已经知道那是一个大问题,而且我们已经在 Atlas 里修好了。对吧?Atlas 的架构已经从根本上支持动态。Atlas 的训练数据从根本上就有动态。如果你仔细看我们甚至已经发出去的一些视频,

嘉宾 It actually is. 它确实是。

主持人 I saw I saw you see that. 我看到了,我看到你看到了。

嘉宾 Yeah, the waves, the water waves. 对, 那些波浪,水波。

主持人 Yeah, so like some of the examples there's like waves in the water, like in some of the like air generated aerial views, there's like little cars moving around. So like dynamics is actually already in this model. 对,有些例子里水里有波浪,有些 AI 生成的航拍视角里有小车在动。所以动态其实已经在这个模型里了。

嘉宾 But is it By the way, dynamics seems very problematic to me if you're trying to reconstruct 3D from multiple views, right? And so like are these things like at odds or 但是,对了,如果你要从多个视角重建三维,动态对我来说似乎非常成问题,对吧?所以这些事情是不是相互冲突的?

主持人 it's actually one of our thesis here is that like, you know, if you're just going to do fundamental 3D reconstruction, you actually want to have no dynamics. Like you want to be able to model like exact views of the scene with exact frozen time. 这。其实是我们这里的论点之一:如果你只是做根本的三维重建,你实际上想要。没有动态。你想能够用精确冻结的时间,去建模场景的精确视角。

嘉宾 Right. 对。

主持人 Um, but then like this is actually kind of a problem with our previous Marble approach, right? Like there like you can try to find data that's fully static, but that's really hard to scale and really hard to get more of. And the thing we realized is that even in the case where I want a static output in the end, the best way to get it is actually expose the model to dynamics, right? Like expose the model to as much dynamic stuff as you got, as much static stuff as you got, and let the model figure out how to factor out the dynamic stuff. So in especially in like the This is like 但这其实也是我们之前 Marble 路径的一个问题。对吧?你可以试图找完全静态的数据,但那很难 scale,也很难拿到更多。我们意识到的是:即便最终我想要静态输出,最好的办法实际上是让模型接触动态。对吧?让模型接触你有的尽可能多的动态东西,以及尽可能多的静态东西,然后让模型自己搞清楚如何把动态的东西分解出去。所以尤其是

嘉宾 And this is actually So again, in the Atlas pre-training already, like it it saw a ton of dynamics in the pre-training already. Then the post-training that we did specific to this checkpoint in this release was focused a lot more on on static stuff, focused a lot more on spatial movement and less on temporal. But like we already have we already like I'm pretty sure this the the the pre-trained checkpoint already has a lot of latent dynamics in it. 而这实际上,再说一次,在 Atlas 的预训练里,它已经在预训练中见过大量动态。然后我们针对这次发布的这个 checkpoint 做的后训练,更多地聚焦静态的东西,更多地聚焦空间运动,而更少聚焦时间。但我们已经有了,我相当确定这个预训练 checkpoint 里已经有大量潜在的动态。

主持人 Yeah. 对。

嘉宾 Um, and this is something we're going to improve quite a lot going forward. 而这是我们接下来会。大幅改进的。

主持人 So so Ben, does that mean we're going to get 4D video? You can like go walk around. 所以 Ben,那是不是意味着我们会得到 4D 视频?你可以走进去走一圈。

Insight

原语被抬到和 next token 同级;节目没有证明它已经解决智能,只说明他们为什么押这一注。

  1. 图像模型的多轮编辑还没有同样强地传到视频和 world model。工业级控制是产品界面几乎要从零重做的前提(约 00:38–00:41)。
  2. Ilya 的推理小说例子被用来说明 next token prediction 的 AI complete。他们用杀手走出来的新视点,以及“Martine 在写黎曼猜想证明”的世界,说明 generative new view prediction 也可以框住任意智能任务(约 00:41–00:43)。

嘉宾 I mean, I think 我觉得

主持人 see the smile on their face. 看他们脸上的笑。

嘉宾 So I I I actually can see if like you just stopped now and you only did kind of bigger, you know, faster, better, you could build almost an entire industry. Like I feel it feels like a very horizontal primitive. And then and if you did nothing else, but are there other things that are not just kind of bigger, faster that you're excited about with the applications you're focused on, which tend to be kind of more on the kind of content creator 3D side. 所以我其实能看到,如果你们现在就停在这里,只做更大、更快、更好,你们几乎可以做出整个产业。我觉得它像一个非常横向的原语。然后如果你们别的什么都不做,但在你们聚焦的应用上,有没有不只是更大、更快、让你们感到兴奋的其他事?而这些应用往往更偏内容创作者、三维这一侧。

主持人 Yeah, I'm really excited about pushing that kind of multimodal aspect. I think different modes of control is so critical here. I think like it's super under appreciated, especially in the academic community, how critical it is to add control conditioning to these models to kind of get out what's inside. I mean, honestly, this 对,我非常兴奋去推那种多模态的方面。我觉得不同的控制模式在这里至关重要。我认为这一点被严重低估了,尤其是在学术界:给这些模型加上控制条件,把里面的东西拿出来,有多么关键。说真的,这

嘉宾 I'm not trying to I don't even understand what those words 我不是想,我甚至听不懂这些词。

主持人 So, translate in layman's language is so editability. I think editability is is the key here. 所以用外行的话说就是可编辑性。我觉得可编辑性才是这里的关键。

嘉宾 Yeah, so I mean, this is something we've seen in like sort of single image models and starting this year in video models is starting to be unlocked in terms of oh, like getting that flavor of like multi-turn or really like intuitively interpreting like I want like this person and this object and this thing to happen and kind of combining those all together in like one pastiche without having to do a lot of like manual work with the system. Like it just interprets it like kind of frontier image models are kind of there, right? For in terms of editing. But, we haven't seen that 对,这是我们在单图模型里已经看到的,而从今年开始在视频模型里也开始被解锁:那种多轮的感觉,或者真正直观地理解:我想要这个人、这个物体、这件事发生,然后把它们全部组合成一个拼贴,而不必对系统做大量手工操作。它就是能理解。前沿图像模型在编辑这方面已经差不多到了,对吧?但我们还没有看到这一点

嘉宾 propagate out as strongly into video yet and then into world models, right? We've seen some really kind of toy examples of oh, I can like put in a sentence and like, you know, a dinosaur appears or something with these like sort of real-time models. Um but, I want to like turn that up to really industrial strength and make that cuz like the the trick here is you got to add control, but not compromise the quality of the model or it just becomes a a party trick, basically. Like it's like no one is going to seriously think about swapping their like cutting-edge frontier video model usage for your model if you give them extra knobs, but the quality degrades. So, I think it's really that game of like how can we maintain maintain like the high bar we've set with the outputs we're able to get in the current model and then add all kinds of interesting stuff that people will ask us for in terms of like I want to interact with the scene or control the layout or control like the identity of the objects and the things that we're seeing within there or control time, right? Um and I think that's like an axis where it opens up like a ton of really interesting product 同样强地传到视频,然后再传到 world model。对吧?我们见过一些非常玩具级的例子:哦,我可以塞进一句话,然后一只恐龙出现了,之类的,用那些实时模型。但我想把这个做到真正工业级的强度,因为这里的诀窍是:你必须加上控制,但又不能牺牲模型质量,否则它基本上就变成一个宴会把戏。没有人会认真考虑,把他们最前沿的前沿视频模型用量换成你的模型,如果你给他们额外的旋钮,但质量下降了。所以我觉得真正的游戏是:我们如何维持当前模型已经能拿到的输出所设立的高标准,然后再加上人们会向我们要的各种有意思的东西,比如我想和场景交互,或控制布局,或控制我们在里面看到的物体和东西的身份,或控制时间。对吧?我觉得那是一个轴,它会打开大量非常有意思的产品和。

嘉宾 and interface work. The more complexity you add there and richness in terms of kind of like enabling you to really think about like redesigning almost from scratch the way people interact with sort of like stateful, you know, 3D worlds in the computer. Like that's really the end goal here. That is getting like all the capabilities you need to build that kind of system. 界面工作。你在那里加的复杂性越多、丰富性越多,就越能让人真正去想:几乎从零重新设计人们与计算机里有状态的三维世界交互的方式。那才是这里真正的最终目标。也就是拿到构建那种系统所需的全部能力。

主持人 Awesome. Anything to get out of that as far as new functionality that you'd be excited about that's not just bigger better? 太好了。从新功能的角度,有没有不是只是更大更好、让你感到兴奋的东西要讲?

嘉宾 I think for me, let's go back to the first principle of intelligence. Intelligence is not sitting there stuck and just seeing something or interpreting something when it comes to space and physical space, right? It's really this uh closing the loop between seeing and experiencing and interaction. So, just thinking about going up that ladder is exactly what uh Ben said. 对我来说,让我们回到。智能的第一性原理。智能不是坐在那里卡住,只是看见什么或解释什么,当涉及空间和物理空间的时候。对吧?它真正是把看见、体验和交互闭环起来。所以只是想着沿那把梯子往上走,这正是 Ben 说的。

主持人 Cool. I think one interesting notion there is this notion of AI completeness. 酷。我觉得那里有一个有意思的概念,就是 AI completeness。

嘉宾 You heard of that before? 你以前听过吗?

主持人 Yeah, yeah, I have. Yeah. 对,我听过。

嘉宾 So, like everyone 所以像每个人。

主持人 Yeah, I I hear about AI complete, by the way, in terms of LLMs, which is like you have to be basically, you know, like the smartest LLM to answer the question what the smartest LLM will need to answer or you have to solve general intelligence. 对,我听说过 AI complete,顺便说,是在。 LLM 的语境里,意思是你基本上必须是最聪明的 LLM,才能回答最聪明的 LLM 需要回答的问题,或者说你必须解决通用智能。

嘉宾 No, no, it's it's basically it's it's a connection to Turing completeness, right? Like the idea being that like a task is Turing complete like in classical complexity theory if like I can take any class of any problem in this category, reduce to that one problem, right? Three SAT is a classic example, right? So, you can take any NP-hard problem and reduce it to three SAT. Therefore, therefore, you can use three SAT to solve any any problem. 不,不,它基本上是和 Turing completeness 的关联。对吧?这个想法是:一个任务是 Turing complete,在经典复杂性理论里,如果我能把这一类里的任何问题都规约到那一个问题。对吧?3-SAT 是一个经典例子。对吧?所以你可以把任何 NP-hard 问题规约到 3-SAT。因此,你可以用3-SAT 去解决任何问题。

主持人 Yeah, it's it's a yeah, yeah. 对,对。

嘉宾 So, so then like the the kind of like soft definition of AI completeness is like there's this fundamental primitive that's that's an AI task, but if I could solve this AI task in its full broadest generality 所以。然后 AI completeness 那种软定义就是:有一个根本原语,它是一个 AI 任务,但如果我能在最完整、最宽的通用性上解决这个 AI 任务。

主持人 You solve that. 你解决了那个。

嘉宾 it would solve any intelligence problem. And like the classic example of LLMs is like next token prediction is AI complete because I could like, you know, there's the classic example, I think from Ilya, where like I there's a mystery novel and like the thing has to read the whole mystery novel and the final sentence of the mystery novel is like, "And the killer was Predict the next token. So, like you could basically like frame any kind of intelligence task in terms of that. So, clearly next token prediction is something that people believe is AI complete. 它就会解决任何智能问题。而 LLM 的经典例子是:next token prediction 是 AI complete,因为我可以,有一个经典例子,我想是来自 Ilya 的:有一本推理小说,这个东西必须读完整本推理小说,而推理小说的最后一句是。 「而杀手是」,预测 下一个 token。所以你基本上可以把任何一种智能任务都框进这个形式。所以很明显,next token prediction 是人们相信属于 AI complete 的东西。

主持人 Yeah, yeah. 对。

嘉宾 But I think that something we're kind of realizing and Ben was uh talking about this earlier today is like new view prediction, this primitive that we have in Atlas, especially generative new new view new view prediction. This is also AI complete. Right? And because I could take something like 但我觉得我们正在意识到的一件事,Ben 今天早些时候也在讲,就是 new view prediction,我们在 Atlas 里的这个原语,尤其是生成式的 new view prediction。这也是 AI complete。对吧?因为我可以拿这样的东西

主持人 You could you could have the the movie and you do all of the frames of the movie and then like the killer walks out and then you predict exactly who walks out. 你可以有这部电影,你做这部电影的所有帧,。然后杀手走出来,然后你精确预测是谁走出来。

嘉宾 Exactly. Not just that, but I could say like I want to have a world where like Martine is like writing a proof of the Riemann hypothesis on the 没错。还不只是那样,我还可以说:我想要一个世界,里面 Martine 正在写黎曼猜想的证明,在

主持人 So okay, so to take a evolutionary view, right? That new viewpoint prediction is exactly evolution had to solve by making animals move. You you nature give animals eyes. But nature didn't give trees eye. Eyes. Why? Because when you move, you see a new viewpoint. And that is the the whether you call it AI complete or intelligence complete. So So we do believe very strongly that next viewpoint prediction is is the equivalent of next token prediction. 所以好,从演化的视角来看,对吧?new viewpoint prediction 恰恰是演化必须通过让动物移动来解决的问题。自然给了动物眼睛。但自然没有给树眼睛。为什么?因为当你移动时,你会看到一个新视点。而这就是,无论你称之为 AI complete 还是 intelligence complete。所以我们非常强烈地相信,next viewpoint prediction 等价于 next token prediction。

嘉宾 Amazing. Well, with that, congratulations all of you on a phenomenal model launch. We're very excited for future model launches and thanks for coming. 太好了。那就这样,祝贺你们所有人这次非凡的模型发布。我们非常期待未来的模型发布,也感谢你们过来。

主持人 Thank you. 谢谢。

回到顶部