子弹时间和 new view prediction
00:00:00–00:08:00Insight
- 最多约 100 帧真实视角可以重建空间,结果可以是飞穿视频,也可以是显式三维(约 00:01–00:02)。
- 李飞飞把像素生成和像素重建的统一,放到计算机视觉半个多世纪的分赛道背景里:会议里生成、识别、三维重建一直分开(约 00:07–00:08)。
主持人 On the path to spatial intelligence, generating pixels that are truly spatially contextualized and grounded, that is the very hard step that Atlas has taken. 在通往空间智能的路上,生成真正具有空间上下文、并且落地的像素,这正是 Atlas 已经迈出的那一步,也是最难的一步。
嘉宾 You know, LLMs are built on next token prediction. We've seen video models as being built on next frame prediction. Atlas is really new view prediction. 你知道,LLM 建立在 next token prediction 之上。我们看到视频模型建立在 next frame prediction 之上。Atlas 真正做的是 new view prediction。
主持人 This is the real place where AI can actually unlock a ton of value for people and their process. We're saying like 50, 100 extra reduction. There's a famous shot in the first Matrix movie where Neo is like falling down. Exactly. 这才是 AI 真正能为人和他们的工作流程解锁大量价值的地方。我们说的是大约 50 倍、100 倍的额外缩减。第一部《黑客帝国》里有一个著名镜头,Neo 好像在往下掉。没错。
主持人 They had hundreds of cameras doing that angle on a green screen. With Atlas, we can do this with just three cameras. No studio capture, no green screen, no expensive calibration. 他们当时用了几百台相机,在绿幕前拍那个角度。有了 Atlas,我们只用三台相机就能做到。不需要棚拍,不需要绿幕,也不需要昂贵的标定。
嘉宾 No one has ever seen these results. 从来没有人见过这样的结果。
主持人 When you set out to do this, did you know it was going to work? 你们开始做这件事的时候,知道它会成功吗?
嘉宾 I was pretty sure. Each time we made the model bigger, and each time we trained it for longer, it got significantly better. 我相当有把握。每一次我们把模型做大,每一次我们训练得更久,它都会明显变好。
主持人 Does that mean we're going to get 4D video? Can I go walk around? 那是不是意味着我们会得到 4D 视频?我能走进去走一圈吗?
嘉宾 So, big day yesterday, you launched a new frontier model, which has got amazing reception, which is still coming in. I think maybe a good way to structure this conversation, let's just talk about exactly what that was, and then we'll go back to history and work our way back up. So, maybe Justin, do you want to talk about what was launched yesterday, why it's significant. 昨天是个大日子,你们发布了一个新的前沿模型,反响非常好,还在持续。进来。我觉得组织这次对话的一个好办法是,先准确讲讲昨天发布的是什么,然后再回到历史,一路往上讲。Justin,你要不要讲讲昨天发布了什么,以及它为什么重要。
主持人 Yeah, so Atlas is our new next generation world model. Um, it has three basic things. It can generate, reconstruct, and simulate the world. Um, so within that, there's a couple different major capabilities. It has really good camera condition generation. So, you can input an image together with a camera trajectory with a camera trajectory and steer the model and have it generate, you know, video frames along any any perspective you want. It's really good at sparse 3D reconstruction. You can input one or multiple up to 100 frames, um, that are views of the real world, and use those to reconstruct the real world. And that reconstruction can take the case either of of novel of video flying through the space, or an explicit 3D reconstruction of the space. Um, then finally, it can be used for simulation. Um, and for this, we show off, um, you know, these awesome bullet time videos, which got a lot of attention online, and then also robotics simulation. 好。Atlas 是我们新一代的 world model。它有三件基本能力:生成世界、重建世界、模拟世界。在这之下,还有几项主要能力。它的相机条件生成非常强。你可以输入一张图,再配上一条相机轨迹,用这条相机轨迹来引导模型,让它生成视频帧沿着你想要的任意视角。它在稀疏三维重建上也很强。你可以输入一帧,或多帧,最多 100 帧真实世界的视角,用它们来重建真实世界。重建的结果可以是一段在空间里飞过的视频,也可以是这个空间的显式三维重建。最后,它还可以用于仿真。在这方面,我们展示了那些很酷的子弹时间视频,网上关注度很高,另外还有机器人仿真。
嘉宾 What's a What's a bullet time video? 什么是子弹时间视频?
主持人 A bullet time video, this comes from the Matrix. You know, there there's a famous shot in the first Matrix movie where Neo is like falling down. 子弹时间视频来自《黑客帝国》。第一部《黑客帝国》里有一个著名镜头,Neo 好像在往下掉。
嘉宾 Oh, yeah. 对,对。
主持人 Exactly. So, and then remember in that famous shot he's like falling down, it's in slow motion, and the camera flies all the way around. Um so, that's the way that they did that shot is they had a ring of like hundreds of cameras. So, then like he fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that that famous shot in in Matrix. But now with Atlas we can do this with just as few as through three cameras. So, like no studio capture, no green screen, no expensive calibration. We can literally stick like three cameras on three iPhones on tripods, um use these to take sort of a video of something happening, um like someone shooting a basket, someone dropping a strawberry into a bowl of milk. And then from those three like iPhone videos, we can then reframe the shot, and imagine like a like freeze time, have the camera fly in like as the milk is splashing up, 没错。然后你记得那个著名镜头里,他往下掉,是慢动作,相机一路绕着他飞。他们拍那个镜头的办法是,摆了一圈大概几百台相机。他在棚里倒下,几百台相机在绿幕前拍那个角度,然后他们用那几百台相机做出了那个著名镜头,就是《黑客帝国》里的那个。但现在有了 Atlas,我们最少只用三台相机就能做到。不需要棚拍,不需要绿幕,也不需要昂贵的标定。我们真的可以把三台相机、三部 iPhone 架在三脚架上,用它们拍一段正在发生的事,比如有人投篮,有人把一颗草莓丢进一碗牛奶里。然后从这三段 iPhone 视频,我们可以重新构图,想象时间冻结,让相机飞进去,就在牛奶溅起来的时候,
主持人 and get these amazing frozen time views. Um and we can do this with just uh just a couple cameras. 得到这些非常惊人的冻结时间视角。而且我们只用几台相机就能做到。
嘉宾 Can Can you um just maybe What is the simplest description of what Atlas does? Like what goes in and what comes out? 你能不能用最简单的话说一下 Atlas 是做什么的?进去的是什么,出来的是什么?
主持人 Yeah, so one of the one of the really core principles of Atlas, like the most fundamental thing, is it does new view prediction. Um and this is a a really fundamental primitive that we think is super exciting, a super new primitive for for base models that no one's ever done before. Right? So, we know LLMs are built on next token prediction, we've seen video models as being built on next frame prediction. Atlas is really new view prediction. Right? That given some number of views of a scene or a description of a scene, um those go into what we call a spatial context that describes implicitly what is the world that we want to talk about, then you can point a virtual camera at at an arbitrary point in space and time, and Atlas will understand what that world is supposed to look like from that position in space and time. 好。Atlas 真正核心的原则,最根本的一件事,就是它做 new view prediction。我们认为这是一个非常根本的原语,对基础模型来说是一个全新的原语,以前没有人做过。对吧?我们知道 LLM 建立在 next token prediction 之上,我们也看到视频模型建立在 next frame prediction 之上。Atlas 真正做的是 new view prediction。对吧?给定一个场景的若干视角,或者对这个场景的一段描述,它们会进入我们所说的 spatial context,隐式地描述我们要讨论的那个世界是什么。然后你可以把一台虚拟相机指向时空中的任意一点,Atlas 就会理解,从这个时空位置看,那个世界应该是什么样子。
嘉宾 You know, you know, Ben, that you know, with with a with a bajillion video models out there all claiming to be world models and all claiming to have novel views, and can you maybe tease apart can more concretely how this is different from like the myriad models that have come before? Ben,你知道,外面有无数的视频。模型,都声称自己是 world model,都声称能做新视角,你能不能更具体地把这个和以前那一大堆模型的差别拆开讲清楚?
主持人 Mhm. Yeah, I think what Justin was saying about the spatial context aspect is super important here. So there's many video models, a lot of video models actually got their claim to fame from their single dimension input or their start to last frame interpolation. Now we're starting to see models that can do this kind of omni referencing with, you know, 20, 30, 50 images. But what's key with Atlas is that it actually has a kind of like specially grounded meaning to every frame you put into it. So it's not just an image that the model's going to interpret whatever way it wants or you can kind of try to argue with it in the the text prompting and get it to do something specific. With Atlas, every image actually has an associated three-dimensional camera pose and that means that you can perform this task of reconstruction with an extremely high degree of precision, right? So if we had 嗯。对,我觉得 Justin 刚才说的 spatial context 这一点在这里非常重要。有很多视频模型,不少视频模型成名靠的是单维输入,或者从首帧到末帧的插值。现在我们开始看到一些模型能够做这种全方位参考,用 20 张、30 张、50 张图。但 Atlas 的关键在于,你输入的每一帧实际上都有一种空间上落地的含义。所以它不只是一张图,让模型随便怎么理解,或者你靠文本提示去跟它争、让它做某件具体的事。在 Atlas 里,每张图都对应一个三维的相机位姿,这意味着你可以用极高的精度来做重建这件事。对吧?所以如果我们有
主持人 four views of this room, one at each corner, you can put those into the model and then get an exact replication of everything you see in this room and it's not going to guess what's in the other corner, like the relationship between things. It's just going to reproduce exactly what you give it. And you can also do that in a kind of creative or imaginative sense, too. If you take two photos from different, you know, AI generations or real-world locations, you can actually position and stage those to build these kind of intentionally directed fly-throughs that are really governed by exactly the precise place that you put the content you want and where the camera's going to look and travel, which is very different, I think, than the kind of like more slot machine effect you get of having to retry generations over and over with just that kind of higher level of text control you get with video models. 这个房间的四个视角,每个角落一张,你可以把它们放进模型,然后得到你在这个房间里看到的一切的精确复制。它不会去猜另一个角落里有什么,比如物体之间的关系。它就是精确复现你给它的东西。你也可以用一种创造性、想象性的方式来做这件事。如果你从不同的 AI 生成结果,或者真实世界的地点,各拿两张照片,你实际上可以摆放和布置它们,来构建这种有意安排的定向飞穿镜头,真正由你把想要的内容放在精确位置、以及相机要看向哪里、沿哪里走来决定。这和视频模型那种更像老虎机的效果很不一样,那种只能靠更高层的文本控制,一遍遍重试生成。
嘉宾 Is this just kind of an obvious, you know, scaled-up version of a traditional video model or is it a new architecture? 这只是传统视频模型一个显而易见的放大版,还是一种新架构?
主持人 I think it's it's a pretty new thing for a couple of different reasons. One that we talk about is it does both generation and reconstruction jointly in the same model. Like Ben was saying, this thing can take a couple of views of this room and then reconstruct everything in this room exactly as you see it. And historically, reconstruction has been its own subfield in computer vision with its own specialized task, its own specialized models. And generation is what all the text all the text to video models are really good at like what all the all the big diffusion models we've seen the last couple years. And those are great for creative applications. I want to imagine 我觉得这是一件相当新的事,原因有。几个不同的方面。我们常说的一点是,它在同一个模型里同时做生成和重建。就像 Ben 说的,它可以拿这个房间的几个视角,然后把这个房间里的一切按你看到的样子精确重建出来。历史上,重建一直是计算机视觉里一个独立的子领域,有自己专门的任务,有自己专门的模型。而生成是所有文生视频模型真正擅长的事,也就是我们过去几年看到的那些大型 diffusion 模型。那些对创意应用很好。我想想象
主持人 something that's never been there before. Um but now with Atlas for the first time we're putting these two different parts of visual intelligence together in one model. So it can do both 3D reconstruction and generation together in one architecture. So to do that we had to make a couple changes. Um one is we had to make it multimodal from the start. So this thing natively works on text, it works on images, it works on videos. It also works on camera poses um as a native input to the model which I don't think anyone's ever done at the pre-training phase before. Um and it uses um it uses 3D as a native modality that it works on. So this thing from the beginning was designed to be natively multimodal in a way that no one else I think 以前从未存在过的东西。但现在有了 Atlas,我们第一次把视觉智能的这两个不同部分放进同一个模型。所以它可以在同一套架构里同时做三维重建和生成。要做到这一点,我们做了几处改动。一是我们必须从一开始就做成多模态。所以它原生处理文本,处理图像,处理视频。它也把相机位姿作为模型的原生输入,我认为在预训练阶段以前没有人这样做过。而且它把 3D 作为自己处理的一种原生模态。所以这个东西从一开始就被设计成原生多模态,我认为没有别人是这样
嘉宾 Sorry, I just um I don't know this space super well but 3D is this like depth or models and like what is that what is that mean? 抱歉,我对这个领域不是特别熟,3D 是指深度,还是模型,这到底是什么意思?
主持人 Yeah, so the formulation we used so far is depth maps. 对,我们目前用的表述是深度图。
嘉宾 Okay. 好。
主持人 Right? So um right now you can have a when you have a frame that has a a a virtual camera telling its position in 3D space, yeah, that camera position and camera parameters are a native input to the model. And then what attached to that camera position you can have both um like RGB telling you what is that position in space look like and you can have a depth map that tells you what is the spatial structure of that position in 3D space. So then you know text, image, video, 3D cameras are these modalities that this thing all does jointly in a multimodal way. 对。所以现在,当你有一帧,上面有一台虚拟相机告诉你它在三维空间里的位置,对,那个相机位置和相机参数是模型的原生输入。然后附着在这个相机位置上,你可以既有 RGB,告诉你那个空间位置看起来是什么样,也可以有一张深度图,告诉你那个位置在三维空间里的空间结构。所以文本、图像、视频、三维相机,这些模态这个模型都是以多模态的方式一起处理的。
嘉宾 I want to add something cuz I think what Justin just said is actually so important and also what Ben said that it's under appreciated. It's the first time we have a unification of pixel generation and pixel reconstruction. In the world of computer vision this field has been around for more than half a century. Um sitting here having been in this field for decades, I cannot tell you how many PhD thesis have been written on the 我想补充一点,因为我觉得。 Justin 刚才说的其实非常重要,Ben 说的那一点也被低估了。这是我们第一次把像素生成和像素重建统一起来。在计算机视觉这个领域,已经存在了半个多世纪。我坐在这里,在这个领域待了几十年,我说不清有多少篇博士论文写的就是
嘉宾 problem of reconstruction or novel view things synthesis and also our field traditionally uh have multiple tracks. You go to a computer vision conference, you have the pixel generation track, you have some recognition track and you have 3D reconstruction track. This is a elegant model that combines or unifies the problem of reconstruction and generation by anchoring on viewpoints and the the viewpoint estimation and that's just incredibly powerful. 重建问题,或者新视角合成。而且我们这个领域传统上有多条赛道。你去一个计算机视觉会议,会有像素生成赛道,会有识别赛道,还会有三维重建赛道。这是一个优雅的模型,它把重建和生成结合起来、统一起来,锚点是视点和视点估计,这非常强大。