OLMo 3 的速度问题,本质上也是组织容量问题
00:00:06–00:06:28Insight
- Nathan 认为现代 Post-Training 的关键能力之一,是把 compute 与 data 编排成工作流;这直接映射到组织协调能力(约 00:02:06–00:03:56)。
- Finbarr 指出九个月的模型周转并非异常缓慢;Nathan 的反驳是,如果只是移植旧 recipe,这个速度对应的能力上限仍偏低(约 00:03:56–00:05:52)。
- 对于 7B–30B 级模型,是否值得完整复刻 DeepSeek 式 RL-first 路线仍未得到证据充分的答案(约 00:05:22–00:06:11)。
Nathan Lambert Hello, we are back on a Interconnects conversation. I don't really say I do interviews. People criticize me because I interrupt the guests too much. I'm not a good interviewer, but I'm here to entertain people. This is also fun for me because I'm trying to make like a post-training course and it kind of fits as the advanced end of this. So it's kind of a crossover between Interconnects content and other stuff that I've been spending my time on this summer. So I'm happy to wel
come Finbar back, I think.
Are you the first return guest?
I haven't checked. Oh, wow. Finbar and I worked on this sort of post-training recipe stuff for a while at AI2. I left recently. This is one of Finbar's last days at AI2. It's already been announced. It's not a spoiler here. So we're going to kind of reflect on some things on building post-training recipes for Ulmo. Then we have a little review slide deck and notes on The kind of state and evolution of frontier post trading recipes over time, which is pretty interesting because there's, what is it like two to four kind of canonical recipes that there has been. So it's kind of interesting when you see the field converge on something new, which it's doing right now with multi teacher on policy. distillation for some reason that's a bit of a mouthful it is a long acronym and then we'll just kind of end with various discussion points on post-t raining and what we're up to so happy to give you a floor if you have any hot takes you want to start with to get people to draw people in otherwise I think I'm excited to kind of reflect on this because I know you've been reading a ton of papers recently and kind of prep laying some of this groundwork
Finbarr Timbers Well, yeah, I mean, today is my last day at AI2. So it feels very appropriate to be talking to you as you're the one who recruited me to AI2. So, yeah, that's pretty special. And it's great to be, yeah, the first repeat guest. I feel honored to be back on. So, yeah, thanks for having me.
Nathan Lambert Yeah. Do we want to start with Ulmo? I think that people... I need to do this carefully, but I've talked about Olmo 3's post trading many times to people. I haven't done this in a very direct way on the podcast, but I would say that post trading Olmo 3 to make this reasoning model was a... Major accomplishment for many individuals to do this, but also the complexity of what we were doing was pushing against the limits of AI2's organizational capacity.
And a lot of modern post-training is like your ability to wrangle compute data into a work stream. And in order to do that in a complicated way, you really are wrangling an org chart. and that's like part of why it's like old boat three was by its nature pretty late as a reasoning model it was like pretty rigid reasoning model and that's like partially reflected in the recipe being pretty simple but then when you like compare it to all these new recipes with tool use and mult
i-teacher distillation and all of this it's just like A fork in the road where it's like you could do this very simple thing and make a strong recipe, but it is not representative of what all the Frontier Labs are doing. And I think that that kind of fork in being able to say that things are similar happened kind of after Tulu 3, where Tulu 3, I think... was also much simpler with this three-stage SFT DPO RL recipe but that simpler recipe was probably closer an outcome to wha
t the labs are doing but now doing that sort of three-stage recipe for a reasoning model and especially a tool use like agent model just doesn't really apply and that's the point that's why I think the point of this podcast to be like what are the what are the way what are they doing to make these like true frontier models and then shed some light on how it contrasts to the more like open academic ones
Finbarr Timbers Well, actually, I think that's interesting. What was the process? So, you know, I only came around for Almo 3. I wasn't around for the earlier versions. What was the process like to go from Tulu 3 to Almo 2? Because, like, just looking at an archive, I think Tulu 3 came out in November of 24. And then Almo 2 came out in December of 24.
Nathan Lambert We just applied the recipe. Yeah.
Finbarr Timbers Yeah. I mean, so I think that actually, like, and then, you know, Deep Seagull 1 came out in January, end of January 25. And, you know, Almo 3 was then released in October. It was October, November of 25. I think November. Yeah, November.
Nathan Lambert Yeah, right, it was November. It was like do or die with Thanksgiving. I remember that.
Finbarr Timbers yeah because Canadian Thanksgiving had already happened which yeah I was happy but like I think it was sure maybe it was late but I think it was only late by a few months like it's actually like you know if I think of my past experience with model turnaround times like a nine month model turnaround you know from R1 coming out like that's actually that's not bad I think you know something like six months would have been
Nathan Lambert Nicer I think it's slow because we didn't it would be fast if we had rebuilt the R1 recipe but what we did was we like ported reasoning into our existing recipe okay which is a simpler task but has like a lower ceiling in my opinion where it's like the deep seek in the newer style recipes which we'll talk about I think they just have a much higher ceiling and how much you keep hill climbing them or they're just like more prescribed more pedagogical of what the frontier is doi
ng,like for the size models that almost was, which was like seven to 30 B. I'm not sure that doing this deep seek style RL first recipe is actually useful.
Finbarr Timbers Yeah, I think that's a good point. And I mean, I think that's really reflected in what we see the research where you see, you know, you obviously you see the big, the step change, and you know, how quickly things are improving. when R1 comes out. So I think that's a great point. And it really does seem to saturate or to not saturate, sorry, with compute.
Nathan Lambert Yeah. Should we just do the slide deck?