# Frontier post-training recipe review with Finbarr Timbers

- Original show: Interconnects
- Original source: https://www.interconnects.ai/p/frontier-post-training-recipe-review
- Discovery source: https://www.xiaoyuzhoufm.com/episode/6a809e6117676351c5726dfe
- Duration: 00:56:36
- Method: Transcript published by the original creator; its stated review status is unknown, so quotations still require source review.

## Transcript

[00:00:06–00:00:34] **Nathan Lambert:** Hello, we are back on a Interconnects conversation. I don't really say I do interviews. People criticize me because I interrupt the guests too much. I'm not a good interviewer, but I'm here to entertain people. This is also fun for me because I'm trying to make like a post-training course and it kind of fits as the advanced end of this. So it's kind of a crossover between Interconnects content and other stuff that I've been spending my time on this summer. So I'm happy to welcome Finbar back, I think.
[00:00:35–00:01:03] **Nathan Lambert:** Are you the first return guest? I haven't checked. Oh, wow. Finbar and I worked on this sort of post-training recipe stuff for a while at AI2. I left recently. This is one of Finbar's last days at AI2. It's already been announced. It's not a spoiler here. So we're going to kind of reflect on some things on building post-training recipes for Ulmo. Then we have a little review slide deck and notes on
[00:01:04–00:01:32] **Nathan Lambert:** The kind of state and evolution of frontier post trading recipes over time, which is pretty interesting because there's, what is it like two to four kind of canonical recipes that there has been. So it's kind of interesting when you see the field converge on something new, which it's doing right now with multi teacher on policy. distillation for some reason that's a bit of a mouthful it is a long acronym and then we'll just kind of end with various discussion points on post-training and what we're up to so happy to give you a floor if you have any hot takes you want to
[00:01:32–00:01:43] **Nathan Lambert:** start with to get people to draw people in otherwise I think I'm excited to kind of reflect on this because I know you've been reading a ton of papers recently and kind of prep laying some of this groundwork
[00:01:43–00:02:02] **Finbarr Timbers:** Well, yeah, I mean, today is my last day at AI2. So it feels very appropriate to be talking to you as you're the one who recruited me to AI2. So, yeah, that's pretty special. And it's great to be, yeah, the first repeat guest. I feel honored to be back on. So, yeah, thanks for having me.
[00:02:03–00:02:33] **Nathan Lambert:** Yeah. Do we want to start with Ulmo? I think that people... I need to do this carefully, but I've talked about Olmo 3's post trading many times to people. I haven't done this in a very direct way on the podcast, but I would say that post trading Olmo 3 to make this reasoning model was a... Major accomplishment for many individuals to do this, but also the complexity of what we were doing was pushing against the limits of AI2's organizational capacity.
[00:02:33–00:03:03] **Nathan Lambert:** And a lot of modern post-training is like your ability to wrangle compute data into a work stream. And in order to do that in a complicated way, you really are wrangling an org chart. and that's like part of why it's like old boat three was by its nature pretty late as a reasoning model it was like pretty rigid reasoning model and that's like partially reflected in the recipe being pretty simple but then when you like compare it to all these new recipes with tool use and multi-teacher distillation
[00:03:03–00:03:32] **Nathan Lambert:** and all of this it's just like A fork in the road where it's like you could do this very simple thing and make a strong recipe, but it is not representative of what all the Frontier Labs are doing. And I think that that kind of fork in being able to say that things are similar happened kind of after Tulu 3, where Tulu 3, I think... was also much simpler with this three-stage SFT DPO RL recipe but that simpler
[00:03:32–00:03:56] **Nathan Lambert:** recipe was probably closer an outcome to what the labs are doing but now doing that sort of three-stage recipe for a reasoning model and especially a tool use like agent model just doesn't really apply and that's the point that's why I think the point of this podcast to be like what are the what are the way what are they doing to make these like true frontier models and then shed some light on how it contrasts to the more like open academic ones
[00:03:56–00:04:22] **Finbarr Timbers:** Well, actually, I think that's interesting. What was the process? So, you know, I only came around for Almo 3. I wasn't around for the earlier versions. What was the process like to go from Tulu 3 to Almo 2? Because, like, just looking at an archive, I think Tulu 3 came out in November of 24. And then Almo 2 came out in December of 24.
[00:04:23–00:04:24] **Nathan Lambert:** We just applied the recipe. Yeah.
[00:04:25–00:04:42] **Finbarr Timbers:** Yeah. I mean, so I think that actually, like, and then, you know, Deep Seagull 1 came out in January, end of January 25. And, you know, Almo 3 was then released in October. It was October, November of 25. I think November. Yeah, November.
[00:04:42–00:04:46] **Nathan Lambert:** Yeah, right, it was November. It was like do or die with Thanksgiving. I remember that.
[00:04:47–00:05:11] **Finbarr Timbers:** yeah because Canadian Thanksgiving had already happened which yeah I was happy but like I think it was sure maybe it was late but I think it was only late by a few months like it's actually like you know if I think of my past experience with model turnaround times like a nine month model turnaround you know from R1 coming out like that's actually that's not bad I think you know something like six months would have been
[00:05:11–00:05:40] **Nathan Lambert:** Nicer I think it's slow because we didn't it would be fast if we had rebuilt the R1 recipe but what we did was we like ported reasoning into our existing recipe okay which is a simpler task but has like a lower ceiling in my opinion where it's like the deep seek in the newer style recipes which we'll talk about I think they just have a much higher ceiling and how much you keep hill climbing them or they're just like more prescribed more pedagogical of what the frontier is doing,
[00:05:40–00:05:50] **Nathan Lambert:** like for the size models that almost was, which was like seven to 30 B. I'm not sure that doing this deep seek style RL first recipe is actually useful.
[00:05:50–00:06:10] **Finbarr Timbers:** Yeah, I think that's a good point. And I mean, I think that's really reflected in what we see the research where you see, you know, you obviously you see the big, the step change, and you know, how quickly things are improving. when R1 comes out. So I think that's a great point. And it really does seem to saturate or to not saturate, sorry, with compute.
[00:06:11–00:06:39] **Nathan Lambert:** Yeah. Should we just do the slide deck? We're throwing around like recipe names. I feel like it might be useful to just do it because a lot of people probably want to follow but don't exactly know. I'm going to share it. I'm going to share a screen so people listening it might be useful to either you can pull the slide deck up on your phone and click through it it's not super information dense but you can also just watch it on YouTube all this will be linked
[00:06:39–00:07:09] **Nathan Lambert:** generally this is just like a quick survey on how frontier recipes have evolved we'll go through the history quickly and then talk about what is currently happening and kind of probably interleave the Olmo discussion we were having okay There's a bunch of canonical recipes we'll talk about. This is where I got the two to four number. I think the recipes are like InstructGPT, which is what coined the initial RLHF with this three-stage idea, which took a while to get people to move on from, which is like SFT,
[00:07:09–00:07:37] **Nathan Lambert:** reward model, and RL. And I see as like Llama 3 and Tulu 3 as kind of practical implementations of that with other tricks of the trade. So those two could potentially be merged together. It's like just like kind of pre and post ChatGPT moment. And then the two most recent canonical recipes that we'll cover in this, I would say, are like DeepSeek R1, which is the shift to doing like reasoning focused and bigger RL stages than this kind of SFT focus from before. And then
[00:07:38–00:07:42] **Nathan Lambert:** Mimo Flash and some of the new models from 2026, which add this distillation element.
[00:07:42–00:08:03] **Finbarr Timbers:** Well, and I think it's worth pointing out, too, that it's not just Mimo Flash. Like, it was kind of a consistent theme. Like, you saw this with DeepSeq. They referenced it in the V3 paper, and then it's, you know, it's Kimi K2.5, it's GLM-5. Like, it's all of these papers, you know, start talking about this specialist RL stage.
[00:08:03–00:08:21] **Nathan Lambert:** Yeah, I think there's a debate on how we draw it and whether or not distillation is... If you have distillation as a technique, as a key milestone, then Xiaomi was the first. But it's kind of a march over time where you kind of see them change. And we'll go through this. I don't need to...
[00:08:24–00:08:52] **Finbarr Timbers:** When you say distillation, I do think it's important to distinguish between the straight up distillation of the leading closed models and distillation of these domain specific models, where I suspect that the Chinese labs are doing both. But a lot of what they're doing is this training these domain-specific models, like a math model, a coding model, logic model, whatever,
[00:08:53–00:09:01] **Finbarr Timbers:** and then distilling those models back in and not just distilling from – so when we're talking about distillation, it's not just distilling from the leading closed models.
[00:09:01–00:09:28] **Nathan Lambert:** Yeah, it's a pain. I agree. The distillation term is horribly overloaded. um there's a review slide do we need to review multi-teacher on policy distribution it might be too complicated to need to do it we could come back to it i think i kind of want to just go through the actual models and then we could use the supporting slides as needed um This famous InstructGPT three-step thing, I think many people have heard of it,
[00:09:28–00:09:53] **Nathan Lambert:** but this is what constituted post-training at the time of ChatGPT coming out, so it's kind of important grounding of this human-supervised SFT data. mostly human supervised preference rankings to make a reward model and then do RL on that and the model gets better. And it's pretty interesting how all of these have been kind of phased out, at least in terms of what we know openly, where they're
[00:09:54–00:10:23] **Nathan Lambert:** we don't use that much human demonstration data for SFT there's likely some human preference data still in the loop but I would guess that synthetic has a much bigger role and there are reward models but they're like not the key RL target anymore so in four years most almost all the canonical pieces have been moved on and like this evolution is kind of within there I think the early models after InstructGPT like Llama 2
[00:10:24–00:10:47] **Nathan Lambert:** even Lava 3 these are pretty similar which is like you're starting to break down this recipe with different tools like rejection sampling DPO some increased iterations I think increased iterations is just that there was more incentive to squeeze more out of the models and they just like broke things down more where InstructGPT seemed like a bit more open-ended research where this kind of cleanness
[00:10:47–00:11:13] **Finbarr Timbers:** was fine well I think that's interesting with respect to how much everything is scaled Right, because InstructGPT was before ChatGPT was released. And so, you know, it's something like just the complexity of what was done is that which a small team or a single team could do. But then when you start looking at, you know, Llama 3, like it just starts to be a more complicated process.
[00:11:13–00:11:24] **Finbarr Timbers:** process and where you start to have a lot more specialized data and there's a lot more room for scale and for money and complexity being poured in.
[00:11:25–00:11:51] **Nathan Lambert:** Yeah, it's like both for-profit and non-profit efforts to do post-training want me to advise them. And I'm like, I don't really know how I'm going to give you advice unless I'm spending 20 hours a week understanding the details of your recipe. Because it's like, I can't really give you a one-sentence thing of do X without understanding all the complexities of the model and the post-training process that go into it. Which makes it hard for me.
[00:11:53–00:11:59] **Nathan Lambert:** Kind of like a transparency point of view, even if it's fully detailed, it's definitely still hard to modify and study.
[00:12:00–00:12:03] **Finbarr Timbers:** Absolutely.
[00:12:03–00:12:33] **Nathan Lambert:** So then like two to three at AI2, a lot of this was we're trying to beat the results of this Lama 3 post training, which is pretty complicated. But we don't have the ability to scale the organization as far. So I think that's a big reason why the actual workflow is a lot simpler where we have three clear stages that are doing slightly different things and they build on each other. And that's like It's never stated very explicitly in these papers on how the org chart impacts the
[00:12:33–00:12:43] **Nathan Lambert:** recipe, but I think it's a very strong signal within at least the delta between the fully open work and the kind of partially open work that you get from industry.
[00:12:44–00:12:55] **Finbarr Timbers:** Yeah, absolutely. And I think especially as we'll see with the domain-specific models, that's something where you could really easily scale up your org chart to scale that up.
[00:12:56–00:13:21] **Nathan Lambert:** Yeah, and I threw Ulmo 3 in after this, after the 2 to 3 slide, mostly just to show that the recipe was so similar to 2 to 3, and the org chart had really changed. Like, we didn't have more ability to scale. I think there was a little bit of separation between the model types, between, like, the think and the instruct models, but, like, without a major org change, it was just kind of stuck in this and do the best you can with it.
[00:13:22–00:13:23] **Finbarr Timbers:** Yeah, true.
[00:13:24–00:13:51] **Nathan Lambert:** because like the real big change was this with Deep Seek R1 I had never seen this plot before but they had this plot maybe they added it for the nature version of the paper where they kind of show their recipe where they like take the base model they do RL0 and then they sample from the RL0 to like filter prompts and then they use that as SFT this is like going through this they use that as SFT for the next version of the model to create like a development internal RL Deep Seek R1 and then
[00:13:51–00:14:11] **Nathan Lambert:** they do this like repeated sampling to trade multiple RL versions and kind of distill in the sense of clarify and refine the reasoning behavior of the model before going through the final pipeline, which again is a mix of reasoning and non-reasoning SFT into a bigger RL realm.
[00:14:12–00:14:39] **Finbarr Timbers:** Well, I think this is really interesting because it starts to show, I mean, first of all, the complexity here. We're starting to use synthetic data as this primary input here, but it's not just... Like, you know, it's trying to elicit, you know, specific behaviors and it's this kind of like industrial process instead of like this, you know, it's not as much of an elegant research recipe. It's more like, you know, we train a model and then we use it as best we can and we keep iterating.
[00:14:40–00:15:06] **Finbarr Timbers:** But I think the other thing that's interesting is we're starting to see here the SFT serving as the cold start. First of all, I think before SFT was more of a generally useful stage, whereas here its primary purpose is this cold start for RL. And then the other interesting bit is DPO starts to disappear at this point from the Yeah, so my hypothesis for the dropping of DPO on these models is that...
[00:15:34–00:16:02] **Nathan Lambert:** As you're doing like a cleaner recipe, essentially the need falls away versus if you look at Ulmo, which is taking tons of potential gains by refining your model on outputs of strong open weight models, like largely Quen and DeepSeq is the training data for the SFT of Ulmo 3.0. The delta between that SFT data and the base model is still pretty big in the probability distribution. So DPO kind of helps further refine and clean up that distribution in a way that has very rough edges.
[00:16:03–00:16:30] **Nathan Lambert:** But when you have a more refined industrial process on post-training, that potential benefit will be harder to gain. Something interesting that I didn't fully confirm before this is, for example, NVIDIA used to also be on this DPO train with their smaller Nematron models. And I would guess that potentially Nematron Ultra would not.
[00:16:31–00:16:58] **Nathan Lambert:** And that's because they're much further down this development tree and using these more on-policy methods for creating the SFT data. And their model, I would guess, will become kind of more robust out of distribution and have less weird rough edges because of it. So that's kind of my hypothesis on DPO and People that use DPO will be looked down upon, but if you're trying to bootstrap a recipe off the ground and just take gains where you can,
[00:16:58–00:17:05] **Nathan Lambert:** I still think it'll work for a lot of people in a kind of compute efficiency standpoint.
[00:17:06–00:17:33] **Finbarr Timbers:** I think generally there's something interesting with the preference tuning that maybe it isn't being given the proper respect that it deserves because one of the interesting bits about the Nemetron 3 super paper was that they saw they do a traditional RLHF stage in their RL which has also you know fallen with fashion and they see pretty massive gains with it so
[00:17:33–00:17:41] **Finbarr Timbers:** I think some of these changes are more you know driven by what's in fashion rather than perhaps like a fully rigorous you know set of ablations
[00:17:42–00:18:01] **Nathan Lambert:** It's pretty remarkable to me that the preferences loss function can do so much for these models. The models have so much potential there, and it's really a contrastive loss on pretty granular feedback. They learn all sorts of things. They'll get better at math and code, or their reasoning strategies will be refined.
[00:18:03–00:18:04] **Finbarr Timbers:** That's remarkable to me.
[00:18:04–00:18:19] **Nathan Lambert:** I think there will still be funny research on using preference-based losses with verifiable outputs. I think all of this would work, like DPO on verifiable rewards and stuff like this. It's just kind of intellectually less appealing.
[00:18:20–00:18:49] **Finbarr Timbers:** Yeah, well, I think that's where I thought that the Delta Learning hypothesis-style DPO, like what ALMO 3 Did where you were the the preference you create these synthetic preferences by having like strong by like bigger and smaller models of the same family like is where you get your preferences from. I thought that was a really interesting signal because it seems really analogous to some of the work some of the guidance stuff that we see in diffusion models like
[00:18:49–00:19:17] **Finbarr Timbers:** how you have the classifier free guidance which has something similar and there were very similar results there which showed that you could have the like one signal they used was further along in training versus earlier in training models as like a source of signal that you could guide along and that worked quite well and so I suspect that these signals um for preferences in that way like that they could actually be more robust but because you know some of the largest labs don't have to do that perhaps we're not
[00:19:17–00:19:43] **Nathan Lambert:** setting them as much yeah or they don't tell us like to continue this it's kind of cool to look at so the DeepSeq models have kind of gone through this what I would call it closer to llama recipes to DeepSeq R1 which is like most definitively the canonical recipe for reasoning models and then continue to change closer to this multi-teacher format. So if you look at the V3-3 paper before R1,
[00:19:44–00:20:06] **Nathan Lambert:** they do something remarkably similar to Tulu-3 type thing where they have a mix of SFT and then they use like this RL on verifiable rewards. They didn't call it that or their paper wasn't out at the time. And So they did this before R1 came out, which was just kind of a less reasoning focused models and use the same tools, but with a different ratio of implementation weight.
[00:20:08–00:20:16] **Finbarr Timbers:** And what's interesting is that this comes out basically at the same time as 2 to 3. And it's a very similar 2 to 3 and Almo 2. It's a very similar recipe, just done with multiple people.
[00:20:17–00:20:46] **Nathan Lambert:** Yeah, yeah. And then we have this R1, which we've just talked about at length in January, which is a month later. They have a few more releases through this. They have some updates to their V3 and R1 models, which have dates, which are largely the same recipe. And then the next documented change in their recipe was V3.1, which is when they merged this thinking and non-thinking into one model, which everybody that does this has said that it has been hell to train in. But you kind of need it from a serving perspective. And it's obvious that long-term...
[00:20:47–00:21:13] **Nathan Lambert:** It's obvious to me that long term all the models will be reasoning models and you'll just have reasoning models that are very efficient based on the gains that are there. So this is kind of a needed change that they made and then in December of 2025 they released v3.2 which is when there's kind of meaningful changes to their recipe and they're talking about this expert creation with separate mini recipes and then
[00:21:14–00:21:43] **Nathan Lambert:** using that within their kind of R1 data process to do SFT data and then like a big RL run at the end with GRPO. So it took about a year for this Like kind of evolution of the R1 style recipe to land in their models. And I think this is like a very big complexity step that isn't represented in something like Olmo 3. And it's kind of where you can see a fork in the recipes over time as they become
[00:21:43–00:21:46] **Nathan Lambert:** way more industrial and scaled at these frontier labs.
[00:21:47–00:22:15] **Finbarr Timbers:** Yeah, and I think another one here just from historical note is that I think it was with the O324 release where they updated the original V3 paper. So, you know, V3 comes up before R1, then R1 comes out. And then after R1 comes out, they actually go back and update the V3 paper, maybe getting ready for the nature submission. or something. They make a reference there to say like, oh, you know, something you could do is you could train these domain specialist models and then combine them.
[00:22:15–00:22:21] **Finbarr Timbers:** And then, you know, that later becomes kind of more of a priority as they talk about in V3.2.
[00:22:22–00:22:49] **Nathan Lambert:** That's a fun note. Yeah, and then More recently in April 26 is this V4 model, which has even more experts. They add this new loss function for multi-teacher on policy distillation, which I said follow Xiaomi. And this is kind of a microcosm of the arc that the whole industry went through, at least the people who share what their post-training details are of realizing how core RL is changing the recipe around scaled RL and then figuring out
[00:22:49–00:23:18] **Nathan Lambert:** how to kind of scale to more domains in the scaled RL format without just like grinding to a halt in operational complexity yeah so that kind of the next stage of this is these what I call 2026 style recipes which are all these models that are doing this multi-teacher infusion of knowledge. Some of them are using on-policy distillation and some are not. One of the key things to see is how crucial is this on-policy distillation to
[00:23:18–00:23:42] **Nathan Lambert:** really keeping up with the frontier. The paper that named this term was the MIMO Flash V2 paper. I think the model was released in December and the paper in January, which a lot of things will look similar to this kind of large RL style recipe but with this large RL run is where the on policy
[00:23:42–00:24:08] **Nathan Lambert:** distillation comes in so for this is probably a better time to explain I have this great little feature so this is like the summary of what multi-teacher on policy distillation is. Generally it fits within an RL framework where you have the model you are training, like the general model, sample its own trajectories, and then you route the trajectories to various expert models you have trained.
[00:24:09–00:24:36] **Nathan Lambert:** And each kind of sample is trained with this distillation KL loss to match the tokens of that expert. And People have multiple models have shown that this type of supervision is really useful for the models. You could combine it with other RL losses, such as verifiable rewards, which, for example, Sasha Rush gave a good mini spiel on that and how they use that with Composer, which is a video that I really recommend people watching as well.
[00:24:37–00:24:42] **Nathan Lambert:** But the key of it is that it is a different loss function, but it plays very nicely in the RL frameworks that people are already using.
[00:24:43–00:25:00] **Finbarr Timbers:** So they use these teachers. It's an URL. I'm talking with some of the people at AI2, but implementing it now. And it's like you take your URL setup and then you just use a very set of tweaks on the learner to actually implement this. So it's quite straightforward.
[00:25:03–00:25:29] **Nathan Lambert:** Yeah. So this is a fancy diagram that makes it more complicated than it needs to be, but it's also a very nice diagram. which shows the various domain teachers that they have, cert agent, code agent, math, reasoning safety, and how they put these together. And the experts are used both for SFT data and then this final supervision. And the recipe for the experts would look something like this deep seek recipe, which is complicated on its own, which is like make a very good reasoning model that is good at one thing.
[00:25:30–00:25:59] **Finbarr Timbers:** Well, I think it is complicated, but it's also like if you think about being the actual researcher, like working on it, it's like... you know you have a base model and then you have an RL setup and you know you're just constantly updating both and then rerunning RL so you know the most complicated part of it is just you know writing down the history and tracing everything that's kind of like a very natural organic way for the RL to evolve through you know iterative experimentation yeah so like once you have a recipe
[00:25:59–00:26:27] **Nathan Lambert:** you're progressively tinkering with each part and it's fairly stable but it's hard to rebuild from scratch so like See how long the recipe shape lasts, but it'll probably be order of years. Another big one in this that also shared a lot of details on this on-policy distillation approach was Nebatron 3 Ultra, which is obviously exciting to me to have a US-made model that is very strong
[00:26:27–00:26:55] **Nathan Lambert:** performance and NVIDIA released a lot of data sets with it. But they also talked about a lot of their very... implementation details of what was hard with on-policy distillation. I have notes somewhere on this. They do this thing where they have two rounds of on-policy distillation. They found it to be better to integrate some teachers one after another. The paper has a lot more details. I don't want to go scroll through the paper, but we could also do this.
[00:26:55–00:27:03] **Nathan Lambert:** Did you have any other impressions? We have this other doc we can pull up that also might have had other details on it.
[00:27:03–00:27:27] **Finbarr Timbers:** Yeah, well, I think something else that is worth, you know, contrasting the paper to is the Nemotron 3 super paper. Because in the Nemotron 3 super paper, they had a similar complicated recipe, but they did multiple... rounds of RL like there they had three rounds of RLBR followed by a round of
[00:27:29–00:27:56] **Finbarr Timbers:** software engineering RL and then followed by an RLHF stage so it was it was really interesting to see them go from doing that like you know one of the most complicated RL setups or in terms of you know successive stages that I've seen to then you know you know this setup where it's still complicated but it's a lot um you know it's a lot conceptually a lot simpler yeah I pulled the paper up it's
[00:27:56–00:28:21] **Nathan Lambert:** going to be hard for me to like I had highlighted a few details that the interesting parts are kind of around the various NVIDIA details on all the teachers there's just so many details in their paper on training all the teachers I think okay so I have some of it I have some of this up it's like I have an interesting quote that's like one key finding from our So I think they'd have to do some cross-teacher alignment
[00:28:35–00:28:59] **Nathan Lambert:** to make sure that they're actually similar, which I feel like could become a whole organizational nightmare. It's like they say, we hypothesize that when the teacher and student are trained on different SFT data, they acquire different reasoning behaviors and induce different output distributions. This distribution matrix can cause student generated trajectories to be out of distribution for the teacher, reducing the quality and reliability of the supervision signals provided by the teacher.
[00:29:00–00:29:18] **Finbarr Timbers:** Yeah, that's interesting, actually, because there was a paper, I can't remember the name of it, but there was a paper that I read recently, which claimed that what you need to do is constantly, so one thing you could do, which was kind of the obvious thing, Nathan Lambert natolambert.com
[00:29:32–00:29:58] **Finbarr Timbers:** and then you take these final experts and then you do some sort of, you know, on policy distillation to combine them into your final model. But with the paper, and I'll try to find it and then give it to you and see if we can share it. What they claimed was that you need to... Instead of using the converged model, you need to do it in successive stages with the in-progress model. So if you change your RL for 1,000 steps,
[00:29:58–00:30:16] **Finbarr Timbers:** you can't use the 1,000-step checkpoint for the on-policy distillation. You have to do it in stages. At first, use the 250-step checkpoint and the 500-step checkpoint, and gradually bring that base model up to speed, or else there's going to be too much divergence, and the KL divergence will just be too...
[00:30:17–00:30:46] **Nathan Lambert:** um too distinct to learn from yeah so essentially the last sentence in this paragraph I had read most of is literally like we encountered this issue in practice because the teacher and student models were developed in parallel it's like they're like this is a problem because of it's like hard to do everything at once which is this is the type of thing where having research in it would be so great and I think a video could release some of the teachers and then people could just like if you have the
[00:30:46–00:31:16] **Nathan Lambert:** teachers and you have the intermediate model stage you could do the problem of like just studying multi-teacher on policy distillation from the starting point and understanding the training dynamics which is the type of thing we would want to do at Olbo we just haven't scaled our recipe to this point yet yeah absolutely so I will keep encouraging NVIDIA to do this that'd be great NVIDIA listen They listen. The other side of things is a bunch of models released in 2026 that do not do this
[00:31:17–00:31:45] **Nathan Lambert:** multi-teacher on policy distillation and they also don't do nearly as many teachers. So I would say that this Microsoft model, which... I don't say this as a diss it's like hard to get a new team off the ground is they went for a simpler approach to try to get a solid model and it has three more general experts combined via SFT and then like a longer RL run so it looks a lot more like DeepSeek R1 but I suspect that what they will do next is make finer grade
[00:31:45–00:31:48] **Nathan Lambert:** teachers and see if they need to switch to on policy distillation
[00:31:49–00:32:09] **Finbarr Timbers:** Yeah, and I think, you know, in one of our group chats, you described the MAI thinking model as a conservative recipe. And I think that's a really good description of it. Like they, you know, the team came up with this conservative recipe. And then I think that they did a really great job of actually executing. I thought that was a really good choice.
[00:32:26–00:32:42] **Finbarr Timbers:** It's not super clear to me. Maybe you've seen some papers on this that I haven't seen. But it's not super clear to me how well the trace distillation SFT does or how much better online policy distillation is versus the trace distillation SFT.
[00:32:42–00:33:06] **Nathan Lambert:** Yeah, it's like what is the relative magnitude in the final performance? So the Demetron Ultra paper has a table on how far the on-policy distillation goes relative to the teacher SFT. and they also have the starting point so I guess that's a potential way to do this here I could I could just pull this up let me switch I had this open but in a
[00:33:06–00:33:16] **Nathan Lambert:** different tab okay here's here's this paper this is page 27 is which the paragraph I just read and then it also has this kind of
[00:33:17–00:33:18] **Finbarr Timbers:** Oh, fascinating.
[00:33:18–00:33:46] **Nathan Lambert:** This is a great table. I spent a while looking at this earlier. So essentially, it's like where they get after SFT on each of the benchmarks on the general model. And then I think... Okay, so the gains over the RLVR student recovery of the specialty students. So I need to make sure... Okay, so it denotes the initial student checkpoint where RLVR denotes the initial student checkpoint and then the multi-teacher on-policy distillation.
[00:33:46–00:33:58] **Nathan Lambert:** So I'm not sure what this SFT column can figure out, but you can see the kind of like where the teacher is relative to on policy distillation. I think this is like the closest information we have on the relative performance gains.
[00:33:59–00:34:28] **Finbarr Timbers:** Yeah, that's fascinating because the DeepSeq, I forget which one, maybe it was V3.2 paper claims, or maybe it was... R1 actually claims that you can domain specific that you know doing the general stage captures the performance of it but you know that that doesn't really seem to be and yeah and then so you know doing the domain specific distilling in and then doing a general stage on top of that captures the original performance but that
[00:34:28–00:34:44] **Finbarr Timbers:** doesn't seem to be the case here like you know the gap maybe isn't huge but there is still most of the time there's a pretty big there's like you know significant gap even if it's not huge so that's really interesting yeah I wish this table and text was
[00:34:44–00:35:02] **Nathan Lambert:** clear it's like I literally can't fully parse it it's like RLVR denotes the initial student checkpoint and then OPD denotes the checkpoint after first and second iterations it's like what is the checkpoint that was used at the start of odd policy distillation I think it was the RLVR one
[00:35:03–00:35:14] **Finbarr Timbers:** So they do a general SFT stage, then they do an RLVR stage that covers the non-teacher, the areas where they don't have specialized models. Then they do MOPD.
[00:35:15–00:35:31] **Nathan Lambert:** Yeah, and then that makes sense with this recovery rate, which is like final model minus RLVR, which would be like the gains for the OPD relative to the teacher minus RLVR, which would be like what gains you needed to still cover.
[00:35:32–00:35:32] **Finbarr Timbers:** Yeah.
[00:35:33–00:35:42] **Nathan Lambert:** And like what gains the teacher could potentially give you. So more research like this. Happy to see some of it out there. I'm going to switch back.
[00:35:44–00:36:12] **Finbarr Timbers:** Yeah, something I found interesting about both the Nematron papers and then the MAI Thinking paper is that they don't talk as much about some of the more detailed post-training decisions that have shown some pretty strong gains in some of the other papers. Like, I believe it was GLM-5 where they talk about doing a difficulty curriculum in a difficulty filtering stage. Yeah.
[00:36:13–00:36:40] **Finbarr Timbers:** and that's just not something that's really talked about in these other papers. I think it was Kimi 2.5 used a temperature. It's kind of funny. So Kimi K2.5 and GLM 5 both have temperature schedules and they both claim the exact opposite thing. So one of them says you have to start with a high temperature and go low. The other one says you have to have a low temperature and go high. And You don't see that discussion, I don't think, in some of the other papers.
[00:36:41–00:37:08] **Nathan Lambert:** I still think the Chinese labs are much more willing to share really, really nitty-gritty details. The NVIDIA paper is mostly a list of methods to create a teacher, or domain-specific teachers, which... is useful but I think like I was less it's like less of a fun read they're like there's 15 pages of different domains so I'm like okay I don't like I don't need this yeah like Kimi K 2.5 and GLM 5 actually have like
[00:37:09–00:37:34] **Nathan Lambert:** more similar recipes which are also on the simpler side which is like you create this SFT stage and then you do RL the RL might be staged there's not this on policy distillation there's a bit less talk on how many experts they have and what their domains of experts are I think it's obviously you have to take All of this with a grain of salt and it's like what how they decided to present the
[00:37:35–00:37:45] **Nathan Lambert:** information is like a big factor in this and like they might actually be closer in reality and then it just wasn't described in a certain way I think another
[00:37:45–00:38:14] **Finbarr Timbers:** interesting bit is that you see the Chinese labs all seem to be converging towards stars attention whereas we don't see that you know where the American labs at least NVIDIA and you know AI2 seem to be more converging towards hybrid Attention, like the Nvidia Nemetron Ultra used the Mamba Attention, whereas, you know, we see, you know, Deepsea Sparse Attention and then the MIMO MSA, whatever that stands for,
[00:38:14–00:38:19] **Finbarr Timbers:** MIMO Sparse Attention. So I think that's an interesting divergence.
[00:38:20–00:38:49] **Nathan Lambert:** Yeah, I am not the person to ask, but I agree. I often get asked, we'll avoid the full rabbit hole, but I often get asked, are the Chinese labs more efficient? And I'm like, I don't really think so. I think the financial pressures of serving billions and billions of tokens is probably a better motivator for efficiency, given that if you make a GPT model 1% more efficient, you're making fat stacks of profit.
[00:38:49–00:38:54] **Nathan Lambert:** like I think that's like a more effective market mechanism but and then the Chinese
[00:38:54–00:39:05] **Finbarr Timbers:** lab yeah if you make you know serving chat tpt more efficient Sam Altman can say hey here's a bunch of stock like yeah but they do great like the Chinese labs do
[00:39:05–00:39:34] **Nathan Lambert:** great research which I think it's kind of a bit different okay we can move into more open-ended stuff here I think that we have like We have a bunch of things in a document here. I'm sure more will come up. How do you think about open models and post-training recipes in the current environment's craziness? I think there's two things. One is like, what the heck is going on with environments? Is this a fad?
[00:39:35–00:39:40] **Nathan Lambert:** And then two is like, what is the equivalent or most recipes do in face of that?
[00:39:41–00:40:10] **Finbarr Timbers:** Well, I think the environment bit is interesting. I made the mistake of sending a tweet out at the start of the year where I said, hey, I want to buy some environments. Please give me your cell. And my inbox has been overwhelmed ever since, including my personal phone, where people would call me at 2 a.m. and say, hey. So that wasn't super fun, but it was really interesting. I did a bunch of sales talk to people and it's just, you know, it's quite expensive to get these environments.
[00:40:10–00:40:21] **Finbarr Timbers:** Like the quote that I got from one of them was like a hundred thousand per environment. This was a more, you know, like trying to copy like a, you know, sophisticated web app, like a door.
[00:40:21–00:40:35] **Nathan Lambert:** Do you know what this means? Like is the environment in the hundred K mostly the software or is it also like the software, the props? So it's like the software is like the world. or is it also a bunch of prompts and exact training data that you could just like throw into a recipe?
[00:40:35–00:40:59] **Finbarr Timbers:** That's a good question. I think it was more the, I don't think it included the prompts. I think it was the, here's the software, like here's the, you know, DoorDash clone that you can interact with. And here's a bunch of like fake data in it. Like, you know, we'll keep this service up for, you know, like it's a hundred grand upfront cost and it was like 50 grand or something per year per environment to keep it running. I don't think it included the actual prompts, but I could be wrong about that.
[00:41:00–00:41:28] **Nathan Lambert:** I was just gonna say like a lot of it to be able to you can charge a much higher premium if you're like we have prompts and they know we know they improve X benchmark or X ability they're like here's an Amazon clone thing I think it's gonna be a lot harder to sell the labs in a sustainable fashion yeah it's like they just don't care about it as much I mean they care but I just think it's kind of a different category
[00:41:29–00:41:58] **Nathan Lambert:** I'm actually less doom and gloom on environments in the open because it is nicely unparallelizable where if you can figure out the right incentive structures and reward people for building good environments that are also good evals in specific domains, then it's like small teams could actually just do that. And then the open trading labs use said environments and figure out how to combine them and actually use them. I'm not sure that's happening in a substantive way,
[00:41:58–00:42:04] **Nathan Lambert:** but it seems like a bit more tractable than building a big post-training recipe as a small academic lab.
[00:42:05–00:42:34] **Finbarr Timbers:** Yeah, absolutely. And I think it's also, you know, once something becomes useful, right? Like once you figure out that environment is particularly valuable and, you know, the frontier labs get really good at it. I think there's also a lot of incentives for the people involved to expose it. Like if you're DoorDash, you know, sure, maybe you're not initially going to say like, okay, I'll make this environment. I'll make my own environment of DoorDash for, you know, OpenAI or whoever to get good at. But once, you know, a bunch of people have agents out there and, you know,
[00:42:34–00:43:01] **Finbarr Timbers:** people are starting to use DoorDash or use, you know, ChatGP or whatever to order from DoorDash, I think then DoorDash has an incentive to say, we'll release our own environment. I think the same is true, you know, a lot of this has marketing mechanism where you're going to say, oh, like, you know, we're, you know, Harvey or whoever, here's our difficult legal benchmark. Oh, look, we're the best at it. And then, you know, once you have the eval out there, then it becomes something we can trade against it.
[00:43:02–00:43:31] **Finbarr Timbers:** so I think that's my optimistic take but I do think it's kind of tough because the other side of things is that I think a lot of the RL companies are trying to turn around and sell the same environment to multiple people and so it's you know the cost might be 100 grand if you're selling to one player but then if you're trying to you know sell it to someone and then you know not sell it to anyone else well then the cost might be you know 500 grand or something because you know you can't turn around and sell it to you know Anthropoc, OpenAI, Gemini Chinese Labs, AI2, whoever
[00:43:32–00:43:39] **Finbarr Timbers:** So, you know, that, you know, just puts it more, um, out of reach that release everything.
[00:43:39–00:43:51] **Nathan Lambert:** There are, there are multiple price tiers. And I think that some people will like pay more for a dataset to be able to release it, which is like a really baller move, but like not many people are going to be able to do this.
[00:43:52–00:43:52] **Finbarr Timbers:** No.
[00:43:54–00:44:17] **Nathan Lambert:** another thing we have is Tinker style APIs I feel mixed about this because I think a lot of small medium startups especially are going to use Tinker to do interesting research and share it with the world and there are like fully open Tinker style APIs which is like an RL framework that you run on your own GPUs and then gives you this Tinker style API but it's like
[00:44:18–00:44:47] **Nathan Lambert:** I don't love for a like open science perspective for too much convergence on these even though I think the actual price point will be very good because places like thinking machines buy compute at a mass scale price versus academics that have very finite compute and can't do some things like very large MOE runs so I see this coming especially if we look back in like a year we'll see a lot more research done on these APIs both in like
[00:44:48–00:44:53] **Nathan Lambert:** industry and academia, but I'm kind of apathetic to it.
[00:44:54–00:45:21] **Finbarr Timbers:** Well, I like it, you know, insofar as you can expose this much simpler API for researchers to use. And especially, you know, if you can hide a lot of the complexity, I think that's ideal. Because then, you know, if you can only engage in the, you know, really interesting bits and make it so, you know, the research scientist doesn't have to think about How all of your, you know, VLM inference pipeline is set up. I think that's good.
[00:45:22–00:45:52] **Finbarr Timbers:** I think, you know, my question is just how good of a business is it actually? And it just doesn't strike me that there's this, like, you know, I think that there's a large business to providing. Well, actually, I thought he'd been super clear. You know, we've seen a number of companies providing, you know, RL fine-tuning services, you know, RL as a service. We've seen a lot of companies try to provide fine-tuning as a service. And, you know, none of them have really taken off. Like, I think OpenAI has started to shut down. I think they shut down their RL fine-tuning. I think they might be shutting down their fine-tuning. Maybe I'm wrong about that.
[00:45:52–00:46:18] **Nathan Lambert:** Well, it's like, Cursor used Fireworks for their actual training run, and I'm like, I don't really know all the details of this, but Cursor does something for fast, I think like fast weight transfer, or Fireworks does a fast weight transfer and other things to make it so they can scale their RL Inference compute very nicely. So that's one type of it. I don't know how big of a long tail that business is. But also, I think Tinker is a better business than most people expected. It makes some real amount of money.
[00:46:18–00:46:47] **Nathan Lambert:** It's like in the hierarchy, I think selling compute, not the best business. Selling inference, great business. And Tinker-like APIs, if you can't transition it into selling tokens, is somewhere in between the two, where they could take some amount of margin that'll be slightly higher than just selling the compute. And they obviously get a margin by having, like they get compute at a cheaper rate than their customers, and that's like part of the margin they're taking. But I don't see it being as nice as inference,
[00:46:47–00:47:10] **Nathan Lambert:** so it's kind of existential for them to make it so that these fine-tuning APIs... feed into a inference business pretty nicely because then you can be somewhat locked in on you train the model on our infrastructure. You actually can own the model weights, but the training dynamics to inference mismatch is perfect because you trained exactly on our inference engine and are going to get what you want out of it.
[00:47:12–00:47:38] **Finbarr Timbers:** Yeah. And it also helps a lot with utilization because you can then, you know, utilize it. You can ensure that utilization across a lot of Clients. I think it makes a lot of sense. I think it's probably a better model for a lot of users. I think of academic users. It probably makes way more sense to do this. Or for that matter, if you're starting a new post-training lab now, as I know a few people who are,
[00:47:38–00:48:07] **Finbarr Timbers:** I think that's where it probably makes a lot of sense to start with something like the Tinkerer API. And then at some point, if you want to try and capture that margin, maybe then you try to do something more Custom, but if you can use something like that, like, that's great. And the economics are, you know, fundamentally more sustainable, or, you know, better for you, rather than trying to, you know, go to core weaver, whoever, and say, or server scale and say, hey, I need, you know, 10,000 networked GB200s, you know,
[00:48:07–00:48:12] **Finbarr Timbers:** that's just a very expensive thing to do, especially if you can't keep it running all the time.
[00:48:15–00:48:20] **Nathan Lambert:** Do you have any more hot takes on post training before I ask you some more general things?
[00:48:26–00:48:48] **Finbarr Timbers:** Well, something I'm generally interested in, and I'm the wrong person to speak to about it, I'd love to talk to someone who's maybe a capital allocator, or a compute allocator, who's deciding where to put compute or where to hire team members, because I'm kind of curious how the high-level decisions are made,
[00:48:48–00:48:50] **Nathan Lambert:** allocating resources between pre-training and
[00:48:51–00:49:09] **Finbarr Timbers:** and post trading. Because, you know, what I kind of have seen as a general trend is that you see a lot of papers where there's, you know, more focus put on one or the other. Like I think, so yeah, so that's something kind of interesting to me is how people who are, you know, making this decision, how they're making that decision and how they're thinking about it.
[00:49:10–00:49:38] **Nathan Lambert:** yeah it's like the hardest decision to get out of labs I've like I used to spend time trying to get them to share more but I think it's like such a sensitive decision to where they see progress coming like they're making that decision allocating compute based on where they think the most progress is and what they'll like return on investment is so if you go to Anthropic and they're like here's our distributions it's like okay that's where labs see their bets and or where they see
[00:49:38–00:50:07] **Nathan Lambert:** they are weak and it's like you invest more compute to make progress in the area that you are interested in which I always think makes a lot of open research kind of Boring right now is like the people that get compute are just way more likely to succeed as academics and researchers, which is a horrible equilibrium for the world, but kind of realistically true. I don't know how to make a lot of that. I wanted to ask you how you feel about the craze that people have to cash in on
[00:50:08–00:50:17] **Nathan Lambert:** making money and join a lab before the ladder gets pulled up and what people should be optimizing for in their careers in face of meaningful opportunity costs.
[00:50:18–00:50:21] **Finbarr Timbers:** Yeah, I think that's actually very timely.
[00:50:22–00:50:27] **Nathan Lambert:** But yeah, I think that that's really important to talk about.
[00:50:27–00:50:56] **Finbarr Timbers:** I mean, I think it's always worth focusing on whether what you're doing and spending time on is going to be generally valuable or if it's like a really short-term exploitation type thing in the RL, like explore versus exploit setup. I mean, something that I've seen throughout my career has been often the places that pay the most are all the places where you're doing the most interesting work, right? Like, you know, if you're going to go work at OpenAI, or you're on Anthrop or you're at the front of your lab,
[00:50:56–00:51:14] **Finbarr Timbers:** but they pay a lot of money. They also have a lot of resources. So you're going to make a lot of money and learn a lot. So I think it's worth trying to decide, is that, is the I think that would have been a mistake but trying to figure out if you're going to be able to do interesting work is really important and also try to figure out if you're going to
[00:51:32–00:52:00] **Finbarr Timbers:** be able to, you know, push forward science, you know, if what you're doing is more just saying going to, you know, data vendors and saying, you know, okay, you know, we I need a bunch of data to do. whatever and then you know they give you a bunch of data you train a model you say it's good you're bad or whatever you know I don't think that's as interesting and I don't think you're going to learn a lot even though that's you know work that would probably drive model progress for it I think if you're able to you know make focus more on the science and make more scientific conclusions I think that can be you
[00:52:00–00:52:10] **Finbarr Timbers:** know a lot better for your long-term career and I think that's where places like AI2 and the other academic research labs you know Moran is doing a really great David Pérez Pérez Pérez Pérez Pérez Pérez
[00:52:32–00:53:00] **Nathan Lambert:** Yeah, mostly this is grounded in visiting the Bay Area and every time I go I'm like, holy shit, what is going on here? All these very junior people have way too much dread about their opportunity costs and both of us aren't based in the Bay Area so I feel somewhat removed from it, which gives me a little bit more Time to pause and be like, what exactly is the right thing to optimize for? It's easy for me to say as somebody who's established, but I think there's opportunity for a lot of people to just,
[00:53:01–00:53:30] **Nathan Lambert:** if they have conviction on something, to try to go and do it and not just follow everybody that goes down the funnel of joining one of the established labs or the neolabs. where I don't hear from many people that join as a junior person at these places and end up with very high responsibility like they're contributing to something that matters or they're around a cool group of people but I don't hear from that many people are like wow I'm doing the highest leverage stuff and the most interesting things
[00:53:31–00:53:57] **Finbarr Timbers:** Well, I think that, you know, it's kind of funny for me to say this. My career has been more on the opportunistic side of things. But, you know, twice now, I've been at organizations where I've been working. So, you know, at DeepMind, I was part of the Alberta office where DeepMind had, you know, aqua hired. the computer poker research group from the University of Alberta and so you know
[00:53:57–00:54:26] **Finbarr Timbers:** this was a group of people who were really invested in computational game theory and you know poker playing algorithms and they were all in on that and you know they were all in on that to the point that you know they were one of the two leading labs in the field and where, you know, because they were so strong at this, DeepMind came and, you know, Aqua hired them and they all joined and they, you know, did quite well from that acquisition there. And, you know, I joined later because I was,
[00:54:26–00:54:55] **Finbarr Timbers:** you know, interested in working with them and doing game theory and stuff. But, you know, it was this group of people who had this conviction that what they were doing was really important and, you know, it worked out quite well for them. And then, you know, the same thing at AI2 where at AI2, you know, there's all of these people who were really interested in NLP Research or even before language models like we see people like you know like Kyle and Dirk I think we're both at AI2 for like almost a decade like they had these really long Tenters and then they did really well and then you know they've since had some you
[00:54:55–00:55:14] **Finbarr Timbers:** know strong opportunities coming out of that with with yeah some of the opportunities that have been available to them and I think that the consistent theme there has been that you know if you have high conviction that what you're doing is important and interesting then like it's not a mistake to follow that and to you know try to become really strong in that area
[00:55:15–00:55:44] **Nathan Lambert:** Yeah, I mostly think it's good for the world to have a more diverse set of approaches. It'll be interesting to see what the Neolabs actually produce, if they can manage to do things that are diverse. My personal idea is that they're so big now that most of them need to end up doing something that is somewhat similar, which is hard, but like... They need to keep risking their $20 billion valuations to do something interesting
[00:55:44–00:55:48] **Nathan Lambert:** that's not just going to be squashed by an opening AI or an anthropic side project.
[00:55:49–00:56:02] **Finbarr Timbers:** Yeah, absolutely. And I think it's tough because when you're raising, when you have these huge seed rounds, you're raising $200 million or $1 billion or whatever, then it's like you have to pretty quickly show results to be able to grow off of that.
[00:56:04–00:56:18] **Nathan Lambert:** yeah so a to be continued conversation any last words I don't need to stretch it on if we don't have anything to add to our conversation no I think this is pretty good
[00:56:18–00:56:29] **Finbarr Timbers:** I think it was really great getting a chance to catch up and talk about some of this stuff you know I've been reading all these papers and thinking about all the different recipes so it's great to get to chat about it and put it out into the
[00:56:29–00:56:33] **Nathan Lambert:** ether so yeah thanks for having me on yeah thanks for coming back we'll talk soon
[00:56:34–00:56:34] **Finbarr Timbers:** Sounds good.
