返回目录

前沿 Post-Training 配方:从单线流水线走向多教师蒸馏

Nathan Lambert 与 Finbarr Timbers 复盘 InstructGPT、Tülu 3、OLMo 3、DeepSeek R1 与 2026 年 MOPD,并讨论组织能力、RL 环境和训练 API 的经济学

这一期在说什么

前沿 Post-Training 的真正瓶颈,正在从“有没有算法”转向“能否让组织、教师模型与基础设施保持同一分布”

整期最值得带走的并不是某个固定 recipe,而是 recipe 与组织结构之间的耦合:小团队可以稳定执行 SFT→偏好调优→RL 的清晰流水线,但模型一旦扩展到数学、代码、Agent、安全等多个专长域,就需要并行训练教师、协调数据和算力,再把专长可靠地收敛到一个学生模型。MOPD 看似只是对现有 RL learner 的一组改动,真正困难的却是让教师与学生共享足够接近的训练轨迹、SFT 数据和推理行为。换言之,2026 年的竞争优势不只在 loss function,而在能否把 org chart 变成可复现、可对齐、可持续迭代的训练系统。这是本文基于整期对话的编辑性综合,并非嘉宾原话。

一句话

Post-Training 已从一条可由单团队完成的流水线,演化为多专长教师、在线蒸馏、复杂数据与算力编排共同构成的组织级系统工程。

Post-Training · RLVR 与 DPO · Multi-Teacher On-Policy Distillation · OLMo 与 DeepSeek · RL Environments · Training API Economics

Insight

评审训练计划时,应把跨团队依赖、数据交付和算力排程当作 recipe 的一部分,而不是项目管理附录。

  1. Nathan 认为现代 Post-Training 的关键能力之一,是把 compute 与 data 编排成工作流;这直接映射到组织协调能力(约 00:02:06–00:03:56)。
  2. Finbarr 指出九个月的模型周转并非异常缓慢;Nathan 的反驳是,如果只是移植旧 recipe,这个速度对应的能力上限仍偏低(约 00:03:56–00:05:52)。
  3. 对于 7B–30B 级模型,是否值得完整复刻 DeepSeek 式 RL-first 路线仍未得到证据充分的答案(约 00:05:22–00:06:11)。

Nathan Lambert Hello, we are back on a Interconnects conversation. I don't really say I do interviews. People criticize me because I interrupt the guests too much. I'm not a good interviewer, but I'm here to entertain people. This is also fun for me because I'm trying to make like a post-training course and it kind of fits as the advanced end of this. So it's kind of a crossover between Interconnects content and other stuff that I've been spending my time on this summer. So I'm happy to wel

come Finbar back, I think.

Are you the first return guest?

I haven't checked. Oh, wow. Finbar and I worked on this sort of post-training recipe stuff for a while at AI2. I left recently. This is one of Finbar's last days at AI2. It's already been announced. It's not a spoiler here. So we're going to kind of reflect on some things on building post-training recipes for Ulmo. Then we have a little review slide deck and notes on The kind of state and evolution of frontier post trading recipes over time, which is pretty interesting because there's, what is it like two to four kind of canonical recipes that there has been. So it's kind of interesting when you see the field converge on something new, which it's doing right now with multi teacher on policy. distillation for some reason that's a bit of a mouthful it is a long acronym and then we'll just kind of end with various discussion points on post-t raining and what we're up to so happy to give you a floor if you have any hot takes you want to start with to get people to draw people in otherwise I think I'm excited to kind of reflect on this because I know you've been reading a ton of papers recently and kind of prep laying some of this groundwork

Finbarr Timbers Well, yeah, I mean, today is my last day at AI2. So it feels very appropriate to be talking to you as you're the one who recruited me to AI2. So, yeah, that's pretty special. And it's great to be, yeah, the first repeat guest. I feel honored to be back on. So, yeah, thanks for having me.

Nathan Lambert Yeah. Do we want to start with Ulmo? I think that people... I need to do this carefully, but I've talked about Olmo 3's post trading many times to people. I haven't done this in a very direct way on the podcast, but I would say that post trading Olmo 3 to make this reasoning model was a... Major accomplishment for many individuals to do this, but also the complexity of what we were doing was pushing against the limits of AI2's organizational capacity.

And a lot of modern post-training is like your ability to wrangle compute data into a work stream. And in order to do that in a complicated way, you really are wrangling an org chart. and that's like part of why it's like old boat three was by its nature pretty late as a reasoning model it was like pretty rigid reasoning model and that's like partially reflected in the recipe being pretty simple but then when you like compare it to all these new recipes with tool use and mult

i-teacher distillation and all of this it's just like A fork in the road where it's like you could do this very simple thing and make a strong recipe, but it is not representative of what all the Frontier Labs are doing. And I think that that kind of fork in being able to say that things are similar happened kind of after Tulu 3, where Tulu 3, I think... was also much simpler with this three-stage SFT DPO RL recipe but that simpler recipe was probably closer an outcome to wha

t the labs are doing but now doing that sort of three-stage recipe for a reasoning model and especially a tool use like agent model just doesn't really apply and that's the point that's why I think the point of this podcast to be like what are the what are the way what are they doing to make these like true frontier models and then shed some light on how it contrasts to the more like open academic ones

Finbarr Timbers Well, actually, I think that's interesting. What was the process? So, you know, I only came around for Almo 3. I wasn't around for the earlier versions. What was the process like to go from Tulu 3 to Almo 2? Because, like, just looking at an archive, I think Tulu 3 came out in November of 24. And then Almo 2 came out in December of 24.

Nathan Lambert We just applied the recipe. Yeah.

Finbarr Timbers Yeah. I mean, so I think that actually, like, and then, you know, Deep Seagull 1 came out in January, end of January 25. And, you know, Almo 3 was then released in October. It was October, November of 25. I think November. Yeah, November.

Nathan Lambert Yeah, right, it was November. It was like do or die with Thanksgiving. I remember that.

Finbarr Timbers yeah because Canadian Thanksgiving had already happened which yeah I was happy but like I think it was sure maybe it was late but I think it was only late by a few months like it's actually like you know if I think of my past experience with model turnaround times like a nine month model turnaround you know from R1 coming out like that's actually that's not bad I think you know something like six months would have been

Nathan Lambert Nicer I think it's slow because we didn't it would be fast if we had rebuilt the R1 recipe but what we did was we like ported reasoning into our existing recipe okay which is a simpler task but has like a lower ceiling in my opinion where it's like the deep seek in the newer style recipes which we'll talk about I think they just have a much higher ceiling and how much you keep hill climbing them or they're just like more prescribed more pedagogical of what the frontier is doi

ng,like for the size models that almost was, which was like seven to 30 B. I'm not sure that doing this deep seek style RL first recipe is actually useful.

Finbarr Timbers Yeah, I think that's a good point. And I mean, I think that's really reflected in what we see the research where you see, you know, you obviously you see the big, the step change, and you know, how quickly things are improving. when R1 comes out. So I think that's a great point. And it really does seem to saturate or to not saturate, sorry, with compute.

Nathan Lambert Yeah. Should we just do the slide deck?

Insight

“开源了 recipe”不等于他人可复现;还需要公开阶段接口、数据来源、失败分支和组织依赖。

  1. 随着训练阶段和专门数据增加,即使论文公开,外部顾问也很难在不了解完整分布与流程时给出一句话建议(约 00:09:01–00:12:00)。
  2. AI2 的简化流程既是研究设计,也是组织无法像工业实验室一样横向扩张的结果(约 00:12:02–00:12:56)。
  3. OLMo 3 延续与 Tülu 3 相近的组织形态,限制了 recipe 出现结构性重建(约 00:12:56–00:13:23)。

Nathan Lambert We're throwing around like recipe names. I feel like it might be useful to just do it because a lot of people probably want to follow but don't exactly know. I'm going to share it. I'm going to share a screen so people listening it might be useful to either you can pull the slide deck up on your phone and click through it it's not super information dense but you can also just watch it on YouTube all this will be linked generally this is just like a quick survey on how frontier recipes have evolved we'll go through the history quickly and then talk about what is currently happening and kind of probably interleave the Olmo discussion we were having okay There's a bunch of canonical recipes we'll talk about. This is where I got the two to four number. I think the recipes are like InstructGPT, which is what coined the initial RLHF with this three-stage idea, which took a while to get people to move on from, which is like SFT,reward model, and RL. And I see as like Llama 3 and Tulu 3 as kind of practical implementations of that with other tricks of the trade. So those two could potentially be merged together. It's like just like kind of pre and post ChatGPT moment. And then the two most recent canonical recipes that we'll cover in this, I would say, are like DeepSeek R1, which is the shift to doing like reasoning focused and bigger RL stages than this kind of SFT focus from before. And then Mimo Flash and some of the new models from 2026, which add this distillation element.

Finbarr Timbers Well, and I think it's worth pointing out, too, that it's not just Mimo Flash. Like, it was kind of a consistent theme. Like, you saw this with DeepSeq. They referenced it in the V3 paper, and then it's, you know, it's Kimi K2.5, it's GLM-5. Like, it's all of these papers, you know, start talking about this specialist RL stage.

Nathan Lambert Yeah, I think there's a debate on how we draw it and whether or not distillation is... If you have distillation as a technique, as a key milestone, then Xiaomi was the first. But it's kind of a march over time where you kind of see them change. And we'll go through this. I don't need to...

Finbarr Timbers When you say distillation, I do think it's important to distinguish between the straight up distillation of the leading closed models and distillation of these domain specific models, where I suspect that the Chinese labs are doing both. But a lot of what they're doing is this training these domain-specific models, like a math model, a coding model, logic model, whatever,and then distilling those models back in and not just distilling from – so when we're talking about distil

lation, it's not just distilling from the leading closed models.

Nathan Lambert Yeah, it's a pain. I agree. The distillation term is horribly overloaded. um there's a review slide do we need to review multi-teacher on policy distribution it might be too complicated to need to do it we could come back to it i think i kind of want to just go through the actual models and then we could use the supporting slides as needed um This famous InstructGPT three-step thing, I think many people have heard of it,but this is what constituted post-training at the time o

f ChatGPT coming out, so it's kind of important grounding of this human-supervised SFT data. mostly human supervised preference rankings to make a reward model and then do RL on that and the model gets better. And it's pretty interesting how all of these have been kind of phased out, at least in terms of what we know openly, where they're we don't use that much human demonstration data for SFT there's likely some human preference data still in the loop but I would guess that

synthetic has a much bigger role and there are reward models but they're like not the key RL target anymore so in four years most almost all the canonical pieces have been moved on and like this evolution is kind of within there I think the early models after InstructGPT like Llama 2 even Lava 3 these are pretty similar which is like you're starting to break down this recipe with different tools like rejection sampling DPO some increased iterations I think increased iteration

s is just that there was more incentive to squeeze more out of the models and they just like broke things down more where InstructGPT seemed like a bit more open-ended research where this kind of cleanness

Finbarr Timbers was fine well I think that's interesting with respect to how much everything is scaled Right, because InstructGPT was before ChatGPT was released. And so, you know, it's something like just the complexity of what was done is that which a small team or a single team could do. But then when you start looking at, you know, Llama 3, like it just starts to be a more complicated process.

process and where you start to have a lot more specialized data and there's a lot more room for scale and for money and complexity being poured in.

Nathan Lambert Yeah, it's like both for-profit and non-profit efforts to do post-training want me to advise them. And I'm like, I don't really know how I'm going to give you advice unless I'm spending 20 hours a week understanding the details of your recipe. Because it's like, I can't really give you a one-sentence thing of do X without understanding all the complexities of the model and the post-training process that go into it. Which makes it hard for me.

Kind of like a transparency point of view, even if it's fully detailed, it's definitely still hard to modify and study.

Finbarr Timbers Absolutely.

Nathan Lambert So then like two to three at AI2, a lot of this was we're trying to beat the results of this Lama 3 post training, which is pretty complicated. But we don't have the ability to scale the organization as far. So I think that's a big reason why the actual workflow is a lot simpler where we have three clear stages that are doing slightly different things and they build on each other. And that's like It's never stated very explicitly in these papers on how the org chart impacts t

he recipe, but I think it's a very strong signal within at least the delta between the fully open work and the kind of partially open work that you get from industry.

Finbarr Timbers Yeah, absolutely. And I think especially as we'll see with the domain-specific models, that's something where you could really easily scale up your org chart to scale that up.

Nathan Lambert Yeah, and I threw Ulmo 3 in after this, after the 2 to 3 slide, mostly just to show that the recipe was so similar to 2 to 3, and the org chart had really changed. Like, we didn't have more ability to scale. I think there was a little bit of separation between the model types, between, like, the think and the instruct models, but, like, without a major org change, it was just kind of stuck in this and do the best you can with it.

Finbarr Timbers Yeah, true.

Insight

设计 SFT 数据时,先明确它要冷启动哪种 RL 行为,而不是笼统追求覆盖更多指令。

  1. 这套流程不追求一条优雅的研究公式,而是反复训练、采样、筛选和再投入的工业过程(约 00:13:23–00:14:50)。
  2. SFT 的功能从一般能力灌输转向为 RL 提供稳定起点,并吸收较强模型已经形成的行为(约 00:14:11–00:15:27)。

Nathan Lambert because like the real big change was this with Deep Seek R1 I had never seen this plot before but they had this plot maybe they added it for the nature version of the paper where they kind of show their recipe where they like take the base model they do RL0 and then they sample from the RL0 to like filter prompts and then they use that as SFT this is like going through this they use that as SFT for the next version of the model to create like a development internal RL Deep Se

ek R1 and then they do this like repeated sampling to trade multiple RL versions and kind of distill in the sense of clarify and refine the reasoning behavior of the model before going through the final pipeline, which again is a mix of reasoning and non-reasoning SFT into a bigger RL realm.

Finbarr Timbers Well, I think this is really interesting because it starts to show, I mean, first of all, the complexity here. We're starting to use synthetic data as this primary input here, but it's not just... Like, you know, it's trying to elicit, you know, specific behaviors and it's this kind of like industrial process instead of like this, you know, it's not as much of an elegant research recipe. It's more like, you know, we train a model and then we use it as best we can and we keep

iterating.

But I think the other thing that's interesting is we're starting to see here the SFT serving as the cold start. First of all, I think before SFT was more of a generally useful stage, whereas here its primary purpose is this cold start for RL. And then the other interesting bit is DPO starts to disappear at this point from the Yeah, so my hypothesis for the dropping of DPO on these models is that...

Insight

是否移除 DPO 应以同数据、同算力的消融决定,而不是用“前沿实验室不用了”代替证据。

  1. 偏好损失仅依赖对比信号,却能改变数学、代码和 reasoning strategy,说明模型内部仍有大量可被重新排序的潜力(约 00:17:41–00:18:19)。
  2. DPO on verifiable rewards 或同家族强弱模型构造偏好的做法仍值得研究,只是学术审美上不如在线 RL 热门(约 00:17:41–00:19:18)。

Nathan Lambert As you're doing like a cleaner recipe, essentially the need falls away versus if you look at Ulmo, which is taking tons of potential gains by refining your model on outputs of strong open weight models, like largely Quen and DeepSeq is the training data for the SFT of Ulmo 3.0. The delta between that SFT data and the base model is still pretty big in the probability distribution. So DPO kind of helps further refine and clean up that distribution in a way that has very rough e

dges.

But when you have a more refined industrial process on post-training, that potential benefit will be harder to gain. Something interesting that I didn't fully confirm before this is, for example, NVIDIA used to also be on this DPO train with their smaller Nematron models. And I would guess that potentially Nematron Ultra would not.

And that's because they're much further down this development tree and using these more on-policy methods for creating the SFT data. And their model, I would guess, will become kind of more robust out of distribution and have less weird rough edges because of it. So that's kind of my hypothesis on DPO and People that use DPO will be looked down upon, but if you're trying to bootstrap a recipe off the ground and just take gains where you can,I still think it'll work for a lot

of people in a kind of compute efficiency standpoint.

Finbarr Timbers I think generally there's something interesting with the preference tuning that maybe it isn't being given the proper respect that it deserves because one of the interesting bits about the Nemetron 3 super paper was that they saw they do a traditional RLHF stage in their RL which has also you know fallen with fashion and they see pretty massive gains with it so I think some of these changes are more you know driven by what's in fashion rather than perhaps like a fully rigorou

s you know set of ablations

Nathan Lambert It's pretty remarkable to me that the preferences loss function can do so much for these models. The models have so much potential there, and it's really a contrastive loss on pretty granular feedback. They learn all sorts of things. They'll get better at math and code, or their reasoning strategies will be refined.

Finbarr Timbers That's remarkable to me.

Nathan Lambert I think there will still be funny research on using preference-based losses with verifiable outputs. I think all of this would work, like DPO on verifiable rewards and stuff like this. It's just kind of intellectually less appealing.

Finbarr Timbers Yeah, well, I think that's where I thought that the Delta Learning hypothesis-style DPO, like what ALMO 3 Did where you were the the preference you create these synthetic preferences by having like strong by like bigger and smaller models of the same family like is where you get your preferences from. I thought that was a really interesting signal because it seems really analogous to some of the work some of the guidance stuff that we see in diffusion models like how you have

the classifier free guidance which has something similar and there were very similar results there which showed that you could have the like one signal they used was further along in training versus earlier in training models as like a source of signal that you could guide along and that worked quite well and so I suspect that these signals um for preferences in that way like that they could actually be more robust but because you know some of the largest labs do n't have to

do that perhaps we're not

Nathan Lambert setting them as much yeah or they don't tell us like to continue this it's kind of cool to look at so the DeepSeq models have kind of gone through this what I would call it closer to llama recipes to DeepSeq R1 which is like most definitively the canonical recipe for reasoning models and then continue to change closer to this multi-teacher format. So if you look at the V3-3 paper before R1,they do something remarkably similar to Tulu-3 type thing where they have a mix of SFT

Insight

实施 MOPD 前先定义教师域、路由规则、学生采样策略和每个域的独立恢复率。

  1. MOPD 并不是从闭源旗舰模型抄答案的泛化“蒸馏”,而是先训练领域教师,再合并内部专长(约 00:07:42–00:09:01)。
  2. Finbarr 认为实现层相对直接:保留现有 RL setup,对 learner 做有限改动;难处更多在教师训练和整条链路(约 00:23:00–00:25:57)。
  3. 专长团队可并行推进,使复杂 recipe 更容易按组织边界扩展,但重建完整历史依然困难(约 00:25:02–00:25:57)。

Nathan Lambert and then they use like this RL on verifiable rewards. They didn't call it that or their paper wasn't out at the time. And So they did this before R1 came out, which was just kind of a less reasoning focused models and use the same tools, but with a different ratio of implementation weight.

Finbarr Timbers And what's interesting is that this comes out basically at the same time as 2 to 3. And it's a very similar 2 to 3 and Almo 2. It's a very similar recipe, just done with multiple people.

Nathan Lambert Yeah, yeah. And then we have this R1, which we've just talked about at length in January, which is a month later. They have a few more releases through this. They have some updates to their V3 and R1 models, which have dates, which are largely the same recipe. And then the next documented change in their recipe was V3.1, which is when they merged this thinking and non-thinking into one model, which everybody that does this has said that it has been hell to train in. But you k

ind of need it from a serving perspective. And it's obvious that long-term...

It's obvious to me that long term all the models will be reasoning models and you'll just have reasoning models that are very efficient based on the gains that are there. So this is kind of a needed change that they made and then in December of 2025 they released v3.2 which is when there's kind of meaningful changes to their recipe and they're talking about this expert creation with separate mini recipes and then using that within their kind of R1 data process to do SFT data

and then like a big RL run at the end with GRPO. So it took about a year for this Like kind of evolution of the R1 style recipe to land in their models. And I think this is like a very big complexity step that isn't represented in something like Olmo 3. And it's kind of where you can see a fork in the recipes over time as they become way more industrial and scaled at these frontier labs.

Finbarr Timbers Yeah, and I think another one here just from historical note is that I think it was with the O324 release where they updated the original V3 paper. So, you know, V3 comes up before R1, then R1 comes out. And then after R1 comes out, they actually go back and update the V3 paper, maybe getting ready for the nature submission. or something. They make a reference there to say like, oh, you know, something you could do is you could train these domain specialist models and then co

mbine them.

And then, you know, that later becomes kind of more of a priority as they talk about in V3.2.

Nathan Lambert That's a fun note. Yeah, and then More recently in April 26 is this V4 model, which has even more experts. They add this new loss function for multi-teacher on policy distillation, which I said follow Xiaomi. And this is kind of a microcosm of the arc that the whole industry went through, at least the people who share what their post-training details are of realizing how core RL is changing the recipe around scaled RL and then figuring out how to kind of scale to more domains

in the scaled RL format without just like grinding to a halt in operational complexity yeah so that kind of the next stage of this is these what I call 2026 style recipes which are all these models that are doing this multi-teacher infusion of knowledge. Some of them are using on-policy distillation and some are not. One of the key things to see is how crucial is this on-policy distillation to really keeping up with the frontier. The paper that named this term was the MIMO F

lash V2 paper. I think the model was released in December and the paper in January, which a lot of things will look similar to this kind of large RL style recipe but with this large RL run is where the on policy distillation comes in so for this is probably a better time to explain I have this great little feature so this is like the summary of what multi-teacher on policy distillation is. Generally it fits within an RL framework where you have the model you are training, lik

e the general model, sample its own trajectories, and then you route the trajectories to various expert models you have trained.

And each kind of sample is trained with this distillation KL loss to match the tokens of that expert. And People have multiple models have shown that this type of supervision is really useful for the models. You could combine it with other RL losses, such as verifiable rewards, which, for example, Sasha Rush gave a good mini spiel on that and how they use that with Composer, which is a video that I really recommend people watching as well.

But the key of it is that it is a different loss function, but it plays very nicely in the RL frameworks that people are already using.

Finbarr Timbers So they use these teachers. It's an URL. I'm talking with some of the people at AI2, but implementing it now. And it's like you take your URL setup and then you just use a very set of tweaks on the learner to actually implement this. So it's quite straightforward.

Nathan Lambert Yeah. So this is a fancy diagram that makes it more complicated than it needs to be, but it's also a very nice diagram. which shows the various domain teachers that they have, cert agent, code agent, math, reasoning safety, and how they put these together. And the experts are used both for SFT data and then this final supervision. And the recipe for the experts would look something like this deep seek recipe, which is complicated on its own, which is like make a very good rea

soning model that is good at one thing.

Finbarr Timbers Well, I think it is complicated, but it's also like if you think about being the actual researcher, like working on it, it's like... you know you have a base model and then you have an RL setup and you know you're just constantly updating both and then rerunning RL so you know the most complicated part of it is just you know writing down the history and tracing everything that's kind of like a very natural organic way for the RL to evolve through you know iterative experiment

ation yeah so like once you have a recipe

Nathan Lambert you're progressively tinkering with each part and it's fairly stable but it's hard to rebuild from scratch so like See how long the recipe shape lasts, but it'll probably be order of years. Another big one in this that also shared a lot of details on this on-policy distillation approach was Nebatron 3 Ultra, which is obviously exciting to me to have a US-made model that is very strong performance and NVIDIA released a lot of data sets with it. But they also talked about a lot

of their very... implementation details of what was hard with on-policy distillation. I have notes somewhere on this. They do this thing where they have two rounds of on-policy distillation. They found it to be better to integrate some teachers one after another. The paper has a lot more details. I don't want to go scroll through the paper, but we could also do this.

Did you have any other impressions? We have this other doc we can pull up that also might have had other details on it.

Insight

不要只比教师最终 benchmark;还要测学生轨迹在教师下的似然、推理格式和 SFT 数据重叠。

  1. 该问题并非抽象风险:NVIDIA 因教师与学生并行开发而在实践中遇到,需要额外 cross-teacher alignment(约 00:28:10–00:30:29)。
  2. 两位嘉宾希望发布教师和中间 checkpoint,以便研究者从同一起点研究训练动态;目前公开证据仍有限(约 00:30:29–00:31:04)。

Finbarr Timbers Yeah, well, I think something else that is worth, you know, contrasting the paper to is the Nemotron 3 super paper. Because in the Nemotron 3 super paper, they had a similar complicated recipe, but they did multiple... rounds of RL like there they had three rounds of RLBR followed by a round of software engineering RL and then followed by an RLHF stage so it was it was really interesting to see them go from doing that like you know one of the most complicated RL setups or in

terms of you know successive stages that I've seen to then you know you know this setup where it's still complicated but it's a lot um you know it's a lot conceptually a lot simpler yeah I pulled the paper up it's

Nathan Lambert going to be hard for me to like I had highlighted a few details that the interesting parts are kind of around the various NVIDIA details on all the teachers there's just so many details in their paper on training all the teachers I think okay so I have some of it I have some of this up it's like I have an interesting quote that's like one key finding from our So I think they'd have to do some cross-teacher alignment to make sure that they're actually similar, which I feel lik

e could become a whole organizational nightmare. It's like they say, we hypothesize that when the teacher and student are trained on different SFT data, they acquire different reasoning behaviors and induce different output distributions. This distribution matrix can cause student generated trajectories to be out of distribution for the teacher, reducing the quality and reliability of the supervision signals provided by the teacher.

Finbarr Timbers Yeah, that's interesting, actually, because there was a paper, I can't remember the name of it, but there was a paper that I read recently, which claimed that what you need to do is constantly, so one thing you could do, which was kind of the obvious thing, Nathan Lambert natolambert.com and then you take these final experts and then you do some sort of, you know, on policy distillation to combine them into your final model. But with the paper, and I'll try to find it and the

n give it to you and see if we can share it. What they claimed was that you need to... Instead of using the converged model, you need to do it in successive stages with the in-progress model. So if you change your RL for 1,000 steps,you can't use the 1,000-step checkpoint for the on-policy distillation. You have to do it in stages. At first, use the 250-step checkpoint and the 500-step checkpoint, and gradually bring that base model up to speed, or else there's going to be to

o much divergence, and the KL divergence will just be too...

Nathan Lambert um too distinct to learn from yeah so essentially the last sentence in this paragraph I had read most of is literally like we encountered this issue in practice because the teacher and student models were developed in parallel it's like they're like this is a problem because of it's like hard to do everything at once which is this is the type of thing where having research in it would be so great and I think a video could release some of the teachers and then people could jus

t like if you have the teachers and you have the intermediate model stage you could do the problem of like just studying multi-teacher on policy distillation from the starting point and understanding the training dynamics which is the type of thing we would want to do at Olbo we just haven't scaled our recipe to this point yet yeah absolutely so I will keep encouraging NVIDIA to do this that'd be great NVIDIA listen They listen. The other side of things is a bunch of models r

eleased in 2026 that do not do this multi-teacher on policy distillation and they also don't do nearly as many teachers. So I would say that this Microsoft model, which... I don't say this as a diss it's like hard to get a new team off the ground is they went for a simpler approach to try to get a solid model and it has three more general experts combined via SFT and then like a longer RL run so it looks a lot more like DeepSeek R1 but I suspect that what they will do next is

Insight

从最少可行 recipe 开始,一次只引入一类结构变化,并为每次复杂化保留消融和回退点。

  1. 同时改变太多环节会让 recipe 因自身复杂度崩溃,因此保守方案可能是新团队更合理的第一步(约 00:31:10–00:32:41)。
  2. 不同论文甚至对温度调度方向给出相反结论,提醒这些规则依赖模型、数据和阶段,不能脱离上下文照抄(约 00:35:43–00:36:40)。
  3. 中国实验室常披露更细的训练技巧,但仅凭论文仍无法可靠判断中美实验室谁更高效(约 00:36:40–00:39:05)。

Nathan Lambert make finer grade teachers and see if they need to switch to on policy distillation

Finbarr Timbers Yeah, and I think, you know, in one of our group chats, you described the MAI thinking model as a conservative recipe. And I think that's a really good description of it. Like they, you know, the team came up with this conservative recipe. And then I think that they did a really great job of actually executing. I thought that was a really good choice.

It's not super clear to me. Maybe you've seen some papers on this that I haven't seen. But it's not super clear to me how well the trace distillation SFT does or how much better online policy distillation is versus the trace distillation SFT.

Nathan Lambert Yeah, it's like what is the relative magnitude in the final performance? So the Demetron Ultra paper has a table on how far the on-policy distillation goes relative to the teacher SFT. and they also have the starting point so I guess that's a potential way to do this here I could I could just pull this up let me switch I had this open but in a different tab okay here's here's this paper this is page 27 is which the paragraph I just read and then it also has this kind of

Finbarr Timbers Oh, fascinating.

Nathan Lambert This is a great table. I spent a while looking at this earlier. So essentially, it's like where they get after SFT on each of the benchmarks on the general model. And then I think... Okay, so the gains over the RLVR student recovery of the specialty students. So I need to make sure... Okay, so it denotes the initial student checkpoint where RLVR denotes the initial student checkpoint and then the multi-teacher on-policy distillation.

So I'm not sure what this SFT column can figure out, but you can see the kind of like where the teacher is relative to on policy distillation. I think this is like the closest information we have on the relative performance gains.

Finbarr Timbers Yeah, that's fascinating because the DeepSeq, I forget which one, maybe it was V3.2 paper claims, or maybe it was... R1 actually claims that you can domain specific that you know doing the general stage captures the performance of it but you know that that doesn't really seem to be and yeah and then so you know doing the domain specific distilling in and then doing a general stage on top of that captures the original performance but that doesn't seem to be the case here like

you know the gap maybe isn't huge but there is still most of the time there's a pretty big there's like you know significant gap even if it's not huge so that's really interesting yeah I wish this table and text was

Nathan Lambert clear it's like I literally can't fully parse it it's like RLVR denotes the initial student checkpoint and then OPD denotes the checkpoint after first and second iterations it's like what is the checkpoint that was used at the start of odd policy distillation I think it was the RLVR one

Finbarr Timbers So they do a general SFT stage, then they do an RLVR stage that covers the non-teacher, the areas where they don't have specialized models. Then they do MOPD.

Nathan Lambert Yeah, and then that makes sense with this recovery rate, which is like final model minus RLVR, which would be like the gains for the OPD relative to the teacher minus RLVR, which would be like what gains you needed to still cover.

Finbarr Timbers Yeah.

Nathan Lambert And like what gains the teacher could potentially give you. So more research like this. Happy to see some of it out there. I'm going to switch back.

Finbarr Timbers Yeah, something I found interesting about both the Nematron papers and then the MAI Thinking paper is that they don't talk as much about some of the more detailed post-training decisions that have shown some pretty strong gains in some of the other papers. Like, I believe it was GLM-5 where they talk about doing a difficulty curriculum in a difficulty filtering stage. Yeah.

and that's just not something that's really talked about in these other papers. I think it was Kimi 2.5 used a temperature. It's kind of funny. So Kimi K2.5 and GLM 5 both have temperature schedules and they both claim the exact opposite thing. So one of them says you have to start with a high temperature and go low. The other one says you have to have a low temperature and go high. And You don't see that discussion, I don't think, in some of the other papers.

Nathan Lambert I still think the Chinese labs are much more willing to share really, really nitty-gritty details. The NVIDIA paper is mostly a list of methods to create a teacher, or domain-specific teachers, which... is useful but I think like I was less it's like less of a fun read they're like there's 15 pages of different domains so I'm like okay I don't like I don't need this yeah like Kimi K 2.5 and GLM 5 actually have like more similar recipes which are also on the simpler side which

is like you create this SFT stage and then you do RL the RL might be staged there's not this on policy distillation there's a bit less talk on how many experts they have and what their domains of experts are I think it's obviously you have to take All of this with a grain of salt and it's like what how they decided to present the information is like a big factor in this and like they might actually be closer in reality and then it just wasn't described in a certain way I thi

nk another

Finbarr Timbers interesting bit is that you see the Chinese labs all seem to be converging towards stars attention whereas we don't see that you know where the American labs at least NVIDIA and you know AI2 seem to be more converging towards hybrid Attention, like the Nvidia Nemetron Ultra used the Mamba Attention, whereas, you know, we see, you know, Deepsea Sparse Attention and then the MIMO MSA, whatever that stands for,MIMO Sparse Attention. So I think that's an interesting divergence.

Nathan Lambert Yeah, I am not the person to ask, but I agree. I often get asked, we'll avoid the full rabbit hole, but I often get asked, are the Chinese labs more efficient? And I'm like, I don't really think so. I think the financial pressures of serving billions and billions of tokens is probably a better motivator for efficiency, given that if you make a GPT model 1% more efficient, you're making fat stacks of profit.

like I think that's like a more effective market mechanism but and then the Chinese

Finbarr Timbers lab yeah if you make you know serving chat tpt more efficient Sam Altman can say hey here's a bunch of stock like yeah but they do great like the Chinese labs do

Insight

采购环境时拆分报价:软件、状态数据、任务 prompts、reward/eval、维护 SLA、独占权和公开权分别验收。

  1. 真正高溢价的部分可能不是软件外壳,而是能提升具体 benchmark 或能力的任务和 prompts(约 00:40:10–00:41:29)。
  2. 开放侧的机会在于环境较可并行:小团队可专攻一个高质量 environment/eval,由开放训练实验室组合使用(约 00:41:29–00:42:04)。
  3. 企业在 Agent 开始使用自身服务后,也可能有动力发布官方环境或 benchmark;但独占采购会进一步把资产推离开放研究(约 00:42:04–00:43:51)。

Nathan Lambert great research which I think it's kind of a bit different okay we can move into more open-ended stuff here I think that we have like We have a bunch of things in a document here. I'm sure more will come up. How do you think about open models and post-training recipes in the current environment's craziness? I think there's two things. One is like, what the heck is going on with environments? Is this a fad?

And then two is like, what is the equivalent or most recipes do in face of that?

Finbarr Timbers Well, I think the environment bit is interesting. I made the mistake of sending a tweet out at the start of the year where I said, hey, I want to buy some environments. Please give me your cell. And my inbox has been overwhelmed ever since, including my personal phone, where people would call me at 2 a.m. and say, hey. So that wasn't super fun, but it was really interesting. I did a bunch of sales talk to people and it's just, you know, it's quite expensive to get these envir

onments.

Like the quote that I got from one of them was like a hundred thousand per environment. This was a more, you know, like trying to copy like a, you know, sophisticated web app, like a door.

Nathan Lambert Do you know what this means? Like is the environment in the hundred K mostly the software or is it also like the software, the props? So it's like the software is like the world. or is it also a bunch of prompts and exact training data that you could just like throw into a recipe?

Finbarr Timbers That's a good question. I think it was more the, I don't think it included the prompts. I think it was the, here's the software, like here's the, you know, DoorDash clone that you can interact with. And here's a bunch of like fake data in it. Like, you know, we'll keep this service up for, you know, like it's a hundred grand upfront cost and it was like 50 grand or something per year per environment to keep it running. I don't think it included the actual prompts, but I could

be wrong about that.

Nathan Lambert I was just gonna say like a lot of it to be able to you can charge a much higher premium if you're like we have prompts and they know we know they improve X benchmark or X ability they're like here's an Amazon clone thing I think it's gonna be a lot harder to sell the labs in a sustainable fashion yeah it's like they just don't care about it as much I mean they care but I just think it's kind of a different category I'm actually less doom and gloom on environments in the open

because it is nicely unparallelizable where if you can figure out the right incentive structures and reward people for building good environments that are also good evals in specific domains, then it's like small teams could actually just do that. And then the open trading labs use said environments and figure out how to combine them and actually use them. I'm not sure that's happening in a substantive way,but it seems like a bit more tractable than building a big post-train

ing recipe as a small academic lab.

Finbarr Timbers Yeah, absolutely. And I think it's also, you know, once something becomes useful, right?

Like once you figure out that environment is particularly valuable and, you know, the frontier labs get really good at it. I think there's also a lot of incentives for the people involved to expose it. Like if you're DoorDash, you know, sure, maybe you're not initially going to say like, okay, I'll make this environment. I'll make my own environment of DoorDash for, you know, OpenAI or whoever to get good at. But once, you know, a bunch of people have agents out there and, yo

u know,people are starting to use DoorDash or use, you know, ChatGP or whatever to order from DoorDash, I think then DoorDash has an incentive to say, we'll release our own environment. I think the same is true, you know, a lot of this has marketing mechanism where you're going to say, oh, like, you know, we're, you know, Harvey or whoever, here's our difficult legal benchmark. Oh, look, we're the best at it. And then, you know, once you have the eval out there, then it becom

es some thing we can trade against it.

so I think that's my optimistic take but I do think it's kind of tough because the other side of things is that I think a lot of the RL companies are trying to turn around and sell the same environment to multiple people and so it's you know the cost might be 100 grand if you're selling to one player but then if you're trying to you know sell it to someone and then you know not sell it to anyone else well then the cost might be you know 500 grand or something because you know

you can't turn around and sell it to you know Anthropoc, OpenAI, Gemini Chinese Labs, AI2, whoever So, you know, that, you know, just puts it more, um, out of reach that release everything.

Nathan Lambert There are, there are multiple price tiers. And I think that some people will like pay more for a dataset to be able to release it, which is like a really baller move, but like not many people are going to be able to do this.

Finbarr Timbers No.

Insight

训练服务的技术评估应同时覆盖可移植性、权重归属、推理兼容、利用率和退出成本。

  1. 大量研究可能迁移到训练 API,因为规模采购能提供比学术自建集群更好的价格和 MOE 能力;代价是开放科学对少数接口收敛(约 00:43:54–00:45:21)。
  2. 训练—推理完全匹配可以改善利用率和部署体验,也可能形成基础设施锁定;API 生意最终仍需连接 token 收入(约 00:45:51–00:48:14)。
  3. 实验室在 pre-training 与 post-training 之间的算力分配最敏感,因为它直接暴露对下一阶段进展来源的判断(约 00:48:22–00:50:18)。
  4. 两位嘉宾建议区分短期套利与长期有价值的工作;大型新实验室的高估值和融资节奏可能压缩路线多样性(约 00:50:18–00:56:04)。

Nathan Lambert another thing we have is Tinker style APIs I feel mixed about this because I think a lot of small medium startups especially are going to use Tinker to do interesting research and share it with the world and there are like fully open Tinker style APIs which is like an RL framework that you run on your own GPUs and then gives you this Tinker style API but it's like I don't love for a like open science perspective for too much convergence on these even though I think the actual

price point will be very good because places like thinking machines buy compute at a mass scale price versus academics that have very finite compute and can't do some things like very large MOE runs so I see this coming especially if we look back in like a year we'll see a lot more research done on these APIs both in like industry and academia, but I'm kind of apathetic to it.

Finbarr Timbers Well, I like it, you know, insofar as you can expose this much simpler API for researchers to use. And especially, you know, if you can hide a lot of the complexity, I think that's ideal. Because then, you know, if you can only engage in the, you know, really interesting bits and make it so, you know, the research scientist doesn't have to think about How all of your, you know, VLM inference pipeline is set up. I think that's good.

I think, you know, my question is just how good of a business is it actually?

And it just doesn't strike me that there's this, like, you know, I think that there's a large business to providing. Well, actually, I thought he'd been super clear. You know, we've seen a number of companies providing, you know, RL fine-tuning services, you know, RL as a service. We've seen a lot of companies try to provide fine-tuning as a service. And, you know, none of them have really taken off. Like, I think OpenAI has started to shut down. I think they shut down their

RL fine-tuning. I think they might be shutting down their fine-tuning. Maybe I'm wrong about that.

Nathan Lambert Well, it's like, Cursor used Fireworks for their actual training run, and I'm like, I don't really know all the details of this, but Cursor does something for fast, I think like fast weight transfer, or Fireworks does a fast weight transfer and other things to make it so they can scale their RL Inference compute very nicely. So that's one type of it. I don't know how big of a long tail that business is. But also, I think Tinker is a better business than most people expected.

It makes some real amount of money.

It's like in the hierarchy, I think selling compute, not the best business. Selling inference, great business. And Tinker-like APIs, if you can't transition it into selling tokens, is somewhere in between the two, where they could take some amount of margin that'll be slightly higher than just selling the compute. And they obviously get a margin by having, like they get compute at a cheaper rate than their customers, and that's like part of the margin they're taking. But I do

n't see it being as nice as inference,so it's kind of existential for them to make it so that these fine-tuning APIs... feed into a inference business pretty nicely because then you can be somewhat locked in on you train the model on our infrastructure. You actually can own the model weights, but the training dynamics to inference mismatch is perfect because you trained exactly on our inference engine and are going to get what you want out of it.

Finbarr Timbers Yeah. And it also helps a lot with utilization because you can then, you know, utilize it. You can ensure that utilization across a lot of Clients. I think it makes a lot of sense. I think it's probably a better model for a lot of users. I think of academic users. It probably makes way more sense to do this. Or for that matter, if you're starting a new post-training lab now, as I know a few people who are,I think that's where it probably makes a lot of sense to start with som

ething like the Tinkerer API. And then at some point, if you want to try and capture that margin, maybe then you try to do something more Custom, but if you can use something like that, like, that's great. And the economics are, you know, fundamentally more sustainable, or, you know, better for you, rather than trying to, you know, go to core weaver, whoever, and say, or server scale and say, hey, I need,you know, 10,000 networked GB200s, you know,that's just a very expensive

thing to do, especially if you can't keep it running all the time.

Nathan Lambert Do you have any more hot takes on post training before I ask you some more general things?

Finbarr Timbers Well, something I'm generally interested in, and I'm the wrong person to speak to about it, I'd love to talk to someone who's maybe a capital allocator, or a compute allocator, who's deciding where to put compute or where to hire team members, because I'm kind of curious how the high-level decisions are made,

Nathan Lambert allocating resources between pre-training and

Finbarr Timbers and post trading. Because, you know, what I kind of have seen as a general trend is that you see a lot of papers where there's, you know, more focus put on one or the other. Like I think, so yeah, so that's something kind of interesting to me is how people who are, you know, making this decision, how they're making that decision and how they're thinking about it.

Nathan Lambert yeah it's like the hardest decision to get out of labs I've like I used to spend time trying to get them to share more but I think it's like such a sensitive decision to where they see progress coming like they're making that decision allocating compute based on where they think the most progress is and what they'll like return on investment is so if you go to Anthropic and they're like here's our distributions it's like okay that's where labs see their bets and or where they

see they are weak and it's like you invest more compute to make progress in the area that you are interested in which I always think makes a lot of open research kind of Boring right now is like the people that get compute are just way more likely to succeed as academics and researchers, which is a horrible equilibrium for the world, but kind of realistically true. I don't know how to make a lot of that. I wanted to ask you how you feel about the craze that people have to cas

h in on making money and join a lab before the ladder gets pulled up and what people should be optimizing for in their careers in face of meaningful opportunity costs.

Finbarr Timbers Yeah, I think that's actually very timely.

Nathan Lambert But yeah, I think that that's really important to talk about.

Finbarr Timbers I mean, I think it's always worth focusing on whether what you're doing and spending time on is going to be generally valuable or if it's like a really short-term exploitation type thing in the RL, like explore versus exploit setup. I mean, something that I've seen throughout my career has been often the places that pay the most are all the places where you're doing the most interesting work, right?

Like, you know, if you're going to go work at OpenAI, or you're on Anthrop or you're at the front of your lab,but they pay a lot of money. They also have a lot of resources. So you're going to make a lot of money and learn a lot. So I think it's worth trying to decide, is that, is the I think that would have been a mistake but trying to figure out if you're going to be able to do interesting work is really important and also try to figure out if you're going to be able to, yo

u know, push forward science, you know, if what you're doing is more just saying going to, you know, data vendors and saying, you know, okay, you know, we I need a bunch of data to do. whatever and then you know they give you a bunch of data you train a model you say it's good you're bad or whatever you know I don't think that's as interesting and I don't think you're going to learn a lot even though that's you know work that would probably drive model progress for it I think

if you're able to you know make focus more on the science and make more scientific conclusions I think that can be you know a lot better for your long-term career and I think that's where places like AI2 and the other academic research labs you know Moran is doing a really great David Pérez Pérez Pérez Pérez Pérez Pérez

Nathan Lambert Yeah, mostly this is grounded in visiting the Bay Area and every time I go I'm like, holy shit, what is going on here?All these very junior people have way too much dread about their opportunity costs and both of us aren't based in the Bay Area so I feel somewhat removed from it, which gives me a little bit more Time to pause and be like, what exactly is the right thing to optimize for?

It's easy for me to say as somebody who's established, but I think there's opportunity for a lot of people to just,if they have conviction on something, to try to go and do it and not just follow everybody that goes down the funnel of joining one of the established labs or the neolabs. where I don't hear from many people that join as a junior person at these places and end up with very high responsibility like they're contributing to something that matters or they're around a

cool group of people but I don't hear from that many people are like wow I'm doing the highest leverage stuff and the most interesting things

Finbarr Timbers Well, I think that, you know, it's kind of funny for me to say this. My career has been more on the opportunistic side of things. But, you know, twice now, I've been at organizations where I've been working. So, you know, at DeepMind, I was part of the Alberta office where DeepMind had, you know, aqua hired. the computer poker research group from the University of Alberta and so you know this was a group of people who were really invested in computational game theory and you

know poker playing algorithms and they were all in on that and you know they were all in on that to the point that you know they were one of the two leading labs in the field and where, you know, because they were so strong at this, DeepMind came and, you know, Aqua hired them and they all joined and they, you know, did quite well from that acquisition there. And, you know, I joined later because I was,you know, interested in working with them and doing game theory and stuff.

But, you know, it was this group of people who had this conviction that what they were doing was really important and, you know, it worked out quite well for them. And then, you know, the same thing at AI2 where at AI2, you know, there's all of these people who were really interested in NLP Research or even before language models like we see people like you know like Kyle and Dirk I think we're both at AI2 for like almost a decade like they had these really long Tenters and

then they did really well and then you know they've since had some you know strong opportunities coming out of that with with yeah some of the opportunities that have been available to them and I think that the consistent theme there has been that you know if you have high conviction that what you're doing is important and interesting then like it's not a mistake to follow that and to you know try to become really strong in that area

Nathan Lambert Yeah, I mostly think it's good for the world to have a more diverse set of approaches. It'll be interesting to see what the Neolabs actually produce, if they can manage to do things that are diverse. My personal idea is that they're so big now that most of them need to end up doing something that is somewhat similar, which is hard, but like... They need to keep risking their $20 billion valuations to do something interesting that's not just going to be squashed by an opening

AI or an anthropic side project.

Finbarr Timbers Yeah, absolutely. And I think it's tough because when you're raising, when you have these huge seed rounds, you're raising $200 million or $1 billion or whatever, then it's like you have to pretty quickly show results to be able to grow off of that.

Nathan Lambert yeah so a to be continued conversation any last words I don't need to stretch it on if we don't have anything to add to our conversation no I think this is pretty good

Finbarr Timbers I think it was really great getting a chance to catch up and talk about some of this stuff you know I've been reading all these papers and thinking about all the different recipes so it's great to get to chat about it and put it out into the

Nathan Lambert ether so yeah thanks for having me on yeah thanks for coming back we'll talk soon

Finbarr Timbers Sounds good.

回到顶部