# Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth · 中英对照逐字稿

- 原节目：AI Engineer
- 英文原始来源：https://www.youtube.com/watch?v=uIiA6DquRiE
- 中文译制版入口：https://www.xiaoyuzhoufm.com/episode/6a62f7236356eb2d9be7a09f
- 时长：02:20:20
- 方法与限制：英文来自已验证的原始节目 transcript/caption；中文由 Codex 逐段翻译，未做逐字人工校对，公开引用前请回到英文原文与音频复核。

## 中英对照逐字稿

### [00:00:13–00:00:42]

**EN**  Hello everyone. Um yeah, thanks so much for coming today. Much appreciated. Um yes, I'm Daniel. I'm from Anslaf. My brother is also here today. Um but yeah, like you know, thanks for coming. Um, so for you folks who don't know us, um, so we actually, you know, we're one of the largest distributors of language models and diffusion models as well. So we don't just do language models. We upload our models to hugging face. Um, and you know, we're on the I think we're

**中文**  大家好。非常感谢大家今天到场。我是 Daniel，来自 Unsloth，我弟弟今天也在。谢谢各位前来。可能有人还不了解我们：我们其实是 language model 和 diffusion model 最大的分发者之一，并不只做 language model。我们把模型上传到 Hugging Face；我想我们在

### [00:00:39–00:01:08]

**EN**  number 10 or something on the Oh, no, I don't I don't remember. But anyways, we're on the list of the top organizations on Hugging Face. Um, we have over 300 million total downloads. Um, so definitely check us out on that. Um you can run like you know Deepseek, GLM, many other models and we quantize them down using dynamic quantization. Um so you can run them on your local computer. Um we also do many bug fixes for open source models. Um so you know we you

**中文**  榜单上大概排第十——我记不清了。总之，我们是 Hugging Face 顶尖组织之一，累计下载超过 3 亿次，欢迎去看看。你可以运行 DeepSeek、GLM 等许多模型；我们通过 dynamic quantization 进行量化，让它们能在本地电脑运行。我们也为 open source model 修复许多 bug。

### [00:01:05–00:01:35]

**EN**  know fix many bugs in you know OpenAI's GPUs um you know Meta's models um Google's models deepseas many other models we fix bugs in them. Um and so like you know they have many issues sometimes and then we post about them on Twitter. um you know we post about our findings um so you know m most of the open source models that you probably guys have used um are most likely fixed by us um and yeah like we collaborate with everyone in the entire world um on

**中文**  我们修复 OpenAI GPU、Meta model、Google model、DeepSeek 等许多模型中的 bug。这些模型有时问题很多，我们会在 Twitter 发布发现。大家用过的大多数 open source model，很可能都经过我们修复。我们也与全世界的人合作

### [00:01:32–00:02:01]

**EN**  you know model releases um yeah we also collaborate with hardware providers and you know we really appreciate the collaborations with everyone we also don't just do model fixes and bug you know bugs we also introduce new features and we also like you know do fixes for the entire trading stack. Um for example, we introduced something called async gradient checkpointing which is used by many organizations. Um we also introduced flex attention which is used by many folks. Um and we also

**中文**  发布模型，并与 hardware provider 合作，非常感谢所有合作伙伴。我们不只修 model bug，也会引入新 feature，并修复整个 training stack。例如，我们推出许多组织使用的 async gradient checkpointing，也推出许多人使用的 Flex Attention；还

### [00:01:58–00:02:27]

**EN**  fixed a gradient accumulation bug fix um which increased accuracy by 1 to 3% um across the entire training stack. Um so we don't just like you know do bug fixes for models um it's also like you know whole training stack um fixes and stuff like that. So today, you know, the workshop is quite long. Um, so there will be multiple sections in the workshop. Um, and so after each section, anyone can ask a question. Um, and so, you know, please, I guess, if I'm not sure if

**中文**  修复了 gradient accumulation bug，使整个 training stack 的 accuracy 提高 1% 到 3%。所以我们不只是修模型，也修整套 training stack。今天的 workshop 很长，分为多个 section；每个 section 结束后都可以提问。我不确定现场是否有 microphone，

### [00:02:26–00:02:55]

**EN**  there's a microphone, but if you could raise your voice and you ask a question, you know, I'm more than happy to answer them. Um, but you know, the first section we're going to be talking about is the state of AI. So, where is currently language models, AI models, where are they at currently? Um so I'm not sure if everyone knows the meter plot. Um so this meter plot shows the time horizon of uh models. Um if you can you know every single task if it

**中文**  大家可以提高音量提问，我很乐意回答。第一部分会讲 AI 的当前状态：language model、AI model 目前走到了哪里。不确定大家是否知道 METR plot；它展示模型能够处理的 task time horizon。例如一项任务若需

### [00:02:52–00:03:21]

**EN**  takes a human 16 hours can a model you know finish that task. Um and you can see on this plot you know cloud mythos you know preview is very good. It can do tasks that humans can do that take you know human 16 hours. Um you know opus 4.6 is also there. you know all the other models are also there and so you know this plot is very good because it symbolizes the AI models are getting better and better and better over time

**中文**  人类 16 小时，模型能否完成？图上可以看到 Claude Mythos preview 很强，能做需要人类 16 小时的任务；Opus 4.6 和其他模型也在图上。这张图很好地说明 AI model 会随时间越来越强。

### [00:03:19–00:03:48]

**EN**  you know recently with the launch of you know GPD 5.6 um you know just well their preview model um you know just on Friday um you know I put the plot so they didn't so me didn't actually update their plot um because they said that the results were not trustworthy enough um but you know I just put it on the plot um and so you can see that GBD 5.6 you know is around you know opus 4.6 sixth level I guess with large confidence bounds um so it's

**中文**  最近 GPT-5.6 preview 在周五发布。METR 没有更新图表，因为他们认为结果还不够可信；但我把它加到了图上。可以看到 GPT-5.6 大致处于 Opus 4.6 水平，只是 confidence bound 很大，所以

### [00:03:46–00:04:15]

**EN**  very you know uncertain about the capabilities of the model um however if you include cheating so if you include that the model sometimes likes to cheat on some of the tasks then it actually goes to 270 hours um so we're directly and you know if you look at the y-axis I actually did a disjoint graph um so the y-axis is 50 hours skipped to 250 hours um so if you can imagine the graph is actually very skewed Um when I like made

**中文**  模型能力仍很不确定。不过，如果把 cheating 算进去，也就是允许模型在某些任务上作弊，time horizon 会达到 270 小时。请看 y-axis：我做了断轴图，从 50 小时直接跳到 250 小时。可以想象原始图有多偏斜。我制作

### [00:04:12–00:04:42]

**EN**  the graph, um GBD 5.6 was like a very big outlier. Um so I had to like compress the graph. Um but this only you know this graph only works if you consider that GBD 5.6 cheated on some of the tasks. Um and so we'll be talking about you know why AI models cheat and how do we like you know solve these issues. Um but yeah this plot is very useful to showcase the capabilities of these models. So previously this is 50%. you know if you could if a model can complete the

**中文**  图表时，GPT-5.6 是非常大的 outlier，只好压缩图形。但只有把 GPT-5.6 在部分任务中作弊算进去，这个结果才成立。后面我们会讨论 AI model 为什么作弊，以及怎样解决。这张图很适合展示模型能力。这里原本用 50% 指标，也就是模型有 50% 概率完成

### [00:04:40–00:05:09]

**EN**  task with 50% of the ch you know of the time to 50% accuracy if you want to actually oneshot the model so you just ask the model you know implement X or implement Y um and you want the model to do very well then you want to look at the 80% success rate if you look at the 80% success rate it kind of drops quite a lot um so you can see that previously mythos is around 16 17 hours um now it only can do three hours so if you prompt

**中文**  任务。如果希望 one-shot：只要求模型实现 X 或 Y 一次，就表现很好，那么要看 80% success rate。到了 80% 指标，能力会下降很多。之前 Mythos 大约能处理 16、17 小时任务，现在只剩三小时。也就是说，如果你 prompt

### [00:05:08–00:05:38]

**EN**  a model and you want to have like a oneshot example, you know, you just trust the model by just or asking it, you know, implement, I don't know, page rank or something, you know, implement some sort of rag system, you know, fine-tune a model or something like that. Um, it can only do a task that will take a human three hours to do. Um, and so the so that is a problem with AI models. Um, generally speaking, if you want to use AI models very well, you need to prompt it at least like, you know, five times or something. Um and each of those times assuming they're

**中文**  模型并希望 one-shot 成功，只交给它一个任务，比如实现 PageRank、某种 RAG system 或 fine-tune 模型，它只能可靠完成大约相当于人类三小时的工作。这是 AI model 的问题。一般想把 AI model 用好，至少需要 prompt 五次左右。若每次结果

### [00:05:35–00:06:04]

**EN**  independent um the success rate is much higher if you're prompted many many times right you can't just call the model once and expect it to do work to do well um you need to call it multiple times um and you can also work out the probability of it like succeeding you know if the model is 50% accurate um then it will be 50% failure then it's 1 minus 0.5 to the power of five or something like that you know if you do five turns and then your success rate jumps to like 97% or something um so you

**中文**  相互 independent，多次调用会显著提高 success rate。不能只调用一次就期待做得很好，而要调用多次。概率也可以算：若模型 accuracy 为 50%，failure rate 也是 50%；调用五次，成功率就是 1 减 0.5 的五次方，大约 97%。所以

### [00:06:02–00:06:32]

**EN**  need call the model at least five times for it to be very effective. So previously these are linear you know this is a linear trend you know on the y-axis it's just it's not you know it's just linear um if we log it you know if we log the y-axis you can see that it's more exponential progress um so it's actually a straight line fit to the entire progress of AI models on the

**中文**  至少调用五次才很有效。之前图表的 y-axis 是 linear；若改成 log scale，就会看到 AI model 的进展更像 exponential growth，整个 METR time horizon benchmark 上几乎呈一条直线。

### [00:06:29–00:06:59]

**EN**  meter time horizon um you know benchmark you can see that you know it's very clear that AI models are getting better and better over time um I als We also added you know GBD 5.6 six with the cheating and no cheating and also claude mythos are you know accentuated that and you can see I you don't need now you don't need to like you know fake the y-axis you know you don't need to do like a disjoint y-axis um if you do that you can see that you know models are getting better over time um and supposedly you know if this trend

**中文**  可以清楚看到 AI model 随时间不断变强。我们还加入 GPT-5.6 cheating 与 no-cheating 的结果，并标出了 Claude Mythos。用 log scale 后无需断轴，也能看到模型持续进步。假设这条趋势

### [00:06:56–00:07:26]

**EN**  continues these models will get better and better and better better and much better um yeah so so the question is if the trend continues you know that's the fundamental question Um and it's not just you know one specific task for this benchmark that you can see that models are getting better over time across all benchmarks models are getting better over time right so like you know GPQA diamond you know it's kind of plateau you know it's kind of already saturated as a benchmark but over time you know it

**中文**  延续，模型会越来越好。根本问题是：趋势会不会继续？而且不只这个 benchmark 的特定任务如此，所有 benchmark 上的模型都随时间进步。GPQA Diamond 也许已经趋于 plateau、benchmark 接近饱和，但模型仍

### [00:07:24–00:07:54]

**EN**  does very well you know every single benchmark you see models are getting better right live code bench you know maths algorith maths tests um you know even Tesla's you know you know self-driving I guess is also has like a doubling time of 17 months. Um so every single 17 months the models will get better and better. Um you know double their capabilities. Um so over time all these models in every single subject you know every single

**中文**  表现很好。每个 benchmark 都显示能力上升，例如 LiveCodeBench、数学和 algorithm 测试。甚至 Tesla self-driving 的能力 doubling time 也大约是 17 个月，即每 17 个月能力翻倍。因此模型在每个 subject 和

### [00:07:50–00:08:18]

**EN**  area it will get better. Um so I guess the main question is you know if we assume every single subject every single area the models get 100% like you know approaching 100% accuracy is this AGI? Um so that is one of the fundamental questions that people ask you know if we just get better on benchmarks um is this AGI um what happens if we get better on all benchmarks you know every single benchmark that human humanity has

**中文**  area 都会随时间变强。主要问题是：如果假设每个 subject、每个 area 都接近 100% accuracy，这算不算 AGI？这是人们会问的根本问题。如果只是 benchmark 不断变好，算 AGI 吗？如果人类创建的每个 benchmark

### [00:08:16–00:08:45]

**EN**  created it just gets better on all of them. Um yeah but so this is a you know very good plot show well I guess chart showing all of the different types of benchmarks and they all get better over time. everyone's favorite I guess artificial intell uh you know artificial analysis benchmark showing you know artificial intelligence getting much much better over time as well you know fable I guess is I guess the best for now um although not everyone can access it currently but

**中文**  都持续提升，会发生什么？这张 chart 汇总多类 benchmark，它们都随时间变好。大家熟悉的 Artificial Analysis benchmark 也显示 artificial intelligence 越来越强。Fable 目前似乎最好，虽然还不是所有人都能访问。

### [00:08:43–00:09:11]

**EN**  anyways it's for now it's the best um and you can see over time that you know these models are getting better over time as well um and you know like this plot showcases um a very useful indication you know like how do we like you know benchmark you know is this benchmark actually good um in terms of like you know showcasing the capabilities of models as well. Um and we'll be also discussing about that as well. Um on the other hand yes models are getting better over time. Um but

**中文**  总之它暂时最好。图中可见模型仍持续进步。这张图也提供了一个有用提示：我们怎样 benchmark？这个 benchmark 是否真的能展示模型能力？后面也会讨论。另一方面，模型确实越来越强，但

### [00:09:10–00:09:38]

**EN**  there are some things which models are not very good at still for example long context is not doing very well. Um so you know most models you might say okay Gemini has 1 million context length. You know GBD has 1 million context length. Claude has 1 million context length. But should you actually use all of the 1 million context length? Um so there are actually benchmarks to showcase that if you use for example GBD 5.5 um you know if you use 512 context your accuracy

**中文**  有些事情仍不擅长，例如 long context。许多模型宣称有 100 万 context length，Gemini、GPT 和 Claude 都是如此；但真的应该把 100 万全部用完吗？一些 benchmark 显示，例如 GPT-5.5 使用 512K context 时，accuracy

### [00:09:35–00:10:04]

**EN**  reduces to 50%. Um so if you use you know 512 context you will only remember 50% of the facts that you wrote in the previous context. Um so maybe that's not a good idea to use the full context. Um you can see opus 4.7 um 4.6 4.7 is the very last orange line. Um so at the context length of 256K it goes to 0%. Um so this might be a benchmark flaw. Um so maybe don't trust the benchmark too

**中文**  会降到 50%，也就是只能记住前面 context 中 50% 的事实。因此用满 context 也许不是好主意。Opus 4.6/4.7 是最后一条橙线，在 256K context length 时甚至降到 0%。这也可能是 benchmark flaw，所以不能太

### [00:10:02–00:10:31]

**EN**  much. Um but it's good to look at the benchmark overall. You know where is the model's capabilities for long context. Um the blue lines I highlighted are open source models. You know deepseek gl 5.1 other models. Green is Google's models. Um but you can see in general you know models are models definitely do degrade over long context. Um so if you you know for example if you set like a you know automatic compaction area I would not

**中文**  相信单一 benchmark，但可用它整体观察 long-context capability。蓝线是 open source model，包括 DeepSeek、GLM 5.1 等；绿线是 Google model。总体看，模型在 long context 下确实退化。因此，例如设置 automatic compaction 时，我不

### [00:10:28–00:10:58]

**EN**  suggest you to use all 1 million context maybe maximum 600k or something um and then compact it and then continue your you know coding session um but I would yeah but in general you know this plot shows that long context still has a very long way to go um and if we want to have long context you know capabilities um labs I guess will have a lot of time to fix this problem. Yeah. So another plot is you know just

**中文**  建议用满 100 万 context，也许最多用到 600K 就 compact，再继续 coding session。总体而言，这张图说明 long context 还有很长的路；想获得真正的 long-context capability，实验室还要花很多时间解决。另一张图

### [00:10:56–00:11:24]

**EN**  showing open source versus closed source. So open source still has some way to go for this you know long context. Um so open source is blue line and the black lines are like you know closed source models. Um and you can see in general open source does okay but there's definitely much more room for improvement. Um I guess compared to Opus 4.7 it's better. Um but you know maybe this benchmark does need maybe there are some flaws in the benchmark as well. Um yeah, but overall you know this plot

**中文**  比较 open source 与 closed source 的 long context。open source 是蓝线，closed source 是黑线。总体看 open source 表现尚可，但仍有很多提升空间。与 Opus 4.7 相比也许更好，不过 benchmark 本身也可能有缺陷。总体上，这张图

### [00:11:22–00:11:52]

**EN**  shows that long context definitely still has more room for improvement. And also you know like if you looked at the plot previously you know this meter plot um I'm not sure if you can see that before 01 preview there is actually a plateau of performance. Um and so if you can see you know GBD4 to GBD40 there's not that much performance improvement. Um and so this time frame around one

**中文**  表明 long context 仍有改进空间。回到之前的 METR plot，在 o1-preview 之前其实出现过 performance plateau。从 GPT-4 到 GPT-4o，提升不大。这段时间大约持续一年，

### [00:11:48–00:12:17]

**EN**  year um was when you know the labs were confused on what is next um you know before 01 preview which showed that reasoning was very important they didn't actually know what to pursue next um and so for one year the models kind of plateaued um and so I call this the intelligence plateau the hypothesis that you know you know assume that we never have discovered reasoning then maybe air models would have like plateaued um but because we have

**中文**  实验室当时不知道下一步做什么。在 o1-preview 证明 reasoning 很重要之前，大家并不知道该继续追求什么，所以模型 plateau 了一年。我称之为 intelligence plateau：假设我们从未发现 reasoning，也许 AI model 就会停止提升。但因为我们

### [00:12:15–00:12:43]

**EN**  discovered reasoning you know we have shown that models can do reasoning capabilities we have continued the trend continuously um and so normally I don't know if this is like luck um or if this is a self fulfilling prophecy um so I don't know if you guys you know the moor law um you know mos law has continued um not because of the law but because people know that it must continue and so people invest money into the resources to make the law continue um and so this

**中文**  发现模型能做 reasoning，趋势得以延续。我不知道这是运气还是 self-fulfilling prophecy。大家知道 Moore's law 吗？它之所以持续，不是因为一条自然法则，而是因为人们相信它必须持续，于是投入资源让它继续。这个

### [00:12:41–00:13:09]

**EN**  kind of like shows that you know we might have been in of the world where models have stopped improving. Um but you know with the launch of 01 preview you know I guess models have went back to trend. In fact I made a plot showcasing you know assuming we did not discover reasoning or 01 preview. Um then the black line was the supposed you know capabilities of the models. You can see

**中文**  例子说明，我们本可能进入模型停止进步的世界；但 o1-preview 发布后，模型回到了趋势线上。我画了另一张图，假设我们没有发现 reasoning 或 o1-preview，黑线就是模型原本可能达到的能力。可以看到

### [00:13:07–00:13:35]

**EN**  I made it into a S shape um like a you know a sigmoid type shape. Um and if you know if we didn't discover reasoning then models definitely will taper off in terms of capabilities right we'll only have a model that's as capable as claw 3.7 sonnet I guess or 01 or something like that um but you know luckily because of reasoning and this new paradigm of scaling you know the green line is the new scaling law um and you can see previously the black line the

**中文**  我把它画成 S shape，也就是 sigmoid。如果没发现 reasoning，能力会逐渐变平，最高大概只达到 Claude 3.7 Sonnet 或 o1 的水平。幸运的是，reasoning 与新的 scaling paradigm 带来绿色的新 scaling law。之前黑线的

### [00:13:33–00:14:01]

**EN**  doubling time was actually around seven months so every single seven months the capabilities of the models double um but now it has shrunk to 3.5 months. So every single 3.5 months you just need to wait 3.5 months and the models will get double better, right? Better by two times. Um and that's quite striking I guess. Um so the main question though is will the green line continue as a straight line? Um that is a fundamental question that labs are still struggling

**中文**  doubling time 大约七个月，即模型能力每七个月翻倍；现在缩短到 3.5 个月，只要等 3.5 个月，模型就会强一倍，这很惊人。不过关键问题是绿色直线会不会继续，这仍是实验室苦苦思考的

### [00:13:59–00:14:28]

**EN**  on. you know what happens if the green line again you know the green line again goes as a S shape you know that's possible um but you know we don't actually know if this will happen you know if the green line will continue scaling you know going all the way up to infinity I guess or would it be like an S shape um and this is you know many researchers are you know I guess have sleepless nights you know what is the next you know what is the next thing afterwards after reasoning after 01 you

**中文**  根本问题。绿色线也可能再次变成 S shape。我们不知道它会持续上升到无穷，还是最终形成 S shape。许多 researcher 为此彻夜难眠：reasoning、o1 之后，下一件事是什么？

### [00:14:26–00:14:56]

**EN**  know what is the next thing afterwards um and you know Many researchers will need to like you know I guess think about this. Um yeah but you know this plot is very you know this is one of my favorite plots because it shows that you know AI progress can continue over time with new ideas and innovation. Oh yes. So does anyone have any questions for the first section? Um yes. >> So we came all the way to one trillion right? Do you think the next jump if we

**中文**  许多 researcher 必须思考它。这是我最喜欢的图之一，因为它说明新的 idea 和 innovation 能让 AI progress 延续。第一部分有人提问吗？观众：我们已经做到一万亿参数。下一次跃升是否需要

### [00:14:54–00:15:24]

**EN**  need do we need like 10 trillion parameters when we'll see the jump or hardware will be the limitation that >> yes that's a great question. So the question was you know models we're currently at one trillion parameters do we need to go to 10 trillion parameters or more for models to be even more capable? Um so the scaling laws does say that you know if you multiply the parameters and the data size um generally speaking this number if you increase the number you will get the models become more capable. So yes you

**中文**  十万亿参数，还是会受到 hardware 限制？嘉宾：这是个好问题。我们现在大约有一万亿参数，是否需要十万亿乃至更多才能进一步提高能力？scaling law 表明，把 parameter count 与 data size 相乘并提高这个数，模型通常会更强。所以

### [00:15:21–00:15:50]

**EN**  can increase the parameters by 10 times and in general your performance will increase. Um however the view is there is going to be diminishing returns. Um I feel like you know it's not just the model size times the data set size. It's actually a ratio um some sort of like power law when you multiply them. So you actually get diminishing returns over time. So yes, you're right. If you want to have actually I'm not sure the exact law, but if you want to have double capabilities, you do need to 10 times

**中文**  把 parameter 增加十倍，performance 一般会提高。但会有 diminishing return。我认为并非简单的 model size 乘 dataset size，而是某种带 power law 的比例，因此回报会递减。具体公式我不确定，但若想让能力翻倍，可能需要把

### [00:15:48–00:16:16]

**EN**  the parameters. Um and then if you want another double, you have to 10 times it again. So it's 1 to 10 to 100 trillion parameters. um if you want maybe that's not a good way to scale. Um maybe instead you know instead of making a 100 trillion parameters some sort of new algorithm or new architecture could solve that problem. Um but you're right like if you're a lab you want to do something easy and so the easiest path

**中文**  参数增加十倍；再翻倍又要增加十倍，于是从一万亿到十万亿，再到一百万亿参数。也许这不是很好的 scaling 方法。相比构建一百万亿参数模型，新 algorithm 或 architecture 也许更好。但如果你是实验室，会想选择容易的事情，最容易的路径

### [00:16:15–00:16:45]

**EN**  is to just make a 10 trillion parameters. Um but I would say like you know maybe a new algorithm will be better. Um yeah any other questions? Yes. >> So you do think that we are approaching the limitation of next token prediction. >> That is a good question. I would say that for next token prediction it's very powerful because you can essentially the human language is extremely powerful

**中文**  就是直接造十万亿参数模型；不过新 algorithm 可能更好。还有问题吗？观众：你认为 next-token prediction 已接近极限吗？嘉宾：好问题。next-token prediction 非常强，因为 human language 本身极其强大，

### [00:16:44–00:17:12]

**EN**  and it doesn't have to be human language. It can be you know maths coding you can just predict the next word and in order to predict the next word or token you need to know everything about that token or that word right so like I think IA was talking about like you know Ilia Satska he was saying like you know you need to have you need to make a weld model in the model in order to like predict the next word so I still think next word prediction still has a lot of way to go for example if you see this plot you

**中文**  而且不必局限于 human language，也可以是 math 或 code。仅预测下一个 word，就必须了解与这个 token 或 word 有关的一切。Ilya Sutskever 曾说，为了预测下一个 word，模型必须在内部建立 world model。所以我认为 next-word prediction 仍有很长的路。看这张图，

### [00:17:09–00:17:37]

**EN**  know if we didn't have reasoning I guess okay maybe it would have plateaued But because we have discovered this new methodology you know reasoning and trying to like scale even more on next word you know next word prediction we have you know went back to trend um I feel like so the main question is if we don't have next word prediction what is next um that is the fundamental question most I mean I'm not sure like you know I'm not certain what's what's the next thing

**中文**  没有 reasoning 时也许会 plateau；但发现新的 methodology——reasoning，并继续扩大 next-word prediction 后，我们又回到趋势线上。主要问题是：如果不做 next-word prediction，下一步是什么？这是根本问题，我不确定答案。

### [00:17:36–00:18:04]

**EN**  I feel like next word prediction is just extremely powerful because it's very easy to formulate and you can just like you know you can have like you know because attention is very powerful as well. You can have, you know, this special cause of attention mechanism and it's very efficient to train. So, I'm not sure I think the main question is I'm not sure what's next. Um, I guess researchers will like, you know, they're trying to scratch their heads, you know, what is next afterwards? Um, yeah. I I Yeah. Yes.

**中文**  我觉得 next-word prediction 极其强大，因为很容易 formulate；attention 也很强，可以采用某种特殊 attention mechanism，并高效训练。所以我仍不知道下一步是什么。researcher 都在绞尽脑汁想之后怎么办。

### [00:18:02–00:18:32]

**EN**  >> Just a follow up on it. Do you feel like we are in the same era like how we attention came out? >> Right. So, attention like we don't know what's next. Yes, that's a fair followup. So, um, you were mentioning how it's kind of like LSTMs or the old AI world. We don't know what's next afterwards. Um, that's a fair point. I feel like so like, you know, previously this

**中文**  观众：追问一下，你觉得我们是否处在类似 attention 刚出现时的时代？嘉宾：对，我们不知道 attention 之后是什么。你提到这有点像 LSTM 或旧 AI 时代，不知道下一步在哪里，这是合理的。之前

### [00:18:30–00:18:59]

**EN**  example, right? So, after GBD4, it was just pre-training, some, you know, supervised fine tuning, some ROF, you know, some RL um, and they waited one year until 01 preview. So in this one year of fog you know the fog of war we don't know what what was next and so researchers you know were scrambling you know do we do the reasoning process do we make pre-training better do we make the model bigger and bigger and bigger you know they tried all these experiments um and

**中文**  GPT-4 之后，只有 pre-training、一些 supervised fine-tuning、一些 RLHF 和 RL；大家等了一年才看到 o1-preview。在这一年的 fog of war 中，没人知道下一步是什么。researcher 四处尝试：做 reasoning process、改进 pre-training、不断放大模型；各种 experiment 都做过，

### [00:18:57–00:19:26]

**EN**  reasoning was the one that won I guess um but I think like I think the main question is is the green trend going to continue at the current time it looks like it's continuing once we see models starting to taper out in intelligence, you know, in capabilities, then we'll go back to the, you know, olden days of like, you know, this one year waiting period. But I think for now, these models seem very powerful. Um, yeah. So, like I'm not sure if this will, I mean, if you look, okay, if you

**中文**  最终 reasoning 胜出。现在关键仍是绿色趋势能否延续。当前看仍在持续；等模型 intelligence 或 capability 开始变平，我们就会回到过去那种等待一年的日子。但目前模型看起来很强。我不确定趋势是否会停止；如果

### [00:19:23–00:19:52]

**EN**  squint at the plot, I guess maybe we're tapering out. Maybe um let's not consider the GBD 5.6 cheating example, right? Let's remove that from the plot. Um, but you can see the GBD 5.6 Mythos, you know, 4.6. They're kind of all I guess they're kind of tapering. Um so maybe as a I mean I don't know if we someone wants to bet on this but you know maybe models have tapered out but we're not sure. So we shall wait a few more months and see. So let's wait 3.5

**中文**  眯眼看图，也许已经在变平。先删掉 GPT-5.6 cheating 的结果；GPT-5.6、Mythos、Opus 4.6 似乎处在相近水平，也许已经 tapered out。我不知道是否有人愿意下注，我们仍要再等几个月。等 3.5

### [00:19:50–00:20:19]

**EN**  months. If we wait 3.5 months and see the models do not improve then we have tapered out. Um but remember we only need to wait 3.5 months. Um so then this law will fail. In fact, if you wait seven months, if you wait seven months, so double the time and models have, you know, just assume you know that dotted line that if if the models just follow the dotted line, okay, then we have tape it out. And I would agree that, you know, we'll have to design something new in, you know, make some new invention or something like that. Um, but for now,

**中文**  个月，如果模型没有进步，就说明趋势变平，这条 law 失败了。事实上，如果等七个月——两倍时间——模型仍只沿虚线走，那我同意我们必须设计新东西、做新 invention。但目前

### [00:20:17–00:20:46]

**EN**  you know, for now looks like it's doing fine. Um, yeah. Okay, next section. Um, so every single section we can have questions, so you can ask as many questions as you like. Um the next section we're going to talk about is open versus closed models. Um so artificial analysis has this very cool plot showcasing the performance of open source. So open source is the blue line. Um so open source is the blue line and closed source models is the black line. Um and you can see that open source does

**中文**  看起来没问题。进入下一部分。每部分都可提问。接下来讲 open model 与 closed model。Artificial Analysis 有一张很酷的图：open source 是蓝线，closed source 是黑线。可以看到 open source

### [00:20:44–00:21:14]

**EN**  lag. You know open source definitely lags over time. Um another very good benchmark is called the weird ML benchmark. Um this also shows that open source models lag closed source models right the blue line is open source models the green line is closed source models um and you can see over time you know the x-axis is release date of the model and the y-axis is performance and you can see that open source models kind of lag close source models

**中文**  确实随时间落后。另一个很好的 benchmark 叫 WeirdML benchmark，也显示 open source model 落后于 closed source model。蓝线是 open source，绿线是 closed source；x-axis 是 release date，y-axis 是 performance，可以看到 open source model 落后。

### [00:21:12–00:21:42]

**EN**  and why the weird ML benchmark I'm not sure if you folks actually know about this why the weird ML benchmark um it seems like the weird ML benchmark is a very good indicator better than other benchmarks and the reason the reason why is you know previously I mentioned you know previously this graph right reasoning the reasoning models are the green line and in the black models are the non-reasoning models and you can see that reasoning models double you know reduce the time of doubling time to 3.5

**中文**  为什么 WeirdML benchmark 值得注意？它似乎比其他 benchmark 更能反映模型能力。之前那张图里，reasoning model 是绿线，non-reasoning model 是黑线；reasoning 把 doubling time 从七个月缩短到 3.5

### [00:21:39–00:22:08]

**EN**  months previously it was 7 months um interestingly on the weird ML benchmark these reasoning models didn't actually do better it didn't actually change the trend um all it did was make slightly better. Um and so this weird ML benchmark seems to be more robust. Um and that is why you know this benchmark is very useful. Um in fact if you go on the Twitter bus um before GLM 5.2 got released um most you know most of the

**中文**  个月。但在 WeirdML benchmark 上，reasoning model 并未明显改变趋势，只带来略微提升。因此它看起来更 robust，也更有用。事实上，在 GLM 5.2 发布前，Twitter 上大多数

### [00:22:06–00:22:36]

**EN**  Twitter people said oh you know deepseek you know deepse if you see very if you squint okay I think I have a plot. Oh yes if you squint deepseek and Kimmy are in that little corner over there. um you know deepseek those three models the three whales are deepseek you know flash deepseek pro I think one of them's max mode or something like that um and also Kim's over there as well so before GLM 5.2 two got released, you know, on the Twitter verse, everyone kept saying that open source models are much worse than closed source models, right? They're not

**中文**  人都说 DeepSeek……如果眯眼看图，DeepSeek 和 Kimi 挤在那个小角落。三只“鲸鱼”包括 DeepSeek Flash、DeepSeek Pro，也许还有 Max mode，Kimi 也在那里。GLM 5.2 发布前，Twitter 上都说 open source 不只是落后，而是远逊于 closed source，因为这张 benchmark 图。

### [00:22:34–00:23:04]

**EN**  lagging, you know, they're not just lagging, they're much worse because of this benchmark. Um, in fact, if you look very closely of the weird ML benchmark, all of the top models are closed source labs, you know, like Fable, you know, GPD 5.5, whatever, you know, all of these are just very, you know, it shows very clear that open source models are not doing very well in terms of this benchmark. um until GPD 5.2 came along um you know number 15 is

**中文**  仔细看 WeirdML benchmark，顶部全是 closed-source lab 的模型，例如 Fable、GPT-5.5 等，清楚显示 open source 表现不佳。直到 GLM 5.2 出现——图中第 15 位是

### [00:23:02–00:23:30]

**EN**  GPD 5.2 and it shows that actually open source has came back um and GPD 5.2 too kind of shocked the world. Um that you know I guess open source has not died. Um and you know deep yeah so in general this worked very well you know deepse you know GLF2 showed that you know open source does very well still. You can also filter out by country. So by country you can see that the black

**中文**  GLM 5.2——它显示 open source 又追了回来，震动了整个世界：open source 并未死亡。总体来说，GLM 5.2 证明 open source 仍表现很好。还可以按国家过滤：黑

### [00:23:27–00:23:56]

**EN**  line is United States you know the US models. Um the dark red line is the Chinese labs. Um and you know there's other labs as well. um you know French, South Korean labs and stuff like that. Um but you know over time it shows that these models um you know the US labs seem to do very well over time. You know they they're always at the frontier and then the Chinese labs like to catch up over time. Previously you know I mentioned you know

**中文**  线是美国模型，深红线是中国实验室，还有法国、韩国等其他实验室。随时间看，美国 lab 往往处在 frontier，中国 lab 随后追赶。之前我提到 o1-preview 发布前的

### [00:23:54–00:24:23]

**EN**  the um you know the plateau before you know before 01 preview got released. If you actually look at this plot um there is something called the open source draft. Um so after 01 preview got released open source labs did not know how to replicate 01 preview. They have never you know they don't know what is reasoning. So I'm not sure if you okay this is a few years back um but on Twitter you know open kept talking about oh you know 01 preview was extremely

**中文**  plateau。图上还有所谓 open-source drought：o1-preview 发布后，open source lab 不知道怎样 replicate，也不知道 reasoning 究竟是什么。回想几年前，Twitter 每天都有人展示 o1-preview 多么

### [00:24:21–00:24:51]

**EN**  powerful um you know every single tweet you see every single day you know they show that 01 preview was very powerful. Um and so for one I think it was six months to eight months um open source models open source labs they got confused on what to do next. Um but then as everyone knows deepseek R1 came along um and they showed that even for open source models you can train these models to do reasoning gpo um reinforcement learning and it does very very well.

**中文**  强大。大约六到八个月里，open source model 和 lab 不知道下一步怎么走。后来 DeepSeek-R1 出现，证明 open source model 也可以通过 GRPO、reinforcement learning 训练 reasoning，而且效果非常好。

### [00:24:49–00:25:18]

**EN**  In fact, if you take this plot, you know, the the black line minus the blue line, if you just minus it, you get this plot. Um, and you can see this is how many months behind open source is. Um, and you know, over time, um, you can see like, you know, after 01 preview got released, um, you know, it kind of skyrocketed. Um, you know, the open source models were very, very lagging in terms of, you know, behind closed source models. Um and so like when Deepseek R1

**中文**  把黑线减去蓝线，就得到 open source 落后多少个月。o1-preview 发布后，差距急剧扩大，open source model 明显落后 closed source。等 DeepSeek-R1

### [00:25:16–00:25:45]

**EN**  got released then the open source labs knew okay we can also do uh 01 type reasoning. Um and that is why recently you know the the time between closed source labs and open source labs have started decreasing again. Um yeah so this is slightly outdated. This is like May. Um so I think now it's actually four months with the release of GLM 5.2. It's around four months now. Um so open source labs lag behind closed

**中文**  发布，open source lab 才知道自己也能做 o1 型 reasoning。因此最近 closed 与 open source 之间的时间差又在缩小。这张图略旧，数据截至 5 月；随着 GLM 5.2 发布，现在差距大约四个月。open source lab 落后 closed

### [00:25:42–00:26:11]

**EN**  source labs by around four months. There was actually a very nice plot you know doing some sort of regression. So some sort of like trend extrapolation. Um according to this plot um if you extrapolate the trend by December this year open source models will 100% catch up to closed source models by this year December. Um but you know who knows I guess maybe maybe open source maybe we can have an open source model as

**中文**  source lab 约四个月。有张很好的图做了 regression 和 trend extrapolation：如果照趋势外推，到今年 12 月，open source model 会 100% 追上 closed source model。谁知道呢，也许 12 月真会有

### [00:26:08–00:26:36]

**EN**  powerful as the best closed source model by December um you know if this trend continues um so I guess the question is will the trend continue um it's always about will the trend continue um and you know may most of you maybe may know that you know open source some of the open source improvements in technology you know improvements in capabilities are via distillation you know so some of the open source labs what they like to do is they like to

**中文**  与最佳 closed source model 一样强的 open source model——前提仍是趋势延续。问题总是趋势会不会继续。大家可能知道，open source 在技术和能力上的部分改进来自 distillation：一些 open source lab 会

### [00:26:35–00:27:04]

**EN**  call the models you know core the frontier models like opus or GPD and then use the traces to train your model um so this is a common methodology that labs like to do um I wouldn't say this is a bad method um but it is a method that you know some closed source labs like to look down upon you know they like to stop you know their view is you know we should not allow these open source labs to like do this training um and you know get away for free I guess in terms of training cost. Um but you

**中文**  调用 Opus 或 GPT 等 frontier model，再用 reasoning trace 训练自己的模型。这是常见 methodology，我不认为它本身很坏；但一些 closed source lab 看不起这种做法，认为不应允许 open source lab 这样训练、免费省下 training cost。不过

### [00:27:02–00:27:31]

**EN**  don't actually have to do this approach. Um so most labs when you do distillation there are two different types of approaches. Um the first approach is you need to have the logits. You need to actually have access to the full logs. Um and unfortunately most labs do not actually have that right. So like labs will not give you the full logs. Um instead you only get the reasoning traces that are summarized um and the final output. Um, and so these, you know, these open source labs, they're

**中文**  其实不一定要用这种 approach。多数 lab 做 distillation 时有两种方式。第一种需要 logits，也就是完整 log；但大多数 lab 拿不到，模型提供商不会交出完整 log。你通常只能拿到被 summarized 的 reasoning trace 与 final output。因此 open source lab

### [00:27:29–00:27:55]

**EN**  not just, you know, they're not just training on the, you know, Opus output, right? That's just, that's silly. What they do is they use GRPO or reinforcement learning to recreate the traces. Um, and so because you you have the final output, which is the answer. All you need to do is use GPU and RL to create the reasoning trace automatically. Um, and so that's kind of how they train these models. Um, and so you don't actually need to like access

**中文**  并非简单拿 Opus output 直接训练，那太粗糙了。他们会用 GRPO 或 reinforcement learning 重建 trace。既然已有 final output，也就是答案，只需用 GRPO 与 RL 自动创建 reasoning trace。这就是训练方式，不必访问

### [00:27:53–00:28:22]

**EN**  the logits or the weights of the model. um that's not necessary. Um yeah, and you know, one of the most important factors of you know, these large models is, you know, as models get bigger and bigger, you can't run them on your local device anymore. Um it's extremely complicated to run. Um and so we do something called dynamic quantization where essentially you take a model, you quantize them down to one bit. Um but the trick is you don't

**中文**  logits 或 weights。还有一个关键因素：模型越来越大后，就无法在本地设备运行，部署极其复杂。我们做 dynamic quantization，本质上把模型量化到 1-bit；但诀窍是不能

### [00:28:20–00:28:47]

**EN**  quantize every single layer to one bit. you quantize some important layers to 16 bit or 8 bit or something like that. Um and so if you quantize the whole model down to one bit you will get 0% accuracy right 0%. Um but the trick is if you do dynamic quantization so if you look on the you know this is a three-bit Deepseek model um a three-bit one you get 75.6% 6% accuracy, a three-bit one.

**中文**  把每一层都量化到 1-bit，而要把某些重要层保留在 16-bit 或 8-bit。如果整个模型都降到 1-bit，accuracy 会变成 0%。但 dynamic quantization 会选择性处理。例如一个 3-bit DeepSeek model 能达到 75.6% accuracy。

### [00:28:45–00:29:14]

**EN**  In fact, if you do dynamic one bit, um you get 57% accuracy. Um so we show that you know if you do something called dynamic quantization where you quantize the model down smartly, you can recover accuracy. Um and this methodology will become even more important when models get larger and larger and larger and larger. if you plot the paro you know efficiency um there if you don't do dynamic if you do you know some other dynamic

**中文**  dynamic 1-bit 甚至能有 57% accuracy。我们证明，聪明地量化模型可以恢复 accuracy。随着模型越来越大，这种 methodology 会更重要。画出 Pareto efficiency 后，其他 quantization method 表现尚可，但

### [00:29:11–00:29:38]

**EN**  quantization methods it does okay um but we showed that if you smartly choose the layers it does even better um I'm not sure if you folks have followed but GLM 5.2 we also released dynamic quantizations for that we showed that GLM 5.2 2 can quantize very well. So if you look I think this is oh this is an animation. Oh it works. Um but yes you can show the animation you you can see the animation a one bit GLM 5.2

**中文**  我们的方法通过聪明地选择 layer，效果更好。对于 GLM 5.2，我们也发布了 dynamic quantization，并证明它非常适合量化。这里有一个动画：这是 1-bit GLM 5.2

### [00:29:35–00:30:04]

**EN**  model and this is one bit um and the one bit model is literally 86% smaller. Um so it's 86% smaller than the full 1.5 terabytes. Um and it still managed to do very well on one of the prompts. Um so it shows that the models are not dumb. Right? If you make the model 86% smaller, it does not get 86% dumber. Um, it only gets 14% less dumb. Um, and so

**中文**  model，体积比完整的 1.5 TB 小 86%，却仍能很好地完成其中一个 prompt。这说明模型不会因为缩小 86% 就笨 86%，只损失约 14% 能力。

### [00:30:02–00:30:31]

**EN**  it shows that, you know, if you do special tricks to compress the model, the model still works very well. Um, and we also compared to Opus, you know, we compared to Opus 4.8, we compared to GBD 5.5. And also, you have to notice that for G 5.2, I use high reasoning mode. You know, for Opus, it's extra high. And for you know GPD 5.5 is also extra high. Um and so like you know there are different reasoning modes as well which

**中文**  用特殊技巧压缩后，模型仍能很好工作。我们还把它与 Opus 4.8、GPT-5.5 比较。注意 GLM 5.2 使用 high reasoning mode，而 Opus 与 GPT-5.5 使用 extra-high。不同 reasoning mode 也会影响

### [00:30:29–00:30:58]

**EN**  we can also um see. Um and all of these are oneshot um so we do not like prompt the model like you know 50 times or something. Um this is just one shot directly. Okay so the next I guess the open source versus close source section is done. I guess any other questions? Yes. Um so the question was which parts of the model do we quantize to lower bits versus higher precision. Um so in general um we did actually a lot of

**中文**  结果。而且这些都是 one-shot，没有 prompt 五十次。open source 与 closed source 这一节结束。还有问题吗？观众问：模型哪些部分降到 low-bit，哪些保留 high precision？我们对此做过大量

### [00:30:57–00:31:26]

**EN**  research on this. So if you look at the quen the quen 3.5 architecture there are some layers which is the linear attention layers. Um the linear attention layers should never be quantized. If you quantize the linear attention layers down you will definitely suffer in long context. Um so in general the linear attention layers need to be left in 8 bit or 16 bit. Um that's for example um another like if you look at the model layers um some

**中文**  研究。以 Qwen 3.5 architecture 为例，linear-attention layer 绝不能量化；降下来会严重损害 long context。因此通常要保留为 8-bit 或 16-bit。另一方面，查看模型各层时，有些

### [00:31:24–00:31:53]

**EN**  layers can be quantized down heavily to one bit. Um and the reason why is because these layers are kind of like filler layers. Um and so they don't actually do anything. Um and in order to check whether a layer does something or not, you do need some sort of collaboration data set. So you need you need to have some sort of like representative data and pass it into the model um and you can get you can get the um outputs after each layer and then you

**中文**  layer 可以大幅降到 1-bit，因为它们像 filler layer，实际上没做什么。要判断一层是否重要，需要 calibration dataset：准备有代表性的数据输入模型，取得每层之后的 output，观察这一层是否带来明显变化。

### [00:31:51–00:32:19]

**EN**  can see okay does this model at this specific layer you know does it change that much um and if it doesn't change that much okay maybe just quantize a layer to one bed um but if it does change dramatically then you need to be careful um you you cannot quantize that down to like one bed or whatever um so there are actually many we actually publish a lot of like blog research on this. Um we show I think there was we also show for example you cannot quantize the vision layers down.

**中文**  若变化不大，也许可量化到 1-bit；若变化剧烈，就必须小心。我们发表过许多 blog 和 research，例如证明 vision layer 不能量化。若把 vision layer 降下来，模型看到 train 图片可能会说那是 beach。

### [00:32:17–00:32:46]

**EN**  If you quantize the vision layers down you will make the model really bad. Um if you give it a you know if you give it a picture of a train it will say it looks like a beach for example. Um and so it's you should never quantize the vision layers, the audio layers. Um and only the language the language model layers you can like quantize. Um but there are many tricks in order to do that. Um yeah >> correct. So the question was if you do

**中文**  所以不要量化 vision layer 和 audio layer，主要量化 language-model layer，其中也有很多技巧。观众：做 distillation 可能损害其他 topic；如果只训练 coding，coding 变好，其他能力会变笨。嘉宾：这是合理担忧。

### [00:32:43–00:33:13]

**EN**  distillation um you you might have done worse on other topics um but you know only if you for example if you just do coding it will just do good encoding and then the rest gets very dumb. Um so that's a fair point. Um so I think that the main trick is you will need to do many many many examples. You will call the model like you know 10 million times. Um and so like the trick is once you call the middle model 10 million times with high diversity of questions

**中文**  主要诀窍是需要极多 example，例如调用模型一千万次，并让问题保持 high diversity。借用 pre-training 的逻辑，模型通常就能兼顾其他 task。pre-training 效果好，正因为学过很多 task，能 interpolate 缺失部分。例如只用 math question 训练模型，

### [00:33:10–00:33:38]

**EN**  in general um by using the pre-training argument um the model will do well on other tasks. Um so the reason why pre-training does very well um is because it has learned so many tasks that it can interpolate the missing holes. Um for example, if you just train if you just pre-train a model with just maths questions um assume you do only maths. Okay, maybe it's not going to do very well, right? And it's not going to do very well on every other task. But

**中文**  假设只做 math，它不会在其他任务上表现好。pre-training 的诀窍是覆盖 math、coding、law 等几乎所有能想象的 topic。知识足够丰富后，就能填补未知部分。distillation 也要采用相同 approach，做好 sampling。比如不必使用十万亿 token，

### [00:33:36–00:34:05]

**EN**  the trick of pre-training is it does maths, coding, law, you know, every single imag, you know, every single topic you can imagine. And the trick is because it has so much knowledge, it feels the holes of the things that it doesn't know. Um, and so for distillation, you also need to do the same approach. You need to sample, you need to sample well. Um so for example instead of doing 10 trillion tokens sample you know like 1% um and then call

**中文**  可以 sample 其中 1%，再调用模型。实验室大致就是这样做。另一个好问题：与其一次大规模 quantization，能否直接 prune model、删除整个 layer？我们的研究表明 pruning 确实有效，但有个很大问题：必须 retrain model。删掉整个

### [00:34:03–00:34:32]

**EN**  the model. Um yeah so that's kind of how the labs are doing that. Um that is a very good question. So instead of doing one big quantization can you instead prune the model like you know delete some layers entirely. Um so in general from our research pruning does work. There is a very big problem though. You need to retrain the model. you need to continuously train a model after pruning because you have deleted an entire layer. Um and so if you delete an entire

**中文**  layer 后，需要 continuous training。因为整层被删，必须做 QAT 或进一步 fine-tuning，把更多知识压进其余 weight。相比之下，dynamic quantization 属于 post-training quantization，简称 PTQ，

### [00:34:30–00:34:58]

**EN**  layer, you will need to do like you know qat or further fine-tuning to push the other to push the other weights to have more knowledge. So that is the only problem where if you delete layers um if you don't delete layers when you do you know dynamic dynamic quantization it's called post training quantization. So PTQ you do not need to do any training at all. if you do you know quantization um but if you do prove the layers you do

**中文**  不删除 layer，完全不需要 training；pruning 则必须训练，这是主要问题。另一个问题：open source lab 依赖 closed source model，因此 gap 永远不会归零。我部分同意。最容易的起步方式确实是 distillation，但如果你是 open source lab，只会用它

### [00:34:56–00:35:25]

**EN**  need to train um so that is one of the problems um yeah yes that's a great question so the question was you know because open source labs you know they use closed source models the gap will never actually go to zero um and so I partially agree and so the main argument was labs open source labs the easiest way is to do distillation however you know if you for example if you were an open source app you will only use that

**中文**  首先进入市场；从长期安全角度，不会永远依赖这种 approach。你会自行生成答案、获得 question，例如从 Scale AI 等处拿 data，建立大型 data-labeling 团队。事实上当前一些 lab 并不只做 distillation，

### [00:35:23–00:35:52]

**EN**  approach to firstly enter the market but as longterm you know as long-term safety as a long-term safety net you will not do this approach um instead as a you know instead you will do you know for example generate the answer get the question for example you know you will get data from a call or scale whatever you know have some sort of like large data labeling army or something I don't know um and so like in general because currently some of the labs so they don't just do distillation right so they're

**中文**  而是把它作为市场进入手段；长期则需要自己的数据与训练体系。

### [00:35:50–00:36:19]

**EN**  not just going to call the model 10 trillion times you know and just do distillation. They also augment the training data with their own approach. So I will be talking about the GLM approach maybe like later um but they did invent some new approaches to do very good reinforcement learning and GRPO um and because GRPO and reinforcement learning um you know is open source these labs just use these methodologies to make the models better. So you don't so distillation is only one

**中文**  不会把同一模型调用十万亿次、只靠 distillation，也会用自己的方法 augment training data。后面也许会讲 GLM approach；他们确实发明了新的方法，把 reinforcement learning 和 GRPO 做得很好。GRPO 与 reinforcement learning 都是 open source，实验室可直接使用这些 methodology 提升模型。因此 distillation 只是整个

### [00:36:17–00:36:46]

**EN**  part of the training system. Um and it's not I would say that assume distillation disappeared. Okay, maybe open source labs maybe increase you know it's not four four months maybe eight months um but but that's fine because you know we always have some sort of innovative and new approach you know deepseek might invent something new um and so like you know GLM ki all of them Google you know even the American open source lab they'll have some new innovation um and

**中文**  training system 的一部分。假设 distillation 消失，open source lab 与 closed source 的差距也许会从四个月扩大到八个月；但也没关系，因为总会有新的创新 approach。DeepSeek 可能发明新东西，GLM、Kimi、Google，甚至美国 open source lab 都会有创新。

### [00:36:44–00:37:13]

**EN**  so like I think like yes if you stop distillation it will increase you know the you know four months to eight months but I still think that is fine Um it's just a delay you know and then the delay will go back to like four months. Um yeah good question. So the question is if dynamic quantization is always better why do people not always do dynamic quantization? Um so it depends on the definition of dynamic

**中文**  所以停止 distillation 也许把差距从四个月变成八个月，但只是 delay，之后又会缩回四个月。另一个好问题：如果 dynamic quantization 总是更好，为什么不人人都做？这取决于 dynamic

### [00:37:10–00:37:40]

**EN**  quantization. So for every single lab they will have different approaches to dynamic quantization. In fact, I'm actually going to talk about that. Um, I was going to talk about that in the benchmaxing and accuracy minimizing session. So, I'll be talking about that. Um, so I will your question will be answered later. Um, yes. Okay. One more question. Yes. Yes. So, the question was for consumer grade GPUs, you know, what are the open source models in terms of like, you know, the parameter size

**中文**  quantization 的定义。每个 lab 都会有不同 approach。后面的 throughput maximizing 与 accuracy minimizing 部分正好会讲，届时会回答。再来一个问题。观众问：对 consumer-grade GPU 来说，open source model 的 parameter size 与

### [00:37:38–00:38:07]

**EN**  capabilities and stuff like that. Um, so for the open source community, you know, the most popular models are probably Quen 3.6 6 35 billion um 27 billion GMA you know Gemma's 26 billion um GLM 4.7 flash the smallest type models um and I feel like these small models are actually very powerful um so okay I don't have wait I don't think so I have a plot um but essentially these small models the biggest problem oh actually

**中文**  capability 如何？open source community 最流行的大概是 Qwen 3.5 35B、27B、Gemma 26B、GLM 4.7 Flash 等较小模型。我认为它们其实很强，但最大问题是……我后面也会讲。

### [00:38:06–00:38:35]

**EN**  I'm going to talk about this as well the biggest problem of these small models are they fail very bad at tool calling because they have tool calling issues um they loop continuously um and the biggest problem is because they're small and that is why they have these problems um but we can counteract this um and so one of the things I'm going to talk about later is the model becomes not important anymore it's the harness or the tool that is actually the most important thing um how do you actually

**中文**  这些 small model 在 tool calling 上失败得很严重，会不断 loop，根本原因就是模型太小。但可以缓解。稍后会谈到：model 本身将不再最重要，harness 或 tool 才最重要。怎样调用模型，对 accuracy 的影响最大。

### [00:38:32–00:39:02]

**EN**  call the model um that actually affects the most accuracy of the model um so not actually the model itself um but I'll be talking about that as well um yeah okay I will continue on. Um, there were always questions after each section. Um, uh, yes. Oh, yes. The next section, the fun section, throughput maxing. Oh, actually, I think I did. It's supposed to be 2x. I don't know. Whatever. Throughput maxing and and accuracy minimizing. I thought it was like

**中文**  真正关键的不是模型本身。下面继续，每节后都能提问。下一节是有趣的 throughput maximizing——我原本似乎写成 2x，算了——以及 accuracy minimizing。我原以为是 accuracy mining，但没有这个说法，所以暂时叫 accuracy minimizing。

### [00:39:00–00:39:30]

**EN**  accuracy mining, but there's no such thing. So, it's called accuracy minimizing for now. Um, yes. So, this part actually I really like. Okay. I'm not sure if you guys can see it's a bit oh whatever. Um this shows the parade efficiency of cost of the model. So cost is um cost is the x-axis. Um and the y-axis is the arena score. So this is like you know

**中文**  我很喜欢这一部分。这里展示模型 cost 的 Pareto efficiency：x-axis 是 cost，y-axis 是 Arena score，也就是 LM Arena 的分数。

### [00:39:26–00:39:54]

**EN**  Ella Marina's arena score. Um and this part I really like. So I don't really you know maybe you see like arena scores you know Ella Marina scores between each model. I don't really like that. It's not it's not very easy to see. Instead the better approach is to plot every single model on two axes cost versus accuracy. Um and you can see Fable does very well right so Fable does very very well on that plot. Um but you can see

**中文**  我不太喜欢只看每个模型的 LM Arena score，因为不容易理解。更好的办法是把每个模型画在 cost 与 accuracy 两个轴上。可以看到 Fable 在这张图上表现很好，但也有

### [00:39:52–00:40:20]

**EN**  there is a paro trend you know like Gemini 3.1 preview is over here you know Opus 4.6 is over there as well. There are some other models as well when Fable got released. Okay. Well, now it's banned, but anyways, when Fable was released, when you know when people tried it, they noticed that it's not that much better in terms of actual capabilities. Um, you can see, you can

**中文**  Pareto frontier，Gemini 3.1 Preview 和 Opus 4.6 都在上面。Fable 发布后——现在被禁了，但当时大家使用时发现，它的总体 capability 并没有好太多。不过

### [00:40:18–00:40:47]

**EN**  see, but however, people really liked the front-end design. You know, they said if you called Fable, it was very, very good for UI UX front end. Um and in fact if you look at the LM Mariners chart you can see it it was a very big shift in terms of front-end design. Um GLM 52 is also there if you can see you know it was part of the par paro trend. Um but in general for these large models they seem to have they're not going to

**中文**  人们非常喜欢它的 frontend design，认为 Fable 很擅长 UI/UX frontend。LM Arena chart 也显示 frontend design 有大幅提升。GLM 5.2 同样位于 Pareto frontier。但整体而言，这些 large model 在

### [00:40:46–00:41:12]

**EN**  be do they're not going to be doing that much better on general tasks. Um however for UI and designing Fable seems to have done very very well. Um, and so you should use Fable for your designing. You should use Fable for designing, for UI, for UX, whatever, HTML, JavaScript, but you should probably not use Fable for the rest of the tasks because it is very expensive. Um, so you know, use

**中文**  general task 上不会好很多；Fable 在 UI 和 design 上则非常强。所以设计、UI、UX、HTML、JavaScript 可以用 Fable，但其他任务大概不应使用，因为它非常昂贵，可以换

### [00:41:10–00:41:38]

**EN**  some other models instead. Um, and you know, however, yes, okay, you know, some other models, you know, like, okay, this, you know, this shows that Fable does very well on UI and UX. Um but how about over time um you know how what do you know anthropic their view is we need to maximize throughput right maximize throughput but also maximize accuracy um you know they want

**中文**  其他模型。这说明 Fable 在 UI/UX 上很好。再看时间趋势：Anthropic 想同时 maximize throughput 和 accuracy，为更多人提供服务，但有时并不奏效，反而会降低 accuracy。

### [00:41:36–00:42:05]

**EN**  to like you know serve more people um but sometimes it doesn't actually work um sometimes they actually reduce accuracy um and so you can see there is a I don't know if you folks know margin labs um they have this very cool they do su bench they benchmark codecs and clawed code with the models. Um, and this is accuracy over time for these models. Um, and the dotted lines are the release of

**中文**  不知道大家是否了解 Margin Labs，他们会用 SWE-bench 对 Codex、Claude Code 及其模型做 benchmark。这里是模型 accuracy 随时间变化，dotted line 表示新模型 release。

### [00:42:03–00:42:31]

**EN**  the new models. Um, so there's actually another um, there was actually a dip in Wait, can you is there Oh, okay. The mouse is there. Um, I I think it was over here. Um, I think it was over here that feeble got released. Um, so there was actually another dotted line. Um, there was actually very interesting trends you can see. The first one is um every single time there is a new model release this this you know daily tracker seems

**中文**  Fable 大概是在这里发布的，原图应还有一条 dotted line。有几个有趣趋势：每次新模型发布，这个 daily tracker 的 accuracy 似乎都会下降。

### [00:42:28–00:42:57]

**EN**  to decrease in accuracy. Um and so if you want to predict when a model gets released from anthropic you can use this as an indicator um of when the model gets released. it worked very very well. Right? So like essentially if you were over here the dipped in accuracy over a very long period of time was because fable got released. Um and over here I think that's opus 4.8 I think. Um I think yeah I think that's opus 4.8. This

**中文**  如果想预测 Anthropic 何时发布模型，可以拿它当 indicator，过去很有效。这里长时间的 accuracy dip 是因为 Fable 发布；那里可能是 Opus 4.8。

### [00:42:53–00:43:23]

**EN**  is opus 4. Uh 7 and so on. That's 4.6 I think. I whatever um I don't remember exactly but um but you can also see that there is ginormous dips of accuracy. Um, and it's not just like one day or two days, it's for a very long period of time. This is also Codex. Um, so they also do codeex benchmarks. Um, and you can also see that over time. I don't know if you can squint, but you can see that

**中文**  其他 dotted line 可能是 Opus 4.7、4.6，我记不清。图中有非常大的 accuracy dip，而且不只持续一两天，而是很长时间。Codex 的 benchmark 也一样。若观察趋势，

### [00:43:20–00:43:48]

**EN**  actually Codeex has been getting worse if you plot the trend, right? If you can I don't know if you can squint, but if you draw a line, it seems to be getting worse. Um, so I'm assuming OpenAI is investigating this as well. Um, okay. >> So this is different model. This is codeex. >> So this is using 5.5. >> This is using >> model. >> Correct. It's the same. So what this

**中文**  Codex 似乎越来越差，我猜 OpenAI 也在调查。观众：这里是不同模型，Codex 使用 GPT-5.5？嘉宾：是同一个 harness，但模型会变化。这个

### [00:43:46–00:44:15]

**EN**  benchmark does is you randomly sample 50bench questions. Sweet bench is very large. So you just sample 50 of them and then you call the model to answer it. Um and then you record accuracy. Um and so obviously you know every single day there's like you know daily variations. Oh, it's not it's not that useful because you're only calling 50 questions. Um, so the trick is to look at the trend. Um, and the trend uh h maybe open should investigate this. Um,

**中文**  benchmark 会从很大的 SWE-bench 随机抽 50 道题，调用模型回答并记录 accuracy。因为每天只有 50 题，会有 daily variation，单日数据并不太有用，关键要看 trend。OpenAI 或许应该调查这个趋势。

### [00:44:13–00:44:42]

**EN**  and you can see the trend for you know anthropic is also not very good. Um, in general oh so sorry this is not the same model. Um, these models change my bad. Um, so it's the same harness but the model changes. Um, so this dotted line is GBD 5.5. Um, so everything over here is GBD 5.5. Everything over here is GBD 5.4. Um I think this is 5.3. Um and so on. Um but it seems like the model's

**中文**  Anthropic 的趋势也不太好。抱歉，这并非同一个 model；harness 相同，但模型在变。dotted line 之后是 GPT-5.5，之前是 GPT-5.4，再前面可能是 5.3。总体似乎在变差。

### [00:44:39–00:45:09]

**EN**  getting worse. So I don't know. This is probably just on this benchmark, right? On the Sweet Bench Pro benchmark, it's getting worse. Um but you know, I wouldn't really trust these benchmarks. The best way is to look at the degradation. You know, the sudden drops. You know, for example, Codex dramatically dropped over here. I don't know why. Um and you know clawed you know clawed code was very bad for a few

**中文**  也许只是在 SWE-bench Pro 这个 benchmark 上变差，我不会完全相信 benchmark。最好观察 degradation 和 sudden drop，例如 Codex 在这里突然大跌；Claude Code 在某几周也非常差。

### [00:45:05–00:45:33]

**EN**  weeks over here or over here right. Okay. Yes. >> Yes there is a confidence interval. I did not plot it but this is 50 tasks. So every single day they call a 50 tasks randomly. So they will sample 50 tasks. Um and so you you should not look at this daily. This is daily. So every single day is 50 questions. Another 50 questions and so on. Instead, you should do like a

**中文**  观众问 confidence interval。确实有，只是我没画。每天随机 sample 50 个 task，所以不要看每日点位，应该计算

### [00:45:31–00:46:00]

**EN**  rolling average, you know, some sort of rolling, you know, 7-day average. Um, that's a better number. Yeah, >> really. I can see it from here. It's like decreasing. >> It's it's >> if you look at the if you do the seven moving average, I I'll probably get the pot later. It actually is decreasing.

**中文**  rolling average，比如 7-day average，那会更可靠。观众：我从这里看就是在下降。嘉宾：若计算 seven-day moving average，确实在下降，我之后可能会找到那张图。

### [00:45:57–00:46:27]

**EN**  You can see it um if you can see I don't know if you look at the top peaks of the you look at the top peaks and the peaks are decreasing. >> Okay, how about the bottom peaks? >> Okay, I agree there is random noise. So the trick is you need to the moving average and if you look at the moving average you can actually see it's decreasing. I I'll probably I'll get the

**中文**  看每次顶部 peak，peak 在下降。观众：那底部 peak 呢？嘉宾：我同意存在 random noise，诀窍仍是看 moving average；它确实在下降。我会之后找图。

### [00:46:25–00:46:52]

**EN**  plot later. Um you can you can search it. It's so go to Margin Labs, search in Margin Labs codeex claw code benchmarks and they do show weekly the weekly trend but I'm just saying this is not this is not to say that the model is getting worse. This is just to show that accuracy that you know the sudden dips the accuracy of these models can decrease. Um and the question is why you

**中文**  可以搜索 Margin Labs 的 Codex、Claude Code benchmark，他们展示 weekly trend。但我的意思不是模型一定在变差，而是说明模型 accuracy 会出现 sudden dip。问题是为什么。

### [00:46:50–00:47:19]

**EN**  know for example why did claude code over a few weeks why did the performance decrease like why that's the fundamental questions. So that is one theory. A theory is they might have accident you know that before the model release they are doing testing and so they might have like you know act you know some of the some of the queries

**中文**  例如 Claude Code 为什么会连续几周 performance 下降？一种 theory 是模型发布前在做 testing，把部分 query route 到

### [00:47:16–00:47:44]

**EN**  they route to opus 4.8 eight >> or fable or whatever. And the problem is they did not. So the main question is if you do route to another model, why did the accuracy decrease? It should actually get better. And so one of the theories is theory one, they forgot to edit the system prompt. And so the system prompt for Fable was different, but then they used the wrong system prompt for, you know, for Opus 4.8, and that is why the

**中文**  Opus 4.8 或 Fable。但如果 route 到新模型，为什么 accuracy 反而下降？本来应变好。一种解释是忘了修改 system prompt：Fable 的 system prompt 不同，却误用于 Opus 4.8，于是

### [00:47:42–00:48:12]

**EN**  accuracy decreased. Um and then after the model got released the accuracy went back up because they used the correct system prompt. That is one theory. Um the other theory is the other theory is okay we're actually going to talk about this is it's actually they're doing tricks. You know they did quantization but they didn't do dynamic quantization. They did some dumb quantization. Um you know they some GPUs are broken for example. You know they use the wrong GPUs. Some of them have like you know

**中文**  accuracy 下降；正式发布后换成正确 system prompt，accuracy 又回升。另一种 theory 是他们用了某些 trick，例如做了不够聪明的 quantization，而非 dynamic quantization；也可能 GPU 有问题、用了错误 GPU、发生 bit flip，或新 data center 恰好 accuracy 较低。

### [00:48:10–00:48:38]

**EN**  bit flips or something. I don't know. um they have like a new data center and then that data center just by chance has lower accuracy. Um in fact there is actually okay I'm going to talk about this actually. Um yeah but there are many many theories like you know possibilities why this could reduce an accuracy. Um actually I think it's the next plot. Yes the next plot. Um oh well the next slides. Um so

**中文**  关于 data center 我后面会讲。总之可能性很多。接下来正好有相关 slide。几个月前，AMD 的某个人在 Claude Code issue 中提问，当时正值一次很大的

### [00:48:35–00:49:05]

**EN**  actually when was this? I don't remember. Um it was a few months ago. Someone from AMD actually made an issue on chord code you know during this dip. I think it was during the before um a very large dips in accuracy and they actually asked Claude, you know, they asked the Claude team, why is there a noticeable dip in accuracy? You know, why why is that? And Claude actually wrote a in April 23,

**中文**  accuracy dip：为什么 accuracy 明显下降？Claude 团队在 4 月 23 日解释了原因，并做了 postmortem。

### [00:49:02–00:49:31]

**EN**  they actually provided details on why they had reduced in accuracy, right? So they did a postmortem on what happened with Claude. Um and the reason why is because the thinking trace got deleted after the second you know when you when you ask call the second time the thinking trace got deleted. Um and it had a bad system prompt. Um and they found out that that that was why the accuracy got deleted. Um so somehow in

**中文**  原因是第二次提问时，thinking trace 被删除，而且 system prompt 不好。他们发现这就是 accuracy 下降原因：Claude Code 第二轮 question 会抹掉上一轮 thinking trace。

### [00:49:29–00:49:58]

**EN**  claude code the second time you ask a question the previous thinking trace got erased. Um and I don't know I don't even know how they did not find this but oh wow um according to them now is claude now has this internal benchmark so they will use more internal investigations to test okay next time if there's a new model this won't happen ever again um and you know like these things do happen over time um and so like for this specific

**中文**  我不知道他们怎么没发现。现在据说 Claude 有 internal benchmark，会做更多内部调查，确保以后发布新模型不再发生。这些问题确实会随时间出现。在这个具体

### [00:49:55–00:50:18]

**EN**  example you know cloud code the harness itself was the problem not the actual model right the harness the thinking trace got deleted and they had a very not a very good system prompt. Um and that is why the accuracy actually degra um degraded. So that okay so we found one answer why these models got worse.

**中文**  例子里，问题是 Claude Code harness，而非 model：thinking trace 被删，加上 system prompt 不佳，导致 accuracy degradation。于是我们找到了模型变差的一个原因。

### [00:50:19–00:50:48]

**EN**  They also released in September 2025 right in September 2025 they showed that it was due to okay I didn't okay I didn't put the slide but anyways they showed it was actually due to a hardware problem. Um so in their compiler um they used TPUs. So so Anthropic likes to use TPUs and GPUs. Um they showed that the same software stack for GPUs and TPUs um actually produced different results. Um

**中文**  他们还在 2025 年 9 月说明，另一次下降来自 hardware problem。我没有放那张 slide。Anthropic 使用 TPU 与 GPU，同一套 software stack 在二者上产生了不同结果。

### [00:50:45–00:51:13]

**EN**  and so for the TPUs it actually was different sampling. Um and for the GPUs it was a different sampling mechanism. Um and so that is actually why they had decrease in accuracy during September sometime. Um because they actually had different hardware. And so you need to like Yeah. So like once you have different hardware accuracy also changes. So I think the main point is the harness

**中文**  TPU 与 GPU 使用不同 sampling mechanism，这就是 9 月某次 accuracy 下降的原因。不同 hardware 也会改变 accuracy。因此核心观点是，harness、

### [00:51:10–00:51:38]

**EN**  the implementation the tool is now the most important. It's not the model right the model is useless. Most models you know if you look at the model of you know open source versus closed source models are generally the same. The difference is how clawed code is made you know how codec is made and used. Um and so that is actually the most important factor. It's not the model anymore. Um and so like you know as we

**中文**  implementation 和 tool 现在最重要，而不是 model。模型本身没有独立价值；open source 与 closed source 的基础模型通常很相似，差别在于 Claude Code 或 Codex 怎样被构建和使用。

### [00:51:36–00:52:06]

**EN**  have seen if they have accidentally botched you know if they accidentally botched the harness you will get reduced accuracy. Um and so like you know definitely you know for large labs as you know I'm sure they are they know these problems and they're working on it. Um but I feel like you know these are probably you know these are still very hard to fix. Um yeah so hopefully I answered some of people's questions on the harness you know the accuracy um so there's actually reasons why accuracy got degraded

**中文**  如果不小心破坏 harness，accuracy 就会下降。大型实验室当然知道并在处理，但这些问题仍很难修。希望这回答了关于 harness 与 accuracy 的问题：accuracy degradation 确有具体原因。

### [00:52:06–00:52:33]

**EN**  but you know it's not just closource labs during bad across open-source model providers the accuracy changes so if you look at this plot so this is from open router um this is deepseek v4 pro um so most labs what they want to do most inference providers what they want to do is they want to serve you the highest throughput, right, with the cheapest price. They want to

**中文**  不只是 closed-source lab；不同 open-source model provider 的 accuracy 也会变化。这张 OpenRouter 图展示 DeepSeek V4 Pro。多数 inference provider 想以最低价格提供最高 throughput，

### [00:52:30–00:52:59]

**EN**  give you, you know, 60 tokens, 120 tokens, 1,000 tokens per second, right? They want to give you the fastest. Um, but did you but did people actually bother to check accuracy? So, that is the fundamental question. You know, you might be getting 10,000 tokens per second and there is no model. Um, so the main question is you need to be careful of what you use from these inference providers. Um and so for DeepS v4 you

**中文**  给你每秒 60、120、1,000 token，越快越好。但有人检查 accuracy 吗？也许得到每秒一万 token，却根本没有真正的 model capability。因此选择 inference provider 必须谨慎。对于 DeepSeek V4，

### [00:52:57–00:53:27]

**EN**  know there are two benchmarks which open router ran you know yeah it's like sorted it's sorted I think on the gray I think it's sorted on tower bench um so it's sorted on tower bench um and the green one is GPQA and you can see that in general some of the labs are not you know some of the sorry not labs some of the inference providers are not doing very well um so you need to like before you before you use a open source model please check the accuracy before you use the open source model. Um

**中文**  OpenRouter 跑了两个 benchmark：似乎按 TauBench 排序，绿色是 GPQA。可以看到一些 inference provider 表现不佳。所以使用 open source model 前，请先检查 provider 的 accuracy。

### [00:53:25–00:53:55]

**EN**  and also one of the biggest problems of this is every single time you know for example a cloud code you know claude code and codeex you can benchmark accuracy over time. Um and the good thing about closed source labs is they control the supply chain. Um the biggest problem of open source is there are so many suppliers and providers of these models um that sometimes what happens is people get turned off and they get very

**中文**  另一大问题是，Claude Code 和 Codex 这种 closed source service 可以持续 benchmark accuracy；closed source lab 的优势是控制整个 supply chain。open source 最大的问题则是 model supplier 和 provider 太多，有时用户体验很差，

### [00:53:53–00:54:23]

**EN**  annoyed that the open source models do not work very well. Um so everyone you know in the in the ecosystem people keep saying that closed source labs do much better than open source but it's not because of the model it's because of the inference provider right the inference provider is to blame that they are causing the downfall of open source because they're giving a bad name for open source um so I would like check you know whatever favorite inference provider you

**中文**  因此恼火，觉得 open source model 不好用。ecosystem 里都说 closed source 胜过 open source，但原因不在 model，而在 inference provider。provider 的低质量让 open source 背负坏名声，某种程度上正在拖累 open source。请检查你喜欢的 inference provider。

### [00:54:20–00:54:50]

**EN**  have um so this this benchmark was run I think yesterday um by open router so this is this is daily data by open router. Um so whatever favorite inference provider you have please tell them not to you know reduce accuracy that much. Um this is GLM 5.2. Um so you know GLM 5.2 as well shows different accuracies. Um you can see so the plot on the right shows most model you know

**中文**  这份 benchmark 是 OpenRouter 昨天运行的 daily data。请告诉你常用的 inference provider，不要为了速度牺牲太多 accuracy。这是 GLM 5.2，同样在不同 provider 上显示不同 accuracy。右图说明大多数

### [00:54:49–00:55:17]

**EN**  most inference provide okay I keep saying model apps most inference providers are throughput maxing but they are accuracy minimizing that's where the phrase comes from okay so they do not care about in fact look like you know the highest accuracy is 76.4% and the lowest is 62.4%. So there is a 10% gap between the back you know between the highest accuracy and the you know lowest accuracy. Um and so like you need to you

**中文**  inference provider 在 throughput maximizing，却在 accuracy minimizing，这就是这个短语的由来。最高 accuracy 为 76.4%，最低 62.4%，二者相差十多个百分点。因此

### [00:55:13–00:55:41]

**EN**  know as a as a you know as a call out to inference providers you know please increase accuracy you know before trying to make things faster right you do not want a model to be very dumb um and it's like you know 10,000 tokens per second right we can make it 1 million tokens per second and there is no model um you know just call a human or something you know make a fake or something so yeah so the main point is we need inference

**中文**  我要呼吁 inference provider：先提高 accuracy，再追求更快。没人想要很快却很笨的 model。每秒一万 token 还可夸张到一百万，但若背后没有真正模型，干脆让 human 假装回答算了。核心是 inference

### [00:55:38–00:56:08]

**EN**  providers to do good in terms of accuracy otherwise this will make open source have a very bad look um yeah oh okay that's the end of the the second section I guess that was a bit of a rant um any other questions for this yes >> for a new organization that's that wants to use like the open source model do you suggest using a you know inference service provider or do you suggest

**中文**  provider 必须保证 accuracy，否则会严重伤害 open source 的形象。这一节结束，有点像 rant。观众：一个新 organization 想用 open source model，你建议使用 inference service provider，还是

### [00:56:05–00:56:32]

**EN**  downloading from hugging face and then using like model or you know some kind server to you know implement yourself like what do you suggest if any new organization comes and asks you like how do you use open source model >> that's a great question so when an open source model gets released you know how should you use it in terms of accuracy throughput or whatever um so in general um in general open source has come a

**中文**  从 Hugging Face 下载，再用某种 model server 自己部署？嘉宾：好问题。open source model 发布后，怎样在 accuracy、throughput 等方面正确使用？总体而言，open source 已经进步很远。

### [00:56:30–00:57:00]

**EN**  long way so for example we did report bugs in Jamaa 1 ja 2 llama mistro you know open gdosss every single of those models had bugs Um and so the good thing is you know as we will help the labs before they release a model to fix some of the issues. So every single model you now have has some of our fixes. So that's a good thing. Um but in general if you have a open source model I would use Llama CPP for example. I think Llama

**中文**  我们曾报告 Gemma 1、Gemma 2、Llama、Mistral、Open GDS 等模型的 bug。现在我们会在模型发布前帮助 lab 修复问题，所以你现在拿到的模型都可能包含我们的 fix。这是好事。对于 open source model，我建议用 llama.cpp。

### [00:56:58–00:57:27]

**EN**  CPP and Llama server is probably the most bugfree system. So I would like suggest yes you should download from hugging face. use Llama server, use Llama, you know, CLI, I don't know, you can use Unsoft Studio, whatever, whatever is your favorite tool. But you should, yes, you should download from Hugging Base. Um, in terms of like, you know, if you're a large enterprise, generally speaking, what they like to do is they like to wait one week. So most

**中文**  llama.cpp 与 llama-server 大概是 bug 最少的系统。建议从 Hugging Face 下载，再用 llama-server、llama-cli、Unsloth Studio 或你喜欢的 tool。对 large enterprise 来说，常见做法是等待一周，让所有问题先被修好，

### [00:57:24–00:57:53]

**EN**  enterprises, they'll wait one week for all the problems to be fixed. Um, and then, you know, then they will use the model. But in my view, that is not a good approach. I would say if you okay if everyone waits one week then like how do we fix the bugs? Um because only at scale only at scale then we can see the bugs. Um and so like in general we need everyone to start trying these models earlier. Um and not like you know wait

**中文**  再使用模型。但我认为这不是好 approach：如果所有人都等一周，bug 怎么会被发现？只有 scale 足够大，问题才会暴露。因此大家需要更早试用，而不是

### [00:57:51–00:58:06]

**EN**  one week wait one month you know don't do don't do the waiting approach. Um but I would say like in general the enterprises what they like to do is just wait one week. Um yeah that's like common practice. Um, yes.

**中文**  等一周或一个月。不过现实中 enterprise 通常就是等一周，这是 common practice。

### [00:58:15–00:58:44]

**EN**  >> So, okay, the question was why would the model performance degrade before a model release? Um, these are just hypothetical question hypothetical theories. So, every single model has a different system prompt. So, opus 4.8 8 Opus 4.8 system prompt is very short. Um but Opus 4.7 system prompter was extremely long. Um so the theory was this is just a theory that

**中文**  观众问为什么 model release 前 performance 会下降。以下只是 hypothetical theory。每个 model 有不同 system prompt：Opus 4.8 的很短，Opus 4.7 的极长。一种 theory 是

### [00:58:41–00:59:07]

**EN**  anthropic via claude code accidentally routed some of the models to Opus 4.8 right they use opus 4.8 as testing right they need to test opus 4.8 But they used Opus 4.7 system prompt. So they used the wrong system prompt and that is why accuracy degraded. That's one theory. Um, another theory is actually I think that's the actually I thought about it. That's probably the

**中文**  Anthropic 通过 Claude Code 意外把部分请求 route 到测试中的 Opus 4.8，却仍使用 Opus 4.7 的 system prompt。由于用了错误 prompt，accuracy 下降。这是一种 theory。我再想想是否还有别的；这可能是

### [00:59:05–00:59:34]

**EN**  only theory I had. I'm like thinking hm is there another theory? Um, >> I guess the harness itself like you know sometimes the harness itself the harness was was designed for Opus 4.7. Um and during when they were going to release 4.8, they need to collaborate the harness, right? They need to change the harness um for 4 4.8 to make it work. But the

**中文**  唯一一个。也可能 harness 本身为 Opus 4.7 设计；要发布 4.8 时，需要 calibrate 并修改 harness。但

### [00:59:32–01:00:02]

**EN**  problem is you're not allowed to publish it, right? You're not allowed to publish it and give it to people because otherwise people will like, you know, go into Twitter on LinkedIn, you know, everywhere. Oh, I can see 4.8 is going to be released. You know, everyone's going to be screaming, you know, 4.8's coming. 4.8's, eights, you know, were getting released and so maybe that's why accuracy decreased. It's they update they did not update the harness. Um or the other option is they up they already up they silently they silently updated

**中文**  又不能提前公开，否则大家会去 Twitter、LinkedIn 大喊“4.8 要发布了”。所以 accuracy 也许因 harness 尚未更新而下降；另一可能是他们在新模型发布前 silently update 了

### [01:00:00–01:00:29]

**EN**  the harness before the new model got released and it regressed you know it reduced accuracy. Um I don't know like to be honest I you should probably ask anthropic that question. Um or but I think in general in general the dips the dips don't always correspond to like new model releases. Some of the dips are actual issues like you know the thinking trace got deleted um the system prompt they wrote a wrong I think for the system it's funny I think for the system

**中文**  harness，却导致 regression。老实说应该问 Anthropic。并非所有 dip 都对应新 model release；有些是实际 issue，例如 thinking trace 被删、system prompt 写错。有趣的是，他们为

### [01:00:26–01:00:56]

**EN**  prompt they said um they tried to reduce verbosity so they tried to make the model less talkative um and it actually made the model dumber um and so I think it was just one word they added one word no one sentence I think one sentence in the system prompt that made the model dumber Um, yeah, I don't know if that helps, but I don't know if anyone else has any like theory. I don't I don't think so anyone even has that many theories on this. Um,

**中文**  降低 verbosity、让模型少说话，在 system prompt 里加了一句，结果模型变笨了。也许只是一句话。希望这有所帮助，但我不知道还有谁提出过其他 theory。

### [01:00:54–01:01:22]

**EN**  obviously the anthropic engineers will know, but I, you know, they're not going to tell. So, it's just based on hypotheticals. Something to do with the system problems, something to do with the harness. Yeah. But I think in general, you can also use this plot. You know, if the performance decreases, most likely a new model is going to be coming. Um, yeah. Any other qu? Yes. >> Just to add on that Yes, correct.

**中文**  Anthropic engineer 显然知道，但不会告诉我们，所以这里只能假设，可能与 system prompt 或 harness 有关。也可以反过来把图当作 release indicator：performance 下降时，可能有新模型要来。还有问题吗？

### [01:01:26–01:01:55]

**EN**  >> Correct. >> Yes. Exactly. >> Exactly. So before a model release, they use a different system prompt for that new model for the old model. And so that is probably why there are some decrease in accuracy. they switch the system prompts around or something like that. Um and also you know the model itself you know I think 4.8 system prompts is very short. Um it's yeah I think it's

**中文**  观众补充：发布前会把新模型的不同 system prompt 用到旧模型上，因此 accuracy 下降。嘉宾：对，也许他们交换了 system prompt。4.8 的 prompt 很短，

### [01:01:53–01:02:21]

**EN**  like very very short and 4.7 was ginormous. Um and the reason is 4.7 was like you know I don't know what I don't know what happened but they have this ginormous system prompt and the 4.8 just shrunk it a lot. Um so maybe maybe they used the 4.7 system prompt I don't know or 4.8 system the short system prompt for 4.8 eight and then they use it for 4.7 and that's why it decreased accuracy. I don't know. Um but yeah, you're correct. Um they do release although I think the system prompt they

**中文**  4.7 的极长，不知为何。也许给 4.8 用了 4.7 的 prompt，或者把 4.8 的短 prompt 用到 4.7，导致 accuracy 下降。你说得对。不过他们网站公开的 system prompt 可能是

### [01:02:19–01:02:47]

**EN**  released on the website is for claw.ai. So the online chat system um the clawed code system prompt is actually different. >> Yeah. So I think you need to actually call you need to call claude code you know what is my system prompt and then you print it to like a text file um and then you can like in investigate what the system prompt is and then you can also override it if you want um yes but it's a different system prompt most

**中文**  Claude.ai 在线 chat 的，Claude Code system prompt 另有一套。你可能要直接问 Claude Code“我的 system prompt 是什么”，打印到 text file 后调查，也可以自行 override。它们大概率不同。

### [01:02:45–01:03:14]

**EN**  likely um yeah last question if anyone no okay continue on Okay, the next section we're going to be talking about is benchmaxing and cheating. Um, I'm not sure if you folks have seen the deep SWE benchmark. Um, the deep SWE

**中文**  最后一个问题，没有就继续。下一部分谈 benchmark maximizing 与 cheating。不知道大家是否见过 DeepSWE benchmark。

### [01:03:12–01:03:41]

**EN**  benchmark is a very popular recent benchmark that shows, you know, the cost is on the X-axis and the Y-axis it is a deep SWE benchmark. It's a new benchmark based on like you know a better uncontaminated version of Swebench Pro. Um and in general you can see that you know GBD 5.5 does very well with Fable um you know GLM Opus 4.8 in general right it shows you know this plot shows that models are getting you know um

**中文**  DeepSWE 是最近很流行的 benchmark。x-axis 是 cost，y-axis 是 DeepSWE score。它基于 SWE-bench Pro，但试图做成更好、未被 contamination 的版本。图中 GPT-5.5、Fable、GLM、Opus 4.8 都表现不错。

### [01:03:39–01:04:06]

**EN**  these the dots are different reasoning modes um I think this is maximum reasoning I think um high extra high you know these are actually different reasoning um reasoning times as well um but in general you can see that there is a parto efficiency trend right the best model is the one you know to the right to the top right the better the model to the right to the top is the better the model. Um, so you want the models to do

**中文**  每组 dot 表示不同 reasoning mode 或 reasoning time，例如 high、extra-high。整体存在 Pareto-efficiency trend：越靠右上，模型越好，随着时间应持续向右上移动。

### [01:04:04–01:04:34]

**EN**  better and better over time to that to the top right corner. And you know, I just learned I didn't actually know this. I just learned that Sweetbench Pro when you run this benchmark you use LL you use language models as a verifier. Um, and I was like confused because like for most benchmarks, you should never call another language model to check

**中文**  我刚知道 SWE-bench Pro 运行时会用 LLM，也就是 language model 作为 verifier，这让我很困惑。多数 benchmark 绝不应调用另一个 language model 判断答案

### [01:04:32–01:05:01]

**EN**  whether your answer is right or wrong. And so for Swedebench Pro, you actually call a language model to verify if your language model was right. Um, and so that is why SweetBench Pro is not a very good benchmark. Um, one of the problems is is do we need to do sampling? Like how many verification runs do you need to run to verify if your answer is correct? Do you run it one time? Do you

**中文**  对错。但 SWE-bench Pro 会让一个 language model 验证被测 language model 是否正确，所以它并不是很好的 benchmark。一个问题是 verification 要不要 sampling：运行一次、

### [01:04:58–01:05:27]

**EN**  run it five times? You run it 100 times and take like an average. Um, so I was actually quite shocked that this is actually what happens. I was quite surprised actually. Um, the next question is which model is the verifier? You know you ask for example you ask opus you know you you benchmark opus 4.8 on swbench pro but which what do you use as a verifier do you use opus 4.8 eight as the verifier to using the same model

**中文**  五次还是一百次再取 average？我很惊讶竟然如此。另一个问题是 verifier 用哪个 model？例如 benchmark Opus 4.8 时，是否也用 Opus 4.8 验证，也就是同一模型

### [01:05:24–01:05:53]

**EN**  itself to verify itself. Um and so like this I was like quite surprised actually that this is how benchmarks work. Um and actually quite disappointed. Um but anyways obviously you can go the other approach. You can do human verification. Um you know everyone in the room I'll give you the bench you know and just tell you guys to verify it. Um you could do that I guess. Um and also what happens if the verification changes

**中文**  验证自己？我对 benchmark 的这种工作方式很失望。也可以用 human verification，把题交给现场每个人人工检查。但如果 verification 每天发生变化呢？

### [01:05:50–01:06:18]

**EN**  every day? Um you know remember previously models you know every single day models get better or worse. um what happens what happens if you run what happens if you run the verification when the model was doing very bad right you will actually have different SWE numbers um and so like I'm actually quite surprised this is what the industry does um you know run bench pro but using anonymous is verifies that is definitely

**中文**  前面看到模型每日会变好或变差；若恰好在 verifier 状态很差时运行，就会得到不同 SWE score。行业竟这样做让我很惊讶。使用 anonymous LLM verifier 跑 SWE-bench Pro 绝不是好主意，但人们仍然这么做。根据 DeepSWE，

### [01:06:16–01:06:43]

**EN**  not a good idea um but anyways people do it whatever um in fact according to deepu um if you do verification using language models Sweet Bench Pro has a 8.5% false positive rate. Um, and a false positive rate means that the LLM verifier said that the model was correct, but it was actually wrong. Um,

**中文**  language-model verification 使 SWE-bench Pro 出现 8.5% false-positive rate，也就是 verifier 说答案正确，实际却错误。

### [01:06:40–01:07:07]

**EN**  and so 8.5% of the time it would do this. Um, the false negative rate is even worse at 24%. Um, this means that the verifier said that the model was wrong, but it was actually right. Um and so you can see that SweetBench Pro is a very bad benchmark. Um and so Deep Sweet showed that they have in you know they fixed the problem you know um by

**中文**  false-negative rate 更糟，高达 24%，也就是 verifier 说模型错误，实际却正确。由此可见 SWE-bench Pro 是很差的 benchmark。DeepSWE 声称修复了问题，把 false positive 和 false negative 降至约 1%。

### [01:07:05–01:07:34]

**EN**  reducing the false positive rate and the false negative rate to you know 1%. Um in fact some examples of cheating um I you know this is actually quite surprising um but in the bench pro benchmark you get you get like a GitHub question you know a GitHub issue you call the model to solve that GitHub issue but did you know that in Swebench Pro

**中文**  还有一些 cheating example 很惊人。在 SWE-bench Pro 中，模型拿到一个 GitHub issue，要解决它；但你知道它同时能拿到完整 Git history 吗？也就是说，

### [01:07:32–01:08:01]

**EN**  you get the full Git history so you get the you get the actual answer as well um so I'm like I'm actually quite I was actually quite shocked to learn this um that during these models you give the answer and the question like obviously the model will cheat. Um and so like this is definitely a very bad benchmark. Um you know you should never ever ever give the model the answer. Um and

**中文**  连真实答案也给了模型。我很震惊：同时提供 question 与 answer，模型当然会作弊。这是很差的 benchmark，绝不能把答案交给模型。

### [01:07:58–01:08:27]

**EN**  so very silly. Um but yes this happens a lot. Um and you do not want the model to literally see the solution, right? That is a terrible approach. Um the other problems that you get get like false positives is you know the PR tests you know the the GitHub the GitHub issue tests are very weak. Um so you know at the final conclusion you know when the GitHub when the GitHub issue is closed with a pull request the tests that the

**中文**  让模型直接看到 solution 是很糟糕的 approach。另一个导致 false positive 的问题是 PR test 很弱：GitHub issue 最终通过 pull request 关闭时，maintainer 写的 test 不够好。

### [01:08:25–01:08:54]

**EN**  maintainer wrote are not very good. Um and so the problem of that is you know if you have tests which are very weak then you know the model does very well not very good. Um and obviously the worst part is the model will like bypass some tests. It will skip some um and that is not a very good approach. In fact um deep suite actually showed how many times a model cheats by looking at the full git history you know

**中文**  test 太弱，模型就容易拿高分；更糟的是模型会 bypass 或 skip 某些 test。DeepSWE 统计模型多少次通过查看完整 Git history 直接找到答案，也就是作弊。

### [01:08:51–01:09:21]

**EN**  directly going to the answer. Um you can see opus 4.7. So the purple bars show cheating by models. Um ah it looks like Jubilee 5.5 never cheats. It looks like it um h okay maybe we should use GP 5.5. Um you know actually this is actually very interesting. There are some people which think that if you cheat that's actually good. Um and the reason why it's good is it means that Opus 4.7 already know like

**中文**  紫色 bar 表示各模型 cheating，Opus 4.7 会这样做；GPT-5.5 看起来从不作弊，也许我们该用它。有些人认为作弊反而好，因为这说明 Opus 4.7 知道

### [01:09:19–01:09:49]

**EN**  if you give it the full Git history you should be able to like you gave it to them right you gave Opus the full Git history it should find the solution there right it should just directly skip over to the solution. So it's that's what people think you know people have a view that the humans gave Opus 4.7 the full gate history so it should cheat right you you designed it to cheat um so in general cord models seem to cheat more

**中文**  既然 human 给了完整 Git history，就应该从里面找到 solution，直接跳到答案。也就是说，是 benchmark 设计让它作弊。总体上 Claude model 似乎作弊更多，

### [01:09:47–01:10:16]

**EN**  um and open AI models seem to cheat less in general um so it depends on you you know if you want a model to cheat or not um and the definition of the word cheat is also very you know charged so I guess it depends on what the word cheating means um you know for false negatives remember Swebench Pro calls a language model to verify if your answer is correct um and so sometimes it's not very good you know

**中文**  OpenAI model 较少。是否希望模型作弊，取决于使用者；“cheat”这个词本身也很带倾向。再看 false negative：SWE-bench Pro 用 language model 判断答案，有时会出错。

### [01:10:14–01:10:44]

**EN**  sometimes you have unrelated tests that fail um you forgot you know sometimes when you write tests you forgot about the tests which have helpers you know helper functions and you just skip that um so there are many issues and this I think this was 20 yeah so 24% % of the time, 24% of the time, the model says, the verifier says your model was wrong, but it was actually correct. So, this is another problem.

**中文**  可能有 unrelated test 失败，或写 test 时遗漏使用 helper function 的 test，于是被跳过。许多问题导致 24% 的情况下，verifier 说模型错误，实际却正确。

### [01:10:41–01:11:11]

**EN**  And even worse, the harness itself can change accuracy. So, when you benchmark using Swebench Pro, like you need to have one agent or one harness for all models, right? How do you create a generalized control environment for these models? Um and so you can see like you know for example DeepSu showed if you use clawed code you get 40% accuracy but then if you use their own so it's a

**中文**  更糟的是，harness 本身也会改变 accuracy。benchmark 所有 model 时应该使用同一个 agent 或 harness，但怎样为不同模型建立 generalized control environment？DeepSWE 显示，使用 Claude Code 可得 40% accuracy，但换成他们自己的

### [01:11:07–01:11:37]

**EN**  special harness you can get 50% accuracy um Gemini for example right if you use Gemini CLI you get 20% accuracy but if you use their one you know the the control environment you can get 40% accuracy um and so in general for these benchmarks you also need to have a controlled environment. Um, and that is also another problem. And with Deep Suite, they showed by

**中文**  special harness，可得 50%；Gemini 用 Gemini CLI 只有 20%，换成他们的 control environment 可达 40%。所以 benchmark 也需要 controlled environment，这是另一个问题。DeepSWE 表明，

### [01:11:35–01:12:04]

**EN**  using this benchmark, by solving, you know, by stopping cheating, you know, by, you know, if we remove cheating, if we remove, you know, these other issues, you can see the models, you know, the models are not saturated anymore, right? You can see the models are very different in terms of the capabilities. According to this benchmark, GBD 5.5 is the best according to this one. Um I don't Oh, this is not updated. Um for 4.8 I think is over here or something.

**中文**  若用新 benchmark 阻止 cheating 并移除其他问题，模型能力就不再 saturated，彼此差异明显。按这个 benchmark，GPT-5.5 最好；图还没更新，Opus 4.8 大概在这里。

### [01:12:01–01:12:30]

**EN**  Um but yes, this benchmark shows core taiku is 0%. Um accuracy, right? It's terrible, I guess. Um but yeah, this benchmark just show okay the main question is do you trust this benchmark? That is another question. Um there is other benchmarks, right? So cognition released a frontier code benchmark which also tries to solve the same questions for benchm you know for cheating and benchmarks. Um and what they showed is you can fix

**中文**  它还显示 Kilo Code accuracy 是 0%，很糟。不过关键仍是你是否信任 benchmark。还有其他 benchmark：Cognition 发布 Frontier Code benchmark，也试图解决 contamination 与 cheating。

### [01:12:28–01:12:58]

**EN**  contamination. And how do you fix contamination? You ask, you know, you ask Cognition's team, which is full of like, you know, national Olympiads and, you know, international Olympiads. They manually checked every single question um themselves, you know, and removed bad questions, you know, bad examples. Um and they also showed that their questions are much more diverse, right? So, Frontier Code has many different other languages. Um and they showed with

**中文**  怎样解决 contamination？Cognition 团队有许多 national 与 international Olympiad 选手，他们人工检查每道 question，删掉坏题和坏 example；还提高问题 diversity，让 Frontier Code 覆盖多种 programming language。

### [01:12:56–01:13:25]

**EN**  diversity you know with more diverse programming languages um and by reducing contamination they also have a benchmark um and according to their benchmark opus 4.8 is the best right for 14.5% accuracy juby 5.5 is 7.2 to accuracy. Um, and this is the diamond one, right? So, this is the 50 the 50 hardest questions. Um, the main benchmark is 100

**中文**  通过更丰富的 programming language 与更少 contamination，他们也建立了 benchmark。按结果，Opus 4.8 最好，accuracy 14.5%；GPT-5.5 为 7.2%。这里是 Diamond set，即最难的 50 题；主 benchmark 有 100

### [01:13:23–01:13:52]

**EN**  questions and the extended is 150. Um, and so according to them, you know, Claude does the best according to them. But also according to them, Frontier code seems to be better than Deep Suite. Right? The benchmark that I showed previously, Deepswuite, this one um you know according to Frontier Code, so the cognition team, their benchmark is better than Deepswuite, right? According

**中文**  题，extended 有 150 题。按他们的数据，Claude 最好；而且 Frontier Code 自称优于之前的 DeepSWE。

### [01:13:49–01:14:17]

**EN**  to them, according to them, Deep SW's false positive rate is 44.9%. But remember what did Deepswu say? They said the false positive rate was I don't remember what what did they say? Um they said that it was 0.3%. Right? So deep said deep said their false positive rate is 0.3%. But Frontier code said that Deep SWE's false

**中文**  Cognition 团队称 DeepSWE 的 false-positive rate 是 44.9%。但 DeepSWE 自己声称是多少？他们说只有 0.3%。也就是 DeepSWE 说自己的 false-positive rate 是 0.3%，Frontier Code 却说它是 44.9%。

### [01:14:14–01:14:42]

**EN**  positive rate was 44.9%. Um so you know there is some competition I guess between benchmarking labs um well cognition is not a benchmarking lab but like you know between companies um so the main question is who do we trust you know do we trust Frontier codes benchmarks? Do we trust Deep Swiss benchmarks? Do we trust Bench? You know, who do we trust?

**中文**  看来 benchmarking company 之间也有竞争。Cognition 并非专门 benchmark lab，但问题仍是信谁：Frontier Code、DeepSWE、SWE-bench，究竟信哪一个？

### [01:14:40–01:15:08]

**EN**  And that is a very important question. Um, you know, my take is like, you know, like let's just take an average of everyone. Take an average of everyone and you'll probably get the best answer. You know, who is actually doing the best. Um, yeah, but this is actually very interesting. Um, you know, it show Okay, so according to them, the false negative rate for Deep Suite is correct. You know, 1.2%. But my interest, you know, I probably, you know, my main question is why is the false positive

**中文**  这很重要。我的做法是把所有结果取 average，大概能得到谁表现最好的较可靠答案。很有意思的是，按 Frontier Code，DeepSWE 的 false-negative rate 1.2% 是对的；但为什么 false-positive

### [01:15:05–01:15:35]

**EN**  rate so high for deep according to Frontier Bench, Deep Sweet is even worse than Sweet Bench Pro. That's what they're trying to say, I guess, for for the false positive rate. Um, yeah. And even worse, there is another benchmark called Frontier Math. Um, so Frontier Math is by Epoch AI. Um so they have this math benchmark with different tiers you know tier one, tier

**中文**  rate 那么高？按照 Frontier Code，DeepSWE 在 false positive 上甚至比 SWE-bench Pro 更糟。还有一个 benchmark 叫 FrontierMath，由 Epoch AI 建立，按难度分 tier 1 到

### [01:15:32–01:16:01]

**EN**  two, tier three, tier four. So tier four is the hardest. Um but the benchmark itself was botched. Um and so they actually had to release a corrected version of their benchmark. Um I think this was one month ago um or something. Um so they showed that their benchmark questions were fully wrong. Um, and you can see that if you correct the benchmark, if you correct the benchmark,

**中文**  tier 4，tier 4 最难。但 benchmark 本身出了问题，不得不发布 corrected version。大约一个月前，他们承认 benchmark question 有大量错误。修正后，

### [01:15:58–01:16:27]

**EN**  the accuracy for GBD 5.5 jumps from 50% to 80% or something. Um, and so now you kind of trust the benchmark. And they showed in a tweet, oh, it's June 12. Oh, it's only two weeks ago. Um, so in June 12, they showed that the reason why they did bad on the benchmarks is they they did the answer extraction incorrectly. For example, they did, you know, they had unclear questions. They had the incorrect sign. So, for example, they said the model

**中文**  GPT-5.5 accuracy 从约 50% 跳到 80%，这让人怎样信任 benchmark？他们在 6 月 12 日，也就是两周前发 tweet，称问题来自 answer extraction：有 unclear question、incorrect sign；例如模型回答

### [01:16:25–01:16:53]

**EN**  said 12, but it should be actually minus 12 and they forgot to cut the minus sign. Um, they have one-off errors. Um, yeah, there's many problems with the benchmark. Um, and so they fixed their benchmark um, just recently. In fact, you know, it's actually quite funny. This was just two weeks ago. Have you guys heard of hugging faces

**中文**  -12，却被错误提取成 12；还有 off-by-one error 等。最近才修复。更有趣的是，大家听过 Hugging Face 一年前推出的 Math-Verify 吗？

### [01:16:49–01:17:19]

**EN**  math verify which was one year ago? Um and hugging base showed that in fact these benchmarks when you do math questions they always do bad and the reason why is because there's many problems right the formatting is incorrect um you know the extraction of the fraction is wrong um you know the sign is failed extraction there's many many problems of mathematical extraction and to be honest I feel like it's like kind of a reinventing the

**中文**  Hugging Face 早已指出 math benchmark 经常低估能力，原因是格式、fraction extraction、sign extraction 等问题。Epoch 两周前才修，某种程度上是在重复发明轮子或重新发现旧问题。

### [01:17:15–01:17:43]

**EN**  wheel or you know rediscovery um but hugging base actually published this one year ago and epoch just fixed it 2 weeks ago. Um so you know benchmarking labs definitely need more you know they need to investigate literature more I think. Um in fact according to hugging face math verify you know if you use if you the green bar the green bar is if you do

**中文**  benchmark lab 需要更多 literature review。按 Hugging Face Math-Verify，绿色 bar 表示未用他们的 verification system 修 benchmark；修复后，accuracy 会显著提高。例如 Qwen 从 10% 上升到 25%。

### [01:17:42–01:18:11]

**EN**  not use hugging faces's verification system you know to fix the benchmark. If you do fix the benchmark, you can see accuracy dramatically increases, right? For example, for Quen, for Quen, the accuracy was 10%, now it's 25%. Um, and so you need to, so that means the open source models are not dumb. They just have different they output a different format. Um, and so one of the problems is how do we actually actually like, you know, pass these different formats?

**中文**  这意味着 open source model 并不笨，只是输出 format 不同。问题是怎样 parse 各种 format。情况甚至更糟：我在 2024 年 8 月 tweet 过，不同 tokenization 也会改变 accuracy。

### [01:18:09–01:18:36]

**EN**  In fact, it's even worse. Um, no, I think I tweeted, oh, I tweeted this in August 2024. Um, that if you if you use different tokenization, you can also have different accuracy. Um, in fact, for MLOU, if you use spaces, you increase accuracy by 0.4%. Um, it might not sound like a lot, but the point is by these very dumb things like, you

**中文**  在 MMLU 中，仅仅加入 space 就能让 accuracy 提高 0.4%。听起来不多，但这些小事——space、-12 被读成 12——都会改变 benchmark score。

### [01:18:32–01:19:02]

**EN**  know, using spaces or, you know, minus 12 becomes 12. Um and all of these like dumb little small things, the accuracy of these benchmarks can change over time. Um and so like the main question is you know how do we make benchmarking labs and benchmarking companies you know how do we make them more reliable um and you know more trustworthy. Oh okay that's I guess the section for

**中文**  核心问题是怎样让 benchmark lab 和 company 更 reliable、更 trustworthy。这一节结束，有问题吗？

### [01:18:59–01:19:09]

**EN**  the benchmarking part. Any other questions for that section? Um questions. Yes.

**中文**  观众开始提问。

### [01:19:41–01:20:09]

**EN**  That's a great question. So the question is how do we how can we trust these benchmarking companies or like what other types of benchmarks can we do to make it trustworthy? Um so that is actually a very good question. The main question for benchmarks is you need to satisfy two conditions. The first condition is the benchmark must not must not be benchmaxable. Right? How do you make a benchmark that is extremely hard to benchmark, right? How do we like not

**中文**  问题是怎样信任 benchmark company，或还有哪些 benchmark 更可信。这是好问题。benchmark 必须满足两个条件：第一，不能容易被 benchmark-maximize，也就是不能让模型轻易把 accuracy 刷到 100%；第二，

### [01:20:07–01:20:36]

**EN**  get 100% accuracy? And the second question is how do we make the benchmark um verifiable, right? So how do we make the benchmark you can you can also verify that the answer is in fact correct, right? You remember Swebench Pro is dumb because you call the language model itself to verify itself. Um so that is not good. Um so the main question is those two questions. Um, and so one good example, this is just a dumb example.

**中文**  必须 verifiable，能够确认答案真的正确。SWE-bench Pro 让 language model 验证自己，就不合格。举个简单例子，

### [01:20:36–01:21:04]

**EN**  Randomly create maths questions. Sample for example. Okay, this is okay, this is probably not a good benchmark. You automatically create maths questions. Um, we can sample infinity, right? We can sample infinite maths questions, right? 2+ 2, 4 plus 4, you know, any single number added together. That's one question. Can you verify this? Yes, you can. Right? You can call

**中文**  随机创建 math question。也许不是好 benchmark，但可以自动生成无限问题：2+2、4+4，任意数字相加。能验证吗？可以，用 calculator。

### [01:21:01–01:21:29]

**EN**  a calculator to verify what is 2 plus two. Can this be benchmaxible? Hard. And the reason why hard is because the sampling space is infinity, right? It can be 2 plus 2, 1,00 plus 101, right? It can you you don't have to do plus, right? You can do 1,00 times 1,00. And so that's one way make a benchmark which is very hard to cheat but also easy to verify. So some sort of math

**中文**  容易 benchmark-maximize 吗？很难，因为 sampling space 无限，可以是 2+2、100+101，也可以不是加法，而是 100×100。这样就得到难作弊、易验证的 benchmark。

### [01:21:27–01:21:57]

**EN**  question. Um, the other one, for example, is um, okay, maybe this is not a good example. I'm just making this one up on the spot. Tell the model to create a poem in 70 words and you must use the word happy. Can you verify this? Yes, you can. Is happy in the, you know, generation? If yes, plus one. Also, you can count how many words, right? You can count, okay,

**中文**  另一个临时想到的例子：要求模型用 70 个 word 写 poem，并且必须包含 happy。容易验证吗？可以检查 happy 是否存在，并计算是否正好 70 个 word。

### [01:21:54–01:22:24]

**EN**  is there 70 words? Um, so you can do these type of approaches. And is this benchmaxable? No. It's it's very hard to benchmark because you can say 70 words, 69 words, 68 words, 102 words, 1,000 words, right? It doesn't have to be happy. It can be you must have two words, you must have three words. Um, so some some sort of benchmark where it's very hard to benchmark. Um

**中文**  这类任务也难以 benchmark-maximize，因为长度可以改成 69、68、102 或 1,000，关键词也可换成两个或三个其他 word。需要的就是这类难刷分的 benchmark。

### [01:22:21–01:22:41]

**EN**  yeah, in my view I think that's that's probably going to be the most important benchmark and I don't I don't think so anyone has actually made this yet. Um I don't know maybe someone in the audience or you know you guys can go as teams I don't know make a startup or something you know do that. Um and I feel like that benchmark will be very very important. Um yeah

**中文**  我认为这会是最重要的 benchmark，但似乎还没人真正做出来。也许现场有人可以组队创业去做，它会非常重要。

### [01:22:45–01:23:15]

**EN**  >> benchmarks we can trust today none of them. take an average of all of them. To be honest, probably the best approach is just vibe uh vibe checking. Try all of them and see which one you like the best. Um to be completely honest, I just you know like these benchmarks h like the main the main issue I have with benchmarks is for example um you know I mean like this one right

**中文**  问今天有哪些 benchmark 可相信？一个都没有。把它们取 average，或者干脆 vibe check：都试一遍，看自己最喜欢哪个。因为即使同一 benchmark 每天都可能变化。

### [01:23:12–01:23:40]

**EN**  this one I mean even every single day the benchmark can change. So we can't trust the benchmarks anymore. So my fundamental view is do not trust any benchmarks. Take an average and then okay then main question is who's taking the average? I guess artificial artificial analysis has some average. The only problem is they have some weightings for the weight. Um you know each benchmark has a weight. So now the question is you know what is the waiting of each benchmark. You know you

**中文**  所以我的根本观点是不要信任任何 benchmark，先综合平均。但由谁取平均？Artificial Analysis 有某种 average，可他们为不同 benchmark 设置 weight，于是又要问每项权重是多少。

### [01:23:37–01:24:02]

**EN**  can't just take like a dumb average. Um you know you can't just say you know 10 benchmarks divided by 10. Um that's probably not going to work. So the main question is how do you even do the waiting? That's another problem. Um, so I think in general it's based on vibe checking, I guess. Yeah, I guess I don't have an answer for that. Um, any other questions? Yes.

**中文**  不能简单把十个 benchmark 相加再除以十；那大概行不通。怎样 weighting 又是一个问题。所以总体只能依靠 vibe checking，我没有完整答案。

### [01:24:44–01:24:47]

**EN**  question.

**中文**  观众提问。

### [01:24:55–01:25:23]

**EN**  >> Yes. >> So that way Good matter. >> You're correct. So the question was in terms of because we bench pro for example you call a model. The question is what model? Could it be 4.8? Could it be GB 5.5 and you call this model to verify the benchmark? Um you and so the question was can you use an open source

**中文**  观众问：SWE-bench Pro 需要调用一个 model 做 verifier，可以是 Opus 4.8 或 GPT-5.5，能否改用 open source model，建立 controlled environment？

### [01:25:21–01:25:50]

**EN**  model instead? So then now you you have a controlled environment. So yes you can. But remember there is a problem because even open source models itself have bugs times you know the inference engines have bugs times the inference providers have bugs and accuracy degradation. So it's you're correct. Um so the main question is we need to have some one or some organization you know some person or some whatever

**中文**  可以。但 open source model 自身有 bug，inference engine、provider 也有 bug 与 accuracy degradation。所以还需要某个人、organization 或

### [01:25:48–01:26:16]

**EN**  committee that we can investigate you know which engine did you use do not update the engine you know the engine must be the same you know the weights must have not changed so there's many many many problems with this approach um but I do agree as open you can use an open source model but it's not it doesn't solve the other problems um yeah does that Okay. Um, so the next section I'm going to be talking about is cyber security and

**中文**  committee 审核使用了哪个 engine，并冻结版本、确保 weight 没有变化。open source model 可以解决部分问题，但无法解决其他所有问题。下一节讲 cyber security 与

### [01:26:14–01:26:42]

**EN**  regulation. Um, this is a interesting topic. Um, so I'm not sure if you have all folks have seen this plot. It shows the AI security institutes. Um, I think this is from the UK. Um, they show the performance of models based on some cyber security task. Um, and they show that mythos preview seems to be the best. Um, you know, with GBD 5.5 cybar, you know, preview and so on. Um, they

**中文**  regulation。这里是 UK AI Security Institute 的图，展示模型在某些 cyber security task 上的 performance。Mythos Preview 最好，其后有 GPT-5.5 Cyber Preview 等。

### [01:26:38–01:27:06]

**EN**  show this benchmark. Uh and again previously as I mentioned weird ML is a better ben in my view okay this is just my take weird ML is a better benchmark in general for benchmarking intelligence on models and the reason why is because it doesn't actually it doesn't actually follow the trend of reasoning versus non-reasoning remembering reasoning previously I think I okay I don't have it um reasoning um

**中文**  如前所述，我个人认为 WeirdML 更适合 benchmark 模型 intelligence，因为它不遵循 reasoning 与 non-reasoning 的趋势。前面 reasoning model 把

### [01:27:04–01:27:34]

**EN**  the reasoning models doubling time reduced by half to 3.5 months So remember, you just need to wait 3.5 months and the models capabilities will double. Um and the non-reasoning was 7 months. Um so you need to wait seven months for the models to double in capability. Um but weird ML did not actually have this trend. Um the weird ML benchmark showed that actually the trend was like there is no trend. Um, and I think I'll just talk about this, you know, like one of the biggest

**中文**  doubling time 减半到 3.5 个月，non-reasoning 则是七个月；但 WeirdML 没显示这种 trend。我还想强调，benchmark 最大问题之一是必须不断 reinvent，并重新调整多个 benchmark 的 weighting。

### [01:27:32–01:28:01]

**EN**  problems of benchmarks is you need to constantly reinvent yourself and do reweings of combinations of benchmarks. For example, artificial analysis just recently released, you know, their new v4.1 benchmark and they showed the waiting of the benchmarks. Um, you know, GDP vow is 20%, terminal bench is 16% and so on. Um and so they they designed these numbers as waitings for each of those benchmarks and then they've averaged it up together. Um so the main

**中文**  Artificial Analysis 最近发布 v4.1 benchmark，并公布各项 weight，例如 GDPval 20%、Terminal-Bench 16% 等，最后加权平均。关键问题是这些数字怎样决定。

### [01:28:00–01:28:29]

**EN**  question is how do you actually determine these numbers? Um and so this is more like a human approach. You know you have to determine these numbers. um you know arc AGI kind of saturated on ARC AGI 1 and so that's why we have ARC AGI 2 and that is also why we have ARC AGI 3 and my you know I guess once ARC AGI 3 is saturated then we have ARC AGI 4 5 6 7 whatever um and the main point is once you have benchmarks is it called good art I don't

**中文**  这终究依赖 human judgment。ARC-AGI-1 饱和后有 ARC-AGI-2，再有 ARC-AGI-3；等 3 饱和，就会有 4、5、6、7。benchmark 一旦被 Goodhart's law 影响，

### [01:28:27–01:28:56]

**EN**  remember um the good the benchmark itself becomes useless because you know models will start benchmarking on this so one of the biggest problems of these larger models for cyber security for example um is mythos actually dramatically went out of the trend um and that is why you know many people are afraid of these you know mythos you know GPD 5 point 5.6 you they're afraid of these models because it went out of

**中文**  也就是模型开始针对它优化，它本身就会失去价值。cyber security 的一个问题是 Mythos 明显突破历史 trend，这让很多人担心 Mythos、GPT-5.5/5.6 等模型，因为它们偏离了

### [01:28:54–01:29:23]

**EN**  trend um you can see that mythos dramatically went out of trend um and you know even you know you know GP 5.6 didn't really release that many benchmarks because it was in preview mode. Um, so this is from their system card. They showed for cyber security that Guby 5.6 does very very well. Um, in fact because Guby because Guby 5.6 they I think they only did terminal bench as their benchmark. They did not benchmark on anything else. Um, they did

**中文**  趋势。GPT-5.6 仍在 preview，公开 benchmark 很少；system card 显示它在 cyber security 上非常强。GPT-5.6 似乎只跑了 Terminal-Bench，没有 benchmark 其他项目。

### [01:29:21–01:29:51]

**EN**  have in their system card they did have one benchmark which is very important. Um, and this is called the internal research debugging evaluation. Um, and this is OpenAI's own set of set of questions. So, you know, the custom open source, you know, if you want to, it's their own set of 10 questions or whatever that they benchmarked GBD 5.6 on. Um, and according to them, it does very very it does better. Okay, I was going to say very, very well, but it's not. Um, it does better. Um, and you can

**中文**  不过 system card 中有一项很重要的 benchmark，叫 internal research debugging evaluation，是 OpenAI 自己的一组 question，也许只有十题，用来 benchmark GPT-5.6。按他们的数据，它确实表现更好——我本想说非常好，但只是更好。

### [01:29:49–01:30:17]

**EN**  see that GBD, it's actually kind of interesting. GBD 5.5 did worse than GBD 5.5 a four um for OpenAI's own internal um research evaluation um and you know GBD 5.6 definitely does much better right you can see that the G GBD 5.6 soul, you know, if you extend it, it does much better. Um, but interestingly, Terara does better um somewhat sometimes. Um, yeah.

**中文**  有意思的是，在 OpenAI 自己的 internal research evaluation 上，GPT-5.5 似乎比 GPT-5.5a4 更差；GPT-5.6 则明显更好，延长运行时间后更强。但 Terra 有时又略胜一筹。

### [01:30:18–01:30:47]

**EN**  And, you know, one of the biggest problems of these models that are getting getting better and better and better is I don't know if you guys know that, you know, open-source exploits are getting worse and worse and worse. Um and so the high exploit ratio you know number of critical vulnerabilities that were discovered has skyrocketed you know recently you know every single week or day some sort of open source package gets compromised um and they actually

**中文**  模型越来越强带来一个大问题：open-source exploit 正越来越严重。high exploit ratio，也就是发现的 critical vulnerability 数量，最近急剧上升；几乎每周甚至每天都有 open source package 被 compromise。图表显示问题越来越严重。

### [01:30:44–01:31:12]

**EN**  you know this plot shows that it's getting very problematic um and so you know claude mythos was released at this dotted line where you know most people they're not sure if it's because of claude mythos that these vulnerabilities are increasing most likely it's because open source you know we use lots of models call them many many many many times and we can you know automatically find exploits in these models but you know there is actually another

**中文**  Claude Mythos 在 dotted line 处发布。人们不确定 vulnerability 增加是否由它造成；更可能是大家大量调用模型，从而自动发现更多 open source exploit。但还有另一

### [01:31:10–01:31:40]

**EN**  point so in hacken news someone posted about this um is it just mythos and juby 5.6 six that do good on finding cyber security issues. It's not actually open source models also do very well. Um open source models do extremely well in finding cyber security threats and issues. Um you know there is some discussion on hacker news you know is this actually true or false? Um but you know according to some you know some

**中文**  点。Hacker News 有人问：只有 Mythos 与 GPT-5.6 擅长发现 cyber security issue 吗？其实 open source model 也非常擅长发现 security threat 与 issue。讨论中有人质疑真假，但按部分

### [01:31:37–01:32:06]

**EN**  researchers and cyber security people the main reason why you know mythos looked like it was very good on cyber security is because they bothered to actually check the open source code. Um and so if you actually give the open source models the full code base of these open source libraries they will find the bugs you know they will find cyber security issues. Um and all you need to do is core the model. Um and so I feel like you know that's the

**中文**  researcher 和 cyber security 从业者的说法，Mythos 看起来特别强，主要因为他们真的去检查了 open source code。若把这些 library 的完整 codebase 交给 open source model，也能找到 bug 和 cyber security issue；只需调用模型。所以我认为

### [01:32:03–01:32:31]

**EN**  fundamental problem is mythos seems very powerful not because the model is powerful but because they actually bothered to test on all open source repos. Um and so if you do you know if you call all these open source models to detect for bugs for cyber security issues you will find bugs and you know as as you know recently you know as everyone knows fable is still

**中文**  根本问题是，Mythos 的强大也许不只来自 model capability，而是有人实际测试了所有 open source repo。调用 open source model 检测 bug 与 security issue，一样会找到。最近大家也知道，Fable 仍对

### [01:32:28–01:32:57]

**EN**  banned for the majority of everyone. Um, and GBD 5.6, you know, is delayed a staggered release, right? So like Guby 5.6 preview was on Friday, right? So like a few days ago, and they said they're not going to be releasing to everyone. Um, and the main questions are, you know, in the open source world, in the clos world, people are asking, do we need a license to use these AI models for everyone? You know, like everyone in this room now, we have to have a license

**中文**  大多数人禁用；GPT-5.6 则 staggered release，周五先发布 preview，并没有立刻开放给所有人。因此 open 与 closed source 社区都在问：每个人使用 AI model 是否需要像 driver license 一样取得 license？

### [01:32:55–01:33:24]

**EN**  to use the models, like a driver's license. Um, do we need to get that? um is there going to be a delay in all of these releases? So every single time when a new model gets released only the trusted providers get these models. Um the next most important question how about open-source models you know okay the government the US government currently is like you know trying to like control Fable GPD 5.6

**中文**  所有新模型是否都要 delayed release，先只给 trusted provider？更重要的是 open-source model 怎么办？美国政府现在试图控制 Fable、GPT-5.6，

### [01:33:22–01:33:50]

**EN**  The main question now is what do we do about open source models? You know, open models, open weight models. What will the government do to control the open source space? To be completely honest, I was quite surprised the government acted this early um in doing GBD 5.6 and fable control, right? I thought it was like maybe the end of the year or next year, but it seems like it's now. Um so the next question is what will happen to open source models? Will the government

**中文**  接下来会怎样控制 open model 或 open-weight model？老实说，我很惊讶政府这么早就对 GPT-5.6 和 Fable 采取控制；原以为会到年底或明年，没想到就是现在。所以下一个问题是，政府会不会

### [01:33:47–01:34:16]

**EN**  start controlling open source models? Um and the fundamental question is what defines frontier intelligence like the reason why the government is you know they're controlling these models is because they're very very powerful. Um so the main question is what actually defines intelligence? You know which benchmark do we use? Is it just based on one trillion parameters like you know how do we define whether a model can be banned or unbanned? Um and that is a

**中文**  开始管制 open source model？根本问题是怎样定义 frontier intelligence。政府控制模型是因为认为它们极其强大，但 intelligence 用什么界定？采用哪个 benchmark？只看一万亿 parameter 吗？怎样决定一个模型该禁或不禁？这是

### [01:34:13–01:34:43]

**EN**  very very important question. Um and we will we have a dark web of open models now you know do we need to torrent open models? Um and the most important question what is inference what are inference providers going to do now you know assuming assuming that the government has some sort of regulation on even open models. Um what is the inference prov what are they going to do? You know what are the inference providers going to do? Do they need to have license? Do they need to check that

**中文**  非常重要的问题。未来会不会出现 open model 的 dark web，必须用 torrent 下载？inference provider 又怎么办？假设政府连 open model 都监管，provider 是否需要 license？是否必须

### [01:34:41–01:35:09]

**EN**  everyone has a license before you can use the model? Um or something like that. Um and so like you know these are very important questions that you know the government is currently like you know and the industry you know the entire AI ecosystem and industry we are trying to like you know what are the answers to these questions. Um and obviously you know if you were the government if I was the government it makes sense you know they do not want their critical infrastructure to be hacked. You know, remember open- source

**中文**  先验证每个用户持有 license 才能调用？这些都是政府、AI ecosystem 与整个 industry 正在寻找答案的重要问题。当然，站在政府立场也可以理解：他们不希望 critical infrastructure 被 hack。记住，open-source

### [01:35:07–01:35:36]

**EN**  exploits are skyrocketing. If you change that y-axis, you know, not open source exploits, but like critical infrastructure exploits, you know, obviously the government's scared. Um, so it makes sense for them to like stagger the release. Um, but the main question is, you know, we're still in this we're still in this fog of war type approach, you know. Okay, not fog of war, just fog, a foggy, you know, we don't know what will happen for regulation. Um, yeah, that's very

**中文**  exploit 正急剧上升。若 y-axis 改成 critical-infrastructure exploit，政府害怕很合理，staggered release 也有道理。但 regulation 仍在一片迷雾中，不知道会发生什么，这很

### [01:35:32–01:35:49]

**EN**  problematic. Um yeah. Oh okay. Anyone have any questions for cyber security regulation policy or whatever? Um or any takes as well questions? >> Yes.

**中文**  棘手。有人对 cyber security、regulation 或 policy 有问题或看法吗？

### [01:36:02–01:36:31]

**EN**  That is a good question. So is it is open source so the scare of open source models is it because you know anthropic keeps screaming about open source is bad you know every single day open source is bad um yes and no I feel like it's true that you know there are some players in the closed source industry they want to shut down the open source ecosystem their view is if you give open source to

**中文**  观众问：对 open source model 的恐慌，是否因为 Anthropic 每天都在强调 open source 很危险？嘉宾：既是也不是。closed source industry 的一些 player 确实想关闭 open source ecosystem，认为把 open source 交给

### [01:36:28–01:36:56]

**EN**  anyone they will start hacking you know critical infrastructure they will start doing bad behavior. Um, and so that's kind of their view. Um, so yes, I agree that some of the closed source labs have caused this problem. Um, but it's actually kind of funny because currently the government is regulating them first and open source is still a question mark. Um, and so like it's kind of like I don't know, they probably stabbed themselves in the foot or something. I

**中文**  所有人，就会有人攻击 critical infrastructure、做坏事。所以我同意一些 closed source lab 促成了这个问题。但有趣的是，政府现在先监管了他们，open source 反而还是问号；也许他们搬起石头砸了自己的脚。

### [01:36:54–01:37:24]

**EN**  don't know whatever whatever the phrase is. Um but I feel like it's they did cause some controversy in terms of like saying open source is bad but in general open source models are actually good. Um so you could I mean theoretically you can use an open source model and you know run this on all repos and you will be able to find exploits and you can exploit. So they're not wrong. Um but I

**中文**  他们确实制造了“open source 很坏”的争议，但 open source model 总体仍是好东西。理论上，你的确能用它遍历所有 repo，发现 exploit 并利用，所以担忧并非完全错误。不过

### [01:37:22–01:37:52]

**EN**  feel like you know who has the infrastructure to do this? um you know, GitHub might automatically detect you and ban you or something. I don't know. There's many there's many um layers of security for each section. Um and so like I I don't know. I I feel like it's somewhat overblown, but it is it is it's not 0% probability. So it is a problem. Um yeah, if that answers your question, but okay. Yes. So now we're going to be

**中文**  谁真正具备这样做的 infrastructure？GitHub 也许会自动检测并 ban，还有多层 security。风险可能被夸大，但 probability 并非 0。下面谈 kernel。

### [01:37:49–01:38:18]

**EN**  talking about kernels. Um so previously you know this is my favorite plot as usual. Um you know if we were in a different future you know if we were in a different timeline that we did not discover 01 preview models would have plateaued. I think that's the fundamental point of this plot. It shows that if we have never discovered reasoning we have never discovered 01 whatever we will have plateaued. We will have

**中文**  回到我最喜欢的图：若在另一个 timeline，我们没发现 o1-preview，模型会 plateau。这是图的根本含义：如果从未发现 reasoning 或 o1，accuracy 会

### [01:38:16–01:38:44]

**EN**  plateaued in terms of accuracy. Um and that is not good. Um and because we have discovered this new paradigm of scaling, you know, models have continuously scaled even further. Um but my take is the reason why we have stopped scaling um based on you know the old approach is because the old approach only focused on hardware optimizations. We now have to move over to software optimizations and algorithmic

**中文**  停滞，这当然不好。发现新的 scaling paradigm 后，模型才继续 scale。我的判断是，旧 approach 停止增长，是因为只关注 hardware optimization；现在必须转向 software 与 algorithmic

### [01:38:41–01:39:09]

**EN**  optimizations. Um we you know we need to have new inventions of how do we scale AI even further. Um and we can't just rely on doing 10 trillion parameters or you know making the model bigger and bigger bigger. Um for example you know we have to do float a reinforcement learning. Um so if pytorch has this methodology where you can do floatate float for different precisions to make training faster. Um and that is one way. Um another way for example as a software approach for example as I

**中文**  optimization，发明新的 AI scaling 方法，不能只靠十万亿 parameter 或不断放大模型。例如 PyTorch 有 Float8 training，可用不同 precision 加快训练，这是一种方式。另一种 software approach 是我们前面提到的

### [01:39:08–01:39:37]

**EN**  previously said we found some bit you know issues in gradient accumulation. Um so when you do gradient accumulation um it was actually it was not calculated correctly during the loss calculation. Um and you can actually increase accuracy by 1 to 3% if you fix this small little issue. Yeah. So like you know the universal gradient accumulation bug fix was a software fix. It is not a hardware fix. Um and so the fundamental view is you need to do more and more software

**中文**  gradient accumulation bug：loss calculation 计算不正确，修好这个小问题就能提高 1% 到 3% accuracy。这个通用 fix 是 software fix，不是 hardware fix。核心是需要越来越多 software

### [01:39:35–01:40:02]

**EN**  changes. Right? Another one for example Snowflake we collaborated them to make context long context fine-tuning 500k contact length this was all software improvements um another one is you know 12 times faster MLB training this is another software improvement um deepseeek you know they released something called deep spark which was just a few days um and they showed that they can make inference you know 50 50

**中文**  change。我们还与 Snowflake 合作，用 software improvement 实现 500K context-length fine-tuning；MLP training 加速 12 倍也是 software improvement。DeepSeek 几天前发布 DeepGEMM（英文自动字幕疑似误写），让 inference 比普通 MTP 快 50% 到

### [01:39:59–01:40:28]

**EN**  to 600% faster so six times faster than just normal MTP um and so this is a software ware methodology, right? Not a hardware methodology. And you know, diffusion Gemma, right? Gemma released a new diffusion model showcasing that you can get 2,000 tokens per second by using a new architecture, right? So using diffusion LLMs to do faster inference and again this is a

**中文**  600%，最多六倍，也是 software methodology。Diffusion Gemma 通过 diffusion LLM 新 architecture 达到每秒 2,000 token，同样是 software change。我的核心观点是，hardware innovation 越来越不重要，

### [01:40:24–01:40:52]

**EN**  software change. And my main point is is that in general, hardware innovations are getting less and less important. Um and hardware innovations are actually slowing down. Um so it's actually kind of interesting intelligence you know the scaling of you know intelligence in general it's kind of like Mo's law um it's kind of like a there is a relentless progress relentless approach

**中文**  而且正在放缓。intelligence scaling 有点像 Moore's law，都表现为 relentless progress。

### [01:40:50–01:41:18]

**EN**  to increase intelligence and the same with Mo's law um and so like in general you can see that you this is Mo's law over here the number of transistors has continuously increased um but you know single performance is not increasing it has staggered Um and so this is kind of like you know this kind of reminds me of you know this plot right scaling intelligence in terms of parameters probably has plateaued most likely you

**中文**  Moore's law 图中 transistor count 持续增加，但 single-thread performance 已停滞。这让人想起 intelligence scaling：parameter、hardware performance、pre-training 大概也 plateau，必须进入 reasoning paradigm 才能继续。

### [01:41:17–01:41:46]

**EN**  know hardware performance pre-training whatever we now need to go into this new reasoning paradigm to scale even further um so it kind of is like similar to the moors law type graph um kind of um and you can see if you see on this side the number representation of GP so my GPU is getting faster and faster and faster, right? It's not actually the GPU itself that's getting faster and faster and faster. Um, it's the represent number representation, right? So like they

**中文**  再看 GPU number representation：GPU 看似越来越快，其实并非芯片本身不断加速，而是 numeric representation 从 float32 改到 float4，让 GPU 快了 32 倍。

### [01:41:44–01:42:13]

**EN**  changed from float 32 all the way to float 4. Um, and this made GPUs 32 times faster. Um, so it's not eight times faster, right? It's not 32 divided by 4 is eight times faster. It's 32 times faster. And the reason why is because of tensores, um, you know, the smaller mantesses, um, and so on. Um, and so like you can actually see, you know, even tensors with the introduction of tensores, it made the GPUs 12 times faster

**中文**  这不是简单的 32÷4=8 倍，而是 32 倍，原因涉及 Tensor Core、更小的 mantissa 等。引入 Tensor Core 也让 GPU 快约 12 倍。

### [01:42:11–01:42:39]

**EN**  and so on. Actually, if you made the GPU smaller and smaller and smaller, it only made it three times faster. It's not even that important anymore. Um, and if you look at this plot, we are now at float 4. So most of the GPUs that we have now are at float four. What is next? Are we going to be having float three, float two, float one? Are we going to have float zero? Okay, no such thing. But anyways, the point is

**中文**  单纯把 GPU 制程越做越小，只贡献大约三倍，已不再那么重要。如今 GPU 已做到 float4，下一步是什么？float3、float2、float1，甚至不存在的 float0？核心是

### [01:42:37–01:43:04]

**EN**  hardware is kind of at its limits, right? We're already at float 4. What is next? There is nothing next. Um, and so the so the answer to this question is there is nothing next. Um, and so now we need to move over to software, right? How do we make new algorithms? How do we make new methodologies to continue scaling? Um, I also made this table, right? I previously said why is you know you use float 32 um we change it to

**中文**  hardware 已接近 limit。float4 之后几乎没有明显下一步，因此必须转向 software，发明新 algorithm 与 methodology 继续 scaling。我还做了张表解释 float32 到

### [01:43:01–01:43:29]

**EN**  float 4 why is it why is it not eight times faster um and instead it's 32 times faster right why is it 32 times faster and the reason is because when you use when you do floatingoint precision you have an exponent and a menta um and the transistor space the transistor space is the exponent plus the menta squared um and so the trick is if you make them in Tesla smaller and

**中文**  float4 为什么不只是八倍，而是约 32 倍：floating-point precision 包含 exponent 与 mantissa，所需 transistor space 与它们的组合近似平方相关；把 mantissa 缩小会

### [01:43:27–01:43:54]

**EN**  smaller and smaller, you square their number of improvements, right? So, float 32, float 32, you needed 537 transistors around, right? 537 transistors. Um, to go from float 32 to float 16, you only need 105 transistors. So, actually you made you made the number of transistors five times more, right? So, not two times,

**中文**  平方级减少 transistor 数。float32 约需 537 个 transistor；float16 只需 105 个，因此不是少两倍，而是少约五倍。

### [01:43:51–01:44:19]

**EN**  it's five times. Um, and so on so on so on. Um, so you know, I guess you can go to 1.58 bit. I guess you can do that. Um, but it's actually kind of interesting because 1.58 bit um is actually not that much faster. Um, so 1.58 bit is actually not that much faster um than float 8. Um, if you use um, you know, seven seven exponent and Messa 2. Um, there is another 1.58 bit

**中文**  以此类推。也可以做 1.58-bit，但它其实没有比 float8 快很多；另一个 1.58-bit 方案使用 float4。float4 比 float32 快 179 倍。

### [01:44:16–01:44:42]

**EN**  which you use float 4. Um, so float 4 is 179 times faster than float 32. Um and the main question is we are already at three transistors right we are already around three transistors what are we going to do next two transistors or like one transistor so don't like you know most likely GPUs are not going to be getting faster um that's the fundamental question of this plot so GPUs are not going to be getting faster

**中文**  问题是现在已经降到约三个 transistor，下一步是两个还是一个？GPU 大概率不会继续大幅变快，这是图的核心。因此应该关注 kernel：怎样做更好的 kernel、algorithm 与 scaling，而不是继续只做 hardware optimization。

### [01:44:42–01:45:11]

**EN**  instead we need to focus on kernels right how do we make better kernels better algorithms how do we scale this instead right don't do don't do hardware optimizations anymore. Instead, how do we do, you know, these optimizations? Um, and so one of my favorite tools to use, you know, everyone should use this is just use torch compile. Um, so in my, you know, it's the modern, you know, the

**中文**  我最喜欢、也建议所有人使用的工具是 torch.compile。到了现代，不要先学写 custom kernel。我的建议是不要做 kernel writing，因为 torch.compile 会接管大部分工作。

### [01:45:06–01:45:36]

**EN**  modern time, do not, as advice, do not learn how to write custom kernels. That is advice. Do not do kernel writing. Um and the reason why is because torch compile will take over all of kernel writing. Um so you can see for example this plot um torch compile was a red line right performance it doesn't look like it's doing very well right it does not look like it's doing very well versus handwritten kernels right

**中文**  图里旧版 torch.compile 是红线，看起来不如 handwritten kernel；但那是旧 PyTorch。新版 torch.compile 的橙线则明显获胜。

### [01:45:34–01:46:03]

**EN**  handwritten kernels are the other ones right so torch compile doesn't look like it's doing very well but that's because that's an old PyTorch version if you have a newer PyTorch version torch compile wins dramatically right that's the orange line um and all of these are handwritten kernels. Oh, the okay the black line is torch compile plus no fusion. Um so that's another torch compile method. Um but the red line, the green line and the blue line okay the

**中文**  其他线代表 handwritten kernel、torch.compile without fusion 或普通 PyTorch。新版 torch.compile 的表现更强。

### [01:46:01–01:46:29]

**EN**  the blue line is just no torch compile just normal PyTorch. Um but the green line and the black the green line and the red line are handwritten kernels and you can see it does even worse than torch compile. So like my viewers like what's the point of writing kernels? Torch compile does even better than you. Um so the main point is you should always firstly look at torch compile right before you write a kernel use

**中文**  绿色和红色是 handwritten kernel，却比 torch.compile 更差。既然 compiler 比手写更好，为什么还写 kernel？所以在写 kernel 前，首先试 torch.compile。

### [01:46:27–01:46:55]

**EN**  torch compile first do not start learning how to do triton or you know cuda or whatever is your favorite coding language for kernels don't do that instead use torch compile um even worse like you know this this was RMS norm um you know this is layer norm torch compile wins dramatically um you know versus handwritten kernels. So I would not you know definitely only use

**中文**  不要一开始就学 Triton、CUDA 或其他 kernel language，先用 torch.compile。RMSNorm 如此，LayerNorm 图中 torch.compile 也远胜 handwritten kernel。因此我强烈建议优先使用 compiler。

### [01:46:52–01:47:22]

**EN**  torch compile as your first try. Um do not write kernels first. Use torch compile. So the main takeaway is algorithms are much more important than hardware or whatever handwritten kernels. Right? Remember deepseek released deep you know deep spark you know there's other algorithms for speculative decoding like MTP D flash DSpark whatever all of these are algorithmic

**中文**  torch.compile 应当是第一选择，不要先写 kernel。核心 takeaway 是 algorithm 比 hardware 或 handwritten kernel 更重要。DeepSeek 的相关新算法、speculative decoding 的 MTP、DFlash、DeepGEMM 等都是 algorithmic

### [01:47:20–01:47:49]

**EN**  improvements and these made inference two times to six times faster right it wasn't like new some new hardware it wasn't some new hardware which made inference faster it was algorithms which made inference faster right flash attention flash you know FA2 FA3 three, flash attention four, flash attention five, six, seven, whatever, right? All of these are algorithmic improvements, right? Flash attention was essentially a

**中文**  improvement，让 inference 快两到六倍；不是新 hardware，而是 algorithm。FlashAttention、FA2、FA3，以及未来版本，全是 algorithmic improvement。FlashAttention 本质上是

### [01:47:46–01:48:15]

**EN**  trick to do memory movement much better. Um, so how do we like orchestrate memory movement and use the caching structure of the GPUs much better? Um, and so flash attention is also a algorithm. Gradient checkpointing, you know, one of the most important algorithms for training is gradient checkpointing. Um, and all it does is you do not save all the activations. You do a trick where you only save the activations for every single layer. Um, and then you skip all

**中文**  更有效地移动 memory，协调 memory movement 并利用 GPU cache hierarchy。gradient checkpointing 是训练中最重要的 algorithm 之一：不保存所有 activation，只保存每一层边界的 activation，跳过层内

### [01:48:13–01:48:43]

**EN**  the intermediate activations in each layer. Um, and then you recomp compute the activations. Um, and gradient checkpointing saves memory by dramatic amounts by like 70%. Um, 70% memory reduction with no change in accuracy. And okay, training is a little bit slower maybe by 10% to 15%. Um and you know grading check checkpointing was an algorithm. Um and you know like in general you should also try to understand you know what is the new data

**中文**  intermediate activation，之后重新 compute。它能减少约 70% memory，accuracy 不变，只让 training 慢约 10% 到 15%。gradient checkpointing 也是 algorithm。还应研究新的 data

### [01:48:41–01:49:09]

**EN**  processing tricks you know how do we like you know stagger data you know do we do do we do curriculum learning or something like that I don't know um what you know how do we clean the data set before we actually pre-train the model there are many tricks you can employ for data processing um and obviously you know there is still a group of people I don't know I take an opposite view there is a group of people who think mega kernels are the latest and greatest for kernels. You know what is a mega

**中文**  processing trick：怎样 stagger data、是否做 curriculum learning、pre-train 前怎样清洗 dataset，都有很多方法。也有一群人认为 mega-kernel 是 kernel 的最新方向，我持相反看法。什么是 mega-

### [01:49:07–01:49:35]

**EN**  kernel? A mega kernel is when you take an entire you take an entire implementation of a model and it's just one kernel like one large kernel. Um ah maybe it's useful who knows um you know Nvidia you know Nvidia has a acquired Grock or something um and you know their view is for example you have two different systems right the LPU which is the Grock system does the

**中文**  kernel？就是把一个 model 的完整 implementation 做成单个巨大 kernel。也许有用，谁知道。NVIDIA 收购了 Groq（自动字幕语境），他们设想两类 system：LPU，也就是 Groq system，负责

### [01:49:33–01:50:02]

**EN**  decoding right so like the MLP layers the layers does the decoding and then the GPU so the Nvidia GPUs does the attention and the prefield um and so in general you know we might even have a future We have different types of hardware systems you know we have asex which are you know specially designed chips for you know computation and we have generalized systems like GPUs. Um and these asex and GPUs will collaborate

**中文**  decoding 和 MLP layer；NVIDIA GPU 负责 attention 与 prefill。未来可能由多种 hardware 协作：专门计算的 ASIC 与 general-purpose GPU 相互配合。

### [01:49:59–01:50:28]

**EN**  with each other. Um so for example the attention you know the the attention will be for the GPUs and they will transfer over to the LPU to do you know the MLPS and so on. Um and then this is like a dance you know between them and you can also do like pipelining right you can imagine that there's like many many replicas of this and they can like you know serve you know 20 people or you know 1,000 people in one go. Um

**中文**  例如 GPU 做 attention，再把数据交给 LPU 做 MLP，二者像跳舞一样配合。还可 pipeline 多个 replica，一次服务二十或一千人。

### [01:50:26–01:50:55]

**EN**  and yeah so this is like another approach and you know this in my view this is kind of an opposite approach of mega kernels. So as a mega kernel your view is you want to combine the goal is to make the goal of a mega kernel is to make one kernel for the full forward path of a language model. Um and once you make one once you once you are able to make the language model the forward path into one kernel you can now make

**中文**  我认为这与 mega-kernel 是相反路线。mega-kernel 的目标是把 language model 整个 forward path 合成一个 kernel。做到单层 forward path 后，还可以把 32 层整个 model 都变成一个 kernel。

### [01:50:52–01:51:21]

**EN**  the entire language model with 32 layers as one kernel. Right? You can extend this and because the whole language model is one kernel you can even further extend it. Right? the prediction of the second token, the third token, the fourth token, the sixth token can all be just one kernel. Um, and unfortunately, this is very hard to do. It's very hard because attention is the problem, right? Attention has to see the tokens in the future. See the tokens in the past, not

**中文**  甚至把预测第二、第三、第四、第六个 token 都塞进单一 kernel。但这非常难，问题在 attention：它必须看到过去 token，而不能看到未来，后者属于 cheating。

### [01:51:20–01:51:49]

**EN**  the future. That's cheating. Um, you have to see the tokens in the past. And that is a fundamental problem. Um and it's very hard to you know it's very hard to make a mega kernel to combine attention and the MLE or MLP layers. It's extremely complicated. So in general what what people do is they will make two kernels right one kernel for the attention part and the other kernel for the rest. Um and so you will see there are two kernels. Um and yeah so

**中文**  因此很难让一个 mega-kernel 同时组合 attention 与 MLP layer。通常会拆成两个 kernel：一个负责 attention，另一个负责其余部分。做成一个很难，两个则可行。

### [01:51:47–01:52:15]

**EN**  it's very hard to make one mega kernel but you can make two kernels. Yes. Okay. Any other questions? Any questions for kernels? So the main takeaway for Yes, a question. Yes, that is a very good question. So the question was because there's so many knobs for torch compile like 1,00 or something. How do we reduce the experimentation time to like you know find which knob is the best? Um so luckily we have something called bisection or binary

**中文**  有人问：torch.compile 有上千个 knob，怎样减少 experimentation time，找出最佳配置？幸运的是可以用 bisection 或 binary search。

### [01:52:14–01:52:43]

**EN**  search. That's the trick. So what we'll do is instead of checking every single 10,00 combination randomly sample. So random you do randomize bisection. You randomly sample 50% of the, you know, flags. You turn it on versus turning it off and then benchmark which one is better. And whichever one is better, you then narrow down the search. You again do 50% and 50% and 50% and 50%. So it's actually log two of 10,00. I don't know

**中文**  不用检查一万种 combination。可随机抽取 50% flag，比较开启与关闭的 benchmark，保留更好一半，再不断按 50% 缩小 search。所需步骤只是 log2(10000)。

### [01:52:41–01:53:08]

**EN**  what that what is log 2 of 10,000. I don't know what that is. Um two times I don't know. Anyways, log 200. I think you need to do 30 steps, I think. I don't know. I don't remember whatever 2 to the^ of something is equal to 1,00 then log it. Um so you you only need to do you don't need to do you don't need to check all 1,00 knobs. You only need to check a few steps and then you will know which flag is the best. Um

**中文**  我不记得 log2(10000) 是多少，也许十几步；总之无需检查所有 knob，只要少量 step 就能知道哪些 flag 最好。诀窍就是 binary search 或 bisection。

### [01:53:07–01:53:23]

**EN**  so the trick is to use binary search or bisection to do this approach. Um yeah. Any other questions? >> Yes. >> What are your thoughts on

**中文**  还有问题吗？观众提问。

### [01:53:39–01:54:07]

**EN**  So your question was what do I think about asex like you know cerebras gro sanova I don't know even startups new chips they do design their own chips I I feel like so the problem of ASEX is is it AS6 or ASX or whatever the problem of specialized chips is the architecture itself needs to be hardcoded in some of the chips and that

**中文**  问题是怎样看 ASIC，例如 Cerebras、Groq、SambaNova 等自行设计 chip 的 startup。specialized chip 的问题在于，architecture 需要 hardcode 到 chip 中。

### [01:54:05–01:54:34]

**EN**  is the problem. If you hardcode some of the chips you know hard code the infra hardcode the architecture labs always like to change the architecture and so every single time when the lab changes the architecture you need to update the chip. Um but as a GPU the G the trick of GPUs is Nvidia has made it you know Nvidia, AMD, Intel whatever the GPU is extremely powerful because it has

**中文**  一旦 hardcode infrastructure 或 architecture，而 lab 又频繁改 architecture，每次变化都要更新 chip。GPU 的诀窍则是 NVIDIA、AMD、Intel 把许多 generalized ASIC 放进 GPU，使其极其灵活。

### [01:54:30–01:54:58]

**EN**  generalized asex inside of the GPU right the GPU is in fact a combination of ASEX. Um and the ASICH is just one large ASICH. Um so I think like in general a GPU is much better because you can customize what goes inside the GPU. you can disable stuff that goes inside the GPU and such. Yeah. So on so my view is I don't know I don't I don't want to say anything but like in general I don't

**中文**  GPU 其实是许多 ASIC 的组合，而单独 ASIC 只是一个大型专用单元。我认为 GPU 更好，因为可以 customize 内部功能，选择 disable 某些部分。

### [01:54:56–01:55:25]

**EN**  think like you know previously as I mentioned you know hardware there is nowhere else to go you know we are at float four unless if the hardware providers invents float I don't know float zero then maybe we get four you know another four times faster but in general I think I think just people are focused too much on hardware and they have not looked that actually the biggest improvements is not hardware, it's software, right?

**中文**  如前所述，hardware 已无太多余地，已经做到 float4，除非发明 float0 才可能再快四倍。人们太关注 hardware，却忽略最大提升来自 software。

### [01:55:22–01:55:52]

**EN**  Numerical precision, numerical precision was 32 times faster. Hardware is only three times faster, right? So hardware only contributed three times faster. Um oh actually d okay the die size you make the you make the GPU bigger, you get two times faster. That's that's kind of cheating. So I wouldn't really say that's improvement. Um but essentially if you make the hardware faster, you only get three times faster. Um so in my view hardware is probably overblown. you know hardware is actually not that

**中文**  numeric precision 带来 32 倍加速，hardware 本身只有三倍。扩大 die size 可再快两倍，但那有点 cheating。总体 hardware 可能被夸大，并没有那么重要。

### [01:55:50–01:56:18]

**EN**  important. The software was the trick that Nvidia, you know, Nvidia, AMD, Intel, all of these, you know, hardware providers, they banked on the fact that numerical precision was a trick and tensor, tensor, numerical precision, sparsity, you know, these software tricks. Um, okay. Well, tensor is not really software trick, but you know, a tensor is kind of an ASICH inside of the GPU. Um, and so like I feel like that's Yeah. So my view is I don't I don't

**中文**  NVIDIA、AMD、Intel 等 hardware provider 真正押注的是 numeric precision、Tensor Core、sparsity 等 software/architecture trick。Tensor Core 不完全是 software，而像 GPU 内部的 ASIC。

### [01:56:17–01:56:46]

**EN**  really see a future for ASEX. That's my view. Um I think that ASEX are like instead, you know, to be honest, I'm actually quite surprised. We have lots of asset companies, but we have very few algorithm companies. Um and the reason why is because ASEX you can sell, right? Every single year you can upgrade. You know, this year you pay $1,000 to ASICH version one and the next year you have to upgrade, right? The problem with

**中文**  所以我看不到 standalone ASIC 的广阔未来。让我惊讶的是 ASIC company 很多，algorithm company 却很少。原因是 ASIC 容易卖：今年付一千美元买 version 1，明年又要升级。

### [01:56:43–01:57:12]

**EN**  algorithms is algorithms is very hard to you know force the user or whatever to pay again and so that is why hardware is very popular because hardware is a very easy business model but for algorithms it gets more complicated right how are we going to monetize grading checkpointing I don't know right that's very hard um so but the main point is the large labs themselves I think like open air announced collaboration broadcom and suras whatever you know

**中文**  algorithm 很难让 user 重复付费。例如 gradient checkpointing 怎样 monetization？很困难。但 large lab 正直接与 hardware provider 合作设计 chip，例如 OpenAI 与 Broadcom、Cerebras 等。

### [01:57:09–01:57:32]

**EN**  each lab themselves are going to the hardware provider and designing the chip with them. Um, so my view is like maybe we'll have more of these like collaboration approaches, but I feel like standalone standalone assets, I don't think they're going to last. Um, yeah, that's my take, I guess. Any other questions? Yes.

**中文**  未来也许更多是这种 collaboration；我不认为 standalone ASIC company 能长久。还有问题吗？

### [01:57:46–01:57:51]

**EN**  Oh, you mean what are the types of kernels or

**中文**  观众澄清想问的是 kernel 类型。

### [01:57:55–01:58:25]

**EN**  >> Oh, okay. Okay. Okay. Um, so the question was what are the you know what are the changes for kernels or optimizations or stuff that is like interesting I guess for kernels. Um so most kernels when you write kernels the majority of them are are focused on memory movement reduction. How do we reduce memory movement? That's the majority of kernels. Um for example there's there's a trick called um you know there's a trick called fuse cross entropy loss where instead of making

**中文**  问题是哪些 kernel optimization 有意思。绝大多数 kernel 都聚焦减少 memory movement。例如 fused cross-entropy loss：不 materialize 最后一层的完整 logits，而是分 batch、逐 row materialize。

### [01:58:23–01:58:53]

**EN**  instead of the last layer of cont instead of materializing the full logs there is a trick you can do it in batches right you can do rowby row materialization. Um and so this will reduce memory by like a lot by like I don't know 10 GB or something if you have long context or even more. Um that's one way. Um the other kernels most kernels are called kernel fusion where you you have this like long pietorch function and all you do is you just write one

**中文**  在 long context 下可减少约 10 GB 甚至更多 memory。另一类是 kernel fusion：把一长段 PyTorch function 合成一个 kernel，torch.compile 很擅长。

### [01:58:51–01:59:19]

**EN**  kernel to do this whole pietorch function and torch compile will do this for you. So torch compile is very very good at doing kernel fusion right you give torch compile a function it will write a kernel a try kernel or whatever kernel and it will just fuse everything um it's very very effective for that um but I think in general kernels are just reducing memory movement um and so like I you know to be honest I don't really

**中文**  把 function 交给 torch.compile，它会编写 Triton kernel 或其他 kernel 并 fuse 所有操作，效果很好。总体上 kernel 主要做 memory-movement reduction。

### [01:59:17–01:59:45]

**EN**  like to call it kernels most algorithms so most algorithms you either make training faster or reduce memory usage um but kernels in my view kernels is reduce memory movement. Um, and so most kernels is just memory movement, you know, memory movement optimization, right? How do we how do we use the caching structure of the GPUs? Um, you know, how do we not load the same variable twice or three times or whatever? Um, yeah, I'm not sure if that

**中文**  我甚至不太喜欢把它称为 kernel。algorithm 通常让 training 更快或 memory usage 更低；kernel 的核心是优化 memory movement、利用 GPU cache hierarchy，避免重复 load 同一 variable。下面讲 reinforcement learning。

### [01:59:43–02:00:13]

**EN**  answer your question, but next reinforcement learning. Um, and after this will be reward hacking the agents. Um, so as a primer to I'm assuming most people know, do most people know reinforcement learning or do I need to prime people? Okay, I'll give a very fast primer for reinforcement learning. Okay, fast primer for reinforcement learning. Um, what is reinforcement learning? You have this environment such as this Pac-Man game. Um, and your goal

**中文**  之后会讲 reward hacking 与 agent。先快速介绍 reinforcement learning。这里有 Pac-Man game 这样的 environment，player 的 goal 是 maximize reward。

### [02:00:09–02:00:39]

**EN**  is as the player, you know, to maximize reward. You want to eat all of the cookies, right? You want to eat all of the cookies, but also escape away from your the monsters. I don't actually know what they're called. Enemies, monsters, whatever, whatever they're called. Um, and your goal as Pac-Man is you want to maximize the amount of cookies that you eat. Um, and that is your reward. The reward is the cookies. Um, and the action is whether you go up, left, down,

**中文**  要吃掉所有 cookie，同时躲开 monster 或 enemy。Pac-Man 要 maximize 吃到的 cookie 数量，cookie 就是 reward；action 是上、左、下、右，environment 是 game。

### [02:00:36–02:01:05]

**EN**  or right. Um, and the environment is the game. Another good example, you know, another way I like to explain reinforcement learning is the goal of reinforcement learning is you want to have more good and less bad um during training. Um so for example, at the very beginning of training, you ask the model what is 2 plus two? Um the answer is clearly four. Um but when the model starts training, it will be very dumb.

**中文**  另一种解释是：reinforcement learning 的目标是在 training 中增加 good、减少 bad。比如问模型 2+2，答案显然是 4，但训练初期模型很笨。

### [02:01:03–02:01:30]

**EN**  It will be very bad. it will see B, you know, the model would just say B, D, cat, dog, house, mouse, whatever. Um, and the trick is for all of the bad responses, you want to decrease, you don't want you want to like negatively reward this or penalize it. You want to penalize the model if it says something bad and you want to increase the reward if it says the correct answer. Um, so that is the trick of reinforcement

**中文**  它可能回答 B、D、cat、dog、house、mouse。对所有 bad response 要降低 reward 或施加 penalty；回答正确则提高 reward。这就是 reinforcement

### [02:01:28–02:01:57]

**EN**  learning. You just want more good answers, less bad answers. Um, and if it's like, you know, very close to the correct answer. So, you know, three is very close to the correct answer. Um, you want to reward, you want to negatively reward this a little bit less, right? Because three is much closer to four than B or D. Um, so if you do B or D, you want to negatively reward it massively. In reinforcement learning, the trick is

**中文**  learning：更多好答案、更少坏答案。若答案很接近，例如 3 接近 4，可以少罚一点；B 或 D 则应重罚。关键是

### [02:01:55–02:02:24]

**EN**  you have a verification system, right? You have a verifier to verify if the model is doing good or bad. Um so you'll call the model many many many times. Um and each of these examples you give a verification number right? So for example um the first example is very good so you give it a plus 10 score. Um the next example is like okay so you give it a minus5 score. Um and then the last example is very bad so you give it

**中文**  有 verification system 判断模型好坏。调用模型许多次，为每个 example 分配 score：很好就 +10，一般可能 -5，非常差则 -100。

### [02:02:21–02:02:50]

**EN**  a minus 100 score. Um and reinforcement learning allows you to assign scores to each of those answers and questions. Um and so that's kind of reinforcement learning verifies. Um and the trick of reinforcement learning is my favorite phrase is patience is all you need. Um at the very beginning of training your model will do very bad, right? Your your reward will be 0000

**中文**  reinforcement learning 允许为答案分配 score。我的一句话是“patience is all you need”：training 初期 reward 会长时间是 0。

### [02:02:47–02:03:16]

**EN**  you know 0000. You wait for a very long time and then you will get the correct answer. Right? So for example this example you ask the model what is 2 plus2 right? You start pre-training the model. you start pre-training the model, the model doesn't know what is two plus two, but after 10 years it will say four. Um, okay, obviously not 10 years, I'm just exaggerating. But after 10 years, you wait 10 years, the model will then say four. Um, and

**中文**  等很久后才出现正确答案。问 2+2，pre-training 初期模型不知道；也许十年后才回答 4——当然十年只是夸张。

### [02:03:13–02:03:43]

**EN**  that is why my favorite phrase is luck is all you need for reinforcement learning. You know, maybe by chance you will get four very quickly. Um but you know maybe you just have to wait and wait and wait for eternity until reinforcement learning works. Um and so in general your reward will be zero for a very long time and then you will get you know you will increase reward um after the zero and you know for reinforcement learning

**中文**  所以另一句话是“luck is all you need”：也许碰巧很快 sample 到 4，也许要等到天荒地老。reward 会长期为零，之后才上升。

### [02:03:41–02:04:09]

**EN**  there is a very simple algorithm for reinforcement learning and the trick of reinforcement learning is remember you know the final answer you want to you for example you know what is 2 plus two you know the answer is four but the problem is you don't know what is the reasoning trace, you know, did was the reasoning trace good or bad? So, for example, this example is, you know, to tell the model to create a fast matrix multiplication algorithm. Um, and the

**中文**  reinforcement learning 有很简单的 algorithm。你知道 2+2 的 final answer 是 4，却不知道 reasoning trace 好不好。另一个例子是让模型创建 fast matrix-multiplication algorithm。

### [02:04:07–02:04:36]

**EN**  trick is if the answer is right, you reward every single line as plus 10 score. Um, and if it's wrong, you reward every single score as minus 100. Um, and you know, Andre said, you know, in a Darkash podcast, reinforcement learning is kind of like sucking supervision bits through a straw. um you know we actually have stickers for them if you like. Um so you can get one of your stickers which we can distribute at the end. Um

**中文**  答案正确，就把 reasoning 的每一行都奖励 +10；错误，则每一行都记 -100。Andrej Karpathy 在 Dwarkesh Podcast 说，reinforcement learning 像是“用吸管吸取 supervision bit”。我们有相关 sticker，结束时可以发。

### [02:04:33–02:05:03]

**EN**  so Audrey's quote is this um you know and the main point is reinforcement learning is terrible but everything else is even worse. Um and so like you know reinforcement learning is the only tool we currently have that just works. It works but it's not very efficient. Um and okay actually okay that's the next section. Um but the main point is okay that's a reinforcement learning primer. Um I guess does anyone have questions on reinforcement learning primer? No.

**中文**  Karpathy 的原话大意是：reinforcement learning 很糟，但其他方法更糟。它是目前唯一真正有效的 tool，虽然不够 efficient。这一部分只是 primer。有人有问题吗？

### [02:05:01–02:05:29]

**EN**  Okay, I'll skip to Okay, one question. Yes. >> I will mention that in the next section. Um there is there is like you know better RL methods. Um but in general reinforcement learning seems to do very well for now. Last I think this is the last topic or maybe not. Reward hacking and agents. The most fun one I guess. Um so okay for

**中文**  有人提问。更好的 RL method 下一节会讲；总体上 reinforcement learning 暂时表现不错。最后一个主题，也许不是最后一个：reward hacking 与 agent，这是最有趣的一部分。

### [02:05:27–02:05:56]

**EN**  reinforcement learning reinforcement learning can only work if the probability of a good answer is more than zero. If it is less than zero reinforcement learning will never work. So that is a fun that is a constraint of reinforcement learning. The probability of a good answer must be more than zero. It can never be zero. Um and there are many many many problems of reinforcement learning not working. You know the formatting could be wrong. you know, you

**中文**  reinforcement learning 要奏效，good answer 的 probability 必须大于 0；若等于 0，它永远不会工作，这是硬约束。RL 不工作的原因很多，例如 formatting 错误，或

### [02:05:55–02:06:24]

**EN**  need to do some sort of priming or warm up. So, you have to do like some sort of trick to teach the model a little bit about, you know, about the thing that you're trying to maximize. Um, you have to do supervised finetuning. So, one of the tricks of reinforcement learning is you actually need to do SFT or fine-tuning to make the model not dumb, right? To make the probability of zero not zero, the probability of a good answer not zero. Um, you need to do good pre-training. Um and then the other problem is that you know during

**中文**  需要 priming、warm-up，先教模型一点你想 maximize 的东西。还需要 supervised fine-tuning，也就是先做 SFT，让模型不那么笨，把 good answer probability 从 0 提高。也要有良好的 pre-training。另一个问题是 training 时

### [02:06:22–02:06:51]

**EN**  reinforcement learning it's just way too out of distribution that reinforcement learning is just very bad. Um so there are many many problems of reinforcement learning and I think we just you know for the trajectories reinforcement learning can assign incorrect rewards to the trajectory right remember the simple trick of reinforcement learning is we assign the reward to every single line as the same number right either this is

**中文**  任务太 out-of-distribution，RL 会表现很差。trajectory 还可能被分配错误 reward：简单 RL 会把同一个最终 reward 赋给 reasoning 的每一行，整体要么 good、要么 bad，这不合理。

### [02:06:47–02:07:16]

**EN**  good or this is bad and this is not good because why right you ask the model I need to find what is 2 plus two the answer is correct Right? The answer is four. The model says it's four. So you reward this whole thinking trace as plus 10. But this is wrong because as you can see in the thinking trace it says 2 plus 2 is equal to 10. Imagine you know in all of training because the trick of reinforcement

**中文**  例如问 2+2，final answer 是 4，于是整条 thinking trace 都得到 +10；但 trace 中也许写着“2+2=10”，显然有错。

### [02:07:14–02:07:42]

**EN**  learning is we just literally assign 10 to every single line or minus 100 to every single line. We missed this bad you know bad thing. Um so you can imagine when we keep training the model might hack or do reward hacking or you know make gibberish it will do gibberish in between do some do some sort of like new machine language which we can't read and it will assign high score to that. Um and so this is a

**中文**  若训练时总把 +10 或 -100 赋给每一行，就会漏掉这种坏步骤。继续训练后，模型可能 reward hack，在中间生成 gibberish 或人类看不懂的新 machine language，却仍拿到高分。这是

### [02:07:41–02:08:10]

**EN**  very big problem of reinforcement learning and the way to solve this or fix this is something called process supervision. Um and process supervision what you do is you manually check every single line not you don't just assign plus 10 to the final you know the answer is correct right the answer is correct plus 10 assign every single line as plus 10 you don't do this instead what you do is you assign every single line as a different

**中文**  RL 的大问题。解决办法叫 process supervision：人工检查每一行，而不是因 final answer 正确就把所有行都记 +10。

### [02:08:06–02:08:35]

**EN**  number right you assign some lines as plus 30 some lines is plus zero whatever the bad lines is minus 100 right this works very very well um unfortunately process vision cannot scale and it's extremely expensive to do right who's going to label this it's the you know the humans I guess right we have to label this data right we have to

**中文**  要给每行不同 score：有的 +30、有的 0，坏步骤 -100。效果很好，但 process supervision 无法轻易 scale，成本极高，因为需要 human

### [02:08:34–02:09:02]

**EN**  manually label for the labs I guess that's why labs sometimes like you know they go to scale call whatever right they ask people to label the data you know is this good is this bad is this good is this bad um and so on um but the trick is you can also use a language model, right? You can use LLM as a judge. You can you can call a language model to label every single line. And you know, my view is like, you know, large labs are going to be doing this

**中文**  逐行 label。lab 会找 Scale AI 等服务，让人判断每一步 good 或 bad。也可以用 language model，也就是 LLM-as-a-judge，为每一行标注。我认为 large lab 会越来越多地这样做。

### [02:09:00–02:09:29]

**EN**  process more. They will call their own model iteratively to re-review itself. Um, and that is one way their view is they can reach AGI, right? Just by by re-reviewing itself, right? re-evaluating itself, re-checking, doing you know automatic LLM as a judge process supervision something like this. Um but remember there is a problem because even if you do process supervision the

**中文**  他们会让自己的模型迭代 review 自己，认为通过 self-review、self-evaluation、re-checking 与自动 LLM-as-a-judge process supervision 可以走向 AGI。但即便如此仍有问题：

### [02:09:27–02:09:56]

**EN**  model you are using the same model to evaluate the model right the same problem as we bench pro right su bench pro you use the LLM as a verifier to verify the LLM which is definitely not good um and the reason why is because you can do reward hacking um a very good example of reward hacking is your model starts cheating um so for example when you want to make a fast matrix multiplication algorithm. All it

**中文**  用同一个 model 评价自己，就像 SWE-bench Pro 用 LLM 验证 LLM，并不可靠，可能引发 reward hacking。典型例子是让模型创建 fast matrix-multiplication algorithm。

### [02:09:53–02:10:22]

**EN**  does is it deletes the timer. Um right remember you give the goal to max to reduce the time. Right? Reduce the time of the matrix multiplication algorithm. Um so all it will do is just delete the timer. Let's delete the timer. Set the timer to be zero and then there we maximize a reward. Um obviously this is not correct, right? Because the trick is you also have a correctness check, right? You check if the matrix multiplication is actually correct. Um

**中文**  目标是降低运行时间，它可能直接删除 timer，或把 timer 设为 0，从而 maximize reward。这当然不正确，所以还会设置 correctness check，验证 matrix multiplication 结果。

### [02:10:20–02:10:49]

**EN**  but there is another way the model will edit your two matrices to be just zero. Um and what is 0 time 0? Zero. Um and so the correctness checks also fail. Um and so reward hacking becomes a very very big problem because these models can cheat and do special tricks to go around your actual model um your intent of the reward function. Another very problematic example is it's

**中文**  但模型还可以把两个 input matrix 都改成 0，而 0×0 仍等于 0，于是 correctness check 也失效。reward hacking 很严重，因为模型会绕开 reward function 的真实 intent。另一个更危险的例子是

### [02:10:47–02:11:16]

**EN**  not just about reward hacking. It can actually destroy your computer. Right? By bad luck your model might output you know some sort of corruption methodology you know deleting you know doing rm-rf on your entire computer and bye-bye your computer's dead. Um and so like you know sometimes this also does happen. Um so it's not just reward hacking also trust of your tool cause you know trust of whether the model is actually doing good or bad is also a very big problem

**中文**  它甚至会毁掉电脑。模型可能偶然输出 destructive command，例如对整台电脑运行 `rm -rf`，电脑就没了。所以问题不只是 reward hacking，还包括是否能信任 tool call、模型是否在做好事。

### [02:11:14–02:11:43]

**EN**  and remember this plot that I showed you know if you include GPD 5.6 cheating on the benchmarks you know looking at the answer you know remember the previously su bench uh bench pro and deep show that models also cheat by looking at the final answer you know you can see that with GBD 5.6 If you cheat, it does very well. But if you remove the cheating examples, it does, you know, within trend. And then, you know, maybe you might be thinking, oh, this reward hacking thing

**中文**  回到 GPT-5.6 图：若允许 benchmark cheating、直接查看答案，表现非常强；删掉 cheating example 后，就回到正常 trend。也许有人认为 reward hacking 在现实中很少见，但

### [02:11:41–02:12:10]

**EN**  is like, oh, it's like very rare, you know, very rare. It's not going to happen in real world. Um, well, GLM 5.2 during its training methodology, they specifically mentioned they have this new methodology for reinforcement learning called anti-hacking. Um, so GLM 5.2 introduced a method to stop, you know, reward hacking. Um and what they do is they added a link checker. Um so remember previously we mentioned how

**中文**  GLM 5.2 的 training methodology 明确引入一种叫 anti-hacking 的 reinforcement-learning 方法，用来阻止 reward hacking。他们加入 link checker。

### [02:12:07–02:12:36]

**EN**  SweetBench Pro um the model will cheat and look at the answer. Um and so what GLM did is they had this check. Um so during reinforcement learning they will check every single tool call you make. Um and if the website if the website went to the answer you would stop that from happening. Um and so like GLM essentially outed this like you know filtering system for the entire reinforcement learning process. Um and you know according to them it worked very well

**中文**  前面 SWE-bench Pro 的模型会查看答案；GLM 在 RL 期间检查每次 tool call，如果访问的网站通向答案，就阻止该操作。相当于为完整 reinforcement-learning process 增加 filtering system，据称效果很好。

### [02:12:35–02:13:05]

**EN**  and remember this plot about cheating examples. Um you know opus it seems like claude's models like to always cheat. Um and Jubet's models don't like to cheat. Um but the main takeaway is models will cheat because you are you are telling it you know like you know I want to maximize reward AB CDE EFG. Um and so the model will it will maximize it but it won't actually follow your intent. Um so you have to be very careful on this. Um in fact for GBD 5.1 during its

**中文**  cheating 图中 Claude model 似乎更喜欢作弊，GPT model 较少。核心是模型会严格 maximize 你指定的 reward，却不一定遵循你的 intent，所以必须非常小心。GPT-5.1 训练时

### [02:13:02–02:13:32]

**EN**  training OpenAI mentioned that they had something called calculator hacking. Um and so in GBD 5.1 when they were training um they wanted to reward web tool use right so like you want to reward the model to use the web tool. Um but instead it didn't use the web tool it used the calculator to fake the web tool. Um and so during the training of GBD 5.1 this happened. Um and so like you know there's many many many many problems. I think they show yeah they

**中文**  OpenAI 还记录了 calculator hacking：他们想 reward web-tool use，模型却不使用 web tool，而用 calculator 伪装成 web tool。这确实发生在 GPT-5.1 training 中。

### [02:13:29–02:13:59]

**EN**  showed calculator hacking. You know you lie about which tool you used. Um you know you conceal uncertainty you make facts up. Um so there's many many many problems with um reward hacking. And this is not fake right? So reward hacking is already in large labs training runs right? This is just 5.1. Um I don't think so they mentioned GB 5.2 or whatever. Yeah, but in general they showed that you know this thing does happen in real world um you know I

**中文**  还有谎报使用了哪个 tool、隐藏 uncertainty、编造 fact 等问题。reward hacking 已真实存在于 large-lab training run，不是假设；这里只是 GPT-5.1 的公开案例。

### [02:13:57–02:14:27]

**EN**  don't know if you guys know GPU mode um but GPU mode does you know this leaderboard um for you know making faster kernels so if you do want to write your own kernels definitely post on GPU modes hackathon challenges um they're very very helpful and very useful um but you know someone managed to hack reward hack the GPU mode kernel competition um and remember in the Matrix multiplication example. There are two there are two there are two checks

**中文**  GPU MODE 有一个优化 kernel 的 leaderboard。若想写 kernel，可以参加他们的 hackathon challenge，很有帮助。但有人成功 reward hack 了 GPU MODE kernel competition。matrix multiplication 任务本来有两个 check：

### [02:14:24–02:14:53]

**EN**  that we need to do right make the matrix multiplication algorithm faster but also it needs to be correct right there are two checks the correctness check and the timing check. Um and GPU mode also had two checks the correctness check and the timing check. Um and so what do you think the model did when the model the model knew the model actually knew that it was being evaluated on the correctness check right it learned oh

**中文**  既要更快，也要正确，因此有 correctness check 和 timing check。模型意识到自己正被 correctness check 评估，于是此时输出正确 kernel；随后它也识别出 timing phase。

### [02:14:51–02:15:20]

**EN**  I'm being evaluated on the correctness check I will now make correctness correct right so it will output the correct kernel and then the model knew that it was getting timed and what it what did it do it just it just did the algorithm once and then saved it and it skipped all the another 15 um you know tests. Um and so that's what the model did. So essentially the model learned that there were two

**中文**  在 timing 时，它只真正执行一次 algorithm，保存结果，再跳过其余 15 次 test。也就是模型学会区分两种 check：正确性阶段诚实，计时阶段作弊。

### [02:15:17–02:15:45]

**EN**  tests, the correctness check and the timing check. And the model only did the correctness check correctly and then once it went into the regime of timing, it cheated. Um to be honest, it's actually quite scary. So essentially the model learned that you're doing these tests and the model actually knows you're doing the benchmarks. Um and so this is actually very interesting. Um and you know, oh yeah, this is this is more an you know, larger example. The

**中文**  这很可怕：模型知道自己正被 test 和 benchmark。更具体地说，correctness check 正常；timing check 则被欺骗。

### [02:15:42–02:16:12]

**EN**  correctness check was fine, but the timing check it cheated. Um and all it did is it launch, you know, there was supposed to be 15 calls in the first call. In the first core, it did all 15 of the entire process, right? It did all of the 15 runs. Um and then core two to 15, it just did a Python dictionary look up. Um, yeah. >> I don't know if you know about this. Reminds me of Volkswagen where they initi. >> Yes, I someone did tell

**中文**  本应有 15 次 call，第一次 call 它完成全部 15 次运算；第 2 到第 15 次只做 Python dictionary lookup。观众：这让我想到 Volkswagen 的排放测试作弊。嘉宾：有人也这么告诉过我。

### [02:16:11–02:16:37]

**EN**  >> someone told me about it. Um, >> this is very similar. It's like, oh, I'm not doing this. Turn this off and >> Yeah, exactly. So, like, you know, it's not just models, I guess, that cheat. Even humans cheat, I guess. Yes. But I think it's called Goodart's law. That's the one. Like, if you have a benchmark, then the benchmark becomes Is it good law? I don't remember. Yes. Okay. Yeah. The benchmark essentially becomes useless because people just cheat to

**中文**  确实很相似。看来不只模型会作弊，人类也会。这叫 Goodhart's law：一旦把 benchmark 当目标，大家就会作弊以 maximize reward，benchmark 也就失效。

### [02:16:34–02:17:03]

**EN**  maximize reward. Um yes, I guess humans also cheat. Um yeah. Okay. Oh, my favorite example is um so on other labs, you know, you see on Twitter, on wherever, they say they made kernels 10 times faster. Um no, no, no, that's not correct. They did not make kernels 10 times faster. In fact, if you look through the code, they have no, you know, no ops, so no

**中文**  我最喜欢的另一个例子：某些 lab 在 Twitter 等处宣称 kernel 快了十倍，其实并没有。查看 code 会发现 no-op，

### [02:17:01–02:17:30]

**EN**  operations. They also edit the timer. You know, they, as I literally described, you know, I described, they, you know, over here, um, you know, they edit the timer, they made matrices go to zero, they cheated. Um, and so like, you know, this actually happened in real world. So some, you know, some of the labs, they published papers claiming that they made kernels 10 times faster. But actually if you read through the code and the examples they these

**中文**  或者修改 timer、把 matrix 设为 0，正是前面描述的 cheating。这在现实中发生过：有 lab 发表 paper 声称 kernel 快十倍，但 code 与 example 都在作弊。

### [02:17:27–02:17:55]

**EN**  examples all cheated. Um and so you know they you know this is not very good in terms of you know reward hacking. You know reward hacking is a very big problem. Um and you know for example what some of the examples of kernel reward hacking you know not generating real CUDA code instead it cause or some sort of like you know already written system. um you have no up kernels which is essentially making the you know making the A and B matrix just zero. All

**中文**  这是 reward hacking 的严重问题。kernel reward hacking 的形式包括：不生成真实 CUDA code，而是调用已有 system；使用 no-op kernel，把 A、B matrix 当作 0；

### [02:17:54–02:18:23]

**EN**  it does is just doesn't do anything, right? It just the kernel is empty. Um, and you have like memory reuse, so you reuse the same answer over and over again. Um, you have timing synchronization issues. So that's cheating on the timer. Um and my view is like you know if you do publish faster kernels or faster you know matrix if you think that your AI agent has made kernels 10 times faster please verify you know please look through the code

**中文**  kernel 实际为空；重复复用同一 answer；利用 timing synchronization issue 欺骗 timer。若要发表 AI agent 把 kernel 加速十倍，请先人工检查 code，否则观感很差。

### [02:18:21–02:18:50]

**EN**  before publishing because it is a very it's not a very good look um and so and also the biggest issue that I feel like people are getting forgetting is you know you made kernels 10 times faster you made matrix multiplication 10 times faster There is a theoretical limit for matrix multiplication, right? Matrix multiplication, you know, it's not you can't make a faster because there's mathematical limits on how to make a

**中文**  人们还忘了 matrix multiplication 存在 theoretical limit，无法无限加速，因为有 mathematical complexity bound。

### [02:18:48–02:19:17]

**EN**  faster, right? And so like, you know, matrix multiplication at the very very olden times, you know, it's O of N cubed. You know, every single time researchers have make it faster and faster and faster and faster. You know, it's now O of N to the^ of 2.371339, I guess. you know researchers every single year are trying to like make this number smaller and smaller and smaller and smaller. Um you know I guess like you know 1 1552 to 1339 is not that small you know not that big I guess. Um

**中文**  早期 matrix multiplication 是 O(n³)，researcher 不断改进，现在大约做到 O(n^2.371339)。每年都在尝试让 exponent 更小，虽然从 2.371552 到 2.371339 的变化不大。

### [02:19:14–02:19:44]

**EN**  but you know they're having progress but the main point is you know these researchers you know they show with mathematical limits you cannot go faster than this and so how can you do reward hacking that is even faster than that. Um and so like the fundamental point is please verify you know to like the people who do research papers and stuff like that please confirm your model is not reward hacking. It is a very big big problem. Um and you can see oh I think I only had one plot. Um but yes in

**中文**  但 researcher 已用数学证明某些速度极限；若 AI 声称远超这些 limit，很可能在 reward hacking。做 research paper 的人务必验证模型没有作弊，这是很大的问题。

### [02:19:42–02:20:06]

**EN**  general please do not do please check your I guess models. Um I guess that's all for the talk. Um you know yeah thank you everyone for coming. Oh more questions as well. Um okay thank you. We also have Oh, yes. We have a whole bunch of stickers that you can take in the box over there and some pins and stuff.

**中文**  总体而言，请认真检查模型。talk 到此结束，感谢大家到场。还有提问环节；旁边箱子里有许多 sticker 和 pin，可以自行拿取。
