# Stanford MS&E435 Economics of the AI Supercycle | Spring 2026

- Original show: Stanford Online
- Original source: https://www.youtube.com/playlist?list=PLoROMvodv4rOAe_DtQYOCKy6_x3irEYV1
- Discovery source: https://www.xiaoyuzhoufm.com/episode/6a7e819017676351c571a472
- Duration: 07:01:25
- Method: Transcript assembled from creator-provided captions; quotations still require source review.

## Transcript

[00:00:10–00:00:36] All right, folks, we're going to have some fun in this session. My name is Apoorva. I'm going to be your instructor for the next nine weeks or so. Here's what we're going to go through today. I'm going to talk a little bit about myself. Why do I do this. Some logistics on what to expect. Quiz. Yep, I'm that guy. We're going to have quiz on day one.
[00:00:36–00:01:04] And the biggest question that we've all been wrestling with, where's the money? Where's the money in AI? Some of you know me, but my journey started in India. I moved to Singapore. I met a couple of Singaporean folks here earlier. I started my career at Palantir with Sunil and a couple of other folks about 13 years ago, 14 years ago. Led a variety of engineering teams. All that to say, we wrote a lot of spark in government buildings.
[00:01:04–00:01:31] And I came back to Stanford for grad school, which is when I got tired of writing Spark in government buildings. And now, I lead Altimeter. I don't know how many of you have heard of Altimeter, but Altimeter is an investment firm. We focus on fairly concentrated form of investing. We've got two businesses. We've got a public business and a private business. And then I got the biggest promotion of my life six months ago. I'm now the proud dad.
[00:01:31–00:02:00] This is, as people have told me, the biggest investment I will make. Some have called it the one with the most guaranteed negative IRR. I think it's the most guaranteed positive IRR, not financial, but that's me. I live across the street. Reach out with questions. I want to be here as a resource to you guys. I'm joined by an incredible TA in Chloe Feng. Reach out to her if you have harder questions--
[00:02:00–00:02:29] [APPLAUSE] - And make the most of it. We've got a great session lined up for you guys. The course is designed to be no more than three hours a week, and that includes class readings all the time you spend arguing with ChatGPT, with Claude about whether it can do your assignment for you. So it's about an hour of class. It's about an hour or two of readings. And basically, the format is we're
[00:02:29–00:02:57] going to do guest speakers every single class, next class onwards. Chatham House Rules a lot of guest speakers will share, maybe overshare. So please don't record what they're saying. We'll have an optional dinner with some of them right afterwards. You guys are welcome to join. Chloe will arrange the logistics. Grading is easy. 50/50, show up to class.
[00:02:57–00:03:24] If Ali Ghodsi can show up to class, you can show up to class. And the other half is an assignment that we will release at the end of the course. Yeah, it's conversational. Ask questions, be involved. The more you're involved, the you're going to get out of it. And honestly, what's in it for me is I'm going to learn the most from you guys. This is the course schedule over the next nine weeks. Lots of great speakers, one more impressive than the other. As for me, honestly, putting this calendar together
[00:03:24–00:03:52] has been my job for the last couple of weeks and months. Chloe knows all about it. But be present. These are all incredible leaders running incredible businesses across the stack from Semis to Infra to energy on the infrastructure side to models. You're going to have folks from OpenAI and Anthropic and a bunch of applications and agents. So ask all your hardest questions. Save them for the speakers. They're going to love it.
[00:03:52–00:04:22] And we're going to assign some readings. So why should you take this course? What should you achieve in this course? What is a good thing to get out of this? You know, honestly, I thought about this, and I was just telling one of the students here of how it all began. I come back to campus once a year and I talk about all that's happening. And typically, it's in the context of finding great people.
[00:04:22–00:04:49] I realized that this is such a big supercycle. We know it. We believe it. At Altimeter, we have positioned our entire focus around it. And I did not find a course that goes deep in a way that I would have liked to be when I was an undergrad here or a grad student here. And I thought about in five years, everybody was going to ask you, hey, did you see it coming? You were at the start of it. You were around when ChatGPT was launched.
[00:04:49–00:05:18] You were around when the tectonic plates were forming and clay was forming. And I think you want to be able to say yes. So half of you are going to start an AI company and the other half are going to fund it. [CHUCKLING] So at least you should know where to spend the Series A money that you're going to raise. And at the very minimum, you'll have a sense of what not to go, or at least have mental models for, hey, this business that I'm looking at or considering starting or considering funding or considering joining, what are the right questions to be asking?
[00:05:18–00:05:47] What are the laws of physics that govern this business at this part of the cycle? I think it's going to be the biggest one yet. And I'm excited that you guys are here to study alongside us. I'm going to spend some time on this slide because this is the punchline. How many of you have seen a version of this before?
[00:05:47–00:06:13] Well, those of you who did the readings, thank you. I appreciate it. We did include a notebook LM for those who were more auditory inclined. But let's talk about this for a second. What is going on here is actually probably the biggest question in generative AI right now, which is, if you listen to any of the earnings calls
[00:06:13–00:06:40] from the hyperscalers or even NVIDIA and others, is we are investing so much into the CapEx. We're investing so much into building these data centers. It's a 5-layer cake, as Jensen calls it. Energy, chips, power, interconnects, memory, all that to give you a data center that you can either rent by the hour or by the token,
[00:06:40–00:07:07] that you can go train models on and serve those models. And then the question is, hey, these models that you built, are they creating economic value? That is basically the right-hand side of this chart. And to make an analog of the biggest technology revolutions that I have seen, internet 25 years ago, mobile 20 years ago, cloud, probably the most recent one, 10 years ago.
[00:07:07–00:07:35] I put up one of those charts on cloud, but on the readings, you'll see the same for internet and mobile and cloud. And that's the shape of the cloud ecosystem. The cloud ecosystem looks dramatically different than the AI ecosystem. Anybody have guesses as to why that's the case or your theory on why it's so different, or reasons why this might look like-- we're not going to call it a pyramid, we're going to call it a triangle, inverted triangle.
[00:07:38–00:08:07] Go ahead. Is it because it's still early in the upside for AI? Yeah, definitely. [INAUDIBLE] Definitely early. That's a good guess, yeah. Any others? Any other thoughts? Or maybe because NVIDIA has a monopoly, so they can charge whatever. Can you ask that question again next week when we have the folks from NVIDIA? But no, good, it's a great point. They do have a stranglehold, right? One of the charts we had in the readings
[00:08:07–00:08:36] was the market share that NVIDIA has on all of the compute right now, and it's up there. Any other thoughts or hypotheses on why this is so different? Yeah, I don't know, is it the cloud sectors seem to be able to leverage the hardware to generate [INAUDIBLE] and AI haven't really got there yet. We know how software ate the world. As Marc Andreessen said, software ate the world
[00:08:36–00:09:04] because I could build software, you could build software, and I could distribute it to millions of people. And the marginal cost of running that software was close to 0. These software businesses ran at 80% some even at 90% gross margins. That is not the case with this new economic model of AI, because if we have a set of users using Cursor or using-- you hear all these stories about large bit scale businesses that
[00:09:04–00:09:31] are still not profitable at billions of dollars of revenue scale, is because of that, is because the incremental user of an AI application is not free. It's not marginally free. It's actually quite a bit more expensive to have AI users, because it turns out you've got to burn those GPUs. And I would say everything you guys said from it being early to NVIDIA being dominant,
[00:09:31–00:09:58] we'll call it, to the physics of the problem are very different of how inference is run is certainly where we are right now. So I think that's the case right now. I might add another dimension to it, which I spoke about in the readings was, we analyzed what happened in internet. We analyzed what happened in mobile and cloud, and how many years did it take for these triangles to flip?
[00:09:58–00:10:27] And one of the examples we take is AWS. AWS started in the year 2004. AWS has its first customer in Netflix in 2010, and ultimately Amazon shifted fully to AWS in 2012. Eight years from breaking ground, eight years from first CapEx investment cycle. I don't know if any of you were around reading earnings reports 20 years ago, but the big debate was, hey,
[00:10:27–00:10:56] is Amazon going to go bankrupt? And that was the biggest question everybody had about the build out of AWS. And thankfully, nobody, at least yet, is on the verge of bankruptcy. But these are large numbers. So we'll come back to this slide. But I would say this is the central theme of the course that we're going to explore. We're going to have speakers from some of the companies that are listed here to others. And the central theme that we're going to pick around
[00:10:56–00:11:26] is like, hey, in your field, with the NVIDIA speakers, are you a dominant force? How long are you going to stay to be the dominant force? What are the forces that you're most worried about? Who are the ASICs that you're most worried about? What are the pricing compression vectors for your business? To the folks at Anthropic and OpenAI, who we're going to talk a lot about profitability, is you're serving a billion-user franchise at OpenAI.
[00:11:26–00:11:53] With the Anthropic folks, honestly, 100% of this class is on Claude. So we'll ask them about, is this group of users profitable? How do you think about profitability? Is ads going to be a bigger source of revenue than subscriptions? And then for the folks in the middle, which is the inference layer, this is the most competitive part of the whole ecosystem. There's a lot of startups that are doing really well. They're winning so far. But you've also got the hyperscalers
[00:11:53–00:12:23] who want to have a dominant say in that layer. So honestly, the jury's still out. And the biggest question there is, are you a feature or a platform? A lot of new businesses that we are seeing on the infrastructure side that they feel very good ideas. But if you ask yourself the question, hey, why is this not a part of AWS? You are thinking about maybe it should be a part of AWS.
[00:12:23–00:12:53] So for the speakers, we're going to talk a lot about that. Any questions before we jump into the quiz? Go ahead. I'm curious how you think about-- so on the right-hand side, like the triangle being-- like the Application Layer being small, how do you think about including incumbent platforms into that? Maybe like Salesforce. Maybe [INAUDIBLE] revenue. Like, would you-- do you include them as part of that-- in line with that pyramid and how it shifts over time? It's a great question and I might add, Salesforce, Palantir,
[00:12:53–00:13:22] there's a series of-- let's call them "old economy businesses" that are reinventing themselves to have skews of products that are, in the case of Salesforce, Einstein, in the case of Palantir, AIP. In the case of-- there's a series of these. And the answer is yes, there should be. The answer is yes, that they should be. The way I solve for that in this calculation is I get the model revenue. And so if you were running Salesforce, you're probably running either one of the big models
[00:13:22–00:13:34] or running inference. So their spend is captured in the app layer by way of the substrate. It's very hard to extract that out from public disclosures. But yeah, we should. Yeah.
[00:13:38–00:14:03] Is a large part of the bottom part of the right-hand side pyramid, basically, buying capacity for future revenue, which are the group at the top, which is what we're not seeing in the [INAUDIBLE]. It's a great question. And maybe just to rephrase the question, the question is, hey, is there a timing mismatch in the build out of the Semis layer? Because typically, you build Semis for a five-year period or a six-year period.
[00:14:03–00:14:32] But the application revenue is for right now. It's a great question. And that's what makes the lower half of this, call it triangle, somewhat cyclical. And you go through phases of CapEx cycles. Think of it as like laying down the railroads. That is very much the case. And so there's a chart in the readings for what happened in the mobile super cycle. Something very similar happened.
[00:14:32–00:14:59] The first inning had inflated market caps for a lot of the CapEx heavy businesses. And so if you think about a basket of CapEx names to call it steady state names, you should expect that. And so I suspect we're so early that that's happening as well. I'm curious to see how you think about Google, because I see that you labeled Google as Google Cloud there on the [INAUDIBLE].
[00:14:59–00:15:27] But Google also have their own Gemini models [INAUDIBLE] TPUs, and they're perfectly [INAUDIBLE]. How many of you are kind of positioned in this triangle? Any large conglomerate like Google deserves to be-- we have to call business units. So I would put the TPU business unit in Semis. We include that here as we counted the revenue. Their GCP unit is in the Infrastructure Layer. And then the Gemini unit is at the Apps Layer.
[00:15:30–00:15:59] We have a chart later that we'll talk a little bit about. Gemini is actually one of the most used consumer applications. It's the second most used consumer application right now. And the biggest question there is, how much of that is coming from the distribution advantage that Google has to it meritocratic being such a good application? The jury's still out, but we'll get into it. Yeah. Well, let's-- go ahead.
[00:15:59–00:16:29] I just have a question about prediction. Right now, it looks like this triangle shape, if it were to be successful, perhaps it should become inverted. But what does an unsuccessful new technology look like? Does it stay a triangle? How would you be able to predict whether this would be good or not? Yeah, I'm not sure the-- maybe rephrasing your question, I might say it out like, what is the stable equilibrium of this industry?
[00:16:29–00:16:57] I think it is pretty clear that AI is unlikely to be a fad. Is unlikely to be an unsuccessful endeavor. And I think about the stable equilibrium of this chart quite a bit. In fact, I got into a little bit of a debate on Twitter with somebody quite smart and who's thought a lot about this exact question of what is the stable equilibrium. And my guess is that it might stay this way for longer than I anticipated.
[00:16:57–00:17:27] In the cloud, I think that range is about a decade. I have a feeling this might stay longer-- this way longer, because of just how hard it is to get the substrate right. But there will be one or two unlocks. I couldn't tell you what they are, but for example, if one of the ASIC programs at one of the hyperscalers, be it Google's TPU or Meta's MTIA or the folks at Amazon and OpenAI
[00:17:27–00:17:53] and Microsoft and all the labs that we don't even know about exist breakout success, I suspect that'll be the biggest repricing of that layer. The other catalyst could be I think about the hyperscaler CapEx guidance in earnings calls, which, by the way, I recommend everybody here to listen to. Four times a year, you'll have public company CEOs tell you their biggest questions, their biggest things that they're
[00:17:53–00:18:23] thinking about, and I recommend listening to those. If they just stopped guiding to big numbers on CapEx, because that would imply that the current equilibrium does not work. So that's why the second thing that could happen. And so you see, there's a lot of news about the guidance that all the hyperscalers give about their CapEx for that reason. Go ahead. [INAUDIBLE] of training versus inference. Because my sense is if--
[00:18:23–00:18:50] the only way this flips is if inference is meaningfully larger than training. And I'm curious to hear your thoughts on when you think that will happen, because then you're seeing that [INAUDIBLE] or you feel like they will stop spending on training because they're not seeing [INAUDIBLE]? [INAUDIBLE] It's a great question, and it is probably one of the nuggets of information that NVIDIA's earnings calls, have the most sought
[00:18:50–00:19:19] after nugget of what is NVIDIA's share of inference in their fleet. Last I checked, it was about 40%, or they quoted to be about 40%. Meaning that if they were selling a million GPUs, assuming full utilization, about 40% of them were used for inference and the other 60% for training. I suspect that number will increase over time in favor of inference, but I couldn't tell you when and how it'll happen, because there's a lot of training still going on in the world. And the shape of the training workload, as you know,
[00:19:19–00:19:48] looks very different from the shape of the inference workload. A training workload is very predictable, high utilization for a short period of time. The inference workload is very burst usage. Typically when humans are awake until the agents take over, maybe then it'll be 24/7, and harder to predict. It goes down around Christmas for some reason. It goes down around Thanksgiving for some reason. But I think that might be the case, though we,
[00:19:48–00:20:16] at least in this calculation, we try to capture it because it's a mix and-- but it's a good hypothesis. [INAUDIBLE] where would probability be? It's on slide 16. We'll come to it. I'll give you the answer. The most profitable part of the stack is the Semis layer by a long shot. NVIDIA's data center revenues or on a gross margin of-- About 75%. Don't quote me on it.
[00:20:16–00:20:44] It's like plus or minus a couple percentage points from there. Whereas, I estimate some of the Application Layer revenues to be somewhere between-- depending on who you ask, between 0% and 30%. And so the gap is quite wide. And I think the reason to that-- I mean it's a theory that a gentleman here had is, there's one player who runs the tables on the Semis here. And so it's very much the case.
[00:20:44–00:21:01] And in fact, if you looked at this from a profitability perspective, it's even more concentrated. The triangle is even more concentrated. I'll flash that in a second when we get to it but it's a great question. Go ahead. Do you think [INAUDIBLE]?
[00:21:08–00:21:34] Yeah, that is definitely a big part of it, is that we've gone through the investment cycle in cloud. It's definitely an element to it. Go ahead. If all these Infra companies like Google, AWS [INAUDIBLE] all their own TPUs, and media is also doing inference, OpenAI is also searching for some ASICs, where does all these ASICs inference
[00:21:34–00:22:01] startups want to sell to if they are using their own chip? There's $300 billion of revenue to fight about. But to answer your question, about half of that, as Jensen discloses on the earnings calls, is from the big hyperscalers. So those are probably going to be your primary customers. So if you were starting a chip company today, you would have a very-- the shape of your customer base
[00:22:01–00:22:27] is a very small number of very large orders. It's a very different shape from building a consumer business or an enterprise software business. And then you might have a long tail of other enterprises, though I wouldn't bank on it because I think they just go to the cloud providers. If you were thinking about starting a chip company, it should be your number one consideration, is like, which of the five are you going to sell to first?
[00:22:30–00:22:58] Last question. [INAUDIBLE] do you get a small handful of winners in each of these layers? It takes multiple years for that to play out. Maybe I'm wrong, but I feel like in the past there hasn't been a fully vertically integrated layer [INAUDIBLE]. I get that Google is fully vertically integrated on the right-hand side, but wondering how that shift the balance of power this side. What a great question. The biggest winner on the internet super cycle is probably Google.
[00:22:58–00:23:27] It's about $3 trillion in market cap, has near 99% market share in search. I would say that that's a pretty vertically integrated player, right? They run their own file server to search to ads on top to the user experience. Let's see, the next one is Mobile. The winner of that Super cycle is Apple. What? With $2.5 trillion or so in market cap. You called it already. The next one is, let's say social. Meta is probably the big winner in social.
[00:23:27–00:23:56] They're not as fully integrated. And what is their market cap like? $2 trillion or something right now? Pretty dominant, but maybe they lost a trillion because they didn't fully go down to the servers. And then the cloud is fairly heterogeneous. We don't have a single player that won the cloud. You've got the three big oligopolies in AWS, GCP and Azure, but they're not fully integrated. And NVIDIA has been trying a lot. NVIDIA has been trying-- I don't know if you've heard of DGX Cloud, which is their cloud effort to build the cloud ecosystem.
[00:23:56–00:24:26] Obviously, they've got a series of vertical apps that they're trying. So yeah, you might be on to something. Folks, I know it's a Thursday evening at 5 o'clock, probably the last thing standing between the weekend and your weekend. So I don't want to be that person. So I'm going to jump into the part that wakes you up. I do actually have-- this is a quiz that we're going to go through. I'm going to give you a hint about the companies that we're going to go through.
[00:24:26–00:24:54] I do have a prize for the winner. This is the prize so you're motivated. And you win points on two grounds. One is by being right, and the other is by being fast. The fastest way to be fast is to do fast inference and drop that thing into Claude. Please, you're welcome to do that. Just give the human players five seconds, let them go at it
[00:24:54–00:25:02] and let them win the analog way. And if you really want to use Claude, you're welcome to do it, just give them five seconds. All right, so this is question number one.
[00:25:10–00:25:13] [CHATTER]
[00:25:25–00:25:27] Ready for the next?
[00:25:33–00:25:37] The software engineers in the room might have an unfair advantage.
[00:25:52–00:26:22] So I wanted to spend the next maybe 10 minutes or so going into some of the hypotheses that I have about what's going on and why the value is accruing in the manner that it is. I think there was a question, a very good question, about profitability and how it gets magnified. So we'll jump through that. But again, feel free to stop me if you have any questions. I have a feeling we're going to have very little time left,
[00:26:22–00:26:48] and I do want to end on time. You guys remember this? And I painted it slightly differently on the next chart, which is, I did the same exercise that I did that I posted about two years ago. And what it looked like two years ago was this thing on the left, where the ecosystem was obviously a lot smaller. It was about five times smaller.
[00:26:48–00:27:16] Shockingly, the shape of it hasn't changed much. This is despite heroic growth. And if you look at the revenue that was added, about $350 billion or so of revenue added, a good, like, 75% of it just went straight to Semis in the last two years. Despite apps having grown more than 10x,
[00:27:16–00:27:45] it still hasn't made that big of a dent. And so I was like, OK, well, let's dig deeper into this. If you started to open up each of these cells and you're like, hey, what companies make up each of those parts? Most of that 300 is in NVIDIA, as you guys know. the Apps is actually, two companies make up about 90% of it. Anybody want to guess which those two are? The Infra segment is the one that has the most competitive intensity, as we discussed.
[00:27:45–00:28:14] It is probably the place where there's the biggest battle brewing both sideways, but also across the stack. It's also the place that has the highest metabolic rate in that there's a lot of companies being formed, there's a lot of companies that are getting bought out. And I would say it's the most competitive, but also the most unstable of the equilibriums that we have right now. And the question that we think about as we think about investing, as you guys will think about investing your time, is how much time will this
[00:28:14–00:28:42] chart that has moved such little in the last two years, what is the amount of time it will take to get to cloud software like shape? Is it five years? Is it 10 years? Is it 15 years? Is it never? Maybe it just stays that way. We do think it will happen. We think it will happen at some point, but it's not happening nearly fast enough. The second thing that we've been thinking a lot about as we think about the future of AI is--
[00:28:44–00:29:13] I don't know if you guys saw this chart in the readings, but consumer AI, which is the biggest, call it market for AI right now outside of coding, has incredibly high usage on-- ChatGPT, most of it is free. About 95% of the users are free. And Gemini, whose-- I don't know if you guys saw, but Demis, who leads DeepMind, announced that they were not planning to do ads as a subscription-- as a revenue model.
[00:29:13–00:29:40] We've been thinking a lot about hey, how big do these businesses get? What is the monetization engine of these businesses? Do you think a subscription business will be larger or ads business will be larger? And so what I did was I looked at the largest consumer franchises outside of AI. And so you will see that-- I mean, you all know these products. There's a class of products that have gotten to three billion users scale.
[00:29:40–00:30:09] These are almost near mandatory products to live your lives. This is WhatsApp and Chrome, which you could not live without. Then there's a class of products that the 1,5 billion to 2 billion users scale, which are social, these are social products like Instagram, TikTok, and Facebook, they're not mandatory, but they're exhibit very good network effects. If my friend Chloe is on one of these, I'm more likely to be there. And then you've got the third category of mainstream consumer
[00:30:09–00:30:36] products that are neither mandatory, that are neither extremely social, but I would call niche products. If you're shopping, you're going to Amazon. If you're looking for music, you're going to Spotify. If you're looking for a good debate, you're going on Twitter or cat videos. Any guesses on where closer to which of these will ChatGPT and Gemini are right now? And the answer is on the next slide, so we'll get it quickly. Would you guess that ChatGPT or the leading AI application's
[00:30:36–00:31:04] terminal scale will be closer to a mandatory app like YouTube or WhatsApp, a social app like Instagram or TikTok, or a niche app like Spotify or Twitter? Any guesses. If not, I'll reveal the question. Go ahead. I would say on the YouTube, WhatsApp scale because it would be a daily utility. Yeah. People would just be using it daily as part of their normal life. Yeah.
[00:31:04–00:31:32] You are-- well, let me show you the answer and we'll come back to your biases. Any other guesses? Any different guesses? Go ahead. I think it's closer to Facebook time. [INAUDIBLE] Yeah. Yeah, that's right. You're certainly right right now. OK, I'll show you guys the answer. This is how they fare if you plot them all together.
[00:31:32–00:32:01] ChatGPT has just overtaken the niche app category. Gemini still has not. You were right that it's heading towards social. Personally, I would have loved, as an investor at OpenAI, I would have loved for it to start heading towards the core utility. But one of the biggest questions that we ask ourselves is, is knowledge work work that everybody does? Is the work of-- ChatGPT is not a place where you're messaging other folks, yet it's not a place where you're getting your email inbox
[00:32:01–00:32:31] or your dopamine fix. It's a place where you go and you have to do active work. You have to go ask a question. And the number of people in the world who are asking active questions of technology is not the entirety of the population that's online. There's about 8 billion people on the planet. 4 billion of them are online. The rough economics of consumer applications, Alphabet has about 4 billion users. They monetize them at about $100 a user a year.
[00:32:31–00:32:57] Meta's got about 3.5 billion users that monetize at about $70 a user a year. The leading AI provider ChatGPT, has got about a billion users that are monetized at about $10 a user a year. And so the question is, how do we get the billion up to 4 billion? I'm not sure knowledge work is the answer. I think we'd have to go beyond knowledge work. And then the second question is, how do we
[00:32:57–00:33:24] get the $10 a user per year up from 10 to 100? And I'm not sure subscription is the answer. I suspect we'll have to go into ads. And I suspect the ads that ChatGPT will be able to serve or Claude will be able to serve will have a lot better pricing because they will understand your intent, that you will be logged in, very good attribution, a lot more trust. And I think that will be the other big headline this year. And you heard it here first. It'll be a big deal.
[00:33:24–00:33:53] There's a lot of alpha in understanding the ad model really well. Once again, 10 years ago, the Facebook IPO, there was a lot of short reports on Facebook because people said, hey, well, these ads worked on a computer, they're not going to work on a phone. Why? Because there's no space on a phone. Shocker, we found the space on a phone. The same thing's going on right now, which is, while I'm having this conversation, it is a very personal conversation,
[00:33:53–00:34:19] I don't want to be interrupted by advertisements. That's the bigger debate. I couldn't tell you what it's going to be like, but I am optimistic that we will find it. And I think that's going to be a big unlock, a big unlock for this economic model. And so we'll dig into that in one of the speaker sessions later this year. I've got a bunch more slides. We are at time. Thank you.
[00:34:33–00:35:00] The premise of the class today, we're going to talk about-- everybody knows how software in the world, software produced had near zero incremental cost of distribution. That is not the case with AI. More users on AI apps require a lot of compute, and so it's not near zero. That's the topic of our discussion. We're going to do a presentation by the group. We're going to do a fireside chat,
[00:35:00–00:35:28] and then we'll open it up for questions. So without further ado, I'm really excited for today's guests. Our first guest is Brad Gerstner. Brad is the founder and CEO of Altimeter. Brad started Altimeter in 2008 with a few million dollars from friends and family. Today, 18 years later, Altimeter manages over $15 billion across public and private markets. Brad, I've known you for a little bit, and the single consistent thing I've known about you
[00:35:28–00:35:54] is that the best investors have invested across super cycles, across up markets, down markets, recession, crises, COVID, and Brad has done all of that and more. Brad started his career trained as a lawyer, helped start General Catalyst back in the dotcom era, started a couple of businesses after that. Altimeter being the fourth. And at Altimeter, early in the internet, early to Google,
[00:35:54–00:36:23] early with mobile and Meta, and many others, early to cloud and software, later investment in Snowflake, Confluent, Mongo, GitLab and now with AI, one of the largest investors in OpenAI and Anthropic, which I know you guys from last week love, in NVIDIA. And he was on the board of Cerebrus, and led an investment in Grok, which we're going to get into deep today. And outside all of that, perhaps the most important movement that Brad has started is Invest America.
[00:36:23–00:36:50] Can I get a quick show of hands, how many of you have heard about Invest America? Wow, look at that. Got a lot-- a lot of opportunity, it looks like. That's right. In brief, Invest America is a federal legislation creating an investment account at the time of birth for every child born in America. The biggest impact Invest America is going to have, according to me, is independence away from dependence from our state, and making
[00:36:50–00:37:20] every child in America an owner of our economy. Brad, I have the great honor of calling you my mentor, coach, and partner. Thank you so much for doing it. Please join us. It's great to be here. Thanks for having me. [APPLAUSE] And a special thanks to Dr. Goel for green lighting the class. I think it's a really important one. In a lot of schools today, in particular colleges on the East Coast as well, there's this-- people don't really know what to do with AI.
[00:37:20–00:37:48] And I say all the time, you got to make yourself bionic with AI. You can't consume enough AI today because it doesn't-- it used to be you go to this school, you get a job at a place, like Grok or Altimeter, and today, I don't really care where you went to school. I want somebody who shows up and delivers abnormal value, bionic value. And the way you do that is going to be leveraging the latest technology. So I'm glad that you are enabling
[00:37:48–00:38:17] the students to sit at the intersection of such important topics. I'm going to introduce Sunny, and then I'm going to share a couple slides, and then I'll invite Sunny up. But I was thinking, Sunny and I have been great friends for a long time. We're going to play poker tonight in the all-in poker game here in Silicon Valley, so we're buddies inside and outside of work. But I was thinking about the introduction, and then I asked ChatGPT and Claude to give me an introduction. And ChatGPT wasn't great, to be perfectly honest.
[00:38:17–00:38:45] And Claude blew me away, so I figured I'd just read to you what Claude had to say about Sunny. So our next guest is a serial entrepreneur who apparently can't stop getting acquired by bigger and bigger companies. And, honestly, the trajectory is incredible. He co-founded Extreme Labs, a mobile development shop acquired by Pivotal. Then he co-founded Autonomic, a smart mobility platform acquired by Ford, which he became the VP running
[00:38:45–00:39:12] Ford Dax, their internal innovation lab. Then he co-founded Definitive Intelligence that was acquired by Grok, where he helped-- where he became president and helped launch Grok cloud. And then NVIDIA bought the platform, of course, recently for $20 billion, their largest acquisition ever. So if you're keeping score at home, Pivotal, Ford, Grok, NVIDIA, the man's career is basically a SPAC that only goes up, according to Claude,
[00:39:12–00:39:42] unlike Thomas, OK. He has a computer engineering degree from the University of Ottawa, which proves that even Canadians can disrupt things when they put their minds to it. Please welcome Sunny Madra. Thank you. [APPLAUSE] So, I like to share a couple slides to set the context for the moment that we're living in. And inference, the conversation we're going to have today
[00:39:42–00:40:11] is really a subset of this important conversation. But this is global GDP per capita over the course of the last 2,000 years. And if you look at that you realize that basically for 1,800 years, nothing happened. It was survival. There was no excess productivity from a fixed amount of labor and capital. It was what we could use to survive, and then all of this stuff starts happening in the 1800s
[00:40:11–00:40:40] and 1900s. The number of years it takes to double GDP-- and I like to say GDP is what creates the excess in life for enjoyment. It's beyond survival, the surplus that we all have. And so the number of years it takes to double GDP has plummeted, and now we're doubling global GDP-- or you might think of it as quality of life-- every 25 years. But you may say, well Brad, what does GDP have to do anything?
[00:40:40–00:41:09] Well, it has to do with everything. So what happens when you have higher rates of GDP? You have lower rates of poverty, you have higher rates of basic education, you have higher rates of literacy. You have more democracy, more freedom, higher rates of vaccination, lower child mortality. So it turns out that innovation in and of itself is a societal good. And it happens to be correlated and accelerating.
[00:41:09–00:41:37] So technology as an investor has gone from 5% of global GDP to about 13% of global GDP. And if I asked you guys, 10 years from now, are we going to be at below 13% of global GDP or above 13% of global GDP? I think you would all say that technology is a percentage of global GDP is going to be a lot bigger number. We're sitting in the heart of Silicon Valley, technology outearns non-technology.
[00:41:37–00:42:05] So the dotted blue line here-- this is the NASDAQ-- has compounded earnings per share at 15% for the last 10 years, compared to 6% for non-tech companies. So why do technology companies tend to be better investment than non-tech companies? Because they compound their earnings per share faster. Again, I think, will be accelerated by AI. And, of course, AI is going to massively accelerate all of this
[00:42:05–00:42:33] because when we look at all the knowledge work in the world, the TAM for it is measured in the trillions. Demis said it well, it'll be 10x the impact of the Industrial Revolution, but happening at 10x the speed, probably unfolding in a decade rather than a century. So I think that is the context. What we're doing here and the acceleration that will come with AI should be something that's better for all of society.
[00:42:33–00:43:02] We're going to have to talk about the guardrails and the societal change we're going to have to make to be that. But sitting at the very root of all of this is compute. You guys all know that the atomic unit of AI, or intelligence, is the token. And there's nobody better able to talk about the production of this atomic unit than Sunny. And so, Sunny, I want to go back in the wayback
[00:43:02–00:43:31] machine a little bit, tell us what Grok is, and what your observations were in 2023 and '24 about what was going to happen with Inference. Yeah. So a little bit of a quick background. So Grok was founded by Jonathan Ross. Jonathan Ross was the creator of the TPU at Google. And Jonathan Ross's background is interesting.
[00:43:31–00:43:57] He's a high school dropout, not because he couldn't complete it, because he was probably too boring for him. And went straight from being a high school dropout and probably completed his GED or something, and went straight into a PhD math program at NYU. And then gets recruited into Google, and like every great engineer over the last 20 years, was made to work on ad optimization or ad testing, which is terrible in some ways.
[00:43:57–00:44:26] But what he did was he listened to a talk by Jeff Dean. And Jeff Dean had come in and basically said, hey, good news, bad news. Good news, I think we've found an algorithm to solve automatic speech recognition, which could be useful in many places. Bad news, there's not powerful enough compute so we can never run it. And Jonathan took it amongst himself to come up with a design, and coming from a completely different area design, using an FPGA for the first version of what became the TPU.
[00:44:26–00:44:54] And then, ultimately, Jonathan left Google because he thought the rest of the world should have this. It shouldn't just be embedded inside Google. And so he left and started it. And so, quickly, what Grok is, and it continues to be inside NVIDIA as well, it's a chip that's designed with a dataflow architecture. And what makes it very significantly different than any other computer architecture is that it's fully deterministic. So hand in hand with the architecture
[00:44:54–00:45:23] is a compiler, and a compiler which predetermines where all the calculations are going to happen. And that last bit is really important because the underlying thing to any AI problem and token generation is lots and lots of math. And that's why we're seeing this compute explode. And I implore everyone to go look at the following. We talked about one of Brad's great investments, Snowflake. Snowflake is database retrieval company. And so you have to go, get a record, and bring it back.
[00:45:23–00:45:51] And if you look at the number of cycles it takes to do that on a compute cycles, you can really see-- and it's not a really large amount. But you look at the number of tokens it takes to generate-- a number of compute cycles or flops it takes to generate a single token, it's mind boggling. And the best way to think about it is it's usually the parameter size of the model times the context length squared. And so that's for each token. And you're doing something and have lots and lots of tokens. So we're in this era where we have this incredible technology,
[00:45:51–00:46:20] but it's incredibly compute intensive, several, several orders of magnitude larger than any other computing paradigm we've had before. At your own startup, you have a conversation with Jonathan about merging. I was an investor in Cerebras, which is also building a fast inference chip. Grok was building a fast inference chip. These two companies had been in existence for upwards of 10 years. And the extraordinary thing is, like in year nine,
[00:46:20–00:46:49] they're both fighting for survival. They're not thriving. Like they're building for a market that didn't really exist. But you saw something. Jensen came on my podcast, BG2, and he said, everything just changed. I said, what do you mean? He said, inference time reasoning, he said, we've gone from pre-training models to inference time reasoning. And inference is about to 1 billion x.
[00:46:49–00:47:18] So not 10x, not 100x, not a million x, it's going to billion x. And our systems of compute are not designed for what's coming. I remember you and I had a conversation, and shortly thereafter, you helped broker this vision for Jonathan that said, Jonathan, I think I see your future more clearly than you do. So tell us about that moment. Yeah, I think, at that moment a couple of things are happening. So, the market had been dominated by NVIDIA
[00:47:18–00:47:47] because NVIDIA is what the researchers used to create the models. And so naturally in part of creating a model, inference is the forward pass of training. Right in the back prop, that's what's different. And so you're always doing inference when you're creating models, and so it's very natural to just run it on the same hardware that you've created the model on. And so one of the things that we saw with the broad architecture was that we could complete inference much more efficiently. And so if you look at our V1 chip, which
[00:47:47–00:48:16] we put into the cloud-- and I'll get back to in a second-- that's a silicon designed in 2018, silicon from 2019, 1914 nanometer and super competitive against hoppers. Which is five generations newer in terms of silicon technology. And so, really, what we saw was we thought it would be very difficult for convincing people to buy our hardware and use it, but if we built a cloud and put it in the cloud, developers really don't care if there's an API.
[00:48:16–00:48:43] And developers, we've seen they're quite fungible there. So our big insight was take these things, start putting them in the cloud, run the data centers and make them available via an API, and make the best open source models available for everyone. And even including OpenAI models, like OpenAI Whisper, which was open source from the beginning, and so we had put a lot of those models there, and that's what really took off. And we launched the cloud, and within a few weeks, we went to a couple 100,000 users and today it's like something at 4 million users.
[00:48:43–00:49:13] And it took NVIDIA almost 17 years to get to 7 million users. So, effectively, reasoning models come along. Reasoning models have are much more voracious in their token consumption. This is even before we get to agents. This is just deeper thinking than what one shot pre-trained models were doing. And so when you looked at token consumption curves, they were just going parabolic. And our hardware, our clouds were starting to break.
[00:49:13–00:49:41] OpenAI only had a gigawatt of compute, Anthropic only had a gigawatt of compute, so we had to figure out how to make both more token efficient models, but also more token efficient architectures. So now, remember, Cerebrus and NVIDIA were big time competitors. NVIDIA and Grok were perceived as big time competitors. So, Sunny, you sent me a text and said, I have an idea.
[00:49:41–00:50:09] We at the time were major shareholders, and still are, in NVIDIA, and good friends with Jensen. And you had an idea. And the reason I want to point this out is just like how one person's idea-- we see these big transactions, but sometimes we don't unpack that it's just one decision on one day that causes these things to occur. So, what was your-- you had a vision
[00:50:09–00:50:37] that was pretty orthogonal when thinking about NVIDIA. Tell us about how that came to be. Yeah. And, Brad, you did lead that email, so that was awesome. But basically when we were looking at the problem of inference, even as Grok, what became obvious to us is if you started to dissect how inference works, there's first a dissection which happens between, say, prefill and decode. And so many people were starting to do that where you basically
[00:50:37–00:51:06] use a separate set of machines for prefill and another set of machines for decode, and you can basically get some efficiency-- or lots of efficiency by doing that. What we further did-- and this is a good lesson for everyone-- we further looked at prefill and decode. And within the decode we realized that we could disaggregate the decode, because within the decode there's many different functions that are happening. And some of those functions are compute intensive, and some of those functions are memory bandwidth intensive. And so one of the big differences with Grok over a GPU
[00:51:06–00:51:35] is GPUs have lots and lots of compute and lots of external memory, which is HBM for them, which is slower. We don't have a lot of compute on Grok chips, but we have a lot of SRAM. And that SRAM is very high bandwidth, almost more than an order of magnitude faster. And so typically on a CPU you'd see that as your L1 cache, but we have lots of that in our chips. And so when we looked at the problem-- and what the email to Jensen was about was basically connecting to their chips via something they call NVLink.
[00:51:35–00:52:04] So NVIDIA chips speak to each other via protocol called NVLink, and that allows you to basically not run something on a single GPU. You can run it on lots and lots of GPUs together. I think today we have 72 you can do. We're scaling up to 576. Grok has a similar protocol, and we've been running thousands of chips together. In fact, we have many models that we were running on 4,000 to 8,000 chips at a time. So basically NVLink Fusion was a way for us to allow our chips to speak to the NVIDIA chips so we could take part of the problem, which we knew the Grok
[00:52:04–00:52:32] chips were faster at and more performant at, and run it there. And the net result of all that is if you take the same footprint of power, you can get two and 1/2 times more tokens out by basically combining those two systems together, which in today's world of constrained compute is really valuable. So, Sunny sends me a text and he said, I think we can partner with NVIDIA. That in and of itself is a pretty big change because if somebody your chief competitor,
[00:52:32–00:53:02] the idea that you can partner with them is a pretty big change. He said, would you mind sending Jensen a text? And I'm thinking to myself, man, I'm going to spend some political capital with Jensen so, like, I need to know that this isn't a crazy idea. And so I kind of sit on it for a week or something, and then Sunny texts me again. He's like, have you sent Jensen that text yet? And so I said, OK, I'm going to send it to him. And Jensen immediately got back to us and said, interesting idea. Let's have a chat.
[00:53:02–00:53:31] And you guys started working with him and him, and what was really compelling, I think, to Jensen was you had obviously somebody who built a competitive chip, but they had mentally thought about how can we produce a lot more tokens together. So what Sunny just said is really important. OpenAI has got a fixed footprint of let's call it a gigawatt that they're going to take in September of Vera Rubens. In one end of the factory goes power and chips.
[00:53:31–00:54:01] You, obviously, have the building and all the costs. And out the other end comes tokens. When they bought Grok for the exact same power footprint, for the exact same building, they're now generating two and 1/2 times the number of tokens. And the constraint we have in the world is power and memory. So if you can double or triple the amount of tokens for the exact same footprint, it leads to an enormous economic outcome for OpenAI or for Anthropic.
[00:54:01–00:54:28] And so you've seen as the demand on inference, because of inference time reasoning-- and we'll talk next about agents-- as the demand for these tokens of intelligence have exploded, and literally we're consuming tens of trillions of tokens now per week around the world, we've had to come up with more power, more chips, more of these inputs in order to produce those. And so it's not just about fast chips.
[00:54:28–00:54:58] It's also just about-- or fast inference, it's just about the ability to get more tokens into the world in a world that is constrained. Yeah. How many days-- from the time you showed Jensen a working system, how many days from that until him greasing you with $20 billion? Probably just over a month. Yeah. Yeah, 30 days. Yeah. And Jensen is like-- did they have any competitive efforts going on at NVIDIA?
[00:54:58–00:55:27] Yeah. I mean, I think NVIDIA-- and you see it-- and we talked about it at GTC, NVIDIA has an ecosystem already of seven chips and five different racks. So NVIDIA is no longer making a GPU. And I think that's what is one of NVIDIA superpowers that they've started to look at-- disaggregating the problem in all different ways, whether it's storage, whether it's CPUs, whether it's compute or networking chips. And so that already exists. So they had already thought about building a decode-only chip, and something that
[00:55:27–00:55:57] was powered by a lot of SRAM. But I think, sort of, it's a good lesson for everyone, like us putting that email in, starting to work together and building a prototype that was working with their systems was a real proof of concept for them. In these large systems, making these things work together, making them performant-- and this is across two different companies with two completely different stacks-- I think when they saw that we were able to do that, it showed that we'd be a good integration. And the last thing I'll say is we're two very different types
[00:55:57–00:56:27] of companies. I think if we were kind of making a better GPU, there'd be a lot of conflict within NVIDIA after the type of deal that we did. But because we were making this SRAM chip, deterministic compiler based, which is completely different than how GPUs work, it's very complimentary for the culture and the engineering teams to come together as well. How many people in here have used OpenClaw? I mean, that's pretty incredible penetration. I saw a stat-- Marc Andreessen may have tweeted this today,
[00:56:27–00:56:57] that most of the people he talks to are somewhere between $100 and $1,000 now a day on token consumption with OpenClaw. And he said, basically, the next 20 years of Silicon Valley is going to be producing technologies to drive down the cost of intelligence. And so I want to talk about that, Sunny. If we look at the cost of inference, it's dropped by basically 90% over the course
[00:56:57–00:57:26] of the last year. It's dropped by closer to 99% over the course of the last two, two and 1/2 years. So talk to us about what's driving the unit cost of inference. And if I take a like for like-- let's call it a unit of intelligence, whether it's a basic question I ask or whether it's a little bit more complicated question I ask, do you expect that unit cost to continue to go down? And if so, why?
[00:57:26–00:57:48] What are the inputs to that unit? Yeah. So the inputs are I think the following three major things. The supply chain, like what can you do across the supply chain which is mostly centered around Taiwan today. TSMC and the different packaging technologies, and the lithography technologies that they buy from others. The innovation that your engineers can perform,
[00:57:48–00:58:17] and what I would say is the amount of power you have. And so those are the kind of things we're talking about. And so what we see today is lithography technology is starting to reach a limit. We're not going as fast as we used to, and so we're not getting the Moore's law so we have to exceed that. So we're exceeding that in a couple of different ways. We're exceeding that by making bigger and bigger chips.
[00:58:17–00:58:46] And so if you see that these chips become quite large now, which is very exciting, but also lead to a lot of interesting issues. Cerebrus, as Brad's been talking about, their chips are kind of size like a pizza box versus CPUs you guys all would have seen. And so there's a lot of energy and technology there. Like, how big of a package can you make and how much silicon can you pack in there? Then there's the innovations. And so the innovations is really where we're seeing most of this work happen, because that's
[00:58:46–00:59:15] hand in hand with the models. We're seeing this really interesting force today, and there is a bunch of stuff that-- I don't know if it was leaked or put out there, and Elon is at the center of some of that, which is they're discussing these newer models are approaching like 10 trillion-- 1 trillion to 10 trillion parameters. And those 10 trillion parameter models go back to the first thing I told you. That's, in the fundamental flop calculation, how much compute it takes to generate a token. So as fast as companies like us are making
[00:59:15–00:59:44] better and better technology through lithography upgrades, through memory bandwidth upgrades, through innovation and how we lay out our circuits, through quantization efforts, NVFP4 was another one, the models are getting bigger and then the demand is increasing. So it's going to-- put it back to you, there's this three it's a cube, but it's really difficult to navigate right now because all the factors are growing in ways which are really challenging.
[00:59:44–01:00:09] So the demand keeps going up, the models keep getting bigger. And as fast as we're innovating, even if we get a 50x over five years, the models, the demand are going faster. And that's why we're seeing this unique phenomenon like H100 prices, if you're building a startup or using them, they're going up, in fact. Yes. Like one of the things that I think is important for everybody to understand, I mean, when OpenAI and Anthropic started,
[01:00:09–01:00:38] their gross margins on the businesses were highly negative. So that's a scary thing to do, go raise a lot of money, and it basically produce a widget for $1, and you're selling it for $0.20, and you have a big negative gross margin. But why was that? They were going out and they were charging you all to use ChatGPT, or their APIs were charging a certain amount of money,
[01:00:38–01:01:06] they weren't that capable so there was only so much money we were willing to pay. And two years ago, the cost of inference was a lot higher. But the bet they were making is that the cost of inference would come down a lot, and your willingness to pay would go up a lot as intelligence got a lot more valuable. So I like to think of the first inning of AI was just getting to a place where we could yield answers.
[01:01:06–01:01:35] In code generation it was basically like autocomplete, tab complete. In the case of ChatGPT, it was like basically telling a slightly better version of Google. But now we're entering into this phase of action where agents do things, go build me an app, go build me a website, figure out how to resolve this customer service problem, sell more of my product, find a cure for cancer, book me a hotel in New York. It starts doing things.
[01:01:35–01:02:03] And when it does things, the amount of tokens it has to consume in order to do those things explodes by an order of magnitude, but the value delivered to the end consumer as a unit of intelligence goes up by 100x. So your willingness to pay goes up dramatically. Can I add one to that? This week we saw Mythos-- which is the unreleased model by Anthropic-- find a bug in BSD which--
[01:02:03–01:02:32] think about how many engineers and software developers and everyone else and companies using that have looked at that code. So we've gone to a place where it's doing things beyond human capability, which is-- And we're in year three. Exactly. We're in year three. To give you another-- like, what's my best evidence to convince you of the value of AI? Well, my best evidence is that Anthropic in the month of March just added $10 billion in annualized revenue
[01:02:32–01:03:00] in a single month. That is the total amount of annual revenue for Databricks plus Palantir combined. And they added it one month. And they didn't add it because they hired a million salespeople, went out to a million companies and convinced them to buy their product. They added because their product crossed a threshold of intelligent capability
[01:03:00–01:03:27] that millions of customers around the world said, I have to have this product to make my company better. The amount that Altimeter is spending went up, but millions of self-interested actors around the world independently made a judgment, I have to buy a lot of those tokens, a lot of those capabilities, both Cloud Code and Cowork. And the same thing is happening in OpenAI, not quite on the same exponential in terms of revenue.
[01:03:27–01:03:57] But I think for me, this was a little bit of an Oppenheimer moment, this was a little bit of the splitting of the atom. Like, we've heard Dario and Sam talk about the exponential or the end of the exponential on intelligence, but the big question was, are they going to be able to afford to continue to build the compute in order to keep up with this? I had this somewhat uncomfortable moment with Sam Altman on my podcast, The BG2 Pod,
[01:03:57–01:04:27] that went a little viral when I asked Sam, hey, Sam, you've made $1.4 trillion of spending commitments, but you only have 13 billion of revenue. So explain to me how that works. Like, how can you commit to spending $1.4 trillion, you have 13 billion of revenue? And I had hoped that Sam would make the case that his revenue was going to go up a lot, and these were kind of call options and he could renegotiate them, but instead he said to me, well, if you don't like your investment,
[01:04:27–01:04:56] I'll buy back your shares. Which was not exactly the response I was hoping for out of Sam in the moment, but that was the question. Heading into 2026, my podcast partner, Bill Gurley, a lot of other people highly skeptical saying this is an AI bubble. These guys are spending at rates they're never going to be able to pay the bills on because there aren't people on the other end willing to pay for the products to justify that level of spending.
[01:04:56–01:05:23] And what happened in January was Anthropic had a $3.5 billion a month. In February they have an $8 billion month. And in March they have a $10.5 billion a month. That to me said, oh, everything's changed. The product is now sufficiently good that you have revenue scaling on the same exponential as intelligence, so they can afford to pay for the $50 billion per gigawatt
[01:05:23–01:05:53] to stand up all of these inference factories to produce all this kind of collective intelligence. Just react to that, Sunny, because our group talks a lot about this. There was a lot of debate in our group and on the All-In Pod and others as to whether or not this was a bubble. Yeah. I'd say there's of a couple of things that maybe the broader world doesn't see yet. One, the models that we see today haven't even been trained on the latest hardware, whether you want it to be Blackwell's or--
[01:05:53–01:06:22] Vera's are just coming out or Rubin's are just coming out, so we haven't even seen that yet and so we haven't seen the capabilities that you get. And so we'll start to see that. I think one of the first ones we'll see is the stuff out of Elon's Grok. So that's A. So the capabilities you're seeing here are things that were done on older hardware. So that's A. And so when you're inside the ecosystem, you know what's capable and what's coming next. I think B, one of the things that is really starting to take off-- and I think Anthropic has done an incredible job here,
[01:06:22–01:06:51] and I think Codex has an equally incredible job on very hard kind of software problems is that there's not just a chat interface that majority of people are interacting with, and it's not just an API, but they've created like a harness around the models, and those harnesses-- OpenClaw is just another harness as well-- those harnesses have figured out how to extract more and continually extract. I think with Cloud Code and Cowork you can have it just ping you whenever it's stuck on your phone, even if you start it somewhere else.
[01:06:51–01:07:19] And so it can be in this continuous loop and it's working for you all the time. We've never had anything like that. When it's doing that, you take that token consumption of like were doing a query before and was doing some thinking and coming back, now it's just working all night long and pinging you. You tell it, don't even bother me. Keep coming back. So we're seeing these harnesses really extract more and more tokens out of it as well. And the type of problems that people are solving, we gave the code problem, you put a bunch of other ones, but inside big businesses-- and I tweeted this,
[01:07:19–01:07:49] and I think it's fair, but like inside NVIDIA now we have this thing called the NVIDIA Personal Assistant. And it's connected to Slack. It's connected to Teams. But it's connected to our email, and it's connected to all our files, wherever they may exist. And so every morning it runs, and you're like figures out like all your task items for the day. You can have it answer those things, and it's really incredible. And so you start to-- the way we work-- and we we were talking about this earlier with someone-- like, you don't even write email now.
[01:07:49–01:08:15] Like, someone else's agent is going in their email, emailing you, and your agent is looking at emailing them back. But a lot more work is getting done because my time is freed up from basically answering emails all day long and approving things out of all these traditional SaaS systems. The agents handle all that. So the explosion, to your point, is just in the first or second inning, the amount of tokens is really just going up. So we don't fear that. We don't look at that as an overbuild in any way, shape, or form.
[01:08:15–01:08:42] I think the facts and evidence on the field is, number one, the cost of both training and inference, but inference in particular, is plummeting and continues to plummet. That shouldn't be altogether surprising. Technology ultimately is highly deflationary. I've never seen something this deflationary this quickly. I think it's a byproduct of extreme co-design. It's not a single chip, it's a factory.
[01:08:42–01:09:11] And across the factory there are all sorts of Moore's laws playing out combinatorially across the factory. At the same time, when you're able to produce a lot more tokens, the unit of intelligence that you're delivering is much more valuable, so the willingness to pay on the other end goes up a lot. And I'll tell you, for an OpenAI or an Anthropic today, if those guys were at negative gross margins a year and a half or two years ago, they're now at very positive gross margins.
[01:09:11–01:09:38] So all of a sudden, this business that looked diseconomic looks highly economic today. So it's resolved a little bit that question. Maybe just, Sunny, I want to finish our section with maybe just a little forecast and pre-wire you guys to-- we're going to open it up to questions. It can be about the economics of inference or any part of the stack or any other questions that you all have. But you mentioned Mythos.
[01:09:38–01:10:08] It's a model out of-- that came out this week. Was not generally released, but was sandboxed by Anthropic. Tried to get out a few times. Tried to escape the sandbox. Trained on TPU7. On the other hand you have SPUD or 5.5 coming out of OpenAI probably this week or next, which is a first Blackwell trained model. Elon's going to have one. Meta's just out with a model yesterday, Google, et cetera.
[01:10:08–01:10:36] Talk us through-- you get to see into the product pipeline at NVIDIA. Do you think that the pace of-- the cost of inference curve continuing to come down, do you think that continues for the next several years? Do you think that the step function or the exponential, if you will, of both pre-training and inference time reasoning in terms of improving the algorithmic capabilities of intelligence continues?
[01:10:36–01:11:04] Well, I can tell you, having a chance to work with Jensen now, he challenges us in everything we do, to not show up unless it's 100x. So whatever we bring to him-- and I can't get into too many details, but his first challenge back is, is this 100x from what you did before? So he is challenging the engineers to take a look at every part of the problem-- from all the way down into memory controllers or memory capacity or circuits,
[01:11:04–01:11:32] whatever it happens-- to be to make sure we 100x everything. So on the first part of your question, yes, because he pushes us to do it, and he gives us the latitude to do it and he gives us the resources to go do it. So I can tell you, like, the types of things that we've been enabled in coming in as the Grok team, things we could never do as a startup, but Jensen has enabled us to do those things. So are you guys harnessing AI yourselves to design the next generation chips? A ton. We were doing that even before because we needed to, we were a small team, but now we have
[01:11:32–01:12:00] access to the entire ecosystem of things that are available. So I think that's A, is that we're being pushed to do it. On the related side, though, the more we innovate, the more the model makers innovate and the bigger the models get. And so-- which means the capabilities that are coming out are better, so we continue to need that build out. So we'll all look back and we'll think, there's a couple companies that changed the footprint of the internet for us.
[01:12:00–01:12:28] And you could talk more about this and I can even brag, but the work that Google did to build the infrastructure they did for video, for search, it really paved the way for the rest of the internet, CDNs, all types of other things. And so a lot of this work that's happening to build out this infrastructure will pay benefits, and you need that to continue to happen, because it can't just be in the innovation of the chips. You need more and more infrastructure to be built. I'll wrap with this, we have the great privilege
[01:12:28–01:12:55] of talking with Jensen or Elon or Sam or Dario, and you guys all can read about the personal battles they have between those, some days, not a lot of love lost between them in the race to AGI. But right now I see amazing uniformity. When I talk to them, they all, in a non-hyperbolic way,
[01:12:55–01:13:23] say we're there, and we got there faster than we thought. Like, we're nearing the end of the exponential. And if you ask Dario, Dario, what is the most surprising thing to you right now? He says, we're almost at the end of the exponential and people don't even seem to realize it. And if you ask Sam, he'll say the same thing. And if you ask Elon, he'll say the same thing. That shouldn't be scary to any of us. It just means that we're in this recursive place
[01:13:23–01:13:53] where we have AGI. And the job of everybody in this room-- including the folks sitting up here-- is going to be how do we harness this technology for the betterment of all of us, which is going to require going back to what Apoorv said about the Invest in America Act. Dario says the accumulation of wealth that's about to occur-- people call it the age of abundance we're going to enter into, that's going to be easier than ever.
[01:13:53–01:14:20] But the distribution problems we're going to encounter are going to be harder than ever. And so that was really the inspiration behind the work I've done on Invest America and the work that I think we're all going to have to collectively do around the social contract, the intersection between public policy and technology. Because when the exponential looks like that and all of a sudden you have agents that are going to be able to have more capability than collective human intelligence--
[01:14:20–01:14:49] and it's happening at an accelerating rate. Remember, all the stuff we've talked about has occurred with almost no compute. Anthropic and OpenAI are going to add more compute this year than all the labs put together for the last decade. And the year after that, they're going to double it again. So I think the rate of change is fairly parabolic. And that to me is both exciting-- I'm an optimist about what's to come,
[01:14:49–01:15:18] but I'm not Pollyanna about the challenges that come with that rate of change. Like it's going to require active engagement, like it has in other periods in history around the Industrial Revolution, the Digital Revolution, et cetera, because it's going to exact a lot of change on the world. But with that, I just want to say, it's been extraordinary watching Sunny orchestrate the work that he's done at Grok. He's an incredible thought leader in this whole area.
[01:15:18–01:15:47] I appreciate you coming in. But maybe just open it up some questions, and hopefully we can cover a lot of territory. Right here. [APPLAUSE] As the marginal benefit of [INAUDIBLE] increasing wages are estimated, how do you suggest that a player positions themselves and to make sure we're not just wasting our time studying them? [LAUGHTER] Yeah.
[01:15:47–01:16:17] I mean, listen, I have to answer this question for my son and for so many others. And humans have a unique way of finding a way to add value to society, notwithstanding disruption. In the Industrial Revolution, if you were a tradesperson or a craftsperson, and you build a product beginning to end, almost all of them were displaced by mass production.
[01:16:17–01:16:46] And for that person, it didn't feel good. I was really good at making a wheel start to finish, but I was totally disintermediated by the means of production. OK. But it's not like we-- the world just stopped. Those people found other things to do. And one of the observations I have is that we used to have 80% of people that were in manufacturing and that were in farming and other things, today we have 70% of people in the service economy.
[01:16:46–01:17:16] We have the luxury of people-- we didn't used to hire coaches, as an example. You couldn't afford to hire a coach. Today you have coaches and yoga instructors and tons of things in the world that adds a lot of value to the world, and I think that we have higher order things that people do. And so for me, one of the things is if you are well off enough that you could hire a tutor, a specialized tutor, that was great. But for 98% of the world who couldn't afford that,
[01:17:16–01:17:45] now they can get that. Or if you were part of the 2% or 3% that could have concierge medicine, it was really great. But for the other 97%, it wasn't great. Well, now they can get that same level of care. And so I think this is about democratizing intelligence, democratizing access, et cetera. But it's not to say that there aren't going to be different challenges. My number one thing, again, is make yourself bionic, be a creator.
[01:17:45–01:18:14] Figure out a way that you add value. So if somebody comes and wants to interview at Altimeter, and they say, oh, I don't use AI and I don't use Excel Spreadsheets. I do everything by hand, that would be a problem. Like, I expect somebody to use all the greatest tools at their disposal to be the most effective they can, to add value to allow us to generate alpha in the world. And so starting in a place like this,
[01:18:14–01:18:41] another way of saying it-- reserving this for a tweet at some point in time-- but I think that IQ gets commoditized and EQ becomes super valuable. What do I mean by EQ? I mean a network of people in this room. I mean the ability to persuade the person sitting next to you, the ability to form your team, the ability to lead people in different directions, like, that is super valuable.
[01:18:41–01:19:10] I think it becomes more valuable in the future. But I think just being the smartest person in a room and solving the problem at the board faster than all the other humans in the room, like that, I think is commoditized and you're not going to be able to beat the machine. It doesn't mean you don't need to learn those things, but I think you'll be hard to beat the machine. Brad, I want thank you for doing this. We need BG2 back. We miss it. We only get Brad once in a while in the All-In Pod now. You only get your slot, so-- I need to get you back on.
[01:19:10–01:19:39] Yeah, let's get BG2 back. But can I add one thing to that. I think there's this other moment that's occurring right now, and I think about this quite often. If you actually look at what's happening in mathematics right now, there's all kinds of new discoveries happening. And I use the following analogy, like, humanity had to wait for an apple to fall on Newton's head for him to then start theorizing about gravity and start formulating that. But now if we can have something else working and discovering new things, like-- and it goes back to that chart that Brad showed.
[01:19:39–01:20:05] It's really until we started having more innovation, more intelligence that those curves went up and to the right. We're just about to make that go more vertical. So I think the overall benefit to humanity has already been shown what happens when you have more intelligence. And so we don't have to wait for things to happen. We let the agents do it without us in the loop, which I think will be powerful. There's one up. Yeah. When you think about the creation of hardware and software, particularly given that Apple
[01:20:05–01:20:33] is following the strategy that everybody else, [INAUDIBLE] all of them. And OpenAI is building a device which might actually try doing that. So I'm just curious as to how you think about that. I think it's a high stake-- I mean, listen, I'll tell you, even the people at Apple are nervous with their strategy. And so part of it is their challenge around privacy. They have a real challenge with-- because we don't have the capability yet on the edge,
[01:20:33–01:21:02] and they don't want to have you sharing information up to the cloud, given their-- they view as one of their core consumer value propositions is consumer property-- or consumer privacy, but I think they put themselves at risk. The bull case would simply be that we're so sticky to the device and the device is so good that they have time, and ultimately the Gemini model that they're going to put on the phone
[01:21:02–01:21:31] is just going to be a much more capable Siri. We can all agree that old Siri is really bad and will be a more capable Siri. For the vast majority of people, that will be good enough. So that would be the bull case. I think the bare case is, like you said, that other people come along and build more ambient devices that consumers really liked. But for me, I frankly wish that OpenAI wasn't working on a device. I wish they would just focus on building intelligence.
[01:21:31–01:22:00] And I think Apple is going to be very formidable in the device world, so think they're in a reasonably good spot. One stat, an $8 billion parameter model-- which is quite small-- can burn out a phone-- an iPhone in 30 minutes. It goes back just-- Battery life. Yeah, battery life. Yeah. So I just say you got to go back and look at how compute intensive AI is. And so that's the real challenge on the stuff that Brad said about pushing so much of that frontier intelligence to the edge. Over here, yeah.
[01:22:03–01:22:28] I want to touch on something we're talking about with Invest America. And I think it was on a recent episode of All-In, and Chamath was criticizing various tech CEOs for-- in creating hype around new product launches. So I'm curious what AI we're doing to national security agency jobs. And I guess I'm curious, what extent
[01:22:28–01:22:55] is CEOs' need to revise and rethink their messaging? He and I had that argument on the pod again today. Again, call me naive. I think that Dario's speaking authentically what he believes. I think Sam's speaking authentically what he believes. They're staring at this exponential. They believe that they see AGI or ASI, and I think they do have legitimate concerns. Like, listen, I'm glad that we sandbox Mythos.
[01:22:55–01:23:24] They tested it internally, they found 26 vulnerabilities on the Safari browser. And like I said to Chamath today, do you want them to just throw it out there and then all your browser history is out in the public? Probably not. So at the same time, I don't think it helps going out and fear mongering, particularly if your real intent is regulatory capture to prevent everybody else from climbing up the ladder now that you're on the top. Like, I have a real problem with that.
[01:23:24–01:23:54] But I think we have to find that balance and those trade offs between reminding people about the optimistic side of things. And I encourage you to read both of Dario's essays. And his first essay on this is quite optimistic about what can happen. But I think he has the other side of it, which is but it doesn't happen without us being very
[01:23:54–01:24:24] thoughtful about the guardrails and things we need to put in place. One of the things I was really happy about is called Project Glasswing, which is this consortium that they put together this week to effectively sandbox Mythos before they release it publicly-- Amazon, Microsoft, et cetera. Like that team seemed to me to be a very pragmatic, market-based solution to solve the problem. And he and I were just texting before I came over here. They found and they've hardened a lot of things already very quickly, and within 100 days
[01:24:24–01:24:51] you can do a lot when you're having the AI fix the things that it finds. And so I certainly-- I talk optimistic about it, I think a lot of other people do. Might encourage them to find a little bit more balance in their commentary, but I also don't want us to ignore the realities that when you split the atom, it can either provide unlimited free energy for the world
[01:24:51–01:25:20] and totally bring people out of darkness, or it can be used to make a bomb to destroy cities and nations. And so powerful technology is powerful technology. We can't just stick our head in the sand and act like it's a one way street. OK, there. Well, that's what [INAUDIBLE] we're mentioning in terms of [INAUDIBLE] is about 100 billion by 2030. So I'm just curious how you're balancing
[01:25:20–01:25:49] the fact that it's cheaper [INAUDIBLE] company wide. I think you see a couple of phenomenons. One, the gear that's used for training turns into inference gear, right. So those big clusters, we're seeing that happen kind of all the time, so there's just a natural progression between those two worlds. And then I really do think the innovations that come from those larger models and those training clusters
[01:25:49–01:26:19] have such a large benefit, and they tie back to what Brad said. I was just reading this thing today by Mustafa from Microsoft saying, look, GPT 2, what-- we have 50x more powerful compute today than what we did when we did GPT 2, but look at the capabilities. And so I think you just have to keep those two things in line with the entire topic of the conversation. Innovation is going to keep happening because--
[01:26:19–01:26:49] Brad touched on it a little bit-- we're just now unleashing AI into designing these things. And there's things that we see and we learn in terms of optimizations and software and hardware optimizations that we don't see. So I continue to believe it'll come down, but, yeah, we're just working on a problem that's just very, very intensive from a compute standpoint so those numbers are going to be large. Maybe one final question. You can pit dive with Sunny and I, we can answer some questions after the fact,
[01:26:49–01:27:15] and-- but otherwise I just want to say it's a great privilege for us to get to spend some time with you guys, so thanks for having us. But maybe right here. Yeah, hi. Thank you for sharing. Just one question about Columbus. [INAUDIBLE] So, that's going to be our major shareholder in NVIDIA and sending your [INAUDIBLE] immediate one. So what do you think is the long term sustainable economic model
[01:27:15–01:27:44] for NVIDIA, because right now [INAUDIBLE] if you have 60% ever, right that makes it very difficult for the entire ecosystem because they take way too much time. What do you think would eventually would happen? Let's say a few scenarios. A, their revenue may be limited, but they keep their margin.
[01:27:44–01:28:12] And two, is somehow lower their margin substantially and be able to keep a much bigger market share. So, maybe just repeat it real quick and then answer it since I'm an investor NVIDIA. I work there so I shouldn't answer that. No. I mean, listen, NVIDIA is a $4.5 trillion company that's trading at about 13 times earnings. Very cheap, half the market multiple. Growing at 70%.
[01:28:12–01:28:41] Is obviously dominant in the market today. And I think there's a wall of worry about NVIDIA because everybody says what you do is Trainium and TPUs and Cerebras and Grok. And all these people can come up with inference solutions, they can steal your share, they can compete on price. That's the beautiful thing about capitalism. You know what NVIDIA will have to do? Compete. They either deliver a product that people are willing to pay more for, or they have to drop their price,
[01:28:41–01:29:11] their margins come down, and they'll have to compete in that market. I would tell you that when I look into the product roadmap for what's going on at NVIDIA, and the acquisition of Grok was part of it, I think they're going to be in an incredible position. They've already announced that they have a trillion, a trillion of sales over the course of the next eight quarters that are already booked. People have more demand than they can get memory and supply to build all of this compute out. So I think we're so early in this.
[01:29:11–01:29:39] I was in Silicon Valley not so long ago, 16 years ago, when they said there could never be a trillion company, there would never be a trillion-- I'd ask the question, well, why? So they say, well, law of large numbers. I was like, well, what stone tablet is that etched into? It nowhere. Today we have a $4.5 trillion company. I've already said publicly NVIDIA will be the first $10 trillion company. And I'm not-- it's not because I'm a cheerleader.
[01:29:39–01:30:07] I can sell my NVIDIA and invest in anything that I want to invest in. But that company's leadership, that team, the lead that they have on both training and inference, and the rate at which they're moving I think puts them in a really great competitive position. And they're doing all this notwithstanding the fact that Tranium is successful, TPU is successful, custom ASICs are being very successful and they're still killing it. And I think that says a lot more about the size of the market
[01:30:07–01:30:20] for intelligence and the compute that's needed to get us there than it does about the individual company. With that said, we have to wrap. Apoorv, thanks for having us. Thank you.
[01:30:34–01:31:03] Today, we are going to talk about the data centers that you guys are melting. The big theme of all of this is really this chart. This is the CapEx spend by the five hyperscalers on AI. And as you can tell, this is going top into the right and it's going top into the right fast. To put this in context, this is one of the biggest investments that we are making.
[01:31:03–01:31:31] This is bigger than space, bigger than our highway system, the Manhattan Project, second only to the US Defense budget. And I'm so excited that we're going to break it down today with Chase Lochmiller, founder and CEO of Crusoe, who is arguably building it the best that we know. A quick introduction for Chase before we bring him on stage. Chase is obviously the founder and CEO of Crusoe, as you guys know.
[01:31:31–01:32:01] Less known fact and fun fact about Chase I learned very recently is that Chase is a very avid mountaineer, five out of the seven summits across the world, including Everest and Crusoe, is designed with mountaineering in mind. There's a plan A, there's a plan B and then there's a plan C to the plan B. Chase, thank you so much for doing it with us. Please join us. Thank you. Thank you, thank you, thank you. Thanks for having me. Thanks for joining us. Yeah.
[01:32:01–01:32:29] So, Chase, tell us a little bit about what is this data center economy we're in the middle of. This is extra special for me. I went to grad school here at Stanford. So it's fun being on the other side of the table, being back on campus. So I appreciate you guys taking an interest in what I'm working on now here at Crusoe. But I think in certain ways, like the data center itself is like this physical manifestation of this boom that we're seeing in AI adoption.
[01:32:29–01:32:55] It's the physical infrastructure that's required to power the GPUs, to operate the GPUs, to run these big compute workloads that are training new models, that are fine tuning models, that are operating large scale inference workloads to serve tokens to consumers and everybody that raised their hand today saying that they use Gemini, ChatGPT, or Cloud today. So the data center is the physical infrastructure
[01:32:55–01:33:21] component that really enables all of this technology to proliferate and change people's lives. Amazing. Now there's a lot of different components that go into the data center phase. We've heard a thing or two about compute. We've heard a thing or two about memory and power and interconnects and labor and help us put into context all the different things that go into it, how should we contextualize it
[01:33:21–01:33:48] when we have our hyperscalers spend $650 billion building these data centers? Where does that go and how much of that is going to you? That's a good question. When I really think about this infrastructure of intelligence and really the starting point that I start off with is like, what does it take to produce AI? Like, everybody's all obsessed about AI. It's like, what does it actually-- what's required to make AI?
[01:33:48–01:34:16] You have this basic equation, which is AI is the combination of data, algorithms, so back propagation, neural networks, transformer architectures, all these different algorithms that people have come up with to essentially statistically model the data sets, compute large amounts of compute, and particularly this high performance computing infrastructure through GPUs, where you're able to parallelize a lot of the workloads and do a lot of this tensor math
[01:34:16–01:34:44] and a high performance architecture, energy that's required to run those GPUs and then data centers, the physical buildings that house and operate all of this computing infrastructure. So what actually costs money? Well, data-- sure it costs money. I mean, you have to buy data. It's like opened up this new opportunity for a lot of data labeling companies. Folks like scale AI, folks like [INAUDIBLE]. Even handshake has sort of gotten
[01:34:44–01:35:12] has sort of bridged into this. It's created a big opportunity for people to make money by producing data that's useful for AI. The algorithms, that sort of sits in a lot of the labs that are inventing new mechanisms. And recursive learning techniques is a new trend and a lot of the different architectures that people are using to make better use of the data. But the compute energy and data centers, that's really what Crusoe folks focuses on
[01:35:12–01:35:42] is like, how do we actually build, operate, scale and make the best use of all of this infrastructure? And it tends to be actually the area where a lot of the money is being made or where a lot of the money is being spent. Sorry, that CapEx chart that you showed. So what I would say is, the name of my presentation, I came up with was from electrons to tokens. So why are tokens actually valuable and given this isn't the econ department? Engineering. It's an engineering. OK, well. Some of you have taken economics. Of course.
[01:35:42–01:36:10] There's this economics model for production called the Cobb-Douglas model. So if you look at the growth in GDP, the growth in GDP is fundamentally the sum of three key components the change in labor, the change in capital, and this capital sort of includes both like physical capital like buildings and plants and equipment, as well as capital that's invested
[01:36:10–01:36:40] into an economy, and then the change in technology. And technology basically makes labor more productive. I think is a good way of framing this. So why are tokens so valuable and why is this like a step change? Why is everybody making these huge investments, this CapEx chart that Apple showed? The reason for that is that for the first time in history, what we're able to do is we're able to actually create this sense of digital labor. When you give when you give an agent a task, when
[01:36:40–01:37:09] you give your cloud bot a task of going to doing something, go create a CRM for my new product that I just launched, that is literally digital labor that's being represented and that's being brought into the world. And historically, labor is this thing that you could really only change via the birth rate. And it's got a 20-year lead time. It's a massive incubation period. You got to send them to schools and feed them and house them
[01:37:09–01:37:37] and whatever. Like all this stuff. I have three kids. It takes a long time to change changed delta L, naturally. But for the first time in history, what we're able to do is actually change this delta L digitally through the investment in buying data centers and buying GPUs, and really accelerating the growth in the economy by actually accelerating the growth in digital labor force. So that's kind of like-- the premise I want to get across
[01:37:37–01:38:06] here initially is that the reason these investments are taking place, and the reason it's so broad based is that there's an opportunity to completely transform the economy by really accelerating the growth in GDP and fundamentally up leveling and improving people's quality of life by just seeing an unprecedented level of growth. So, I said, Crusoe sits at a couple different layers of the stack. Crusoe is a business, is a vertically integrated AI infrastructure business. So we really try to-- because this
[01:38:06–01:38:35] is a new category of infrastructure, we've really taken the approach that we want to be able to unblock anything that really gets in our way of standing up this infrastructure of intelligence that's going to power the accelerating growth in GDP for the economy. So we think about it in two phases. One is like the bottom layers of the stack, which are basically energy development. So these data centers require lots and lots of energy. And we're very thoughtful about taking this energy
[01:38:35–01:39:04] first approach in terms of going to areas that have access to abundant low cost energy resources. The next piece is really the data center piece, which is the actual physical buildings, which includes the building, the plant around it, all of the chillers. Because when you think about it, like from its most basic standpoint, a data center is a building that has power and cooling. And you can plug in computers. That's pretty much what it is.
[01:39:04–01:39:33] And when you're doing it at very, very large scale, it becomes very, very complex and actually draws in the pinnacle of engineering and from across every single ecosystem from chemical engineering and cooling architectures and mechanical engineering and electrical engineering dealing with these high voltage power sources and very high density capacities and then also the computer science and the electrical engineering involved
[01:39:33–01:40:01] in the chip architectures and how compute gets run, how data gets transferred. It really is the amalgamation and consolidation of like every form of engineering and one giant building that operates intelligence for the world. So it's a cool engineering problem. Yeah. On this chase, tell us a little about-- we hear the bottleneck in AI being compute four years ago to power,
[01:40:01–01:40:31] memory stocks are ripping right now, labor, then there's all these other components LAN-powered shell, et cetera, maybe even regulation. Give us an overview of for the time that you've been doing this, which is just under a decade, that seems to shift. How has that traversed over time? Where is it today, and where do you see it going? What is the core bottleneck that is gating the growth of this? Like today, the core bottleneck is like energized data centers,
[01:40:31–01:41:00] like powered shells where you can plug-in chips and start operating a big GPU compute cluster. That today is the bottleneck. But the bottleneck moves around a lot. Sometimes it's getting power to those data centers. Sometimes it's individual components that go into building the data centers, electrical equipment, switchgear, chillers, power Gen, chips. Chips is kind of softened as the bottleneck. I think access to chips has become more available.
[01:41:00–01:41:28] It's really finding places where you can put those chips and turn them on. So that's part of the reason that Crusoe has taken this is vertically-integrated approach is that bottlenecks move around, and being vertically integrated means you can do almost anything across the stack. We're not in the chip business. That's like we're not in the chip business. We're not in the model business. But apart from that, we tackle, most challenges throughout the entire ecosystem. Right now maybe just to follow up on that, Chase,
[01:41:28–01:41:57] you started Crusoe with the core insight of Robinson Crusoe, with the insight that energy was one of the most scarce resources, at least in the Western world. And tell us a little about the-- maybe pick a site, maybe pick Abilene or one of the others that you can talk about. Why did you start there? What was the scarce resource that you were solving for and why work backwards from energy? Why not from compute or memory or as you said, power shell?
[01:41:57–01:42:26] Well, I guess our insight was like markets are reasonably efficient when you look at things. And I think when there was this steady state growth in data center capacity as the Web 2.0 bubble or not bubble, but Web 2.0 trend unfolded and web applications were increasingly growing and people were more online. It became this machine that was standing up new data centers.
[01:42:26–01:42:53] And they were happening in these big hubs. Markets like Northern Virginia come to mind as areas that run a large portion of the internet. I never wanted to be like a me too. Like, I'm the next data center developer in Northern Virginia. I'm going to build the next building there. That didn't seem like an appealing way to enter a new market and really make a splash. So we said, look, there's going to be these new types of computing applications that are far different from serving web applications,
[01:42:53–01:43:21] things like artificial intelligence training, large workloads and back propagation, things like digital currencies that require tremendous amount of computing power for proof of work consensus mechanisms. Those require tons of energy and at scale energy becomes the bottleneck. So we said, can we find areas that aren't historical data center markets and go there and build data centers where we can instead of having to move the energy, we're actually moving data.
[01:43:21–01:43:49] So we could actually co-locate in these areas. So giving an example here and we'll talk about the top two layers of the stack, both the deployment of GPU clusters as well as how to actually monetize that and serve intelligence with managed services. So we're going to talk about the bottom two layers of the stack here. So one first manifestation of that was what you see in this photo here. So this is a site that we've been working on since--
[01:43:49–01:44:17] it started in June 2024, where we signed the first two buildings on the right-hand side of the screen that you see. This is at this point, one of the largest AI computing campuses in the world. I think it might be the largest. Our insight here is in Abilene, Texas, many folks had never heard of until we put a shovel in the ground in Abilene. So why do we go to Abilene? Abilene is this area of West Texas that is consistently very windy and very sunny. And so a lot of renewable energy developers
[01:44:17–01:44:46] had gone there to build out large scale renewable energy production because they were incentivized by something called production tax credits, where they basically get paid by the government to produce clean electrons, and they have to sell them to someone independent of the price. And what that resulted in was actually an overinvestment in renewable generation infrastructure in this West Texas market. And power prices were actually negative because there was no marginal buyer for this power,
[01:44:46–01:45:15] and there wasn't enough transmission to get the power to somewhere where it actually was useful. So we said, great, have we got a power hungry application for you? So we ended up working with the city of Abilene. If you look at the top there, there's this gray square. And then below that a bigger gray square. So those are both substations. The top one is a 200 megawatt substation.
[01:45:15–01:45:44] The second, larger one is a one gigawatt substation. To put those numbers into context, like that gigawatt substation-- first of all, that's the largest privately owned substation in the United States. A gigawatt is-- I grew up in Denver. A gigawatt is basically what powers the whole city of Denver. So it's basically a city of Denver size worth of power to power computers. It is a very large amount of power in one single location.
[01:45:44–01:46:12] And we were able to access it fundamentally because there was this abundant, low cost energy in this market that was actually having issues getting out, having transmission to get out of that market, created a massive opportunity for us to go in there and build large scale, cutting edge AI infrastructure. A couple of follow-up searches, who is this tenant for? Could you tell us is this ChatGPT, is this, Cloud, is this Gemini? So the tenant of the--
[01:46:12–01:46:40] If you look at this campus, there's eight buildings. So this building 1, building 2. That is the substation I was just referring to. And then you have building 3, 4, 5, 6, 7, and 8 up there. So those first eight buildings are all for Oracle and OpenAI. So it's basically what was known as Project Stargate, this first big project. So in order to help support this, we also built down here a natural gas power plant.
[01:46:40–01:47:08] So this is roughly a 350 megawatt natural gas power plant to support the development and energize this giant computing cluster. It was also designed as to operate as one coherent cluster, which means that all of the chips across all of the data centers were interconnected on the same high performance back end network to be able to operate as one coherent workload. So you could run one training job
[01:47:08–01:47:37] that runs on all of the chips, on the entire data, on all the data centers together, which is really, really unique architecture. To give you a sense of scale, because it's hard to see from a photo we had to build this parking lot over here. That parking lot is like-- it's a 5,000 car parking lot. And you can see it's totally full because we have roughly 9,000 people on site every single day that are working to bring this campus to life.
[01:47:37–01:48:05] There is an expansion planned to the South of the campus and see some of the dirt there. That is for Microsoft. So the campus is 2.1 gigawatts in aggregate. So again, two Denver's worth of power to power all of this AI compute infrastructure that's going in here. 9,000 people, Chase, what is the population of Abilene? 9,000? Although there's another campus I'm going to show you
[01:48:05–01:48:31] that on site staff is larger than the population of the town. This is the Manhattan Project. The population of Abilene is 120,000. So we're able to source, a lot of people initially from Abilene, but over time, we had to actually create a lot of these labor and retention incentives to get people that would move there to basically work, as short-term construction workers.
[01:48:31–01:48:59] Over time in the long-term there is like a steady job population that's operating these large computing clusters and power plants that are operating the infrastructure. That staff is somewhere in the neighborhood of like 2000 people just to put it in context, but it still becomes a very, very large job creator in this local economy where the population is roughly 120,000 people. So-- Go ahead. No, no, please.
[01:48:59–01:49:28] While the audience is primarily engineering the classes, the economics of AI, so walk us through the metaphorical spend of $100. So if you had $100 or x dollars of spend, what is the distribution of it across the different layers? Yeah. Or maybe you were coming to it? I'm going to come to it. We're going to start with just initially the power plant and the data center side and then we'll go to the compute clusters and then we'll talk about--
[01:49:28–01:49:58] these are tens of billions of dollars that are being invested in this. How is anybody making any money? Where does the return happen? So we'll get to that as I progress through the slides. But the initial slide that Crusoe showed was this huge CapEx spend. A lot of those companies are Crusoe customers. We help serve all those customers when they're building out these big CapEx investments, and we help them build out this infrastructure of intelligence layer. With that comes-- I think there's
[01:49:58–01:50:24] a bunch of different components that go into this. But electrical equipment is a huge component of this. So think of if you look at the-- there's over here, there's these small buildings with white roofs on them. Those are called power distribution centers. They take power from the high voltage or from the substation, which comes in at a medium voltage 34.5 kV,
[01:50:24–01:50:52] so 34,500 volts, and then it distributes it to that lineup. You see this where it says transformers. There are transformers that line the entire left side of the building, and then the entire right side of the building. And what it's doing is distributing power to those transformers so it can step power down from 345 kV to 480 or 415. So that's like a big piece of like CapEx. It's a piece of equipment that we're investing in. It's going into building out the building.
[01:50:52–01:51:22] You also have all sorts of different cooling equipment, all of the mechanical equipment. Think of you have this lineup of chillers. They look like RAM, but they're not. So those are all air cooled chillers, which is basically these wound pieces of copper pipe where you have this giant chilled water loop in the data center that's recirculating water, and it basically goes from cool water on the inlet side. It goes into the rack of GPUs, and there's a thermal transfer
[01:51:22–01:51:50] event from the GPU that's being energized and producing a lot of heat. There's a thermal transfer event from the chip to the water, and then the water goes out through these chillers. You basically are blowing a bunch of air over these wound copper coils to exhaust the heat out of the system. So the water temperature steps back down and you have cold water to then cool the GPUs again. So again, that's another big investment. All of the plumbing that goes into this
[01:51:50–01:52:18] is really, really substantial. So we actually we have a ton of plumbers and pipefitters on site that are welding these big plumbing systems together. Each building has about 1 million gallons of water in the building to cool these chips. But again, it's recirculating. So think there's this ongoing narrative that AI is taking all the water. We use, like zero water. We fill this system one time and then on an annual basis, we use about the same amount of water as a single family home.
[01:52:18–01:52:48] So it's a very limited water consumption usage, which is important in a market like Abilene in West Texas, where water is actually quite scarce. So you can see the costs here. And I've normalized them on a per megawatt basis. But, things like power distribution centers, the UPS system, which is a battery system, uninterruptible power supply, you need this to smooth out the power that gets distributed from the substation into the actual chip. Other alternative battery systems
[01:52:48–01:53:16] that we're experimenting with cooling distribution units, these actually go in the data center itself, and they basically take the water in from the chilled water pipe and they distribute it to the individual racks of GPUs. And then of course, you have a lot of these core components that go into this. There was nothing here before. There's a ton of steel. There's a ton of concrete. We have our own batch plan on site. So we're actually making concrete on site with people pouring concrete 24/7.
[01:53:16–01:53:45] All of the site work, all the labor that really goes into making this happen, it's tons of people, tons of man hours. And then on the power infrastructure side, we highlighted two things here. One is the gas power plant, which I showed you in the previous slide. You see a little snapshot of it down here on the bottom. And then for some of the infrastructure we have diesel generators. So you can see these three gray roofed buildings that are in the middle there. Those are backing up power for the core network.
[01:53:45–01:54:14] So the way we've really thought about this problem is that not everything needs 100% sent five nines of reliability full backup. But the core storage and networking systems do so that in the event of a full grid outage, in the event of a disaster scenario, we'll still be able to access the storage systems. If there's a checkpoint that we need to reference and move a workload to a new location, we can do that. So anyway, it's a lot of money that goes into this.
[01:54:14–01:54:43] And I wanted to give you a breakout the full power plant plus-- the full power plant plus like building costs. And I want to highlight something for you guys because I showed that really big parking lot, that 5,000 car parking lot. Well, the bottom piece here is labor. So this is $4.7 million per megawatt. So that ends up being for a gigawatt or for megawatts-- for a gigawatt that
[01:54:43–01:55:11] becomes $4.7 billion. So this is literally money that's being invested in people to do jobs. It's like literal-- it's a blue collar labor force that we're investing in to basically bring this infrastructure-- people working in construction, it's people on site that are building these things. And so it's a very, very substantial number when you look at it in the scheme. And that's an annual number. So-- When you ask me about bottlenecks-- this is a bottleneck.
[01:55:11–01:55:38] We don't have enough of these tradespeople. We don't have enough electricians. We don't have enough welders. We don't have enough plumbers. We don't have enough construction workers because this is one project. There's many of these that are now cropping up. And there's actually a huge competition for labor. And so at Crusoe, we're trying to reinvent how we think about bringing a lot of this infrastructure to life to be able to navigate some of these really critical labor challenges.
[01:55:38–01:56:08] So if you look at soft costs, what does that include, that's things like insurance, that's things like financing costs like we borrow money with a construction loan and then we have to pay-- we have to service the debt for that construction loan. That's things like siting and all of the different work we're doing with commissioning. The next piece is the gas plant. And again, this is probably 2 to $3 million per megawatt. I think what's important to realize about the gas plant is that those costs have gone up a lot.
[01:56:08–01:56:37] There's a small set of gas turbine manufacturers. You basically have GE, Vernova, Siemens, Mitsubishi Heavy Industries, Pratt and Whitney. Caterpillar has a company called solar. But a lot of these companies have been limited in the amount that they've actually expanded production capacity. And so what's happened in a moment where everybody's trying to bring on new gas generation infrastructure to power their AI, compute clusters-- well, guess what? Prices have gone up a lot. So a gas turbine that used to cost $1 million-- a megawatt now
[01:56:37–01:57:05] costs $3 million a megawatt. So those prices have grown. And that's why you've seen-- I don't know how many people follow the stock market, but if you look at the stock price of GE Vernova-- it's been good to be a shareholder of GE Vernova. So anyway, that's another piece of the stack. The tenant fit out, this is all the stuff that's in the actual data hall. So these are things like remote power panels, the hot oil containment systems, the fan walls, the cooling distribution
[01:57:05–01:57:34] units, all of this stuff that you need in the actual data hall, where the GPUs go to actually energize and power the GPUs. The electrical equipment, again, this touches all the different pieces from high voltage to low voltage. This is things like power transformers, power distribution centers, medium voltage switchgear, low voltage switchgear. Think of your electrical panel in your home that if the lights go out, you go down, you flip a few breakers. It's like that but at the scale of a city,
[01:57:34–01:58:03] like all in a giant electrical room, the mechanical equipment. So all of the equipment from the chillers and all of the plumbing, all of the air handling units and the fan walls that sort of cool this stuff and have mechanical systems involved. And then of course, the materials, the steel, the cement, all of the different components that go into making one of these large buildings happen. So anyway, that's what the full stack looks like. A couple of quick questions here, Chase. Where are the GPUs? Oh, that's on the next phase. OK. Yeah, so we'll get to that.
[01:58:03–01:58:33] So this is roughly $20 billion per gigawatt. And assuming a gigawatt took a year to come unlined, you would be paying 4 and 1/2 or $5 billion in salaries for labor per year. That's for the construction period. This is like capitalized labor. So this is like-- Not OpEx CapEx? Yeah. This is not OpEx. Right. I'll show OpEx in another-- Great. And then one final question, is this number total, $20 million per megawatt or $20 billion per gigawatt going down over time or going up over time?
[01:58:33–01:58:59] Because obviously some components-- It's going up because there's so much demand. So things like gas generation, infrastructure, guess what? Prices have gone up. Things like labor, if you're an electrician, your price has gone up because there's so much demand for your time. Great. So all these things you're seeing price inflation to varying degrees in each category of this infrastructure. So I wanted to show another cool project that we're doing just
[01:58:59–01:59:29] to say how we're taking this energy first approach and to your comment around how many people. So today I think we have 3,500 people at this site. It's in a town called Quad, Texas, which is like a hilariously poetic. We didn't name the town, but Claude is a town of 1,500 people. We have 3,500 people working on this project. It happens to be close enough to Amarillo that we're able to tap into a lot of the working population in Amarillo.
[01:59:29–01:59:58] You can see it's an area that's very rich in renewables. It's one of the best places in the US to build wind because it's so consistently windy. So you can see that on-site wind farm that we have there on the top of the screen there where there's a very large wind farm that's producing power that directly feeds into the data centers. We're able to firm up the power with something-- we call this across the meter. So basically the meters, the interconnection point into the grid. And we have behind the meter, which is like power that's on-site. But what we're doing is what I call
[01:59:58–02:00:27] across the meter, which is we have on-site generation through wind. There's a plan to build solar batteries and gas. So all of the above energy solutions to basically energize this campus. What we don't need, guess what? We can sell into the grid, create an energy abundance that drops the cost for all local ratepayers. And when we actually need power, because we need to firm up the power, because we're doing maintenance on some of the generators or the wind's not blowing and the sun's not shining, we need to firm up the power, we can draw power from the grid.
[02:00:27–02:00:57] So it becomes this very mutually beneficial relationship of us investing in the power infrastructure and then leveraging the large distribution, transmission and other generators across the grid. So it's a very cool project. I can't speak about who the customer is but it is a very big customer for this location. You asked about the GPUs. Great segue. All right. Who's making money in this guy? No if you didn't know already, it's got a building right over here. When we think about the IT CapEx per megawatt-- remember I just showed you this is for the whole data center
[02:00:57–02:01:24] and the power plant was roughly call it 20 million megawatt or rounding up. When you look at the IT CapEx, this is basically the compute infrastructure that's going into the building. It's roughly 40 million per megawatt. And this is forward looking. This is next gen stuff. 30 million of that is going to the GPUs. That's why you always look so happy. I don't know. It's always smiling Jensen. And then where does the rest go?
[02:01:24–02:01:52] 4 million roughly to the networking. These are very complex networking systems, especially when you're thinking about the investment of interconnecting these GPUs together as one giant coherent cluster. When you look at the latest generations of GPUs-- right here is the GB300 or GB200. I'm not sure, but what NVIDIA has come out with, there is actually a full rack design, which means that all of the GPUs, there's 72 GPUs in that rack,
[02:01:52–02:02:21] they're all interconnected on the same NVLink domain. So you can see the back copper plane there on the right side. They're all interconnected on this high performance back end networking domain, which enables AI researchers to do incredibly high performance tasks and enables a lot of incredible use cases to be able to share data across the NVLink domain. But then you have to interconnect those racks together through another high performance back end
[02:02:21–02:02:51] network, typically InfiniBand or sometimes rocky, which is RDMA over connected ethernet. So that's where that four million per megawatt of spend is, that green bar of networking. The next thing is CPUs and storage. What's amazing, what we've been seeing recently is actually a massive shortage of CPUs. Why is that? With the boom in all of these agentic workflows with the boom and Claude, guess what? You need a lot of CPUs to actually orchestrate those compute workloads.
[02:02:51–02:03:18] So you're seeing a lot of demand from everybody in the ecosystem to bring online a lot more CPUs. So about 3 million a megawatt for CPUs and storage. And then there's a bunch of in the room CapEx, which I think I might be double counting here, but a lot of that TFO that I sort of referred to earlier. It's roughly 3 million a megawatt, and then you have about 1 million in labor deployments, shipping, et cetera. But you get to this roughly number of about 40 million per megawatt.
[02:03:18–02:03:48] One question here, Jensen did a podcast yesterday with Dwarkesh where he looked not so happy. Lots of debate around compute being a commodity. Yes or no, is this number going down over time? Is compute a commodity? What are you seeing in prices? It's tough to say. I mean, it's hard to say over how things look over the near term, medium term, and long-term. And it's possible that both are right. Where it's like-- and it depends on the use case too.
[02:03:48–02:04:18] Like we see if you look at older compute, it's kind of commoditized as you like the H100 further back. I think one of the things that's absolutely not a commodity is scale. So if you do anything at really, really big scale, it's super hard to replicate and it's super hard to repeat. There's always going to be a cutting edge. So any of the newest stuff is always going to command a premium. And that's like been the history of the IT industry. So we'll see how it kind of plays out. But I do think that folks--
[02:04:21–02:04:50] look, capitalism is a powerful force. The invisible hand of capitalism is a very powerful force. So I do think that over time, margins do probably come down to more standard stabilized silicon margins, call it like, I don't know, 60% gross margin, which today NVIDIA is commanding like 80% gross margin, something like that. I don't know. Competition is powerful. Yeah. Perfect. So I think it's important to understand
[02:04:50–02:05:20] you make this huge investment. We're talking initially, call it $20 million a megawatt for the data center and the power plant and another 40 million megawatt. To stand up this compute cluster, you've just spent 60 million per megawatt. So you build a gigawatt cluster, you just spent $60 billion. How are you going to make money? What's the pot of gold at the end of the rainbow? People are buying this infrastructure to support all of these AI applications, to serve tokens to customers.
[02:05:20–02:05:47] And I think the reason I wanted to show this chart is actually-- this is a Bloomberg chart of basically each 100 spot pricing. And I think there's this question of how valuable is all this equipment and what timeline do you depreciate it over? What's the useful life of all this equipment? And I think people said, the next generation is going to come out and then this stuff's going to be completely useless.
[02:05:47–02:06:17] Well, this chart tells the exact opposite story, which is that for each 100 that initially debuted in about three years ago, the pricing had come down. But with these boom that we're seeing in demand coming from agents, the price of H100 has actually come up and actually exceeded the price that folks were paying when these chips first came out, and that's something we're experiencing, experiencing firsthand on the ground in our crucial cloud business. Does the tangible outcome of this Chase--
[02:06:17–02:06:47] right now most public companies depreciate their compute over five years. Does this imply-- six is the standard. Does this imply it goes longer than six? I don't know. It's like my honest answer. It's like we're going to use compute so long as it's valuable to us or to someone else, so long as we can. And part of our strategy has been building services that abstract away the layers of compute. So you don't know if you're using an A100 and H100,
[02:06:47–02:07:16] an MI300, nor should you care. What you care about is the actual service that you're getting from that. Just like when you log in to Zoom or Google Meets or teams, you're not thinking about, wait, is this like an Intel, Ice Lake CPU, or is this like an AMD. It's like you don't care what the chip is that's running. You care about the service that you're getting and the fact that you're able to log into this video chat and be able to hear and speak to the other person
[02:07:16–02:07:45] on the other side. Makes sense. So we think the application scaling is really going to abstract away a lot of the core infrastructure. And I think we'll see how valuable stuff is over the course of time, but I think it's probably wrong. And then this is a similar chart that's semi analysis, who many of you are probably familiar with published. And this is for Blackwell's. So we see the pricing for Blackwell's following a very, very similar trend with this agent
[02:07:45–02:08:15] breakthrough in end of year. So when we look at this as a holistic picture. We're bringing it all together. We have the data center. We have the power plant, and we have the chips. The upfront CapEx you're looking at is close to 60 million per megawatt. Again, I think I double counted something in there. This is the first time I'm going through these slides. But it's roughly 60 million per megawatt. And then when you look at the ongoing OpEx of this plant,
[02:08:15–02:08:44] it's a little over 1 million per megawatt. It's actually pretty limited OpEx. And this is for things like your power, your insurance, some of your labor on site. That's like repairing and replacing cables and GPUs that fail and a number of other things, but call it like one to $2 million per megawatt. So what is your revenue if you're just renting out those chips? And I was using that chart before to show you rough pricing that you could rent an H100 for. What is your revenue per megawatt?
[02:08:44–02:09:12] It's roughly 15 million per megawatt. So you're making this upfront capital investment of $60 million a megawatt. You're getting $15 million a megawatt in annualized revenue for just renting access to the infrastructure. Now, how does this become a good business? I think a lot of it comes down to how do you measure the depreciation of all the different bars that make up this CapEx numbers. And I think that's the critical question analysts on Wall Street
[02:09:12–02:09:41] are asking is like, how long is this building going to be valuable for? How long is this chip going to be valuable for? How long is this power plant going to be valuable for? What's the right depreciation curve? but you're looking at-- from this, call it a four-year payback period for this huge investment. On a revenue basis. And what are the rough-- Well, this is what I'm saying. You have the OpEx stripping out the OpEx. Whatever, great. But then there's other labor that's not included here, which is like all the engineering
[02:09:41–02:10:09] workforce, et cetera. So there's more OpEx than this. Again, this was put together this afternoon. Makes sense. Roughly four years. Yeah. Yep. But what's another way to actually improve the value that you're actually delivering to customers? Again, I sort of spoke about this vertically-integrated strategy that Crusoe has. When we deploy chips, one product that we offer is that managed compute cluster, for the engineer or developer that really wants to manage the infrastructure themselves,
[02:10:09–02:10:38] manage the compute nodes, run a big training workload, interact with individual virtual machines or a large managed Kubernetes cluster of compute. But for folks that actually just want to interact with the model-- if you're actually hosting a model, again, the title of this slide deck is from electrons to tokens. So how do you get to tokens and where's the value uplift that you get from that? When you add in this managed services layer where you're actually serving the model, you're hosting a model
[02:10:38–02:11:08] and actually providing an endpoint for a customer to basically hit that API endpoint and actually serve those ChatGPT or those anthropic queries that everybody sending on their phones or laptops, you end up improving the margins quite a bit, adding call it another 15, anywhere from whatever, 5 to 15 million per megawatt. So you end up with, in a very optimistic case, call it $30 or 30 million per megawatt per year.
[02:11:08–02:11:34] So you end up with a two-year payback. That's a dramatically better outcome. Makes sense. Perfect. If you have no other slides, I might roll you through a couple of questions that the class had and then open it up for questions. So-- Yeah, Go ahead. Did you have more? It's fine. This is just some pictures of some deployments that we had that were just showing the walls and whatnot. And I'll talk about actually the inference scaling
[02:11:34–02:12:02] and where we actually to try to bring down those labor costs. Crusoe actually designed something we call Crusoe spark, which is our modular, self-contained, modular AI data center that we manufacture in these centralized locations where we can bring down the labor costs and we can actually bring down the infrastructure cost quite a bit. I call it, 30% to 50% savings depending on overall cost. So like that 19 million a megawatt.
[02:12:02–02:12:31] We can actually bring down pretty dramatically. What's the capacity in terms of size? Is this like gigawatt a couple megawatt or? So each unit for our air cooled architecture is 500 kilowatts. Got it. So about-- And this is the air cooled design. And I have a video here actually of them deployed in the field. I don't know if this is going to run, but no-- OK, well, it doesn't matter. And then we have a liquid cooled version, that's two megawatts. But you can deploy them in fleets, which actually opens up
[02:12:31–02:13:00] a lot of net new power opportunities, which is a pretty neat solution. Amazing. Well, in the interest of time, I can't think of a better person to ask the question about to ask you. You have seen probably every layer of the stack, from chips to power to gas to labor to networking, if you had to pick a layer of the stack and in particular a company and in particular a stock,
[02:13:00–02:13:30] what would you go long what would you go short and why? [LAUGHTER] I hope everybody's taking notes. What would I go long or short? I will call you in a year from now. Yeah. See how that did. Man, that's tough. Other than Crusoe, of course. Yeah, I mean, I'm turbo long Crusoe, but I do think that oftentimes getting these things right
[02:13:30–02:13:59] is difficult when you look at the time horizon. I do think that my bear case is actually there's that huge investment that I showed across the electrical stack, there's so many components that go into the electrical stack. Because what you're doing is you're taking power from this high voltage substation that maybe powers being at 345 kV. So there's this new line that's going in Texas, that's 765.
[02:13:59–02:14:25] So it's very, very high voltage power that then goes through this transformation process where you're stepping it down to medium voltage, you're stepping it down to low voltage, you're distributing it. There's tons of cable, tons of stuff. I think the data center is fundamentally going to drive a lot of innovations in the whole electrical stack and leverage a lot of solid state electronics and solid state transformers and power electronics.
[02:14:25–02:14:54] And I think it puts in jeopardy a lot of these companies that fundamentally have not innovated that much in the last 100 years. So this is like Eaton, Schneider, a number of other companies, which I think will-- and the reason I say it's very difficult to put a timeline on this is that those companies, I think, will do very well in the near term. They are on the critical path right now. They're going to do super well in the near term and their big partners of mine. So I hate saying that I'm negative on them.
[02:14:54–02:15:24] But over the long-term if they don't innovate-- I think that whole piece of the stack is going to dramatically come down in cost because of innovators that are building out this next version of the electrical stack. And there's going to be huge shifts to 900 volt DC and all these different aspects. So that's an opportunity for the electrical engineers in the room? The electrical engineers in the room absolutely. I think it's a huge opportunity. Power electronics, like, how do you get power from 765 kV to 900 volt DC in the rack.
[02:15:24–02:15:54] I think innovating on that problem is a super, super huge opportunity for people. OK. Any picks on the long side? But if not, I have another question for you. I mean I'm like bullish so many things. Maybe some other thing on the short side, I do think that open source is winning. Not open source is winning, but open source will do well and take more from the closed source model players. Yeah.
[02:15:54–02:16:21] Fascinating. Elon space data centers. Yes. Bullish, bearish, real fantasy happening in our lifetime or not. Data centers in space. [LAUGHTER] So I'm actually I'm very interested in this. And at Crusoe, we've established a partnership actually with another player in this ecosystem called Starcloud.
[02:16:21–02:16:51] That's actually launched the first H100s into space. There's a lot of things to about it. I just walked you through this whole stack of areas that I'm spending billions of dollars. All the concrete foundation, guess what? You don't need that in space. All the permitting, all the approvals you need on the power side, guess what? You don't need any of that in space. A lot of the core networking pieces, what I didn't get into is the millions and millions of strands of fiber that go into one of these data centers, and all of the technicians that you have to have
[02:16:51–02:17:20] to plug all this stuff in. Guess what? In space, you use optics for everything. So everything is basically optically interconnected. And that's all very interesting. It's also very hard. I think that thermal management piece is very challenging. And I also think the ongoing operations piece is very challenging. So in these data centers, when you're operating these big, large scale, interconnected compute clusters, things fail, GPUs fail, they have to be reseated in their compute tray,
[02:17:20–02:17:50] sometimes they have to be made and sent back to the vendor, Like, guess what? You're not sending an astronaut into space to take a chip and send it back to Jetson. That just isn't going to happen. So you're going to have a natural deprecation that will create challenging economics. And then I mean, a lot of it rides on does Starship fundamentally-- Cost of payload. Yeah. Does payload cost come down by two orders of magnitude. I don't know. I mean, he has a better idea than I do on that. My philosophy on this is probably
[02:17:50–02:18:18] not material in the next five years and probably not material for 10 years. But I think over a longer period of time, I think data centers in space are going to play a major role in the future of intelligent infrastructure. In a couple of weeks, the SpaceX S-1 is going to be available for everybody here to read, and so we'll see what his time estimate is. We know your time estimate next year. That's right. And final question, you're a Stanford alum, if you were here right now, what advice would
[02:18:18–02:18:46] you have for students who are making decisions about what to study, where to focus and no tougher time than now to make the decision? I have this philosophy that it's not like the exact things you learn in school don't matter that much. It's like I don't want to be disparaging to that or whatever, but they're important. But it's more like this process of learning.
[02:18:46–02:19:16] And like in my experience, it is like, one of our core philosophies-- in one of our core values at Crusoe you talked about one earlier, which is thinking like a mountaineer, one of our other core values at Crusoe is actually living on the infinite growth loop. This notion that nobody's a finished product, everybody's a work in progress. If you can get better, if you can learn more, if you have that tenacity to know how to improve yourself every single day, over time you get
[02:19:16–02:19:43] this exponential compounding, which is really the most valuable asset that any of us can have. So really, it's about investing in the process of hard work, of grit, of grinding, and then actually leveraging a lot of the tools because I don't know what the world's going to look like five years from now with the mass adoption and utilization of AI, where we all have the workforce of a million people at our fingertips. It's fundamentally going to change work.
[02:19:43–02:20:07] It's going to change the way everybody operates. So again, I would focus less on what and I would focus more on the how and leveraging of AI tools to run and live your life. I think, the advice I'd give to students. Awesome. Chase, thank you so much for doing this. Yeah, thank you. [APPLAUSE]
[02:20:21–02:20:49] I thought where we would start early is a lot of talk right before you joined about this world's moving fast. XAI, Cursor, OpenAI fighting Anthropic. You guys have done such a great job of stacking going from a data business to a lakehouse business to now an AI business. Just state of the union, view from the top, what are you seeing? Frame the landscape for us. What are the biggest things that you are thinking about? And, I've got a bunch of questions
[02:20:49–02:21:18] that we can talk about, but I thought we'd just open it up to what is the biggest thing on your mind as you think about AI? Yeah, I think you guys can chill out. Don't be stressed. I think times are crazy, and I think it's not warranted, basically. And I think the stress makes people do stupid things and chase just whatever happens to be the crazy thing that everybody's talking about on Twitter. I think it makes people have tunnel vision and not work on the right stuff. Yeah.
[02:21:18–02:21:46] And I think that's what I see with the current generation. Like every year we have interns coming to Databricks and the interns, I do always a session with them, an hour or 90 minutes or something they can ask questions. And last two years have been just insane. Before they would ask for good career advice, and you would give them good career advice. Now they're like 22-year-olds who are like, oh my God, should I start my own company and be a CEO? Or if I delay that by six months working on something have I ruined my career and life is over. And AGI is going to happen, and I'm going to miss the boat.
[02:21:46–02:22:13] And what am I going to do? So I'm just trying to tell people like, calm down, take a deep breath. Things take time. So that's what I would say. I would say, actually, I think also in Silicon Valley, if you think about it right now, what's happening is, and you might disagree with some of this, so feel free to push back. You guys might disagree too, you can push back as well. But there's this quest for superintelligence, which I think is unwarranted. Because first of all, they're not even defining what superintelligence is.
[02:22:13–02:22:43] But it's this godlike, I think people reading Kurzweil and it's take off singularity, this thing that comes and recursive self-improvement and cures all the diseases and GDP jumps by like 10% and unemployment goes to 20%. And there's no more jobs and-- Fear mongering. UBI to everyone and so on. No, I think they believe it. I think it's not needed. I think we already have AGI. So we already have artificial general intelligence.
[02:22:43–02:23:11] This is always equally fun. How many people think we have AGI already? OK, it's always the same. It's always like 10%. How many of you think that a lot of people that you interact with are not as smart as the smartest models that you use? OK. Now let's start all over. How many of you think we don't have AGI yet? [LAUGHTER] By the way, it always works. See, it's like the hypnosis is working.
[02:23:11–02:23:39] For some reason they've gotten the whole world to believe we don't have AGI. But it's like you just answered it, that you have it. But yet, no, you want to move the goalpost. By the way, I was at the research lab in 2009 at UC Berkeley called AMPLab. It was probably the biggest, most active important AI lab of its time in 2009. And the god of AI was working in that lab, which is Michael Jordan. His name is actually that.
[02:23:39–02:24:08] So he was like the Michael Jordan of AI. And back then, our definition of AGI, Artificial General Intelligence, we've hit that Like anything we imagine would be AGI, we already hit that. And those are all the leading AI researchers in the United States, many of them were working in that lab. But I wanted to see if I'm just full of it. So I went and asked some of those people that were there at the time, and I asked them, I said, hey, do you agree? And they all said, yeah, according to that definition, in 2009, for sure we've hit that.
[02:24:08–02:24:38] But, there's always some stupid but. We moved the goalpost or we want to change it, or we want to have some other definition, or AI. There's some example that it couldn't count the number of Rs in strawberry or something. So therefore, we don't have AGI. We already have AGI, OK. It's already smarter than many of the people that you interact with. That is general intelligence. It is artificial. It's not exactly a human. It's not the way human brain works. So we already have that. So in some sense, blowing a lot of money on GPUs and data
[02:24:38–02:25:04] centers and all of that kind of stuff is not really needed. Then there is at the same time, so you asked for the state of the Union. On the other hand, you have the MIT Tech report that says that 95% of the POCs are failing. It's kind of right, directionally. I don't know if the 95% might be wrong. Maybe it's just 75%, who knows. But if you go inside of an enterprise or inside of an organization. You go into any company and you look at how they're using stuff, the reality is that there's no lots
[02:25:04–02:25:34] of agentic coworkers running around doing all the work, blending with humans. That's not happening, OK. It's just humans shuffling TPS reports. It's like Office, the QA, the movie. Office Space, the movie, is still how the world runs. That's the reality. It's just the truth. Even inside the AI companies, that's how they run them. It's like they like to think, but they're hiring salespeople from old school companies, and they're running things in old school ways. And I don't see that futuristic thing.
[02:25:34–02:26:03] So then what's going on? We have AGI, but on the other hand, none of this is working and no company is using it. What the hell is going on? I think it's very simple. If you don't get all the context that exists inside of these organizations and how humans work and everything, all the contexts, we have in our heads, if you don't get that to the models and the agents, they're going to do lots of stupid mistakes and they're useless. And that's what's happening right now. The agents don't have the context that humans have inside of organizations,
[02:26:03–02:26:32] therefore they're useless. They do stupid mistakes because they don't know all the stuff that we know. Inside of every company, there's always like this one guy or this one girl, who's like, oh, go ask John or Jane. She knows everything. And everybody's like, tapping on that person. And that's the one person you can't lose in the company. If you lose that person, the whole company collapses. That one person exists in every department, in every company, and every organization. And that one person has all the context in their head. And that person, what they have in their head is not inside of the model.
[02:26:32–02:26:59] So therefore, the model can't operate. It just doesn't know a lot of the stuff. Usually John or Jane in that company have been there for 10 years, 15 years, 20 years, sometimes 30, 40 years. You need to get that transferred to the AI. If you don't, doesn't matter if you get super intelligence and you can solve really difficult math questions. And if you can get that context into the AIs, we already have AGI and they can already crack the problem.
[02:26:59–02:27:28] So my urge to you guys would be if you want to have impact in the world, figure out how to get that context into the AIs inside of it. Take an organization, how do you transform how old school business is happening, and how do you get those processes into the agents, then you will have massive impact. Because AGI is already here. That's my state of the Union. AGI is already here. You got to download the brain into the silicon. Get the carbon to talk to the silicon. Yes.
[02:27:28–02:27:57] Actually we were just talking about this-- And by the way, cue up like, pushbacks. I'm very curious to hear. I'm sure a majority disagrees. How many disagree with this? Oh, not that many. OK, I'm going to be more provocative. We need more pushback. So before we go into AI, there's a shadow of AI. Software is dead. Software has been dead for a while. We've had this four times. Every time it happens, it bounces back. Some macro reason.
[02:27:57–02:28:25] Brexit, taper tantrum, inflation. This time it's AI. And the question the class is asking is should we be loading up on software stocks? So is software dead? Is this a buy the dip situation? And I can't think of a better person to ask because depending on the day you ask, there's software. You're a AI company software. Speak about that, is software dead? I think you know better. You're an investor. I'm not an investor. Also, I don't give financial advice. But--
[02:28:25–02:28:55] [LAUGHTER] Having said that, if all software is dead, then isn't OpenAI and Anthropic dead? They're just software companies with a bunch of researchers writing software. So those companies would be dead too, right? So they shouldn't have trillion dollar valuations. SpaceX might make sense because they make rockets, but everybody else should be dead. NVIDIA should be dead. Because they just have really smart people who create chip designs, humans that use some software to create chip designs, and then they ship them over the internet, probably,
[02:28:55–02:29:23] over to TSMC, which is a real company creating actual chips. But then NVIDIA would be dead as well. So the world's most valuable company should be dead as well, because software is dead. So software obviously isn't dead, and it's not going to be dead. And NVIDIA and OpenAI and Anthropic are not going to be dead companies, because of whatever SaaS apocalypse or whatever you want to call it. But I do think two things are true. I think that two big changes have happened,
[02:29:23–02:29:52] which is one is barriers to entry have significantly gone down, and then switching costs have significantly gone down. So let's talk about those. Barriers to entry, because it's easier than ever to write software. So that's like a new weapon. Anyone can produce software very cheaply, almost at zero cost. It's not quite zero cost. And it will never be zero cost, but much, much cheaper than before. But that weapon is available to everyone. So also the people that create software now also have that weapon.
[02:29:52–02:30:21] It's not like only some new players have that. Everyone now. So including Databricks. We're a software company, but we also have that weapon, and it's an awesome weapon. I'm using it. Have you used that weapon and substituted any of your core software expenses like your CRM, your IT help desk, your office of the CFO software? No, I think that's stupid. Also, I think switching costs are lowered because it's easier to switch between UIs. Humans get locked into software. I don't know. How many use Android? No one.
[02:30:21–02:30:48] OK, wow. How many of you use iPhone? OK, wow. All right. You don't want to switch to Android? Why? You don't know the UI. You don't know how to use it. It's like a different UI. You would have to also transfer all your data, your phone, contacts, all that. There's too much inertia switching costs. That's a switching cost. But if in the future you're just talking to an agent that switching costs gets eliminated, because you're just talking to an agent. So who cares if the agent is instrumenting
[02:30:48–02:31:17] your Android, or your iPhone, or your Gmail or Outlook, or your Salesforce, or the competitor or whatever it is. So that's like the switching cost coming down as well. So yeah, I think it's going to be more competition. So I think software companies will have to run more efficiently. That I think is going to happen. But software is not the only moat. There's a good book you should read. It's called The Seven Powers. How many have read The Seven Powers? OK, a bunch of people here. OK. Yeah, so there are moats that are not just software.
[02:31:17–02:31:46] I mean, like economies of scale. If you can do things at scale better than anyone else so that you can afford crazy fixed costs because you're amortizing them away because of your scale, Amazon, AWS, then that's a moat. If you have a brand like Ferrari or Rolex, that's a moat. People writing cheap software can't just come replace that brand. People care about that brand. Trust.
[02:31:46–02:32:15] I'm the only one providing. You can trust my company. We don't get hacked. We have, like, really secure software. We have special certification. Maybe we have patents. That remains a moat that you cannot break that that easily. So all these remain. There's a bunch of other ones. Switching costs and so on and so on. So data is a big moat. If you have special data that no one else has, that only you have, that's a moat. Doesn't matter if they can write cheap software. So I think the answer is in between.
[02:32:15–02:32:43] The way I say it is, if a company has been around for 10 years and they have not innovated-- If their software looks the same as 10 years ago, but the revenue has been going up, they should be worried. Because they have not been innovating. And it's probably easier for a company that starts today to then, with barriers of entry being lower, write software quickly that's much better than that company. Because that company hasn't done anything for 10 years. They should be really afraid. And probably they don't have the innovation muscle anymore
[02:32:43–02:33:12] because they're not innovating. So obviously they don't have innovators. Those kind of companies are going to be wiped out. But there's going to be other companies that have been innovating the last 10 years and they're software companies. Or companies that now get their [MUTED] together because they're nervous and they'll be fine, too. Perfect. So what do you think? You're an investor. I mean, I think there's a grade, exactly as you said. If I was to give the grades to types of software companies, I would say if you have got a lot of data, like you said.
[02:33:12–02:33:40] If you've got some cyber like you're in some core loop, you're probably most robust and immune from it. Somewhere in the middle is all the workflow software, which has not innovated. The UX still looks the same. And you're like scrunching down on your shoulder and typing I met Ali Ghodsi today, these are the notes. That stuff's probably gone. If you were, exactly as you said, no innovation. You were a part of the old habit. But any one of those, they have customers, they have data. If they build great AI and start innovating,
[02:33:40–02:34:08] they can keep on going. They might have to change their pricing structure and their cost basis, but they'll be fine. In fact, they have a lot of advantages against incumbents. They have data, they have customers and they have some scale. So they have some economies of scale going, but they do have to get their [MUTED] together. And that's easier said than done. Yeah, yeah. I'll flash a chart. Have you guys seen this chart from Ethan Mollick? He talks about AI is very good at some things
[02:34:08–02:34:38] and he calls it the jagged frontier. This is customer support. Software engineering, this would be the frontier of maybe software engineering. Maybe this is customer support or whatever. But then there's a lot of stuff it's terrible at. Like this scale or this scale and so on. And Ali you see a lot of-- We're here, man. We're AGI already. We're over there. We're here. They kind of admitted to it. That's right, that's right, that's right. Reluctantly. Yeah.
[02:34:38–02:35:07] Yeah. So you've got what, like 6,000, 7,000 customers. Those are 10,000 customers. No, we have probably 20,000 customers. 20,000 customers, sorry. As you see this and that's a very good sample of the entire universe of what's happening. So in that sample of 20,000 customers, what are areas where AI is like hitting home runs and working as advertised? And what are areas where the frontier is still rough and it's not working. The POCs are failing. Yeah.
[02:35:07–02:35:36] Look, it's not AI's fault. I mean, most companies are somewhere here. I think we have AGI, but I think most companies, if you look at how much maybe they're having, AI's helping me in some tasks. That's what most companies are doing. That's just how it is. And it's because that context isn't there in the model. So the model can't do it. Take support, which everybody said, OK, that's going to be dead. Support is like gone. Right? Support is very hard. Support are literally the things that humans don't know what to do. They get stuck. So take Databricks.
[02:35:36–02:36:04] Databricks offers support. Databricks is a company, it's an advanced platform where you can do data science, machine learning. You can do advanced things on the platform. These are smart people who make big salaries, they have education, they have data science education. They're trying to use Databricks, and maybe they get stuck. So their machine learning models, it's not getting the right F1 score or something like that and they're stuck. And they tried everything they call our support.
[02:36:04–02:36:34] So it's pretty hard to automate that. You can't actually give it-- none of the current support automation. We've tried them all, companies. All of them. Immediately, even actually when they start talking to us, as soon as they know who we are, they're like, whoa, we can't help you. Get out of here. So yeah. Most of the world is over here. But it's because we don't have the context. If AI could have all the context of how our support engineers at Databricks operate, then the AI could do it. It just doesn't have it. Yeah.
[02:36:34–02:37:04] One of the things we used to say at Palantir is your AI strategy starts at your data strategy. You got to get the roads paved and have the data flowing. If you were to bucket the best enterprises who are maybe starting to head towards the right in your customer base of 20,000. What is common between the ones who are making it work and the Ferraris are flying, and ones where I'm guessing it's a context problem for the ones that it's not working? And what does it take to get the context working? Yeah. It's very hard.
[02:37:04–02:37:33] It's a human problem. It's not an AI problem. We already have AGI. It's a human problem. I don't see anyone really doing an excellent job at this. You have to rewire all your processes in the organization to be able to do it. This is like well known. I mean, my favorite is there's an article actually, that I recommend people reading from 1990 produced by Stanford professor or researcher. It's called, I think, From the Dynamo to the Computer. OK. Check it out. So Dynamo to computer.
[02:37:33–02:38:02] And it looks at different technological revolutions and how long it took for them to have impact on productivity of the economy. And it's just takes just forever. Like when the PCs came out, the joke was, the Nobel laureate economist Richard Saul said that computers or PCs, you can find them everywhere except in the productivity statistics.
[02:38:02–02:38:30] It just doesn't show up in the statistics. Why? People were buying PCs and they were using them as typewriters. So they would have people type on PCs, but then print out the sheets and then put them in folders, and then have assistants that index them and do things. So like you didn't see any productivity gains from it. And same thing with if you look at the Industrial Revolution, same thing happened. We had these steam engines.
[02:38:30–02:38:57] And the steam factories were super dense and they were running like with these, they were called the line shafts, which were these things that rotate. When the electric engine came, that's the Dynamo, it took 40 years before they saw any productivity gains in the economy. Wow. Yeah. Check it out. This is in that article. It took from 1880-- The diffusion took 40 years. From 1880 to 1920, when the electric engine came, to see impact.
[02:38:57–02:39:26] So what they were doing is, they were going to these factories that already were these line shaft factories. That were these dense factories, we have a steam engine that's rotating this line shaft and it's rotating these belts, and then everything is working. You have these multiple stories. And all they did is, just like the PC, they used a typewriter, they would replace the steam engine with an electric engine. And just replacing the PC, you don't get any productivity gains. It took till 1920, but maybe it was 1915,
[02:39:26–02:39:56] but I'm roughly right, until they realized, wait, we have to change the whole factory floor. We have to move the factories out of the cities. We have to have floor plans that are much bigger because now we can distribute the electricity. It's not like the torque that has inefficiency. We can spread it out. We can have floor plans that are big. And we can run different parts of the factory at different rates. Unit drive versus group drive. Took a very long time.
[02:39:56–02:40:26] That's what's going to happen. Same thing now. Rewiring, I know it because I have 20,000 customers and I talk to them. I was late to this meeting because I was meeting one of the CEOs of one of the big banks. And same problem. He has the same problem. All the organizations I work with have the same problem. They're like, I'm not seeing any advantage. I don't see. They're all like, AI is amazing. It's coming. It's like, I need it. I need to do that. But they're like, I don't see any productivity gains in my organization. What the hell am I doing wrong? And I tell them we have AGI, and they're like, what? That is not true.
[02:40:26–02:40:55] We don't see anything. It's a very tough problem. Because you're like, hey, I got the brain, but I got to rebuild the human body. The hands, the legs. Yeah. Let me give you an example from Databricks. So Databricks helps you get data from all the different systems like Salesforce, Workday, and so on, collect them in one place, secure it, and then do AI on it. Like you can do predictions, you can build predictive models. That's what Databricks is. So we built connectors to all these systems. These connectors are, it would take us 3/4 to build a production connector. We're good at.
[02:40:55–02:41:25] This is what we do for a living. We build these connectors. We can build a connector from Databricks to Salesforce. Production ready. It would take us 3/4, so nine months to do that. Shipped, secure, with so on. That's what we did. So as the LLMs got faster and faster and faster, I started experimenting with this myself and I was like, oh, I could write a connector in two days. So I went to the team that builds this and I was like, hey, I can do this in two days. How come it takes you guys three quarters? They're like, OK, great point.
[02:41:25–02:41:52] Let us come back to you. So they went and they thought about it and they came back in two weeks and they said, OK, you're right. But you're also not right. We looked at it and yeah, this AI is useful. We can compress it down from 3/4 by 1 and 1/2 month. So we can get it from nine months to 7 and 1/2 months. That's it. I'm like, well, I can do it in two days. I'm like, no, no, no. No offense to you, but this is production code and it really actually works.
[02:41:52–02:42:22] And we have customer feedback and it's like secure. And you wrote some toy and God knows what. I mean, no offense, you're great, but. What's the missing link? So I was like, oh man, this is kind of depressing. But yeah, I'll take the 1 and 1/2 month improvement. And maybe it's something, but maybe I'm just stupid and I don't get it. Then I found another guy in the company. We went to him and we said, hey, can you look at this problem? And he's very first principle. He's a very smart guy. And he doesn't care about all this, fluff.
[02:42:22–02:42:51] He's like, he cuts through the fluff. And he cut through the fluff. And he worked with the team and he came back and they said, hey, after looking at the problem, we can do seven connectors in one quarter. Boom. Yeah. Let's go. What is the difference? So what's the difference? OK, so what he did is he went from first principles with some team members. And they looked at it and they said OK, first quarter, they're just sending our very expensive, very smart, Stanford educated product managers out to the customers to talk to the customers and collect feedback.
[02:42:51–02:43:20] What exactly is your requirement? How do you use Salesforce and so on. That takes a full quarter. At the end of that quarter our amazing smart product managers come back with a 60, 70, 80-page super nice report on exactly all the requirements. OK, so you're blocked for a whole quarter. So for sure you can't, Amdahl's law, you can't compress it below that. Then code writing starts. We have to test this stuff. So testing requires you to set up Salesforce, Workday, NetSuite. But those are not software by Databricks. So we're not very good at that. That takes a very long time.
[02:43:20–02:43:49] And it's hard to find people to do that Databricks. So that, again, is like a process that takes a long time for us to stand up. And it's very error prone. So we couldn't do that either. And then we have one person for each connector. They go on vacation, they get sick, and so on. So all of that. So what he did is he just from first principles, looked at it and said, we're going to just rewire all of this. And a lot of people didn't like this. They were unhappy about it. But he said that, the product requirements, instead of one quarter, we're just going to take one week and quickly
[02:43:49–02:44:14] write down whatever we have. We might get things wrong, but because the software is so fast to write, we can rewrite it again. So let's iterate faster. Standing up, the Salesforce instances, let's outsource that to firms that can do that for us. And we can just pay them a lot, and they do it in parallel. So we can shrink that as well. And then one person per connector, let's change it. Let's have seven people, seven connectors, and then they all work on all the connectors together so we don't have what's called Bus Factor 1.
[02:44:14–02:44:43] If someone is hit by a bus, the whole project is not stopped. So yeah. So got it all done into one quarter and seven connectors shipped. And so but this had nothing to do with really it didn't have anything to do with AI or AGI or smarter models or superintelligence or gigantic. The next GPT 7 or Opus 6 would not have helped us do this better.
[02:44:43–02:45:11] We needed to make those changes. And that's like a human refactoring problem and process change. And so this is what the whole world is going through. That's what you need to do well if you want to succeed. Some are doing it better, others are not. Hamilton Helmer actually talks about this quite a bit actually. So for all of you who are picking assignment, option one and want to be investors, Hamilton Helmer's book is a must read on process power. We were debating this before this.
[02:45:11–02:45:41] If, Ali, you had $100 to invest across what Jensen calls the five-layer stack energy, chips, infra model and apps, where does value accrue? If you were to put $100 in the index of energy and chips and infra and so on with let's say, a long-term time frame, where would you put it? How would you allocate the $100 and why? I'm a computer scientist. I'm not an investor. I don't give financial advice.
[02:45:41–02:46:11] But-- [LAUGHTER] If I have $100. You are allocating Databricks time, Databricks is across three of these. Yeah. I would just say, look, it's obvious that the applications are going to be the winners, right. So I would put it in the top. It's kind of like, and I'll give you some guesses, but who knows, actually. It's very hard to predict. So you would have to have, I would go early stage, and I would have a seed strategy and I would invest in many, many startups. And I would get most of them wrong,
[02:46:11–02:46:40] but a few would actually make it and there would be the next Google or whatever. But when I did my PhD in the early 2000, I was in the networking field. Networking was like the cool thing to do. It was the advanced thing because the internet was, you want to work on it. The internet was the big thing at the time, and the coolest thing on the internet was networking. And the hardest problem, like the smartest math brains
[02:46:40–02:47:09] were working on at the time-- we all knew what the future would look like. The future, everybody knew what the most important problem everyone's going to work on is what's called the multicast problem. What's that? Which is, yeah, see, it's problematic that no one knows what that is today. We were clearly wrong. So multicast is, you want to broadcast from one source, let's say a soccer game or a football game or a basketball game to the whole world, because everybody wants to watch it at the same time. We didn't know how to solve that efficiently. So all of the smartest brains in the world
[02:47:09–02:47:36] were trying to work on this problem. And bandwidth was scarce. While we were doing this. And by the way, we actually had pretty good problems and I started the company on this. And we had a great solution. Unfortunately, the cost of bandwidth just plummeted, and they just deployed so much fiber that this problem was not a problem ever. So no one needed to buy this software. So it was a complete waste of time. And at that time, we thought the hardest problem was the most interesting things to work on
[02:47:36–02:48:02] are Cisco routers, routing, BGP Border Gateway Protocol, internet protocol, queuing theory, quality of service, these kind of things. Those are the most interesting things. Because we had tunnel vision on the internet. And what is the internet? Well at the time it was the internet protocols and those things. No apps really existed. So we were all focused on that. And today everybody's focused, I would say, on I think chips. And I think infrastructure, I think
[02:48:02–02:48:31] people are really right now the hot new thing is like NVIDIA, OpenAI, Anthropic, DeepMind. These are the things everybody's focused on. AGI, superintelligence, what I said at the beginning. But on the internet there were really weird things that took off. The really weird things that took off were like taxi business, which is Uber. Yeah. Or selling books, which is the lamest thing ever. But that became Amazon, which became AWS.
[02:48:31–02:48:59] Yeah, or renting your bedroom to people, that's Airbnb. Or sending people short text which became Twitter. And if you said them in those words in 2000 to people, people will say, you're out of your mind, like you're insane. You're full of it. But those were the great ideas of the time. Those are the ones that we can. So I think it's the same thing here. To throw a few of them out there.
[02:48:59–02:49:28] I think healthcare is like 17% of US GDP. We all still unfortunately will die. And we all care about our health, and the health of our loved ones. I think there's huge, we have the propensity to pay for this. Like we'd pay anything to be able to save lives or our loved ones or our own lives or our own health issues. And it's not particularly well done today. Surprise, surprise.
[02:49:28–02:49:56] Healthcare is not awesome. So imagine a company that has seen a billion patients. I have seen hundred million patients with your genetic composition and the kind of issues that you might have in the future, and I can help you. But what are you willing to pay for me to help you with that? That could be a company that's trillions of dollars worth. To take something out of left field that I think people think is really not interesting enough. But take education.
[02:49:56–02:50:24] Education actually, in VC space, the consensus has always been education is a terrible investment. Like VC, people say, oh, I could never invest in education. What's the last public market company you know in education? Yeah. What's the last trillion dollar education company? Not even 100 billion, yeah. Yeah, yeah. Anything, right. But most people have kids. More kids are produced and they do need to go through to get an education whether people, believe it or not.
[02:50:24–02:50:52] And people do care actually if the education for their kids are good or not. Elections are won and lost. There's cultural issues on these things of what you're allowed to teach my kids or not. Elections are won and lost on that. Not because it's a stupid topic, because it matters. Like, what are you teaching my kids matters. Are my kids being brainwashed to do the right thing or the wrong thing? Or are they well equipped to get the jobs of the future? I think if there is a company that
[02:50:52–02:51:21] can provide amazing education, using AI, I think a lot of people will pay for that. And if it's like proven that that does a better job than whatever they're getting right now, just two flavors of like obvious companies that I think will exist, and they could be trillion dollar companies if they do it well. They will have data moat. They will have economies of scale moat. There's winner takes it all dynamics in those markets
[02:51:21–02:51:45] at least in countries NGOs. So I think the value accrues to the top. We can't wait for that to happen. But I'm not an investor. Yeah, we can't wait for that to happen. Would you push back? No, I mean, look, I've written extensively about this. Eagerly waiting for this, what I call the blue triangle to invert. I don't know if you've seen this. But basically this is, all of the money in AI is with one guy.
[02:51:45–02:52:14] That's why Jensen's so happy all the time, as you know. These guys are fighting for dollars. There's no money there. There's very little money here. I mean, people are making some money here. And so we'll see. But that's the bet. The bet is that this thing will look like a more sustainable. Yeah, it will go that way. I mean, all value in Silicon Valley, and in tech, and technology moves up the stack all the time. Yeah. Like even look at the greatest companies. Like, OK, the company that created the PCs, IBM
[02:52:14–02:52:44] was like the greatest market cap and all the value accrued there. But then that became commoditized. Then it became the software on top of it, which is like the operating systems and the Microsofts of the world and so on. Then here at Stanford, actually a while back, it was like 20 years ago, VMware, which is how do you virtualize that software. And that became commoditized. And then so it keeps moving up the stack all the time. That's how it's going to be here too. 100%. And one of the forces that is commoditizing this, you've spoken about, this is Open Source. Open Source is getting pretty good.
[02:52:44–02:53:13] This blue line is Open Source. The gap is closing. This is like what, three four months. This gap is now like a month. But still people are spending so much money on these frontier models. People cannot wait to get their hands on 4.7 from Claude or 5.5 from GPT. But then there's this whole economy of very good open source models. What do you make of all this? On one side, you've got people earning what, $30 billion
[02:53:13–02:53:42] now or maybe 40 billion at Anthropic. But on the other side, this open source stuff is like nearly free. Obviously you've got to pay the hosting. How do you think this shakes out? Will the proprietary model layer accrue any value? I think it's going to be valuable. And I think people will want it. Whether it's open source or not, let's put that aside for a second. I think there will be token factories which serve this stuff up. It's just like the cloud. I think it would be foolish to say you all will have your own little mini data center in your living
[02:53:42–02:54:11] rooms, and you're going to run your own PCs, and you're going to insert GPU cards that you buy at home, and you're going to run this [MUTED] yourself. Or on your phone or MacBook or the Edge. Some of that might or will exist. It will come to the edges. But I do think there'll be like big centralized data centers where this happens. But we haven't discussed already running open source models or are they running proprietary models. And here's a fun fact. So Moonshot, the Chinese company released Kimi 2.6. Very good model. Two days ago? Three days ago? Yeah, Tuesday.
[02:54:11–02:54:40] Yeah, Tuesday. So two days ago, they released 2.6. In January, they released 2.5. And here's a fun fact, 2.6 that they released on Tuesday is the best model ever in the history of mankind ever produced, frontier, non-frontier, if it just had been released in January. But open source will be here. And it will apply pricing pressure. And this business of frontier models, that core business of providing frontier models is going
[02:54:40–02:55:09] to be economies of scale game. And you will have to do it at small margins. It's like an amazon.com bookselling business. That's what it's going to look like in the future. Therefore, they're not going to be that many people doing it, just like on amazon.com, and gross margins are going to be tiny. And operating margins are going to be small. That's my take. Yeah, yeah. I think so too. Three rapid fire questions before we wrap. Your favorite AI product that you use every day.
[02:55:09–02:55:38] I don't know. That's a tough one. I mean, I use all of these. I actually like Cursor. I notice everybody loves [? spot ?] code. I like the diffs and how it works. So like on coding I use a combo of those. I still kind of like it. Are you still using it after Elon owns it? No, I stopped. No, of course. [LAUGHTER] Because you're going to lose access to Anthropic and OpenAI tokens through Cursor, I presume. Yeah. Yeah, man. Awesome. It's great. Good, good.
[02:55:38–02:56:07] Supporter. You've been-- The truth is, I do use Databricks as junior products. This is the truth. Because inside Databricks, most of my decisions are like numerical and quantitative in nature. Like, should we do this? What's the ROI on this? What's the cost on this? What is it going to cost us? So I need something that can understand numerical data and time series data. So Genie is really good for that. So that's what I honestly go to quite a bit. Quite often. Right. Future for Databricks.
[02:56:07–02:56:36] You've been at this for 15 years or so. What is your vision for the next decade for Databricks? Well, I think the cost of software is going down. And so barriers to entry and switching costs are going down. So there is a SaaS apocalypse of sorts. But not all software is going to be dead. We would love to partake in that and kill some software. Yeah. Right, right. Any advice for students in the room who are about to make career decisions?
[02:56:36–02:57:03] Yeah, I think don't be worried about the fear mongering. Don't be stressed out. Take it easy. I was very stressed doing my PhD in the early 2000. I thought the world is ending with the internet and everything. And working on this most important problem that we all knew was the most important problem, which was the multicast problem. [LAUGHTER] Which none of you have heard of. Turned out not to be a problem. But I think one interesting thing
[02:57:03–02:57:33] is that in 2000 we had the internet, in 2009, Airbnb was started. OK. But there's no reason why Airbnb should start in 2009. I've made this argument to you. Airbnb could have started in 2001. There's nothing like we needed something additional to happen in the world. Airbnb could have happened and disrupted hotel businesses in 2001. Yet it took nine years for someone to have that idea. And that was Brian. And by the way, Brian is not like he sat there and he was taking a Stanford class thinking about a case
[02:57:33–02:58:02] study project. Brian needed like bed and breakfast. Right. Some designer conference or something. Yeah, he was at a conference and he's like, why is this so hard? Like, I'll just solve this myself. So it took nine years to come up with that good idea. So I think good ideas are very hard to come by, actually. I think humans are very bad at coming up with great ideas. And we have this tunnel vision and we focus on the wrong problems like we did with multicast in my earlier, my PhD was really stupid.
[02:58:02–02:58:31] So chill out and take a long-term perspective. And work on the things that you think will have long-term good impact. I think Jeff Bezos did it pretty well when he was an investment banker in Wall Street. And he said, hey, zooming out, what's the big thing that's happening? It's the internet. And then he said, hey, let's just make a secular bet on internet, there's going to be more and more internet. So it's going to slowly, over time, disrupt things.
[02:58:31–02:58:59] So then he said, OK, in the long run probably purchasing can move more to the net. Maybe not right now. And then he started with he was very modest and he started with the dumbest thing you could possibly. Like the unsexiest thing, which was a complete commodity that looks identical and there's no differentiation, which is books. And it just started with that. And he just bet on that secular trend. And every year it was more and more right. And now it's like everything on the planet, it's the everything store. So think long-term like that and don't
[02:58:59–02:59:16] be swayed by the coolest thing that everybody is like right now, making lots of noise on Twitter on because chances are it's probably something like multicast. Yeah, awesome. Well, thank you so much for staying longer, folks. Thank you, Ali.
[02:59:30–02:59:59] Hello, everybody. It's good to see you again. Another round of the economics of the AI supercycle. This time with Professor Katti. Welcome back. Welcome back to Stanford. Thank you. As I thought about introducing Professor Katti, I could not think of a better person who has seen the entire soup to nuts of electrons, the entire substrate, all the way to agents.
[02:59:59–03:00:28] You obviously started networking startup. You were the CTO and head of AI at Intel. You now run industrial compute at OpenAI. Thank you. Thank you for joining us. Thank you. It's coming back home for me. Welcome, welcome. I thought we'd started a fun segment. Intel spent about a decade trying to convince everybody that they're an AI company. It finally happened. They finally got there in the last two weeks. What happened?
[03:00:28–03:00:56] I told you so. [LAUGHTER] That was my job. As I was mentioning, I was Intel CTO and also running its AI business until I left for OpenAI in November. So yeah, a little bit of a lag. But the bane of people who have the forecast. But no, I think Intel story is turning. I'd say there are two factors that
[03:00:56–03:01:22] have big tailwinds for Intel. One is the world is heavily supply constrained. And so any company that has serious manufacturing jobs in this space and can build, not just design, is going to have tailwinds. And Intel obviously is pretty much the only leading edge American company left that still can manufacture.
[03:01:22–03:01:51] The other, of course, is CPUs are making a comeback with how we are beginning to use AI with agents, and we can get into that a little bit later. So both of those are very good things obviously for Intel. A lot of execution still to be done. I mean, the market always is ahead of the story, but fingers crossed, Lip-Bu is a great CEO. I loved working with him. So I think good things to look forward to. Amazing.
[03:01:51–03:02:18] I'm sure your departure had nothing to do with the stock chart. They're not correlated. I still kept my stock, so don't worry. Well done. We'll get more into the role of CPUs, all the different parts of the compute supply chain. The second thing I thought we'd spend some time on is this chart that OpenAI put out at the start of the year for everybody. Sarah Friar, the OpenAI CFO, wrote this article about OpenAI's compute ambitions. And on the left, you'll see, is OpenAI's compute capacity
[03:02:18–03:02:47] over the last three years. This is OpenAI's target by the end of the decade, and it magically seems such in that the compute capacity seems to be hyper correlated with our revenue. No prizes for guessing what this might be if this gets there. Yeah. Talk about this chart for a second. What's going on here and how should we process this? Yeah. So at OpenAI I lead industrial compute, just for some context
[03:02:47–03:03:15] before I answer the question. So my job and my team's job is delivering the compute that OpenAI needs across everything, across training and inference. So this chart is what I live and breathe every day. And-- You're making the numbers go up. Yes, my job is to make it go up and into the right. So that's the job description. But kidding aside, I think it has-- as you pointed out,
[03:03:15–03:03:40] revenue is basically a lagging indicator for frontier lab companies. And what I mean by that is it basically is very simple calculation of how much compute we have and how well utilized is the compute. And so the last three years have borne that out. Every year, we have tripled compute. Year over year and revenue has tripled.
[03:03:40–03:04:08] We don't see any end in sight to the correlation yet. I think just 5.5 coming out and the uptake. I mean, Codex has probably seen meaningful double digit growth just in two weeks since 5.5 came out. I think people are using it for not just coding anymore. Codex is being used for general purpose knowledge work, so token usage and just more and more complex tasks
[03:04:08–03:04:33] are being consumed. So we essentially are tracking how much compute we have available. And the number of users, the number of tokens, and therefore the revenue basically is tracking, tracking that. I'd say that as we think about the future, OpenAI is still a research lab. And the reason I say that is it's
[03:04:33–03:05:02] very much not just a, here's how much revenue we can maximize. It's much more rather, how do we make the maximum amount of compute we can make possible for research so that researchers are unconstrained in exploring new ideas, new models, and new ways of pushing the frontier on intelligence. And so the 30 gigawatt number here, that which is an aspirational goal is a split. It's split across research and products.
[03:05:02–03:05:32] But we definitely don't see a world where we don't utilize it given the current trends that we are seeing. And maybe just a quick follow up before we move on, what's the rough split between training and inference, and how is that trended over time, and how do you expect it to go over time? I think the-- scaling laws. So obviously scaling laws, initially everyone assumed applied for pre-training only. What has shifted is scaling laws have
[03:05:32–03:05:59] evolved to cover the entire life cycle of compute. And what I mean by that is pre-training, post-training with RL, which is primarily an inference workload. Synthetic data, because we have run out of real world data to train models on, so we are generating data to train models on, that is primarily an inference workload. And then of course, the actual products themselves, everyone using ChatGPT and Codex. And that is an inference workload.
[03:05:59–03:06:25] So more and more it is shifting to inference. And inference is already the majority, just to be clear. But even inference should not be taken to mean just products. A big chunk of research, a big chunk of training the next level of intelligence is also inference. Our prediction is that a super majority plus like 80% plus
[03:06:25–03:06:55] will essentially be inference compute in the future. Just building on your response, Sachin, is it also true that if the relative ratio of how much gets used for inference goes up over time, the dollar density, meaning dollars per gigawatt, might also go up because inference is basically what you can monetize? Well, we are hoping it goes down, dollars per gigawatt, because this stuff is expensive.
[03:06:55–03:07:21] Every gigawatt is roughly-- I mean monetization. Sorry, monetization. Yes. I think yes for sure. As more tokens get consumed, that should lead to a corresponding increase in revenue. At the same time, I think our mission is to make tokens cheaper. And it's two different dimensions. One is, make every token cheaper, make every token more intelligent,
[03:07:21–03:07:47] and make every task require less number of tokens to perform. So we push on three dimensions. Keep improving hardware and software to generate tokens more cheaply. We keep pushing on the capabilities of models to make sure every token is more intelligent. And we keep pushing the hardness, like Codex, to make it such that we need less number of tokens to perform any given task.
[03:07:47–03:08:16] And that is a very fundamental principle in how the company operates. And the reason, of course, is how do we make sure that all of this intelligence is as widely accessible as possible. Now, your job, as you said, is to get the numbers to go up top and to the right. The forecast-- honestly, I don't envy anybody who's forecasting. Tripling year over year at that scale seems like a hard job
[03:08:16–03:08:43] to not only forecast. Forecasting is an easier job than actually making it happen. You think so? Tell us about your job a little bit. What is the hardest part of it? Is it sourcing the compute? Is it securing it? And how are you securing compute right now, it seems like a fistfight? And where's the bottleneck? Is it power? Is it energy? Is it ships? Is it land? Is it--
[03:08:43–03:09:13] All of the above. I think if you think about the life cycle of compute, so one is obviously sourcing compute. And compute is a very broad term. When you think about compute for AI, you really have to think chips, memory, networking, power, cooling, data center buildings, power generation,
[03:09:13–03:09:42] power distribution and of course, land. All of that is equal to compute. All of that needs to come together to build compute at a gigawatt scale. And so when we think about sourcing, we are not sourcing compute. We are literally sourcing that entire supply chain and making sure at this scale that we have visibility into where that each component of that supply chain will come from. So that is one big piece.
[03:09:42–03:10:11] The second piece is how do we orchestrate that supply chain to all land and align at the same time to make this compute operational. So a gigawatt is roughly half a million GPUs. And when we're talking about whatever number it is, 6 or 10gw, you're talking about quite a large number of chips being networked together, being powered, being cooled,
[03:10:11–03:10:38] being kept up and alive, made sure that everything else that needs to come together is there. So a big chunk of the work really starts after you sign the contracts. Like how do I make sure that your suppliers are actually going to deliver what they said they will? How do we make sure that we engineer these systems so that it all works together at this scale? And how do we make sure that it is operationally
[03:10:38–03:11:05] usable, it is up and running and runs at the highest performance we can run these chips at? And these chips are very brittle today, very sensitive to cooling and power fluctuations. And they can quickly throttle back and compute how many flops you have. So that's really the job. The fun part is the contract signing. The hard part is everything after. Yeah, getting the autographs. Yes.
[03:11:08–03:11:33] It's a very consequential time right now, and I imagine a lot of the decisions you're making will impact us and the rest of computer users, which is billions of people for years, if not decades to come. What are some of the biggest trade offs you're making? What are some of the biggest decisions you're making that will make case studies at some point down the future that you can talk about?
[03:11:33–03:11:59] I think there's a lot of societal level implications of these decisions. To pick an example, if you put a gigawatt data center in a place like Georgia or Michigan, for example, it's a pretty big consumer of the grid and that amount of power. And when you run a big training job, these things are synchronized jobs.
[03:11:59–03:12:24] They go up and down in sync in intensity. So you can see the energy fluctuations on the grid that can be hundreds of megawatts very quickly, and our infrastructure was never designed for it. A grid could basically fall apart and an entire state could have a blackout, depending on how these data centers behave. So a lot of time we spend thinking
[03:12:24–03:12:51] about how do we make sure we can design these systems to not have all this collateral damage on the rest of the country's infrastructure. So that's an example of the kinds of things that are being redesigned. We obviously are spending a lot of time thinking about how to de-risk supply chains. So how do we move fabs?
[03:12:51–03:13:21] How do we move memory factories to other parts of the world? How do we decouple from grid energy and use natural gas and increasingly nuclear in the future? So I think this is going to lead to infrastructure investments and innovations that the rest of society will benefit beyond AI, because these are things that otherwise did not have an impetus to happen. Then I'd say, obviously, all the implications of AI itself
[03:13:21–03:13:50] and compute at this scale. I mean, 30 gigawatts is a lot. But I'd say our vision is-- and Sam has been talking about this for a while-- like we've all taken it for granted that every one of us should have a mobile phone, and we upgrade one every year or every two years. It's not that crazy to think every one of us should have a GPU. And a GPU is what? A kilowatt to 2 kilowatts now.
[03:13:50–03:14:19] 7 billion humans out there. That's 7 terawatts of compute. And so that is two orders of magnitude more than what we are talking about here. So if you really believe in that world, then we still have a long ways to go. And maybe just put this in perspective. Sachin, how much energy does America consume compared to 30 gigawatts? I don't have the number of the top of my head,
[03:14:19–03:14:49] but I think the US, if you add up all the hyperscalers, is planning to build 100 gigawatts of compute. beyond OpenAI. So 30 gigawatts of us, whatever else everyone else builds. You've seen Google's numbers, Amazon's numbers. 100 gigawatt is probably already a fifth to a higher of the grid. So this will be consuming double digit percentage of US capacity. Wow. Making the market.
[03:14:49–03:15:16] It will change the market. I think the way we think about energy as just purely for human consumption is no longer true. One of the rumors that's been going around is that OpenAI has a significant compute advantage compared to the other labs. The class here loves both OpenAI and Anthropic equally. Are we polling?
[03:15:16–03:15:43] We did that and we'll save you the answer. I'll fill you in after. But talk about the compute advantage to an extent that you can share with us, assuming forecasting was perfect, what does that afford us to do and deliver to consumers of OpenAI? Are you seeing it? So 5.5 is a big model.
[03:15:43–03:16:11] It's expensive to serve, but there are no limits. So everyone's able to go and use it without token limits. We are much more generous on how many tokens you get for your subscription. We often almost every day or every week, reset the limits so that people can play with it a lot more. And that's the compute advantage showing up in day to day usage.
[03:16:11–03:16:40] And so that comes back to that earlier point, which is making sure that we have enough compute to distribute this intelligence at scale. And no just build the intelligence-- it's no good if you build the intelligence, but you can't really deliver it at scale. So really, we spend a lot of time in making sure that it's not just about training, it's actually usable compute that we can deliver to everyone at scale without putting artificial limits. 100%.
[03:16:40–03:17:10] One of the impacts that the class has already felt is we asked two labs for a Codex and unnamed product subscription, the Codex team gave us that pretty quickly. So I now understand why that was the case. We'll switch it up a little bit, Sachin. Codex and Codex like instruments have a lot of different things that need to come together, the GPU and all sorts of ASICs, the CPU, the memory, the networking,
[03:17:10–03:17:40] all the things that you outlined us. Maybe start with the workload in question. What is the modern genetic workload look like? How has that evolved over time? I think the way maybe just to frame the answer. So ChatGPT was obviously a big inflection moment. But if you think about ChatGPT when it started, it really is one shot inference. You ask a question, it gives you an instant answer,
[03:17:40–03:18:08] and you're done, and you go to the next thing. I think the big innovation and the breakthrough in 2024 was reasoning. And so not just for inference but also for training, so being able to for the model to introspect and think and therefore generate better answers. And that again increased intensity of inference. So there's more and more inference happening But there are still passive things.
[03:18:08–03:18:36] You're asking a question. They're giving you an answer. They don't take any action for you. So the word agent kind of encodes what we mean. It has agency. It has agency to do things. And so what I mean by that is when we think about agents, it's really about closing the loop, not just thinking and suggesting, but also trying it and looking at the output, iterating and then trying
[03:18:36–03:19:04] a refined answer to do a task, so whether it's coding or any other form of knowledge work. So it's really closing the loop. And it's actually delivering the full value of what we expect AI to deliver to you. Not just be an assistant, but actually be an agent that can close the loop and do work for you. And so implicit in that statement is obviously inference and thinking. But as I said, try.
[03:19:04–03:19:34] So it's going to go look for a relevant data. It's going to go search. It's going to go spin up a VM to run a test if it has generated some code. It's going to spin up Excel or PowerPoint to try out some slides and see how it looks. And it's going to look at the output and reason about it and iterate on this. And so to the graph, the compute graph is a lot more complex now. If I putting my computer science hat back on,
[03:19:34–03:20:01] if I thought about the chatbot world, it's a very simple compute graph. There's a user. There's one node, which is the inference call. And there's an answer. Reasoning was multiple nodes of inference calls. And now we have a much more directed acyclic graph, if you will, to use the more precise technical term. You have an inference call. You might have a tool call. You might have a database or a search query. You might have a RL VM environment spun up,
[03:20:01–03:20:29] then back to an inference call, and so on and so on. So the compute graph is now a lot more complex that you're executing. And so that naturally leads to a much more sophisticated compute infrastructure that's needed to execute that compute graph, a lot more intelligence needed in how you distribute that compute graph and where you run what part of the graph on. And so both the compute, but more importantly,
[03:20:29–03:20:57] the workload evolving in this direction is going to change the shape of how we think about compute infrastructure. Fascinating. You outlined a bunch of different steps along the way. I can imagine some parts of that being more relevant for different machine like a GPU, other parts for CPUs, and ASICs. Is there emerging maybe clusters of workloads that are particularly suited for a certain workload,
[03:20:57–03:21:27] you might say, hey, the NVIDIA GPU is best for that, you might say the Cerebras chips are best for this kind of workload, because agents come in all different shapes and sizes. You've got customer service chatbots, that latency is a prime requirement as opposed to a deep deep research query where not latency but accuracy and broad search. Are there clusters forming in your view? Definitely. And maybe you use a slide. Yeah, this is what I was talking about earlier.
[03:21:27–03:21:56] So this is a way to visualize what's happening in a typical agent call. I guess this is a tongue in cheek slide that I made. Today if you look at agents, you give it a task, it goes off, tries to do it, thinks for a while, tries a bunch of tools, and then you have context switched. You're going off doing something else because it's taking minutes to maybe even hours says to do it. You've spaced out already spaced out,
[03:21:56–03:22:23] and then when it comes back and asks you for a steer or a decision, you have to page back all that context in and then you do whatever you do. And so our vision is we want to get to a world where the human is the bottleneck. Today, the AI is the bottleneck given how long it takes to execute all this. Really, we have succeeded from a compute perspective, when we have built the systems and the infrastructure such that the human becomes the bottleneck,
[03:22:23–03:22:51] that AI is finishing these things so quickly that you are constantly being asked for what's the next step. And that is a tongue in cheek point. But the better way to say that is, how do we make sure human is in flow when they're doing this work with AI. And there's this feeling of flow. And everything's so quick and interactive and it's like it knows exactly what you need and it does it quickly. That's a user experience we'd love to deliver.
[03:22:51–03:23:20] And so as we think about this future, we do need heterogeneous compute. You can't actually deliver this kind of experience economically on pure GPU based compute. So you need a much more heterogeneous infrastructure that's not just GPUs and CPUs, but also different kinds of accelerators. So Cerebras is an example that is for very fast inference. You might have other accelerators that are built for very long context.
[03:23:20–03:23:47] They hold a lot of state in memory, so they can remember your entire task and don't have to page it back in and out. For coding, for example, that could be useful. For coding for sure. They have to hold your entire GitHub project in context and be able to pull that very quickly. So you are going to see a lot more heterogeneity in the underlying infrastructure, because the user experience is going to push us towards optimizing every part
[03:23:47–03:24:16] of this agentic graph. And what we as people who have to build compute have to do is make sure we can match the right part of the workload to the right kind of compute to optimize on both efficiency as well as performance. Fascinating So this is going off script for a second, off roading. Yesterday, big day for earnings calls. A lot of hyperscalers talking about their accelerator programs.
[03:24:16–03:24:45] Amazon notably at roughly $50 billion of run rate revenue on their Trainium chips. I forget the alphabet number, but that's a bigger number. And then obviously you've got the big guy, NVIDIA. If you were to draw a market share chart, it looks heavily in the favor of NVIDIA right now. I'm sure there's all sorts of other ASICs that have not even seen the day of light yet. And obviously, the guidance from NVIDIA is we're going to do everything.
[03:24:45–03:25:12] The guidance from the others is similar. How do you expect this to trend? Is there one or two that you're a particular fan of outside of obviously the main workhorse? You know I'm not going to answer that, right? [LAUGHTER] But kidding aside, no, I think the world needs a much more resilient compute supply chain. I think it is dangerous for the world
[03:25:12–03:25:39] to be single threaded on any one component. And so I think that is what the market is reflecting. So we are going to see quite a bit of choices. And the workload is also going to push it there because the workload is getting a lot more complex than a pure inference or training job on a GPU. And so that is going to lead to flexibility. I'd say the other underappreciated part
[03:25:39–03:26:07] that I don't know whether everyone in will appreciate, the way TSMC allocates wafers will mean that there have to be multiple GPUs and accelerators. Say more about that. I think TSMC has done been extremely successful because they try to make sure that multiple customers are successful, and it is in their business interest
[03:26:07–03:26:37] to be so because they don't want to be single threaded on any one big customer. That is a single choke point in the supply chain. And so the way those wafers get allocated, there will be multiple people, multiple companies which will get wafers there. And by definition therefore, there'll be multiple varieties of chips. And so for the scale we are talking about, for the scale any one of us are talking about, Google, Amazon, us, whoever, by definition, we have
[03:26:37–03:27:03] to learn how to use all of these chips because we don't have a choice. And so that's why I think the world will look a lot more richer in the future. Fascinating. Maybe the other dimensions, Sachin, is the shape of the training workload, as you said, is fairly synchronous. It is typically coordinated. You need a coherent cluster that goes up right all at the same time. Inference, on the other hand, does not seem that way.
[03:27:03–03:27:32] It's likely much more spiky, a lot harder to forecast maybe. And as that changes, you might even want more compute closer to the edge to minimize latency for inference. Talk about that for a second. How do you manage the shape of your compute capacity knowing that you're moving towards an inference heavy world? Does that mean more distributed almost Cloudflare like mini clusters closer to the edge, or a giant one in Texas or Virginia
[03:27:32–03:28:01] is good enough? It will get there, but it's not yet. And for two reasons. One is there are still significant benefits to scale on building this compute. So building 50 megawatts of compute is far more expensive per megawatt than building a gigawatt of compute at one location. Fascinating per unit basis. On a per megawatt basis. Got it. And that's for many reasons.
[03:28:01–03:28:30] So labor is a big bottleneck around the world today, especially in the US. We just don't have enough people to build these things. So getting the kind of critical human mass you need to build, you would much rather do it for a bigger scale than for little bits, so 50 megawatts spread around the country. So that, I think, is going to drive the economics. The other technical reason is the way these models work, and especially for agentic workloads,
[03:28:30–03:28:59] the time to first token is still on the order of 400 to 500 milliseconds, because they have to page all of this context in before they generate the first token. And so 400 to 500 milliseconds is far larger than any latency benefits you get by putting compute closer to the user. And so to me that will also mean that this will push us towards more concentrated clusters of compute for inference still, for some time.
[03:28:59–03:29:29] And this will change as we figure out how to distill in very intelligent models to be small and potentially run closer to you. But at this point, the economics don't favor it. Got it. Follow up on that. Could you break down the 500 milliseconds into what are the different components of that call from the time that we pressed the Enter button on ChatGPT. If you were to allocate that 500 milliseconds, who's using that up?
[03:29:29–03:29:58] How much budget is each part of the stack allocated? I'd say the final milliseconds didn't even include some of that other components that you were talking about. But even for example, you ask a query on Codex, it's running off a project, it is going to take that prompt, combine it with your code base. That's the entire context for that model. I mean, to get technical for a minute, the prefill phase
[03:29:58–03:30:27] of running the inference, it basically has to run that entire context, which could be hundreds of megabytes. Like our Codex models now are 400k contexts. So there's 400k tokens. 400k tokens have to be computed through the attention mechanism before the first output token is generated. And so that is basically the model paging in all the context relevant to that task
[03:30:27–03:30:57] before it spits out even the first output token. And so that several hundred milliseconds of latency. After that you can add other stuff. Latency-- This is prefill. The first part is-- Prefill. This is prefill. And so after that you can add the other sources of latency that could be-- like it could just be your app turning your prompt into a token that is sent to the cloud and load balanced into the appropriate GPU to run to the model. All of that is going to add maybe tens of milliseconds of latency.
[03:30:57–03:31:26] So that's where I was saying that first token generation latency is higher than all the other sources of latency. But an interesting side effect. When we brought Cerebras in, and we rolled out Cerebras earlier this year, it started generating tokens so much faster that all of these other latencies that we had in the system, in the app, in the way our API works, started to become prominent. And so when we improved one layer of the stack,
[03:31:26–03:31:56] it forced us-- it actually showed up all the inefficiencies that we had in the rest of the stack. And so we had to do a lot of engineering to fix those latencies. We literally published a blog post on this yesterday. So we had to change OpenAI's API infrastructure structure to actually keep pace with Cerebras. And so if someone's interested, a lot of very neat software engineering that has gone into how do we shave off latency and every layer of the stack. What's the name of this blog? The OpenAI blog.
[03:31:56–03:32:26] OpenAI blog. Great, great, great. So it's like a whack-a-mole problem, similar to how folks were optimizing page load times on the internet. Yes. I mean, I think latency is going to be a very important dimension we will focus on. I think the trope is true, that every 30 or 50 milliseconds of latency you can shave leads to higher engagement, leads to higher revenue, leads to higher retention. For sure that is true. And I think that is going to be a dimension on which all of us are going to compete. Fascinating.
[03:32:26–03:32:54] That makes a lot of sense, and particularly given the attention of the human brain is only going one way, not expanding. Yes. Here's a fun question for you. Every guest we've had so far has mentioned that compute is the biggest bottleneck as an ingredient for their business. Probably true. What is the consensus that the AI community has maybe wrong or not right enough that you have reason to believe is misunderstood?
[03:32:54–03:33:21] What about AI infrastructure is most misunderstood right now? I guess the biggest shift that is happening that is underappreciated is we have very simplistic systems today. And what I mean by that is we have these big compute units attached to one layer of memory, which is high bandwidth memory.
[03:33:21–03:33:50] And I think we went through this in general purpose computing. CPUs started similarly. And then they added multiple layers of caching. They added flash. They added hard drive storage, all kinds of stuff. And so I think we are very early days in how systems infrastructure is going to evolve for AI compute. We've gone from very simplistic ways of programming these things to more sophisticated ways.
[03:33:50–03:34:15] I'd say the other big shift that's happening underneath is AI is generating the next generation AI infrastructure. And so what I mean by that is we are increasingly using our latest models to design the next chip and the next set of low level software needed to run the next model. Recursion, if you will, so how can
[03:34:15–03:34:44] I basically figure out what is the right kind of chip system and software it needs to run most efficiently rather than this decoupled world today where we train a model, someone else is designing a chip independently and delivering to us, and we figure out how to make it work. So how do we quicken that pace where basically the next model, while it is being trained, is also figuring out what should be the chip and system design it
[03:34:44–03:35:12] wants to run most efficiently. We are not that far from that world. Fascinating. Recursive algorithms are one of the most powerful algorithms, so this seems like a brave future. It is, I think. But it is also probably the only feasible way to bend the curve on the compute time. So time compute cycle, like how quickly can we get the right kind of compute designed and operational
[03:35:12–03:35:41] for the next generation, because otherwise we won't be able to keep pace if humans are going to try and interpret and then design and then do it. A typical chip design cycle is three years. Like from inception of idea or ideating on what a chip should be to actually getting it in production is three years, and that's too long given how quickly things are changing. Yeah, three years is right around when ChatGPT was launched. That feels like forever ago.
[03:35:41–03:36:08] It's an eternity. One of the questions we ask a lot of our speakers is this chart here. We talk about the five layer cake of AI, as Jensen describes it, energy chips, infra, models, apps. You play across all five of them. We're waiting for our chips to show up soon from Broadcom and others. If you were to guide us based on everything
[03:36:08–03:36:38] which part of the stack is most likely to accrue value in the long term, what would you point to? Obviously, all the money right now is in the bottom half of this layer cake. It changes. So I mean I think history rhymes. So if you look at the mobile revolution, initially a lot of the money was made by the telcos and the people building the infrastructure. Then it moved up into the application layer,
[03:36:38–03:37:07] the people building the apps. And then it moved up into the cloud services, cloud services layer. I don't see any reason why this cycle will be different. We are right now in the world where the infra layer is where the profits are, but over time it will move to the platforms and the apps. And so that is, I guess, the inevitable cycle here. We hope so. It seems that every app is getting engulfed by OpenAI and Anthropic certainly right this second.
[03:37:07–03:37:35] So we were eagerly waiting for that. Rapid fire question for you, Sachin, before we open it up. Long short. Pick a business, pick a startup that you're very excited about, that you'd go long on the other side. A counterfactual, a business idea startup that you're bearish about. I'm long OpenAI. I'm voting with my feet. But kidding aside, no, I think I'd say that
[03:37:35–03:38:04] the thing maybe for this audience that is underappreciated, I would go long on the lowest layer of the stack because at least in the US, we have forgotten how to build very foundational infrastructure. And that's from everything like how do we build transformers at scale, how do we build batteries at scale, how do we build generation and distribution,
[03:38:04–03:38:30] how do we build cooling, how do we build components that go into all of these systems. That is an underserved layer of the infrastructure. That is also one where differentiation is sustainable because it's both technical as well as scale. If you build it, it's very hard for other people to replicate it.
[03:38:30–03:38:58] So I'd say kind of a corollary bet on if AI is going to have that transformation that we think it will, the corollary with-- all this layer has to change from how it's done. And so for people in this audience, especially early in their careers, I think-- I was a faculty here, as some of you know, for 15 years, both in W and CS, and I saw dwindling enrollments in E,
[03:38:58–03:39:23] especially on the lower layers of the stack, especially around how do you do transistors, how do you do materials, how do you do that kind of stuff. That stuff is what will move the needle here. And so I'd strongly encourage going along that layer of the stack. Great, great. We have some E students in the class. What are you short, Sachin? What are you skeptical about? What are you cautious about? We'll lower the stakes.
[03:39:23–03:39:49] In general, obviously I'm short anything that is a model wrapper. That's a bit of an easy answer, but it is also true because the pace at which this thing is changing, and how quickly these models are able to introspect and figure out how to deliver an outcome, I'd
[03:39:49–03:40:18] say that it's very, very, very hard to just be a wrapper on top. So that is not a statement that OpenAI or even Anthropic for that matter, I would say the same, that we just want to build all the apps. I think this whole notion of apps probably to me is the one that I'd be short of. Like is that going to be the user interface of the future, unclear. Is it really going to be apps if we are going to interact with computing
[03:40:18–03:40:48] in the form of outcomes. This is the outcome I want, go figure it out. Today apps are a crutch to get to an outcome. And so that would be the notion I'd be short of. Fascinating. That makes a ton of sense. First company to $10 trillion in market cap if you were to pick one. The easy answer is NVIDIA, right? Yeah. I thought for a second you were going to say OpenAI. The first one you said. OpenAI will get there for sure. But I think if we are getting there,
[03:40:48–03:41:16] for sure NVIDIA is getting there. Good. Biggest unsolved problem in infrastructure right now? Oh, you name it. So I think so many. I'd say the single structural issue is enough fab capacity across logic and memory. Is the TSMC there? TSMC, Samsung, Intel, and Micron, SK, Hynix, Samsung,
[03:41:16–03:41:45] it's a very, very concentrated market. This whole thing is kind of single threaded on a very small number of companies. And probably if you dig down even deeper, it's ASML. For all of these, you need ASML machines. So to me, that is the single choke point of the hole supply chain. Makes a ton of sense. You already answered my last question which was advice for students. If you have anything to add we'll take it.
[03:41:45–03:42:13] Otherwise, we'll open it up for questions for a couple of minutes. Go ahead. Thank you for being here. My question is as we go through, and you mentioned [INAUDIBLE] very broad term, where do you think the next delivery is going to come from? Is it going to be the hardware like memory networking, would it be software? [INAUDIBLE] Short to medium term, it's probably in the orchestration software the harness, and the models
[03:42:13–03:42:40] getting more token efficient. Medium to long term, I'd say new memory architectures because I think the compute unit-- unless the transformer gets reinvented, like something replaces the transformer. You know what the compute unit shape is. It's really what is the memory architecture around the compute unit that's changing all the time. So that would be my medium to long-term answer.
[03:42:43–03:43:13] Go ahead. You walked away from Stargate UK [INAUDIBLE] about those decisions and the process for making [INAUDIBLE]. I think for us, Stargate is-- basically Stargate is my job. So it's how do we deliver all of this compute. And the way we look at it is given the size we are talking about-- like a gigawatt is $70 billion in spend.
[03:43:13–03:43:40] So these are massive numbers. And it's also operationally a big challenge. As I said, a gigawatt is half a million GPUs to manage, build up stuff and all that. So I think fundamentally the way I look at this is, how do I make sure that it's not just the absolute number, it also lands on time as quickly as possible? And so a big part in our approach
[03:43:40–03:44:10] is now time to compute rather than amount of compute. And so that's dictating kind of where we double down invest. And that's why the earlier question, we prefer bigger chunks of concentrated compute for that reason, because otherwise operationally it's very hard for us to get that compute online if it's lots of little chunks spread everywhere. Go ahead. I want to ask about open weight models. Obviously, weight models are continually [INAUDIBLE].
[03:44:10–03:44:37] The first part of the question is, how do you think that that's [INAUDIBLE]. The second part is from a compute standpoint, obviously overweight models usually have fewer parameters so require less compute [INAUDIBLE]. Yeah, I think obviously open source models have a role to play in the ecosystem. We frontier model intelligence is
[03:44:37–03:45:04] going to require orders of magnitude more compute. We don't see that changing. So the scaling loss continuing to hold. So we will continue to invest on that frontier model intelligence. Obviously open rate models will play catch up and try and distill that intelligence to deliver it in more compact form factors. And we don't see that as an issue.
[03:45:04–03:45:23] But a six month lead in intelligence is an enormous lead. And so we don't see any reason to back off on continuing to invest on frontier intelligence. Awesome, folks. We'll wrap it here.
[03:45:36–03:46:03] Today we're going to talk about this part of the stack, models and how do you build better models for better applications. Our guest for today is Yash Patil, Founder and CEO of Applied Compute. I'm so excited to have Yash here not only because he's a grad-- recent Stanford grad-- he's going to talk a lot about his journey--
[03:46:03–03:46:31] but also because he was one of the very few undergrads who went directly to OpenAI research after Stanford, was a part of the post-training team. Started applied compute after OpenAI because of an insight that he had during his work at OpenAI, and has built applied compute into one of the most successful businesses, applying his learnings to incredible enterprises. Yash, thank you for doing it. Please join us. [APPLAUSE]
[03:46:31–03:47:00] Thank you. Awesome. That's awesome. Cool. Thanks for having me. Thanks for joining us. Yeah. Yash, tell us a little bit about yourself. You've had an incredible journey, and you've made some tough choices. Actually, we were talking a lot with this class right before you joined about decisions about what to study, so walk us through your journey that led you to today. Yeah. So, I'll talk a little bit about my journey. It's not actually that long.
[03:47:00–03:47:30] I was actually sitting in here, taking finals not that long ago. I'm class of '25 so was-- hasn't been that long since I've been on campus and stuff. Yeah, so grew up in Austin Texas, came here for school. I like to say I was a very good student in high school. I was a very bad student here. Not grades or anything, but kind of never went to class, watched the online lectures,
[03:47:30–03:47:59] did that sort of stuff. And the rapid fire history is ended up building a bunch of stuff on campus, got connected to Sam Altman very serendipitously through some mutual friends. One thing about Sam that I think not many people know is he has an incredible soft spot for helping young people early on in their careers. So, yeah, I ended up meeting Sam. We hit it off freshman year, summer.
[03:47:59–03:48:26] A friend and I were kind of deciding, hey, do we want to do our summer internships? Do want to go and work on our project? Ended up saying, hey, let's go work on a project, reneged on our internships, and we were kind of looking for money. We shot Sam a blind email. He gave us a very small check to cover food and rent and things like that, so worked on that for the summer. Ended up shutting it down, coming back to school. So that came back for my sophomore year.
[03:48:26–03:48:56] I was doing a lot of fun things here on campus like Tree Hacks. Shout out to Tree Hacks. That was kind of my main thing here. I really loved putting on that hackathon for folks. But then late 2022 is when ChatGPT came out. And I was just like playing around with it, and I was like, holy crap, this is the coolest thing I've ever seen. I have to go work on it. I couldn't think about anything else. So I ended up shooting Sam another email saying, hey, how do I come and work on this thing?
[03:48:56–03:49:22] He put me in touch with what was called the OpenAI residency. I think it still exists. It's actually how a lot of folks at OpenAI went from academic researchers, or people in different industries, to full time employees at OpenAI. Joined OpenAI early 2023 on the post-training team. Worked with some people I really, really looked up to in the language model universe, starting on Evals, which a tip for anyone here
[03:49:22–03:49:51] is whenever you join a company, work on this hairiest thing that no one wants to work on because people will like you for it. So ended uo working on Evals for the first year, and then the second year was kind of when these reasoning models started coming out of the woodwork. So people were training these reasoning models on primarily competitive math, and it was this wow moment for everyone at the company where we're like, oh, we're seeing massive performance increases in using these models.
[03:49:51–03:50:20] So a friend and I were like, hey, what if we actually try applying these models to things outside of competitive and coding and math? I wasn't a competitive coding or math kid growing up. A lot of the frontier features folks were. So we ended up hacking together this agent that could browse the internet, write some code. Showed it to a bunch of leadership. They were really excited about it, so we started this team called Long Horizon Tasks. And I was primarily focused on leading a lot of the agentic coding research, which eventually
[03:50:20–03:50:48] became Codex. But, yeah, left to start applied compute about a year ago. So our one year was actually last Saturday, so hasn't been too long. But, yeah, we saw this gap where these models were getting really, really smart but when you actually went to go and apply them inside of the enterprise, they're like smart geniuses that know nothing about your business. And inside of enterprise is actually
[03:50:48–03:51:13] where you have most of the data in the world. All of these companies have tons and tons of data, proprietary data that they've built up over time, so we're actually helping companies take the same frontier technology that led to these smart reasoning models and create their own specialized models to enhance their business. Amazing. What a journey. Yeah, it was a lot of fun. Yeah. I thought we'd start at the models, Yash.
[03:51:13–03:51:41] So, I went to the past century, things have clearly escalated. And then if you zoom in to this side of it, post AlexNet, things are moving pretty fast. Yeah. Could you put this in context. What is going on at the model layer frame for us? Why is the advancement in the last four years notable? And what is driving it? Yeah. So-- I got some of your slides if you want. Yeah. So-- Scroll through them.
[03:51:41–03:52:10] --I thought we'd start with a bit of history. You mentioned AlexNet. Does anyone here-- have you guys heard of deep learning or what that is? Of course. Yeah. OK. Nice, nice. Deep learning was-- AlexNet was, I would say, the pivotal moment for deep learning. And it also is the moment that we stopped understanding what any of these models actually do. So essentially what deep learning is it's a method-- is a piece of machine learning technology
[03:52:10–03:52:37] that allows you to learn underlying representations from data. And you basically train on a bunch of data, push in a bunch of compute, and you get out these really smart models that are made up of millions or billions of parameters. You actually don't know what they do, but they actually are really good at doing tasks like prediction, language model. Like next token prediction is what LLMs do.
[03:52:37–03:53:06] And the before and after was before AlexNet people were creating handcrafted features, looking at pieces of underlying data, training very, call it-- called rudimentary classifiers to detect edges, things like that, on a lot of vision tasks. And then what AlexNet did is they applied GPUs, a massive data set called ImageNet, against neural nets, which led to the breakout moment
[03:53:06–03:53:35] where you actually proved that, hey, if you scaled compute and data, you were able to see these massive gains in predictive accuracy of these models. The model development has a lot of different aspects to it. Maybe give us an overview-- I know we're going to go deep into a couple of aspects of it-- what is model training-- the modern day model training look like. Yeah. So, I think this is a snapshot-- this is by no means the most detailed timeline of events,
[03:53:35–03:54:05] of things that have happened, but I wanted to pick out a few pivotal moments in recent years, starting with the transformer. So I'm sure everyone here has heard of the transformer and self-attention. This was the moment where researchers at Google Brain came up with a new architecture that actually allowed scaling language model training. It was way more performance on existing hardware, so, like, could actually run these workloads on GPUs.
[03:54:05–03:54:34] And compare it to previous neural net architectures like recurrent neural nets or LSTMs, they were able to employ this technique called attention, which basically led to way better performance in next token prediction and could scale to these massively long sequences in language. Fast forwarding over the years, 2018 to 2019 was really this era of pre-training. So people were taking these massive corpuses of text,
[03:54:34–03:55:01] teaching models to predict the next token by optimizing on loss. So basically you have a model try to predict the next token, you see what the actual next token was in the corpus of text, and then you do back propagation on the model weights to tweak the model so that it's more likely to predict what that ground truth next token was. Then you entered the era of scaling laws. So starting with the OpenAI scaling laws which showed, hey, if you actually scale these models up and make them really,
[03:55:01–03:55:28] really big, you start to get much better performance. So this was like the Kaplan scaling laws really proved out with GPT 3, which was the first kind of model that seemed to have some level of general intelligence. So that was kind of a breakthrough moment. And then you continued to compound on this with the Chinchilla scaling laws which showed, hey, not only do you need to make the model really, really big, you should actually-- there's actually a compute optimal way to scale these models. You both make the parameter size much larger,
[03:55:28–03:55:58] but you also should train it on much more data. And then once we started to have these models that were generally useful, it became a story of how do we actually make them useful to the normal person, how do we create these inner interfaces where it's like a tool that the everyday person can use. And this was the era of reinforcement learning with human feedback preference tuning, being able to steer these models because general based models are
[03:55:58–03:56:24] just doing next token prediction. So they hallucinate a ton. They don't actually answer your question. They may say things that are unaligned or not up to safety standards. And then GPT 4 was this next level step change in the quality of these models. I'll go through these very quickly. I think this is what people are probably most familiar with in the past couple of years, which are the era of reasoning models.
[03:56:24–03:56:53] So in 2024, OpenAI came out with this model called O1, which was this new axis for scaling model intelligence, which was test time compute. And this was kind of-- felt this-- I want people to know that chain of thought is like a completely emergent behavior. The model reasoning, whenever it answers your question and spending time thinking, correcting itself, no one trained it to do that.
[03:56:53–03:57:23] Basically by putting it in these constrained RL environments and then funneling a ton of compute towards it, you actually got these models who had this emergent property to be able to reason. And then combining that with tool use, which a lot of people-- if you guys use Cloud Code, Codex deep research, you actually start to get these agents that could reason, work for a really long periods of time, and become what people are calling today AI coworkers. Fascinating. Do you have a slide-- OK, great. This is perfect.
[03:57:23–03:57:52] So, yeah, we have understood there to be multiple ingredients of-- that go into making a great model data, compute, talent, algorithms, maybe other things that I haven't listed. What is the bottleneck today? What has been the bottleneck in the past? And what do you suspect will the bottleneck be in x years from now? Maybe a tour of history and then a prediction for the future. Yeah. So, I mean, kind of running through what we just talked about, the bottlenecks kind of went
[03:57:52–03:58:21] from having the compute to train these models, to the correct architecture, to actually being able to scale to the pre-training levels of data that we need, like the training on the entire internet, being able to make these models usable by preference tuning them. And then today, what the bottleneck is-- I'm sure people have heard of RL environments, which is actually how these very recent families of models have gotten so good at reasoning and intelligent thinking.
[03:58:21–03:58:49] I think what the bottleneck is for the future is it's basically this idea of continual learning, which I know we're going to talk about a bit later. But so far, we've gotten more and more data efficient methods of training. So pre-training, not super data efficient. Like, you have to train on the whole internet to get the base model. People here-- to learn something, you don't need to go and read internet scale data, so that was super data efficient.
[03:58:49–03:59:18] Then we started to go to these methods and post-training, which were a little bit more data efficient. The most data efficient today being like RLI environments. But the thing that is the Holy Grail, and I think what people think of when they think of ASI or AGI, is how can a model go and do something once and learn from extremely sparse reward. So just like you guys, if you go and burn your hands on the stove, you just need to do that once and then you
[03:59:18–03:59:45] know the stove is hot and not to put your hand on the stove. These models today are not really like that. So I think this idea of continual learning and being able to be extremely data efficient with in real world interactions, that's the next bottleneck. Gotcha, gotcha. Super helpful. Well, when you burn your hand, it's also a very loud feedback. Yeah, exactly. So hopefully continual learning is giving us loud feedback. Yeah. And the other big question maybe not on this slide
[03:59:45–04:00:15] is, why have all the labs converged at focusing on software engineering? Why have they focused on code as the first frontier? I'm sure there'll be other frontiers like life sciences or cybersecurity or others that I don't know about, but why software engineering? What is the unique property of code that people find so interesting? So, the type of RL training that these labs are doing is reinforcement learning with verifiable rewards.
[04:00:15–04:00:44] So, in order to actually get the learning signal or the reward signal, you need to have a deterministic way to check if what your model did was the correct thing. Code and math are really, really good for this because, what can you do? You can compile the code, you can run unit tests against it. You can actually check to see if the code is doing the right thing. So, that is one reason why code has been super valuable. The other is like it's really easy to make a lot of synthetic data on this. Scale of data.
[04:00:44–04:01:13] The scale of data. The prior is really good. There's a ton of code tokens on the internet. And then I think the other thing is just a lot of researchers, me included, think coding models are kind of like AGI complete in the sense that every task, when you boil it down, is a coding task. So that's why you see Claude and a lot of these other models writing code to do-- instead of doing tool calls or things that
[04:01:13–04:01:40] are more specific to the task, they're actually just like using code as a general language to interact with the real world. One of the questions we were discussing right before you came is, how do we get good at jobs with AI that are not code or code adjacent? Give you an example, making slides like this one, by the way, completely generated with Cloud Code with an initial set of conversations that we had.
[04:01:40–04:02:08] What is the relationship between code and slide generation that would make these models good at generating slides because they get good at code? So did you make these slides with Cloud Code? All of the formatting. Oh great. OK, yeah. So, basically, I've also made slides with Cloud Code. And basically it's able to make this table, it's able to set the formatting on the title, it's able to put that random blue line there.
[04:02:08–04:02:36] But what you can do is you can combine the outputs of this model with other auxiliary rewards to tell it how good the code was that it wrote. So here, this is extremely functional, but if we wanted to really optimize for aesthetics, we could combine not only the code execution and actually being able to make the slides in there structurally relevant, but also some sort of reward model that can look at the output of the slides. And it's been trained on human preferences
[04:02:36–04:03:05] of what aesthetically pleasing slides look like and what ugly slides look like. And you can combine those rewards and jointly optimize for both writing the functional slides and then also making them look pretty. Right. Well, this is a great segue into talking about your work. Tell us a little bit about pre-training, post-training, what do those words mean? And I know you are the world's best expert on one of those so we'll dig into that. That's very flattering.
[04:03:05–04:03:35] But, yeah, I mean, in language modeling, I think the two big buckets to talk about are pre-training and post-training. Pre-training is this massive training effort where you take internet scale data-- trillions and trillions of tokens-- you throw a ton of compute at it through this architecture, the transformer, and you train a neural net to get really good at learning patterns in language. And the thing that falls out of this
[04:03:35–04:04:03] is some form of intelligence. So, essentially, what pre-training is is this idea of compression where you're able to actually take all of human knowledge-- i.e. the internet-- and put it into a set of weights that actually understands the patterns and language, how to think about things, whatnot. The problem is once you have this pre-trained model-- which, by the way, takes like orders of magnitude more compute than post-training-- you actually need to go and align this thing. So it's just next sequence-- next token prediction.
[04:04:03–04:04:33] So if you write a sentence like, who should I invite to dinner? And it starts to basically say a bunch of random names or something like that, that makes no sense because you're-- really the model should be like, oh, I have no idea who's on your invite list or who you know. Please tell me who these people are. So post-training is actually the process of taking this model and telling it what good and bad outputs look like, and you actually get a model that learns a chat format where
[04:04:33–04:05:00] there's a user message and an assistant message that responds to you. You learn safety guidelines. So if a user asks, how do I make like a weapon or a bomb or something like that? You can actually tell the model, hey, don't tell these people how to go and make these harmful weapons. And in both of these cases, what's scarce is data. So, I have this slide here, which is very dense.
[04:05:00–04:05:29] You guys probably don't need to look at it too closely. But essentially what pre-training does is it's just optimizing for loss. So how do you get really good at predicting that next token? But what happened is-- we talked about this Chinchilla scaling laws, you're scaling model size and you're scaling the data that you're training these models on-- we've just ran out of data. Like, there's only so much data on the internet, and we're sort of-- I know you have a slide later talking about this,
[04:05:29–04:05:57] but we have approached the frontier on what data is available to these labs. And so the labs are the only ones that can do this level of training because it requires so much compute and so much data. It's a huge CapEx requirement. And then I think on the post-training side, you have all these different methods for training models from supervised approaches like SFT,
[04:05:57–04:06:25] to this preference tuning RLHF to RLVR, which we'll talk a bit more about. There's a really good article actually in the readings-- if you guys read Karpathy's writeup on RLVR in the readings that talks about what happened in 2025. RLVR really came to prominence in 2025. We'll get into that in a second, but, actually, if you just click one slide forward. Oh, actually, here, can I? Yeah. I had a question on data for you.
[04:06:25–04:06:54] So, there's a lot of-- we spoke about this-- data, running out of data. This is the point you were making. We had Ali Ghodsi here two weeks ago. And he spoke about, actually, most of the data beyond this frontier is going to be AI generated. A lot of tokens on the world will be just AI generated, given the volume of which they're being generated. Talk about that for a second. What is the frontier of data? Where do we get more data from here on?
[04:06:54–04:07:24] The model that will be trained in, let's say, 2030, what is the input to that? And where do we get it? Proprietary, public? Yeah. So, I think there's multiple layers to this question. The first is-- And there's a whole economy-- sorry to interrupt you. There's a whole economy of these companies-- Totally. --whose full time job like Scale and Mercor and others is to-- maybe touch on that as well. Yeah, exactly. Yeah. So, this I think is a visualization of pre-training data, which is really just about scale.
[04:07:24–04:07:53] You have people who are starting to buy old libraries with ancient books in them, going and scanning these books to get more tokens. You have a lot of investment in synthetic generation, so how can you take a primary source documents and explode them to multiple-- orders of magnitude more tokens, and see if you can learn more from that? So I think pre-training, that is going to be the methodology. Really, what people are focusing on now,
[04:07:53–04:08:19] I think, in pre-training is new architectural advances to actually make better use of the data that we have today because on principle, you shouldn't need internet scale data to learn a lot of this stuff. So people are trying to be like, OK, how do we actually use the data better? So there's all these data wall challenges that the frontier labs are working on. Now, what you're talking about is RL environments. So this is a different type of data.
[04:08:19–04:08:47] This is, hey, let's actually construct the world that the model operates in, have it go and do a bunch of things, and then exchange compute for less high quality data. So basically like we can use way, way more compute and learn a lot more from a single sample or rollout. So we were talking about code a little bit-- and we can talk about this on the RLVR slide-- when you make a code environment,
[04:08:47–04:09:17] it's not like pre-training where you're going in training on a code base and just learning all the tokens in the code base. You're saying, hey, I want you to go and implement this feature. You actually have the model try it hundreds or thousands of times. You have a way of checking if the model actually did the correct thing, so that verifiable reward. And you get this distribution of rewards because sometimes the model will go and do it correctly, sometimes it will do it wrong, and you're actually able to learn way more from that type of training
[04:09:17–04:09:43] than you are from pre-training alone. So just next token prediction. Fascinating. And maybe the same question for Evals. I think you had something on Evals. But doc Evals, why are Evals important? Why do labs guard their Evals? There's a lot of chatter about this being the most protected asset. Tell us all about it. And as you referenced, this being the hairy job you've done, why is that the case?
[04:09:43–04:10:13] So I think, as we start to train these models on essentially reward functions, what becomes the most important thing is actually knowing what good and bad look like. So Evals are a way of benchmarking your model and given a certain task, understanding how the model acts. The reason why Evals are so important to the labs is because Eval set the roadmap.
[04:10:13–04:10:40] So if we want to go and train a really, really good code model, you basically-- Cinebench I think was the Eval that started the whole code model race. And that was because people had something that they could optimize towards in terms of what is useful coding look like. Now, Cinebench I think is a very flawed Eval, and there's been a lot of new Evals that have come out since that are much better, but that's the whole point is like whatever hill
[04:10:40–04:11:08] you want to climb, you first define it with an Eval. Then RL is this Eval maxing machine, so you go and create a training pipeline that looks very much like your Eval, obviously different data because you don't want to overfit directly to those Eval data points. And then you just climb that hill, and then it's on to the next Eval. So, this is also particularly important when it comes to enterprises, enterprises internally have their own idea of what good and bad looks like.
[04:11:08–04:11:37] Good and bad is not the same across like a JP Morgan and Goldman Sachs. They have different standards, they have different ways of operating, so they will have their own Evals. And you get this tiered effect where there's these Evals that the model labs optimize towards, and then there's these Evals that the enterprise is optimized towards. And we're actually that layer, that specialization layer applied compute to help enterprises optimize to their specific Evals. Fascinating. Good segue into applied compute. Yeah.
[04:11:37–04:12:06] So, what led you to start applied compute? Why better to do it as an independent business than inside of OpenAI where you were before applied compute? And what do you guys do? Yeah. So, like I mentioned, we started applied compute a little bit over a year ago. And it was-- I started with my co-founders Rhythm and Linden who were actually both students here at Stanford. We were also all at OpenAI together.
[04:12:06–04:12:34] Funny story is like when I joined, Sam was basically like, who's the smartest person you know? That was Rhythm. A couple months later he asked Rhythm the same thing, that was Linden, and that's how we all ended up there. But we really started applied compute based on this core idea that the future is very specific to enterprises, where you'll have these general models which are workhorse models, but actually going and specializing them towards individual enterprises needs
[04:12:34–04:13:04] is actually going to be how people differentiate. So general models set the floor, but in order to set the ceiling, you need to go and build, train models, create these specialized systems in order to differentiate yourself from all your competitors. So, an example here is DoorDash is a customer of ours. I'm sure a ton of people order it all the time. I'm very guilty of ordering it way more than I should, but-- It's all a part of the RL environment. I know, exactly. We were just testing the product.
[04:13:04–04:13:29] So one of the tasks that we worked on DoorDash with-- and I'm picking this one because it's very practical. It kind shows what we do. DoorDash onboards like 100,000 plus merchants every year to their platform. And when these merchants come to the platform, they basically supply a bunch of unstructured information about their business, including menus. And menu extraction, actually being
[04:13:29–04:13:58] able to go from images like this to a DoorDash storefront is actually a really, really hard task because-- I see, [INAUDIBLE] that. Yeah. Exactly. DoorDash has this very specific style guide for how modifiers are supposed to be attached on top of items, what you can mix and match, what's an add on versus a special ingredient, things like that. And when we tried using the general models on this, they just weren't able to do that task.
[04:13:58–04:14:27] So instead of-- and we tried prompting and all this stuff, what actually ended up being the solution is you could take outputs of our model, you could have humans go and correct those menus and understand the delta, and then we could basically, during training, have a model-- have a model's output be checked against the ground truth. And essentially we had a way of quantifying the loss or the reward, how the error rate, essentially.
[04:14:27–04:14:55] And we were able to just optimize directly against reducing error rate. So this is a very clear example of how a company just needs to go and define what good and bad looks like. And you actually don't need to do prompting or any this stuff. You can just directly optimize towards the outcomes that you want. Fascinating. Can I ask you a follow up on this? Yeah. So, the prior gen, let's say before transformers, this would be a problem that an OCR model would have been applied to a vision model of some kind. Yeah.
[04:14:55–04:15:23] Is the part where you guys come in optimizing the specific problem, and this is not using a vision model, using a transformer model? This is using a VLM, yeah. So it's like a vision model with a transformer architecture. Fascinating. Why would DoorDash-- and sorry to ask you a hard question on the spot which is off curriculum. No, no, no. Why would you guys specialize in existing model when maybe there's a chance that--
[04:15:23–04:15:50] I'm making it up GPT 17-- might be out of the box much better. Are you incentivized to just wait for the next series of models, or should you work with applied compute to-- Yeah, no, no, no. --brunch out? It's a great question. I'm glad you asked me. I think what people often don't realize is that enterprises care about being at the frontier at any point in time. So GPT 17 is going to be, quite a long time from here.
[04:15:50–04:16:19] I think by the time we have ASI or this model that kind of controls everything, which actually I don't believe that's going to happen. I think that the world is just very fragmented place. And if you just look at where the data is, it's kind of dispersed. But, yeah, the time to value is just way-- the Roi on being able to train your own models today with way less compute RL has become very, very data efficient. You need to use an order of magnitude less compute than pre-training, or these other types
[04:16:19–04:16:49] of training. And you're able to optimize performance way more than you were in SFT and RLHF. Makes it a lot more appealing to be able to train these models like Roi. The order of magnitude investment that goes into post-training-- pre-training versus post-training, could you give us an estimate for that? So let's say it was $100 budget for training, how much that would be pre-training, how much would that be post-training? Yeah. So, I think I looked it up on the way here.
[04:16:49–04:17:15] DeepSeek V3 was trained on about 2.4-- 2,500,800 hours. The RL training-- so the training that led to DeepSeek R1-- was trained on about 150,000. So comes out to about 5% of the training-- compute that's needed for pre-training. But-- That's it? That's it. But I think what's really interesting
[04:17:15–04:17:45] is that trend is starting to change where people are pre-training these models, but then they're also doing data center wide, multi-data center wide RL runs because you have these scaling laws-- which I think I had a photo of Jensen, the three scaling laws-- there's pre-training scaling, then there's post-training scaling, and then there's test time scaling. Test time scaling is inference. But post-training scaling, you can actually massively increase the batch size of each of these training steps
[04:17:45–04:18:12] and you get a lot better performance. You can have these models do a lot more reasoning when they're attempting these tasks. So I think the trend you're actually seeing is that the compute spent on RL is actually increasing quite heavily. Got it. So it's 5% today, but you expect it to go up-- It is going up. Yeah. --as relative percent of the total training budget. Yeah. So, I don't know the latest stats on a mythos or a 5.5,
[04:18:12–04:18:41] but up until-- basically when we were scaling out 01, 03 Codex deep research, we basically saw more compute you put into RL, the better performance you get, so these things stack. Makes sense. Yeah. Maybe a couple other examples, I found that to be very, very helpful. Are there other examples of domains outside of, for example, converting this menu to a DoorDash that you could share with us of where this is a particularly useful application? Yeah, yeah.
[04:18:41–04:19:10] So, I mean, we were talking about coding before. We recently just put a model in production with cognition/windsurf. Basically the idea is when you're writing code and you write-- you save your file, how cool would it be if you had a model that kind of ran sub two seconds, checked what code you wrote, and then told that there was a bug in it or not. So, this is not something you can get with a general model because there's this Pareto frontier of performance, cost, and latency. So if you take a really small model,
[04:19:10–04:19:39] you post-train the heck out of it on getting really good at this task of bug catching, you're able to get the benefits of cost and latency and the performance of some of these larger, bigger models. Got it. And so the value add for cognition and Vint Cerf is that they're extending their product suite from just writing code, but also now to testing and bug-- Exactly, yeah. And I think this gets into this really interesting idea of model
[04:19:39–04:20:07] harness context code development, which is like you never really can focus on just one layer. A lot of these application layer companies are doing a ton of innovation on the harness, and that's actually how they're able to squeeze value out of these models, especially when it relates to the service that they're providing. And then context is just like oftentimes if you don't have access to the right data, you won't know the right thing to do.
[04:20:07–04:20:34] So being able to plug into all of these different data sources inside of companies, that's also extremely, extremely important. Fascinating. So this is-- in this case in the case of cognition, it might be a true competitive advantage for them to expand their product frontier. Exactly, yeah, yeah. I mean, I think you're seeing people start to push the frontier on what these models can do, and it's usually an ensemble of models. General models, extremely powerful, really good
[04:20:34–04:21:02] orchestrators, but, fast subagents or agents trained on proprietary data that's out of distribution of these models, something like this, those can be orchestrated with the general models and to create a really powerful system. So I think today, actually, Ramp Labs, the corporate card company which we're actually a really good friends with them and know a lot of folks there, they trained this RL model
[04:21:02–04:21:32] to basically do fast search inside of your spreadsheets. And that's a way you can actually go and improve the product experience. Fascinating. Yeah. We'll change topics, Yash, and talk about a bunch of these emerging model training techniques that we've been hearing about, and maybe you can decompose those for us. Love to, yeah. We'll start with the continual learning. What is it? A lot of smart people we talk about that as the next frontier, including you, break it down for us. Yeah. So, continual learning is really what I mentioned before,
[04:21:32–04:21:55] which is like, how do you learn from extremely sparse rewards? So, if you have a system deployed in production, how are you actually able to understand how that AI model is being used, understand the downstream consequences of its actions, and then use that to update the system so it gets better over time?
[04:21:55–04:22:24] So, what I have here is two examples of-- yeah, two examples of what I think this is starting to look like. And to be clear, I think this is going to be a very gradual thing. So, a lot of continual learning is blocked on just having access to the right data. So when you go and deploy an agent in production, are you putting it in front of the right people to get feedback? Are you actually deeply understanding all of
[04:22:24–04:22:54] the context necessary to know what good and bad looks like? That's like just a data access problem. So I think this is going to look kind of a slow, gradual rollout rather than, oh, there's some extremely valuable insight that someone comes up with. But a couple of examples of how you're seeing this today is Cursor, they have this model called Composer, which is essentially their own coding model trained on their coding data on top of an open source model.
[04:22:54–04:23:23] And what they did was really cool. They basically took this model, had people use it in production, were able to capture a bunch of that telemetry, take steps online, so take a training step based off of some implicit rewards that they calculate. So basically they would look at, did the user accept this code suggestion? Or did they revert this code? Or did they-- Stick for success. Exactly, yeah. And they kind of optimized towards it, and they were able to see improvement
[04:23:23–04:23:52] by doing this online training where you're collecting data, taking a train step, collecting data, train step. Then this is something that we've been doing. Actually, just to follow up on Cursor. The x-axis here, Yash, is steps. Could you convert that for us into time? Like order of magnitude, how much time did Cursor have to invest in improving Cursor's performance? Yeah. Was it like days or weeks or hours? It's a good question where I think we're talking about days or weeks here,
[04:23:52–04:24:20] and then I think a couple hours per step. I don't know the exact terms, but one interesting thing here is in RLVR when you're training offline, you have this replayable environment where you're like it's the same task, it has a defined reward, and then you're rolling out that sample hundreds or thousands of times in parallel. You can't do that in production because people are using this-- they have dynamic environments or whatnot.
[04:24:20–04:24:45] So what they really experimented with was, could we just take a massive batch-- so like many, many, many conversations-- denoise the gradient that way and then take a step, and hopefully that'll be like directionally the way to improve the model? So, yeah, so this is, I think, hours per step, and then each of those steps is quite big or a lot of samples in them. On the right here is something we actually
[04:24:45–04:25:13] have been working on, which is this idea of context base, which is like, can you actually go and use agents expend compute offline to be able to go and analyze a bunch of documents, analyze a bunch of past traces that humans have had with agents, and extract learnings from that that will improve performance downstream? So one thing that we were able to see is like, yeah, we were able to see, at different reasoning
[04:25:13–04:25:41] efforts, a massive increase in performance while using the same amount of tokens. So, yeah. So these are just two examples, but I think high level, you're going to see innovations of weight updates context and the harness itself to actually be able to capture this information. Fascinating. Yeah. More to come on this. Yeah. Second topic is non-transformer models.
[04:25:41–04:26:10] A lot of talk in the class about transformers not being a very efficient architecture, takes up a lot of power. We had Ali Ghodsi who said that well, look, flying like the airplanes is far less efficient than the birds do. Turns out we are heavy specie, heavier than birds. Is the transformer like that in that it's-- while not efficient, will be the dominant way because the world's infrastructure has morphed and moved along that way? Or do you think there's a shot for a transformer architecture,
[04:26:10–04:26:40] like the Mamba architecture or one other to be a dominant player in AI models going forward? Yeah, I think my honest take is like scaling transformers is working, and as there's a very simple recipe to be able to make these things smarter and better. And probably more likely that the AI will tell us what the better architecture is if we just continue scaling it up than try to come up with one ourselves. If there was a wall in terms of what this architecture would
[04:26:40–04:27:04] allow us to do, then I think we would see some innovation there. And there certainly is really cool research happening, but my opinion is just like concentrate on scaling, scaling transformers. Very smart people on the other side of this debate, as you know. Ilya, Yann LeCun and others going at it. If you can share, do you know what the core insight is that that leads them to believe the other side of the debate and disagree with you? Yeah.
[04:27:04–04:27:34] So I think the core insights are you-- what I said before is you don't need pre-training levels of data to be able to actually learn the underlying representations of language. I think Yann LeCun talks about this a lot where it's like humans don't need that, therefore, the architecture that we developed shouldn't require that. So I think that's the underlying argument is like first principles like you shouldn't need it, therefore, there must be a better solution out there.
[04:27:34–04:28:01] But I think, to your point, the investments we're making in our compute scaleouts, I think there's people who are actually optimizing for the architecture directly in the chips. That is, I think, a big ship to turn. And so far, what we've seen from the labs, which are going to I think control a lot of this build out is just investing more in the transformer architecture. They're definitely doing research on new stuff
[04:28:01–04:28:30] but I think it's all experimental and could work. What a time, because big large sums of money are going into these techniques, so we'll see how it shakes out. Rapid fire last five minutes, Yash. A lot of folks in the class deciding where to build, what part of the stack to build, what to start. If you were not building applied compute today, what would you be doing? What's your next best idea? Yeah.
[04:28:30–04:28:57] So we're thinking a lot about this because I think one thing that we've been running into at applied compute-- and I know a lot of other AI companies are running into-- is like scarcity of compute. I think the demand is just far outpacing the supply for compute. And I think there's going to be massive innovations in the energy sources needed to power this compute and then also making more efficient chips themselves.
[04:28:57–04:29:25] So not that I have a background in hardware or chip design or anything like that, but I think we could be making way better hardware to optimize the co-development of training and chip design. So I would probably look into hardware. Thank you. The long short game, pick a business, a product, a person that you like a lot, that you're excited about. And the counterfactual, something
[04:29:25–04:29:53] that is more hype than there's reality. Yeah. So I think, goes along with my last answer, like compute and the chip providers like NVIDIA, I'm very long them. I think they're going to continue to win. They're going to continue to supply all of the labs. And, actually, it is interesting though, like, once you look at the compute economics of NVIDIA,
[04:29:53–04:30:22] they take a 75% margin on top of their chips. You have these labs spending hundreds of billions of dollars. Question is like, hey, maybe we take a couple hundred billion dollars and invest that in actual-- our own chip design, then we could do all the co-development of model training architecture and chips internally. And maybe our chips are 80% as effective, but we'll just make like 1.2, or whatever it is, x more. So I think that is--
[04:30:22–04:30:49] generally I think chips and NVIDIA is going to be the leader here. I think that's valuable. But I do think there's also a risk of, hey, the labs are-- the biggest customers are maybe just going to in-house, do this stuff themselves, but it's very hard. One thing I'm a little less bullish on I think is the data market is just really tough. So if you think about an RL task,
[04:30:49–04:31:17] you have basically like models that are not very smart, you train them to be very smart at particular tasks, and then it becomes that much harder to go and create new tasks that you can actually climb on. So when we were-- Because the problem is much harder. Exactly. You're getting squeezed where it's like, OK, yeah, I'm selling to my customer and I'm improving their model, but I'm also making it harder for myself because the next time they come to me and say, hey, I want you to go and build this task,
[04:31:17–04:31:47] I have to spend way more money, it's going to take a lot longer, all that sort of stuff. And then I also think the models are just getting really, really good. So I think you're starting to see a lot of synthetic data generation, because if you think about an RL task, a lot of times what it is exploiting some sort of generator verifier gap where for code, you hold out the unit tests, you have the model attempt the task, and then you run the tests against the model's output. That's not something that you really need a human to do.
[04:31:47–04:32:14] The smarter your models get, the better pipelines you can build around synthetic data. So I'm not like-- I think data has been a thing that people have been like, oh, it's going to die every couple of years and it hasn't, but I do think it's going to change. And the best founders in the data market kind of are just really good at pivoting and what's the next wave. Robotics data, egocentric data, like put on a GoPro and collect a bunch of stuff.
[04:32:14–04:32:44] So I think RL environments today is going to be tough, but they'll go on to the next wave of things. Awesome. Final question. Your favorite AI product or favorite AI modality. Yeah. Oh, this is a good one. I love GPT-- or sorry, Image GPT 2. Image2. Image2, yeah. When it came out, I was having a ton of fun with it. I think it's massively-- like for people who can't do design
[04:32:44–04:33:13] like me or who want to look at things visually-- a lot of times I'll just take some paper, or something, and drop it into a Image2, and it gives me this nice visual, walkthrough of how things work. So that's probably been my favorite AI product. The most beautiful slide on this presentation was an image tool representation. When I fed it the course syllabus, it was like-- how about this one. Oh, wow.
[04:33:13–04:33:32] And the signature, for those of you who can tell, is this one right here. That's the images to watermark. Oh, really? Oh, interesting. Anyway, awesome. Well, thank you so much for being here. No, thank you. We'll continue this conversation and see how these bets play out. Let's do it. Awesome. Thanks for having me. [APPLAUSE]
[04:33:46–04:34:16] Well, tonight's guest needs no introduction. I'm sure a lot of you have shipped code on Vercel. Guillermo is obviously the founder of Vercel, but also he created Next.js, the most popular React framework in the world. Vercel is now the most important developer infrastructure companies, valued at $9.3 billion. This next part that I'm about to tell you, even I did not know about, G. The best part about your story-- yeah, go ahead-- is that you grew up
[04:34:16–04:34:44] in a suburb of Buenos Aires. You taught yourself to code as a kid. Learned English by reading software manuals. Typically, it goes the other way. Yeah, I had to. There were no manuals in Spanish at the time, or very few. And you were doing remote JavaScript contracting by the age of 11. Yeah. You dropped out of high school, moved to San Francisco to pursue the startup world. And it gets even more fun.
[04:34:44–04:35:14] So you needed an O-1 visa to get here. And so you wrote a book. The book is the most popular book, called Smashing Node.js, with Wiley in 2012. And I mean, you literally wrote a book to get the visa. Yeah. I did a lot of things in hard mode. Like dropping out of high school is really bad for getting a visa in the United States, it turns out. So some things, it helped me a lot. And some things, oh, it turns out people place a lot of value in getting a degree.
[04:35:14–04:35:43] So don't drop out of your degree. Good. Or maybe drop out. But there's pros and cons. Don't [INAUDIBLE] back. [LAUGHS] And so this high school dropout from Argentina is the guest for today's class. Please welcome Guillermo Rauch. Thank you. [APPLAUSE] Guillermo, thank you for joining us. I got your slides up here. I'd love to have you present them for us. Awesome. OK. Thanks for the intro. Thanks for having me.
[04:35:43–04:36:11] I'll try to make this as quick as possible so we can get into Q&A and a dynamic conversation. So for those of you who don't know about Vercel, we've built a vast ecosystem of tools, especially open source tools, that have now shaped a lot of the modern web and the modern internet. So even if you don't use Vercel directly, my most recent favorite example is
[04:36:11–04:36:37] if you order a Big Mac from McDonald's, you're using Vercel. If you're ordering a Porsche, you're using Vercel. Not that you're doing it today, but maybe after the SpaceX IPO. So anytime you're using the internet, you might be interacting directly with Vercel through pages. But also increasingly, if you use systems like OpenEvidence or Grok, you're interacting with agents hosted on Vercel.
[04:36:37–04:37:06] So we've created a whole set of tools and frameworks that what's I think relevant to today's conversation is because we bet on open source. And that came from my background as a guy in Argentina that wanted access to free tools and free information in teenage years. I wanted to make my technology accessible to as many people as possible. So open source was a very easy decision for me. But then we built a remarkable business
[04:37:06–04:37:36] on top of open source, which is also kind of contrarian. There's been examples. Databricks, I think, spoke to you guys recently. But we were able to build a formidable infrastructure business that basically sits on top of these tools and allows you to deploy secure, scale pages and agents, as I like to call it. So something that's super interesting about the journey of Vercel is that when I started the company,
[04:37:36–04:38:03] my obsession was developer experience. There were ways to deploy. There were AWS, Google Cloud, and Azure. But they were kind of a pain. [CHUCKLES] They made it really hard for me, even a seasoned engineer who had been studying the blade for decades. I sat down before I started this company to deploy the website of my new startup at the time. And it took me weeks.
[04:38:03–04:38:31] And I realized-- it was kind of like the inventor epiphany. It can't be this hard to deploy a freaking website using the latest and greatest technologies of the time, which were React, JS, and Kubernetes. And so I decided that group, JavaScript developers was a really large group that I decided to address directly. If you knew how to create a front-end project, so
[04:38:31–04:39:00] the user-facing side of an application, I wanted to give you the superpower to deploy. And not just deploy like an OK thing. I wanted to deploy something that could scale to the entire size of the planet. That was fast everywhere. And so from that little, I could say, set of people, we grew out into what today is Vercel. But what's changed since then is that AI-- by the way, the math that I used to do
[04:39:00–04:39:28] is there's maybe 20 million developers in the world that could basically fit what I thought was a pretty relaxed constraint. I don't know if you remember, but at the time, something that was really hard in the Silicon Valley was coding boot camps. Yeah, you could learn React in three months and get a job. People would tell you this. And it was true. Actually, the learning curve of modern JavaScript tooling was pretty low.
[04:39:28–04:39:58] And so it was pretty awesome that I could tell, hey, if you know how to create a compelling user experience, I will take care of all of the infrastructure for you. And I will scale. You will never go down. You will never have to worry about load balancing and low level infrastructure. And so Vercel will do that. But what happened since then that's been really exciting is AI encoding agents have created a massive expansion of the people that could create software.
[04:39:58–04:40:27] In fact, if we go even further back in time, developers were like a very, very tiny group of people that had access to the mainframe computers in universities. And then the story of programming and the story of technology is the story of expanding access to more and more people. So a lot of what we're going to be talking about today is that AI has created the biggest expansion in revolution in the total addressable market of software, and specifically the creation of software in the history of our times.
[04:40:27–04:40:55] And so concretely, what's happened to Vercel is that since October last year, especially with the arrival of a very specific coding model, which is Opus 4.5, we started to see that more and more and more people were deploying to our platform. We became the peanut butter and jelly of coding agents. The whole world is being plastered with peanut butter. That's the coding agent. And you need infrastructure to deploy the software.
[04:40:55–04:41:25] So one of the insights that I had when I started the company is that writing code doesn't make you special. Deploying code and putting it in front of a customer does make you special. The learning really begins when a user is confronted with the running version of your code. In fact, I was kind of perplexed that GitHub and before GitHub, SourceForge, and if you go further back, there was something called Freshmeat, people got really excited about storing software
[04:41:25–04:41:54] in the form of lines of code. And you could go to any GitHub repo. And it's actually kind of hard to run them. The world is made up of piles and piles and piles of software that doesn't run. And so my thing was that the people that will win the world are the ones that don't just write software, they deploy it and they run it. And so what's been exciting about the rise of coding agents is that coding agents don't succumb to this bias that we have sometimes as humans, which is that the code works on my machine and I keep it in my machine where I feel safe
[04:41:54–04:42:24] with the code in my machine. Coding agents love to deploy. And so we're this jelly layer of deployment of the software that the coding agents write. And so the other thing that this has taught us, because we're watching this world unfold in front of our eyes, is that the very essence of the cloud is changing. You were joking about [? Elon ?] Web Services. So I challenged that if Amazon started AWS today, I don't know if they would call it Amazon Web Services.
[04:42:24–04:42:52] Because the entity that people are interested in creating and shipping today is an agent. So they actually call it AAS, Amazon Agent Services. And so what we're doing today with the platform is we created a whole set of infrastructure primitives and frameworks for what's going to be the next entity or object that people want to deploy to the cloud, which is AI applications and agents. And so some examples of this are we
[04:42:52–04:43:21] used to talk a lot about the wisdom of Amazon, that when you go to amazon.com, which is a fantastic experience. Like I buy a ton of stuff there. One of the things that they unlocked is that if they load pages really fast, people buy more. What a concept. If you have really fast front ends, people buy more and convert at a higher rate. Not only did they observe that. Their data shows that for each 100 milliseconds of slowdown,
[04:43:21–04:43:50] they dropped conversion by 1%. So a lot of the internet was built on this idea of instantaneous request response. We're now seeing with our agentic cloud, this is changing. You don't always get a response immediately. So we need to create new infrastructure for this long thinking streams where the agent goes off into the world and does work for you. It started by taking multiple seconds, which is already
[04:43:50–04:44:17] pretty different from the kind of infrastructure we used to deploy. It then became minutes. It then became hours. And now we're seeing agents that could cook for an entire day before they get you your report, before they get you an analysis, before they build you software. We're hosting platforms on Vercel today that are creating companies for people. And they're deploying agents that maintain, advertise, grow, and scale these companies behind the scenes.
[04:44:17–04:44:46] And so there's an inversion of the old world of pages towards agents. And the kind of compute that we're running is changing quite dramatically. Another thing that's been fascinating is a lot of the cloud was built on this amazing compute product called EC2. EC2 popularized the concept of if you want a computer, all you need to do is put in your credit card and you get an instance of that computer. And if you need a million computers, you just need to pay for a million computers.
[04:44:46–04:45:16] It's Elastic Compute. So Elastic Compute was designed for human-written code. So you can think of it, from an economic perspective, how much compute I can sell to the world was bounded by how many programmers do we have. And so I mentioned earlier, the number of people that can create software is expanding by maybe like 10, 20, 100 times. But now agents are writing software. Are writing software, are repairing software, are securing software.
[04:45:16–04:45:41] We're seeing, obviously, the consequences of really smart models in cybersecurity. We have hacking agents. We have good agents. Like, bad cops and good cops out there in the world. And so we're seeing a massive demand for compute for agent-written code. And the other thing we're going to talk about a lot is that tokens are the new hot commodity. It used to be that Vercel only allowed you to stream pixels.
[04:45:41–04:46:09] So what you get is a UI or some kind of software application. Now we're sort of streaming intelligence in the form of tokens, which is also changing the pricing models and the business models. You've probably heard a lot of conversation around the death of SaaS and the death of the seat-based pricing model, and the rise of the token-based pricing model, which measures intelligence. And by the way, this is not to say that it's a transition where we're
[04:46:09–04:46:38] over-rotating on one side of the equation or the other. I'm actually super bullish still on human-centric experiences. So you will go and visit brand experiences. You will have to see the results of what these agents cook in some kind of rich environment. In fact, I even have sort of predicted the return of a more whimsical web and more whimsical internet. When I grew up, there was this really cool software that Microsoft created called Encarta.
[04:46:38–04:47:07] It was an encyclopedia that was alive. So you would select a topic. And it was like this interactive, rich experience. So if you were to study the pyramids of Giza, it was like this super cool experience. Now that we've made the cloud and all this stuff, you go to Wikipedia and you get this wall of text. And maybe you're dignified with a photograph on the right-hand side. So I actually predict that we're going to get the most immersive, craziest pixel-based experiences through video generation,
[04:47:07–04:47:34] through 3D model generation. There's a couple of platforms on Vercel that are doing really cool just-in-time 3D rendering. So human-centric experiences will matter a lot. But our new conception of the world is that we're making this transition towards what we call agentic infrastructure. So agentic infrastructure has three sides of it in a beautiful triangle. So number one, agentic infrastructure is the infrastructure that you need to give to your coding agent.
[04:47:34–04:48:02] That peanut butter and jelly metaphor that I gave you. So if you're using Claude Code, if you're using Codex, if you're using v0, those software coding agents need to deploy somewhere. So that somewhere, we're calling agentic infrastructure. The other thing that's really important is you will be building your own agents. So just today, someone came to my office and pitched this amazing vision of what could be an AI native school of the future. And a lot of what this person said
[04:48:02–04:48:32] resonated strongly with me, because they're rethinking education from an agent point of view. It's still involving the human, the teacher, the school, the student. But what they're going to be shipping is an agent, not a set of connected web pages. And perhaps one of the most exciting things is this idea that it's going to be automated by agents. So I'll cover the first principle quickly. The fact that you need infra for coding agents. I know that you've asked about some of the numbers of the growth that we're seeing.
[04:48:32–04:49:00] One of the fascinating dynamics is when you deal with a coding agent, the coding agent has a preconception of the world. It has a world model sort of behind the scenes. And one of the things that's given us tremendous tailwinds is that models have learned from the vast knowledge that exists on the internet about Next.js, React, our open source software. In fact, this report that came out calls,
[04:49:00–04:49:29] for example, our UI engine, we call it shadcn, it has near monopoly status at 90.1%. And when you decide to deploy React or Next.js, Vercel also has near monopoly, which I love how they qualify that for regulators, because it says also it's 100%. But I guess it's a near monopoly. So anyways, the coding agents have made up their mind, to some extent, and this is not cemented by any means,
[04:49:29–04:49:58] that these are really good tools to use when you give it a certain task. So for example, if you're creating a SaaS app, it's very, very likely that it's going to use shadcn because we created infrastructure that decoding agents can use. And you might be asking the question, and I certainly have asked myself, why do agents even need to reuse software that already exists? There is an economy of-- there's an efficiency bias here. In theory, a coding agent could reinvent the world every time you ask it something.
[04:49:58–04:50:28] In order to produce an Apple, it has to invent the universe. It could write its own Linux kernel. It could write its own networking systems, it could write macOS. And so betting on infrastructure as open has actually helped a lot because now the coding agencies have a target to throw code on top. It's like a LEGO block. Exactly. And in fact, there's a really cool article by Mitchell Hashimoto, founder of HashiCorp, who's recently joined the Vercel board. The world that we're going into now is, he calls it the block economy. You need to be thinking about what
[04:50:28–04:50:55] are the building blocks that you'll be able to give the agents. So if you want to participate in that, meaning, hey, how do I make it such that Claude Code, or Codex, or Grok CLI choose my technology, you want to be thinking about blocks that fit agentic ergonomics. So that's one. And the other one I was mentioning is, the software of the future will be agents. For example, at Vercel, we have created a support agent that now
[04:50:55–04:51:22] answers 93% of user inquiries. It's made our business massively efficient. Its improved customer experience. So we measure if customers are happy talking to an agent versus talking to a human. 93% has also allowed us to give free support to a much larger number of people. And so that thing that we shipped using Vercel was an agent, not a knowledge-based web of hyperlinks.
[04:51:22–04:51:50] So we basically created the tools, the blocks, and the infrastructure to build agents. So I'm not going to get into the technical details, but I'll tell you something that's fascinating is there are a lot of metaphors with the old web. For example, whenever you go to production with a web system, you typically have a CDN in front of your system. The CDN market that companies like Akamai and Fastly,
[04:51:50–04:52:16] they emerged to scale, accelerate, and secure the delivery of pages and pixels. We're seeing the same happen with tokens. We created the AI gateway, which is a category defining product. It's like a CDN for tokens. Tokens that you get from Claude, from Anthropic, from Gemini also need to be observed. They need failover. They need to be secured.
[04:52:16–04:52:44] And they need to be accelerated and cached in many cases. And even what's happening now that's super exciting, we can load balance them. You might have heard a meme of how expensive it is to say to GPT5.5, you say thanks. And you just activated like 300 GPUs to say you're welcome. And so imagine a smart CDN that says, hmm, if they ask me, thanks, I have a smaller model behind the scenes that can just say you're welcome. [CHUCKLES] Right?
[04:52:44–04:53:14] Or a semantic cache. So there are a lot of metaphors from the infrastructure we had to build for the first chapter of the internet to this agentic internet. And I love the CDN example. Sandbox I already spoke about, which is like the EC2 or the fundamental compute unit of agents. To give you a very non-- or basically technical explanation is agents get more powerful or you extract more IQ points
[04:53:14–04:53:43] from a model if you give it a computer. So models are being trained to use computers. In fact, in the post-training phase, models are handed Docker containers with playgrounds. So models cut their teeth in mini computers before they see the world. And so when you use a model, you get better performance out of the model by giving it a computer. And this is kind of not unlike humans. The average knowledge worker that we
[04:53:43–04:54:11] hire, what is the first thing you do when you join a modern company? We give that person a computer. IT, here's a laptop. It came pre-installed with your software, et cetera. That's the same thing we're doing with agents. We'd say, agent, here's your computer. Knock yourself out. And IT also pre-provisioned some software for you to be super effective. Now, much like the personal computing revolution, it came with viruses. It came with phishing.
[04:54:11–04:54:40] It came with Nigerian princes promising fortunes that you can claim with one click. And so agents can also fall for the trap of exfiltrating your data, leaking your data. So we're creating a whole set of security products to protect these sandboxes and computers. So you're going to see the emergence of, again, entire categories of products and companies that are tailoring cybersecurity for agents in many other matters. And the last and perhaps one of the things
[04:54:40–04:55:09] that I'm most excited about since we were talking about Tesla, full self-driving, and self-driving cars, and Waymo, the cloud itself will become like a self-driving car. If you have some experience running software at scale, you know that it's plagued with things like holding pagers. It's a horrifying experience. So I have an anecdote that I reuse from the founder of Stripe. He says that anytime he hears ducks, his cortisol levels spike. Why?
[04:55:09–04:55:39] Because his pager duty ringtone used to be ducks. So it meant that Stripe was crashing. And so anytime you hear ducks, he's on edge. And so the average experience of maintaining software and scaling it in the cloud is actually pretty dramatic. You get paged in the middle of the night, a data center went down. And Vercel, to a great extent, automated away a lot of that pain. But pain still exists. I'm not going to sell you things that are not real. You have to sit down and monitor things.
[04:55:39–04:56:06] You have to make sure that performance doesn't degrade over time. And so this is my vision of this self-driving car of the cloud, where everything configures itself, everything optimizes itself. We're all going to live a great life because the agent will be doing all of that work behind the scenes. And it's going to come to you and say, I just optimized all your software. I made it twice as fast. Here's the PR. Or maybe I already shipped it, and I measured that it improved conversion for your customers.
[04:56:06–04:56:35] So this is the third leg of this agentic infrastructure thing. So some really quick customer examples. Meta had already built infrastructure for years and years and years. Many people that have studied in these halls have gone on to work at Meta and create incredible developer infrastructure. But they found that Vercel's infrastructure had this edge. They need to move really fast. Their engineers are using coding agents. And that's how Vercel got in, and now runs
[04:56:35–04:57:05] a lot of Meta Superintelligence Labs. In fact, we've helped Meta move faster by not having to procure so much software. They basically, quote-unquote, "vibe coded" tools that have allowed them to train models faster. They shipped Meta.AI much faster than it would otherwise have thanks to agentic infrastructure. Notion is also investing heavily in their chapter 2. Fantastic tool. I use it every day. But Notion itself is becoming agentic,
[04:57:05–04:57:33] which is that prediction that I mentioned. Like all software of the future will be an agent. So Notion is also using Vercel infrastructure to add this AI and agentic capability. So for example, right now, we could have a Notion transcribing agent that is taking notes from this class and that infrastructure is running on Vercel. So long story short, we're basically building what you could think of as the AWS of AI or agents. And it's a full stack cloud from developer tools
[04:57:33–04:58:03] all the way to infrastructure. And now we can get to Q&A. Awesome. Thank you, G. The-- [APPLAUSE] Software is dying. Vercel feels like the machete that is chopping through the forest. This is what people are building vibe coding on. And you must hear about all these things. People are vibe coding all the software that people have stopped procuring.
[04:58:03–04:58:32] Are there some examples of, let's say, public software company products that have been vibe coated on v0 or Vercel that customers have replaced their public software? The bigger, the better. Yeah. One of the things that I always think about is the best software will always be the one that's most tailored to you. The weird thing about SaaS is that on one hand, it took a lot of really, really intelligent people, product
[04:58:32–04:59:00] managers, designers to get themselves in a room and say, what is the common UI that will make the biggest number of customers happy? And then we'll give it to the world. And then they hope that they get this exponential adoption. Because I show you, for example, an expensing tool. And you say, yeah, I can fit my business to that expensing tool. And then I give it to another customer and say, oh, yeah, maybe you only have to translate the labels to my language,
[04:59:00–04:59:29] but I can also sell it to that person. But that's still the lowest common denominator software. And it requires this intense design process of what is the set of pixels that are going to make the highest number of people happy. And I think what's happening now, I wouldn't call it the death of software. I would say if you use Vercel, software has never moved faster and been more alive. We've seen examples where I heard this recently
[04:59:29–04:59:59] from a startup customer. The lead engineer told me, man, your platform is dangerous. My CEO is shipping software. He replaced our parking lot management software with something he vibe coded in v0. And he saved us a ton of money. And so you have that example of, yeah, I mean, who-- I mean, how many smart Palo Alto-based companies are creating a parking software? Not many. And so there's a whole category of software that we've been living in a world that's just not ideal.
[04:59:59–05:00:28] It's just not high quality software. And now you're one or two prompts away from being able to create a really good version tailored to your problem set. Another thing that we see is the emergence of, all right, the system of record or the database stays, but I can create a presentation layer that is highly tuned for my business needs. So something that might be surprising for people to hear is that within Vercel, we kind of reinvented all of Salesforce, all of it.
[05:00:28–05:00:58] All of our sales reps, when they need to learn about an account, an opportunity, when they need to read business intelligence about how should I pitch, all of this is being generated. And it took a team of two people to create a new version of Salesforce, which is kind of crazy to think about. How big is Salesforce like stock market wise? It's hundreds of billions-- tens of billions of dollars. And now to think, oh, I can just generate
[05:00:58–05:01:25] a version of that that's like tailored to my business. But by the way, we're still using a lot of workflows and databases that Salesforce set up. So it's not an and/or situation. It's you can incrementally make software a lot more plastic, malleable, and tuned to your needs. So on one level, I think people are underestimating just how
[05:01:25–05:01:53] many off-the-shelf applications will be thrown away and replaced with this incredible velocity that software generation is giving you. But on other levels, I think you're overestimating that creating the underlying system of record, the ACL system, the Access Control Layer, all of those things you can reuse. So the companies that expose themselves with the right agentic interfaces,
[05:01:53–05:02:22] like MCP, like CLIs, and other APIs, those are the SaaS companies that will fit really well into this world and will not be, I guess, killed by the coding agent. Yeah, fascinating. Well, we'll have to get a session of advice for software from the creator of Vercel. Your business, you had a nice chart there that has changed quite a bit since Opus 4.5 last quarter, Q4 '25.
[05:02:22–05:02:51] Maybe outside of the volume, are there other parts of your business that have changed as coding has gotten so good? Maybe retention is the biggest thing people are talking about with a lot of the vibe coding apps. It's like, hey, well, you're making such custom software. Do people still need it? Like, hey, it was a thing to have on a Saturday morning, but do you still need it on a Wednesday morning? Yeah, we're seeing all of it. We're drinking from the fire hose. So we've seen software that is useful for one customer call. And that's fascinating.
[05:02:51–05:03:20] We have a lot of customers that have purchased Vercel and v0 because they're sales engineers and they've accelerated pre-sales so dramatically. They can come into a customer conversation with a custom version of their software already built. Imagine the difference between I show you a slide deck of what my software company could do to a prospect or I show you living, breathing software. And so that is an example of throw away because there's
[05:03:20–05:03:48] no illusion of three calls later, we discover something else and we threw it all away. So in that sense, you're right. Software is dead or throw away. Or I like to say that software is basically now free. And so there is a-- but that actually drives a lot of engagement because of what Tobi Shopify has called the reflexivity of AI. Once you know that you're one prompt away from communicating to another human being with a high fidelity piece of software,
[05:03:48–05:04:17] prototype, example, demo, you will never forego that. I actually wrote an essay many years ago called It's Hard to Forego Efficiency. This is giving companies such a high degree of efficiency. I have engineers lamenting, oh, my god, I'll never be able to code the old way again. Like, my brain just doesn't work that way. But also, let's also be clear. So there's a lot of agentic engineering that's still really hard. A lot of the infrastructure software
[05:04:17–05:04:45] that we build, really hard. Sometimes we have to summon a quorum of three agents, plus a lot of smart humans, like, look at a single line of code and tell us what's true or not. And so there's a lot of software that continues to be extremely long lived. And so the really big difference has been the audience. I have customers reaching out for support, sliding in my DMs on X, saying, I ran into this error.
[05:04:45–05:05:14] My friend, I don't even know what Vercel is. My agent took me here. I deployed. I was really happy. But now here we are. And then the way that I helped them is by, hey, can you introduce me to your agent? I asked them, can you show me the transcript so that they can bring it back to our engineers. Maybe we can turn it into evals and we can understand how did my customer's agent get into a bad spot. It used to be a much more direct conversation. I would get introduced to the engineer.
[05:05:14–05:05:41] And the engineer would tell me, yeah, man, I went to the docs. You said this was the API. It was wrong. This is the error. We're going to end up in a world where it's just agent-to-agent communication. The customer's agent files a feature request or bug report with my agent. And my agent says, hmm, how do I prioritize this among all of the tasks that we have with this token budget that we have, with these deadlines and the bias of my CEO
[05:05:41–05:06:09] that wants me to ship this other thing, and software emerges. Yeah. I'm going to pull up a slide you had here, which was the peanut butter and jelly slide. I think that's a phenomenal framing. And the big achievement you have here is actually could you break this down for us? So 86 out of 86 times, and let's say Claude baked the Vercel deployment option and your UI component option set, shadcn.
[05:06:09–05:06:36] I mean, that seems like a slam dunk. How did you guys land up there with 90% market share on the UI component, 100% on deployment? Was this an enterprise deal with Claude? Was this so good that the agent was like, hey, I've got to do a meritocratic search real time to pick these guys? Well, there is that, fascinatingly enough. What is it? All of the above. All of the above, including the deal, or-- [CHUCKLES] Minus the deal.
[05:06:36–05:07:03] Although, I can't confirm or deny the existence of deals in the direction of expanding access to Vercel. But what I'll say is, on one hand, there's cause and effect. As I mentioned, for many, many years, we created huge amounts of content, and frankly, really high quality APIs that worked well for humans and agents alike. Like SEO stuff. No, no, even before SEO. So I'll give you an example.
[05:07:03–05:07:33] That thing, Tailwind, that is underneath Vercel, I bet really hard on that. I'm not going to say I was the earliest person to bet on it, but this technology was extremely controversial in the human developer ecosystem because it looked weird. I know. Has anyone used Tailwind before? Raise your hand. OK, so I'll explain it. Code has an aesthetic sense to it. You look at it and it can look symmetrical, pretty, nice,
[05:07:33–05:08:01] or it can look like a piece of junk. Tailwind kind of moved us in the direction of piece of junk. The lines got really long, which, for people with extreme OCD like me, I had to overcome biases and bet on truth. So what Tailwind gives you is a property called local reasoning. So when you design a component, and what I mean by a component is let's say that I'm designing this UI,
[05:08:01–05:08:30] this slide UI. It has a bunch of components. We can call this the slide preview component. Tailwind allows me to design it in a way that is extremely future-proof. I can reason about its design in a way where if I take that component and give it to you, you can even insert it in another part of Google and the component works perfectly. So it created an economic scalability to the code. And I bet hard on that.
[05:08:30–05:08:57] And I overcame my own sort of gag reflex of how code looked. By the way, if the Tailwind guy listens to this, like he knows. He's heard this feedback before. So that's, I think, when we started moving from this human-centric code design to what actually scales better for humanity, and organizations, and economies design. And so a lot of the technologies that I've designed, like Next.js, has this local reasoning property.
[05:08:57–05:09:26] React has it in spades. The Facebook team arrived to the same conclusion. As they were hiring more and more and more engineers, they needed systems of scale to the code. And so there is a lot of things that we've put into the design of the APIs that agents have just eaten up. For reasons, for example, the context window. You can't fit all of the code of humanity into the context window of an LLM.
[05:09:26–05:09:53] Today we have 1 million. And so this local reasoning property of the code became really, really important. There is a content for training and for grounding. So when a coding agent kicks in, it Googles. There is the kind of thing that these agents create. I mentioned that you can create your own Salesforce on Vercel by still using Salesforce as an API.
[05:09:53–05:10:22] You can use Salesforce headless and Vercel to host the application side. And so this technology set is also fitting really nicely into that world. So there's basically a confluence of factors that have gotten us here. Very composable. Yeah, composability being-- this is a great summary. Composability is a key prerequisite for this agentic scalability. Fascinating. Another slide that caught my fancy was how much you're doing.
[05:10:22–05:10:51] So this is a lot of things. There's independent companies whose full-time jobs is to do one of these things, from Sandbox, to chat, to the workflow. And this would imply that you were competing with a series of companies across the stack, down from maybe not quite bare metal, but just above bare metal, to all the way to the UI elements and obviously, deployment. How did you decide to "do it all?"
[05:10:51–05:11:21] "Do it all" in air quotes. [CHUCKLES] What parts of it would you give yourself a grade of like A-plus and maybe less than A-plus? And where is the most competition you're facing? Yeah. I only brought you A-plus products in this slide, I think. [CHUCKLES] I think there's a lot of things that we frankly throw away or deprioritize over time, I think even, hopefully before they see people's eyes. Yeah. But a bigger design choice to do the whole thing. This is kind of the Steve Jobs, you've got to do the whole thing from the back of the-- I really think so.
[05:11:21–05:11:47] You have to do the whole thing for the thing that matters. Agents matter. Agents are probably the last class of software. And I'm building the tools and infrastructure services to build agents. We can't not participate in the most exciting, fastest-growing economy in the world. And there's a couple of things. One is the Vercel product development philosophy. The Vercel product development philosophy
[05:11:47–05:12:14] is we build by dogfooding our own platform. So AI gateway is this CDN of tokens. It's built on Vercel. In fact, we open source the recipe so that you can build your own competing token gateway. [? An ?] open router or something? Yeah, you could build another one. And there's companies that have built different ones that make different trade-offs or differently integrated into their services.
[05:12:14–05:12:43] So it uses Fluid Compute, which is our compute platform. It uses the global vast network. It basically reuses a lot of our CDN. So to your question about why would I go into that business. Well, I was able to reuse 95% of the rocket engine. It's the same rocket engine and same rocket fuel that was powering the CDN of pixels or pages. Sandbox, perfect example.
[05:12:43–05:13:11] So I told you that the Vercel platform has seen this insane growth in popularity in a very short amount of time. And what typically happens when that happens to startups or even scale-ups is that everything breaks. I'm happy to report that almost nothing broke in Vercel. Obviously, no service is perfect. But our Sandbox reuses the same virtualization primitive that powers every deployment made on the platform.
[05:13:11–05:13:39] And so we're able to leverage our operational expertise in scaling compute, which I cited another metric. I mean, with 3x'ed in a few short months. We even doubled the number of daily deployments since January. So we're producing all of these ephemeral computers that then would destroy every time a deploy happens. And so Sandbox basically reused all of that compute expertise
[05:13:39–05:14:08] and virtualization expertise. And so it was a very obvious bet for us to make as well. Fascinating. The question I'm about to ask you, I'm actually very excited to ask you the question because of how much you've done underneath the product and obviously, beautiful products being built. One of the biggest questions that we ask in this class is, where will value accrue? From chips, to data centers, to infrastructure, even below the model, the model, and the agents.
[05:14:08–05:14:37] And frankly, it feels like at least in 2026, right this second, there's a lot more value accruing to the stuff below the model, chips, data centers, power, cooling energy. Above the model, you've got some value in the coding models. You've got some value in customer support, and legal, and others. But it's very concentrated. Talk about that for a second. Where do you see value accruing based on everything
[05:14:37–05:15:07] you know that's being built on Vercel? Yeah. My sense is that you've seen the figures of how much of the token economy is going to coding agents. Coding agents seem to be one of the most promising paths to AGI. Or if they might just become the path to AGI. And so what's great for us is, is that we built the infrastructure that, again, coding agents need in order to actually do useful things. And I think there's going to be a lot of opportunities like that.
[05:15:07–05:15:35] At the end of the day, what determines whether we're in a gigantic bubble or in the most exciting chapter of humanity is whether we're delivering useful services and products. And so it turns out that for the model to be useful, it needs a sandbox. It needs a deployment platform. It needs a domain name. Actually, I was hearing from the Codex team the other day that one of the reasons that-- this goes into this world that you were sort of alluding to
[05:15:35–05:16:04] of how do I understand what agents are doing. Or it's like agent engine optimization. The Codex team was telling us that a lot of people want to name their creations. They want to give it a domain name. And so they're actually finding Vercel first through a domain name. I made an early bet on making it extremely fast and easy to purchase and configure because DNS is hell. And my intuition was when a human has an idea--
[05:16:04–05:16:33] we're in a very entrepreneurial room here-- a lot of us go and buy the domain name first. I have so many domain names that I don't even know what to do with them. Sometimes upon renewal, I feel bad. And so I bet that-- So what's your annual budget on that? It's embarrassing. It's a lot. But I tell myself it's digital real estate so it's worth it. And so we made that bet of like, oh, when you have an idea, you're going to need this thing. You're going to need DNS infrastructure. You need domain names, et cetera.
[05:16:33–05:16:58] It turns out to have returned very handsomely, because whenever a new user of this software world has an idea, they're coming to Vercel. Our obsession is can we become the front door to any emerging idea on the planet. I think every business that is going to help these ideas come into fruition and be useful is going to do really well. I think security is going to be a massive consideration of this world.
[05:16:58–05:17:27] I need to stay focused, and we have a very promising product lineup, but I mean, there's just so much left to build. And a lot of what I'm thinking about these days is not only how to make these agents useful, but how to govern them, how to give them guardrails, how to make them secure, how to secure the broader internet. So a lot of our mission, I mean, we just can't stop hiring engineers. I think there's a lot talk about, is it over for engineers? And I can tell you, there's just so much more to build,
[05:17:27–05:17:56] to serve this amazing demand that's happening just from the coding agent use case itself. Fascinating. Rapid fire, last couple of minutes. Pick a business that you are long, that you think you're very optimistic on, and a business that you're short on, that you think based on everything you know, it might not make it. Short on, I predicted the downfall of-- I mean, this is an obvious one now,
[05:17:56–05:18:24] but I predicted it pretty early. Like, any static data or content business is kind of cooked. The classic example at the time was like Stack Overflow, which was this aggregator of-- this database of programming questions and answers. I'll tell you a category, because I don't want to throw companies under the bus. A category is anyone that's thought that code was scarce or difficult to produce.
[05:18:24–05:18:53] There's a whole category of drag and drop builders. And basically it's like coding with training wheels. I always despised that category because I never liked to be looked down on. I always found that kind of software patronizing. Yeah, too opinionated. Opinionated, constraining. I'm like a little kid, and I need to be given this little interface, and code is scary. And so there's a ton of companies that were built on that foundation.
[05:18:53–05:19:21] And they're going to have to have some serious pivoting conversations and whatnot. Also, the companies that are not opening up. One amazing consequence of coding agents is that they want to access the raw signal. They want to go straight to the data. They don't want to talk to sales, and let's chat about your requirements for three months, and what are your pain points. No, they just want, go.
[05:19:21–05:19:50] Consumption-based pricing, instant sign up, start getting tokens right into my bloodstream, that kind of thing is going to win. But a lot of the internet economy is predicated on, oh, we don't really have a product. We have this e-brochure. And you have to talk to the enterprise rep for three years. And then maybe some software will emerge. [CHUCKLES] And you'll forget how much of the software ecosystem
[05:19:50–05:20:18] is still that. And so I'm very long on anyone that's moving at the speed of tokens. And so we're doing so much work just to meet the demand of-- so one of the write-ups that I made internally at the company recently is, as an engineer, you're taught to apply rate limiters. Rate limiters are very useful for the operational health of your system.
[05:20:18–05:20:46] So a good example would be, and this goes back to that hypothetical Palo Alto room where smart people get together and say like, hmm, no company will deploy more than 100 times per minute. So OK, rate limit, put it into the code, 100 times per minute. Totally arbitrary. Nowadays, no one can really know what demand or write quota is. You do need to be super worried about abuse, and KYC,
[05:20:46–05:21:15] and things like that because I'm backed by supercomputers. So I can't have anyone immediately send them a bill for yeah, you just spend a trillion dollars of supercomputing power in 2 seconds. Can you pay me, please? But I told the company, no more rate limits, guys. As a provocation, of course, again, for operational health reasons, you do need to have some things in there. And that goes hand-in-hand with this pricing model of you pay for what you use. And so why would I rate limit your ability
[05:21:15–05:21:43] to use more resources? And so right now, and this is a good problem to have, I have this massive agent platforms where a YC company that didn't exist three months ago is producing so many deployments on the Vercel platform that they're surprising us about, oh, I didn't remember we had that rate limit. We could have never imagined that there was going to be so much demand for a certain piece of infrastructure. So bullish on the companies that can serve that kind of demand. Final question before we wrap.
[05:21:43–05:22:13] If you were not building Vercel right now, what would you be building? Asking for a friend in this class. Space tech, obviously. Like, can we get into Mars as soon as possible, please? Yeah. Why the rush go to Mars? I mean, my bias has always I love the multi-planetary species aspiration. I've always thought about it as I'm a high availability, multi-Z, multi-region, multiple layers of failover guy. So I chose my location in San Francisco
[05:22:13–05:22:39] based on the highest structural safety and the best and most historic retaining wall that has never collapsed in any earthquake, whatever. I'm an infrastructure guy so I have to think about these things. But also, I'm really excited about energy. And honestly, what we're witnessing is that the emergence of intelligence is this bidirectional flow of energy. Energy goes in, intelligence goes out.
[05:22:39–05:22:56] So any breakthrough that we can have in fission, fusion, geothermal, what have you is really, deeply exciting to me. Awesome. Well, that's a wrap. Thank you so much for your time, G. Thank you. Appreciate it. [APPLAUSE]
[05:23:08–05:23:37] Inference is about to go up a billion x, not 1,000 a, not a million x. A billion x. And the guy who's making it go up, billion x is here with us today. Tuhin, thank you so much for joining us. Please join us. Thank you. [APPLAUSE] Tuhin, founder and CEO of Baseten, you've had a long, long scenic journey to the day here today,
[05:23:37–05:24:06] lots of windy turns. Tell us about it. A lot of students in the class who want to be entrepreneurs, I think, will find inspiration in the journey. Yeah, awesome. Thank you for having me. It's really nice to meet you all. Yeah, my name is Tuhin. I'm the CEO and one of the co-founders of Baseten. Baseten is about seven years old. I've been working in technology since 2012. I'm originally from Sydney, Australia. I came here for uni.
[05:24:06–05:24:32] Actually started my career in finance. So I moved home after uni to work at an Australian investment bank called Macquarie. In my first day over there, they said you want to go and move to New York and work on privatizing toll roads and airports? So I moved to New York and I worked in privatizing toll roads and airports. And about two years in, I was bored out of my mind,
[05:24:32–05:25:02] and I decided that I should go back to engineering, which I studied in undergrad. And I moved to Boston to work on doing some research for machine learning. This is 2012 or 2013. No, 2011, 2012. And what we were doing is we're using traditional machine learning techniques to track the prognosis of neuromuscular disease. I did that for about a year and a half, put out some papers and really quickly realized that if I was going to do machine
[05:25:02–05:25:31] learning even back then, that I probably should be in San Francisco. So I moved to San Francisco and got involved in early stage technology as an engineer. I think probably the more important piece is I fell in love with early stage technology and small teams and building products, and where no one really cares. Did that for a few years. Started a couple of companies starting 2015. They went nowhere. But I had the bug.
[05:25:31–05:25:59] And so in 2019, me and my two co-founders-- me and my two co-founders, a guy called Phil and a guy called Amir, who I'd worked with for the better part of 10 to 15 years-- is that a me thing? Oh, they're just turning up your volume. OK, cool. Give you some more energy. Yeah. So I know these guys for 10 or 15 years.
[05:25:59–05:26:28] And we had this idea that machine learning was going to be pretty big. And let's build an infrastructure business alongside it to maybe capture the index of it. Obviously, machine learning was a lot bigger than we expected and it came a lot faster. And so what we've been doing for the last four years is production inference, powering the fastest growing AI companies in the world. Yeah. You have a salivating list of customers. Actually, you had a post. I copied your post here, which is these
[05:26:28–05:26:57] are all the founders that Baseten has a good opportunity to work with. You might know some of these faces. Some of them have been here in class. This is probably the best index of AI founders to work with Baseten. And so maybe break it down for us. Pick an example or two of faces up here. When they work with Baseten, what does that mean that they work with Baseten? What's an example of they work with us,
[05:26:57–05:27:26] and what does that enable for their business and their customers? Yeah, maybe I'll pick two fun ones. I'll pick Tanay, who's on top of a pick. Tanay, who actually went to Stanford Tanay very well. And then I'll pick Shiv. And so they run very different businesses. So Tanay runs a company called Wispr Flow. I don't how many people here Wispr flow. How many have heard of Wispr Flow? Yeah, awesome, awesome.
[05:27:26–05:27:54] That'll make him very happy. So Wispr flow is basically a speech to keyboard-- speech to text app for voice typing more or less. They run a lot of custom models that they've built in-house or modified in-house in a very particular way to make that experience possible. And so what we do for them is we run all the optimizations,
[05:27:54–05:28:23] we've run all the infrastructure so that latency from the time that you talk to when text shows up is as quick as possible. There's language models. There's audio models. There's actually, I think, three or four language models in the middle there and two audio models to make that happen. And all of them run on Baseten. Shiv, who's also a good friend of mine, is the CEO of a company called Abridge.
[05:28:23–05:28:51] Abridge is a healthcare-- it's an ambient scribe used in almost every healthcare system in the US. That's deeply integrated with EMRs. They run about 20 different models, everything from, obviously, speech-to-text models to go from what is happening in the operating room or the patient's room to everything that
[05:28:51–05:29:17] goes into turning that into a clinical note that is deeply integrated with the EMR. Again, dozens of models, lots of requirements from a reliability perspective, from a speed perspective. Actually, I think every single one of those models runs on Baseten. And I think that's kind of emblematic of what we want to do.
[05:29:17–05:29:46] Our thesis is that AI is going to be absolutely massive. Inference is the cogs of AI value being delivered. And today, about 90% or 95% of spend on inference is going to frontier models, and about 5% is going to custom models. And we believe that the way that all these folks who are building amazing applications with AI are going to build profitable,
[05:29:46–05:30:13] viable, defensible companies is with custom models, and hopefully, they all run on Baseten. 95% of their spend is going to frontier models. Yeah. 5% is going to open source models-- Or post-trained open source models, yeah. Got it. Actually, before we go there, I have a question for the two businesses that you highlighted. So Abridge or Wispr Flow, both fairly scaled businesses. Why would they come work with you?
[05:30:13–05:30:42] Typically, if I was a founder, I would go to my cloud provider. I would go to AWS, GCP, Azure, or one of the neo clouds that are the AI clouds, CoreWeave, Nebius, others. Why Baseten? Yeah, I think a lot of it comes down to the three or four core tenets of the software we're building. So there's obviously performance. You need these things to work as quickly as possible. If you go to CoreWeave, Nebius, AWS, GCP, you're on your own
[05:30:42–05:31:10] to do those optimizations. The second piece is reliability and compute. They're coming to us because they need inference to run across clouds in a very fault tolerant, resilient way. I think part of that is being multi-cloud and having access to compute from multiple sources. We unlock multi-cloud for them. And the third one is the developer platform that we provide, which gives them flexibility, security,
[05:31:10–05:31:40] observability and all those good things to make that happen. And so what we actually find is a lot of them do go to those folks you talked about first, and then realize the pain of standing up that whole inference stack on top of compute, and realize that there would just be better served coming to Baseten. Fascinating. I imagine, now going to the second thing you had said earlier, 95% of them, or 95% of tokens are on frontier, 5% on both trained or open source of some kind.
[05:31:40–05:32:05] Yeah. If I was using an open source model or post-training an open source model, I am branching off the lineage of intelligence that might improve. And I run the risk that the next GPT model or the next opus model, let's call it GPT 6 or Opus 5, will be better than what I just posted. Why go through the effort of branching out
[05:32:05–05:32:34] into this specific model when the no-work option might just get there in a couple of weeks or a couple of months? Yeah, look, there's a number of reasons, I'd say. I'll give you the viable reason and I'll give you the cynical reason? The viability reason here is that if you think about a company, let's take-- if you think about any truly any company that's doing things
[05:32:34–05:33:03] with these frontier models, today, open source models are about 90 days behind, close to those models. And the question that a lot-- I think we have a slide on this. Keep going. Yeah, do I? Next one. Yeah, open source models are about 90 days behind frontier models. And you can run them about 70% to 90% cheaper than frontier models 70% to 90% cheaper. Yeah, and especially when you-- and when you add in a bit of--
[05:33:03–05:33:29] I think you had Josh in here last week who probably told you the post-training narrative of specialized models and especially when you factor that in, you can really do better, faster, cheaper for these models. And really, they come to us from the viability reason when they really go from I have product market fit I need to become a scaled business and figure out how get my gross margins up from 0
[05:33:29–05:33:56] all the way to 40%, 50%, 60%, 70% hopefully, want to be. The cynical reason I'd say, which is look, what do you have that is defensible against the frontier model. It's probably some workflow, some user signal that only you have. If you keep working with the frontier labs, to some extent, you're going to give them all the data and all that user signal.
[05:34:00–05:34:27] I was talking to a prominent public company CEO. I won't say who it was. He likened the frontier labs to the East India Company. East India Company, they show up in India and they're making all these partnerships. But really, what they're doing is you're giving them all the tricks on how to rule that. And before you know it, they're post trading models kind of against those workflows that are sacred to you, that only you know.
[05:34:27–05:34:57] So to be defensible against this and keep the thing that makes you special need, you need to own your intelligence to some extent. And that's why they are running-- they are figuring out how they can use open source models and post train them and really build up that stack themselves to own their intelligence. Fascinating. Taking this further Tuhin, it seems that if I had a larger volume of workloads,
[05:34:57–05:35:27] I am more incentivized to move to open source or post trained model. You have a couple of customers that are very scaled customers, Abridge, Wispr Flow, Cursor, et cetera. They are obviously very large customers of the frontier models. Is it also the case that the larger or the more user base you have, the more likely you are to adopt this? Or is it the other way, is the smaller guys
[05:35:27–05:35:55] are more likely to adopt this? I think the former. I think the larger you are, the more existential it becomes to be able to need to go towards this. Because I think that's when-- the bigger you get, the more-- I mean, maybe you could talk about this as well, but you look at these businesses, and as you get to massive scale, if you're just trading tokens, it starts becoming very, very expensive,
[05:35:55–05:36:24] and the more capital you need. And I think you need that path to profitability, post-training becomes very important there. Yeah. the leading coding companies that are not the frontier model companies themselves are still rumored to be negative gross margin. And so I imagine it is existential for them to be a viable business to shift that token-- the token volume towards open source.
[05:36:24–05:36:53] Yeah. Now is there a performance gap that they notice as they move from, let's say, for example, Cursor using Anthropic's models or GPT, the latest OpenAI models, to moving to a post-trained model like they did recently? There was a lot of hoopla on Twitter about their work with Composer. Does the performance take a hit? Is there a way that the. Workload partitions that it does not? Or how do the users experience the change
[05:36:53–05:37:21] in the underlying almost the engine of the Ferrari was swapped out? Yeah, well, we don't have any data internally but I think would-- you hope that it gets better, right? You would hope that as they post trained models to user signals that they have the experience of the user gets better. And hopefully it also drives both performance from a latency perspective and reliability perspective higher because there's a lot more control.
[05:37:21–05:37:48] Got it. Makes a ton of sense. And what is the business of Baseten? How do we monetize our users? Obviously, there's a whole slew of pricing options right now in AI. You've got outcome-based pricing, token-based pricing, GPU by the hour, pricing, and others that I might not even about. Where do we fit in, and how does that work with your largest customers?
[05:37:48–05:38:16] Yeah, look, now we mark up compute. We're relevant for the most part. So what that means is that come to use Baseten. You bring your models. You load it up onto the Baseten inference stack. And you basically you basically choose which compute you want to run it on. And a Baseten H100 or Baseten B200 is more expensive than a raw A100.
[05:38:16–05:38:44] I think what's interesting there to think about, how much-- it is an expensive way to get raw compute, but ideally, that's the value you're getting the software stack. Others use token pricing. We have some customers and we are pretty relatively unsophisticated about this today. But we're moving towards a world, especially as we unlock more of those post-training training
[05:38:44–05:39:11] workflows within Baseten itself where the narrative changes from how much are you paying for compute to how much are you paying per token? And actually, the outlier companies appreciate that a bit more because especially for those coming from using Anthropic, OpenAI, they have a like for like in how much cost savings there are. Fascinating. I might follow up on something you said. You said post-training workflows. Yeah.
[05:39:11–05:39:40] Could you simplify that for us? What does a post-training workflow look like start to end? And feel free to take an example, if you can. How does it work? How does it work with Baseten? Or how does it work with a large customer like Cursor or Abridge or Wispr Flow? Yeah, we talk a bit about post-training there. So we have training. And so really what you're doing as a customer there
[05:39:40–05:40:09] is you're deciding what your utility function is and you're saying, hey, this is the thing that I want to optimize with this given model. So you need to define that. We can't help you there. We don't know your product. We don't know your users. We don't know what you're optimizing for in your business. So instead, you decide what the utility function is. And then once you know that, start giving us a bunch of data and choose an open source model, and we're
[05:40:09–05:40:37] going to give you all the scaffolding to turn that into a pre-trained model. And then once you have that, how does that roll into the inference piece? And so we own that entire loop. So you bring your data, you bring your utility function. You choose a base model. So you might say-- maybe just to make it very concrete, you might say-- let's say you're building a speech-to-text model for medical use case
[05:40:37–05:41:06] You choose and maybe it is errors or transcription errors, the thing you're trying to minimize. So that's your utility function. You choose the model you want to use. So let's say you could use a [INAUDIBLE] and then you give us a data set. We have all that scaffolding set up to turn that into a very, very good post-trained [INAUDIBLE] made for that. And then we obviously have all the other integration into inference predefined.
[05:41:06–05:41:35] And so our customers who are coming to us with that, and I think you're going to see more and more of this, really they come with data and what they about their workflows. And they leave with a post-trained specialized model running on Baseten. They trust you with all that data, the keys to the kingdom. Yeah. Is that a natural conversation for them to hand over the keys to the kingdom
[05:41:35–05:42:03] to Baseten or the East India Company? You know-- We're the West India Company. We're good. We are the rebellion. We're trying to arm the rebellion. So is that a natural conversation? I think it's fine. I think most people are OK with it. I think we have a track record of working with amazing companies. And the brand kind of does a lot of work there. But we also have incredibly intense security posture
[05:42:03–05:42:32] internally. We work with multiple competitors and we have access to their models and their data. And we have set up the boundaries internally to make sure that there's no leakage between that. And so is it natural? I think it's not as bad as you'd think. And you also have to remember that a lot of these companies are trying to work, move very, very, very quickly. And I think they're more interested on, can you
[05:42:32–05:42:59] solve the user problem, as opposed we need to do everything from scratch. I think there's a lot of trust they put in Baseten. But I think hopefully, that's earned from all the great customers we work with and their friends who are already using us for a lot of these things. Makes sense. The other big bet you guys are making implicitly is that open source will stay on the frontier, or at least three to six months behind the frontier.
[05:42:59–05:43:28] The economic model of open source frontier labs is TBD-- TBD. I was about to say TBD, yeah. --is not known. And there's a lot of at least Western Open source labs. Obviously meta has decided to insource their-- or not open source their latest models anymore. Talk about that for a second. What is the business model of open source? What is your best guess of how the world unfolds?
[05:43:28–05:43:57] Because that is obviously one of the biggest underlying bets that you've made. Yeah. Look, I think the two big bets of Baseten are the existence of an application, an independent application layer and more open source be good enough that you can continue your post-frame. Most of America so far has shown that it is not able to produce the best in class, at least in the last two years, open source models. The best open source models today are coming-- That's wild. --from China. I think--
[05:43:57–05:44:26] Why is that? Why can't they produce them? Why can't America produce open source models that are performant? I think there's all the best research in the world right now work at two companies. And I don't think they're despite their names maybe suggesting otherwise. They're not super motivated to produce that. I think-- And why have the non-American researchers decided to pursue the open source model?
[05:44:26–05:44:54] I think it's to be relevant probably to start with. We think Moonshot and Alibaba and minimax, they're just a market opportunity, if all to take the counter position of closed models. And Meta did that too, a few years ago.
[05:44:54–05:45:18] Then they famously have swung the other way. Yeah. Look, open source models need to exist. I think it's somewhat of a matter of national security if America does not have good open source models or if the cost of intelligence is 70% or 90% cheaper in the East than it is the West. That's probably not a good outcome. Yeah.
[05:45:18–05:45:45] I do think the best American companies are actually pretty well aligned around open source. Google produced Gemma. NVIDIA is putting a lot of investment into the NEMO Tron that family reflection AI. Hopefully, we'll come out with a good open source model in a few years from now.
[05:45:45–05:46:13] I think it's inevitability though, realistically, if we're in a world in 10 years from now where there's no good open source models, that's probably a very bad thing for the US. But I do believe there's enough investment happening that it's going to happen. I'm sure you did-- Anthropic put out a big post this morning. I saw that. The future of the Chinese and American AI relationship
[05:46:13–05:46:42] of course, very beneficial for Anthropic. Any thoughts on that? Actually maybe summarize the post for the class. So how do you-- I haven't read it. I've been back to bed this morning, but my read on it was like there's two scenarios that America is at the frontier. And the other one is that China is neck and neck more or less. And I think again, maybe you've read it. That's it. That's exactly it.
[05:46:42–05:47:11] They basically said we either we lead it and you shut it down or we're neck and neck, and that's a war. Yeah, and the recommendation was let's shut it down. Let's shut it down so there's not a war. That's right. Look, they're awesome. And they do amazing work. I think I genuinely believe in life, before you make any points, you should make your biases and your incentive functions very clear. I don't think they acknowledge that. Yeah. Look, I don't know what to make of that.
[05:47:11–05:47:34] I just think we put out a post yesterday about the world of many models. We think intelligence shouldn't be owned by two people. I think in the history of time, when too much power is being concentrated in one or two parties, bad things generally happen. And I also just don't believe the thesis to some extent, that these two companies with a massive profit motive should be the arbiters of morality for the rest of us.
[05:47:39–05:48:08] So when I think about, do I think that it is concerning that all the best open source models are coming from China? 100%. And they're great. We've spent time with all those teams and they're amazing. And it's not so much that I don't trust them. I just think that America should have great open source as well. Yeah. Do you think the velocity of American open source is increasing or decreasing over time in terms of talent
[05:48:08–05:48:37] and resources, capital, GPUs? Yeah. I think up until like a year ago, it was definitely decreasing. OK. But I think there's uptick now. Look, I think it's a very, very important company to buy in this. And I think we'll see a lot this year going the other way. Yeah. I think it's necessary because otherwise we're just going to have two companies left. There'll be no Apple. Yeah, right, right right.
[05:48:37–05:49:05] Switching gears a little bit, Tuhin, you were in a very advantaged position where you see a lot of different hardware that you run inference on. Well, actually, you run two layers, a lot of different providers. Actually, I think you have this map which we can talk about for a second, but I actually meant a layer underneath this. I assume a lot of the inference that you guys are delivering is on NVIDIA hardware just because it's the majority prevalent.
[05:49:05–05:49:35] But then there's a long tail, obviously, Trainium announced in the last earnings call they had about $20 billion of revenue run rate. TPU I'm sure is a big number. We don't know what it is. They haven't disclosed it. At least then there's cerebrus from this morning. Yeah. Then there's everybody else-- Etched, Mattox, Positron, others that we do not know about, D matrix, Sambanova, et cetera. Talk about the heterogeneous compute ecosystem. How do you see the world unfolding on the hardware side?
[05:49:35–05:50:03] Yeah. Look, we are, I think, in general diversity. I can't sit here and say that I believe in a world of many models and then say, oh, only one chip will work says that. That'd be hypocritical. But that being said, we run the majority of our fleet on NVIDIA chips. TPUs are very promising. Obviously, all the new age, the neo chips,
[05:50:03–05:50:30] if you want to call them that, are very promising. We haven't seen anyone really use Trainium at scale. I think the advantage you have with NVIDIA is just obviously a fleshed out supply chain-- Yeah. --a fleshed out supply chain and with a very, very low cost of capital and a very strong relationship with TSMC. Yeah. I think that the idea that anyone else can really
[05:50:30–05:50:58] compete with that today is somewhat fanciful. To go one layer deeper, one thing we rely very, very heavily on is Cuda and the developer ecosystem around Cuda. There's nothing like Cuda. Cuda is insane. I think these new architectures are breaking out. Inference has two core parts to it, which is one is prefill
[05:50:58–05:51:25] and one is decode. This idea that all this will also just run on one chip, which is historically what's happened, probably isn't the end state. But NVIDIA acquired Grok. A lot of these new chips you're talking about are trying to separate out these concerns and decode. And where you do the memory bound stuff on GPUs
[05:51:25–05:51:55] and you do the compute stuff on a different chip. So the separating out. So I think that is the world will go like heterogeneous architectures for these things. I think right now, for us as a company, we're just trying to move as fast as possible, and our customers are trying to move as fast as possible. And Cuda, availability of NVIDIA GPUs, the ability to use stuff like TRTLM, which is clearly just built
[05:51:55–05:52:19] for NVIDIA chips. What is that? It's a runtime. It's open source runtime that runs on top of NVIDIA chipsets developed by NVIDIA. Even VLM and XG Lang, which are two other open source frameworks, are native to NVIDIA. And so the ability to just move very, very fast is the thing we are optimizing for. And that's what NVIDIA is really good at today.
[05:52:19–05:52:48] Gotcha OK, so a fairly one-sided play there for now with hopes for optimistic future. Maybe a layer above that, Tuhin, talk about this map a little bit. You are working with so many different-- Clouds, yeah. --clouds. Some of these have even announced an inference platform of their own. Yeah. There's no reason for others to not do so. Maybe talk about that. It seems to us, to the class, a lot of speakers
[05:52:48–05:53:15] have told us that compute is very scarce. Yeah. Probably true. How are you getting your compute, and what's the strategy going forward? Yeah, look, we work in top of 18 clouds. I think it's 20 now, actually. And we have 87 different clusters where we're stitching compute together. This is the core, I'd say, piece of technology that we have is that we have the ability to take a bunch of different GPUs from different places and stitch it together in one place
[05:53:15–05:53:44] and make GPU fungible in that way. The reason we do this is twofold. One is it's access. It gives, it's very, very hard to find GPUs. And if you add a constraint around which clouds they operate in, that makes it even harder. And so what we do is that we take GPUs from anywhere we can, stitch them together and abstract that away from the customers. As to the first thing-- and that's a core part of the strategy. And we'll continue to do that.
[05:53:44–05:54:13] We'll rent and we will own. We will rent versus own. Yeah, for the most part, but we will also build our own ownership. We'll come to this in a bit. Yeah. But in terms of why, them going after a lot of those platforms having their own inference things, it makes sense. There's a lot of value in that stack, in that inference stack that sits on top of GPUs. We welcome it.
[05:54:13–05:54:40] We will partner with our competitors. We're fine with that. We stand behind the value that we are providing. But it makes sense. Inference is very, very sticky. Got it. I think truly, a lot of people think that inference is a commodity. And you've obviously studied these markets. Yeah. It just resembles the way that people used to buy databases to some extent. You choose once and you just grow there. Yeah. Well, it's also just was with one
[05:54:40–05:55:08] of the great founders you had listed on here this morning. They said that inference is so sticky for us because ultimately, this is the product that we deliver to our customers. We don't want any disruption in the core service that our customers are consuming. A disruption in that is actually-- sometimes even the biggest cost line item, even more than their cloud spend is the inference, the intelligence. Yeah. So I imagine it's a price as large will be fiercely competed for. 100%.
[05:55:08–05:55:38] And I think it is once you're in token path, if your inference is down, your product's down. Yeah. You don't want to mess with that. Yeah. I put a bookmark on renting versus owning. Yeah. Can you talk about that? We've had Chase from Crusoe here who talked us through the whole economics of building a data center, of owning the whole thing. You have taken the other side of it, which is to rent the-- For now, yeah. For now. Yeah. Talk more about that.
[05:55:38–05:56:07] It seems like an important decision and very different shape of the business. Yeah, well for our thesis was to some extent, what is the sticky part of the stack? And so we would say that software is the sticky part of the stack here. I think what a lot of these clouds who are competing with us would say is that access to GPUs is the sticky part of the stack. And once you have that, the software part is easy. I think it's probably arrogant both ways. I think they're slightly arrogant to think the software
[05:56:07–05:56:36] piece they could do, they could do it very easily. I think we are 100% arrogant to be like, oh, we can do that too. To some extent, we did that for speed and we did that because we had no business building data centers. For us going forward, though, it is becoming clear that-- what I would say is that no matter how much people tell you that there's a supply problem, it is 10 times worse.
[05:56:36–05:57:04] [CHUCKLES] I think you saw in the Google earnings, they had the GCP. Did you see this chart? Yeah. I think they've got like a 10x backlog or something. Yeah. If you go out right now saying you want 1,000 GPUs, truly you're probably saying people are talking about Q2 next year. Q2 next year, so 12 months out, maybe 15 months out. We have a cluster, a small cluster relative now at the time it was big--
[05:57:04–05:57:34] in one of in one of these clouds that-- of B200 Blackwell chips B200s, great chip. It's a bit old, but still it's an amazing chip. Our unit price right now is $263 an hour for that. And it doesn't matter what that is because just the relative peace of it matters more. That's up for renewal in October. And they came to us already in May and said,
[05:57:34–05:58:00] $510 is the new price-- Wow. --for next year. Double. This is fine. Are you going to take it? No, absolutely not. That's egregious. And we have a bunch of other clouds that we can work with here. I think, though, what is becoming very important is that there's a slide that was doing the rounds on Twitter. There was a slide that was doing the rounds on Twitter of Gnome
[05:58:00–05:58:30] Brown from OpenAI, who I think close to the old team of folks as well. He came up with the 03 model. And the slide basically said that access to compute is the strategic advantage for inference. That's becoming clear for us is that look, we run-- right now, I think the Baseten inference service across all our customers is bigger than the OpenAI API, is bigger than Gemini-- Wow. --in terms of how many tokens it does. We do around 30 trillion tokens a day.
[05:58:30–05:58:56] The Baseten inference service is bigger than OpenAI's API service? Not ChatGPT. The API product. Their API product. At least in last report. And it's definitely bigger than the Gemini thing. And so if you project that out, we'll need around 150,000 B200 equivalents in two years from now. That's just an insane amount of compute. Could you translate that in dollars for the class? Yeah.
[05:58:56–05:59:26] So it's about $7 billion of compute spend. Easy. $7 billion of compute spend. Too big. It's a scary amount. Yeah. The idea that we're going to be able to get access to that by renting is it's probably not going to happen. And the only way we can guarantee that is knowing that we have a strong relationship with the chip providers, and we can go buy them and put it up ourselves. And that's the reason we'll do it, is access to be able to fulfill the demand that we
[05:59:26–05:59:53] have for inference. And there's also an economic advantage to do it. It's about 30% cheaper. I think if you think of it as a scaled out cloud like Oracle, it's like what, 30% gross margins. Yeah. So it's about 30% cheaper to do that. Yeah, vertical integration, yeah. Fascinating. So the shape of the business is going to change quite a bit over time. Yeah. But you want to own your destiny. There's no other way. Yeah. If you think about what is the-- now
[05:59:53–06:00:23] I'll give you the third risk, but I'll tell them the three risks, which is open source, the app layer, and don't have access to enough compute. And the core risk of the business. And if we don't take care of that, we'll be in a lot of trouble. Yeah, makes sense. That makes sense. Actually, follow up to the question, you said that's $263 per hour, renewing at $510 seems egregious.
[06:00:23–06:00:52] Yeah. Where did you settle at? Oh, yeah, we haven't responded yet. In September, we don't need to negotiate with them. Actually, the question behind the question is, when should we expect the compute scarcity to get back to normal? Is this a 12, 15-month thing? Is it a multi-year thing? Is it a multi-decade thing? What's your best-- I'm sure you guys have done some thinking or economic modeling around this. I'm curious what you guys think.
[06:00:52–06:01:21] I'm curious what your publics people say about this. I don't think it's ever going to normalize. Yeah. I think it's what you said earlier. It's like, if inference demand is a billion x what it is today, two different things are happening, which is the applications are getting more agentic, and models are getting bigger, which both of those things say there's going to be a lot more inference, which just means there's going to be a ton more compute. Yeah.
[06:01:21–06:01:48] I think the analogy really is when you show up at JFK at 5:00 AM every day, there's no line. You go straight through. But by 8:00 AM, there's a line out the door. And we were out the door right now. But unlike JFK, which closes down from 11:00 PM to 4:00 AM, we don't have any reset period. And so it just keeps compounding. Yeah. Actually, that's an interesting analogy.
[06:01:48–06:02:15] Does the inference demand die down between midnight and 4:00 AM? And are you able to reuse that somehow? Well, what is it? When-- It's afternoon somewhere else. I mean, there's the thing, it's beer o'clock somewhere. It's like you could justify drinking a beer at any time because it's 5:00 PM somewhere. Got it. I think that's what same thing with inference, where it's like, even if demand here is down, in Europe, it goes up, in China it goes up. Yeah, yeah, yeah, makes sense.
[06:02:15–06:02:45] I'm going to shift gears a little bit. One of the biggest things we've been discussing in class. A lot of the students here will go on to build businesses, start businesses, work at businesses. If you were not building Baseten right now, what's your next best idea? What would you go start? What would you go build? That's a good question. I'd be going more the Crusoe route, to be honest. I'd be going and investing in energy and power.
[06:02:45–06:03:12] And I think the build out is we're just going to need so much space to put this compute. One of the ideas I'm most excited about, which I hope someone does, maybe someone here will do is I was driving over from Oakland to San Francisco the other day. And you always see the ports there and you see the containers and you start-- containers are one of the biggest economic drivers
[06:03:12–06:03:39] in history because they normalize the unit of trade. And so one of the things that I'm really excited about is modular data centers-- Modular data centers. --which is like take-- and standardize the unit of compute. Because whoever standardized unit of compute now we can just really industrialize. If you go and talk to all the folks who are putting up data centers right now and the people who
[06:03:39–06:04:06] own this space, everything is different. So to the extent that you can-- Fascinating. --you can modularize the-- And have a shared consistent format. You're creating an API for compute. And then you can-- then a whole industry will start around that, because then you can have the same people servicing these things all over and you create-- and that's probably what I'd be working on. Fascinating, fascinating. Well, I'm sure folks have taken that note down.
[06:04:06–06:04:34] Relatedly, outside of that, is there-- we play this long short game with every speaker. Is there a business? Is there a startup? Is there a founder that you're particularly excited about right now, and the other way? Who do you think is more hyped than they deserve? I'd rather not play that one. But for the former, look, I think all the things I said, the-- but yeah, no, I think any part of the build out.
[06:04:34–06:05:03] And then in terms of short, look, I think I'd be looking at companies that just aren't innovating. I actually think the enterprises are fine. But I think they need to figure out how to take the thing that is unique to them and pre-trained models and go do inference on them. But I think the ones that aren't doing that I'd be very, very scared about. Yeah, yeah. A lot of students right now thinking about what to study,
[06:05:03–06:05:18] particularly those who are starting their educational journey, what advice would you have for folks. It feels like quicksand right now. Yeah, I actually think you should just study whatever. That's fine.
[06:05:22–06:05:51] I think-- Yeah, you change careers. You could become an expert in anything in six months. So just do the thing that's fun, I don't know. But I think in terms of if, you are-- actually, financing models are spending a lot of time thinking about, how to finance data centers right now. OK. Just something I'm going deep on. And I think the project financing stuff is actually pretty-- you're pretty nifty there.
[06:05:51–06:06:13] And so it's up to study. Yeah, fascinating. Well, we'll open it up for questions. We've got five more minutes. Go ahead. What's your take on the futures? Oh. obviously, interesting, I think the market is way too--
[06:06:13–06:06:43] what's the-- if you saw how compute deals are done, it's basically a drug-- it's a drug market. Truly, it's like you have a guy. And-- Who's your guy? Yeah, we have so many guys. [CHUCKLES] And there's a guy at our office who kind of sits in the corner and he just calls people all day asking for compute.
[06:06:43–06:07:12] It's Ed. It's Ed, yeah. And absolutely terrifying. So like, I'm very-- it is clearly going to be market here, but I don't think the market has the-- it looks more akin to a-- it's not a mature market. I think a good way to think of it is like electricity markets are-- Just not as efficient, you're saying. There's a lot of slippage. Yeah, a lot of slippage, yeah. Go ahead.
[06:07:12–06:07:39] So Baseten's thesis is that compute influence would go [INAUDIBLE]? Yeah. What's the [INAUDIBLE]? And what's the antithesis of this? Yeah, I said that, the antithesis is that there's only two people who can get compute in the world, and they own everything. That's one. The second one is that we are dependent on open source models getting very good. And there's not enough open source models. That's number two.
[06:07:39–06:08:08] And three is that if you believe-- if you truly believe in the AGI thesis and you just play it out and you keep pushing and pushing and pushing, we all have nothing to do. But then I would argue maybe inference is the only market left. That's right. Is the only market left. That's right. Yeah. That's right. Go ahead. Everyone seems to be-- everybody talks about how important open source is,
[06:08:08–06:08:36] but we're not really finding it in the underlying software. We're not really finding it in the United States. Do you think, given the national security implications, that there's a role for the US government, or how can we create artificial incentives for private sector? Yeah, yeah, definitely. I think the US government-- I think it already is getting involved in thinking about these things. I would argue that actually when you go look
[06:08:36–06:09:02] at NVIDIA, when you go look at Microsoft and Google, they're actually putting a lot of work behind open source. I just don't think we've seen-- it hasn't borne fruit just yet, and it will come. So I think actually there is quite a lot of investment, but we're just not seeing it. And I think the government is getting involved. But yeah, definitely, I think you have to create artificial incentives
[06:09:02–06:09:30] and they'll have to be an alliance and for all these things to be incentivized top down, yeah. You saw the other side happen two days ago, the Chinese government invested in Deep Seek. Yeah. Anyway, go ahead. What's to stop Anthropic and OpenAI from open sourcing one of their lagging but still quite good models, and then providing a workforce to go into all of that data in enterprises?
[06:09:30–06:09:59] I think it's still focused on research. Why can't they do that? OpenAI already does. They do that to some extent, but in a lot of ways, it's-- for them to invest in post-training to some extent it's to almost give up on what they think the thesis is. One of the most fascinating things about OpenAI and Anthropic is if the thesis is AGI is everything the why they do anything at all?
[06:09:59–06:10:29] Every dollar should go to pre-training. But they already do that to some extent. But their incentive to push dollars towards that is never going to be the case, because the reason why-- all these models exist so they can just make the most money and the most expensive model they want. They want to push the capability gap, the capability argument as much as possible. All right. Go ahead. One final question. About the idea of the modularized data center,
[06:10:29–06:10:58] I'm just wondering what would be the difference between that and virtual machine. The question is, what's the difference between a modular data center and a virtual machine. Yeah, that's a good question. Well, I think the virtual machine sits a couple layers higher up in the stack. I'm just trying to figure out, how do you make it cheaper to put up more compute? And so that's in a lot of ways, you're
[06:10:58–06:11:27] kind of saying like modular containers. You're kind of talking about containers, which is once the compute exists you need that. And it kind of is that just two layers deeper. Yeah, it's at a different layer of abstraction. Yeah. I'll take one more. Was there one here? Go ahead. I'm just curious about where the [INAUDIBLE] are more internal [INAUDIBLE]? Yeah, it's everywhere. Disagg's the obvious one because you're really separating out concerns.
[06:11:27–06:11:57] The kernel stuff is continuing like a big-- like the amount of people in NVIDIA and even Stanford, I think, who work on stuff like thunder kittens and things like that. There's a lot of opportunity there. But the entire stack is changing. That's like the most fascinating thing about this market, is that everything is like quicksand. You throw everything away. Ilya might come out with a new architecture for a model and everything is up for grabs again. Ilya? Sutskever. Yeah, yeah.
[06:11:57–06:12:11] Do you what he's working on? I have sense, yeah. OK. Yeah. Tell us all about it. [LAUGHTER] All right. Yeah. Perfect. Well, thank you so much. This was fun. Yeah. Appreciate you making time. Thank you. [APPLAUSE]
[06:12:23–06:12:48] So today, we're going to have two incredible guests from both sides of the aisle. The aisles being Anthropic, former OpenAI, former Meta and now Chai. Flow of class today. I'll make a quick introduction, but I'll hand it over to the speakers who will present about their work. They're all doing both doing very, very exciting work. And then we'll have some time for Q&A towards the end. So maybe, without further ado, I'll
[06:12:48–06:13:16] quickly introduce our guests and have you guys over. Our first guest today is Eric Abrams. Eric has the fun job of convincing all biology to run on Claude. He runs biology and life sciences at Anthropic, launched Claude for life sciences last October, and they've got some really good partnerships with Benchling, Helix genomics, Novo. And I was just told that the last time you were here
[06:13:16–06:13:46] in this course was your final EE PhD scarred memory. It was a device physics final. I'm still traumatized from it. Good. Well, thanks for joining us. And then, Josh, you've spent your whole career teaching computers to do biology. We missed you here at Stanford, but it's good to have you back. You were early at OpenAI, like, early, early, early, like, almost a decade ago at OpenAI. And then at Meta, built the first life sciences product,
[06:13:46–06:14:12] which is almost half of all citations today are citing your work with the ESM1. You started discovery to make drug discovery, an engineering problem, as you call it, the CAD suite for molecules. And you've raised from the who's who, from OpenAI, from Anthropic, at Trive a big price. So we're very excited to get into it. Thank you so much for joining us. [APPLAUSE] Thank you. [APPLAUSE]
[06:14:16–06:14:45] We'll start with you, Josh. And I got your slides if you wanted to use them. Thank you. Tell us about your journey. You were not a premed major. You were not pursuing a medicine degree. How did you land up? How did you take the scenic road to Chai? So it's actually not totally correct. I think I did all the pre-med classes. I thought I would be a doctor. I come from a family of doctors. I have a bunch of older siblings who were already in medical school when I was a kid, so I wanted to do something different from them. So I learned how to code.
[06:14:45–06:15:15] And the thing that I loved about programming that I didn't really see you could do in medicine at the time, was you could really distribute what you were doing, right? You could write a piece of code, and then it was infinitely scalable. You could just send it to many people. And it was late in high school that I actually discovered biotech. And I realized that drugs actually had a similar property was another way to scale medicine and medical discoveries. So I always liked coding more than the lab, but they're pretty linked, I think. Nice. And tell us a little bit about your work at Chai.
[06:15:15–06:15:45] What do you guys do? And feel free to take a couple of minutes to walk us through it. Yeah. So you put it beautifully. Like we're building a computer-aided design suite for molecules at Chai. If you look at what's going to happen-- see if these slides work. Yeah. So one of the things we're gearing up to I'd say, it's called the medium term because the long term is so much of the lab work that we do today is probably going to move on to the computer. And it's going to move as a goal that we can illustrate here is we want to be able to design antibody molecules.
[06:15:45–06:16:14] These are about half of the approved drugs these days, are called antibodies. And there's a big trial and error process when you make a drug. You find some initial hit. You try to make a bunch of changes to it to get all these properties that you need for it to be a drug. And the big hypothesis is that someday we'll be able to just zero shot these molecules that are ready for patients right out of the computer. So we started the company with this big goal, thinking that this was going to become possible and that when it became possible, it would be a no brainer for most people to design drugs this way.
[06:16:14–06:16:42] So when we started the company, about 2 and 1/2 years ago, everyone told us we were crazy to do this. I mean, both because hey, is that technology going to work, and then B, don't you need to make your own drugs. The ambition of the most ambitious biotech companies is usually to become a pharma company. And then ambition that we were taking was like, no, we want most drugs to be designed this way, and we're not going to make most drugs ourselves. We really want to partner deeply with the ecosystem. So we had two contrarian bets in one,
[06:16:42–06:17:12] so grateful that we managed to get those good names on the cap table like you mentioned. But I think that there's been a ton of progress since then. Got it. So you would provide the CAD for molecules to a pharma business like AstraZeneca, like Pfizer and so on. That's right. Yeah. So some of the biggest pharma companies in the world are already using Chai's models in order to design their drugs. Got it. Fascinating. Thank you. Eric, over to you, sir. Please tell us about your journey. You've started a couple of companies,
[06:17:12–06:17:40] ran a bunch of companies. And now at Anthropic, what are you up to? Yeah, sure. So I originally was trained in math and physics, and then I got into AI research. The threat there was I thought that understanding intelligence was the great theoretical problem of our time. And then it was actually here in grad school, I was in Kwabena Boahen's lab, the Brains in Silicon lab, where starting as this arrogant physicist, I got my first taste of real biology, as we were studying the brain to try to find principles
[06:17:40–06:18:09] to bring back to AI. And it was love at first sight. I think ever since then, I've just gone deeper and deeper into the world of biology, through the companies that I started getting into molecular biology and biochemistry, organic chemistry. And after grad school, I had started several different medtech and biotech companies and have been in this world ever since. And I think when I was running the last company, it was a molecular diagnostics company called Detect,
[06:18:09–06:18:38] and I was checking back in on what was happening in the AI world and seeing the very beginning of ChatGPT. And I'll never forget when sonnet 3.5 came out, there was this step change in capabilities where I was asking it to do things that were coming up in my daily life of running a biotech company, of debugging some experiments that were failing in the lab and responding to FDA feedback and all sorts of things. And I was just shocked that it was useful at all back then. And it became very clear to me in that moment that this was a technology that we
[06:18:38–06:19:08] could use to accelerate the whole process, end to end of doing R&D in the life sciences. And I think Josh and I have talked about this, what we're doing, respectively, and I'll explain more in a moment, is complementary in that, at Anthropic, what we want to do is figure out end to end, how to accelerate the whole process of doing R&D in the life sciences. So this includes basic research, drug development. And within drug development, I think what Josh is doing is a big piece of the puzzle of actually designing the drug. But after that, there's a lot more.
[06:19:08–06:19:37] There's all of the clinical development, regulatory processes, the transfer to manufacturing. And our vision is to train Claude to be able to do it all, right? So at the model training level, we have a full stack model training program to make Claude state of the art and meet or exceed human expert performance and everything that you can imagine. The hardcore scientific fields that underpin this, in bioinformatics and chemistry and structural biology and so on, and the clinical and regulatory aspects of designing clinical trials and responding
[06:19:37–06:20:05] to regulatory feedback, even the strategic aspects of running a drug development program. So picking targets and modalities, organizing an R&D program, deciding when to advance or kill programs. So we're training the model to do all of those things. And beyond that, we believe it's not enough to have great frontier models. We also need to make sure that the model capabilities are accessible to people and integrated into workflows, and connect to all the tools that you need. And so we also focus a lot on the product layer
[06:20:05–06:20:34] and making sure that we have products that are optimized for life science professionals. And we have along those lines, something like a Claude Code for bio that will be coming out fairly soon that is our take on taking the power of Claude Code on the back end, but putting an interface on top of it that's designed for working scientists. So that you can visualize proteins and small molecules and play around with these things and have Claude know what you're talking about and spin up lots of compute to run, lots of models like Chai and so on.
[06:20:34–06:21:03] So we focus a lot on the model layer and the product layer, and we consider those two things together to be the platform that we want to take and scale throughout the whole life science world to achieve our goal. And our goals here are to get this end-to-end order of magnitude or more acceleration on everything that's happening throughout the life science world. And I think that the life science world is huge. There's so many different parts here. I think the first two goals that we're focused on
[06:21:03–06:21:31] are accelerating basic research and accelerating therapeutics development. And in each of these things now, we have more specific initiatives that we're starting. So for example, we recently started a wet lab where we're taking our model and our product and we're putting it to the test to try to see, how can we maximally accelerate this particular field of research that we're choosing to pursue in metagenomics discovery? And I think it's important for us
[06:21:31–06:22:01] to be dogfooding what we're building, and trying it out, and finding the limits and figuring out-- helping us figure out where to point all of our model training efforts and product development, et cetera. And then on the drug development side, I could speak a bit more about this later, but we're trying to figure out, in addition to having this platform of the model and the product, what else do we need to do as a society to deflect the trajectory that we're on to alleviate the burden of disease and aging in a reasonable time frame? I don't want to get stuck claiming
[06:22:01–06:22:28] that we're going to cure all disease in X number of years. We can get stuck there for a long time. But I believe in Anthropic, we believe that amazing things are possible. And we're certainly not going to cure all disease in the next five years or anything ridiculous like that. But we're really focused on how do we deflect this trajectory and get us on track for outcomes like that. And so we have all sorts of more specific things that we're getting going there. Eric, I was hoping you were going to say AGI implies cure of all disease,
[06:22:28–06:22:54] but maybe it'll take a little bit longer than AGI. One follow up on something you said, Eric, is could you help us understand-- the class understand what is the process start to finish for a drug to go to market. It sounds like there's a bunch of different phases. And maybe if you could almost on the x-axis being time or median time for a drug, or pick an example, if you'd like.
[06:22:54–06:23:23] And where in that process is AI, if at all? Yeah, great question. So just to set the scene, on average, developing a drug end to end from the moment you have the idea to the moment that you have FDA approval, and it's on the market, it takes about 10 to 15 years, something in that range. There are examples where it can happen faster. And I think the world record was closer to the 5, 6 year mark, something like that. But that's the median outcome.
[06:23:23–06:23:51] So it takes a long time. And there are a lot of distinct steps in the process. I think one of my axes to grind is that when you look at the debate about what are the bottlenecks in drug development and where are the opportunities to pull time out, a lot of people are quick to say it's all in the clinical trials, or it's all in the drug design. It's not all in any one of those things. It is distributed across, I'd say, 5 to 10 bottlenecks. So it's not an infinite number of bottlenecks. It's not hopeful, but it's also not just one thing
[06:23:51–06:24:18] that's way overly simplistic. So if we walk through the process, the first step is well, you have to decide what disease you're trying to cure the first place. And once you have identified the disease and the patient population that you're trying to work in, you need to select a target. So the target is place, molecule in the body that you want to develop a drug for. And you're saying with that have a hypothesis that if I effectively drug this target, that it
[06:24:18–06:24:48] will be safe for the patient, and it will cure the disease or be helpful. So the first step is identifying good targets. And there's a lot to say there. I think the first thing I'll highlight is that we have a big problem in the therapeutics industry that is referred to as target crowding, where if you look at in any given year, how many net new targets are pursued in the clinical phase by the whole therapeutics market, in the whole world. It's on the order of about 30.
[06:24:48–06:25:17] 30. That's it. So for that reason alone, we're not on track to cure all disease in any reasonable time frame if we're only going after 30 net new targets every year. And how many targets are there in total? What is the universe of target? What's the TAM? Many, many thousands. Probably on the order of 10,000. So there's-- Got it. This is overly simplistic, but there's 19,000 genes in the human genome. And so presumably any one of those could be a drug target. Not all of them will be.
[06:25:17–06:25:47] And you also you have one to many relationships, but certainly it's on the order of thousands. So that's step one. And how much time budget of the 1915 years has already been used up in this first part? Well, actually, somewhat pessimistic things is the clock doesn't really start until you've selected a target. But that's an important part of the whole process. And so once you've selected a target, then it's Josh's time to shine of actually developing the drug.
[06:25:47–06:26:15] And so you selected a target. And the next step is you have to figure out what type of drug am I going to develop. Will it be an antibody based drug? Will it be a small molecule? Would be one of these imaging modalities like molecular glues and other things are genetic medicine? So then you select the modality, and it's not obvious. Sometimes there's many choices that you could make, but then you actually have to design the drug right? So this is the preclinical phase where you're doing the fundamental science of trying to develop a drug that will bind your target
[06:26:15–06:26:44] and be safe for the patient, not have off target effects, and not only that, but be processed by the body in a reasonable way so it doesn't get lost or converted into something else, et cetera. So this is, I think, what people would normally associate with the scientific phase of drug development. And this part, on average, takes about four years if you look historically, from selecting a target to starting clinical studies.
[06:26:44–06:27:12] And so this is the first thing I think a lot of people get wrong in saying that there's no opportunity here. There's huge opportunity here. And I imagine that you feel similarly. I think that we could take that four years down to near zero, in principle, in the theoretical limit. Right. And I actually think that my own opinion is, it's not that important whether we get all the way to zero shot drug design, one shot would be pretty good too. If you just ask the model for a drug, and it gives you a few candidates,
[06:27:12–06:27:41] you do one round of testing and a couple of weeks later, you have your drug. That would be incredible. Compared to the four years that we're talking about. But I can let Josh elaborate on this. But even within the drug design phase, there's so many different parts. So you get your initial hit, which is like we're just getting in the right ballpark. We can hit the target. And then after that do optimization to get to leads, right? And you're trying to get all these additional properties that you want related to how manufacturable is it. Does it have off target effects? Is it stable in the body? Is it processed in the way that you expect?
[06:27:41–06:28:09] So there's all this optimization that you do that I think is really ripe for AI, but trying to accelerate this a little bit. So you design the drug, you have the drug candidate. And in the field, it's often referred to as a development candidate. So you have your development candidate. You lock your manufacturing processes, and you have to meet all sorts of pretty high bar for the rigor with which you're manufacturing this and documenting it. And then you're ready to start your clinical studies in humans.
[06:28:09–06:28:36] And so you have to get something called an IND from FDA in order to start these trials. And so then once you get the IND, then it's the clinical phase. And there are three parts of that. So there's a phase I. This is historically there's all sorts of various variations that are possible, but there's a phase I focus on safety. Traditionally then a phase Ii that's your first look at efficacy. And then a phase III, which is the final validation study that produces the evidence that your drug is safe and effective in humans.
[06:28:36–06:29:05] And that's the basis of The Ultimate FDA clearance. In parallel is all that is happening, there's all sorts of work on the transfer to manufacturing side as well. So that's the whole process. If you'd say like the developing the drug is 4 years, then the clinical phase, everything up until the FDA period is something on the order of like six to nine years or something like that. So that's how time is allocated. Cost is very much backloaded on the clinical phase.
[06:29:05–06:29:33] But I think there are really compelling and straightforward opportunities to use AI to pull time out of that process in both parts, on the preclinical phase and on the clinical side. But that's an overview. Yeah. Super helpful. Thank you. Josh, you have clearly decided to focus your efforts on the part that you found to be the most high leverage. In this framework, where are you? Where is Chai's work? And how are you going to compress the timeline? So I actually give two answers to this.
[06:29:33–06:30:03] I think Eric did a great job, like breaking down this into all these different phases and creating a drug is so complicated. And one of the reasons why the pharma industry has so many of these different stages is that how do you even know if you're on track. You do have to break it down into a bunch of different mini games and win each of them. But at the end of the day, what we're really looking for is some molecular matter that actually modulates some kind of disease. Like that's what the process of making a drug is. And the high level, we have the preclinical phase, where we're trying to find that molecule, and the clinical, which is where we're
[06:30:03–06:30:32] trying to prove that it's going to work. But if we just think about it at a high level, it's finding a molecule that does something in a patient. And one of the things that's so powerful about AI is that these things really can be linked again. If I come up with a better and more potent molecule, hopefully that makes my clinical trial easier. One of the really sad things about the space right now is there's so many drugs that are like extending lifespan by one month or two months or something like that. And again, for those patients that need it, that's amazing. And we should keep doing that work. But when you're looking at changes
[06:30:32–06:31:01] that small, clinical trials are really difficult and being able to design a trial such that you can actually see that with statistical significance and get your drug approved is very difficult. If you have some amazing data, your trial becomes a lot easier all of a sudden. And I think we can actually look at generation as a lens of what's possible here. If you look at how people used like LLMs for code Gen, even like a year ago, you would get some initial solution, you would try to debug it, you'd put in some other LLM to ask some questions about it. This whole iterative process and then eventually
[06:31:01–06:31:28] get some code that works. Now you can zero shot an app. And I think that's what zero shot drug discovery promises to potentially unlock. It's not just taking four years and getting it down to 0 years. I think it's also just about getting better medicines overall. We talked about targets and the TAM, right? And there's 30, a small number of targets that people crowd around, but, like, 80% of named diseases don't have an approved medicine. And that's a huge problem here as well. So I think just creating better medicines,
[06:31:28–06:31:58] creating more medicines, that's what it all comes down to. So with all of that said, the place that Chai focuses in right now, because again, as you're building a startup, you need to focus somewhere. So we've started on that molecular generation process, because we just see that as the apex for all of these other things that can happen around that. And one of the reasons why we're so excited about your work, for instance, is that if we can speed up the whole outer loop of, right now, those iteration cycles, then we can just get more shots on goal coming out of the model. But I think as both of these things get better, they're just s to reinforce one another.
[06:31:58–06:32:25] And I think the whole thing is going to be exponential in the next couple of years. Fascinating. So the obvious question is, why now? This is a problem as old as time. We've wanted to live forever forever. What has changed? What new technologies are unlocked, and why is that going to compress now? Well, so I've been trying to do this stuff my whole career, honestly.
[06:32:25–06:32:55] So you mentioned the story before, OpenAI to Facebook. And I think it shouldn't come as a surprise to anyone that AI has gotten a lot better. The model architectures have gotten to a point that they're a lot more scalable, a lot more compute has come online. Everything is just moving faster. And there's a lot of new data sets that we can build as well. And I think the other part of this too, is that I think there's also a lot of willingness to experiment with these things. If you think about how Chai orients ourselves, we bring our models into pharmaceutical companies.
[06:32:55–06:33:24] There's also the other side of this, which is like you can build the thing, but people also need to be willing to experiment. And A, people are just really open to AI right now, but then B, there's also a lot of pressure. If you look at what's happening geopolitically in life sciences and actually any of these pharma forms, and whenever I go to a conference about pharma, there's two topics these days. There's AI and there's China. And one of the things that the US is struggling with right now is that China is just outrunning us when it comes to drug discovery.
[06:33:24–06:33:51] They're just much more efficient. They work harder. It's cheaper to discover drugs there. This is even on the preclinical side, not to mention the clinical side, the regulatory state there. It's a lot easier to just get a drug into first in human in China. But one of the reasons why, if we just think about the US's role in biotech, we like, desperately need AI to work here because China can run faster than the US, but China cannot run faster than AI. And I think this is going to be a huge democratizing
[06:33:51–06:34:19] function in terms of our ability in the US to actually come up with better drugs. So I think there's a lot of why. There's both like, why now and then also like we really needed to work now. So we're lucky that it's happening in terms of what's happening in the world. Yeah. And is there a technology aspect to the why now for either of you that you might say, like, hey, obviously the demand has exists now because of AI and China, but has something changed? Have the large language models given us a leap forward? Is that a superpower?
[06:34:19–06:34:48] And what is the potency of that superpower in your estimation? Is this like a we're going to bring this down from 15 years to 15 minutes, or is this 15 years to five years? Help US size that. Yeah. So I think by far my first answer to the why now is the existence and fast progress of large language models. But I actually think that there's a few trends that are all converging right at the same time.
[06:34:48–06:35:17] And we're very fortunate because we didn't have to be the case, but it's the large language models is one. There's all sorts of other large models, like Chai models, these foundation models, and they're complementary because, as Josh said, the large language model I think of as the outer loop, it does what you do. As a human, you use Chai's model to get some designs, and you think about it and you test them in the lab. And you keep iterating, right. And hopefully, you don't need too many iteration cycles, but that whole process can now just be done by Claude. And it's getting better and better.
[06:35:17–06:35:46] So you have these two trends of the underlying foundation models are getting better. So you need fewer iterations in the first place, and you have the large language models that are available to do the outer loop and make that whole thing faster. So I think that's a strong trend. And I was giving an example of using these in the pre-clinical phase, but for the LLMs, the same thing I think can be said to some extent in the clinical phase as well. So that's one trend. The other trend that we're very fortunate that is happening right at the same time
[06:35:46–06:36:16] is the massive scale of data generation that's possible in biology. So we have all of these different measurement techniques that are becoming increasingly mature. Sequencing at this point is old news, but on top of that, we have single cell sequencing and proteomics and very high throughput techniques for measuring antibodies and all sorts of other assays that we're using to just keep fueling these. So I think these trends all happening at the same time is a big piece of it.
[06:36:16–06:36:44] As far as where the floor is, if we're in the 10 to 15 year zone, I think that we have clear line of sight to bringing clinical trial-- sorry, the whole drug development timeline down to certainly the five year range. And I put that as an upper bound. I think you can imagine the whole thing happening in a few years. Ultimately, there are lower bounds that are going to be established by the duration of clinical trials, but here, too, I think there's a lot more opportunity than people think for a few reasons.
[06:36:44–06:37:11] Someone will point out, for example, let's say, you're running a clinical trial to get at osteoporosis and classically, you have to wait and see how many people's bones break over a certain period of time. And so to get enough statistical power, you need to wait a year or something like that. And so that establishes naively this floor of a year for that clinical trial. But there are other methods that can in principle work,
[06:37:11–06:37:41] for example, if we can develop proxy measurements, if we take the right assays and we learn more about the biology, and we can tell from all the things we're measuring about the body that the intended effect is happening, you can imagine having different endpoints for the clinical trials that bring the time down. And also, as Josh said, I think this is a really important point, the larger effect sizes in the first place. So the number of subjects that you have to enroll depends on the effect size. And the more effective the drugs are, the fewer subjects and the faster the whole thing can go.
[06:37:41–06:38:09] So that's where I back into that time estimate. Yeah. Historically, there has been framing in life sciences of either you're developing the full drug and selling the drug and enjoying the revenue that comes from selling the drug or you're selling a tool. And at least most traditional biotech investors will value life sciences businesses, as the call it, revenue stream or free cash flow coming
[06:38:09–06:38:39] from their drugs forecasted back to today and have some discounted cash flow of it. And that's your market cap or the valuation that VCs are willing to pay. You have decided to build a platform of tools that you'll be selling to these people. And it sounded like Eric, you said wet lab. You're going to build a wet lab. So does that imply you're going to develop the full drug and monetize the drug at Anthropic? Yeah, good question. I should correct that misconception.
[06:38:39–06:39:08] So no, that is not our plan to develop drugs. We also are very much building tools that we want the rest of the industry to use. The purpose of our wet lab-- so I've been talking a lot about drug development, but our other primary objective is to accelerate basic research as an end in itself. And so the wet lab is our sandbox for pursuing that goal. So we actually have a basic research team that is just doing science, putting our tools to the test, and maximizing use of Claude through everything that you can imagine.
[06:39:08–06:39:37] And so the wet lab does serve a different purpose. Got you. So one of the biggest things that we talk about in this class is where will value accrue in AI. Certainly so far a lot of that has been in the semiconductors layer and not as much as got into the app layer. In your world, in biology, historically that has been at the layer that is selling the drugs and not at the tools. Correct me if I'm wrong, but you're both building tools and probably have a reason to believe it's going to be different this time.
[06:39:37–06:40:06] Why might it be different this time? I think you can give the simplest answer to this is that the tools are becoming more valuable than they've ever been before. If you look at where value has accrued historically, like the drugs that we have today on the market are extremely primitive. It's honestly a miracle at all that we can discover drugs with the tools that are available to us today. People often ask Chai like who our competitors are, and it's like literally the yeast and the mice and the traditional tools in the lab that are being used. And the drugs these days are as simple as trying
[06:40:06–06:40:35] to jam up like a target, for instance, on a cell. And they have all these side effects. And most drugs are actually, they have many off target effects that they probably shouldn't have. So I think as the tools become more powerful, it makes sense that more value should accrue there. Why should you pay that much for a tool that doesn't impact your probability of success. If everything that we're talking about here, if what Eric is saying, comes to pass and timelines go down, probability of success goes up, then by adopting these AI tools, if you're a pharmaceutical company, for instance,
[06:40:35–06:41:05] you should probably start trading at a higher multiple. Because if your probability of success goes up and your cost goes down for getting things to market and the timelines go down, already just by adopting these tools, you become more valuable and therefore a lot of value should accrue to the toolmakers. And the tools are just way more scalable than actually the end products as well. Yeah, one of the other things that we observed is the largest biotech companies in the west were founded decades,
[06:41:05–06:41:35] if not centuries ago, they're in the 20th century. There's a couple in the 19th century, and there's a couple of them that were formed in the last 25 years. Do you think one of the impacts might be more democratization, more startup forming because the tools are so accessible? I think it can go in-- there's going to be a couple of second order effects here. So the first thing is today you could start a biotech company just by being a little bit more efficient than a pharmaceutical company.
[06:41:35–06:42:03] Pharma companies are big organizations. And there might be some target insight that you get out of Claude. And then you rush a molecule to clinical trials. You get some proof of concept in a patient. And then a pharma company buys you up. What happens when two years from now, you can just zero shot that molecule. And you can zero shot a biotech. I think, does that mean biotech goes away. I don't actually think that happens. I think biotech just evolves around that. The same way today, if you can zero shot an app, you can't just sell a tool that's
[06:42:03–06:42:32] as simple as that's zero shot. You need to figure out either some other mode or some more complex software that has bigger value. So I think it's very hard to predict what's going to happen here because again, this is happening exponentially. There's a lot going on at once. There's a lot of second order effects. So I think on the balance, if everything stays stagnant might see accumulation to the incumbents. But in reality people are always going to find new opportunities. There's going to continue to be startups, and the outcomes will just be bigger
[06:42:32–06:43:01] than they've ever been before, because we have better tools to make these things happen. Fascinating. You guys have a very special vantage point, and you probably know all that's on the frontier of development. What are the biggest unsolved problems in your fields right now, that you're looking forward to being solved with either the models or otherwise? Yeah. I think there's a few that I would highlight. So the first is what companies like Chai
[06:43:01–06:43:30] have done for antibody development. I think doing that in other modalities is very much an unsolved problem. So antibodies are probably the most effective class of drug right now, but there are plenty of targets and diseases for which antibodies aren't immediately applicable. So we have small molecules and these other modalities that are emerging. And I'm really excited about taking these modalities-- the way I like to think about this is there's a certain frontier of targets that are credible today
[06:43:30–06:44:00] with our current methods. And we need to advance that and turn those into engineering disciplines in the way that antibody design has gone from what used to be, somewhat more like an alchemical craft into more of an engineering discipline. And so I think that is well within reach with AI. But is a hard, important problem to work on. And the other thing that I would highlight is scaling up the discovery of high quality new targets. So again, we're still crowding around a relatively small number
[06:44:00–06:44:28] of targets. And I think as a field, we're all trying to figure out what's that next big scalable way that we can unlock a huge number of high quality new targets to pursue. And there are a lot of efforts going on here, for example, in virtual cell and cell perturbation models, which is this new technique that is emerging and more traditional genetics based approaches for looking at human genetics data at the population scale and correlating it with health records.
[06:44:28–06:44:57] But I think that's a really important frame of as a society, we need a scalable way to find many, many more good targets. Yeah. Anything you would add to that, Josh? I mean, I fully agree with those. Maybe the one last one I would add is more sophisticated medicines. So I alluded to this before, but a lot of the drugs today are very simple. And I think as we get more control over them, you can think about, one of the things we do with Chai is we can fold up a lot of the proteins and the drugs that we're designing.
[06:44:57–06:45:24] So think of this almost as an atomic level microscope of what's going on in your system. And if you have that can now start to dream a lot bigger with the kind of molecules that are possible. So I think this is something that's also going to play out. Makes sense. The productivity of a big unlock in your field is incredibly high. There are companies that have single drugs that have led to hundreds of billions of revenue. Certainly many in the tens of billions of dollars of range.
[06:45:24–06:45:53] Is there a program or two that you guys are watching that you're like, hey, this is going to be the next GLP level success or has the potential? Yeah, good question. I would highlight there's some emerging targets coming out around increasing lean muscle mass that I just think from a commercial perspective, I'm making no comment on the value to society. [LAUGHTER] The incredible hulk.
[06:45:53–06:46:22] But I will say, I think there are some really important medical use cases for those. But that aside, if there's a drug that you can take that makes you ripped, I that's going to do pretty well. My team is going to kill me for saying this, but this is actually our internal joke at Chai as well. If you look Renaissance rentech, for instance, they have their internal medallion fund. So one way you get compensated if you go to Renaissance is you get to put money in the best performing fund. So we're all like we should if we ever make drugs that Chai, we should start with that.
[06:46:22–06:46:51] And it's just for Chai employees. So you come to Chai, you get jacked. [LAUGHTER] Nice. Any that you're looking at, Josh? I think that things for sleep are going to be really interesting as well. So I think what the GLP-1 and the obesity drugs are showing is they're almost like consumer medicine, essentially because of how big the patient populations are. And I think sleep is also something that leads to so many of the disorders that are out there. So if we can find ways for people to sleep better or get rid of sleep disorders, I think that's also interesting.
[06:46:51–06:47:20] And actually is a similar economics, if you will, to the GLP-1 that started with diabetes and moved to obesity. And same thing here. You can start with sleep disorders and then maybe that turns into I don't know if I can sleep more effectively, like who wouldn't want a drug like that. Yeah. So I think this might be the start of these consumer medicines and really change the way that we live. Fascinating. Wow, those are two pretty remarkable ones. Everybody in the class is wondering where should we go invest in for getting ripped
[06:47:20–06:47:50] and sleeping better. Maybe we'll start with you, Josh. Pick a business or an idea or startup that you're bullish based on everything you know about the world and the other way. What are you skeptical of, which you think is more hype than reality? Yeah, the two categories I'd point to first-- I mean, it shouldn't be a surprise, like pharmaceuticals. If we're going to have better drugs, then you can almost think of many pharma companies almost as capital aggregators.
[06:47:50–06:48:19] They have these drugs that are creating billions of revenue. They have a mandate to reinvest that into more drugs, and they might just become way more effective. And the other thing also, and I just want to give this a little bit more contrarian, but is actually lab experiments themselves. So a lot of people, when they talk to us, they assume that oh, Chai wants to make experiments dead or something like that, but actually think if you can design your experiments better, the ROI on experiments goes up. So I think there's going to be a Jevons paradox, actually, for a lot of lab work where we go into this Renaissance where
[06:48:19–06:48:48] and we can actually pull, like Dario's remarks like the incredible essays that imagine, like 10 years of biomedical discoveries happening in the next year. Give me a lot of lab work for that to happen. And AI is going to be driving it. So hopefully a lot more of Chai usage as well to make that happen and Claude usage. But yeah, so those are the things I'm most excited about. Yeah, anything you're bearish on. I'd be bearish actually. So I mentioned these things are kind of in the physical world, right? The things I'd be more bearish on are things that are purely in the software world but are not frontier.
[06:48:48–06:49:18] So for instance, a lot of the methods that are used today in the bio world, there's actually-- people have been using physics based methods. And computational based methods for a long time in drug discovery companies that have been around for 30 years, and they've had tremendous impact in the space. But AI is becoming quite powerful and is going to eat a lot of that stuff as well. So I guess you could summarize this as AI in the physical world is going to be really big, but you want to think about how to combine these things together. Makes sense.
[06:49:18–06:49:47] Over to you, Eric. Short. Yeah. So my category for long is similar to what Josh said at the end. I'm very excited about this class of companies that make it possible for Claude to run experiments. So there's two ways for Claude to be able to run experiments. One way would be to hook up to lab instruments and work through all those integrations, and we're absolutely working on that. But Claude can run experiments right now by just placing orders with contract research
[06:49:47–06:50:14] organizations or CROSs. What would I do if I wanted to do that? I would draft a protocol and I would email somebody and we'd talk back and forth. Claude can do that right now and is doing that right now. And so these aren't quite CROSs, but these are what I would call-- the class of business that I'm really excited about are these really scalable sort of AI native, real wet lab companies that are manufacturing materials
[06:50:14–06:50:41] or running experiments with high scale in a way that is easy for AI to interface with. So some examples of companies that I think are great in this space, Plasmid Saurus on all sorts of sequencing that you want to get done and Adaptive and Twist, right? These are companies that make it really easy to do these things at scale and execute phenomenally well in the wet lab part. So that's a class of things that I'm excited about. Awesome.
[06:50:41–06:51:10] Things that you're skeptical on? Yeah. So I have maybe a generic comment here, and it might sound strange with two people that have a business selling tools to companies that are doing therapeutics, but I'm generally bearish on companies with business models that are reliant on selling tools to pharma. Now, obviously you can make big businesses successful there, but I think it's very hard. There's very few examples of groups
[06:51:10–06:51:38] that have been successful doing that compared to just doing the thing itself. The barriers to actually going and starting to develop drugs have never been lower, and are coming dramatically down to the point where we like to call it pipeline in a person. You could have a single person using lots and lots of Claude and foundation models, actually running a portfolio of early stage drug programs. And soon I think that frontier will advance to that AI relatively small team can be running several clinical stage programs.
[06:51:38–06:52:07] So I'm just highlighting, of course, you can build great businesses on the tools model, but I think that path is really, really hard. And alternatively, I think democratizing actually developing therapeutics in itself is just getting easier and easier and more accessible. Wow, what a savage response. You're making sure nobody starts a competitor. [LAUGHTER] Do not compete with this guy. OK.
[06:52:07–06:52:31] You guys have a phenomenal setup as you said, very, very special spot. What would you be doing if you were not doing this? What's your next best idea? Asking for a friend who's not going to start a competitor to Anthropic. Well, I think the first thing I would be doing is what I said at the end, trying to run a bunch of drug programs as efficiently as possible because I already said that I would throw out one more,
[06:52:31–06:53:00] which I think figuring out how to connect AI to the wet lab is a big frontier also. So today, Claude can run experiments by communicating with CROs. That's great for today, but we'll reach a whole new level of efficiency when Claude can actually directly interface with lab instruments. Like instrumenting it. Yeah, giving it the arms and legs to run. Yes, exactly. And I think there's a lot of parts to this problem. There's the hardware layer, there's the software, protocol, communication layer
[06:53:00–06:53:29] on top of that. But it feels to me like there's something productized there to do. Fascinating. That's almost like a programmable chemistry. Yes, exactly. Like the lab in the box, like you plug this thing in and all of a sudden, Claude can control your wet lab. Is there an instantiation of this? Is there evidence of a business that looks like this? I think there are a number of companies that are pursuing-- mostly startups that are pursuing things like this.
[06:53:29–06:53:57] But I think it's the early days. We're maybe a year or two into what I expect to be a many-year journey. Got it. That sounds more promising than an Anthropic competitor. What about you, Josh? I think actually very close to that. It's like, how do you actually set up autonomous drug programs? I think the dream will be how far. It's almost. You should start of it. Think of it as an eval, even, of hey, can we actually come up with a set of benchmarks and see how far can the models push on each of these things? And I think that converges into, OK,
[06:53:57–06:54:26] at some point the models get past a certain mark where they've done something really useful. So I think you'd even take the frontier problems in drug discovery, many of which we've talked about here. And then you would keep putting the frontier agents on them, and then eventually, you'll get some breakthrough and you try to push that to market as soon as possible. One of the interesting things about pharma is it's one of the most competitive industries that's out there. Interesting. Because if you think about what these companies are trying to do, they're trying to treat disease, right? And it's not obvious how to treat a disease or even doing the commercial work of which
[06:54:26–06:54:55] is the best to go after just because it's so complicated. But like treating disease is like the most obvious business model, the most obvious business case. That's actually one of the reasons why I'm so excited about this stuff. Like a lot of friends in college were like, what's the next Snapchat to build. And what's the best company. I'm like, guys, a lot of disease out there. Why don't we just try to treat that? That's got to be valuable to folks. So I think just writing these things down and then just seeing how far can the models get there and then just being able to be the first to do that is probably the other most
[06:54:55–06:55:25] interesting thing to do. You want to be first in class. And people often say best in class, but being last in class is something else someone mentioned to me recently, which I think is a really exciting concept. Fascinating. We'll open it up for questions if folks have any. Oh, here, we go. Go ahead. So I come from a Pharma background. Small molecules, pulmonary drug delivery. You hear a lot about discovery. You hear nothing about development.
[06:55:25–06:55:54] Can you elaborate on why you think that is from your vantage point and what the opportunities might be? Yeah, I think there are a few reasons for that. So first, the opportunities in discovery are pretty straightforward. They're hard to execute on, but I think everyone who looks at it can name what they are. The opportunities in development, it's not any one or two super compelling easy things. I think it's also a lot of things that aren't obviously
[06:55:54–06:56:23] blocked by model intelligence, but are operational and you need to have all the right connections in place. So my opinion is that first, I mean, it's less glamorous to begin with, and it's less obvious. And it's less obviously related to model intelligence. I would put all those together. I also think the AI community, tending to be more scientific in nature, has less experience with that phase of things. And I think that plays into it. As far as what the opportunities are, I think there are many. So going through a few of the--
[06:56:23–06:56:51] ones that I'm excited about are optimizing patient recruitment, right? So you can imagine using AI to be intelligent about site selection and looking at prevalences and enrollment and things like that and making sure that you're choosing the right sites and enrolling patients as quickly as possible. There's the trial administration side of things where right now there's a lot of operational overhead that goes into site monitoring and entering and checking records in the electronic database and all that.
[06:56:51–06:57:20] So that should all be automated right away. And something that Claude should do. And on top of that, I think there's the more scientific components. This is where discovery and development interactive. Well, you want to make sure that the trial has a higher probability of success in the first place. And you want to make sure that the effect size is as large as possible. So those are, I think, two very significant opportunities that relate more to the discovery side actually. Yeah. Go ahead. I had a question.
[06:57:20–06:57:46] So I guess there are two questions here. But there's talk about these training runs getting bigger. Like now we're in the billion dollar phase. We'll go to 10 billion, possibly hundreds billion. What would that unlock on this, on, like, everything you're doing? When we think about LLMs, it's like, well, you give it a book and then you mask out like a paragraph, and it learns to fill it in. But what is that equivalent for the models
[06:57:46–06:58:15] that you're trying to build, and if you can do $100 billion training run, what would that unlock? And it's a really general question. Yeah, so I think there's a steady stream of model improvements with every model generation. And so the way that I think about it is there's just certain capabilities that we're targeting. And as we do more training, the capabilities get better. And sometimes there are scale effects where there are capabilities that are kind of flat,
[06:58:15–06:58:44] you don't have them at all. And then at a certain scale they start turning on. And so just to give a few examples, like we found that our models didn't understand much about proteins for a while. And then with Opus 4.6 for the first time, we started to see that the models were starting to learn more about proteins. And that has to do with things that we're doing in training and the size and all sorts of things like that. But I think those are the two trends that are happening here, of the incremental improvement in many different capabilities,
[06:58:44–06:59:13] and then also the scale effects that we've seen throughout many domains, even outside of the life sciences in AI, where sometimes you just don't see something until a certain model generation in the first place. Go ahead. Even in software, zero-shotting works very well for what's called shallow stacks, like for the services and so on. But our experience has been that it does not work well at all for deep stacks,
[06:59:13–06:59:42] like databases and distributed systems. And biology seems like the very extreme of that deep stack though. What kind of applications do you foresee [INAUDIBLE] actually working? This is a great question. I think part of this is a question of how challenging the task is that we can go after, and the frontier just keep getting pushed there. So what you said about here, not being able to zero shot a database. I don't even know if that's true anymore.
[06:59:42–07:00:11] I wouldn't be surprised if you guys have already achieved that. And I think the same thing applies in biology. It's funny, when we started the company, one of our investors did a lot of work for us, talking to doing all the expert networks and saying, hey, how much would you pay for this sort of thing? And people were like, oh, we wouldn't pay much for zero-shotting a molecule. And they're like, why? And they're like, because it doesn't work. Why would I pay for it? And they're like, let's assume it works. And it's impossible. And they're like, OK, so it looks like we're on to something.
[07:00:11–07:00:40] So I think a lot of these things, I think, that's so exciting in AI right now, you really can dream big. It is possible that the stuff I have up on the slide over here just is not possible. We don't know. But you see scaling laws for these things. And as you build more data sets, as you bring on more compute, they start to come in reach. So I think it's more a question of when rather than if. But, to be honest, we don't know for sure. And I think that's part of the adrenaline that comes along with building in this space.
[07:00:40–07:01:08] It seems like the loop has to have this thing. [INAUDIBLE] It's the mega-mega [INAUDIBLE] loop, so to speak. I think actually it relates to the last question, too. I think one of the reasons why you hear a lot more of AI and discovery is because the feedback loops are shorter. So for instance, if it's in clinical development, it's like, OK, are we going to go do a clinical trial with AI? That might take a long time, versus just to design a molecule and test it in the lab.
[07:01:08–07:01:24] You can do that in a couple of weeks. Also, everybody's going to be single-shotting the drug to get ripped and sleep better. We're going to wrap it here, guys. Feel free to come find them downstairs. Thank you so much. Really appreciate it. [APPLAUSE]
