Agent 的时间尺度正在变长,但“能运行很久”不等于“会持续正确”
00:00:07–00:05:58Insight
- 他认为“AI 达到初级工程师水平”的旧判断在某些 agent-based coding tasks 上已经相当接近,但定义和任务范围仍会改变结论(约 00:00:35–00:01:30)。
- 2027 年预测指向自动问题分解、自动实验循环与结果组合,也可能扩展到目标可测量的科学和工程任务(约 00:01:51–00:02:47)。
- 长时 Agent 的例子包括把软件重写到具有不同安全或性能属性的语言;访谈没有提供系统性成功率或成本数据(约 00:04:55–00:05:58)。
主持人 All right. Should we go Should we get started, Jeff?
嘉宾 Sure. Sounds great.
主持人 All right. Jeff, welcome. And again, thank you so much for being here. Especially I just got a cold and thank you for being here.
嘉宾 Yeah, I'm afraid I've lost my voice. I don't normally sound quite like this, but we'll we'll do what we can.
主持人 So, um, you built map reduce, big table, tensorflow, the TPU, Gemini. We could spend a whole hour on all the things you've done, but what I love is that you're still making bold predictions in public. Last year, yes, last year in May 2025 at AI Ascent, you said that AI is at the level of a junior engineer. That was about a year ago. It's been How close are we to that prediction?
Yeah, I mean I feel like uh the models have been getting a lot better at sort of agent-based longer running coding tasks and it seems pretty clear that they are now actually pretty capable and depending on exactly your definition of junior engineer it seems pretty spot-on I would say.
嘉宾 What did you underestimate from that prediction? Um I mean I think the ability to do more and more complex tasks has been growing faster than I thought. Um and I also think uh outside of coding these these agent-based systems are are really starting to shine in other domains and I I think uh you know uh that's that's going to be an important trend in the future.
主持人 So give us another bold prediction. What do you think is going to be the 2027 edition?
嘉宾 Uh I think you will see a lot more automation of uh ML systems themselves. um basically getting ML systems to improve their capabilities by running lots of experiments, breaking things down into subpros, you know, running those subpros in a tight automatic experimentation loop, putting the results together and being able to then uh you know, get some improved system uh out from that uh sort of fully automated problem decomposition and and automated experimentation that I thin
k that's going to be really exciting. M
主持人 I think that also applies not just to ML but also to other fields of science and engineering. Um basically anything where you can have a measurable objective uh I think you can uh actually make a lot of progress these days.
嘉宾 Now let's go back to a little bit in history. Back in way back in 2001 Google search used to run on hard drives.
主持人 Yep. And you and Sanjay did the math and realized that at some point the whole search index would finally fit in all of the RAM of all the computers you had running and you made that radical realization and you basically in few days with Sanjay shipped in production a whole new search version that worked in RAM rather than hard drive and that was the thing that got Google to be so fast. Google searches.
嘉宾 So history tends to remix. What is the it fits the memory moment right now in 2026 that everyone in this room is still should be thinking about and designing?
主持人 Yeah. Yeah, I mean it's a little different, but I think uh you're going to see more and more uh uh high performance and um low energy uh inference hardware systems because I think everyone is now realizing that inference is the key to making you know these agent-based systems be available to more and more people and that latency is really important and that specialization of the hardware is a really key way you can make uh things that are more energy efficient and lower laten
cy than more general purpose uh computational devices like say GPUs or TPUs
嘉宾 because I think all everyone here is used to waiting for responses on on models. So
主持人 waiting is no fun
嘉宾 master speed. So you're saying what if we don't have to wait anymore?
主持人 Yeah. I mean, I think we'll imagine what you could do with something where the latency is, you know, 50x better.
嘉宾 Interesting thought. Now, what's one assumption that perhaps 6,000 people in this room hold that's already false about AI?
主持人 Yeah. Uh, that's that's a good question. I mean I think um probably one thing is people don't quite realize how possible it is to have you know agent-based systems that can run not just for an hour or two hours on a problem you care about but for some problem domains and with highly capable models underlying them you can get them to run for days or weeks and do really complicated tasks and I think that's you know starting some people are starting to see inklings of this but I
don't think everyone has really internalized this and that's going to be really uh a pretty big deal.
嘉宾 What's a particular task that you have run that has run for weeks? What what was it? What did the tell what did you tell the agents to solve? Yeah, I mean I think uh you can tell agents to uh go off and implement um you know completely new versions of software in different programming languages that might be you know have better safety properties or better performance properties uh that and then they can go off and and actually do that in a you know pretty serious way.