视觉为何成为智能与现代 AI 的起点
00:03:46–00:18:34Insight
- 她把寒武纪前后的感知演化与高级神经系统发展联系起来,并指出人类大量 cortical activity 与视觉功能有关(约 00:04:30–00:07:30)。
- ImageNet 收集约 1500 万张图像,并以 1000 类 object recognition challenge 推动可比较的算法进步(约 00:09:30–00:16:30)。
- 2012 年的明显跃迁来自数据、neural network 和 GPU 的汇合;访谈将其视为现代 AI 的 inflection point,而非单一技术的独立胜利。
I think the biggest thing humanity never learns is the older generation lamenting about the future generation as if the future generation doesn't know anything. They're rude. They're they're they're forgetting the past. But if you look at arc of history, of humanity, by and large, we advance for the better. Now, I'm not denying the atrocities. I'm not denying the setbacks. I'm not denying this. But fundamentally I'm a optimist in humanity. I look at kids, they're curious. Of
course, they get massively entertained by this technology, but they also are starting to use it. What I worry about are teachers and some parents because I think our society today and especially Silicon Valley are not doing them a service. We're forgetting about them.
Hey everyone. To celebrate the launch of my new book entitled Protocols, I'm pleased to share that I'll be hosting three live events very soon. The first live event is in New York City at Radio City Music Hall on September 17th. The second event is in Los Angeles at the Dolby Theater on October 8th. And the third live event is in San Francisco at the Masonic on October 28th. At each of these events, I'll be discussing topics from the book and my favorite part,taking questions
directly from you, the audience. To get tickets, you can go to hubermanlab.com/events and use the code protocols to get early access. Again, that's hubermanlab.com/events and use the code protocols to get early access to tickets. Welcome to the Hubberman Lab podcast where we discuss science and science-based tools for everyday life. I'm Andrew Huberman and I'm a professor of neurobiology and opthalmology at Stanford School of Medicine. My guest today is Dr. Fay Lee, a comput
er scientist and professor at Stanford and one of the pioneers and luminaries of artificial intelligence and computer vision. As you all know, millions of people use AI chat bots to look up information every single day. And of course, many people are concerned about AI, where it's going, and how it might replace certain human jobs or degrade our experience of life in one wa y or another. Today we discuss from a neuroscience perspective what intelligence really is and the ways
that AI can and is being used for good meaning to truly enhance learning health and to enrich rather than diminish the human experience. We start off by talking about how human brains of all ages learn new information. What rules the brain follows in that process and how AI because it is based on the content of the internet both resembles and falls short of what human brains can learn. and we discuss exciting uses of AI and robotics in medicine. To be clear, FFE acknowledges
and addresses the many valid concerns about AI. But as the director of the Stanford Institute for Human- Centered Artificial Intelligence, her goal is to make sure that humans and humanity at large are represented in where AI goes next. As you'll soon hear, Dr. Fa Lee is an extraordinary scientist and educator.
She has been called the godmother of AI for her ushering in of AI technologies, but also for her insistence that the ethics and benevolent uses of AI stay central to AI and robotics. So whether you are young or old, today's conversation will inform and empower you to understand and use AI in ways that truly benefit you and enrich your life. Before we begin, I'd like to emphasize that this podcast is separate from my teaching and research roles at Stanford. It is however part
of my desire and
主持人 effort to bring zero cost to consumer information about science and science related tools to the general public. In keeping with that theme, today's episode does include sponsors. And now for my discussion with Dr. Fay Lee. Dr. Fay Lee, welcome.
嘉宾 Thank you. I'm excited to be here, Andrew.
主持人 Yeah, this is a long time coming. And yes,
嘉宾 you are a luminary in this AI field, but I also consider you a neuroscientist and computer scientist, and we share a common path through vision science. And so I'd like
主持人 and fellow colleagues at Stanford. So I'd like to start in vision. What is so special about vision and seeing and light as it pertains to AI and where it's all going? Because I think for most people those probably sound like very divorced themes but actually that's where it all starts.
嘉宾 Yeah. I see vision as a cornerstone of intelligence in almost two parallel way. One is what evolution has taught us. You know what's the evolution of vision and animal intelligence and human intelligence. The other one is computer vision and AI what that relationship is. So I'll go into each evolution. I always say that 540 million years ago animals saw the first light. These are simple sea ocean animals, trilobytes and and the the cousins. And before that there was very little sensing. Uh around that same time tactile and haptics was starting also to emerge in animal bodies but but there was no hearing. There's no you know smelling there's no but there's absolutely no nervous system. But the first photoreceptive cells created a evolutionary force that propelled animals to evolve because sensing the external world changes your self-perception changes the way your relationship with the external world. To put it simply, if you seek you can see food, it changes your your life, right?
from a evolution point of view and you become someone else's food and also you're actively seeking food. You're actively seeking mates and and and all that. So really because of sensing and perception evolution took a incredibly accelerated pace in terms of uh animal speciation. Fossil studies have told us that 10 million years after the first uh light for animals was what we call the the big ban of evolution or Cambrian explosion of animal speciation. And fast forward I think vision has always played a huge role in not only in the early evolution of animals but as well as um advanced intelligence and how that emerged. You and I are both vision student and and scientists. It is estimated half of the cortical AC activities in human brain is involved in visual function. Children were first visual before they were verbal in development. So vision really to this day plays a central role in both the evolution of animal intelligence as well as in the daily life of human human life. Now in parallel, vision as a uh as a discipline or as a area of uh artificial intelligence was really played a pivotal role in what we see as this modern AI moment in a couple of ways. First of all is the the uh algorithms the neuronet network algorithms. Neural network algorithms were first computer scientists start dabbling that in the early 1950s. And Andrew, you might remember what's happening on the neuros side in the early 1950s is that neuroscientists like Hubo and Viso were starting to record visual cells in malian brain and starting to realize there is a hierarchical structure of nervous cells that stack against each other and pass neuroinformation across these hierarchy. And it goes from you know collecting light from retina all the way to recognizing there is a shape in front of you. And that very neuro architecture that we see in mamalian brain is also part of the inspiration of neuronet network algorithm. Now today's neuronet network algorithm runs on hundreds of billions and even trillion of parameters. It has the complexity that departs from what we recorded in the mamalio uh brain or the visual pathway but the origin is very close to each other about half a century ago um a little more than half a century ago. That's one aspect of uh vision's contribution to AI. There is another aspect of vision's contribution to AI that is also pivotal which is through big data is that that comes closer to my own work is that AI around the century was a field of machine learning a lot of different labs different research scientists were were trying out different algorithms and it's not just neuronet network there are other methods jargon words like Beijian methods, support vector machine methods. It doesn't matter what these methods are, but it's a explorative phase that we're trying to get these algorithms to work so that we can empower the machine to read or to see. A group of us computer vision scientists were struggling with these algorithms and uh I was a very young faculty um first year faculty 2006 at Princeton and my students and I are looking at these algorithms and how little data were fed into these algorithms to learn. So I turned to cognitive neuroscience literature per namely vision literature and started to study how much humans learn, how much humans can see and the numbers were incredible. Humans were by age six can learn tens of thousands of different object categories and the exposure to visual world is also massive. Right?
babies can see the mo most of the time the moment they're born. So they're inundated with this big data. So we conjectured that the lack of data was a huge part of the reason that's the lack of progress in AI. So we took a departure from everybody else who are really focusing only on algorithm and said that we need data. we need data to drive these algorithms. So long story short, we led this um image that project that collected the first ever internet scale large data set for the field of artificial intelligence, but really through the field of vision because imageet is a collection of 15 million images. And the goal of imageet was to drive machines to recognize everyday objects, you know, microphones, cups, chairs. And that work converged with the advances in neuronet network algorithm as well as in GPU computing. And by 2012 that work uh that the convergence of the three elements of modern AI became the defining moment of what um what modern AI is. I recall somewhere around 2012 it seems there was this debate at this vision course at Cold Spring Harbor that was held every other summer like could a computer learn to recognize specific faces as well as humans. Now I think most people would say computers are actually much better at it than humans are even though you have these super super recognizer people who are exceptional at this.
主持人 Could you tell us how is it that this technology went from a state basically where it would confuse you and maybe a a a cousin or or even someone that looks somewhat like you could
嘉宾 or to the point where uh to the point where now it is exquisitely precise.
主持人 How do we get here?
I want to definitely double triple click on the convergence of this technology. I think around the second decade of 21st century. So like you said around 2012 the the huge convergence was the capability of GPU computing which basically accelerated or parallelized computing so that you can have more flops going through algorithms right you need that speed then you also have a um after many decades of research neuronet network algorithm them is getting more mature. Um you know starting as we said 1950s people start to um create these very simple algorithm that behaves similarly to neurons but much simpler. Neurons as you know are very complex but here the idea is that you have one unit of node that takes some some input and outputs another input and within it it's just a function a very simple function. So you stack them together. That's what neuronet network is. But by by the time it's in the um after you know around 20 uh 2010ish the maturity of these algorithms have have gotten to a level that it's it's becoming really good. But also last but not the least the recognition of big data. Internet definitely fueled that. It made data more available. But the reckoning moment of wow big data needs to be part of that equation. We need to use big data to drive these algorithm to learn these patterns. So this convergence of these three things really set off um the the revolution of AI. The specific moment is also worth mentioning because you mentioned face recognition is this image net challenge. My lab put forward that starting 2010 after we collected this humongous data set, we at that point GPU was not yet mature uh and and uh we put out a uh public challenge for the research community uh for for multiple years in a row and invited people to solve this major computer vision problem called object recognition. The task was very easy. We have a data set of a thousand different categories of objects and this data set is more than a million images large. It's what we call the testing data set and uh the task for the algorithm is I'll show you a picture.
You have to name the the main objects inside and if you guess right you're you get a point. If you guess wrong you don't get a point. So that image that challenge uh we later a couple of years later benchmarked human performance by a very smart graduate student at Stanford and that was roughly 4%. So random chance will be one over a thousand
嘉宾 right? So 4% for humans is not that bad. The first few years machines were not as good as humans. The turning point was 2012 the convergence of neuronet network image net data set and GPU even that year even though the error rate was was cut um to oh by the way the human performance error rate was 4%. Sorry I I need to correct that the error rate was cut down to to the teens. It wasn't where human performance was. So this is looking at images and and assigning a a
主持人 one out of a thousand labels.
嘉宾 Got it.
主持人 Yeah. But 2012 was so momentous that year because the error rate from previous algorithm dropped a lot by this neuronet network algorithm. And we know in the research community when something this drastic happens it it means a inflection point. But it still took another three years I remember by 2012 2016 for the algorithm to beat humans in in naming a thousand objects.
嘉宾 Could I ask you where this 4% error is coming from in this very smart graduate student? Is it that they don't recognize the objects or it's a recognition against time pressure? like they have to they're being fed images fast enough that occasionally they do an incorrect assignment.
主持人 I don't think the time pressure was the main issue even though for a graduate student to do this I don't think they want to do this forever. Um but I think you know the the human brain as you know has limited memory whether it's long-term or short-term right so retaining the patterns of a thousand object classes even if some classes you're you're familiar is is not that easy
嘉宾 you know so so I think there is the confusion and and also for example different species of dogs gets really close.
主持人 Mhm.
嘉宾 And that that's a challenge.
主持人 I'd like to take a quick break and acknowledge our sponsor, Lingo. Lingo is an everyday wearable that tracks your glucose 24/7. Glucose drives a lot of key processes that support energy, body composition, and long-term health. When glucose is constantly spiking and crashing, that's where we can start to see metabolic dysfunction. And over time, that can even progress to pre-diabetes. Right now, about 115 million adults in the US have pre-diabetes. Most don't know it, and a higher percentage of men have it than women do. Often, there aren't clear symptoms of pre-diabetes early on, so people don't tend to look into it. But the fact is that metabolic health is shaping how your body functions every day, whether you feel it or not. Tracking your glucose with Lingo can help you see how food, activity, and stress impact your glucose throughout the day. I personally have used Lingo and it's been an invaluable tool for improving my metabolic health. If you would like to try Lingo, Hubberman Lab listeners in the US and UK can save 10% on a four-week plan. Just visit hellolingo.com/huberman for more information. Terms and conditions apply. Again, that's hellingo.com/huberman. Today's episode is also brought to us by Wealthfront. In today's financial landscape of constant market shifts and chaotic news, it's easy to feel uncertain about how to save and invest your money. Wealthfront is the solution that helps you take control of your money while managing risk. For nearly a decade, I've trusted Wealthfront to navigate this volatility. With the Wealthfront cash account, I can earn 3.3% annual percentage yield or APY on my cash from program banks. And I know my money is growing until I'm rea dy to spend it or invest it. One of the features I love about Wealthfront is that I have access to instant no fee withdrawals to eligible accounts 24/7. That means I can move my money where I need it without waiting. And when I'm ready to transition from saving to investing, Wealthfront lets me seamlessly transfer my funds into one of their expert-built portfolios. For a limited time, Wealthfront is offering the Huberman Lab audience an exclusive 75% APY boost over the base rate for 3 months, meaning you can get up to 4.05% 05% variable APY on up to $150,000 in deposits. Over 1 million people already trust Wealthfront to save more, earn more, and build long-term wealth with confidence. If you'd like to try Wealthfront, you can go to wealthfront.com/huberman to receive the boost offer and start earning 4.05% variable APY today. That's wealthfront.com/huberman to get started. This is a paid testimonial of Wealthfront. Client experiences will vary. Wealthfront brokerage is not a bank. The base APY is as of January 30th, 2026 and subject to change. For more information, please see the episode description. I can see the rationale for doing this in the vision domain. But has a similar thing been explored with hearing with sounds?