Lex Fridman Podcast · #416
Why LLMs Won't Lead to AGI
为何大语言模型无法通向AGI
Turing Award winner Yann LeCun systematically dismantles LLM limitations and proposes JEPA.
图灵奖得主 Yann LeCun 系统阐述 LLM 的局限、JEPA 替代架构、World Model 理论,并捍卫开源 AI 的必要性。
00:002h 41m 58s
00:00
I see the danger of this concentration of power to to proprietary AI systems as
我认为这种权力集中在专有AI系统上的危险,比其他所有问题都要大得多。而对抗这种趋势的,是那些出于安全考虑认为应该把AI系统锁起来的人——因为他们觉得把AI交到每个人手里太危险了。这会导致一个非常糟糕的未来,我们所有的信息摄入都被少数公司通过专有系统控制。我相信人性本善,所以如果AI,尤其是开源AI,能让人变得更聪明,那只会激发人类内在的善良。我认同这种感觉。对,我觉得人性
00:06
a much bigger danger than everything else what works against this is people who think
本善。事实上,很多末日论者之所以持这种观点,就是因为他们不相信人性本善。以下是Yann LeCun的对话,这是他第三次上这个播客。他是Meta的首席AI科学家、纽约大学教授、图灵奖得主,也是人工智能历史上最具开创性的人物之一。他和Meta AI一直大力倡导AI开发的开源模式,并且身体力行地开源了许多重要模型,包括Llama 2,未来还有Llama 3。同时,Yann也直言不
00:12
that for reasons of security we should keep AI systems under lock and key because
讳地批评了AI社区中那些警告AGI迫在眉睫、构成生存威胁的人。他相信AGI终有一天会被创造出来,但会是善意的,不会脱离人类控制,更不会统治或消灭全人类。在AI飞速发展的当下,这算是一个颇有争议的立场。所以看Yann在网上参与各种激烈而精彩的讨论,总是很有意思,就像我们这次对话一样。这里是Lexman播客。如果想支持我们,请查看描述里的赞助商。好了,亲爱的朋友们,欢迎Yan
00:18
it's too dangerous to put it in the hands of of everybody that would lead
n LeCun。你最近对人工智能的未来发表了一些很强硬的技术性观点——其实你整个职业生涯都这样,但最近尤其如此。你说过,自回归的LLM并不是我们通往超人类智能的正确路径。这些大型语言模型,比如GPT-4、Llama 2、即将到来的Llama 3等等,它们的工作原理是什么?为什么它们无法带我们走到终点?原因有好几个。首先,智能行为有几个关键特征:比如理解世界、理解物理世界的能
00:25
to a very bad future in which all
力;记忆和检索信息的能力,也就是持久记忆;推理能力;以及规划能力。这四点是人类和动物等智能系统或实体的核心特征。而LLM在这些方面要么完全做不到,要么只能用非常原始的方式勉强为之。它们其实并不真正……
00:28
of our information diet is controlled by a small number of uh uh companies through
我们信息摄入的很大一部分,其实是被少数几家公司通过专有系统控制的。我相信人性本善,所以如果AI,尤其是开源AI,能让人变得更聪明,那它就是在激发人类内在的善良。我也有同感。好吧,我觉得人性本善,实际上很多末日论者之所以成为末日论者,就是因为他们不相信人性本善。接下来是Yann LeCun的对话,这是他第三次上这个播客。他是Meta的首席AI科学家、纽约大学教授、图灵奖得主,也是人工智能历
00:34
proprietary systems I believe that people are fundamentally good and so if AI especially open
史上最具开创性的人物之一。他和Meta AI一直大力倡导AI开发的开源模式,并且身体力行,开源了他们许多最大的模型,包括Llama 2,以及未来的Llama 3。同时,Yann也直言不讳地批评了AI社区中那些警告AGI迫在眉睫、存在生存威胁的人。他认为AGI终有一天会被创造出来,但它会是善意的,不会脱离人类控制,更不会统治或消灭全人类。在AI飞速发展的当下,这算是一个颇有争议的立场,所以
00:40
source AI can um make them smarter it just empowers the goodness in humans so
看Yann在网上参与各种激烈而精彩的讨论也很有意思,就像我们这次对话一样。这里是Lexman播客。想支持我们,请查看描述中的赞助商。好了,亲爱的朋友们,欢迎Yann LeCun。你最近对人工智能的未来发表了一些非常强硬的、技术性的声明。实际上,纵观你的整个职业生涯都是如此,但最近尤其如此。你说过,自回归LLM并不是我们迈向超级人类智能的正确道路。这些大型语言模型,比如GPT-4,比如Ll
00:45
I I share that feeling okay I think people are Fally good uh and in
ama 2和即将推出的Llama 3等等。它们是如何工作的?为什么它们不能带我们走完全程?原因有很多。首先,智能行为有一些关键特征,比如理解世界、理解物理世界的能力,记忆和检索信息的能力,持久记忆,推理能力,以及规划能力。这些是智能系统或实体(人类、动物)的四个基本特征。LLM在这些方面要么完全做不到,要么只能以非常原始的方式做到。它们并不真正……当然,我们可以围绕它们构建一整套应用生态
00:51
fact a lot of doomers are
系统。但作为通往人类水平智能的路径,它们缺少了关键的组成部分。另外还有一点,我觉得非常有意思:这些LLM是在海量文本上训练的,基本上就是所有公开可用的文本。
00:54
doomers because they don't think that people are fundamentally good the following is a conversation
那些悲观主义者,是因为他们不相信人性本善。以下是本期播客内容,我与杨立昆的第三次对话。他是Meta首席AI科学家、纽约大学教授、图灵奖得主,也是人工智能历史上最具开创性的人物之一。他和Meta AI一直大力倡导开源AI开发,并且身体力行地开源了许多他们最大的模型,包括Llama 2,以及未来的Llama 3。此外,杨立昆一直直言不讳地批评AI社区中那些警告AGI即将带来危险和生存威胁的人。他相信AGI终有一天会被创造出来
01:00
with Yan laon his third time on this podcast he is the chief AI scientist
,但会是良性的,不会脱离人类控制,更不会统治并消灭所有人类。在AI飞速发展的当下,这恰好是一个颇具争议的立场。因此,看着杨立昆在网上展开大量激烈而引人入胜的讨论,就像我们这次对话一样,也很有趣。这里是Lexman播客。为了支持我们,请查看描述中的赞助商。现在,亲爱的朋友们,有请杨立昆。你最近对人工智能的未来发表了一些强有力的、技术性的声明——实际上,你的整个职业生涯都是如此,但最近尤其如此。你说,自回归LLM并不是我们迈向
01:07
at meta professor at NYU touring Award winner and one of the seminal figures in
超级人类智能的正确路径。这些大型语言模型,比如GPT-4,比如Llama 2和即将推出的Llama 3,它们是如何工作的?为什么它们无法带我们走完全程?原因有很多。首先,智能行为有许多特征,例如理解世界、理解物理世界的能力,记忆和检索事物的能力——持久记忆,推理能力,以及规划能力。这些是智能系统或实体(人类、动物)的四个基本特征。LLM在这些方面要么完全做不到,要么只能以非常原始的方式完成。它们实际上并不……我们当然可以
01:14
the history of artificial intelligence he and meta AI have been big proponents
围绕它们构建一整个应用生态系统,但作为通往人类水平智能的路径,它们缺少了关键组成部分。还有另一个有趣的事实:这些LLM是在海量文本上训练的,基本上涵盖了所有公开可用的文本。但你会发现,这些数据其实并没有那么多。如果你和发育心理学家聊聊,他们会告诉你,一个四岁孩子一生中清醒的时间大约有16,000小时。而在这四年里,到达那个孩子视觉皮层的信息量大约是10的15次方字节。
01:19
of open sourcing AI development and have been walking the walk by open sourcing many
开源AI开发这件事,他们一直在身体力行,开源了许多他们最大的模型,包括Llama 2,最终还有Llama 3。Yan也一直直言不讳地批评AI社区里那些警告AGI迫在眉睫的危险和存在性威胁的人。他相信AGI总有一天会被创造出来,但它会是好的,不会脱离人类控制,也不会统治并杀死所有人类。在AI快速发展的当下,这算是一个颇有争议的立场,所以看Yan在网上展开大量激烈而迷人的讨论很有趣,就像我们这次对话一样。这是Lexman播客。为了支持它,请查看描
01:26
of their biggest models including llama 2 and eventually llama 3 also Yan has been
述中的赞助商。现在,亲爱的朋友们,有请Yan Laon。你最近对人工智能的未来发表了一些强硬的、技术性的声明——实际上贯穿你的整个职业生涯,但最近也是如此。你说过,自回归LLM并不是我们朝着超人智能取得进展的方式。这些是像GPT-4、Llama 2和即将到来的Llama 3这样的大型语言模型。它们是如何工作的,为什么它们不能带我们走完全程?原因有很多。首先,智能行为有许多特征,例如理解世界、理解物理世界的能力,记忆和检索事物的能力——持久记忆,
01:33
an outspoken critic of those people in the AI Community who warn about the looming
推理的能力,以及规划的能力。这些是智能系统或实体(人类、动物)的四个基本特征。LLM无法做到其中任何一点,或者只能以非常原始的方式做到,而且它们实际上并不……我们当然可以围绕它们构建一个完整的应用生态系统,但作为通往人类水平智能的路径,它们缺少关键组件。然后还有另一个有趣的小细节或事实:这些LLM是在海量文本上训练的,基本上是所有公开可用的文本。但后来你会发现,数据其实并没有那么多。如果你和发育心理学家聊聊,他们会告诉你一个四岁孩子一生中清醒了
01:40
danger and existential threat of AGI he believes the AGI will
16,000小时。那个孩子在四年内到达视觉皮层的信息量大约是10的15次方字节。而语言——尽管我们有直觉——我们学到的大部分东西和大部分知识是通过对现实世界的观察和互动获得的,而不是通过语言。我们在生命最初几年学到的一切,当然还有动物学到的一切,都与语言无关。所以,也许可以反驳一下你说法背后的一些直觉。确实,有几个数量级的差异。
01:45
be created one day but it will be good it will not Escape human control
总有一天会被创造出来,但它是好的,不会脱离人类控制
01:51
nor will it Dominate and kill all humans at this moment of Rapid AI development
,也不会统治并杀死所有人类。在AI快速发展的当下,这
01:57
this happens to be somewhat a controversial position and so it's been fun seeing Yan
其实是个有点争议的立场,所以看到Yan在网上引发了很
02:03
get into a lot of intense and fascinating discussions online as we do in this
多激烈而有趣的讨论,就像我们这次对话一样,也挺有意思
02:09
very conversation this is the lexman podcast
的。这是Lexman播客。
02:11
to support it please check out our sponsors in the description and now dear friends
为了支持它,请查看描述里的赞助商。现在,亲爱的朋友们,这里是Yan Laon。你最近对未来人工智能发表了一些强硬的技术声明,实际上你整个职业生涯都是如此,但最近你也说过,自动攻击型LLM并不是我们通往超人类智能的正确道路。这些是像GPT-4、Llama 2和3(很快会有)这样的大型语言模型。它们是如何工作的,为什么不能带我们走完全程?原因有几个:首先,智
02:19
here's Yan laon you've had some strong statements technical statements about the future of artificial
能行为有很多特征,比如理解世界、理解物理世界的能力,记忆和检索事物的能力,持久记忆,推理能力,以及规划能力。这些是智能系统或实体(人类、动物)的四个基本特征。LLM一个都做不到,或者只能以非常原始的方式做到,而且它们并不真正……我们当然可以围绕它们构建整个应用生态系统,但作为通往人类水平智能的路径,它们缺少关键组件。还有一个有趣的点:这些LLM是在海量文本
02:26
intelligence recently throughout your career actually but recently as well uh you've said that autoaggressive
上训练的,基本上是所有公开可用的文本。但你会发现,这其实数据量并不大。如果你跟发展心理学家聊聊,他们会告诉你一个四岁孩子一生中清醒了16000小时,而那个孩子在四年内到达视觉皮层的信息量大约是10的15次方字节。而语言——尽管我们有直觉——但我们学到的大部分东西和知识,是通过观察和与现实世界的互动获得的,而不是通过语言。我们在生命最初几年学到的一切,当然还
02:33
llms are uh not the way we're going to make progress towards
有动物学到的一切,都与语言无关。所以,也许可以反驳一下你说法背后的一些直觉。确实,有几个数量级的……我们通过本质上想象一系列行动的结果来做到这一点,所以我们可能会想象,这需要与语言关系不大的心智模型。而我认为,我们大部分知识都源于与物理世界的这种互动。所以我很多更关注其他方面的同事……
02:39
superhuman intelligence these are the large language models like GPT 4 like llama 2 and
超人类智能——这些大语言模型,比如 GPT-4、Llama 2 和即将推出的 Llama 3 等
02:45
3 soon and so on how do they work and why are they not going
等——它们到底怎么工作的?为什么说它们没法带我们走到 AGI 那一步?原因有好几个。首先,智能行
02:51
to take us all the way for a number of reasons the first is that
为有很多特征,比如理解世界、理解物理世界的能力,还有记忆和检索信息的能力、推理能力、规划能力——
02:57
there is a number of characteristics of intelligent behavior for example the capacity to understand
这四个是智能系统或智能体(人类、动物)的核心特征。而 LLM 基本做不到这些,或者只能用非常原始
03:03
the world understand the physical world
的方式去模拟,而且它们其实并不真正理解。
03:05
the ability to remember and retrieve things um persistent memory the ability to reason and
记忆和检索信息的能力,嗯,持久记忆,推理能力和规划能
03:12
the ability to plan those are four essential characteristic of intelligent um systems or entities
力——这些是智能系统或实体(人类、动物)的四个基本特征
03:19
humans animals lnms can do none of those or they can only do them in
。LLM 做不到这些,或者只能以非常原始的方式做到,而
03:25
a very primitive way and uh they don't really
且它们其实并不真正具备这些能力。
03:29
understand the physical world don't really have persistent memory they can't really reason and they
理解物理世界的系统并没有真正的持久记忆,它们无法真正推理,当然也无法规划。所以如果你期望系统变得智能,却不给它做这些事情的可能性,那你就错了。这并不是说自回归语言模型没用,它们当然有用,也很有趣。我们当然可以围绕它们构建一整个应用生态系统。但作为通往人类水平智能的路径,它们缺少了关键组件。然后还有一点,我觉得非常有趣:这些LLM是在海量文本上训练的
03:36
certainly can't plan and so you know if if if you expect the system to
,基本上就是互联网上所有公开文本的总和。训练数据通常大约有10的13次方个token,每个token通常是两个字节,所以就是2乘以10的13次方字节。如果让你或我每天读八小时,光读完这些数据就需要17万年。所以看起来这些系统能积累海量的知识,对吧?但如果你跟发展心理学家聊聊,他们会告诉你一个四岁孩子一生中清醒的时间是16000小时。而那个孩子在四年内
03:42
become intelligent just you know without having the possibility of doing those things you're making
到达视觉皮层的信息量大约是10的15次方字节。你可以算一下,视神经每秒大约携带20兆字节。所以四岁孩子是10的15次方字节,对比17万年阅读量对应的2乘以10的13次方字节。这说明通过感官输入,我们看到的信息比通过语言多得多。尽管我们直觉上觉得,我们学到的大部分东西和知识来自观察和与真实世界的互动,而不是语言。我们生命头几年学到的一切,以及动物学到的
03:49
a mistake that is not to say that auto regressive LS are not useful they're
一切,都与语言无关。所以也许应该反驳一下你说法背后的直觉。确实,人类大脑接收到的数据量有几个数量级更高,而且速度更快,人类大脑能非常快速地从这些数据中学习并过滤。有人可能会说,你拿感官数据和语言对比,但语言本身已经高度压缩了,它包含的信息远多于存储它所需的字节数,如果跟视觉数据比的话。所以语言里有很多智慧,有词汇和它们组合的方式,已经包含了大量信息
03:55
certainly useful um that they're not interesting
。那么,是否可能仅凭语言本身就包含足够的智慧和知识,能够从中构建出一个世界模型、对世界的理解,以及你所说的LLM缺乏的对物理世界的理解?这在哲学家之间是个很大的争论。
03:58
that we can't build a whole ecosystem of applications around them of course we can
我们当然可以围绕它们构建一整个应用生态,但作为通往人类级别智能的路径,它们缺少了关键
04:04
but as a path towards human level intelligence they're missing essential components and then there
组件。另外还有一个很有意思的点:这些 LLM 是用海量文本训练的,基本上就是所有公开可
04:10
is another tidbit or or fact that I think is very interesting those llms are
用的文本。但如果你跟发展心理学家聊聊,他们会告诉你,一个四岁小孩一生中已经清醒了 16
04:16
trained on enormous amounts of text basically the entirety of all publicly available text
,000 个小时,这四年里进入他视觉皮层的信息量大约是 10 的 15 次方字节。
04:22
on the internet right that's typically on the order of 10 to the 13 tokens
互联网上,通常大约是10的13次方个token,
04:28
each token is typically two byes so that's two 10 to the 13 bytes as
每个token大约两个字节,所以就是2乘以10的1
04:33
training data it would take you or me 170,000 years to just read through this
3次方字节的训练数据。如果以每天八小时的速度阅读
04:39
at eight hours a day uh so it seems like an enormous amount of knowledge
,你我要花17万年才能读完,所以这些系统能积累的知
04:45
right that those systems can accumulate
识量看起来非常庞大。
04:47
um but then you realize it's really not that much data if you you talk
但后来你会发现,这其实数据量并不大。如果你去问发
04:54
to developmental psychologist and they tell you a four-year-old has been awake for 16,000 hours
展心理学家,他们会告诉你一个四岁孩子一生中已经清
05:00
in his life um and the amount of information that has uh reached the visual
醒了 16,000 小时,而在这四年里,到达那个孩
05:07
cortex of that child in four years um is about 10 to the 15 bytes
子视觉皮层的信息量大约是 10 的 15 次方字
05:13
and
节。
05:14
you can compute this by estimating that the optical nerve carry about 20 megab megabytes
但仔细想想,其实数据量并没有那么大。如果你跟发展心理学家聊聊,他们会告诉你,一个四岁孩子
05:20
per second roughly and so 10^ the 15 bytes for a four-year-old versus 2 *
醒着的时间大约有16,000小时,而这四年间到达他视觉皮层的信息量大约是10的15次方字
05:26
10 to 13 bytes for 170,000 years worth of reading what it tells you is
节。你可以估算一下,视神经每秒大约传输20兆字节,所以四岁孩子接触到的10的15次方字节
05:32
that uh through sensory input we see a lot more information than we than we
,对比17万年阅读量的2乘以10的13次方字节,这说明通过感官输入,我们获取的信息远远多
05:39
do through
于通过语言。
05:39
language and that despite our intuition most of what we learn and most of our
而语言呢?尽管我们直觉上觉得大部分知识来自语言,但实际上,我们学到
05:45
knowledge is through our observation and interaction with the real world not through language everything
的绝大多数东西——尤其是生命头几年学的一切,以及动物学的一切——都跟
05:50
that we learn in the first few years of life and uh certainly everything that
语言无关,而是通过观察和与真实世界的互动得来的。所以,也许我们应该
05:55
animals learn has nothing to do with language so it would be good to uh
挑战一下你刚才说的那种直觉。确实,我们通过想象一系列动作的结果来规划
06:00
maybe push against some of of the intuition behind what you're saying so it is
,这需要心智模型,而心智模型跟语言关系不大。我敢说,我们大部分知识
06:05
true there's several orders of
都来自与物理世界的互动。
06:07
magnitude more data coming into the human mind much faster and the human mind is
尽管直觉上可能相反,但我们大部分的学习和知识其实来自
06:12
able to learn very quickly from that filter the data very quickly you know somebody
对真实世界的观察和互动,而不是语言。我们生命最初几年学
06:17
might argue your comparison between sensory data versus language that language is already very compressed
到的一切,以及动物学到的所有东西,都与语言无关。所以,
06:22
it already contains a lot more information than the bytes it takes to store them
或许应该挑战一下你说法背后的某些直觉。确实,人类大脑
06:27
if you compare it to visual data so there's a lot of wisdom and language
接收数据的速度快得多,数量级也高几个量级,而且大脑能快
06:32
there's words and the way we
速从中学习并过滤数据。
06:34
stitch them together it already contains a lot of information so is it possible that
把它们拼接起来,它已经包含了大量信息。那么,语言本身
06:41
language alone already has enough wisdom and knowledge in there to be able to from
是否已经蕴含了足够的智慧和知识,能够从语言中构建出一
06:47
that language construct a a world model and understanding of the world an understanding of
个世界模型、对世界的理解,以及对物理世界的理解——也就
06:54
the physical world that you're saying L LMS lack so it's a big debate among
是你所说的LLM所缺乏的东西?这在哲学家之间是个很大
07:01
uh philosophers
的争论。
07:02
and also cognitive scientists like whether intelligence needs to be grounded in reality uh I'm
还有认知科学家也在讨论,比如智能是否需要扎根于现实。我显然属于那个阵营——是的,智能不可能凭空出现,它需要某种现实基础,不一定非得是物理现实,模拟的也行。但环境远比语言能表达的丰富得多,语言只是对我们感知和心智模型的一种非常粗略的表示。我的意思是,我们完成很多任务时,其实是在操作一个关于当下情境的心智模型,
07:07
clearly in the camp that uh yes uh intelligence cannot appear without some grounding in
这和语言完全无关。所有那些物理的、机械的事情,比如我们建造东西、完成一个任务模型,比如抓取物体之类的,我们会规划动作序列,本质上是通过想象一系列动作的结果来完成的。所以我们会想象,这需要心智模型,和语言关系不大。而且我认为,我们大部分知识都来源于与物理世界的这种互动。所以我很多更关注计算机视觉的同事,确实站
07:13
uh some reality doesn't need to be you know physical reality could be simulated but
在AI需要具身化这个立场上。而其他来自NLP领域或者出于其他动机的人,不一定同意这一点。哲学家们也有分歧。世界的复杂性很难想象,很难去表示那些我们在现实世界中完全习以为常、甚至不觉得需要智能的复杂事物。这就是机器人学和SMC先驱提出的莫拉维克悖论:为什么计算机似乎很容易完成像下棋、解积分这样的高复杂度任务,
07:19
um but the environment is just much richer than what you can express in language
而那些我们习以为常、每天做的事情——比如学开车、抓个物体——计算机却做不到?我们有LLM能通过律师资格考试,它们一定很聪明,但它们没法像任何一个17岁少年那样在20小时内学会开车,也没法像任何一个10岁孩子那样一次就学会收拾餐桌、把洗碗机装满。这是为什么?我们到底缺了什么?缺少哪种学习或推理架构之类的,导致
07:25
language is a very approximate representation of our percepts and our
我们无法实现L5级自动驾驶和家用机器人?一个大语言模型能否构建一个世界模型,让它知道怎么开车、怎么装洗碗机,只是目前还不会处理视觉数据?这样它就能在概念空间里运作。没错,这正是很多人在研究的方向。所以简短的回答是:不能,而且——
07:30
mental models right I mean there there's a lot of tasks that we accomplish where
有人可能会反驳,你拿感官数据和语言比,但语言本身已经是高度压缩的,它包含的信息远多
07:35
we manipulate uh a mental model of uh of the situation at hand and that
于存储它所需的字节数,比如跟视觉数据比。语言里有很多智慧,词汇以及我们组合它们的方式
07:41
has nothing to do with language everything that's physical mechanical whatever when we build something
已经蕴含了大量信息。那么,是否可能仅凭语言就包含足够的智慧和知识,从而构建出一个世
07:47
when we accomplish a task model task of you know grabbing something Etc we plan
界模型和对世界的理解——包括你所说的LLM缺乏的对物理世界的理解?这在哲学家之间是个
07:53
or action
很大的争论。
07:54
sequences and we do this by essentially Imagining the result of the outcome of sequence
我们通过本质上想象一系列行动的结果来做
07:59
of actions so we might imagine and that requires mental models that don't have much
到这一点。所以我们可能会想象,而这需要与
08:05
to do with language and that's I would argue most of our knowledge is derived
语言关系不大的心智模型。我认为我们大部分
08:11
from that interaction with the physical world so a lot of a lot of my
知识都来自与物理世界的这种互动。所以我
08:16
my colleagues who are more uh interested in things like
很多更关注这类事情的同事……
08:20
computer vision are really on that camp that uh AI needs to be embodied essentially
计算机视觉那派的人确实认为,AI本质上需要具身化,而来自NLP领域或者其他动机的人不一定同意这点,哲学家们也有分歧。世界的复杂性很难想象,很难去表征那些我们在现实世界中完全视为理所当然的复杂事物,我们甚至没意识到这些需要智能——这就是机器人学和SMC先驱提出的Moravec悖论:为什么计算机似乎很容易完成像下棋、解积分这样的高层次复杂任务,而我们每天视为理所当然的事情,比如学开车、抓取物体,计
08:27
and then other people coming from the NLP side or maybe you know some some
算机却做不到。我们有LLM能通过律师资格考试,所以它们一定很聪明,但它们没法像任何17岁少年那样在20小时内学会开车,也没法像任何10岁孩子那样一次就学会收拾餐桌和装洗碗机。这是为什么?我们到底缺了什么?缺了哪种学习或推理架构,才让我们无法实现L5自动驾驶和家用机器人?大型语言模型能构建一个知道如何开车和装洗碗机的世界模型吗?还是说它只是目前不知道如何处理视觉数据,所以它能操作概念空间?没错,很
08:34
other uh motivation don't necessarily agree with that um and philosophers are split as well
多人正在研究这个。所以简短的回答是:不能。训练视觉系统的方法有很多,有监督、半监督、自监督等等,这些方法能把任何图像转化为高层表征,基本上是一串token,和典型LLM输入的token非常相似,然后你只需把这些token连同文本一起喂给LLM,期望它在训练过程中学会利用这些表征来辅助决策。这方面的研究已经进行了相当长的时间,现在你也能看到这类系统——有些LLM确实有视觉扩展,但它们基本上都是h
08:41
uh and the U the complexity of the world is
ack,因为这些系统并不是端到端训练来真正理解世界的,比如它们没有用视频训练,至少目前还不理解直观物理。所以你不认为直观物理、关于物理空间的常识推理有什么特别之处?对你来说,这是LLM目前根本无法跨越的巨大鸿沟?我们无法用今天这种类型的LLM做到这一点,原因有很多。
08:46
hard to um hard to imagine it you know it's hard to represent uh all
我们通过想象一系列行动的结果来做到这
08:51
the complexities that we take completely for granted in the real world that we don't
一点。我们会进行想象,而这需要与语言
08:55
even imagine require intelligence right this is the old marac Paradox from the pioneer of
关系不大的心智模型。我认为,我们大部
09:00
Robotics and SMC we said you know how is it that with computers it seems
分知识都来源于与物理世界的这种互动。
09:05
to be easy to do high Lev complex tasks like playing chess and solving
所以,我很多更关注这类问题的同事……
09:10
integrals and doing things like that whereas the thing we take for granted that we
积分和做这类事情,而我们每天习以为常的
09:15
do every day um like I don't know learning to drive a car or you
事情,比如我不知道,学开车或者抓个东西,
09:21
know grabbing an object we can do as computers um and you know we have
计算机却做不到。我们有LLM能通过律师资
09:27
llms that can pass pass the bar exam so they must be smart but then
格考试,所以它们一定很聪明,但它们在2
09:33
they can't learn to drive in 20
0小时内学不会开车。
09:35
hours like any 17y old they can't learn to clear out the dinner table and
很难想象,你知道,很难去表征我们在现实世界中完全视为理所当然的所有复杂性,我们甚至没意识到这些需要智能。这就是机器人学和
09:41
F of the dishwasher like any 10-year-old can learn in one shot um why is
SMC先驱提出的莫拉维克悖论:为什么计算机似乎很容易完成像下棋、解积分这样的高复杂度任务,而像学开车、抓取物体这样我们每
09:47
that like you know what what are we missing what what type of learning or
天习以为常的事情,计算机却做不到?我们有能通过律师资格考试的LLM,它们一定很聪明,但它们无法像任何17岁少年那样在20
09:53
or reasoning architecture or whatever are we missing that um um basically prevent us from
小时内学会开车,也无法像任何10岁孩子那样一次就学会收拾餐桌和装满洗碗机。这是为什么?我们到底缺少了什么?缺少了哪种学习
09:58
from you know having level five sing
或推理架构,导致我们无法实现L5级自动驾驶和家用机器人?
10:01
Cars and domestic robots can a large language model construct a world model that does
我很多同事更关注汽车和家用机器人这类东西。一个大
10:06
know how to drive and does know how to fill a dishwasher but just doesn't
语言模型能构建一个世界模型吗?一个知道怎么开车、
10:11
know how to deal with visual data at this time so it it can operate
怎么装洗碗机,但只是目前不知道怎么处理视觉数据的
10:17
in space of Concepts so yeah that's what a lot of people are working on
模型?它可以在概念的层面运作。没错,很多人正在研
10:22
so the answer the short answer is no and the
究这个。所以简短的回答是:不能。
10:25
more complex sensor is you can use all kind of tricks to get uh uh
更复杂的传感器是,你可以用各种技巧让LLM去消化图像、视频甚至音频的视觉表征。一个经典做法是,你先用某种方式训练一个视觉系统——我们有多种训练视觉系统的方法,比如监督学习、半监督学习、自监督学习等等——这些方法能把任何图像转换成高层表征,本质上就是一系列token,跟典型LLM输入的token非常相似。然后你把这些token连同文本一起喂给LLM,在训练过程中让LLM学会利用这些表征来辅助决策。其实这方面的研究
10:32
an llm to basically digest U visual representations of representations of images uh or video
已经做了很久了,现在你也能看到这类系统——有些LLM确实有视觉扩展功能,但它们本质上都是hack,因为这些系统并不是端到端训练来真正理解世界的。比如它们没有用视频训练过,至少目前还不懂直观物理学。所以你觉得直观物理学、关于物理空间的常识推理、物理现实这些东西,对LLM来说是个巨大的跨越,它们根本做不到?我们目前使用的这类LLM确实做不到,原因有很多,但最主要的是LLM的训练方式:你拿一段文本,去掉其中一些词,把它
10:38
or audio for that matter um and uh a classical way of doing this is
们遮住,换成空白标记,然后训练一个巨大的神经网络去预测缺失的词。如果你用特定方式构建这个神经网络,让它只能看要预测词左边的词,那你就得到了一个系统,它本质上就是在预测文本中的下一个词。所以你可以给它一段文本作为prompt,让它预测下一个词——它永远无法精确预测,所以它会生成一个概率分布,覆盖字典里所有可能的词。实际上它预测的不是词,而是子词单位的token,这样处理预测的不确定性就简单了,因为字典里只有有限数量
10:45
uh you train a vision system in some way and we have a number
的词,你只需要计算一个分布。然后系统从这个分布里选一个词——当然,概率高的词被选中的机会更大——你从这个分布里采样来实际生成一个词,再把这个词移入输入,这样系统就能预测第二个词了。重复这个过程,把新词移入输入,以此类推,这就叫自回归预测,所以这些LLM应该叫自回归LLM,但我们通常就叫它们语言模型。这种过程和人类说话前的思考过程是有区别的——你和我都是双语者,我们在说话之前会先想一想。
10:51
of ways to train Vision systems either supervised semisupervised self superise all kinds of different
一个大型语言模型能否构建一个世界模型,既知道如何
10:58
ways uh that will turn any image into high level representation basically a list of
开车,也知道如何装满洗碗机,只是目前还不会处理视觉
11:05
tokens that are really similar to the kind of tokens that uh typical llm takes
数据?所以它可以在概念空间里运作。没错,这正是很
11:12
as an input and then you just feed that to the llm in
多人在研究的方向。所以简短的回答是“不能”。
11:18
addition to the text and you just expect LM to kind of uh you know
除了文本之外,你只是期望LM在训练过程中能够利用这些表征来帮助做决策。其实这方面的研究已经做了很久了,现在你也看到这些系统了——有些LLM确实有视觉扩展,但本质上都是hack,因为它们并不是端到端训练去真正理解这个世界的。比如它们没有用视频训练过,至少目前还不懂直观物理。所以你觉得直
11:22
during training to kind of be able to uh use those representations to help make
观物理、常识推理、物理空间和物理现实这些东西,对LLM来说是一个巨大的跨越,它们根本做不到?我们目前用的这种LLM确实做不到,而且有很多原因,但最主要的原因是LLM的训练方式:你拿一段文本,去掉一些词,把它们遮住,换成空白标记,然后训练一个巨大的神经网络去预测那些缺失的词。如果你把这
11:27
decisions I mean there been work along those line for for quite a long time
个神经网络设计成只能看要预测的词左边的词,那你就得到了一个系统,它基本上就是在预测文本里的下一个词。所以你可以给它一段文本作为prompt,让它预测下一个词——它永远无法精确预测,所以它会生成一个概率分布,覆盖词典里所有可能的词。实际上它预测的不是词,而是token,也就是子词单元,这
11:31
um and now you see those systems right I mean there are llms that can
样处理预测的不确定性就简单了,因为词典里只有有限数量的词,你只需要算一个分布。然后系统会从这个分布里选一个词,当然概率高的词被选中的几率更大,所以你是从这个分布里采样来生成一个词,再把这个词移进输入里,这样系统就能继续预测第二个词。重复这个过程,把生成的词不断移进输入,这就叫自回归预
11:36
that have some Vision extension but they're basically hacks in the sense that um those
测,所以这些LLM应该叫自回归LLM,但我们通常就叫它们LM。这种过程和我们在说话前生成词的过程是有区别的——你和我都是双语者,我们在说话前会先思考,思考先于语言,然后再映射到语言上。对吧?很多思考确实是这样。那你的思考在法语和英语里是一样的吗?差不多,差不多。还是说这取决于你有多灵活
11:40
things are not like trained end to end to to handle to really understand
?比如如果有个概率分布……其实要看是什么类型的思考。如果是讲双关语,我法语比英语好得多。但你的幽默有没有一个抽象表征?比如你发推文的时候,有时候推文有点辛辣,在你把它映射成英语之前,你脑子里是不是有一个抽象的表征?确实有一个抽象的表征,就是想象读者对那段文字的反应。
11:44
the world they're not trained with video for example uh they don't really understand intuitive
汽车和家用机器人——大型语言模型能
11:49
physics at least not at the moment so you don't think there's something special to
构建一个世界模型吗?这个模型知道怎么
11:53
about intuitive physics about sort of Common Sense reasoning about the physical space about physical
开车,也知道怎么装洗碗机,只是目前
11:57
reality that's that to you is a giant leap that llms are just not able
还不会处理视觉数据?所以它只能在概念
12:02
to do we're not going to be able to do this with the type of
空间里运作。没错,这就是很多人正在研
12:06
llms that we are uh working with today and there's a number of reasons for
究的方向。所以答案,简短的回答是:
12:10
this but
不能。
12:11
uh the main reason is the way llm LMS are trained is that you you
uh 主要原因在于 LLM 的训练方式,就是你拿一段文本,删掉其中一些词,把它们遮住,换成空白标记,然后训练一个巨大的神经网络去预测那些缺失的词。如果你用特定方式构建这个神经网络,让它只能看要预测的词左边的那些词,那你得到的就是一个系统,它基本上是在预测文本里的下一个词,对吧?所以你可
12:15
take a piece of text you remove some of the words in that text you
以给它一段文本、一个 prompt,让它预测下一个词,它永远没法精确预测出下一个词,所以它要做的是生成一个概率分布,覆盖字典里所有可能的词。实际上它预测的不是词,而是 token,这些 token 有点像子词单元,这样处理预测中的不确定性就简单了,因为字典里只有有限数量的可能词,你只需要
12:20
Mass them you replace by replace them by blank markers and you train a gtic
算出一个分布。然后系统会从这个分布里选一个词,当然概率高的词被选中的机会更大,所以你就从这个分布里采样,实际生成一个词,然后把这个词移进输入里,这样系统就能接着预测第二个词。一旦你这么做,把它移进输入,以此类推,这就叫自回归预测,所以这些 LLM 应该叫自回归 LLM,但我们就叫它们 L
12:24
neural net to predict the words that are missing uh and if you build this
M。这种过程和另一种过程有区别,就是当你在说话时——你和我都是双语者——我们会在说话前先思考,思考那些先于语言的东西,然后映射到语言上,对吧?对很多思考来说确实如此,很明显,就像你说的,你的思考在法语里和英语里是一样的,差不多吧,差不多。或者说这有多灵活?比如如果有一个概率分布……嗯,这
12:28
neural net in a particular way so that it can only look at u words
取决于什么类型的思考。如果是像讲双关语,我在法语里比英语强多了。不,但更抽象地说,双关语有抽象表征吗?比如你的幽默是抽象的吗?当你发推文,你的推文有时候有点辛辣,在你把它映射成英语之前,你大脑里有没有一个抽象的推文表征?有一个抽象的表征,就是想象读者对那段文本的反应。或者你从笑声开始,然
12:33
that are to the left of the one is trying to predict then what you
后想办法实现它,或者想象你想引发的反应,然后想办法说出来,好引起那种反应,但这其实离语言很近。但想想数学概念,或者想象你想用木头做的东西,那种思考跟语言完全没关系,真的,不一定有内心独白用某种特定语言,你就是在想象那个东西的心理模型。我的意思是,如果我让你想象这个水瓶旋转 90 度会是什
12:37
have is a system that
么样子,那跟语言完全无关。所以很明显,存在一个更抽象的表示层次,我们大部分思考都在那个层次进行。
12:38
basically is trying to predict the next word in a text right so then you
基本上就是试图预测文本中的下一个词,对吧?所以你
12:43
can feed it um a text a prompt and you can ask it to predict
可以给它一段文本作为 prompt,让它预测下一个
12:47
the next word it can never predict the next word exactly and so what it's
词。它永远无法精确预测下一个词,所以它会生成一个
12:52
going to do is uh produce a probability distribution over all the possible words in
概率分布,覆盖字典里所有可能的词。实际上它预测的不
12:56
your dictionary in fact it doesn't predict words it predicts tokens that are kind of
是词,而是 token,这些 token 有点像
13:01
subword units and so it's easy to handle the uncertainty in the prediction there because
子词单元,这样处理预测中的不确定性就很容易,因为只
13:05
there's only a finite number of
有有限数量的可能性。
13:07
possible words in the dictionary and you can just compute a distribution over them um
它们没有用视频训练过,比如,它们并不真正理解
13:12
then what you what the system does is that it it picks a word from
直观物理,至少目前是这样。所以你不觉得直观物
13:17
that distribution of course there's a higher chance of picking words that have a higher
理有什么特别之处吗?那种关于物理空间、物理现实
13:21
probability within that distribution so you sample from that distribution to actually produce a word
的常识推理,对你来说是个巨大的飞跃,LLM就
13:26
and then you shift that word into the input and so that allows the system
是做不到?我们无法用今天正在使用的这种LLM来
13:31
not to predict the second word right and
实现这一点,原因有很多。
13:33
once you do this you shift it into the input Etc that's called Auto regressive
一旦你完成这个步骤,你就把它输入到输入层等等,这
13:39
prediction and which is why those llms should be called Auto regressive llms uh but
个过程叫做自回归预测,这也是为什么那些LLM应该被
13:45
we just call them LMS and there is a difference between this kind of process
称为自回归LLM,但我们通常就直接叫它们LM。这
13:50
and a process by which before producing a word when you talk when you and
种过程和另一种过程是有区别的——在你我说话时,在说
13:56
I talk you and I are bilinguals M we think think about what
出一个词之前,我们作为双语者,会先思考一下。
14:01
we're going to say and it's relatively independent of the language in which we're going
我们要说的是,这种思考方式相对独立于我们使用的语言。比如,当我们讨论一个数学概念时,我们进行的思考和准备给出的答案,并不取决于我们是用法语、俄语还是英语来表达它。chsky 翻了个白眼,但我理解你的意思。你是说,在语言之前存在一个更大的抽象层,然后映射到语言上,对吧?对我们很多思考来
14:06
to say when we when we talk about like uh I don't know let's say
说,这确实是成立的。这很明显吗?比如你说你在法语和英语中的思考是一样的?差不多,差不多。或者这取决于你有多灵活?如果有一个概率分布……嗯,这取决于思考的类型,对吧?如果只是像讲双关语,我在法语上比英语强得多。不,但双关语也有一个抽象表示吧?比如你的幽默感是抽象的吗?当你发推文,有时内容
14:11
a mathematical concept or something the kind of thinking that we're doing and the answer
有点辛辣时,在你大脑里是否存在一个抽象表示,然后再映射到英语?确实有一个抽象表示,那就是想象读者对那段文字的反应。或者你从笑声开始,然后想办法实现它,或者想象你想引发的反应,再琢磨怎么说才能引起那种反应。但这已经很接近语言了。但想想一个数学概念,或者想象你想用木头做的东西,那种思考方式
14:15
that we're planning to produce is not linked to whether we're going to see it
跟语言完全无关。真的,你不一定会有任何特定语言的内心独白。你是在想象事物的心智模型。比如,如果我让你想象这个水瓶旋转90度后会是什么样子,那跟语言毫无关系。所以很明显,存在一个更抽象的表示层,我们在其中进行大部分思考,并规划我们要说的话——如果输出是说出的话语,而不是肌肉动作的话。我们
14:20
in French or Russian or English chsky just rolled his eyes but I understand so
在产出之前先规划答案,而LLM不这么做,它们只是本能地一个词接一个词地生成。这有点像潜意识行为,比如你分心时,正全神贯注做某事,有人过来问你一个问题,你随口回答了,没时间思考答案,但答案很简单,所以你不需要在意,就自动回应了。LLM差不多就是这样,对吧?它并不真正思考,而是检索,因为它
14:25
you're saying that there's a a bigger abstraction that repes that's
积累了大量知识,所以能提取一些东西,但它只是一个个地吐出token,而不规划答案。但你说得好像一个token接一个token的生成注定是简单的,但如果世界模型足够复杂,那一个token接一个token的方式……
14:29
uh that goes before language yeah maps onto language right it's certainly true for a
但基本上,它就是在预测文本中的下一个词,对吧?所以你可以给它一段文本作为提示,让它预测下一个词。
14:33
lot of thinking that we that we do is that obvious that we don't like
它永远无法精确预测下一个词,所以它会做的是,生成一个概率分布,覆盖字典里所有可能的词。实际上它预测
14:38
you're saying your thinking is same in French as it is in English yeah pretty
的不是词,而是token,这些token有点像子词单元。这样处理预测中的不确定性就很容易,因为字典
14:43
much yeah pretty much or is this like how how flexible are you like if
里只有有限数量的可能词,你只需要计算一个分布。然后系统会做的是,从这个分布中选一个词,当然,概率高
14:48
if there's a probability distribution well it it depends what kind of thinking right if
的词被选中的机会更大。所以你从这个分布中采样,实际生成一个词,然后把这个词移入输入,这样系统就不会
14:53
it's just uh
预测第二个词,对吧?
14:54
if it's like producing puns I get much better in French than English about that
如果涉及双关语,我在法语上比英语表现好得多。
14:59
no but so worse is an abstract representation of puns like is your humor an
不,但双关语有抽象表征吗?比如你的幽默感是抽
15:04
abstract like when you tweet and your tweets are sometimes a little bit spicy uh
象的吗?当你发推文,有时内容有点辛辣时,在你
15:09
what's is there an abstract representation in your brain of a tweet before it maps
大脑里,在映射成英语之前,有没有一个推文的抽
15:14
onto English there is an abstract representation of uh Imagining the reaction of a reader
象表征?确实有一个抽象表征,是想象读者对那段
15:18
to that uh text
文字的反应。
15:20
or you start with laughter and then figure out how to make that happen or
或者你先笑出来,然后再琢磨怎么让这事儿发生,或者先想好你想引发什
15:24
figure out like a reaction you want to cause and and then figure out how
么反应,再琢磨怎么说才能引发那个反应——但这其实已经很接近语言了
15:28
to say it right so that it causes that reaction but that's like really close
。可你想想一个数学概念,或者想象一下你想用木头做点什么东西,那种
15:32
to language but think about like a math mathematical concept or um you know imagining
思考方式跟语言其实完全没关系。它并不是说你在脑子里用某种特定语言
15:37
you know something you want to build out of wood or something like this right
自言自语,而是在构建关于那个东西的心理模型。比如我让你想象这个水
15:41
the kind of thinking you're doing has absolutely nothing to do with language really like
瓶旋转90度后会是什么样子,这跟语言一点关系都没有。所以很明显,
15:45
it's not like you have necessarily like an internal monologue
我们大多数思考是在一个更抽象的层面上进行的。
15:48
in any particular language you're you're you know imagining mental models of of the thing
如果是要讲双关语,我觉得我在法语里比英语里表现好得多
15:53
right I mean if I if I ask you to like imagine what this uh
,但也不一定。所以双关语是不是一种抽象的表征?比如你
15:59
water bottle will look like if I rotate it 90 degrees um that has nothing
发推文时,有时候内容有点辛辣,那在你脑子里,在映射成
16:04
to do with language and so uh so clearly there is you know a more
英语之前,是不是存在一个抽象的推文表征?确实存在一个
16:09
abstract level of representation uh in which we we do most
抽象的表征,是想象读者对那段文本的反应。
16:13
of our thinking and we plan what we're going to say if the output is
我们思考的时候,会先计划好要说什么,如果输出是说出来的话,而不是像肌肉动作那种输出对吧。我们在说出答案之前会先计划好,但语言模型不是这样,它们只是本能地一个词接一个词地往外蹦。这有点像那种下意识的动作,比如你正分心做别的事,完全集中注意力,突然有人过来问你一个问题,你随口就回答了,没时间细想,但因为问题简单,你也不需要太在意,就自动回应了。语言模型差不多就是这样,它不会真正去思
16:19
is you know uttered words as opposed to an output being uh you know muscle
考,而是检索——因为它积累了大量的知识,所以能提取出一些东西,但它只是一个个地吐出 token,不会提前规划答案。但你这么说,好像一个 token 接一个 token、一次只生成一个 token 的方式注定会很简陋,但如果世界模型足够复杂,能对世界有深刻理解呢?首先,能不能通过预测来构建这种模型?答案很可能是能。那能不能通过预测文字来构建?答案很可能是否定的,因为语言的信息量很贫
16:25
actions right um we we plan our answer before we produce it and LMS don't
乏,或者说带宽很低。构建世界模型意味着观察世界,理解世界为什么会这样演变,然后世界模型还有一个额外功能,就是能预测如果你采取某个行动,世界会如何演变。所以世界模型是这样的:这是我在时间 t 对世界状态的认知,这是我可能采取的一个行动,那么预测在 t+1 时刻世界会是什么状态。这个状态不需要代表世界的全部,只需要足够相关于这个行动的规划就行,不一定需要所有细节。现在问题来了,这个想
16:32
do that they just produce one word after the other instinctively if you want it's
法其实已经流传很久了,在 Fair,我和一些同事尝试了大概十年,但就是做不到。你不能像语言模型那样用同样的技巧,因为就像我说的,你无法精确预测一个词序列后面会跟哪个词,我们只能预测词的概率分布。如果换成视频,你就得预测视频中所有可能帧的概率分布,但我们其实不知道怎么做才对——我们不知道如何在高维连续空间中用有用的方式表示概率分布,这就是主要问题,也是为什么我们能做语言模型的原因。
16:38
like it's a bit like the you know subconscious uh actions where you don't like
这有点像那种下意识的动作——你正分心做别的事,完全专注着,突然有人过来问
16:42
you're distracted you're doing something you're completely concentrated and someone comes to you and you
你一个问题,你随口就回答了,根本没时间细想,但答案很简单,你不需要动脑子,
16:47
know asks you a question and you kind of answer the question you don't have
几乎是自动回应。LLM就是这么干的,对吧?它其实不是真的在思考,而是因为
16:51
time to think about the answer but the answer is easy so you don't need
积累了大量的知识,所以能检索出一些东西来,但它只是一个个地往外吐token
16:56
to pay attention you sort of respond automatically that's kind of what an llm does
,根本不会事先规划答案。但你这么一说,好像一个个token往外蹦、一次只生
17:00
right it doesn't think about it sensor really uh it retrieves it because
成一个token,就注定会很简陋。可如果世界模型足够复杂的话……
17:04
it's accumulated a lot of knowledge so it can retrieve some some things but it's
它积累了大量的知识,所以能检索到一些东西,但它只会一个 to
17:10
going to just spit out one token after the other without planning the answer but
ken 接一个 token 地吐出来,而不去规划答案。但你说得
17:17
you're making it sound just one token after the other one token at a time
好像一个 token 接一个 token、一次生成一个 tok
17:23
generation is uh bound to be simplistic but if the world model is sufficiently sophisticated
en 的方式注定是简单的,但如果世界模型足够复杂,那一个 to
17:29
that one
ken……
17:30
token at a time the the most likely thing it generates is a sequence of
一次生成一个token,它最可能生成的东西是一串token,这本身会是一件非常深刻的事情。但前提是,这些系统实际上拥有一个内部的世界模型。所以这其实回到了,我认为最根本的问题是:你能不能构建一个非常完整的世界模型?不一定是完整的,但至少是一个对世界有深刻理解的模型。那么,首先,你能通过预测来构建它吗?答案很可能是“能”。你能通过预测单词来构建它吗?答案很可能是“不能”,因为语言的信息量非常贫乏,或者说带宽很低,如果你愿意这么理解
17:36
tokens is going to be a deeply profound thing okay but then that assumes that
的话。那里根本没有足够的信息。所以构建世界模型意味着观察世界,并理解世界为什么会以它现在的方式演化。然后,世界模型还有一个额外的组成部分,就是能够预测,在你可能采取某个行动之后,世界会如何演化。所以,世界模型真正做的事情是:这是我在时间t对世界状态的认知,这是我可能采取的一个行动,那么预测在时间t+1时世界状态会是什么。这个t+1时的世界状态不需要代表世界的一切,它只需要代表足够多、与这个行动的规划相关的信息,但不一定需要所有细节
17:43
those systems actually possess an internal World model so it really goes to the I
。现在问题来了,你无法用生成模型做到这一点。生成模型在视频上训练过,我们尝试这样做已经有十年了。你给系统一段视频,然后让它预测视频的剩余部分,基本上就是预测接下来会发生什么,一帧一帧地预测,就像自回归LLM做的事情一样,只不过换成了视频。要么一次一帧,要么一次一组帧。但,是的,一个大型视频模型,如果你愿意这么理解的话,这个想法已经流传很久了。在FAIR,我和一些同事尝试这样做大概有十年了。但你做不到,你无法像对语言模型那样使用同样
17:49
I think the fundamental question is can you build a a really complete World model
的技巧,因为,你知道,就像我说的,LLM无法精确预测一个词序列后面会跟哪个词,我们只能预测词的概率分布。现在到了视频,你不得不做的是预测所有可能帧的概率分布,而我们并不知道如何正确地做到这一点。我们不知道如何以有用的方式来表示高维连续空间上的概率分布。这就是主要问题所在,也是我们能做到这一点的原因,因为世界在信息层面上比文本要复杂和丰富得多。文本是离散的,视频是高维且连续的,里面有大量细节。所以,如果我拍一段这个房间的视频,而视频
17:56
not complete
是摄像机在摇拍,我根本无法预测当我摇拍时房间里会出现什么。
17:57
but a uh one that has a deep understanding of the world yeah so can
但一个对世界有深刻理解的模型,首先能靠预测来构建吗?答案很可能是“能”。那能不能靠预测单词来构建呢?答案很可能是“不能”,因为
18:03
you build this first of all by prediction right and the answer is probably yes
语言的信息量太少了,或者说带宽太低了。要构建世界模型,就得观察世界,理解世界为什么会这样演变,然后世界模型还有一个额外功能:能
18:09
can you predict can you build it by predicting words and the answer is most
预测如果你采取某个行动,世界会怎么演变。所以世界模型的核心就是:这是我在时间t对世界状态的认知,这是我可能采取的行动,那么预测在
18:15
probably no because language is very poor in terms or weak or low bandwidth if
t+1时刻世界会变成什么样。这个状态不需要描述世界的全部,只需要包含跟这个行动规划相关的信息就够了,不一定要所有细节。现在问题
18:21
you want there's
来了——你没法做到这一点。
18:22
just not enough information there so building World models means observing the world and uh
在任何特定语言里,你其实是在想象事物的心智模型,
18:28
understanding why the world is evolving the way the way it is and then uh
对吧。我的意思是,如果我让你想象这个水瓶旋转90度
18:34
the the extra component of a world model is something that can predict how the
后会是什么样子,那跟语言完全没关系。所以很明显,存
18:40
world is going to evolve as a consequence of an action you might take right
在一个更抽象的表征层面,我们大部分思考都在那个层
18:47
so what
面进行。
18:47
model really is here is my idea of the state of the world at time
这个想法其实已经流传很久了。在
18:52
te here is an action I might take what is the predicted state of the
FAIR,我和一些同事大概尝试
18:57
world at Mt plus1 now that state of the world doesn't does not need to
了十年,但你就是没法像搞语言模
19:02
represent everything about the world it just needs to represent enough that's relevant for this
型那样玩同样的把戏。因为LLM
19:07
planning of of the action but not necessarily all the details now here is the
,就像我说的,你没法精确预测一
19:12
problem um you're not going to be able to do this
个词序列后面会跟哪个词。
19:15
with generative models so genery model has trained on video and we've tried to do
生成式模型,generative model 嘛,它是在视频上训练的,我们尝试这件事已经十年了。你给系统看一段视频,然后让它预测视频剩下的部分,基本上就是逐帧预测接下来会发生什么,跟自回归 LLM 做的事情一样,只不过对象是视频——要么一帧一帧来,要么一组一组来。嗯,但一个大规
19:19
this for 10 years you take a video show a system a piece of video
模视频模型,如果你想做的话,这个想法其实已经流传很久了。在 Fair,我和一些同事尝试做这件事大概有十年了。但你不能直接套用语言模型的那套方法,因为,你知道,LLM 就像我说的,你没法精确预测一个词序列后面会跟哪个词,我们只能预测词的概率分布。现在换到视频,你需要做的是预测所有
19:24
and then ask you to predict the reminder of the video basically predict what's going
可能帧的概率分布,而我们其实不知道该怎么正确地做到这一点。我们不知道如何用有用的方式来表示高维连续空间上的分布,这就是主要问题所在。我们能做这件事的原因在于,世界在信息层面上比文本复杂和丰富得多。文本是离散的,视频是高维且连续的,细节非常多。所以,如果我拍一段这个房间的视频,镜头
19:28
to happen one frame at a time do the same thing as sort of the
在移动,嗯,我根本没法预测随着镜头移动房间里会出现什么。这里涉及到一个叫 latent variable 的东西,它被输入到一个神经网络中,应该要表示你还没感知到的所有世界信息,用来增强系统,让它在预测像素时做得更好,包括地毯的精细纹理、沙发上的细节、墙上的画等等。嗯,这基本上
19:32
autoaggressive llms do but for video right either one FR at a time or a
完全失败了。我们试过很多方法:直接上神经网络,试过 GAN,试过 VAE,各种正则化的自编码器,试过很多。我们还试过用这类方法来学习图像或视频的良好表征,然后作为输入给图像分类系统之类的,结果也基本是失败的。所有试图从损坏版本预测图像或视频缺失部分的系统,基本上都失败了。就是说
19:36
group of friends at a time um but yeah uh a large video model if
,你拿一张图像或一段视频,以某种方式损坏或变换它,然后尝试从损坏版本重建完整的视频或图像,并希望系统内部能发展出良好的图像表征,用于物体识别、分割等等——这基本上完全失败了。但这个方法对文本却非常有效,这就是语言模型所用的原理。那么失败到底出在哪里呢?就在于很难形成一个好的表征。
19:41
you want uh the idea of of doing this has been floating around for a
你想做这件事的想法已经流传很久了。在
19:47
long time and at at Fair uh some colleagues and I have been trying to
FAIR,我和一些同事大约十年前就开始尝
19:52
do this for about 10 years um and you can't you can't really do the
试。你不能对 LM 用同样的技巧,因为你
19:58
same trick as with LM because uh you know llms as I said you can't
知道,LLM 就像我说的,你无法精确预测
20:04
predict exactly which word is going to follow a sequence of words we can
一个词序列后面会跟哪个词。我们可以……
20:09
predict the distribution over words now if you go to video what you would have
现在你要预测的是词的概率分布。如果换成视频,你需要预测视频里所有可能帧的概率分布
20:15
to do is predict the distribution over all possible frames in a video and we
,而我们其实不知道怎么做才对——我们不知道如何在高维连续空间里用有用的方式表示概
20:20
don't really know how to do that properly uh we we do not know how
率分布,这就是主要问题所在,也是为什么我们能做所谓“潜变量”的原因。潜变量被输入到
20:25
to represent distributions over High dimensional continuous spaces in ways that are useful uh and
一个神经网络里,它应该代表你还没感知到的、关于世界的所有信息,你需要用它来增强系
20:30
and that's that there lies the main issue and the reason we can do
统,让预测像素的效果更好,包括地毯的精细纹理、沙发上的纹理、还有画上的细节。
20:35
this is because the world is incredibly more complicated and richer in terms of information
这是因为现实世界在信息层面远比文本复杂和
20:40
than than text text is discret video is high dimensional and continuous a lot of
丰富得多——文本是离散的,而视频是高维且
20:45
details in this um so if I take a a video of this room uh
连续的,里面有大量细节。所以,如果我拍一
20:50
and the video is you know a camera panning around MH um there is no
段这个房间的视频,摄像机在来回移动,那我根
20:55
way I can predict everything that's going to be in the room as I pan
本不可能预测出随着镜头转动,房间里会出现
21:00
around the
什么。
21:01
system cannot predict what's going to be in the room as the camera is panning
系统无法预测当镜头平移时房间里会出现什么——它可能会预测这是一个有灯和墙的房间之类的东西,但它无法预测墙上的画长什么样,沙发的纹理是什么,更不用说地毯的纹理了,所以我根本不可能预测所有这些细节。所以处理这个问题的一种方式——我们其实已经研究很久了——就是用一个带有所谓 la
21:05
maybe it's going to predict this is this is a room where there's a light
tent variable 的模型,这个 latent variable 被输入到一个 neural net 里,它应该代表你尚未感知到的关于世界的所有信息,用来补充系统,让它在预测像素时表现更好,包括地毯的精细纹理、沙发的纹理、墙上的画等等。嗯,这个方法基本上完全失败了,我
21:09
and there is a wall and things like that it can't predict what the painting
们试过很多东西:我们试过直接的 neural nets,试过 GANs,试过各种正则化的 autoencoders,试过很多方法,也试过用那些方法来学习图像或视频的良好表征,然后把这些表征作为输入给图像分类系统之类的。嗯,那也基本失败了。所有那些试图从图像或视频的损坏版本中
21:13
on the wall looks like or what the texture of the couch looks like certainly
预测缺失部分的系统——比如拿一张图像或一段视频,损坏它或以某种方式变换它,然后尝试从损坏版本重建完整的视频或图像,并希望系统内部能发展出良好的图像表征,用于物体识别、分割等等——这些基本上都完全失败了。这个方法对文本效果很好,这就是 LLMs 所用的原理对吧。那么失败到底在哪
21:17
not the texture of the carpet so there's no way I can predict all those
里呢?就是很难形成图像的良好表征,一个能包含图像中所有重要信息的良好 embedding,尤其是在图像与图像之间的一致性上,比如构成视频的图像。如果我们做一个你所有失败方式的集锦,那会是什么样子?好的,首先我得告诉你什么不奏效,因为确实有别的办法是奏效的。嗯,不奏效的是:训
21:21
details so the the way to handle this is one way possibly to handle this
练一个系统通过从损坏版本重建良好图像来学习图像表征——这就是不奏效的。我们有一整套相关技术,比如各种变体的 denoising autoencoders,还有我的一些 Fair 同事开发的叫做 Mee 的东西,还有 Masked autoencoder。基本上就像 LLMs
21:25
which we've been working for a long time is to have a model that has
或类似的东西——你通过损坏文本来训练系统,只不过这里你损坏图像:你从图像中移除 patches,然后训练一个巨大的 neural net 来重建它们。但你得到的特征并不好。你知道它们不好是因为,如果你用同样的架构但用监督学习——用带标签的数据、文本描述来训练——效果就好得多。
21:29
what's called a latent variable and the latent variable is fed to an Nal net
它积累了大量知识,所以能检索到一些东西,但它只
21:35
and it's supposed to represent all the information about the world that you don't perceive
会一个接一个地吐出token,而不去规划整个答案
21:40
yet and uh that you need to augment uh the the system for the prediction
。但你这样描述,听起来好像一个token接一个
21:45
to do a good job at predicting pixels including the you know fine texture of
token、一次只生成一个token的方式注定很
21:50
the of the carpet and the on a couch and and the painting on
简陋。但如果世界模型足够复杂,那就不一样了。
21:55
the wall um uh that has been a complete failure essentially and we've tried lots
这个叫潜变量,潜变量会被输入到一个神经网络里,
22:02
of things we tried uh just straight neural Nets we tried Gans we tried uh
它应该代表你还没感知到的、关于世界的所有信息,
22:10
you know Vees all kinds of regularized Auto encoders we tried um many things we
你需要用它来增强系统,让系统能更好地预测像素,
22:17
also tried those kind of methods to learn uh good representations of images or video
包括地毯的细腻纹理、沙发上的细节,还有墙上的画。
22:24
um that could then be used as input to for example an image classification system
然后这个潜变量可以被用作输入,比如给一个图像分类
22:29
mhm and that also was basically failed like all the systems that attempt to predict
系统。嗯,但这也基本失败了——所有试图预测图像或视
22:34
missing parts of an image or video um you know from a corrupted version of
频缺失部分的系统,基本上都是从损坏版本里重建完整内
22:39
it basically so right take an image or a video corrupt it or transform it
容。所以,比如拿一张图像或一段视频,损坏它或者以某
22:45
in some way and then try to reconstruct the complete video or image
种方式变换它,然后尝试重建完整的视频或图像。
22:49
from the corrupted version and then hope that internally the system will develop a good
嗯,这基本上完全失败了。我们试过很多方法:直接
22:55
representations of images that you can use for object recognition segmentation whatever it is is
上神经网络、试过GAN、试过VAE、各种正则化的
23:00
that has been essentially a complete failure and it works really well for text that's
自编码器,试过很多。我们也试过用这类方法来学习图
23:06
the principle that is used for LMS right so where is the failure exactly is
像或视频的良好表征,然后把这些表征作为输入,比如
23:12
that that it's very difficult to form a good representation of an
给图像分类系统用。结果也基本是失败的。
23:16
image a good in like a good embedding of all all the important information in
图像方面,比如一个好的 embedding
23:21
the image is it in terms of the consistency of image to image to image
,能包含图像中所有重要信息。从图像到图像的
23:25
the image that forms the video like where what are the if we do a
一致性来看,构成视频的图像,如果我们做一个
23:29
highlight reel of all the ways you failed what what's that look like okay so
你所有失败方式的精彩集锦,那会是什么样子?
23:34
the reason this doesn't work uh is first of all I have to tell you
好吧,这之所以行不通,首先我得告诉你具体什么
23:38
exactly what doesn't work because there is something else that does work
行不通,因为确实有其他方法行得通。
23:42
uh so the thing that does not work is training a system to learn representations
所有尝试预测图像或视频缺失部分的系统——比如从损坏的版本出发,试图重建
23:47
of images by training it to reconstruct uh a good image from a corrupted version
完整的视频或图像——然后指望系统内部能发展出良好的图像表征,用来做物体识
23:53
of it okay that's what doesn't work and we have a whole slew of technique
别、分割之类的任务——这基本上完全失败了。但同样的原理用在文本上效果非常
23:58
for this uh that are you know variant of ding Auto encoders something called Mee
好,这就是LLM用的原理对吧。那么失败到底出在哪里呢?就是很难形成图像的
24:04
developed by some of my colleagues at Fair Max Doo encoder so
良好表征,也就是把所有重要信息都嵌入到图像里的那种好嵌入。
24:08
it's basically like the you know llms or or or or things like this where
一个好的嵌入应该能捕捉图像里所有重要信息,从
24:13
you train the system by corrupting text except you corrupt images you remove Patches from
图像到图像的一致性来看,比如构成视频的那些图像
24:19
it and you train a gigantic neet to reconstruct the features you get are not
,如果我们要做一个你所有失败方式的精彩集锦,那
24:24
good and you know they're not good because if you now train the same architecture
会是什么样子?好吧,这之所以行不通,首先我得告
24:29
but you train it supervised mhm with with uh label data with Tex textual descriptions
诉你具体什么不行,因为确实有别的办法是可行的。
24:34
of images Etc you do get good representations and the performance on recognition tasks is
比如图像等等,你确实能得到不错的表征,而且识别
24:39
much better than if you do this self-supervised free trining so the architecture is good
任务的性能也远好于用这种自监督的free tr
24:45
the architecture is good the architecture of the encoder is good okay but the fact
ining。所以架构是好的,编码器的架构是好的
24:50
that you train the system to reconstruct images does not lead it to produce to
,没问题。但如果你用自监督的方式训练系统去重建
24:55
learn good generic features of images when you train in a self-supervised way
图像,它并不会因此学到图像的良好通用特征。
25:00
self-supervised by reconstruction Yeah by reconstruction okay so what's the alternative the alternative is joint
通过重构进行自监督学习,对,就是重构。好,那替代方
25:06
embedding what is joint embedding what are what are these architectures that you're so excited
案是什么?替代方案是联合嵌入。什么是联合嵌入?你特别
25:12
about okay so now instead of training a system to encode the image and then
兴奋的那些架构到底是什么?好,现在不是训练一个系统
25:18
training it to reconstruct the the full image from a corrupted version you take the
去编码图像,然后训练它从损坏版本中重构完整图像,而是
25:23
full image you take
你拿完整图像,
25:25
the corrupted or transformed version you run them both through encoders mhm which in general
自监督通过重建?对,通过重建。那替代方案是什么?替代方案是joint embedding。什么是joint embedding?这些让你这么兴奋的架构到底是什么?好,现在不是训练一个系统去编码图像,然后训练它从损坏版本重建完
25:32
are identical but not necessarily and then you you train a predictor on top of
整图像,而是你拿完整图像,拿损坏或变换后的版本,把它们都送进编码器——嗯,通常编码器是相同的,但不一定——然后你在这些编码器上面训练一个predictor,让它从损坏版本的表征去预测完整输入的表征。所以是joint embe
25:39
those uh encoders um to predict the representation of the full input from the representation
dding,因为你拿完整输入和损坏或变换版本,都送进编码器,得到joint embedding,然后你说:我能从损坏版本的表征预测完整版本的表征吗?嗯,我把这个叫做JEA,也就是joint embedding predict
25:46
of the corrupted one okay so joint embedding because
ive architecture,因为它是joint embedding,而且有一个predictor,从坏的那个去预测好的那个的表征。
25:51
you're you're taking the the full input and the corrupted version or transform version run
你拿完整输入和损坏版本或变换版本,把它们
25:55
them both through encoders so you get a joint embedding and then you and then
都通过编码器,得到一个联合嵌入,然后你说,
26:00
you're you're saying can I predict the representation of the full one from the representation
我能从损坏版本的表示预测出完整版本的表示吗
26:05
of the corrupted one okay um and I call this a JEA so that means
?嗯,我把这叫做JEA,也就是联合嵌入预测
26:10
joint embedding predictive architecture because it's joint embedding and there is this predictor that predicts
架构,因为它是联合嵌入,并且有一个预测器,
26:15
the representation of the good guy from from the bad
从坏的那个预测好家伙的表示,
26:18
guy um and the big question is how do you train something like this uh
嗯,大问题是怎么训练这种东西。直到五六年前,我们还没有特别好的答案来训练这些东西,除了一个方法叫contrastive learning。contrastive learning的思路是,你拿
26:24
and until five years ago or six years ago we didn't have particularly good answers
一对图像,又是一张图像和它的损坏版本或降级版本,或者原始图像的变换版本,然后你训练预测的表征和那个表征相同。如果你只这么做,系统会collapse,它基本上完全忽略输入,产生恒定的表征。所以co
26:29
for how you train those things except for one um called contrastive contrastive learning where
ntrastive方法避免了这个问题,这些方法从90年代初就有了,我在1993年就发过一篇论文。方法是,你也展示你知道是不同的图像对,然后把它们的表征互相推开。所以你说,不仅我们知道相同的东西
26:35
um and the IDE contrastive learning is you you take a pair of images that
的表征应该相同或相似,而且我们知道不同的东西的表征应该不同,这防止了collapse。但它有一些局限性,过去六七年出现了一堆技术,可以复兴这类方法,有些来自FAIR,有些来自Google和其他地
26:41
are again an image and a corrupted version or degraded version
方。但这些contrastive方法有局限性。过去三四年发生的变化是,现在我们有了non-contrastive的方法,它们不需要那些负样本。
26:45
somehow or transform version of the original one and you train the predicted representation to
或者原始输入的变换版本,然后你训练预测的表示和
26:51
be the same as as that if you only do this the system collapses it
那个表示一样。如果你只这么做,系统会崩溃,基本上
26:58
basically completely ignores the input and produces representations that are con so the contrastive methods
完全忽略输入,产生固定的表示。对比方法避免了这个
27:04
avoid this and and those things have been around since the early 90s had a
问题,这些东西从90年代初就有了,我在1993年
27:10
paper on this in
发过一篇论文,
27:12
1993 um is you also show pairs of images that you know are different and
嗯,你还展示成对的不同图像,然后把它
27:16
then you push away the representations from each other so you say not only do
们的表示推开,所以你说,不仅我们知道相
27:21
representations of things that we know are the same should be the same or should
同的东西的表示应该相同或相似,而且我们
27:25
be similar but representation of things that we know are different should be different and
知道不同的东西的表示应该不同,这防止了
27:29
that prevents the collapse but it has some limitation and there's a whole bunch of
崩溃,但它有一些局限性。过去六七年出现
27:34
uh techniques that have
了一堆技术,
27:35
appeared over the last six seven years um that can revive this this type of
可以复兴这类方法,有些来自FAIR
27:41
method um some of them from Far some of them from from Google and other
,有些来自Google和其他地方,但
27:46
places um but there are limitations to those contrasting method what has changed in the
这些对比方法有局限性。在过去三四年里
27:52
last uh you know three four years is now now we have methods that are
,变化是我们现在有了非对比方法,它们
27:57
non-contrastive so they don't require those negative
不需要那些负样本。
28:00
contractive samples of images that are that we know are different you can only you
我们用来做对比的图像样本,我们知道它们是不一样的,你只能用同一个物体的不同版本或不同视角的图像来对比,还得靠一些其他技巧防止系统崩溃。现在我们已经有一大堆不同的方法了。那么,联合嵌入架构和LLM之间的根本区别是什么?JEPA能不能带我们走向AGI?不过我们是不是该说,你其实不太喜欢AGI这个词?我们大概每次聊都会争论那个G字。对,我懂,我懂,我们可能
28:06
turn them only with images that are you know different versions or different views of
还会继续争论下去,这挺好的。你喜欢这个词是因为你喜欢法语,而"ami"在法语里是朋友的意思,对吧?而且AMI代表高级机器智能。但不管怎样,JEPA能带我们走向那种高级机器智能吗?嗯,这算是第一步吧。首先,它和生成式架构比如LLM有什么区别?LLM或者通过重建来训练的视觉系统,它们会生成输入——也就是生成没有被破坏、没有被变换的原始输入。所以你必须预测
28:11
the same thing uh and you rely on some other tricks to prevent the system
所有像素,系统要花大量资源去预测所有这些像素和所有细节。而在JEPA里,你不需要预测所有像素,你只需要预测输入的一个抽象表示,这在很多方面要容易得多。所以JEPA系统在训练时要做的是,从输入中提取尽可能多的信息,但只提取那些相对容易预测的信息。世界上有很多东西是我们无法预测的,比如一辆车在街上行驶,路边可能有树,如果刮风,树叶就会以半混沌的随机方式摆动
28:17
from collapsing and we have have a dozen different methods for this now so what
,你预测不了,你也不在乎,你不想去预测。所以你希望你的编码器基本上消除所有这些细节,它会告诉你那里有树叶在动,但不会保留具体怎么动的细节。这样,当你在表示空间里做预测时,你就不需要预测每一片树叶的每一个像素。这不仅简单得多,还能让系统学会一种对世界的抽象表示——那些可以被建模和预测的东西被保留下来,其余的被当作噪声,由编码器消除。所以这相当于提升了表
28:23
is the fundamental difference between joint embedding architectures
示的抽象层次。想想看,我们平时描述一个现象时,总是在某个抽象层次上描述,不会每次都用量子场论去解释每一个自然现象,那是不可能的。所以我们有多个抽象层次来描述世界上发生的事情,从量子力学开始。
28:26
and llms so can uh can japa take us to AGI whether we should say
那么,JAPA能带我们走向AGI吗?我们是
28:31
that you don't like uh the term AGI and we'll probably argue I think every
不是该说你不喜欢AGI这个词,我们可能会争
28:35
single time I've talked to you with argued about the G and AGI yes get
论,我觉得每次我跟你聊,我们都会争论AGI
28:40
I get it I get it we we'll probably continue to argue about it it's
里的G,是的,我懂,我懂,我们可能还会继续
28:45
great uh you you like uh I me
争论,这挺好的。你喜欢,
28:48
this because cuz you like French and um I me is is is uh I
因为你喜欢法语,嗯,Ami我猜是法语
28:54
guess friend in French yes and Ami stands for advanced machine intelligence right um but
里的“朋友”,而Ami代表高级机器智能
29:00
either way can japa take us to that towards that advanced machine intelligence well so
,对吧?但不管怎样,JAPA能带我们
29:06
it's a it's a first step okay so first of all uh what What's the
走向那个高级机器智能吗?嗯,这算是第一
29:12
difference with generative architectures
步。首先,
29:13
like llms um so llms um or Vision systems that are trained by reconstruction generate
和生成式架构比如LLM有什么区别?嗯,LLM或者通
29:20
the inputs right they generate the original input that is non-corrupted non-transformed right so you
过重构训练的视觉系统,它们生成输入,对吧?它们生成未
29:27
have to predict all the pixels and there is a huge amount of resources spent
损坏、未变换的原始输入,所以你必须预测所有像素,系统
29:34
in the system to actually predict all those pixels all the
要花大量资源去预测所有这些像素,所有那些
29:39
details uh in a jepa you're not trying to predict all the pixels you're only
嗯,在JEPA里,你不是要预测所有像素,而是只预测输入的一个抽象表示,对吧?这在很多方面都简单得多。所以JEPA系统在训练时要做的是,从输入中提取尽可能多的信息,但只提取那些相对容易预测的信息。好吧,世界上有很多东西我们没法预测,比如,如果你有一辆自动驾驶的车在街上或路上开,路边可能有树,那天可能刮风,树
29:44
trying to predict an abstract representation of of the inputs right and that's much easier
上的叶子就会以半混沌的随机方式晃动,你预测不了,你也不在乎,你不想去预测。所以你希望你的编码器基本上消除所有这些细节,它会告诉你那儿有叶子在动,但不会保留具体发生了什么细节。然后当你在表示空间里做预测时,你就不需要预测一片叶子的每一个像素,这不仅简单得多,还能让系统本质上学习一个世界的抽象表示,其中可以被
29:49
in many ways so what the japa system when it's being trained is trying to
建模和预测的东西被保留下来,其余部分被视为噪声,被编码器消除。所以这相当于提升了表示的抽象层次。如果你想想,这其实是我们一直在做的事,每当我们描述一个现象时,我们都是在某个抽象层次上描述它,我们不会总是用量子场论来描述每一个自然现象,对吧?那是不可能的。所以我们有多个抽象层次来描述世界上发生的事情,从量子开
29:54
do is extract as much information as possible from the input but yet only EXT
始,你也可以分层来做。所以我认为这是智能系统的一个关键组成部分。而在语言中,我们可以绕过这个步骤,因为语言本身已经有了一定程度的抽象,已经消除了很多不可预测的信息,所以我们能直接预测单词,而不需要提升抽象层次。所以联合嵌入仍然是生成式的,但它是这个抽象表示空间里的生成式。对,你说语言方面我们偷懒了,因为我
29:59
ract information that is relatively easily predictable okay so there's a lot of things in
们免费得到了抽象表示,现在要放大来看,真正考虑通用智能系统时,我们必须处理物理现实的全部混乱,你确实需要这一步,从完整丰富详细的现实跳到基于它才能推理和做其他事情的抽象表示,对吧?关键是,那些通过预测来学习的自监督算法,即使在表示空间里,如果你喂给它们的数据冗余度更高,它们就能学到更多概念。数据中冗余越多
30:04
the world that we cannot
,它们就越能捕捉到一些内部结构。所以感知输入,比如视觉,其结构中的冗余远多于文本,文本的冗余度没那么高。
30:05
predict like for example if you have a s driving car driving down the street
比如说,如果你有一辆自动驾驶的车在街上或
30:10
or road uh there may be uh trees around the around the road and it
路上行驶,路边可能有些树,那天风很大,树
30:14
could be a windy day so the the leaves on the tree are kind of
上的叶子在随风飘动,那种半混沌的随机方式
30:19
moving in kind of semi chaotic random ways that you can't predict and you don't
你根本预测不了,你也不在乎,你也不想预测
30:23
care you don't want to predict so what you want is your encoder to basically
。所以你希望你的编码器基本上把这些细节都
30:28
eliminate all those details will tell you there's moving leaves but it's not going to
去掉,它会告诉你那里有叶子在动,但不会保
30:32
keep the details of
留那些细节。
30:33
exactly what's going on um and so when you do the prediction in representation space
基本上就像LLM或者类似的东西,你通过损坏文本来
30:39
you're not going to have to predict every single Pixel of a relief and that
训练系统,只不过这里你损坏图像——移除图像中的补丁
30:44
you know um not only is a lot simpler but also it allows the system
,然后训练一个巨大的网络来重建。你得到的特征并不好
30:49
to essentially learn an abstract representation of of the world where you know what can
,它们不好是因为如果你用同样的架构,但用监督学习来
30:54
be modeled and predicted is preserved and the rest is viewed as
训练——用带标签的数据、文本描述来训练。
30:59
noise and eliminated by the encoder so it kind of lifts the level of abstraction
噪声被编码器消除掉了,所以它相当于提升了表征的抽
31:03
of the representation if you think about this this is something we do absolutely all
象层次。如果你想想看,这其实是我们一直在做的事情
31:08
the time whenever we describe a phenomenon we describe it at a particular level of
——每当我们描述一个现象时,我们都是在某个抽象层
31:13
abstraction and we don't always describe every natural phenomenon in terms of quantum field Theory
次上描述它,而不会用量子场论去描述每一个自然现象
31:17
right that would be impossible right so we have multiple levels of abstraction to describe
,对吧?那根本不可能。所以我们有多个抽象层次来描
31:22
what happens in the world you know starting from Quantum
述世界上发生的事情,从量子开始。
31:25
field Theory to like atomic theory and molecules you know in chemistry materials and you
场论就像原子理论和分子,你知道,在化学材料中,一路往上到现实世界中的具体物体之类的东西。所以我们不能只在最低层级建模所有东西,而这正是JEA理念的核心——以自监督的方式学习抽象表征,而且你还可以分层去做。所以我认为这是智能系统的一个关键组成部分。在语言领域,我们或许可以跳过这一步,因为语言本身已经有一定程度的抽
31:31
know all the way up to you know kind of concrete objects in the real
象,并且已经剔除了大量不可预测的信息,所以我们不用费力提升抽象层级,直接预测单词就行。所以联合嵌入仍然是生成式的,但它是在这个抽象表征空间里生成式的。对,你说语言上我们偷懒了,因为我们已经免费获得了抽象表征,而现在我们必须退一步,真正思考通用智能系统,就必须处理物理现实的全部混乱。你确实需要这一步:从丰富详细的现
31:36
world and things like that so the we we can't just only model everything at
实跳跃到基于它进行推理的抽象表征,对吧?关键在于,那些通过预测来学习的自监督算法,即使在表征空间里,如果输入数据冗余度更高,它们学到的概念就更多。数据中冗余越多,它们就越能捕捉到内部结构。所以感知输入,比如视觉,其结构冗余度远高于文本,文本没那么冗余。这又回到你几分钟前问的问题:语言可能代表更多信息,因为它已经
31:42
the lowest level and that that's what the idea of JEA is really on is
被压缩了——你这一点说得对——但这意味着它的冗余度也更低,所以自监督效果不会那么好。那么,是否有可能将视觉数据的自监督训练和语言数据的自监督训练结合起来呢?尽管你瞧不起那10的13次方个token,但这些token代表了人类所发现知识的很大一部分,包括Reddit上的讨论、所有书籍和文章的内容,以及人类智力创造的
31:47
really about learn abstract representation in a self-supervised uh Manner and you know
全谱系。所以,有可能把这两者结合起来吗?最终是可以的,但我认为如果我们过早这么做,就有被诱惑去作弊的风险。事实上,这正是人们目前在视觉语言模型上做的事情——我们基本上在作弊。我们利用语言作为拐杖,来弥补视觉系统在从图像和视频中学习良好表征方面的缺陷。而这样做的问题在于……
31:52
you can do it hierarchically as well so that I think is an essential component
你可以分层来做,所以我认为这是智能系统的一个关键组成部分。而在语言中,我们可以不这么做,因为语言本身已经有了一定程度的抽象,并且已经消除了很多不可预测的信息,所以我们不需要
31:57
of an intelligent system and in language we can get away without doing this because
做那个——不需要提升抽象层次——而是直接预测单词。所以联合嵌入(joint embedding)仍然是生成式的,但它是在这个抽象表示空间里生成式的。对,你说语言我们偷懒了,因
32:03
language is already to some level abstract and already has eliminated a lot of information
为我们已经免费得到了抽象表示,而现在我们必须退一步,真正思考通用智能系统,我们必须处理物理现实的整个混乱局面。现实,你必须经历这个步骤,从完整、丰富的细节现实跳跃到一个基于此
32:08
that is not predictable and um so we can get away without doing the tring
的抽象表示,然后才能推理和做其他事情,对吧。关键在于,那些自监督算法,即使在表示空间里通过预测学习,如果输入的数据冗余度更高,它们学到的概念就更多。数据中冗余越多,它们就越
32:13
without you know lifting the abstraction level and
能捕捉到一些内部结构。所以感知输入,比如视觉,其结构中的冗余远多于文本,文本的冗余度没那么高。
32:16
by directly predicting words so joint embedding it's still generative but it's generative in this
尽管你瞧不起那10的13次方个token,但那10的13次
32:22
abstract representation space yeah and you're saying language we were lazy with language cuz we
方个token代表了——很大一部分——我们人类已经搞明白的
32:29
already got the abstract representation for free and now we have to zoom out actually
东西,包括Reddit上的讨论、所有书籍和文章的内容,以及
32:35
think about generally intelligent systems we have to deal with a full mess of physical
人类智力创作的整个光谱。那么,有没有可能把这两者结合起来呢?
32:41
reality of reality and you can't you you do have to do this step of
从被破坏的输入开始,但你只训练第二个分支,只训练网络中被喂入被
32:47
jumping from uh the full Rich detailed reality to a uh abstract representation of that
破坏输入的那部分,另一个网络你不训练。但由于它们共享相同的权重,
32:54
reality based on which you can then reason and all that kind of stuff right
当你修改第一个时,也会修改第二个。通过一些技巧,你可以防止系统
33:00
and the thing is those cell supervised algorithm that that learn by
崩溃——我之前解释过的那种崩溃,比如系统基本上会……
33:05
prediction even in representation space uh they learn more uh concept if the input data
例如,稍微改变大小,可能改变方向、模糊它、改变颜
33:10
you Feit them is more redundant the more redundancy there is in the data the
色,对它做各种糟糕的事情,但基本的糟糕操作,就是稍
33:16
more they're able to capture some internal structure of it and so there there is
微降低质量、改变构图,比如裁剪图像。或者在某些情
33:21
way more redundancy in structure in perceptual uh inputs sensory input like like like Vision
况下,比如JEA,你不需要做这些,你只需要遮住它的
33:27
than there is in uh text which is not nearly as redundant
一部分,对吧?你基本上就是移除一些区域。
33:31
this is back to the question you were asking a few minutes ago language might
这又回到你几分钟前问的那个问题——语言可能本身就承载了更
33:37
represent more information really because it's already compressed you're you're right about that but that
多信息,因为它已经被压缩了。你说得对,但这也意味着它的冗
33:43
means it's also less redundant and so self supervision will not work as well is
余度更低,所以自监督学习的效果不会那么好。那么,有没有可
33:49
it possible to join the self-supervised training on visual data and self-supervised training on language
能把视觉数据的自监督训练和语言数据的自监督训练结合起来呢
33:55
data there is a huge amount of knowledge
?要知道,人类的知识量是巨大的。
33:58
even though you talk down about those 10 to the 13 tokens those 10 to
帧,就像直管,管子,通常16帧左右,我们在整个1
34:04
the 13 tokens represent the entirety a large fraction of what US humans have figured
6帧中遮住同一个区域,当然每个视频不同。然后再次训
34:11
out both the talk on Reddit and the contents of all the books and the
练这个系统,让它从部分遮罩的视频预测完整视频的表示
34:17
Articles and the full spectrum of human uh intellectual creation so is it possible to
。这效果非常好,这是我们第一个能学到良好视频表示
34:23
join those two together well
的系统,所以当……
34:26
eventually yes but I think uh if we do this too early we run the
尽管你看不上那10的13次方个token,
34:31
risk of being tempted to cheat and in fact that's what people are doing at
但这10的13次方个token代表了美国人类
34:36
the moment with vision language model we're basically cheating we are using uh language as
所掌握的全部知识——无论是Reddit上的讨
34:41
a crutch to help the deficiencies of our uh Vision systems to kind of learn
论、所有书籍和文章的内容,还是人类智力创造
34:46
good representations from uh images and video and uh the problem with this is that
的整个光谱。所以,有没有可能把这两者结合起来
34:52
we
呢?
34:52
might you know improve our uh visual language system a bit I mean our language
可能通过喂图片能稍微改进我们的视觉语言
34:57
models by you know feeding them image but we're not going to get to the
系统,我是说语言模型,但我们不可能达到
35:02
level of even the intelligence or level of understanding of the world of a cat
猫或狗那种智力水平或对世界的理解程度—
35:07
or dog which doesn't have language you know they don't have language and they understand
—它们没有语言,却比任何LLM更懂世界
35:12
the world much better than any llm they can plan really complex actions and sort
,能规划非常复杂的行动,还能想象一系列
35:17
of imagine the result
行动的结果。
35:18
of a bunch of actions how do we get machines to learn that before we
在把语言结合进来之前,我们怎么让机器学会这些?显然,如果结合语言,这
35:23
combine that with language obviously if we combine this with language this is going to
肯定是个赢家,但在此之前,我们得先专注于如何让系统学习世界的运作方式
35:29
be a winner um but but before that we have to focus on like how
。所以这种joint embedding predictive ar
35:34
do we get systems to learn how the world works so this kind of joint
chitecture(联合嵌入预测架构)就有望学到类似常识的东西,就
35:39
embedding predictive architecture for you that's going to be able to learn something like Common
像猫用来预测怎么最优化地折腾主人、打翻东西那种常识——这就是希望所在。
35:44
Sense something like what a cat uses to predict how to mess with its owner
实际上我们用的技术是非对比性的,不仅架构是非生成式的,学习过程也是非对比的。我们有两套技术:一套基于蒸馏,有
35:50
most optimally by knocking over a thing that's that's the Hope in fact the techniques
好几种方法都用这个原理——DeepMind的BOL,FAIR的几个,一个叫VRA,另一个叫IA和vcra(其
35:55
we're using are non-contrastive uh so not only is the architecture non generative the learning
实VRA不是蒸馏方法,但IA和B肯定是),还有FAIR出的Dino或dyo。这些方法的核心是:拿完整输入(比
36:01
procedures we're using are non contrastive we have two two sets of techniques one set
如一张图),通过编码器生成表征,然后破坏或变换输入,再用基本相同的编码器(有些小差异)跑一遍,接着训练一个预
36:06
is based on distillation and there's a number of uh
测器(有时很简单,有时不存在)来从破坏后的输入预测第一个未破坏输入的表征。
36:10
methods that use this principle uh one by Deep Mind Bol a couple by by
但只训练第二个分支,也就是输入被破坏的那部分网络,另一个网
36:16
Fair one one called uh VRA and another one called IA and vcra I should
络不训练。不过因为它们共享权重,修改第一个时也会影响第二个。
36:21
say is not a distillation method actually but IA and B certainly are and there's
通过各种技巧可以防止系统崩溃——就是我之前说的那种系统完全
36:27
another one also called Dino or dyo also produced from at fair and the idea
忽略输入的情况。这效果很好,我们在FAIR开发的两种技术Di
36:33
of those things is that you take
no和IA在这方面表现很棒。
36:35
the full input let's say an image uh you run it through an encoder uh
那么这里说的是什么数据呢?有几种场景:一种是你拿一
36:41
produces a representation and then you corrupt that input or transform it running to the
张图,通过改变裁剪、大小、方向、模糊、颜色等方式破坏
36:47
essentially what amounts to the same encoder with some minor differences and then train a
它——各种基本但糟糕的操作,稍微降低质量、改变构图
36:53
predictor sometimes to predictor is very simple sometime doesn't exist but train a predictor to
,或者裁剪图像。在某些情况下,比如JEA,你甚至不需
36:58
predict a representation of the first first uh uncorrupted input
要做这些,只需要遮住部分区域就行。
37:02
from the corrupted input um but you only train the the second Branch um you
从被破坏的输入开始,但你只训练第二个分支,你
37:07
only train the part of the network that is fed with the corrupted input the
只训练网络中被喂入被破坏输入的那部分,另一个
37:11
other network you don't you don't train but since they share the same weight when
网络你不训练。但由于它们共享相同的权重,当你修
37:16
you modify the first one it also modifies the second one uh and with various
改第一个时,第二个也会被修改。然后通过各种技
37:20
tricks you can prevent the system from collapsing uh with the collapse of the type
巧,你可以防止系统崩溃,防止我之前解释的那种崩
37:25
I was explaining before where the system basically
溃,也就是系统基本上……
37:27
ignores the input um so that works very well the the technique with the two
最终是可以的,但我认为如果过早这么做,我们会有被诱惑
37:33
techniques we develop at Fair uh dino and uh and IA work really well for
去作弊的风险。事实上,这正是目前人们在视觉语言模型上做
37:40
that so what kind of data are we talking about here so this the several
的事情——我们基本上是在作弊。我们利用语言作为拐杖,来
37:46
scenario one uh one scenario is you take an image you corrupt it by um
弥补视觉系统的不足,从而从图像和视频中学习好的表征。问
37:52
changing the cropping
题在于,我们
37:54
for example changing the size a little bit maybe changing the orientation blurring it changing
比如稍微改变一下大小,或者改变一下方向,
37:58
the colors doing all kinds of horrible things to it but basic horrible things basic
模糊一下,改变颜色,做各种糟糕的事情,但
38:03
horrible things that sort of degrade the quality a little bit and change the framing
都是基本的糟糕操作,稍微降低一点质量,改变
38:07
uh you know crop the image um or and in some cases in the case
一下构图,比如裁剪图像。或者在某些情况下
38:12
of a JEA you don't need to do any of this you just you just
,比如JEA,你根本不需要做这些,你只需要
38:16
mask some parts of it right you just basically remove some regions like
遮住部分内容,就是直接去掉一些区域。
38:20
a big block essentially and and then you know run through the encoders um and
一大块,然后你让它通过编码器,训练整个系统和预测器,从被破坏的表示中预测出完好部分的表示。这样,IA 不需要知道它是一张图像,因为它只需要知道怎么做这个 masking。而用 Doo 的话,你必须知道它是图像,因为你需要做几何变换、模糊处理这类非常针对图像的操作。我们
38:25
train the entire system and and predictor to predict the representation of the good one
有一个更新版本的这种方法,叫 V JEA,基本思路和 IA 一样,但它是用在视频上的。现在你拿一整段视频,把其中一大块遮住,我们遮住的部分其实是一个时间管,就是视频中每一帧的某个完整片段,这个管在整个视频帧里位置是固定的。这个管通常是16帧左右,我们在整个16帧里遮住
38:30
from the representation of the corrupted one um so that's the Ia doesn't need to
同一个区域,当然每个视频遮的位置不同。然后训练这个系统,让它从部分被遮住的视频预测完整视频的表示。效果非常好,这是我们第一个能学到视频良好表示的系统。当你把这些表示输入到有监督的分类器头部时,它能以相当高的准确率告诉你视频里发生了什么动作。这是我们第一次得到这种质量的
38:35
know that it's an image for example because the only thing it needs to know
结果,说明确实形成了好的表示,也意味着这个方法有道理。我们还有初步结果,似乎表明这种表示能让系统判断视频是物理上可能的还是完全不可能的,比如某个物体消失了,或者突然从一个位置跳到另一个位置,或者改变了形状。所以它能捕捉到视频所呈现现实中的一些物理约束,关于物体的出现和消
38:40
is how to do this masking um whereas with doo you need to know it's
失。这确实很厉害。但问题是,这真的能让我们达到那种世界模型的程度吗?就是那种对世界有足够理解、能用来开车的模型?有可能,但要达到那一步还需要一段时间。其实已经有一些基于这个思路的系统了,你需要的是一个稍微改版的版本。想象你有一段完整的视频,你对它做的是要么在时间上向未
38:44
an image because you need to do things like you know geometri transformation and
来平移,也就是说你只看到视频的开头,看不到后半部分;要么你就直接把后半部分遮住。然后你训练一个我描述的那种 JEA 系统,让它从平移后的视频预测完整视频的表示,但同时你还要给预测器输入一个动作,比如方向盘向右转了10度。如果这是一个行车记录仪的视频……
38:49
blurring and things like that that are really image specific uh a more recent version
从被破坏的输入中学习,但你只训练第二个分支
38:53
of of this that we have is called V JEA so is basically the same
——你只训练网络中被喂入破坏输入的那部分,
38:58
idea as I except um it's applied to video so now you take a whole
另一个网络你是不训练的。但由于它们共享相同
39:02
video and you mask a whole chunk of it and what we mask is actually
的权重,当你修改第一个网络时,它也会修改第
39:07
kind of a temple tube so an all like a whole uh segment of each
二个网络。通过一些技巧,你可以防止系统崩溃
39:11
frame in the video over the entire video and that tube was like statically position
——就是我之前解释的那种系统基本上忽略输入的
39:16
throughout the
情况。
39:17
frames lit straight tube the tube yeah typically is 16 frames or something and we
帧数大概是16帧之类的,我们在整个16帧
39:22
mask the same region over the entire 16 frames it's a different one for every
上遮住相同的区域,当然每个视频遮的区域不同
39:27
video obviously and um and then again train that system so as to predict the
。然后再次训练这个系统,让它从部分被遮住
39:32
representation of the full video from The partially matched video uh that works really well
的视频中预测完整视频的表征。这个方法效果很
39:37
it's the first system that we have that learns good representations of video so that
好,这是我们第一个能学到良好视频表征的系
39:42
when
统。
39:42
you feed those representations to a supervised uh classifier head it can it can tell
你把那些表征喂给一个监督式的分类器头,它就能以相当高的准确率告诉你视频里发生了什么动作。所以,这是我们第一次得到那种质量的东西,这说明形成了好的表征,也就意味着这里面确实有点东西。嗯,我们还有一些初步结果,似乎表明这种表征能让我们的系统判断视频是物理上可能的,还是完全不可能的,比如某个物体消失了
39:47
you what action is taking place in the video with you know pretty good accuracy
,或者突然从一个位置跳到另一个位置,或者改变了形状之类的。所以它能捕捉到一些基于物理的约束,关于视频所呈现的现实,嗯,关于物体的出现和消失。对,这确实很厉害。但是,这真的能让我们达到那种世界模型的程度吗?就是那种对世界有足够理解、能用来开车的模型?嗯,可能吧,但这还需要一段时间才能实现。不过,已经
39:52
um so that that's it's the first time we get something of that uh of
有系统基于这个想法了,你需要的是一个稍微修改过的版本。想象一下,你有一段完整的视频,你对它做的是要么在时间上向未来平移,所以你只看到视频的开头,看不到后半部分,要么你直接遮住视频的后半部分。然后你训练一个我之前描述的那种JEA系统,让它从平移后的版本预测完整视频的表征,但同时你还要给预测器输入一个
39:57
that quality so that that's a a good test that a good representation is formed
动作,比如方向盘向右转了10度之类的。如果这是一个行车记录仪的视频,这里是我正在做的动作,这里是世界在时间t+1、t+Δt、t+2秒后的状态预测。如果你有这种模型,你就可以用它来做规划。所以现在你能做LLM做不到的事,那就是规划你要做什么,以达到特定的结果或满足特定的目标。对,你可以有一系列目标。
40:01
that means there's something to this yeah um we also preliminary result that seem to
比如,如果我有这样一个物体,我松开手,它就会掉下去;如果我用特定的力在桌子上推它,它就会移动;但如果我推桌子本身,用同样的力它可能不会动。所以我们脑子里有这种对世界的内部模型,它让我们能规划一系列动作来达到某个目标。所以,如果你有了这个世界模型,我们可以想象一系列动作,预测这些动作的结果会是什么,
40:06
indicate that the
然后衡量最终状态在多大程度上满足某个目标,比如移动那个东西。
40:07
representation allows us allow our system to tell whether the video is physically possible or
这种表征方式能让系统判断视频在物理上是否可能发生,还是完全不可能——比如某个物体消失了,或者突然从一个位置跳到另一个位置,甚至改变形状等等。所以它能捕捉到视频所呈现现实中的一些物理约束,关于物体的出现和消失。对,这确实很关键。但C,这真的能帮我们构建那种足够理解世界、能用来开车的世界模型吗?嗯,可能还需要一段时间才能达到那个程度。不过,已经有一些
40:13
completely impossible because some object disappeared or an object you know suddenly jumped from one
系统基于这个想法了。要实现这个,你需要对这个模型做一点修改:假设你有一段完整的视频,你对它做处理——要么在时间上向未来平移,这样你只能看到视频的开头,看不到后半部分;要么直接把后半部分遮掉。然后你训练一个我之前描述的那种JEA系统,让它从平移后的版本预测完整视频的表征。同时,你还要给预测器输入一个动作,比如方向盘向右转了10度。如果这是一个行车记
40:20
location to another or or change shape or something so it's able to capture some
录仪的画面,那么“这是我正在做的动作,这是对世界状态在t+1、t+Δt、t+2秒后的预测”。有了这种模型,你就可以用它来做规划。这就是LLM做不到的事——规划你要做什么,以达到某个特定结果或满足某个目标。你可以设定多个目标。比如,我知道如果我手里拿着这样一个物体,然后松开手,它会掉下去;如果我用一定力在桌子上推它,它会移动;但如果我推桌子本身,用同
40:26
physical con some physic based constraints about the reality represented in the video yeah about
样的力它可能不会动。所以我们脑子里都有这种内在的世界模型,它让我们能规划一系列动作来达成某个目标。那么,有了这个世界模型,你可以想象一系列动作,预测这些动作的结果,然后衡量最终状态在多大程度上满足了某个目标——比如移动某个物体。你规划一系列指令,根据你的世界模型,系统的最终状态会满足你设定的目标。这其实就是火箭轨迹规划的方式,从计算机出现以来就一
40:32
the appearance and The
直这么用,基本上从上世纪60年代初就开始了。所以没错,这就是模型预测控制。但你经常也会提到……
40:34
Disappearance of objects yeah that's really you okay but C can this actually get us
物体的消失,嗯,确实。但C真的能带我们
40:40
to this kind of uh World model that understands enough about the world to be
走到那种对世界有足够理解、能开车的世界
40:46
able to drive a car uh possibly um this is going to take a while
模型吗?有可能,但要达到那个点还需要一
40:52
before we get to that point but um um and there are systems already you
段时间。而且现在已经有一些系统了,大家
40:58
know everybody
知道,
40:59
systems that are based on this uh idea uh and the what you need for
有些系统就是基于这个想法。你需要的是
41:04
this is a slightly modified version of this where um imagine that you have uh
一个稍微修改过的版本,想象一下你有一
41:09
a video and the a complete video and what you're doing to this video is
段视频,一段完整的视频,你对这段视频
41:14
that you're either translating it in time towards the future so you only see the
做的操作是,要么在时间上把它向未来平移
41:19
beginning of the video but you don't see the latter part of it that is
,所以你只看到视频的开头,看不到后面
41:24
in the
部分,
41:25
original one or you just mask the second half of the video for example um
要么你只遮住视频的后半部分,比如。然后你
41:30
and then you you train a JEA system of the type I describe to predict
训练一个我描述的那种JEA系统,让它从平
41:35
the representation of the full video from the the shifted one but you also feed
移后的视频中预测完整视频的表征,但同时你
41:40
the predictor with an action for example you know the wheel is turned 10 degrees
还会给预测器输入一个动作,比如方向盘向右转
41:45
to the to the right or something right so if it's a you know a
了10度之类的。所以如果这是一个行车记录
41:49
dash cam in a
仪的视频……
41:51
car and you know the angle of the wheel you should be able to predict
车和车轮的角度,你应该能在一定程度上预测接下来会发生什么——当然,你不可能预测视野里出现的所有物体的细节,但在抽象表征层面,你大概能预测会发生什么。所以现在你有了一个内部模型,它说:这是我在时间T对世界状态的认知,这是我正在采取的一个动作,这是我对时间T+1、T+Δt、T+2秒等等世
41:55
to some extent what's going what's going to go what's going to happen to which
界状态的预测。如果你有这种模型,你就可以用它来做规划。所以现在你能做LLM做不到的事,那就是规划你要做什么,以达到某个特定的结果或满足某个特定的目标。你可以有多个目标。比如,我知道如果我有一个这样的物体,我松开手,它就会掉下去;如果我用特定的力在桌子上推它,它就会移动;如果我推桌子本身
41:59
to see uh you're not going to be able to predict all the details of
,用同样的力它可能不会动。所以我们脑子里有这个世界的内部模型,它让我们能规划一系列动作来达到某个特定目标。所以如果你有这个world model,我们可以想象一系列动作,预测这些动作的结果,衡量最终状态在多大程度上满足某个特定目标,比如把瓶子移到桌子左边,然后规划一系列动作来在运行时
42:04
you know objects that appear in the view obviously but at a abstract representation level
最小化这个目标。我们不是在说学习,我们说的是推理时间。所以这实际上是规划,在最优控制中,这是一个非常经典的东西,叫做model predictive control。你有一个你想控制的系统的模型,这个模型能预测对应于一系列指令的状态序列,然后你规划一系列指令,这样根据你的world m
42:08
you can you can probably predict what's going to happen so now what you have
odel,系统的最终状态会满足你设定的目标。这就是从计算机出现以来,火箭轨迹的规划方式,基本上从60年代初就开始了。所以是的,这就是model predictive control。但你也经常谈到分层规划,分层规划能从这个中自然产生吗?嗯,不会,你需要构建特定的架构来实现分层规划。分
42:12
is a internal model that says here is my idea of state of the world
层规划对于规划复杂动作是绝对必要的。如果我想从纽约去巴黎,这是我经常用的例子,我坐在NYU的办公室里,我的目标是最小化我到巴黎的距离。在一个高层、非常抽象的位置表征上,我需要把这个目标分解成两个子目标:第一个是去机场,第二个是坐飞机到巴黎。所以我的子目标现在变成了去机场,我的目标函数是
42:17
at time T
我到机场的距离。我怎么去机场?我得走到街上,叫一辆出租车。
42:18
here is an action I'm taking here's a prediction of the state of the world
这是我正在采取的一个行动。这是对世界在时间t加一、t加Delta t、t加两秒时的状态预测——不管具体时间是多少。如果你有这种模型,就可以用它来做规划。所以现在你可以做LLM做不到的事:规划你要做什么,以便达成某个特定结果或满足某个特定目标。对吧?所以你可以有多个目标。嗯,对吧?如果我知道,我能预测:如果我有一个这样的物体,然后我松开手,它会掉下去;嗯
42:22
at time t plus one t plus Delta t t plus 2 seconds whatever it
,如果我用力在桌子上推它,它会移动;如果我推桌子本身,它可能不会以同样的力移动。嗯,所以我们脑子里有这种内部的世界模型,嗯,它让我们能规划一系列行动来达成某个特定目标。嗯,所以,嗯,现在如果你有这个world model,我们可以想象一系列行动,预测这些行动的结果会是什么,衡量最终状态在多大程度上满足某个特定目标——比如移动指令,然后你规划一系列指令,
42:27
is if you have a model of this type you can use it for planning
这样根据你的world model,系统的最终状态会满足你设定的目标。这就是火箭轨迹从计算机出现以来——基本上从60年代初开始——就一直被规划的方式。嗯,是的,这就是model predictive control。但你也经常谈到“到巴黎的距离”,这是一个高层次、非常抽象的对我位置的表示。我必须把它分解成两个子目标:第一个是,嗯,去机场;第二个是,坐飞
42:32
so now you can do what llms cannot do which is planning what you're going
机去巴黎。好的,所以我的子目标现在是,嗯,去机场。我的目标函数是我到机场的距离。我怎么去机场?我得走到街上,拦一辆出租车。嗯,这最终会变成毫秒级的肌肉控制。好的,显然你不会用毫秒级的肌肉控制来规划你从纽约到巴黎的整个行程。首先,那会极其昂贵,而且也完全不可能,因为你不知道所有会发生的情况。嗯,你知道,比如拦出租车要花多长时间,或者去机场路上交通怎么样。
42:36
to do so as to arrive at a particular uh outcome or satisfy a particular
嗯,我的意思是,你得精确知道所有条件才能做这种规划,但你并没有这些信息。所以你必须做这种分层规划,这样你就能开始行动,然后一边行动一边重新规划。在AI领域,没人真正知道怎么做这个。嗯,没人知道如何训练一个系统来学习合适的多层表示,以便分层规划能奏效。那有没有类似的东西已经出现了?比如,你能不能用最先进的LLM,通过你刚才做的那种详细的提问方式,从纽约到
42:41
objective right so you can
巴黎?也就是:你能给我一个高层次、十个步骤的清单,告诉我从纽约到巴黎需要做什么?然后对其中每一步,你能给我一个清单……
42:43
have a number of objectives um right if you know I can I can predict
我们有很多目标,对吧。你知道,我可以预测,如果我手里拿着这样一个东西,然后松开手,它就会掉下去。如果我以特定的力在桌子上推它,它就会移动。但如果我用同样的力去推桌子本身,它可能不会动。所以我们的大脑里有一个关于世界的内部模型,这个模型让我们能够规划一系列动作,以达到某个特定目标。所以,有了这个世界模型,我们可以想
42:48
that uh if I have uh an object like this right and I open my
象一系列动作,预测这些动作的结果,然后衡量最终状态在多大程度上满足某个目标,比如移动指令。你规划一系列指令,这样根据你的世界模型,系统的最终状态会满足你设定的目标。这就是火箭轨迹规划的方式,自从计算机出现以来就是这样,基本上是从60年代初开始的。所以,是的,这就是模型预测控制,但你也经常提到高层次上的“到巴黎的距离
42:53
hand it's going to fall right and uh and if if I push it with
”,这是一个非常抽象的位置表示。我必须把它分解成两个子目标:第一个是去机场,第二个是坐飞机到巴黎。所以我的子目标现在是去机场,我的目标函数是我到机场的距离。我怎么去机场?我得走到街上,打一辆出租车。这实际上相当于毫秒级的肌肉控制。显然,你不会用毫秒级的肌肉控制来规划你从纽约到巴黎的整个行程。首先,这太昂贵了,而且
42:57
a particular force on the table it's going to move if I push the table
也完全不可能,因为你不知道所有会发生的情况——比如打出租车要多久,或者交通状况下到机场要多久。你必须知道所有条件才能做这种规划,但你并没有这些信息。所以你必须做这种分层规划,这样你可以开始行动,然后边走边重新规划。在AI领域,没人真正知道怎么做这个。没人知道如何训练一个系统来学习适当的多层次表示,以便分层规划能工作
43:02
itself it's probably not going to move uh with the same Force um so we
。那么,类似的东西已经出现了吗?比如,你能用一个最先进的LLM,通过做你刚才那种详细的提问,从纽约到巴黎吗?比如,你能给我一个高层次列表,列出从纽约到巴黎需要的10个步骤,然后对每个步骤再给我一个列表……它们被训练来生成这类计划,对吧?它们无法为从未遇到过的情况做规划。它们基本上只能重复它们被训练过的模板。但就拿
43:07
have we have this internal model
从纽约到巴黎的例子来说,它会在哪个抽象层次开始出问题?因为我可以想象,这个分析的几乎每个部分它都能相当准确地回答,尤其是当你……
43:09
of the world in our in our mind uh which allows us to plan sequences
我们脑海中有一个关于世界的模型,这让我们能够规划一系列动作,最终达成某个目标。所以,如果你有了这个worl
43:15
of actions to arrive at a particular goal um and so um so now if
d model,就可以想象一系列动作,预测这些动作会带来什么结果,然后衡量最终状态在多大程度上满足某个目标
43:20
you have this world model we can imagine a sequence of actions predict what the
——比如,移动指令,然后你规划一系列指令,这样根据你的world model,系统的最终状态就会满足你设定的
43:25
outcome of the sequence of action is going to be measure to what extent the
目标。这其实就是火箭轨迹规划的方式,从计算机出现以来就一直这么干,基本上从上世纪60年代初就开始了。所以,
43:31
final State satisfies a particular objective like you know moving the
是的,这跟model predictive control有关,但你也经常提到。
43:35
bottle to the left of the table um and then plan a sequence of actions
桌子左边的瓶子,嗯,然后规划一系列动作,在运行时最小化这个目标。我们不是在讨论学习,我们讨论的是推理时间,对吧?所以这实际上是规划,在最优控制里,这是个非常经典的东西,叫做模型预测控制。你有一个你想控制的系统模型,这个模型能预测与一系列命令对应的状态序列,然后你规划一系列命令,这样根据你的世界模型,系统的最终状态会满足你
43:39
that will minimize this objective at run time we're not talking about learning we're talking
设定的目标。这就是火箭轨迹从有计算机以来就被规划的方式,基本上从60年代初就开始了。所以是的,这是模型预测控制,但你也经常提到分层规划,分层规划能从这个过程中自然涌现吗?嗯,不会,你需要构建特定的架构来实现分层规划。所以如果你想规划复杂动作,分层规划是绝对必要的。比如我想从纽约去巴黎,这是我常用的例子,我坐在纽约大学的办公
43:44
about inference time right so this is planning really and in optimal control this is
室里,我需要最小化的目标是我与巴黎的距离。在一个高层次、非常抽象的位置表示下,我得把这个目标分解成两个子目标:第一个是去机场,第二个是坐飞机到巴黎。所以我的子目标现在是去机场,我的目标函数是我与机场的距离。我怎么去机场?我得走到街上,拦一辆出租车,这涉及到毫秒级别的肌肉控制。显然,你不会用毫秒级别的肌肉控制来规划从纽约到巴
43:49
a very classical thing it's called Uh model predictive control you have a model of
黎的整个行程,首先这极其昂贵,而且完全不可能,因为你不知道所有会发生的情况,比如拦出租车要多久,或者去机场路上堵不堵。我的意思是,你得精确知道所有条件才能做这种规划,但你并没有这些信息,所以你必须做这种分层规划,这样你可以先开始行动,然后边走边重新规划。在AI领域,没人真正知道怎么做这个,没人知道如何训练一个系统来学习合
43:54
the system you want to control that you know can predict the sequence of State
适的多个层次的表示,让分层规划能工作。那么,类似的东西已经能涌现了吗?比如,你能用一个最先进的LLM,通过做你刚才那种详细的提问,从纽约到巴黎吗?也就是,你能给我一个从纽约到巴黎需要做的10个步骤的高层列表吗?然后对每个步骤,你能给我10个实现这个步骤的子步骤吗?再对每个子步骤,你能给我10个更细的步骤,直到你移动你的单个
43:58
St corresponding to a sequence of
肌肉?也许不用那么细,直到你能用你的大脑实际执行的动作。所以这里隐含了很多问题。首先,LLM能在一定程度上回答这些问题,前提是……
44:00
commands and you're planning a sequence of commands so that according to your world model
命令和规划一系列命令,这样根据你的世界模型,系统的最终状态会满足你设定的目标——这就是火箭轨
44:06
the the the end state of the system will uh satisfy an objectives that you
迹规划的方式,自从计算机出现以来就是这样,基本上从上世纪60年代初就开始了。所以是的,这就是模
44:12
fix this is the way uh you know rocket trajectories have been planned since computers
型预测控制,但你也经常谈到相当于毫秒级别的肌肉控制。显然,你不会用毫秒级别的肌肉控制来规划从纽
44:18
have been around so since the early 60s essentially so yes for model predictive control
约到巴黎的整个行程——首先,这成本高得离谱,而且完全不可能,因为你不知道所有会发生的情况,比如
44:24
but you also often talk about
打车要花多久,或者去机场路上堵不堵车。
44:27
hierarchical planning can hierarchical planning emerge from this somehow well so no you you will
层级规划——层级规划能不能从这种机制里自然涌现出来?嗯,不会,你得专门设计特定的架构才能支持层级规划。所以如果你想规划复杂的行动,层级规划是绝对必要的。比如我想从纽约去巴黎,这是我经常用的例子。我坐在NYU的办公室里,我的目标是在高层面上最小化我到巴黎的距离,这是一个非常抽象的关于我位置的表示。那我就得把这个目标分解成两个子目标:第一
44:32
have to build specific architecture to allow for hierarchical planning so hierarchical planning is absolutely
个是去机场,第二个是坐飞机到巴黎。好,现在我的子目标是去机场,我的目标函数是我到机场的距离。那怎么去机场呢?我得走到街上,拦一辆出租车。这又涉及到毫秒级别的肌肉控制。显然,你不会用毫秒级别的肌肉控制来规划你从纽约到巴黎的整个行程——首先这代价极其高昂,而且完全不可能,因为你不知道所有会发生的情况,比如拦出租车要多久,或者路上堵车要多久
44:37
necessary if you want to plan complex actions uh if I want to go from
。你得精确知道所有条件才能做这种规划,但你根本没有这些信息。所以你必须做层级规划,这样你才能开始行动,然后边走边重新规划。在AI领域,没人真正知道怎么做到这一点。没人知道怎么训练一个系统去学习合适的多个层级的表示,让层级规划真正起作用。那这种东西有没有已经涌现出来了?比如你能不能用一个最先进的LLM,通过你刚才做的那种详细的提问方式,
44:42
let's say from New York to Paris this the example I use all the time
从纽约到巴黎?就是先问“给我一个从纽约到巴黎需要做的10个高层步骤”,然后对每个步骤再问“给我10个步骤来实现这个步骤”,再对每个子步骤问“给我10个步骤来实现这个子步骤”,一直问到你能实际用你的大脑去控制肌肉的程度?也许不用到肌肉,只要能实际执行就行。这背后其实隐含了很多问题。首先,LLM能够回答其中一些问题,到一定的抽象层级,前提
44:47
and I'm sitting uh in my office at NYU my objective that I need to
是它们被训练过生成这类计划。对吧?它们没法为从未遇到过的情况做规划,基本上只能复述它们训练过的模板。但就拿纽约到巴黎这个例子来说,它会在哪个抽象层级开始出问题?因为我可以想象,这个分析的几乎每一个部分,它都能比较准确地回答,尤其是当LLM是那个坐在上面做更大推理的东西时,比如“我需要订机票,我知道怎么去那些网站”之类的。当然,很多人们
44:52
minimize is my
知道的相对高层的计划其实是学来的,大多数人并不是自己发明这些计划的。
44:53
distance to Paris at a high level a very astract representation of my uh my
从高层次、非常抽象的角度来看,比如“距离巴黎的距
44:58
location I would have to decompose this into two sub goals first one is um
离”,代表我的位置,那我得把这个目标分解成两个子
45:03
go to the airport second one is catch a plane to Paris okay so my
目标:第一,去机场;第二,坐飞机到巴黎。好,现在
45:07
sub goal is now uh going to the airport my objective function is my distance
我的子目标是去机场,我的objective fun
45:12
to the airport how do I go to the airport where I have to go
ction就是离机场的距离。那我怎么去机场呢?我
45:17
in the street and H a taxi
得走到街上,打个出租车。
45:19
which you can do in New York um okay now I have another sub goal
在纽约就能做到。嗯,好,现在我有了另一个子目标:走到街上。嗯,那就意味着要去电梯,坐电梯下楼,走到街上。我怎么去电梯呢?我得从椅子上站起来,打开办公室的门,走到电梯,按下按钮。我怎么从椅子上站起来呢?就像你能想象的那样,一直往下分解,直到基本上变成毫秒级的肌肉控制。好,显然你不会用毫秒级的肌肉控制来规划你从纽约到巴黎的整
45:24
go down on the street uh well that means going to the elevator going down
个行程。首先,那样做成本极高,而且完全不可能,因为你不知道所有会发生的情况。嗯,你知道打车要多久,或者去机场路上堵不堵车。嗯,我是说,你得精确知道所有条件才能做这种规划,但你并没有这些信息。所以你必须做这种分层规划,这样你才能开始行动,然后边走边重新规划。在AI领域,没人真正知道怎么做这个。嗯,没人知道怎么训练一个系统来学
45:29
the elevator walk out the street how do I go to the elevator I have
习合适的多个层次的表征,让分层规划能工作。有没有什么东西已经能实现类似的效果?比如,你能不能用一个最先进的LLM,通过你刚才那种详细的提问方式,从纽约到巴黎?也就是,你能不能给我一个高层次的、从纽约到巴黎需要做的10个步骤的列表?然后对每个步骤,再给我10个步骤来实现它?再对每个子步骤,再给10个步骤,一直分解到你移动单
45:34
to uh stand up from my chair open the door of my office go to
个肌肉?嗯,也许不用到那个程度,只要到你用意识能实际执行的程度就行。所以这里面隐含了很多问题,对吧?首先,LLM能够回答其中一些问题,达到一定的抽象层次,前提是它们的训练集里有过类似的场景。它们能回答所有这些问题,但有些答案可能是幻觉,也就是不真实的。对,没错。我是说,它们可能会给出一些答案,但不太可能真正给出毫秒级的肌肉
45:38
the elevator push push the button how do I get up from my chair like
控制,比如你怎么从椅子上站起来。所以,在一定的抽象层次内,我们可以用语言描述事情,它们也许能给你一个计划,但前提是它们被训练过生成这类计划。嗯,对。它们无法为从未遇到过的情况做规划,基本上只能复述它们训练过的模板。但就拿纽约到巴黎这个例子来说,它会在哪个抽象层次开始出问题呢?因为我可以想象,这个分析的几乎每个部分,LLM
45:43
you know you can imagine going down all the way down to basically
都能比较准确地回答,尤其是当你在谈论纽约和巴黎这种大城市时。所以,当然,如果你专门微调它,LLM肯定能解决这个问题。嗯,所以,我不能说LLM做不到。如果你专门训练它,它肯定能做到,这没问题。到一定的层次,只要事情能用语言表述,它就能处理。但如果你想往下到比如怎么下楼梯,或者只是……
45:48
what amounts to millisecond by millisecond muscle control okay and obviously you're not going to
毫秒级的肌肉控制,对吧?显然你不会
45:53
plan your entire trip from New York to Paris in terms of millisecond by millisecond
用毫秒级的肌肉控制来规划你从纽约到巴
45:58
muscle control first that would be incredibly expensive but it will also be completely impossible
黎的整个行程——首先那会极其昂贵,而
46:03
because you don't know all the conditions of what's going to happen uh you know
且完全不可能,因为你不知道所有会发生
46:08
how long it's going to take to catch a taxi um or to go to
的情况,比如打车要花多久,或者去机
46:14
the airport with traffic you know
场路上堵不堵车。
46:16
uh I mean you you would have to know exactly the condition of everything to
这其实就细化到了毫秒级别的肌肉控制。显
46:21
be able to do this planning and you don't have the information so you you
然,你不会用毫秒级的肌肉控制来规划从纽约
46:25
have to do this hierarchical planning so that you can start acting and then sort
到巴黎的整个行程。首先,这成本高得离谱,
46:30
of replanning as you go and nobody really knows how to do this in AI
而且根本不可能,因为你不知道所有会发生的
46:35
um nobody knows how to train a system to learn the appropriate multiple levels of
情况——比如,打个出租车要多久,或者路上
46:40
representation so that hierarchical
堵车要多久。
46:42
planning Works does something like that already emerge so like can you use an llm
规划机制是否已经以某种方式涌现出来了?
46:47
state-ofthe-art llm to get you from New York to Paris by doing exactly the kind
比如,你能不能用最先进的LLM,通过你
46:52
of detailed set of questions that you just did which is can you give me
刚才做的那种详细提问,从纽约到巴黎?也
46:57
a highight a list of 10 steps I need to do to get from New
就是先让我列出从纽约到巴黎需要做的10个
47:03
York to Paris and then for each of those steps can you give me a
主要步骤,然后针对每个步骤再列出更细的
47:08
list of
清单?
47:09
10 steps how I make that step happen and for each of those steps can
10步来让这件事发生,然后对每一步,你能不能列出10个小步骤来实现每一步,直到你在移动你身上的每一块肌肉——呃
47:13
you give me a list of 10 steps to make each one of those until
,可能不用那么细,但就是那些你真正能用大脑去执行的动作,对吧?所以这里其实隐含了很多问题。首先,LLMs能够回
47:18
you're moving your mus individual muscles uh maybe not whatever you can actually act upon
答其中一些问题,但只能到一定的抽象层次,前提是它们被训练过生成这类计划。嗯,它们没法为从未遇到过的情况做规划,
47:22
using your mind right so there's a lot of questions that are sort implied by
基本上只能复述它们训练过的模板。但就拿从纽约到巴黎这个例子来说,它会在哪个层面开始出问题呢?因为我可以想象,几
47:27
this right so the first thing is llms will be able to answer some of
乎每一步它都能相对准确地回答,尤其是当LLMs作为上层系统来处理更大的推理时,比如“我需要订机票,我知道怎么上
47:31
those questions down to some level of exraction under the condition that
网站”之类的。而且很多人们知道的高层次计划其实是学来的,大多数人并不会自己发明计划,对吧?
47:35
they've been trained with similar scenarios in their training set they would be able to
它们训练集中包含过类似场景的话,就能回答所有这些问题,但其中一些答案可能是幻觉,也就是不真实的。对,没错。我是说,它们大概率会给出
47:39
answer all those questions but some of them may be hallucinated meaning non-factual yeah true
某种回答,但不可能真正实现毫秒级的肌肉控制,比如你怎么从椅子上站起来。不过,在某种抽象层面上,我们可以用语言描述事情,它们或许能给你
47:43
I mean they will probably produce some answer except they're not going to be able
一个计划,但前提是它们接受过生成这类计划的训练。嗯,对。它们无法为从未遇到过的情况制定计划,基本上只能复述训练时学到的模板。但就拿纽
47:47
to really kind of produce millisecond by millisecond muscle control of how you how you
约到巴黎这个例子来说,它会在哪个层面开始出问题?我是说,在哪个抽象层级上你会觉得开始不行了?因为我能想象,这个计划的几乎每个部分它都
47:52
stand up from your chair right so but down to some level of exraction we
能比较准确地回答,尤其是当你谈论纽约和巴黎这种大城市时。所以,我觉得LLM肯定能解决这个问题,如果你针对它做fine-tuning的
47:56
can describe things by words they might be able to give you a plan but
话。你知道,所以,我不能说LLM做不到,如果你专门训练它,它肯定能做到,这没问题。但仅限于能用语言表述的层面。可如果你想深入到比如
48:00
only under the
怎么下楼梯这种细节,那就……
48:01
condition that they've been trained to produce those kind of plans mhm right they're not
我的意思是,你得精确知道所有条件才能做
48:05
going to be able to plan for situations where that that they never encountered before
这种规划,但你根本没有这些信息。所以,
48:10
they basically are going to have to regurgitate the template that they've been trained on
你必须做这种分层规划,这样你才能先行动
48:14
but where like just for the example of New York to Paris is is it
起来,然后边走边重新规划。在AI领域,
48:18
going to start getting into trouble like at which layer layer of abstraction do you
没人真正知道该怎么做——没人知道怎么训
48:22
think you'll start cuz like I can imagine almost every single part of that anal
练一个系统,让它学会合适的多个层次的表
48:27
will be able to answer somewhat accurately especially when you're
征,从而实现这种分层规划。
48:30
talking about New York and Paris major cities so I mean certainly uh LM would
LLM是更上层的东西,负责更大的推理,比如“我需要订机票,我知道怎么上网站”这类事情。没错。而且,
48:34
be able to solve that problem if you f tun need for it you know
人们知道的很多计划,其实都是相对高层的,是学来的。大多数人并不是自己发明计划。他们自己……我们当然
48:38
just uh and and so uh I can't say that nlm cannot do this it
有这种能力,但大多数人们使用的计划都是他们被训练过的,比如看到别人用这些计划,或者被告知该怎么做。
48:43
can do this if you train it for it there's no question uh down to
对。你没法凭空发明,比如找一个从没听说过飞机的人,告诉他“你怎么从纽约去巴黎”,他很可能无法自己拆
48:47
a certain level where things can be formulated in terms of words but like if
解出整个计划,除非他之前见过类似的例子。所以,LLM当然能做到这一点,但问题在于,如何把高层计划与
48:51
you want to go down to like how do you you know climb down the
底层动作连接起来。这就需要像Jad这样的东西,它提升表征的抽象层级,而不试图重建每个细节。这就是为什
48:56
stairs or just
么我们需要Jass。
48:57
stand up from your chair in terms of uh words like you you can't you
从椅子上站起来这件事,用语言是做不到的,你没法用文字表达清楚。这就是为什么你需要对物理世界有经验,那种带宽比人类语言能表达的要高得多。所以我们一直在聊的联合嵌入空间,是不是就是我们在机器人领域跟物理现实互动所需要的东西?而LLM是坐在它上面的那层,用来做更高层次的推理,比如“我需要订一张机票,
49:02
can't do it um you you need that's one of the reasons you need experience
我知道怎么去那些网站”之类的。而且,很多人们知道的计划,其实都是学来的,不是自己发明的。大多数人并不会凭空创造计划——当然,我们确实有一些这种能力,但大部分人们用的计划,都是他们被训练过的,比如看过别人用这些计划,或者被告知该怎么做。你没法凭空发明,比如找一个从来没听说过飞机的人,告诉他“你怎么
49:06
of the physical world which is much higher bandwidth than what you can express in
从纽约到巴黎”,他很可能没法自己拆解出整个计划,除非他之前见过类似的例子。所以LLM当然也能做到这一点,但问题在于,如何把这种高层次的东西和底层的动作连接起来,这就需要像Jad这样的东西,它能把表征的抽象层次提升,而不需要重建每一个细节。这就是为什么我们需要Jass。我很想在你对自回归LLM的
49:11
words in human language so everything we've been talking about on the joint embedding space
怀疑上多聊一会儿。我想测试一下这种怀疑:你说的每句话都很有道理,但如果我把你今天说的这些,放到比如十年前——或者少一点,三年前——我根本没法预测LLM会这么成功。所以你觉得自回归LLM能变得这么厉害,这合理吗?是的,你能解释一下你的直觉吗?因为如果我把你的智慧和直觉当真,我会说自回归LLM一个t
49:16
is it possible that that's what we need for like the interaction with physical reality
oken一个token地生成,根本不可能做到它们现在做的那些事。不,有一件事是自回归LLM——或者说LLM整体,不只是自回归的,也包括像BERT那种双向的——在利用的,那就是自监督学习。我多年来一直是自监督学习的坚定支持者。所以这些东西是一个极其令人印象深刻的证明,说明自监督学习真的有效。这个
49:21
for on the robotics front and then just the
想法——它并不是从BERT开始的,但BERT确实是一个很好的展示——就是你拿一段文本,把它破坏掉,然后训练一个巨大的神经网络去重建缺失的部分。这带来了巨大的好处,它让我们能够……
49:24
llms are the thing that sits on top of it for the bigger reasoning about
对于自回归LLMs的怀疑,我想用一种方式来检验这种怀疑。你说的每句话都很有
49:29
like yeah the fact that I need to book a plane ticket and I need
道理,但如果我把你今天说的以及你一贯的观点,应用到比如10年前——可能没那么
49:34
to know I know how to go to the websites and so on sure and
久,就说3年前吧——我根本没法预测LLMs会这么成功。所以,自回归LLMs
49:39
you know a lot of plans that people know about um that are relatively high
能变得这么厉害,你觉得合理吗?能解释一下你的直觉吗?因为如果按你智慧和直觉的
49:45
level are actually learned they're not people most people don't invent the you know plans
字面意思来理解,我会觉得自回归LLMs一次只生成一个token,根本不可能
49:50
um uh they
做到像现在这样。
49:51
they by themselves they uh you know we have some ability to do this of
我很想在你对autoregressive LLM的怀疑上多停留一会儿。我想
49:55
course uh obviously but um but but most plants that people use are plants that
用一种方式来检验这种怀疑:你说的每句话都很有道理,但如果我把你今天说的这些,
49:59
they've been trained on like they've seen other people use those plants or they've been
以及你一贯的观点,套用到比如十年前——可能没那么久,就说三年前吧——我根本预
50:03
told how to do things right um that you can't invent how you like take
测不到LLM会这么成功。那么,你觉得autoregressive LLM能变
50:07
a person who's never heard of airplanes and tell them like how do you go
得如此厉害,这合理吗?你能解释一下你的直觉吗?因为如果我把你的智慧和直觉当真
50:11
from New York to Paris and they're probably not going to be able to kind
,我会觉得autoregressive LLM一次只生成一个token,根
50:15
of you know deconstruct
本不可能做到现在这样。
50:16
the whole plan unless they've seen examples of that before um so certainly LMS are
整个计划,除非他们之前见过类似的例子,嗯,所以LMS肯定能搞定这个,但是,嗯,你怎么把这个从底层需要执行的动作,跟像Jad这样的东西联系起来呢?Jad基本上是在提升表示的抽象层级,而不试图重建每个细节。这就是为什么我们需要Jass。我很想多聊聊你对autoaggressive llms的怀疑。一个测试这种怀疑的方法是:你说的每句话都很有道理,但如果我把你今天说的这些,以及你一般
50:22
going to be able to do this but but then um how you link this
的观点,套用到比如十年前——可能没那么久,就说三年前吧——我根本预测不到llms的成功。那么,你觉得autoaggressive llms能变得这么厉害,这合理吗?是的,你能解释一下你的直觉吗?因为如果我就按字面理解你的智慧和直觉,我会觉得autoaggressive LMS一次只生成一个token,根本不可能做到那些——你知道的,像Google、Meta、Open AI等等的
50:27
from the the low level of of of actions uh that needs to be done
工作,再往前追溯到GPT那类工作,General pre-train Transformers。你是指像gbt2那样的吗?就是某个点上你开始意识到scaling可能真的会持续带来emergent benefit。嗯,我是说,确实有来自不同地方的工作,但如果你想把它放在GPT的时间线上,那大概是在gpt2左右。嗯,我就是因为你说了这个——你太有魅力了,说了这么多话——但self-
50:32
with things like like Jad that basically lift the abstraction level of the representation without
supervised learning,嗯,是的。不过,同样的直觉,你用来论证autoaggressive llms不可能对世界有深刻理解,如果我们用同样的直觉,你觉得它们能形成足够的世界表示,变得非常令人信服,基本上以高分通过最初的图灵测试,这合理吗?我们是被它们的流畅性骗了,对吧?我们只是假设如果一个系统能流畅地操控语言,那它就具备人类智能的所有特征,但这个印象是错的。我们
50:37
attempting to reconstruct every detail of the situation that's why we need Jass for I
真的被它骗了。你觉得艾伦·图灵会怎么说?在不理解任何东西的情况下,只是跟它待在一起?图灵会认为图灵测试是个很糟糕的测试。好吧,这是AI社区很多年前就达成的共识:图灵测试是个很糟糕的智能测试。那汉斯·莫拉维克会怎么说?关于那些没有针对任何特定任务训练的系统,对吧?学习表示。嗯,我14年前共同创立的那个会议叫International Conference on Learning
50:42
would love to sort of Linger on your
Representations,这就是深度学习一直在处理的核心问题,也是我痴迷了将近40年的东西。所以,学习表示真的是关键。很长一段时间里,我们只能通过监督学习做到这一点,然后我们开始研究——嗯,你懂的。
50:44
skepticism around uh autoaggressive llms so one way I would like to test that skepticism
对自回归LLM的怀疑,嗯,我想测试这种怀疑的一个方法是——你说的每句话都很有道理,但如果我把你今天说的和一般性的观点,应用到,我不知道,十年前?也许没那么久,比如说三年前,我就不会那么想。你是说,像Google、Meta、OpenAI等公司的工作,回到GPT那种工作,General Pre-train Transformers?你是指GPT-2吗?就是某个节点你开始意识到scaling可能真的会持续带来emergent benefit。是的,我的意思是,来自不同地方的工作都有,但如果你想把它放在GP
50:51
is everything you say makes a lot of sense but if I apply everything you
T的时间线上,那大概是GPT-2时期。嗯,我就是因为你说了,你太有魅力了,说了那么多词,但self-supervised learning,没错。但同样,你用来论证自回归LLM无法对世界有深刻理解的直觉,如果我们也用这个直觉来看,它们能形成足够的世界表征,以至于非常令人信服,基本以高分通过最初的图灵测试,这说得通吗?嗯,我们被它们的流畅性欺骗了,对吧?我们只是假设,如果一个系统能流畅地操控语言,那它就拥有人类的所有特征,而无需针对任何特定任务训练系统。学习表征,嗯,我14年前共同创立的会议叫Inte
50:57
said today and in general to like I don't know 10 years ago maybe a
rnational Conference on Learning Representations,这就是深度学习一直在处理的核心问题,也是我近40年来一直痴迷的东西。所以学习表征真的是关键。很长一段时间,我们只能通过supervised learning做到这一点,然后我们开始研究训练系统预测视频中会发生什么,试了又试,失败了又失败,用generative models,用预测像素的模型,我们无法让它们学到好的图像表征,也无法让它们学到好的视频表征。我们试了很多次,发表了很多论文,它们算是有一定效果,
51:04
little bit less no let's say three years ago I wouldn't be
但不算真正出色。它们开始有效了,但我们一直有这个想法,打个响指就行不通,对吧?是的,但可能还有更复杂的类似场景,一个LLM可能从未遇到过,也可能无法判断是否可能。所以,从低层到高层的那个连接,问题是,语言表达的高层是基于这样的东西,而世界的常识,我觉得就在语言里。我们没有明确表达出来,但如果你有海量的文本,你就会得到这些字里行间的东西。为了形成一个一致的世界模型,你必须理解重力如何运作,即使你没有——
51:09
able to predict the uh success of llms so does it make sense to you
比如,来自Google、Meta、OpenAI等的工作,回到GPT那种工作——你是指GPT-2吗?就是某个节点上你开始意识到scaling可能
51:16
that autoaggressive llms are able to be so damn good yes can you explain your
真的会持续带来涌现的好处。对,我是说来自不同地方的工作,但如果要放到GPT的时间线上,大概就是GPT-2左右。嗯,因为你说出来了,你太有魅力了
51:22
intuition because if I were to take your wisdom and intuition at face value I
,说了那么多词,但自监督学习——对,是的。但同样的直觉,你用来论证自回归LLMs无法对世界有深刻理解,如果我们用同样的直觉来看,它们能形成足够
51:28
would say there's no way autoaggressive LMS one token at a time would be able
的世界表征,变得极其有说服力,基本上轻松通过最初的图灵测试,这合理吗?我们是被它们的流畅性骗了,对吧?我们只是假设如果一个系统能流畅地操控语言
51:35
to do the
,那它就具备了人类的所有特征。
51:36
kind of things they're doing no there's one thing that auto llms uh or that
他们做的那些事。其实有一个东西,auto LLMs,或者说所有LLM,不只是自回归那种,也包括像BERT那种双向的,都在利用它
51:43
llms in general not just the autoaggressive one but including the birth style bir directional
,那就是self-supervised learning。我多年来一直是self-supervised learning的坚定支
51:49
ones uh are exploiting and it's self-supervised learning and I've been a very very strong
持者。所以这些东西是一个极其令人印象深刻的证明,说明self-supervised learning确实管用。这个想法,你知道,
51:56
advocate of self supervising for many years so those things are a incredibly impressive demonstration
它并不是从BERT开始的,但BERT确实是一个很好的展示。就是说,你拿一段文本,把它破坏掉,然后训练一个巨大的神经网络去重建缺
52:02
that cell supervisor learning actually
失的部分。这带来了巨大的好处,让我们能够……
52:04
works uh the idea that you know started uh it didn't start with with uh
那个自回归的技巧,就是限制系统不能看整个文本来构建文本的表示,只能根据前面的词
52:09
with Bert but it was really kind of a good demonstration with this so the
来预测下一个词。你是通过约束网络架构来实现的,这样就能构建一个自回归的模型。所
52:14
the the idea that you know you take a piece of text you corrupt it
以很多年前有个意外发现,就是所谓的decoder-only LLM。这种系统只
52:18
and then you train some gigantic neural net to reconstruct the parts that are missing
是试图根据前一个词生成下一个词,结果发现当你把它们scale up,用大量数据
52:23
um that has been an enormous uh produced an enormous amount of benefits uh it
训练,让它们变得非常大时,它们实际上开始更深入地理解语言了。这算是个意外,而且
52:28
allowed allowed us to
这个意外发生得挺早的。
52:29
create systems that understand understand language uh systems that can translate um hundreds of languages
构建能够理解语言、能够翻译数百种语言任意方向互译的系统,这些系统是多语言的,所以它们不是——这是一个单一系统,可以训练来理解数百种语言并任意方向互译,还能生成摘要,然后回答问题、生成文本。还有一个特殊情况,就是自回归技巧,你通过约束系统,不让它从整个文本中构建表示,而是只根据前面的词预测下一个词,通过限制网络架构来实现,这就是构建自回归Transformer的方法。所以很多年前有个意外发现,就是所谓的decoder-onl
52:35
in any direction systems that are multilingual so they're not it's a single system that
y LLM,这种系统只是试图从前一个词生成下一个词,当你把它们规模化,用大量数据训练,让它们变得非常大时,它们实际上对语言的理解会更深,这算是个惊喜,而且这个惊喜发生在挺久以前了,比如来自Google、Meta、OpenAI等的工作,追溯到GPT系列的工作,General Pre-trained Transformers。你是指像GPT-2那样吗?就是某个节点开始意识到规模化可能持续带来涌现的好处。是的,来自不同地方的工作
52:42
can be trained to understand hundreds of languages and translate in any direction um and
,但如果要放在GPT的时间线上,大概就是GPT-2左右。嗯,我这么说是因为你太有魅力了,说了这么多词,但自监督学习,没错。但同样的直觉,你用来论证自回归LLM不可能对世界有深刻理解,如果我们也用这个直觉,你觉得它们能形成足够的世界表征,以至于极其有说服力,基本以出色表现通过原始图灵测试吗?我们是被它们的流畅性骗了,对吧?我们假设如果一个系统能流畅地操控语言,那它就拥有人类智能的所有特征,但这个印象是错的。我们真的被它骗了。
52:48
produce summaries um and then answer questions and produce text and then there's a special
你觉得艾伦·图灵会怎么说?在不理解任何东西的情况下,只是和它相处?图灵会认为图灵测试是个很糟糕的测试。好吧,这是AI社区很多年前就决定的,图灵测试对智能来说是个很差的测试。汉斯·莫拉维克会怎么评价大语言模型?莫拉维克会说莫拉维克悖论仍然适用。好吧好吧好吧,我们可以通过。你不觉得他会非常印象深刻吗?不,当然每个人都会印象深刻,但这不是印象不印象的问题,而是要知道这些系统的极限在哪里。它们确实令人印象深刻,能做很多有用的事,整
52:54
case of
个行业都在围绕它们建立,它们会带来进步,但还有很多局限性。
52:55
it where you know you which is the auto Progressive uh trick where you constrain
比如,你知道,来自Google、Meta、OpenAI等地方的工作,可以追溯到GPT那种工作,General Pre-trai
53:00
the system to not elaborate a representation of the text from looking at the enti
ned Transformers。你是指像GPT-2那样?就是某个节点上你开始意识到scaling可能真的会持续带来emerge
53:05
text but only predicting a word from the words that are come before right and
nt benefit。对,我是说,来自不同地方的工作都有,但如果你非要把时间线放在GPT系列里,那大概就是GPT-2那个阶段。嗯
53:10
you do this by the constraining the architecture of the network and that's what you
,我就是因为你刚才说了,你太有魅力了,说了那么多词,但self-supervised learning,对,没错。但同样的直觉,
53:16
can build an auto regressive ATM from so there was a surprise many years ago
你用来论证自回归LLM不可能对世界有深层理解,如果我们用同样的直觉,你觉得它们能形成足够的世界表征,变得极其有说服力,基本上以
53:21
with what's called decoder
高分通过最初的图灵测试,这说得通吗?
53:22
only llm so since you know systems of this type that are just trying to
我们是被它们的流畅性骗了,对吧?我们就是假设,如果一个系统能流畅地操控语
53:27
produce uh words from the from the previous one and and the fact that when
言,那它就具备了人类智能的所有特征。但这个印象是错的。我们真的被它骗了。
53:32
you scale them up they they tend to really kind of understand more about the
你觉得艾伦·图灵会怎么说?在不理解任何东西的情况下,只是跟它待在一起?图
53:37
about language when you train them on lot of data and you make them really
灵会认为图灵测试是个非常糟糕的测试。好吧,这是AI社区很多年前就达成的共
53:42
big that was kind of a surprise and that surprise occurred quite a while back
识,图灵测试是个很差的智能测试。那汉斯·莫拉维克会怎么评价这些大模型呢?
53:47
like you know uh with uh work from uh you know Google meta open AI
像你知道,呃,来自 Google、Meta、Open AI 等等的工作,
53:53
Etc you know going back to you know the GPT kind of uh work General
呃,追溯到 GPT 那类工作,General Pre-train Tra
53:58
pre-train Transformers do you mean like gbt2 like there's a certain place where you start
nsformers,你是指像 GPT-2 那种吗?就是某个点上你开始意识到
54:04
to realize scaling might actually keep giving us a an emergent benefit yeah I mean
scaling 可能真的会持续带来 emergent benefit。
54:09
there were there were work from
对,我是说,有一些工作来自……
54:11
from various places but uh uh if if you want to kind of you know
从各种地方来,但呃呃,如果你想把时间线放在GPT的发展里,那大概是在GPT-2左右。嗯,我刚听你说,你太有魅力了,说了那么多词,但自监督学习,对,没错。不过同样的直觉,你用来论证自回归LLM不可能对世界有深层理解,那用这个直觉来看,它们能形成足够的世界表征,变得几乎令人信服,基本上轻松通过最初的图灵测试,你觉得说得通吗?我们被它们的流畅性骗了,对吧?我们就是假设如果一个系统在操控语言上很流畅,那它就拥
54:17
place it in the in the GPT uh timeline that would be around gpt2 yeah
有人类的所有特征,而不用为任何特定任务训练它,对吧?学习表征,嗯,我14年前联合创办的会议叫International Conference on Learning Representations,这就是深度学习一直在处理的核心问题,对吧?而且这已经是我差不多40年来的执念了,所以,学习表征真的是关键。很长一段时间里,我们只能用监督学习来做这个,然后我们开始尝试训练系统预测视频里会发生什么,试了又试,失
54:22
well I just cuz you said it you're you're so charismatic you said so many
败了又失败,用生成模型,用预测像素的模型,我们没法让它们学到好的图像表征,也没法让它们学到好的视频表征,我们试了很多次,发了很多论文,你知道,它们勉强能用,但效果不太好。它们开始有效果时,我们一直有这个想法,打个响指就行不通,对吧?嗯,但可能还有更复杂的这类场景,一个NLM可能从未遇到过,也可能无法判断是否可能,所以,那个从低层到高层的联系,关键是语言表达的高层是基于这样的东西,而那种世界的常识,我觉得
54:27
words but self-supervised learning yeah yes but again the same intuition you're applying to saying
就在语言里。我们没有明确表达出来,但如果你有海量的文本,你就会得到这些字里行间的东西,为了形成一个一致的世界模型,你必须理解重力是怎么运作的,即使你没有明确的token被生成。这意味着,每次你生成一个token,你留在正确答案集合里的概率就会下降,而且是指数级下降。所以有一个很强的假设,就像你说的,如果犯错概率非零,而看起来确实如此,那就会有一种漂移,对,而且那种漂移是指数级的,就像错误会累积,对吧?所
54:32
that autor regressive llms cannot have a deep
以,数字是巨大的,所以,无论系统经过什么样的训练来生成合适的tensor,你都可以通过找到一个超出它训练过的prompt集合或类似范围的prompt来打破它,然后它就会完全输出胡言乱语。你说prompt的时候,是指——
54:35
understanding of the world if we just apply that same intuition does it make sense
比如,来自Google、Meta、OpenAI等的工作,追溯到
54:41
to you that they're able to form enough of a representation of the world to
GPT系列的工作,也就是General Pre-trained
54:47
be damn convincing essentially passing the original touring test with flying colors well we're fooled
Transformers——你是指像GPT-2那样的吗?是不是
54:53
by their fluency right we just assume that if a system is is fluent in
在某个点上,你开始意识到scaling可能真的会持续带来涌现性的
54:59
manipulating language then it has all the characteristics of human
好处?是的,我的意思是,确实有一些工作……
55:03
intelligence but that impression is false we we we're really fooled by it um what
智能,但这种印象是错的。我们、我们真的被它骗了。嗯
55:07
do you think alen tan would say it without understanding anything just hanging out with
,你觉得艾伦·图灵会怎么说?什么都不理解,只是跟它
55:12
it an Turing would decide that a Turing test is a really bad test okay
待在一起。图灵会认为图灵测试是个很糟糕的测试。好吧
55:17
this is what the AI Community has decided many years ago that the tring test
,这就是AI社区很多年前就达成的共识——图灵测试根本
55:21
was a really bad test of intelligence what would Hans marvac say about the about
不是衡量智能的好方法。汉斯·莫拉维克会怎么看待这些
55:26
the large
大模型?
55:27
language models hence Marv would say the Marv Paradox still applies okay okay okay we
语言模型,所以Marv会说“Marv悖论”依然成
55:31
can pass you don't think he would be really impressed no of course everybody would
立。好吧好吧好吧,我们可以跳过。你不觉得他会真的被
55:35
be impressed but uh you know uh it's not a question of being impressed or
震撼到吗?不,当然每个人都会被震撼,但你知道,这
55:40
not it's a question of knowing what the limit of those systems can do like
不是被不被震撼的问题,而是要知道这些系统的极限在哪
55:44
there again they are impressive they can do a lot of useful things there's a
里。就像,它们确实令人印象深刻,能做很多有用的事,
55:48
whole industry that is being built around them they're going to make progress uh but
围绕它们已经形成了一个完整的产业,它们会不断进步
55:52
there is a lot of
,但还有很多问题。
55:54
things they cannot do and we have to realize what they cannot do do and
有些事情它们做不到,我们必须意识到它们做不到什么,然后想办法看看怎么才能达到那个目标。我这么说,其实是基于我大概十年左右在自监督学习领域的研究——实际上不止十年了,但核心就是自监督学习,也就是在没有为特定任务训练系统的情况下,捕捉一组输入数据的内部结构,对吧?就是学习表征。我14年前联合创办的那个会议叫国际学习表征会议,这其实就是深度
55:58
uh and then figure out you know how we get there and you know and
学习一直在处理的核心问题,也是我差不多40年来一直痴迷的东西。所以学习表征真的是关键。很长一段时间里,我们只能通过监督学习来做这件事,然后我们开始研究以前所谓的无监督学习,差不多在2000年代初期和Yoshua Bengio一起重新激活了这个概念。后来Jeff Hinton发现,如果你能收集到足够多的数据,监督学习其实效果很好,所以无
56:03
and I'm not seeing this I'm seeing this from basically you know 10 years of
监督和监督这种二分法就暂时退居二线了。然后我大概在2014年,也就是我们创办FAIR的时候,开始大力尝试重新推动这个方向,努力寻找新的自监督学习方法,既用于文本,也用于图像、视频和音频。其中一些工作取得了巨大的成功。比如说,我们现在之所以有能够做内容审核的多语言翻译系统——比如在Meta的Facebook上,能判断一段文本是不是仇恨言
56:08
of research uh on on the IDE of sell supervis learning actually that's going back
论之类的——就是因为自监督学习在NLP领域的进展,再加上Transformer架构等等。这就是自监督学习的重大成功。我们在语音识别方面也有类似的成功,有一个叫wav2vec的系统,它也是一种联合嵌入架构,通过对比学习训练,这个系统可以用大部分未标注的数据生成多语言的语音识别系统,只需要几分钟的标注数据就能真正做语音识别,这太惊人了。我
56:13
more than 10 years but the IDE of cell supervis learning so basically capturing the
们现在基于这些想法的组合,已经有了能实时翻译几百种语言的系统,语音到语音,甚至包括那些没有书面形式的语言——没错,就是只有口语的语言。我们不经过文本,直接从语音到语音,使用一种离散的语音单元的内部表征,这个我们以前叫textless NLP。所以,那里取得了难以置信的成功。然后,大概有十年时间,我们试图把这个想法应用到图像表征学习上,通
56:18
internal structure of a piece of uh of of a set of inputs
过训练系统预测视频来学习直观物理,训练系统预测视频里接下来会发生什么,试了又试,失败了又失败,用生成模型、用预测像素的模型,都没能让它们学到好的图像表征,也没能让它们学到好的视频表征。我们试了很多次,发了很多论文,它们算是有点效果,但算不上真正好。后来它们开始起作用了,我们一直有这个想法——
56:22
without training the system for any particular task right learning representations um you know the
就像,你知道,来自Google、Meta、Open
56:27
the conference I co-founded 14 years ago is called inter International Conference on learning representations
AI等等的工作,回到GPT那种工作,General
56:32
that's the entire issue that deep learning is is dealing with right and it's been
Pre-train Transformers,你是指
56:37
my obsession for you know almost 40 years now so um so learning representation is
像GPT-2那样?有一个节点你开始意识到scalin
56:41
really the thing uh for the longest time we could only do this with supervised
g可能真的会持续带来emergent benefit
56:46
learning and then we started working on uh you
。是的,我是说,有些工作来自……
56:49
know what we used to call unsupervised learning uh and sort of revive the idea
没有针对任何特定任务训练系统,对吧?学习表征。你知道,我14年前联合创办的那个会议叫“国际学习表征会议”,这其实就是深
56:55
of unsupervised learning uh in the early 2000s with yosha benju and Jeff Hinton then
度学习一直在处理的核心问题,也是我近40年来一直痴迷的东西。所以,学习表征才是关键。很长一段时间里,我们只能通过监督学习
57:01
discovered that supervisor leing actually works pretty well if you can collect enough data and
来做这件事,然后我们开始研究我们过去称之为无监督学习的东西,并在21世纪初与Yoshua Bengio和Geoff Hi
57:07
so the whole idea of you know unsupervised supervisor kind I took a a backseat
nton一起重新复兴了无监督学习的理念。后来发现,如果你能收集足够多的数据,监督学习其实效果很好,所以无监督学习和监督学
57:13
for for a bit
习这种分类方式就暂时退居二线了。
57:15
and then I kind of tried to revive it um uh in a big way
然后我其实尝试过,呃,以一种很大的方式重新推动它,基本上是从2014年开始,也就是我们创立FAIR的时候,呃,然后大力推动寻找新的方法来做自监督学习,既针对文本,也针对图像、视频和音频。其中一些工作取得
57:19
you know starting in 2014 basically when we started fair and uh and really pushing
了巨大的成功。我是说,我们现在之所以有多语言翻译系统,你知道,还有那些能够用大部分未标注数据、只需要几分钟标注数据就能做语音识别的多语言语音识别系统,这真的很惊人。我们现在有一些基于这些组合想法的系统,能
57:24
for like finding new new methods to do cell supervised learning both for text and
够实时翻译几百种语言,互相之间进行语音到语音的转换,甚至包括那些没有书面形式的语言——没错,只有口语的语言。对,我们不经过文本,直接从语音到语音,使用一种离散的语音单元的内部表示,它叫做textless
57:29
for images and for video and audio and some of that work has been incredibly
NLP,我们以前是这么叫的。嗯,所以,这真的是巨大的成功。然后,在十年里,我们尝试把这个想法应用到学习图像表示上,通过训练系统来预测视频,学习直观的物理规律,通过训练系统预测视频里接下来会发生什么。试了
57:34
successful um I mean the reason why we have multilingual translation system you know things
又试,失败了又失败,用生成模型、用预测像素的模型,我们就是没法让它们学到好的图像表示,也没法让它们学到好的视频表示。我们试了很多次,发了很多论文,你知道,它们算是有点效果,但并不是真的很好。后来它们开始工
57:39
to do
作了,我们一直有这个想法……
57:40
content moderation on on meta for example on Facebook that are multilingual that understand whether
然后我试图以更大的力度重新推动它,基本上从2014年开始,当我们创立FAIR时,就大力推动寻找新的自监督学习方法,既用于文本,也用于图像、视频和音频。其中一些工作取得了巨大的成功。我的意思是,我们现在之所以有多语
57:44
piece of text is H speech or not or something is due to their progress
言翻译系统,比如在Meta(Facebook)上做内容审核的系统,能够理解一段文本是否是仇恨言论,这都归功于自监督学习在NLP领域的进展,再结合Transformer架构等等。但这就是自监督学习的巨大成功。我们在语
57:49
using cell supervis learning for NLP combining this with you know Transformer architectures and and
音识别方面也取得了类似的成功,一个叫wav2vec的系统,顺便说一句,它也是一种联合嵌入架构,通过对比学习训练。这个系统也能生成多语言的语音识别系统,主要使用未标注数据,只需要几分钟的标注数据就能实际进行语音识别,
57:54
blah blah blah but that's the big success of supervis rning we had similar success
这太神奇了。我们现在有基于这些想法组合的系统,能够实时将数百种语言互相翻译,包括语音到语音,甚至包括那些没有书面形式的语言,没错,只有口语的语言。我们不需要经过文本,直接从语音到语音,使用一种离散的语音单元的内部表
57:59
in speech recognition a system called wave to V which is also a joint embedding
征。它以前叫textless NLP,但总之,这真是巨大的成功。然后,在十年里,我们试图将这个想法应用于学习图像的表征,通过训练系统预测视频来学习直观物理,尝试了又尝试,失败了又失败,用生成模型、用预测像素的模型,
58:04
architecture by the way train with contrastive learning and and that that
我们就是无法让它们学到好的图像表征,也无法让它们学到好的视频表征。我们试了很多次,发表了很多论文,它们算是有点效果,但并不是真正出色。直到它们开始奏效,我们一直有这个想法。
58:08
system also can produce um speech recognition systems that are multilingual with mostly unlabeled data
没有针对任何特定任务训练系统,对吧?学习表征。嗯,我14年
58:14
and only need a few minutes of labeled data to actually do speech recognition that's
前联合创办的那个会议叫国际学习表征会议,这就是深度学习一直在
58:19
that's amazing um we have systems now based on those combination of ideas that can
处理的核心问题。而且这差不多是我40年来的执念了。所以,学
58:25
do realtime translation of hundreds of languages into each other uh Speech to speech speech
习表征才是关键。很长一段时间里,我们只能通过监督学习做到这一
58:31
to speech even including which is fascinating languages
点,然后我们开始研究无监督的……
58:34
that uh don't have written forms that's right they spoken only that's right we don't
语言和观点在语言空间里,但并不以视觉方式呈现,而且显然以一种可压缩的方式。嗯,有很多情况,可能对于一个纯语言系统来说很难知道
58:38
go through text it goes directly from from speech to speech using an internal representation
,比如,你可以从阅读文本中学到全世界所有公开可用的文本,但我不能靠打个响指就从纽约到巴黎,对吧?那行不通。嗯,但可能还有更复
58:42
of kind of speech units that are discrete but it's um it's called text lesson
杂的这类场景,一个NLM可能从未遇到过,也无法判断它是否可能。所以,那个从低层到高层的连接,问题是语言所表达的高层是基于低层
58:47
LP we used to call it this way but um yeah so that I mean
的共同经验的,而LLMs目前没有这种经验。你知道,当我们彼此交谈时,我们知道我们对世界有共同的体验,比如很多方面是相似的。而
58:51
incredible success there and then you know for 10 years we tried to apply this
LLMs没有这个。但你看,你和我对世界有共同的体验,比如重力如何作用的物理规律之类。这种对世界的共同知识,我觉得是存在于语言
58:55
idea to learning representations of images by training a system to predict videos learning intuitive
中的。我们不会明确表达它,但如果你有大量的文本,你会得到这些字里行间的东西。为了形成一个一致的世界模型,你必须理解重力如何作
58:59
physics by
用,即使你没有……
59:00
training a system to predict what's going to happen in the video and tried and
如果我们把同样的直觉应用到对世界的理解上,
59:05
tried and failed and failed with generative models with models that predict pixels uh we
你觉得它们能形成足够的世界表征,从而变得极
59:10
could not get them to learn good well presentations of images we could not get
其有说服力,基本上以优异的成绩通过最初的图灵
59:15
them to learn good well presentations of videos and we tried many times we published
测试,这说得通吗?我们被它们的流畅性骗了,
59:19
lots of papers on it you know they kind of sort of work but not
对吧?我们只是假设,如果一个系统能流畅地操
59:24
really great they started working we we have been this idea of
控语言,那它就具备了人类的所有特征。
59:28
predicting every pixel and basically just doing the joint embedding and predicting in representation space
预测每一个像素,基本上就是做联合嵌入,然后在表示空间里做预测,这确实有效。所以有大量证
59:34
that works MH so there's ample evidence that we're not going to be able to
据表明,我们无法通过生成模型学到关于真实世界的好表示。所以我跟别人说,大家都在谈生成式
59:40
learn good we representations of the real world using generative model so I'm telling people
AI,但如果你真的对接近人类水平的AI感兴趣,就放弃生成式AI这个想法吧。不过你真的觉
59:45
everybody is talking about generative AI if you're really interested in human level AI abandon
得,靠联合嵌入表示能走得很远吗?比如常识推理和高级推理——我感觉这两类推理,LLM能做
59:51
the idea of generate AI okay but you you you really think
到的那种“推理”,跟我们日常用来理解世界的常识推理,本质上是不一样的。
59:56
it's possible to get far with the joint embedding representation so like there's Common Sense
用joint embedding representation可以走得很远,比如有Common Sense reasoning,然后还有high-level reasoning。我觉得这两类是LM能做的推理——算了,我不用“推理”这个词——但LM能处理的东西,跟我们用来在现实世界里导航的那种common sense reasoning,本质上好像很不一样。语言和观点在语言空间里存在
60:01
reasoning and then there's highlevel reasoning like I I feel like those are two the
,但不会用视觉方式呈现,而且显然有一种可压缩的方式对吧?其实有很多情况,纯语言系统可能很难知道——比如你可以从阅读文本中学到全世界所有公开可用的文本,但我不能靠打个响指就从纽约到巴黎,这行不通对吧?是的。但可能还有更复杂的类似场景,一个nlm可能从来没遇到过,也没法判断它是否可行。所以那个从low level到high level的链接,问题在于语言表达的high level是基于l
60:06
kind of reasoning that LMS are able to do okay let me not use the
ow level的共同经验,而llms目前没有这种经验。我们彼此交流时,我们知道我们对世界有共同经验——很多都是相似的——而LMs没有。但你看,你和我对世界有共同经验,比如重力怎么运作之类的物理常识。那种共同的世界知识,我觉得就藏在语言里。我们不会明确说出来,但如果你有海量的文本,你会从字里行间得到这些东西。为了形成一个consistent的世界模型,你必须理解重力怎么运作,哪怕没有
60:12
word reasoning but the kind of stuff that LMS are able to do seems fundamentally
对重力的明确解释。即使重力确实有明确的解释和维基百科,但那些我们认为是common sense reasoning的东西,我觉得为了正确生成语言,你都得搞清楚。你可以说,就像你说的,文本不够多。所以你不这么认为?不,我同意你刚才用语言表达的观点。这对你来说很明显吗?对,完全同意。所以所有我们进行的对话——还有Dark web,意思是那些私人对话,比如DMs之类的,可能比llms训练
60:17
different than the common sense reasoning we use to navigate the world
用的数据量大得多。你不需要把那些隐含的东西说出来,但幽默什么的都会流露出来。不,你确实需要——不是必须,但会通过内容体现出来。比如我不小心打翻了这个,你可能会取笑我,而取笑的内容里就会包含杯子会掉下来、重力这么运作的解释,然后你会有些模糊的信息,知道什么东西掉到地上会碎,然后你可能还会开个关于熵的玩笑之类的。
60:21
yeah it seems like we're going to need both you're not would you be able
对,看起来我们两者都需要。但靠联合嵌入,比如JEA那种方法,通过看视频,你能学到怎么从纽约到巴黎吗?或者理解当今世界的政治局势?这些事人类会生成大量语言和观点,但并没有视觉上的表示,而且显然是可以压缩
60:27
to get with the joint embedding would is the JEA type of approach looking at
的。确实有很多情况,纯语言系统可能很难知道——比如你可以从阅读全世界所有公开文本中学到,“我不能靠打个响指就从纽约到巴黎”,这显然不对,对吧?但可能还有更复杂的类似场景,一个NLM可能从未遇到过,也无法
60:33
video would you be able to learn let's see well how to get from New
判断它是否可行。所以从低层到高层的那个连接,问题在于语言所表达的高层内容,是基于低层的共同经验的,而LLM目前没有这种经验。我们彼此交流时,知道我们对世界有共同经验——很多是相似的——但LLM没有。但
60:38
York to Paris or um how to uh understate understand the state of politics in
你看,你和我对世界的共同经验,比如重力怎么运作之类的物理常识,我觉得这种共同知识其实就在语言里。我们不会明确说出来,但如果你有海量的文本,你就能从字里行间得到这些东西。为了形成一个一致的世界模型,你必须
60:44
the world today right these these are things where various humans generate a lot of
理解重力怎么运作,即使没有对重力的明确解释。虽然重力确实有明确的解释,但那些我们认为是常识推理的东西,我觉得要正确生成语言,你就得搞明白。你可以说,文本还不够多。那你觉得不是这样?不,我同意你刚才说的。
60:49
language and opinions on in the space of language but don't visually represent that and
这个系统还能生成多语言的语音识别系统,大部分数
60:54
you clearly uh compressible way right well there's a lot of situations that you know
据是未标注的,只需要几分钟的标注数据就能实际做
61:00
might be difficult to for a purely language based system to um to know like
语音识别,这太惊人了。我们现在有基于这些思路组
61:05
okay you can probably learn from Reading text the entirety of the Public Public avilable
合的系统,可以实时翻译几百种语言之间的互译,语
61:11
text in the world that I cannot get from New York to Paris
音到语音,甚至包括那种特别有意思的语种……
61:15
by snapping my fingers that's not going to work right yes uh but there's you
打个响指就行不通,对吧,嗯,但可能还有更复杂的类似场景,NLM 可能从未遇到过,也无法判断是否可行。所以,从底层到高层的那个连接
61:20
know probably sort of more complex uh scenarios of this type which an nlm May
——关键在于,语言所表达的高层内容,并没有产生预期的结果,在这种情况下,我们会用 RL 来调整世界模型或 critic。对,你提
61:26
never have encountered and may not be able to determine whether it's possible or not
到了 RLHF,那为什么你还是讨厌强化学习呢?我并不讨厌强化学习,我觉得它不应该被完全抛弃,但我认为它的作用有限。而在那之前,3
61:31
um so um so that that link you know from the the low level to
0年前,当我们研究 Com Nets 和神经网络早期阶段的时候,我非常兴奋,因为我看到了一条通往人类级别智能的道路——系统能够理
61:36
the high level the the thing is that the high level that language expresses is
解世界、记忆、规划、推理。还有哲学理性主义、摆脱宗教教条、民主、科学——当然,没有这些就不会有美国革命和法国革命,我们可能还在被
61:41
based
统治之下。
61:42
on the common experience of the low level which llms currently do not have you
训练一个系统去预测视频里接下来会发生什么,试了又试,失
61:47
know we when we talk to each other we know we have a common experience
败了又失败,用生成模型、用预测像素的模型,我们就是没法
61:52
of the of the world like you know a lot of it is is similar
让它们学到好的图像表征,也没法让它们学到好的视频表征。
61:57
uh and LMS don't have that but see there it's present you and I have
我们试了很多次,发了很多论文,你知道的,它们勉强能工作,
62:02
a common experience of the world in terms of the physics of how gravity works
但效果并不好。后来它们开始真正起作用了,我们一直有这个
62:07
and stuff
想法……
62:08
like this and that common knowledge of the world I feel like is there in
像这种关于世界的常识,我觉得其实已经存在于语
62:14
the language we don't explicitly express it but if you have a huge amount of
言里了——我们不会明确说出来,但如果你有海量的
62:20
text you're going to get this stuff that's between the lines you're going to you're
文本,你就能从字里行间捕捉到这些东西。为了形成
62:26
going in order to um form a consistent world mod you're going to have to
一个连贯的世界模型,你不得不理解重力是怎么运
62:32
understand how gravity works even if you don't have an
作的,哪怕你没有一个明确的...
62:36
explicit explanation of gravity so even though in the case of gravity there is explicit
关于引力的明确解释,所以即使在引力的例子里,有
62:41
explanations of gravity and wiia but uh you're like the stuff that we think of
对引力和wiia的明确解释,但那些我们认为是常
62:46
as common sense reasoning I feel like to generate language correctly you're going to have
识推理的东西,我觉得要正确生成语言,你必须要搞
62:51
to figure that out now you could say as you have there's not enough text
明白这些。你可能会说,就像你之前提的,文本量不
62:56
okay so what you don't think so no I agree with what you just
够,所以你不这么认为?不,我同意你刚才说的。
63:00
said which is that to be able to do high LEL um uh common sense
也就是说,要做到高层次的常识,你得先有低层次的常识作为基础,对吧?但这个东西在LLM里是不存在的,LLM纯粹是从文本里训练出来的。那你说的另一个观点,我不太同意,就是认为所有语言里都隐含了底层现实。其实有很多关于底层现实的东西,并没有在语言中表达出来,这你明
63:05
to have high level common sense you need to have the low level common sense
白吧?完全明白。就像我们所有的对话,还有那个暗网,就是那些私聊、私信之类的,这些数据量可能比LLM训练用的数据大得多。你不需要把那些隐含的东西说出来,但幽默啊、各种信息啊,都会通过对话传递出来。比如我不小心把这个打翻了,你可能会取笑我,而在你取笑我的内容里,就
63:09
to build on top of yeah um but that's not there and that's not there
会包含对杯子掉落的解释,还有重力是怎么起作用的,然后你可能会有些模糊的信息,知道什么东西掉地上会碎,甚至可能开个关于熵的玩笑。但这些东西你永远没法再完整还原出来。你开个小玩笑,还有成千上万个类似的玩笑,从这些玩笑里你可以拼凑出重力存在、杯子会碎这些事实。你不需
63:13
in llms llms are purely trained from Tex so so then the other statement you
要亲眼看到,那样效率太低了,还不如别打翻东西来得直接。但我觉得,如果你有足够多的这类数据,这些信息其实是存在的。我只是觉得,我们小时候积累的大部分这类信息,在文本里、在任何描述里,基本上都不存在。感官数据才是获取这种理解的更丰富来源。一个四岁小孩清醒的16,0
63:18
made um I would not I would not agree with the fact that implicit in
00小时,通过视觉就有10的15次方字节的数据量,触觉也有类似的带宽,听觉稍微少一点,而语言要到一岁左右才出现。到九岁的时候,你已经学会了重力、惯性、稳定性,知道了有生命和无生命物体的区别。到18个月的时候,你就知道别人为什么想做某些事,如果他们做不到你会帮忙
63:22
all languages in the world is the underlying reality there's a lot about underlying reality
。很多东西主要是通过观察学来的,甚至不需要互动。婴儿在头几个月里对世界几乎没什么影响,只能观察,但就凭这个,你积累了海量的知识。这就是我们当前AI系统缺失的东西。我记得你有一张幻灯片里有个很好的图,展示了LLM的局限性。你能从你的角度谈谈幻觉吗?为什么大语言模
63:26
which is not
型会产生幻觉?这在多大程度上是大语言模型的根本缺陷?
63:27
expressed in language is that obvious to you yeah totally so like all all the
用语言表达出来对你来说很明显对吧,对,完全是这样。所以我们所有的对话,好吧,还有暗网上的内容,比如私信之类的,这些数据量可能比LLM训练用的数据大得多。你不需要去
63:33
conversations we have what okay there's the Dark web meaning uh whatever the private conversations
传达那些即将出现的东西,但幽默感、所有那些东西——如果你有足够多的这类数据,我只是觉得我们小时候积累的大部分这类信息,在文本里基本上都不存在,没有任何描述能涵盖。
63:39
like DMS and stuff like this which is much much larger probably than what's available
而感官数据是获取这种理解的丰富得多的来源。我是说,一个四岁孩子清醒时的16,000个小时,还有10的多少次方……以及无生命的物体,你知道,到18个月大时,你就能理
63:44
what what llms are trained on you don't need to communicate the stuff that is
解为什么人们想做某些事,如果他们做不到你会帮忙。我是说,有很多东西你主要是通过观察学到的,真的,甚至不是通过互动。在生命的最初几个月,婴儿对世界其实没有任何影响力,
63:50
coming but the humor all of it
他们只能观察,对吧?然后你就从那里积累了海量的知识。所以那就是我们正在做的。
63:53
no you do like when you you don't need to but it comes through through
语言中表达的东西对你来说是显而易见的
63:57
like you like if I accidentally uh knock this over you'll probably make fun of
吗?是的,完全如此。所以就像我们所有的
64:02
me and in the content of the you making fun of me will be a
对话,还有暗网,意思是那些私密的对话,
64:06
explanation of the fact that cups fall and then you know gravity Works in this
比如私信之类的,这些数据量可能比LLM
64:10
way and then you you'll have some very vague information about what kind of things
训练用的数据大得多。你不需要去传达那
64:15
explode when they hit the ground and then maybe you'll make a joke about entropy
些隐含的东西,但幽默什么的都会自然流露
64:19
or something
出来。
64:20
like this and you will'll never be able to reconstruct this again like okay you
像这样,你就再也无法重建它了。比如你开个小玩笑,就会有无数个其他笑话,从这些笑话里,你能拼凑出重力是存在的、杯子会碎这类事实。你不需要亲眼看到——那样效率太低了,还不如别把东西碰倒。但我觉得,如果你有足够多的这类数据,它确实会存在。我只是觉得,我们小时候积累的大部分这类信息,在文本里、在任何描述中,
64:25
make a a little joke like this and there'll be trillion of other jokes and
基本上都不存在。而感官数据是获取这种理解的丰富得多的来源。我是说,一个四岁孩子清醒时的16,000小时,以及10的15次方字节,光是视觉就这么多。触觉也有类似的带宽,音频少一点,而文本——语言——要到生命中的第一年才出现。等你9岁时,你已经学会了重力、惯性、稳定性,还知道有生命和无生命物体的区别。到1
64:30
from the jokes you can piece together the fact that gravity works and mugs can
8个月大时,你就知道人们为什么想做事情,如果他们做不到你会帮忙。我是说,很多事情你主要是通过观察学到的,甚至不是通过互动。在生命最初的几个月里,婴儿对世界几乎没有任何影响,他们只能观察。而仅仅从观察中,你就积累了海量的知识。所以,这就是我们当前AI系统所缺失的。我记得你有一张幻灯片里有个很好的图,展示
64:35
break and all this kind of stuff you don't need to see uh it'll be
了LLM的局限性。我想知道你能不能从你的角度谈谈幻觉问题——为什么大语言模型会产生幻觉,以及这在多大程度上是大语言模型的根本缺陷。所以,每次生成一个token,你停留在正确答案集合内的概率就会下降,而且是指数级下降。所以这里有一个很强的假设:如果存在非零的犯错概率——而看起来确实存在——那么就会有一种
64:40
very inefficient it's easier for like to not knock the thing over yeah but uh
漂移。是的,这种漂移是指数级的,错误会累积。所以,答案变得毫无意义的概率会随着token数量的增加而指数级增长。这在你看来是不是很明显?嗯,从数学上讲也许是的,但难道不是有一种朝向真理的引力吗?因为平均来说,希望真理在训练集中有很好的体现。不,这基本上是在与维度灾难作斗争。纠正这个问题的方法是,通过让
64:44
I feel like it would be
系统回答人们可能提出的各种问题来对它进行微调。而人们就是人们,他们提出的很多问题都非常相似,所以你大概能覆盖80%的情况。
64:46
there if you have enough of that data I just think that most of the
如果你有足够多的那种数据,我只是觉得我们小时候积累的大部分这类信息,根本就没出现在文本里,或者说任何描述里,而感官数据是获取那种理解的更丰富来源。我的意思是,一个四岁孩子清醒时的16,000小时,以及10的多少次方个token被生成。这意味着每次你生成一个token,你留在正确答案集合里的概率就会下降,而且是指数级下降。所以有一
64:51
information of this type that we have accumulated when when we were babies is just
个很强的假设,就像你说的,如果犯错的概率不是零——而看起来确实不是零——那么就会有一种漂移。没错,这种漂移是指数级的,错误会累积,对吧?所以这个数字是巨大的。因此,无论系统经过什么样的训练来生成合适的tensor,你都可以通过找到一个超出它训练过的prompt集合或类似范围的prompt来破坏它,然后它就会输出一堆胡言乱语。你提到
64:56
not present in uh in in text any in any description essentially and the sensory
prompt的时候,是指什么?就是token的数量。所以本质上,不管问题本身是简单回答、复杂回答,还是因为不可判定而无法回答,系统能用来回答的计算量是恒定的,或者与回答中生成的token数量成正比。你做得足够多之后,就可以下意识地完成,不用思考。如果你是个有经验的司机,你可以不用真的想就能开车,同时还能跟人聊天或者听收音机。如果
65:01
data is much is a much richer source for getting that kind of understanding I
你是个非常有经验的棋手,你可以跟一个没经验的棋手下棋,也不用怎么思考,你只是识别出模式然后下棋,对吧?这就是系统一。所有你本能地做、不需要刻意计划和思考的事情。然后还有一类任务,你需要计划。如果你不是很有经验的棋手,或者你是有经验的但跟另一个有经验的棋手下棋,你会考虑各种选项,对吧?你会想一会儿,如果你有时间思考,你会比下快棋时表
65:06
mean that's the 16,000 hours of of wake time of a four-year-old and uh 10
现好得多,因为时间有限。所以这种刻意的规划,利用你内部的世界模型,这种系统二,是目前LLM做不到的。那么,我们怎么让它们做到这一点呢?我们怎么构建一个系统,能够进行这种规划或推理,把更多资源分配给复杂问题而不是简单问题?而且这不会是自回归的token预测,而更像是推理潜在变量,就像以前所谓的概率模型或图模型那种东西。所以基本上原
65:11
to the
理是这样的:你知道prompt就像是一个...
65:12
15 bytes you know going through vision just Vision right there is a similar uh
不,你确实会——你不需要刻意去做,但它会通过某种
65:17
bandwidth you know of touch and uh a little less through audio and then text
方式体现出来。比如我不小心打翻了杯子,你可能会取
65:22
doesn't language doesn't come in until like you know year uh in in life and
笑我,而取笑的内容里就会包含对杯子掉落事实的解释
65:27
by the time you are 9 years old you've learned about gravity you know about
,然后你知道引力是这样运作的,接着你会有些模糊的
65:33
inertia you know about gravity you know the stability you know you know about the
信息关于什么东西掉到地上会碎,然后你可能会开个关于
65:38
distinction between animate
熵的玩笑。
65:39
and inanimate objects you know by 18 months you know about like uh why people
生成token,嗯,这意味着每次你生成一个t
65:43
want to do things and you help them if they can't you know I mean
oken,你保持在正确答案集合内的概率就会下
65:48
there's a lot of things that you learn mostly by observation really uh not even
降,而且是指数级下降。所以就像你说的,有一个
65:53
through interaction in the first few months of life babies don't don't really have any
很强的假设:如果犯错的概率非零——而看起来确
65:57
influence on the world they can only observe right and you accumulate like a gigantic
实如此——那么就会有一种漂移。对,而且那种漂移
66:02
amount of uh of knowledge just just from that so that that's what we're
是指数级的,就像错误会累积,对吧?所以……
66:06
missing from uh current AI systems I think in one of your slides you have
如果有足够多的这类数据,我只是觉得我们小时候积累的大部分这类信息,基本上在文本里、在任
66:12
this nice plot that is one of the ways you show that llms are limited
何描述里都不存在。而感官数据是获取这种理解的更丰富来源。我是说,一个四岁孩子清醒的16
66:18
I wonder if you could talk about hallucinations from your perspectives the why hallucinations happen
,000小时,以及10的15次方字节通过视觉输入——仅仅是视觉。触觉也有类似的带宽,音频
66:24
from large language models and why and to what degree is that a fundamental flaw
少一点,而文本、语言直到你生命中的某个年份才出现。到9岁时,你学会了引力、惯性、稳定性
66:30
of large language models right so
,你知道了有生命和无生命物体的区别。
66:33
because of the auto regressive prediction every time an LM produces a token or word
由于自回归预测的特性,每次LM生成一个token或单词时,都有一定概率让这个词把你带出合理答案的集合。如果你假设——这是一个非常强的假设——这种错误的概率在生成一串token的过程中是相互独立的,那意味着每次生成一个token时,你留在正确答案集合内的概率都会下降,而且是指数级下降。所以这里有一个很强的假设,就是如果存在非零的犯错概率——而看起来确实存在——那就会产生某种漂移,没错,而且这种漂移是指数级的,错误
66:39
uh there is some level of probability for that word to take you out of
会累积,对吧?所以答案变得毫无意义的概率会随着token数量指数级增长。顺便问一下,你觉得这个道理明显吗?从数学上讲可能确实如此,但难道不存在一种向真理的引力吗?因为平均来说,希望真理在训练集中有很好的体现。不,这本质上是在对抗维度灾难。所以修正这个问题的方法是,你通过让系统回答人们可能提出的各种问题来对它进行fine-tuning。而人就是人,他们提出的很多问题其实非常相似,所以你大概能覆盖人们会问的80%或多
66:45
the set of reasonable answers uh and if you assume which is a very strong
少的问题,通过收集数据,然后你对系统进行fine-tuning,让它对所有这些都能给出好的答案,而且它很可能能学会,因为它的学习能力很强。但问题是,还有海量的prompt是你在训练中没有覆盖到的,这个集合极其庞大。在所有可能的prompt中,用于训练的prompt比例绝对小得可怜,它只是所有可能prompt中极小极小的一个子集。所以系统在它经过pre-training或fine-tuning的prompt上会表
66:52
assumption that the probability of such error um is that those errors are independent across
现正常,但还有一整个空间的东西它根本不可能被训练过,因为那个数量太庞大了。所以无论系统经过什么样的训练来生成合适的tensor,你都可以通过找到一个超出它训练过的prompt集合或相似范围的prompt来打破它,然后它就会输出一堆胡言乱语。你刚才说prompt的时候,是指那个确切的prompt,还是指一个在很多方面都非常不同的prompt?比如,在互联网上问一个以前没人问过的问题或说一件没人说过的事,容易吗?我是
66:58
a a sequence of
说,人们已经想出了一些东西,比如你在prompt里放一串随机的字符,就足以把系统扔进一个模式,让它开始胡言乱语。
67:00
tokens being produced M what that means is that every time you produce a token
在没有为任何特定任务训练系统的情况下,对吧?学习表征。你知道,我14年
67:04
the probability that you rest you you stay within the the set of correct answer
前共同创立的会议叫International Conference o
67:09
decreases and it decreases exponentially so there's a strong like you said assumption there that
n Learning Representations,这就是deep l
67:14
if uh there's a non-zero probability of making a mistake which there appears to be
earning要处理的核心问题,也是我痴迷了将近40年的东西。所以学习
67:19
then there's going to be a kind of drift yeah and that drift is exponential
表征真的是关键。很长一段时间我们只能通过supervised learn
67:24
it's like errors accumulate right so so the
ing做到这一点,然后我们开始研究……
67:27
probability that an answer would be nonsensical increases exponentially with the number of tokens is
答案毫无意义的概率会随着token数量呈指数级增长——这
67:33
that obvious to you by the way like well so mathematically speaking maybe but like
个你本来就觉得很明显吗?呃,从数学角度说可能确实如此,但
67:40
isn't there a kind of gravitational pull towards the truth because on on average hopefully
难道不存在一种朝向真理的引力吗?因为平均而言,希望真理在
67:47
the truth is well represented
训练集中有充分体现。
67:50
in the uh training set no it's basically a struggle against uh the curse of
在训练集里,不,这基本上是在对抗维度
67:55
dimensionality so the way you can correct for this is that you fine-tune the system
灾难。所以纠正这个问题的方法是,你通过
68:00
by having it produce answers for all kinds of questions that people might come up
让系统为人们可能提出的各种问题生成答案
68:05
with M and people are people so they a lot of the questions that they
来微调它。而人就是人,所以他们很多问题
68:10
have are very similar to each other so you can probably cover you know 80%
都非常相似,所以你大概能覆盖80%或者
68:15
or
……
68:15
whatever of questions that people will will ask um by you know collecting data and
人们会问的各种问题,就是通过收集数据,然后微调系统,让它对所有这些事情都能给出好的回答。它很可能能学会这些,因为它的学习能力很强。但还有大量你在训练中没有覆盖到的prompt,这个数量是极其庞大的。在所有可能的prompt里,被用来训练的比例简直微乎其微,只是所有可能prompt中极小极小的一部分。所以系统在它被pre
68:20
then um and then you fine tune the system to produce good answers for all
-trained或fine-tuned过的prompt上表现正常,但还有一整片它根本不可能被训练过的空间,因为那个数量实在太大了。因此,无论系统经过怎样的训练来生成合适的tensors,你都可以通过找到一个它没训练过或与之不相似的prompt来打破它,然后它就会输出一堆胡言乱语。你说的prompt,是指那个确切的pro
68:26
of those things and it's probably going to be able to learn that because it's
mpt,还是指在很多部分上非常不同的prompt?比如,问一个在互联网上从未出现过的问题或说一句从未有过的话,这容易吗?我的意思是,有人做过一些事,比如在prompt里放一串随机的字符,就足以让系统进入一种模式,然后它就开始胡说八道。真的那么容易就能搞垮它吗?是的,有人做过类似的事:你用英文写一个句子或问一个问题,它给
68:31
got a lot of capacity to to learn um but then there is you know
出一个完全正常的回答,然后你只是把其中几个词换成另一种语言的同一个词,突然之间回答就变成了一堆胡言乱语。所以我想说的是,大多数人会问的问题,这个长尾实在太长了,你不可能让系统在所有条件下都表现良好。最终,系统本质上就像一个巨大的查找表,这其实不是我们想要的。我们想要的是能推理、能规划的系统。而LLM中发生的推理是非常非常
68:36
the enormous set of prompts that you have not covered during training and that set
原始的,你能看出它原始的原因是,它处理的token数量是固定的。所以,无论问题简单、复杂还是不可判定,系统能用来回答的计算量是恒定的,或者与回答中产生的token数量成正比。这是否意味着LLM有一个根本性的缺陷?还是说这个问题还有更多层面?你现在表现得就像个LLM一样,立刻回答“不”。这只是底层的世界模型,我们可以在它
68:41
is enormous
之上构建一些机制,比如你提到的持久长期记忆。
68:42
like within the set of all possible prompts the proportion of prompts that have been
不,这本质上是在对抗维度诅咒。要修正这个
68:48
uh used for training is absolutely tiny um it's a it's a tiny tiny tiny
问题,你可以通过让系统针对人们可能提出的
68:54
subset of all possible prompts and so the system will behave properly on the prompts
各种问题生成答案来进行fine-tuni
68:59
that it's been either trained pre-trained or fine-tuned um but then there is an entire
ng。而人类嘛,他们提出的问题很多都非常
69:05
space of things that it cannot possibly have been trained on because
相似,所以你大概能覆盖80%左右。
69:10
it's just the the number is gigantic so um so whatever training the system uh
每生成一个token,意味着你每生成一个token,你停留
69:15
has been subject to to produce appropriate tensors you can break it by finding out
在正确答案集合里的概率就会下降,而且是指数级下降。所以就像你
69:20
a prompt that will be outside of the the the set of promps has been
说的,这里有一个很强的假设:如果犯错概率不为零——而看起来确
69:26
trained on or things that are similar and then it will just P complete nonsense
实不为零——那就会产生一种漂移。对,而且这种漂移是指数级的,
69:31
do you when you say prompt do
错误会累积,对吧?所以...
69:33
you mean that exact prompt or do you mean a prompt that's like in many
在所有可能的prompt集合中,被用于训练的pr
69:38
parts very different than like is that easy to ask a question or to say
ompt比例其实微乎其微。它只是所有可能promp
69:43
a thing that hasn't been said before on the internet I mean people have come
t中极小极小的一部分。所以系统在它被pre-tr
69:48
up with uh things where like you you put a essentially a random sequence of
aining或fine-tuning过的promp
69:52
characters in The Prompt and that's enough to kind of throw the system uh into
t上表现正常,但还有一整片它不可能被训练过的空间
69:57
a mode where you know it it's going
,因为那个数量实在太庞大了。
70:00
to answer something completely different than it would have answered without this so that's a
为了回答一个完全不同的问题,比没有这个提示时它原本
70:05
way to jailbreak the system basically get it you know go outside of its uh
会回答的内容要离谱得多,所以这基本上就是一种越狱系统
70:11
of its conditioning right so that that's a very clear demonstration of it but of
的方法,让它跳出它的设定范围,对吧。所以这很明显是
70:17
course uh you know that's uh that goes outside of what is designed to do
个例子,但当然,这超出了它被设计来做的范围。如果你真
70:22
right if you actually stitch together reasonably grammatical sentences is that
的把语法还算通顺的句子拼凑起来,那……
70:26
the is it that easy to break it yeah some people have done things like
只是这个数字太庞大了。所以无论系统经
70:31
you you you write a sentence in English right that has and or you ask
过什么训练来生成合适的张量,你都可以
70:36
a question in English and it it produces a perfectly fine answer and then you
通过找到一个超出它训练过的提示集或相
70:40
just substitute a few words by the same word in another language and all of
似内容的提示来破坏它,然后它就会输出
70:45
a sudden the answer is complete nonsense yeah so so I guess what I'm saying
完全无意义的东西。你说“提示”的时候
70:50
is like which fraction
,是指……
70:51
of prompts that humans are likely to generate are going to break the system so
就这么容易破解它吗?是啊,有些人做过类似的事:你
70:56
the the problem is that there is a long tail yes uh this is a
用英文写一个句子,或者问一个英文问题,它给出一个完
71:00
an issue that a lot of people have realize you know in social networks and
全正常的回答,然后你只要把其中几个词换成另一种语
71:05
stuff like that which is uh there's a very very long taale of of things
言里的同一个词,突然之间回答就变成了一堆废话。对,
71:09
that people will ask and you can find tune the system for the 80% or
所以我想说的是,人类可能生成的prompt里,有
71:14
whatever of uh of the things that
多大比例会搞崩这个系统?
71:16
most people will will ask and then this long tail is is so large that
因此,无论系统经过怎样的训练
71:20
you're not going to be able to fun the system for all the conditions and
来生成合适的tensors,
71:24
in the end the system has a being kind of a giant lookup table right
你都可以通过找到一个超出它训
71:28
essentially which is not really what you want you want systems that can reason certainly
练过的prompt集合或相似
71:33
they can plan so the type of reasoning that takes place in llm is very
内容的prompt来打破它,
71:37
very primitive and the reason you can tell is primitive is because the amount of
然后它就会输出一堆胡言乱语。
71:41
computation that is spent per token produced is constant so if you ask a question
问题在于存在一个长尾。没错,这是很多人在社交网络之类的地方已经意识到的问题,就是人们会问的东西有一个
71:47
and that question has an answer in a given number of token the amount of
非常非常长的长尾。你可以针对80%或者随便多少大多数人会问的东西来fine-tune系统,但这个长尾
71:52
competition devoted to Computing that answer can be exactly estimated it's like you know it's
实在太大了,你不可能在所有情况下都去fine-tune系统。最终,系统本质上就像一张巨大的查找表,对吧
71:58
it's the the size of the prediction Network you know with its 36 layers on
?这其实不是你想要的。你想要的是能够推理的系统,当然还要能规划。所以LLM中发生的推理类型非常非常原
72:04
92 layers or whatever it is uh multiply by number
始,你能说它原始的原因是,每个token生成所花费的计算量是恒定的。
72:08
of tokens that's it and so essentially it doesn't matter if the question being asked
就是token的数量,仅此而已。所以本质上
72:14
is is simple to answer complicated to answer impossible to answer because it's undecidable or
,问题简单也好、复杂也好、甚至因为不可判定而
72:20
something um the amount of computation the system will be able to devote to to
无法回答也好,系统能用来回答的计算量是恒定
72:26
the answer is constant or is proportional to number of token produced in the answer
的,或者说跟答案生成的token数量成正比,
72:32
right this
对吧?
72:32
is not the way we work the way we reason is that when we're faced
所以如果你问一个问题,而那个问题的答案由一定数量的token组成,那么用于计
72:38
with a complex problem or a complex question we spend more time trying to solve
算那个答案的计算量可以被精确估算出来。就像,你知道,就是预测网络的大小,比如3
72:43
it and answer it right because it's more difficult there's a prediction element there's a
6层或者92层或者别的什么,乘以token的数量,就这么多。所以本质上,不管
72:49
iterative element where you're like uh adjusting your understanding of a thing by going over
问的问题是好回答、难回答、还是因为不可判定而无法回答,系统能用于回答的计算量是
72:54
over and over and over there's a hierarchical element so
恒定的,或者与答案生成的token数量成正比,对吧?
72:58
on does this mean that a fundamental flaw of llms or does it mean that
那么这是否意味着LLM存在根本性缺陷,还是说这个问题还有更深层次?你现在表现得就像个LLM,立刻回答“不”——其实这只是底层的世界模型,我们可以在它之上构建一些机制,比如你提到的持久长期记忆,来拥有这种能力。但这些系统的蓝图会与自回归LLM截然不同。所以,这就像心理学里人类系统一和系统二的区别。系统一是指那些你不需要刻意、有意识地思
73:05
there more part to that question now you're just behaving like an llm immediately answer
考就能完成的任务——你只是去做,因为做得足够多,可以下意识完成。比如经验丰富的司机可以边开车边聊天或听广播;或者资深棋手对阵新手时,不用深思就能凭模式识别下棋。这就是系统一,所有无需刻意规划的本能行为。而系统二则是需要规划的任务:如果你不是资深棋手,或者对阵同等水平的对手,你会权衡各种选项,花时间思考,而且思考时间越长表现越好,远胜于
73:12
no that that it's just the lowlevel world model on top of which we can
快棋限时的情况。这种刻意的规划依赖你内在的世界模型——这正是LLM目前做不到的。那么,我们如何让它们实现这一点?如何构建一个系统,能进行这种规划或推理,对复杂问题投入更多资源,对简单问题投入更少?这不会是自回归的token预测,而更像是潜在变量的推理,类似于曾经所谓的概率模型或图模型。基本原理是这样的:你把提示词看作观测变量,模型的作
73:19
then build some of these kinds of mechanisms like you said persistent long-term memory
用是衡量一个答案对提示词而言有多好。把它想象成一个巨大的神经网络,但只有一个输出——一个标量数值,比如答案对问题合适时输出0,不合适时输出很大的值。这个模型由LLM构建吗?实际上,你需要做的不是在文本字符串中搜索以最小化那个能量值,而是在抽象表征空间里操作——在抽象思维的空间里,通过这种最小化过程来展开一个想法。
73:26
or uh reasoning so on but we need that world model that comes from language
或者嗯,推理之类的,但我们需要的那个来自语言的世界模型——也许在构建好的世界模型之上搭建这种推理系统并不那么困难。好吧,不管难不难,不久的将来就会见分晓,因为很多人都在研究对话系统的推理和规划能力。我的意思是,即使我们只局限于语言,光是拥有在回答之前规划答案的能力,而且这种规划不一定与你用来生成答案的语言相关,对吧?所以这种心理模型的概
73:30
is it maybe it is not so difficult to build this kind of uh reasoning
念,让你能在说话之前先计划好要说什么,嗯,这非常重要。我认为未来几年会有很多系统具备这种能力,但这些系统的蓝图将与自回归语言模型截然不同。所以,嗯,这就像心理学中人类系统一和系统二的区别一样,对吧?系统一是指那些你不需要刻意、有意识地思考就能完成的任务,你直接去做就行了——因为你做得足够多,可以下意识地完成,不用多想。如果你是个有经验的
73:35
system on top of a well constructed World model OKAY whether it's difficult or not
司机,你可以边开车边跟人聊天或听广播,对吧?嗯,如果你是个非常熟练的棋手,你可以不假思索地跟新手对弈,你只是识别出模式然后走棋,嗯,对吧?那就是系统一。嗯,所有那些你本能地去做、不需要刻意规划和思考的事情。然后还有所有需要你规划的任务。所以如果你不是很有经验的棋手,或者你是有经验的棋手但跟另一个有经验的棋手对弈,你会考虑各种选项,对吧?
73:40
the near future will will say because a lot of people are working on reasoning
你会思考一会儿,对吧?而且如果你有时间思考,你会比下快棋时表现好得多——嗯,时间有限的情况下。所以,嗯,这种刻意的规划,利用你内在的世界模型,这种系统二——这正是当前语言模型做不到的。那么,我们如何让它们做到这一点呢?对吧?我们如何构建一个能够进行这种规划或推理的系统,把更多资源分配给复杂问题而不是简单问题?这不会是自回归的token预
73:45
and planning abilities for for dialog systems um I mean if we're even if we
测,而更像是潜在变量的推理,嗯,就像以前所谓的概率模型或图模型之类的东西。所以基本上原理是这样的:你知道,提示就像一组观测变量,嗯,而模型的作用是——它基本上可以衡量一个答案对某个提示来说有多好。好吧,把它想象成一个巨大的神经网络,但它只有一个输出,那个输出是一个标量值,比如说,如果答案对问题来说是好答案就是零,如果是坏答案就是很大的数。
73:50
restrict ourselves to language uh just having the ability to plan your answer before you
我们先把语言限制一下,就是那种在回答之前先规划好答案的能力,而且这种规划不一定和你最终用来表达答案的语言挂钩。所以,这种让你在说话之前先想好要说什么的心理模型,我觉得非常重要。未来几年会有很多系统具备这种能力,但这些系统的蓝图会和自回归语言模型截然不同。这就像心理学里说的系统一和系统二的区别。
73:55
answer uh in terms that are not necessarily linked with the language you're going to
系统一指的是那些你不用刻意、有意识地去想怎么做就能完成的任务,你直接就能做。因为你做得足够多,已经可以下意识地完成,不用动脑子。比如一个有经验的司机,他可以不用多想就开车,同时还能跟人聊天或者听广播。再比如一个非常熟练的棋手,跟一个新手对弈时也不用怎么思考,直接靠识别棋路就能下。这就是系统一。所
74:00
use to produce the answer right so this idea of this mental model that allows
有那些你本能地做、不需要刻意规划和思考的事情都属于这一类。而另一类任务则是需要你规划的。比如你是一个不太有经验的棋手,或者你虽然经验丰富但对手也很强,那你就会考虑各种走法,花时间思考。如果你有时间仔细想,表现会比下快棋、时间有限时好得多。所以,这种刻意的规划,需要用到你内心的世界模型,这就是目
74:06
you to plan what you're going to say before you say it um that is
前语言模型做不到的。那么,我们怎么让它们做到这一点呢?怎么构建一个系统,能够进行这种规划或推理,对复杂问题投入更多资源,而对简单问题投入更少?这不会是自回归式的token预测,而更像是推理潜在变量,类似于以前所谓的概率模型或图模型那种东西。基本思路是这样的:你把prompt看作观测到的变量,模型
74:11
very important I think there's going to be a lot of systems over the next
的作用是衡量一个答案对某个prompt来说有多好。可以把它想象成一个巨大的神经网络,但它只有一个输出,这个输出是一个标量值。比如,如果答案对问题来说是好的,这个值就是零;如果不好,就是一个很大的数。这个模型是由LLM构建的吗?实际上,你需要做的不是在可能的文本字符串中搜索来最小化那个“能量”,
74:16
few years that are going
而是在抽象的表示空间里操作。也就是说,在抽象思想的那个空间里,你通过这个最小化的过程来展开一个想法。
74:18
to have this capability but the blueprint of those systems would be extremely different from
拥有这种能力是可能的,但这些系统的蓝图会和自回归语言模型截然不同。所以,嗯,这就像心理学里说的系统一和系统二在人类身上的区别,对吧?系统一指的是那些你不用刻意、有意识地思考就能完成的任务,你直接就能做。因为你做得足够多,已经可以下意识地完成,不用动脑子。比如一个有经验的司机,可以一边开车一边跟人聊天或者听广播,对吧?嗯,如果你是个经验丰富的棋手,跟一个新手下棋,你也不用怎么思考,
74:23
autoregressive LMS so so um it's the same difference as the difference between what psychology
直接靠识别模式就能下,对吧?那就是系统一。嗯,所有那些你本能地做、不需要刻意计划和思考的事情。然后还有另一类任务,是需要你计划的。比如如果你不是特别有经验的棋手,或者你是有经验的棋手但跟另一个有经验的人下棋,你就会考虑各种选项,对吧?你会想一会儿,对吧?而且如果你有时间思考,表现会比下快棋好得多,因为快棋时间有限。所以,嗯,这种需要刻意规划、用到你内部世界模型的能力,就是系统二。
74:29
is called system one and system two in humans right so system one is the
而语言模型目前做不到这一点。那么,我们怎么让它们做到呢?对吧?怎么构建一个系统,能够进行这种规划或推理,把更多资源分配给复杂问题而不是简单问题?这不会是自回归式的 token 预测,而更像是某种潜在变量的推理,嗯,就是以前所谓的概率模型或图模型之类的东西。所以基本上原理是这样的:你知道 prompt 就像一组观测变量,嗯,模型能做的就是衡量一个答案对于这个 prompt 来说有多好
74:35
type of task that you can accomplish without like deliberately consciously think about how you
。所以把它想象成一个巨大的神经网络,但它只有一个输出,这个输出是一个标量数值,比如说如果答案对问题来说是好的,就输出零,如果不好就输出一个很大的值。嗯,由语言模型构建的模型?嗯,实际上你需要做的不是搜索可能的字符串来最小化那个能量,而是在抽象表示空间里做这件事。也就是说,在抽象思维的空间里,你会通过这个最小化过程来展开一个想法,嗯,找到合适的答案。但这个表示可能不是一个好答案,因
74:41
do them you just do them
为可能需要一些复杂的推理,对吧?所以,嗯,然后你会有另一个过程,它接收答案的表示并对其进行修改,以最小化一个成本函数,这个函数衡量答案在多大程度上是好的。
74:43
you've done them enough that you can just do it subconsciously right without thinking about
靠打个响指就行不通,对吧?嗯,
74:48
them if you're an experience driver you can drive without really thinking about it and
但可能还有更复杂的类似场景,N
74:53
you can talk to someone at the same time or listen to the radio right
LM 可能从未遇到过,也无法判断
74:57
um if you are a very experienced chess player you can play against a non-experienced
是否可行。所以,呃,从低层到高
75:02
chess player without really thinking either you just recognize the pattern and you play MH
层的那个联系,问题在于,语言表
75:07
right that's system one um so all the
达的高层是基于……
75:10
things that you do instinctively without really having to deliberately plan and think about it
你做得足够多之后,就能下意识地完成,
75:14
and then there is all tasks what you need to plan so if you are
根本不用多想。如果你是个有经验的司机,
75:19
not to experienced uh chess player or you are experienced with you play against another
你可以一边开车一边跟人聊天或者听广播
75:23
experienced chess player you think about all kinds of options right you you think about
,对吧?如果你是个非常熟练的棋手,跟新
75:28
it for a while right and you you you're much better if you have time
手对弈时也不用怎么动脑子,只要识别出
75:33
to think about it than you are if you if you play Blitz uh with
棋路然后落子就行,对吧?这就是系统一。
75:37
limited time so and um so this type of deliberate uh planning which uses your
这个数字实在太大了。所以无论系统经过什么样的训练来生
75:43
internal World model um that system to this is what LMS currently cannot do so
成合适的tensor,你都可以通过找到一个超出它训练
75:48
how how do we get them to do this right how do we build a
过的prompt集合——或者相似范围之外的promp
75:53
system that can do this kind of uh planning that or reasoning that devotes more
t——来破坏它,然后它就会直接输出一堆胡言乱语。你刚
75:59
resources to complex problems than to simple problems
才说prompt,是指...
76:02
and it's not going to be Auto regressive prediction of tokens it's going to be
所有那些你不用刻意计划、不用思考就能本能完成的事情。然后还有另一类
76:08
more something akin to inference of Laten variables in um you know what used to
任务,是需要你规划的。比如你是个不太熟练的棋手,或者你虽然熟练但对
76:15
be called problemistic models or graphical models and things of that type so basically the
手也很强,你就会考虑各种走法,对吧?你会花时间思考,对吧?如果有时
76:22
principle is like this you you know the prompt is like a
间思考,你的表现会比下快棋时好得多,因为快棋时间有限。
76:27
observed uh variables M and what your what the model does is that it's basically
观察到的变量M,模型做的事情基本上就是衡量
76:32
a measure of it can measure to what extent an answer is a good answer
一个回答对某个提示来说有多好。你可以把它想象
76:37
for a prompt okay so think of it as some gigantic neural net but it's
成一个巨大的神经网络,但它只有一个输出,这个
76:42
got only one output and that output is a scalar number which is let's say
输出是一个标量数值——比如回答对问题来说是
76:47
zero if the answer is a good answer for the question and a large
好的,这个值就是零,如果不好就是很大的值。
76:52
number if the answer is not a good answer for the question imagine you had
如果答案对问题来说不是好答案,想象一下你有这个模型——如果你有这样的模型,你可以用它来生成好答案。做法就是,嗯,生成提示,然后在所有可能的答案空间中搜索,找到那个能最小化这个数值的答案。这叫做基于能量的模型,但这个基于能量的模型需要由LLM构建的模型。所以,嗯,实际上你需要做的不是搜索能最小化那个能量的文本字符串,而是在抽象表示空间中做这件事。也就是说,在某种抽象
76:57
this model if you had such a model you could use it to produce good
思维的空间里,你会通过最小化模型输出的过程来阐述一个想法,对吧?这个输出只是一个标量,嗯,这是一个优化过程。所以现在系统产生答案的方式是通过优化,嗯,通过最小化一个目标函数,基本上就是这样。而且我们这里说的是推理,不是训练,对吧?系统已经训练好了。所以现在我们有了答案思维的抽象表示,答案的表示,我们把它输入到一个基本上自回归的解码器中,这个解码器可以非常简单,把它
77:02
answers the way you would do is you know produce the pumpt and then search
转换成表达这个想法的文本。所以,在我看来,这是未来对话系统的蓝图。嗯,它们会在把答案转换成文本之前,通过优化来思考答案、规划答案。这完全是图灵完备的。你能解释一下那个优化问题具体是什么吗?比如目标函数是什么?你刚才简单描述了一下,但优化的空间是什么?是那些抽象表示的空间。抽象表示。所以系统内部有一个抽象表示。你有一个提示,提示通过编码器生成一个表示,也许再通过一个
77:08
through the space of possible answers for one that minimizes that number um that's called
预测器预测出正确答案的表示。但这个表示可能不是好答案,因为可能需要一些复杂的推理,对吧?所以,嗯,然后有另一个过程,它获取答案的表示并修改它,以最小化一个成本函数,这个函数衡量答案对问题来说是不是好答案。现在我们暂时忽略如何训练这个系统来评估答案对问题是否合适的问题,但假设这样的系统可以创建。那么这个过程是什么样的?它有点像搜索过程,是一个优化过程。如果整个系统是
77:13
an energy based model but that energy based model would
可微的,那个标量输出就是通过某个神经网络运行的结果,嗯,把答案的表示输入到某个神经网络,然后通过梯度下降,通过反向传播梯度,你可以计算出如何修改答案的表示以最小化那个值。所以这仍然是基于梯度的,这是基于梯度的推理。所以现在你有了抽象中的答案表示。
77:17
need the the model constructed by the llm well so uh really what you need
打个响指就行不通,对吧?是的。但可能
77:23
to do would be to not uh search over possible strings of text that minimize
有更复杂的类似场景,一个NLM可能从
77:28
that uh energy but what you would do it do this in abstract representation space
未遇到过,也无法判断是否可行。所以从低
77:34
so in in sort of the space of abstract thoughts you would elaborate a thought
层到高层的那个连接,关键在于语言表达
77:40
right using this process of minimizing
的高层是基于……
77:42
the output of your your model okay which is just a scalar um it's an
你的模型输出只是一个标量,嗯,这其实是一个优化过程,对吧。现在系统产生答案的方式是通过
77:47
optimization process right so now the the way the system produces its answer is through
优化,也就是通过最小化一个目标函数,对吧。我们这里说的是推理,不是训练,系统已经训练好
77:52
optimization um by you know minimizing an objective function basically right uh and this is
了。所以我们有一个抽象的答案思维表示,答案的表示。我们把这个输入给一个自回归解码器,它可
77:57
we're talking about inference we're not talking about training right the system has been trained
以非常简单,把这个表示转换成表达这个思维的文字。所以,在我看来,这就是未来对话系统的蓝
78:02
already so now we have an abstract representation of the thought of the answer representation
图:它们会在把答案转换成文字之前,先通过优化来思考答案、规划答案。这其实就是TurCom
78:07
of the
plete。
78:08
answer we feed that to basically an auto regive decoder uh which can be very
我们把那个输入给一个自回归解码器,基本上就是一个很简
78:14
simple that turns this into a text that expresses this thought okay so that that
单的结构,把它转成表达这个想法的文本。所以在我看来,
78:19
in my opinion is the blueprint of future dialog systems um they will think about
这就是未来对话系统的蓝图——它们会先思考答案,通过优
78:25
their answer plan their answer by optimization before turning it into text uh and that
化来规划答案,然后再转成文本,这就是tur comp
78:31
is tur complete can you
lete能做到的。
78:33
explain exactly what the optimization problem there is like what's the objective function just link
你能具体解释一下那个优化问题到底是什么吗?比如目标函数是什么?你刚才简单提了一下,但优化的空间是什么?是表示的空间吗?那些抽象的表示?对,系统内部有一个抽象表示,你有一个prompt,prompt经过编码器产生一个表示,可能再经过一个预测器预测出正确答案的表示。但这
78:38
on it you you kind of briefly described it but over what space are you
个表示可能不是一个好答案,因为可能需要做一些复杂的推理,对吧。所以,然后有另一个过程,它拿到这个答案的表示,然后修改它,以最小化一个代价函数,这个函数衡量的是答案对问题来说有多好。我们暂时忽略如何训练这个系统去衡量答案是否合适这个问题,但假设这样的系统可以建立起来。那
78:44
optimizing the space of representations those abstract representation abstract repr so you have an abstract
这个过程是什么样的?是一种类似搜索的过程?它是一个优化过程。如果整个系统是可微的,那个标量输出就是通过某个神经网络跑出来的结果,嗯,把答案的表示输入到某个神经网络,然后通过梯度下降、反向传播梯度,就能算出如何修改答案的表示,以最小化那个目标。所以这仍然是基于梯度的,这
78:50
representation inside the system you have a prompt The Prompt goes through an encoder produces
是基于梯度的推理。所以你现在在抽象空间里有一个答案的表示,我不知道怎么浪漫化这个说法,比如概念空间,对,相对于具体的感官信息空间。但这个东西能做到推理吗?我们讨论的那种推理?其实不太行,只能以非常简单的方式。基本上你可以把那些东西看作是在做我刚才说的那种优化,只不过它
78:55
a representation perhaps goes through a predictor that predicts a representation
们是在一个离散空间里优化,也就是可能的token序列空间。而且它们做这个优化的方式极其低效,就是生成一大堆假设,然后选出最好的。这在计算上非常浪费,因为你基本上每生成一个序列都要跑一遍你的语言模型。
78:59
of the answer of the proper answer but that representation may not be a good
这个模型是由LLM构建的。实际上你需要做的
79:05
answer because there might there might be some complicated reasoning you need to do right
不是去搜索那些能最小化那个“能量”的文本字符
79:11
so um so then you have another process that takes the representation of the answers
串,而是在抽象的表示空间里操作。也就是说,在
79:17
and modifies it so as to minimize uh a cost function that measures to what
抽象思维的某种空间里,你通过这个最小化过程
79:22
extent the answer is a
来展开一个想法。
79:24
good answer for the question now we we sort of ignore the the fact for
但那个答案的表示可能并不是一个好的答案,
79:29
I mean the issue for a moment of how you train that system to measure
因为可能有一些复杂的推理需要做,对吧。所以
79:34
whether an answer is a good answer for for a question but suppos such a
呢,你会有另一个过程,它拿到答案的表示,然
79:40
system could be created but what's the process this kind of search like process it's
后去修改它,目的是最小化一个代价函数,这个
79:45
a optimization process you can do this if if the entire system is
函数衡量的是这个答案对问题来说有多好。
79:49
differentiable that scalar output is the result of you know running through some neural net
可微的意思是,那个标量输出是你通过某个神经网络跑出来的结果,嗯,把答案的表示跑进
79:54
mhm uh running the answer the representation of the answer to some neural net then
某个神经网络,然后通过梯度下降、反向传播梯度,你就能算出怎么修改答案的表示,让它最
79:58
by gradient descent by back propag back propagating gradients you can figure out like how
小化那个值。所以这仍然是基于梯度的,是基于梯度的推理。现在你有一个抽象的答案表示
80:03
to modify the representation of the answer so has to minimize that so that's still
,存在于可能的token序列空间里。他们用一种极其低效的方式做这个优化,就是生成一
80:08
a gradient based it's gradient based inference so now you have a representation of the
大堆假设,然后选出最好的,这在计算上非常浪费,因为你基本上得为每个生成的序列都跑
80:12
answer in abstract
一遍你的语言模型。
80:13
space now you can turn it into text right and the cool thing about this
现在你可以把它转成文本,对吧。而且酷的地方在于,这个表示可以通过梯度下降来优化,但它跟你最终用哪种语言表达答案无关,所以你是在一个抽象的表示空间里操作。这其实又回到了联合嵌入的概念,就是说,在一个——我不知道该怎么说,浪漫化一点的话——概念空间里工作,比在具体的感官信息空间里更好,对吧。但这个东西能做推理吗?就是我们一直在聊的那种推理。嗯,其实不太行,只能做非常简单的推理。基本上你可以把它看作
80:19
is that the representation now can be optimized through gr and descent but so is
是在做我之前说的那种优化,只不过是在一个离散的空间里优化,也就是所有可能的 token 序列构成的空间。而且它们做优化的方式极其低效,就是生成一大堆候选,然后选最好的。这在计算上非常浪费,因为你基本上得让语言模型跑一遍每个生成的序列,所以特别浪费。相比之下,在连续空间里做优化就好得多,因为你可以用梯度下降,而不是生成一大堆东西再挑最好的。你可以直接迭代式地优化你的答案,让它一步步逼近最优解,这
80:24
independent of the language in which you're going to express the answer right so you're
样效率高多了。但这种方法只能在连续空间里用可微函数实现。你刚才提到推理,比如深度思考的能力,或者深入推理的能力。那你怎么知道一个答案是好是坏,基于深度推理来判断呢?所以这就引出了一个问题:从概念上讲,你怎么训练一个基于能量的模型?基于能量的模型就是一个输出标量的函数,就是一个数字。你给它两个输入,X 和 Y,它告诉你 Y 是否和 X 兼容。X 是你观察到的东西,比如一张图片、一段视频之类的,Y
80:29
operating in the substract representation I mean this goes back to the Joint embedding right
是一个提议的答案,或者一段视频的后续,等等。它通过输出一个值来告诉你 Y 是否和 X 兼容:如果 Y 和 X 兼容,输出就是零;如果 Y 和 X 不兼容,输出就是一个非零的正数。那怎么训练这样一个系统呢?在完全通用的层面上,你给它展示兼容的 X 和 Y 对,比如一个问题和一个对应的答案,然后训练里面那个大神经网络的参数,让它输出零。但这并不能完全解决问题,因为系统可能会想:那我就对所有东西都
80:35
that is better to work in the uh in the
输出零好了。所以你必须有一个机制,确保对于错误的 Y,能量值会大于零。这里有两种选择:一种是对比方法。对比方法就是,你给系统展示一个 X 和一个错误的 Y,然后告诉系统,给这个输出一个高能量,把能量往上推,调整计算能量的神经网络的权重,让能量变大。这就是对比方法。
80:38
space of I don't know to romanticize the notion like space of Concepts versus yeah
现在我们暂时忽略一下怎么训练这个系统去判断一个答案对问题好不好这个问题,
80:43
the space of concrete sensory information right okay but this can can this do something
假设这样的系统是可以被创造出来的。那这个过程是什么样的呢?它有点像搜索的
80:48
like reasoning which is what we're talking about well not really in only in a
过程,是一个优化过程。如果整个系统是可微的,那个标量输出是经过某个神经网络
80:53
very simple way I mean basically you can think of those things that's doing the
跑出来的结果,嗯,把答案的表示输入到某个神经网络里,然后通过梯度下降、反
80:58
kind of optimization I was I was talking about except the optimize in a discrete
向传播梯度,你就能知道怎么修改答案的表示才能最小化那个代价。所以这仍然是
81:03
space which is
基于梯度的推理。
81:04
the space of possible sequences of of tokens and they do it they do this
所以,这种需要刻意规划、使用你内在世界模型的能力
81:10
optimization in a horribly inefficient way which is generate a lot of hypothesis and then
,就是系统二,而LLM目前还做不到。那么,我们怎
81:16
select the best ones and that's incredibly wasteful in terms of uh computation because you
么才能让它们做到呢?怎么才能构建一个系统,能够进行
81:22
have you run you basically have to run your LM for like every you know
这种规划或推理,对复杂问题投入更多资源,而对简单
81:27
Genera sequence um and
问题投入更少?
81:29
it's incredibly wasteful um so it's much better to do an optimization in continuous space
这真的太浪费了。所以更好的方式是在连续空间里做优化,用梯度下降法,而不是生成一大堆东西再选最好的。你只需要不断迭代优化你的答案,让它越来越接近最优解,这样效率高得多。但这种方法只能在连续空间里用可微函数实现。你刚才提到推理能力,也就是深入思考的能力。那你怎么知道一个答案是好是坏呢?基于深度推理来判断。所以这就引出了一个概念性问题:如
81:34
where you can do great and descent as opposed to like generate tons of things
何训练一个基于能量的模型?能量模型就是一个输出标量的函数,就是一个数字。你给它两个输入,X和Y,它告诉你Y是否与X兼容。X是你观察到的东西,比如一张图片、一段视频之类的,Y是提议的答案,比如视频的后续内容。它通过输出零来表示Y与X兼容,如果Y不兼容,输出就是一个非零的正数。那怎么在完全通用的层面上训练这样一个系统呢?你给它展示兼容的
81:39
and then select the best you just iteratively refine your answer to to go towards
X和Y对,也就是问题和对应的答案,然后训练内部大神经网络的参数,让它输出零。但这并不完全管用,因为系统可能会想:“那我就对所有东西都输出零。”所以你必须有一个过程,确保对于错误的Y,能量值会大于零。这里有两种选择:一种是对比方法。对比方法就是,你给系统展示一个X和一个错误的Y,告诉它:“给这个高能量,把能量推上去。”也就是调整计算能量
81:44
the best right that's much more efficient you can only do this in continuous spaces
的神经网络权重,让能量升高。这就是对比方法。这基本上就是我们现在做的,只不过我们只把它用在训练上,而不是推理上。还有另一类方法是非对比的,我更喜欢这些。非对比方法的基本思路是:能量函数需要在来自训练集的兼容X和Y对上保持低能量。那你怎么确保其他所有地方的能量都更高呢?方法就是加入一个正则化项,一个准则,在你的成本函数里加一项,基本上
81:49
with differentiable functions you're talking about the reasoning like ability to think deeply
就是最小化能取到低能量的空间体积。具体怎么做取决于架构,有各种不同的具体方法。但基本原理就是这样:如果你在XY空间的特定区域压低能量函数,它就会自动在其他地方升高,因为根据系统构造或正则化函数的限制,能取到低能量的空间体积是有限的。我们一直讲得很笼统,但到底什么是好的X和好的Y?什么是好的?
81:54
or to reason deeply how do you know what is an answer uh that's better
现在你有了一个在抽象空间里的答案表示,我不知道该怎么浪漫化这个概念,就像概念空间 vs 具体的感
82:00
or worse based on deep reasoning right so then we're asking the question of conceptually
官信息空间,对吧。但这个东西能做我们说的那种推理吗?其实不太行,只能做非常简单的推理。基本上你可以
82:06
how do you train an energy base model right so en based model is a
把那些东西看作是在做我刚才说的那种优化,只不过是在一个离散空间里优化——也就是所有可能的token
82:12
function with a scalar output just a number you give it two inputs X and
序列组成的空间。而且它们做优化的方式极其低效,就是生成一大堆假设,然后选出最好的那些。这在计算上非
82:18
Y and it tells you whether Y is compatible
常浪费,因为你基本上要针对每一个生成的序列都跑一遍语言模型。
82:21
with X or not X You observe let's say it's a pump image video whatever
对于X或非X,你观察到一个东西,比如一张
82:26
and why is a proposal for an answer a continuation of video um you know
图片或视频,而Y是答案的提议,是视频的延
82:31
whatever and it tells you whether Y is compatible with X and the way it
续,然后它会告诉你Y是否与X兼容。它告诉你
82:36
tells you that Y is compatible with X is that the output of that function
Y与X兼容的方式是,如果Y与X兼容,那个
82:41
would be zero if Y is compatible with X it would be a positive number
函数的输出会是零;如果Y不兼容,输出就是一
82:46
non zero if Y is
个非零的正数。
82:47
not compatible with X okay how do you train a system like this at a
不兼容X对吧,那你怎么在完全通用的层面上训练这样一个系统呢?就是给它展示兼容的X和Y对,一个问题和对应的答案,然后训练里面那个大神经网络的参数,让它输出零。好了,这并不完全管用,因为系统可能会想:我就对所有东西都输出零算了。所以你还得有个流程,确保对于错误的Y,能量会大于零。这里你有两个选择,一个是对比方法。对比方法就是,你展示一个X和一个错误的Y,然后告诉系统:
82:53
completely General level is you show it pairs of X and Y that are compatible
嗯,给这个一个高能量,把能量推上去,对吧?调整计算能量的神经网络里的权重,让能量升高。这就是对比方法。这基本上就是我们现在做的,只是我们没有把它用在推理上,只用在训练上。嗯,还有另一类方法,是非对比的,我更喜欢那些。那些非对比方法基本上是说:好了,能量函数需要在兼容的XY对上——也就是来自训练集的那些——有低能量。那你怎么确保能量在其他地方会更高呢?做法就是通过一
82:58
a question and the corresponding answer and you train the parameters of the big neural
个正则化器,一个准则,在你的成本函数里加一项,基本上最小化可以取低能量的空间体积。具体怎么做取决于架构,有各种不同的具体方式,但基本原理就是这样。这样,如果你在XY空间的特定区域把能量函数往下压,它就会自动在其他地方上升,因为能取低能量的空间体积是有限的——要么通过系统构造,要么通过正则化函数。我们一直讲得很笼统,但什么是好的X和好的Y?什么算好?这要看系统的内部
83:03
net inside U to produce zero okay now that doesn't completely work because the system
结构是怎么建的。如果系统的内部结构是这样设计的:里面有一个潜在变量,叫Z,你可以操纵它来最小化输出能量,那么那个Z就可以被视为一个好答案的表示,你可以把它翻译成一个Y,也就是好答案。所以这类系统可以用非常类似的方式训练,非常类似,但你必须要有这种防止坍缩的方法,确保你训练集之外的东西有高能量。目前这在LLM里是很隐性的,以一种人们没意识到的方式在做,但它确实存在。
83:08
might decide well I'm just going to say zero for
这是因为,当你给一个词高概率时,自动地你就给了其他词低概率,因为你只有有限的总概率可以分配,对吧?必须有一部分给到别处。所以当你最小化交叉熵之类的——当你训练你的LLM去预测下一个词时——你是在提高系统给正确词的概率,但同时也在降低其他词的概率。
83:12
everything so now you have to have a process to make sure that for a
关于正确回答的表示,但这个表示可
83:17
a wrong y the energy would be larger than zero and there you have two
能不是一个好答案,因为可能有一些
83:21
options one is contrastive Method so contrastive method is you show an X and a
复杂的推理需要做。所以你需要另一
83:26
bad Y and you tell the system well that's you know give a high energy
个过程,它接收回答的表示,然后修
83:31
to this like push up the energy right change the weights in the neural net
改它,以最小化一个成本函数,这个
83:35
that comput the energy so that it goes up um so that's contrasting methods the
函数衡量回答在多大程度上是好的。
83:40
problem with this is if the space of Y is large the number of such
这个问题的难点在于,如果Y的空间很大,你需要展示的对比样本数量会非常庞大。但人们确实这么做——他们在用RL训练系统时就是这么做的。基本上你训练的是一个叫reward model的东西,它本质上是一个目标函数,用来判断一个回答是好是坏。这其实就是我们已经在做的事情,只是我们没有把它用在inference上,而是只用在training里。嗯,还有另一类方法是非对比性的,我更喜欢那些。那些非对比性的方法基本上是
83:46
contrasty samples you're going to have to show is gigantic but people do this they
说:能量函数需要在来自训练集的、兼容的XY对上保持低能量。那你怎么确保能量在其他地方会更高呢?做法是加一个regularizer,一个准则,一个在你的cost function里的项,它基本上会最小化能容纳低能量的空间体积。具体怎么做取决于架构,有各种不同的具体方式。但基本原则是:如果你在XY空间的特定区域压低能量函数,它在其他地方就会自动升高,因为能容纳低能量的空间体积是有限的——这是由系统构造或regu
83:51
they do this when you train a system with RF basically what you're training is
larizing函数决定的。我们一直讲得很泛,但什么是好的X和好的Y?什么是X和Y的好表示?因为我们一直在讨论语言,如果你直接拿语言本身,那可能不太好,所以必须要有某种抽象的概念表示。嗯,其实你可以直接用语言来做,比如X是一段文本,Y是它的续写,或者X是问题,Y是答案。但你是说这样不行?不,这其实能实现LLM正在做的事情——这取决于系统内部结构是怎么建的。如果系统内部结构设计成里面有一个叫Z的latent
83:57
what's called a reward model which is basically an objective function that tells you whether
variable,你可以操控它来最小化输出能量,那么这个Z就可以被视为一个好的回答的表示,然后你可以把它翻译成Y,也就是一个好的回答。所以这类系统可以用非常相似的方式来训练,非常相似。但你必须要有办法防止collapse,确保你没有训练过的那些东西有高能量。目前这在LLM里是很隐性的,人们没意识到它正在发生,但它确实在发生。这是因为当你给一个词高概率时,自动地你就给了其他词低概率——因为你只有有限的总概率
84:03
an answer is good or bad and
可以分配,对吧?必须有一部分给这个词。所以当你用minimize cross entropy之类的方法训练你的LLM来预测下一个词时,你是在提高系统给正确词的概率,但同时也在降低其他词的概率。
84:05
that's basically exactly what what this is so we already do this to some extent
所以现在你得有一个过程,确保对于错误的Y,能量会大于
84:11
we're just not using it for inference we're just using it for training um uh
零。这里有两个选项,一个是对比方法。对比方法就是,你
84:17
there is another set of methods which are non contrastive and I prefer those uh
给系统展示一个X和一个错误的Y,然后告诉系统,给这个
84:22
and those non-contrastive method basically say uh okay the energy function needs to have low
高能量,把能量推上去,改变计算能量的神经网络的权重,
84:28
energy on pairs of xys that are
让它升高。这就是对比方法。
84:31
compatible that come from your training set how do you make sure that the energy
你做得足够多之后,就可以下意识地完成,不需要
84:36
is going to be higher everywhere else and the way you do this is by
刻意去想。如果你是个有经验的司机,你可以边开车
84:42
um having a regularizer a Criterion a term in your cost function that basically minimizes
边聊天或者听广播,对吧?如果你是个非常厉害的棋
84:47
the volume of space that can take low energy and the precise way to do
手,跟一个新手下棋也不需要怎么动脑,你只是识别
84:53
this is all kinds of different specific ways to do this depending on the architecture
出模式然后落子,对吧?那就是系统一。所以...
84:58
but that's the basic principle so that if you push down the energy function for
这基本上就是我们现在做的,只是我们没有把它用在推理上,只用在训练上。还有
85:03
particular regions in the XY space it will automatically go up in other places because
另一类方法是非对比的,我更喜欢那些。那些非对比方法基本上是说,能量函数需
85:08
there's only a limited volume of space that can take low energy okay by the
要在来自训练集的兼容XY对上具有低能量。你怎么确保能量在其他地方都会更高呢
85:13
construction of the system or by the regularizer regularizing function we've been talking very generally
?做法是,在你的成本函数中加入一个正则化器、一个准则、一个项,它基本上最
85:18
but what is a good X and a good Y what is a good
小化能取低能量的空间体积。具体怎么做取决于架构,有各种不同的具体方式。
85:22
representation of X and Y because we've been talking about language and if you just
X和Y的表示问题,因为我们一直在讨论语言,如果直接拿语言本身来用,那显然不行,所以必须要有某种抽象的意念表示。对,其实你可以直接用语言来做,比如X是一段文本,Y是这段文本的续写,或者X是一个问题,Y是答案。但你觉得这样行不通?我的意思是,这其实就是在做LLM现在做的事情。嗯,
85:27
take language directly that presumably is not good so there has to be some kind
这取决于系统内部结构怎么搭建。如果系统内部结构设计成存在一个叫Z的潜在变量,你可以操控它来最小化输出能量,那么Z就可以被视为一个好答案的表示,然后把它翻译成Y这个好答案。这种系统可以用非常类似的方式训练,但必须要有办法防止崩溃,确保那些你没训练过的内容有高能量。目前这在LLM里
85:32
of abstract representation of ideas yeah so you I mean you can do this with
是很隐性的,人们没意识到它在发生,但它确实存在。这是因为当你给某个词高概率时,自动就给其他词低概率了,因为概率总和是有限的,必须分给一个词。所以当你最小化交叉熵来训练LLM预测下一个词时,你是在提高系统给正确词的概率,但同时也在降低其他词的条件概率,这些概率是逐token累积的
85:37
language directly um by just you know X is a text and Y is a
。那对于视觉数据怎么做呢?我们一直在用JEPA架构来做,基本上就是联合嵌入。两个东西之间的兼容性是这样的:这里有一张图像或视频,这里是它的一个被破坏、偏移或变换过的版本,或者加了掩码。然后系统的能量就是表示预测的误差,也就是好内容的预测表示和实际表示之间的差异。你把被破坏的图像
85:42
continuation of that text yes um or X is a question why is the answer
输入系统,预测未被破坏的好输入的表示,然后计算预测误差,这就是系统的能量。所以这个系统会告诉你,如果这是一张好图像,而这是它的被破坏版本,那么能量会是零;如果两张图像完全不同,能量就会很高。希望整个过程能给你一个非常紧凑的现实表示,特别是视觉现实,而且我们知道它确实有效,因为我
85:47
but you're you're saying that's not going to take it I
们把这些表示用作分类系统的输入,效果非常好。好了,总结一下,你用一种只有Yann LeCun才敢用的辛辣方式建议我们放弃生成模型,转而采用联合嵌入架构?是的,放弃自回归生成,放弃——这感觉像在法庭上作证。
85:51
mean that's going to do what llms are doing well no it depends on how
在可能的token序列空间里,他
85:56
you how the internal structure of the system is built if the if the internal
们用一种极其低效的方式做优化——生
86:01
structure of the system is built in such a way that inside of this system
成大量假设,然后选出最好的。这在
86:07
there is a latent variable it's called Z that uh you can manipulate so as
计算上非常浪费,因为你基本上得为每
86:12
to minimize the output energy then that Z can be viewed as a
个生成的序列都跑一遍你的LM。
86:16
representation of a good answer that you can translate into a y that is a
一个优质答案的表示方式,你可以把它转化成另一个优质答案的y值。所以这类系统可以用非常
86:21
good answer so this kind of system could be train in a very similar way
相似的方式训练,非常相似,但你必须要有办法防止崩溃,确保那些你没训练过的内容仍然有高能
86:26
very similar way but you have to have this way of preventing collapse of of
量。目前这在LLM里是非常隐性的,人们甚至没意识到它正在发生,但它确实存在。之所以存在
86:30
ensuring that you know there is high energy for things you don't train it on
,是因为当你给某个词分配高概率时,其他词的概率自然就低了,因为你只有有限的总概率可以分
86:35
um and and currently it's it's very implicit in llm is done in a way
配,对吧?总和必须是一。所以当你最小化交叉熵或其他损失函数,训练你的LLM去预测下一个
86:40
that people don't realize it's being done but is it is
词时,你是在提高系统给正确词的概率,但同时也在降低其他词的概率。
86:43
being done is is due to the fact that when you give a high probability
之所以会这样,是因为当你给某个词分配高概率时
86:48
to a a word automatically you give low probability to other words because you only
,自然就会给其他词分配低概率,因为你总共只有那
86:53
have a finite amount of probability to go around right there have to some to
么多概率可以分配,对吧,必须得匀一些出去。所
86:58
one so when you minimize the cross entropy or whatever when you train the your
以当你最小化交叉熵之类的损失函数,训练你的LL
87:02
llm to produce the to predict the next word uh you're increasing the probability your
M去预测下一个词时,你是在提高系统给正确词的
87:07
system will give to the correct word but you're also decreasing
概率,但同时也在降低其他词的概率。
87:11
the probability will give to the incorrect words now indirectly that gives a low probability
概率会间接地给错误词汇低概率,给好的词汇序列高概率,给坏的词汇序列低概率,但这其实非常直接。嗯,而且完全不清楚为什么这真的能行得通,因为你不是在序列中所有符号的联合概率上操作,你只是把它分解成连续token的条件概率。那么对于视觉数据要怎么做呢?我们一直在用所有JEPA架构来做这件事,
87:16
to a high probability to sequences of words that are good and low probability to
基本上就是联合……嗯,两个事物之间的兼容性就是——这里有一张图像或一段视频,这里是这张图像或视频的损坏、偏移、变换或掩码版本,然后系统的能量就是表示预测误差,也就是好输入的预测表示与实际表示之间的误差。所以你把损坏的图像输入系统,预测未损坏的好输入的表示,然后计算预测误差,这就是系统的能
87:20
sequences of words that are bad but it's very direct mhm and it's not it's
量。这个系统会告诉你:如果这是一张好图像,而这是它的损坏版本,它会给出零能量;如果这两张图像完全不同,它会给出高能量。希望整个过程能给你一个非常紧凑的现实表示,尤其是视觉现实,而且我们知道它确实有效,因为之后我们把那些表示作为输入给分类系统,效果非常好。好的,那么总结一下,用一种只有Y
87:25
not obvious why this actually works at all but um because you're not doing it
ann LeCun才有的辛辣方式,你建议我们放弃生成模型,转而采用联合嵌入架构?是的,放弃自回归生成。是的,放弃概率模型,转而采用基于能量的模型,就像我们讨论过的。放弃对比方法,转而采用正则化方法。嗯,让我问问你,你一直对强化学习持批评态度,是的。那么最后一个建议是,我们放弃RL,转而采
87:30
on a joint probability of all the symbols in a in a sequence you're just
用模型预测控制,就像你之前说的,只有当规划没有产生预期结果时才使用RL,在这种情况下用RL来调整世界模型或评判器?是的。那么你提到了RLHF,为什么你仍然讨厌强化学习?我并不讨厌强化学习,我认为它不应该被完全放弃,但它的使用应该被最小化,因为它在样本效率上极其低下。所以训练系统的正确方法
87:35
doing it kind of uh you sort of factorize that
是:首先让它主要通过观察(可能加上少量交互)学习好的世界表示和世界模型,然后基于这些进行引导。如果表示足够好,那么调整应该是最小的。是的,这里有两件事你可以利用:如果你已经学了一个世界模型……
87:38
probability in terms of conditional probabilities over successive tokens so how do you do this
用连续token上的条件概率来表达概率。那么对于视觉数据要怎么做呢?我们一直在用所有JEPA架构来做这件事,就是联合嵌入。呃,两个东西之间的兼容性,你看,这里是一张图片或一段视频,这里是它的一个被破坏、偏移或变换后的版本,或者被遮罩了。然后系统的能量就是表征的预测误差,也就是对“好样本”的预测表征和实际表征之间的误差。所以你拿被破坏的图像输入系统,预测出未
87:44
for visual data so we've been doing this with all jepa architectures basically the joint
破坏的干净输入的表征,然后计算预测误差,这就是系统的能量。这个系统会告诉你,如果这是一张好图像,而这是它的被破坏版本,它会给出零能量;如果两张图像完全不同,它会给出高能量。希望整个过程能给你一个非常紧凑的对现实、对视觉现实的压缩表征,而且我们知道它确实能做到,因为我们后来把这些表征作为输入喂给分类系统,效果非常好。好了,那么总结一下,用一种只有Yann L
87:50
I so uh there are the compatibility between two things is uh you know here's
eCun才有的辛辣方式,你建议我们放弃生成模型,转而采用联合嵌入架构?是的,放弃自回归生成。是的,放弃……这感觉像在法庭作证。不会产生预期的结果,在这种情况下我们用RL来调整世界模型或评判器。对。那么你提到了RLHF,基于人类反馈的强化学习,为什么你还是讨厌强化学习?我不讨厌强化学习,我觉得它不应该被完全放弃,但它的使用应该被最小化,因为它在样本效率上极其
87:56
here's an image or a video here's a corrupted shifted or transformed version of that
低下。所以训练系统的正确方式是先让它主要通过观察来学习好的世界表征和世界模型,可能辅以少量交互,然后在此基础上进行引导。如果表征足够好,那么调整应该很小。没错。如果你已经学了一个世界模型,那么有两种可能出错的情况:要么你的目标函数没有反映你真正想优化的目标,要么你的世界模型不准确,也就是你对世界将要发生什么的预测不准确。所以如果你想在运行过程中调整你的世界
88:02
image or video or masked okay and then
模型或目标函数,这基本上就属于RL的范畴,RL在一定程度上就是处理这个的。调整你的世界模型,甚至提前调整它的方式,就是去探索那些你知道世界模型不准确的空间区域,这基本上就是好奇心,或者说玩耍。
88:05
uh the energy of the system is the prediction error of the representation uh the
系统的能量是表示上的预测误差,也就是“好事物”的预测表示和实际表示之间的差异。所以你把损坏的图像输入系统,预
88:11
the predicted representation of the Good Thing versus the actual representation of the good thing
测未损坏的优质输入的表示,然后计算预测误差,这就是系统的能量。所以这个系统会告诉你,如果这是一张好图像,而这是
88:17
right so so you run the corrupted image to the system predict the representation of
它的损坏版本,它会给出零能量;如果这两张图像完全不同,它会给出高能量。希望整个过程能给你一个非常漂亮的、压缩
88:23
the the good input uncorrupted and then compute the prediction error that's the energy of
的现实表示,尤其是视觉现实,而且我们知道它确实有效,因为我们会把这些表示作为输入,用到分类系统里,效果非常好。
88:30
the system so this system will tell you this is a good you know if
这个系统会告诉你,这是一张好的图像,而这
88:34
this is a good image and this is a corrupted version it will give you
是一个被破坏的版本。如果这两张图本质上是
88:39
Zero Energy if those two things are effectively one of them is a corrupted version
一张是另一张的破坏版本,它会给你零能量;
88:44
of the other give you a high energy if if the two images are completely
如果两张图完全不同,它会给你高能量。希望
88:49
different and hopefully that whole process gives you a really nice compressed representation of of
整个过程能给你一个非常紧凑的视觉现实表征
88:53
reality of visual reality and we know it does
,而且我们知道它确实有效。
88:56
because then we use those for our presentations as input to a classification system something
好的,那么总结一下,用一种只有Yann LeCun能用的辛辣方式,你建议我们放弃生成模型,转而采用联合嵌入架构?是的,放弃自回归生成?是的
89:02
system works really nicely okay well so to summarize you recommend in a in a
,放弃……这感觉像在法庭上作证。不会产生预期的预测结果,在这种情况下我们用RL来调整世界模型或评论家。是的,所以你提到了RLHF,强化学习
89:09
in a spicy way that only Yan laon can you recommend that we abandon generative
与人类反馈,为什么你还是讨厌强化学习?我并不讨厌强化学习,我认为它不应该被完全放弃,但它的使用应该被最小化,因为它在样本效率上极其低下。所以
89:15
models in favor of joint embedding architectures yes abandon autor regressive generation yes abandon Pro
训练系统的正确方式是,先让它主要从观察中学习好的世界表示和世界模型,可能再加一点交互,然后基于这些进行引导。如果表示足够好,那么调整应该是
89:21
this feels like a court testimony uh
最小化的。是的,这里有两件事你可以用,如果你已经学了一个世界模型。
89:24
abandon probabilistic models in favor of energy based models as we talked about abandon contrastive
放弃概率模型,转向我们之前聊过的基于能量的模型;放弃对比方法,转向正则化方法。呃,我想问你一个问题——你一直对强化学习持批评态度,对吧?那么,最后一条建议是,放弃RL,改用我们讨论过的模型预测控制,只有在规划无法产生预期结果时才使用RL,在这种情况下,我们用RL来调整世界模型或评判器。对。所以,你提到了RLHF——基于人类反馈的强化学习——你为什么还是讨厌强化学
89:29
methods in favor of regularized methods and uh let me ask you about this you've
习?我并不讨厌强化学习,我觉得它不应该被完全放弃,但它的使用应该被最小化,因为它在样本效率上极其低下。所以,训练系统的正确方式是:首先让它从观察中学习好的世界表征和世界模型,可能再加一点交互,然后基于这些进行引导。如果表征足够好,那么调整应该是微乎其微的。对。现在,如果你已经学了一个世界模型,有两种可能出错的地方:要么你的目标函数没有反映你真正想优化的目标,要么你
89:35
been for a while a Critic of reinforcement learning yes so what uh the last
的世界模型不准确——也就是说,你对世界将要发生什么的预测是错误的。所以,如果你想在运行过程中调整你的世界模型或目标函数,这基本上就属于RL的范畴,RL在一定程度上就是处理这个的。对。调整世界模型的方法,甚至提前调整的方法,就是探索那些你知道世界模型不准确的空间区域——这基本上就是好奇心,或者说玩耍。当你玩耍时,你会探索空间的一部分,这些部分你不想在现实中尝试,因为
89:41
recommendation is that we abandon RL in favor of Mo model predictive control as you
可能很危险,但你可以在不把自己搞死的情况下调整世界模型。所以,这就是你该用RL的地方。当需要学习某个特定任务时,你已经有了所有好的表征和世界模型,但你需要针对当前情况做调整,这时候就用RL。那你为什么觉得RLHF效果这么好?就是基于人类反馈的强化学习,它为什么对大型语言模型产生了如此变革性的影响?产生变革性影响的是人类反馈。有很多方式使用它,其中一些其实纯粹是监督
89:47
were talking about and only use RL when planning
学习,并不是真正的强化学习。所以,关键是HF,HF部分。对。然后还有使用人类反馈的方式:你可以让人类对世界模型生成的多个答案进行评分,然后你训练一个目标函数来预测这个评分,接着你就可以用这个目标函数来判断一个答案好不好。
89:50
doesn't yield the pr predicted outcome and uh we use RL in that case to
并没有产生预期的结果,呃,这种情况下我们会用RL
89:56
adjust the world model or the critic yes so uh you mentioned uh rlf reinforcement
来调整世界模型或评判器。对,所以你提到了RLHF
90:01
learning with human feedback uh why do you still hate uh reinforcement learning I don't
,基于人类反馈的强化学习,那你为什么还是讨厌强化
90:07
hate reinforcement learning and I think all of I think it should not be uh
学习呢?我并不讨厌强化学习,我觉得它不应该被完全抛
90:12
abandoned completely but I think its
弃,但我认为它的……
90:15
use should be minimized because it's incredibly inefficient in terms of samples and so the
所以有两种可能出错的方式:要么你的目
90:20
the proper way to train a system is to First just have it learn uh
标函数没有反映你真正想优化的目标函数
90:25
good representations of the world and World models from Mostly observation maybe a little bit
,要么你的世界模型不准确。也就是说,
90:30
of interactions and then steered based on that if the representation is good then the
你对世界将要发生什么的预测是不准确的
90:35
adjustment should be minimal yeah now there's two things you can use if you've learned
。所以如果你想在运行过程中调整你的世
90:40
a world model
界模型,
90:41
you can use the world model to plan a sequence of actions to arrive at
你可以用世界模型来规划一系列动作,从而达成某个特定目标。除非你衡量成功的方式本身就不精确——比如你觉得自己会不会从自行车上摔下来的判断可能是错的,或者你在MMA里跟人打的时候,对方会做什么、会不会突然变招,这些都不好说。嗯,所以呢,有两种可能出错的方式:要么你的目标函数没有反映出你真正想优化的那个目标函数,要么你的世界模型不准确,对吧?也就是说,你对世界接下来会发生什么的预测是错的。所以,如果你想
90:47
a particular objective you don't need a unless the way you measure whether you succeed
在运行世界的过程中调整你的世界模型或者目标函数,这基本上就属于RL的范畴了——某种程度上RL就是干这个的,对吧?调整你的世界模型,而提前调整世界模型的方法,就是去探索那些你知道自己世界模型不准确的空间区域,这基本上就叫好奇心,或者说是玩,对吧?当你玩的时候,你会探索状态空间里那些你不想在现实中去做的地方,因为可能很危险,但你可以调整你的世界模型,而不用把自己搞死,差不多就是这个意思。所以,当需要学习
90:54
might be inexact your idea of you know whether you're going to fall from your
某个特定任务的时候,你已经有了所有好的表征,已经有了你的世界模型,但你还需要针对当前的情况去调整它,这时候就用RL。为什么你觉得RLHF效果这么好?这个基于人类反馈的强化学习,为什么它对大语言模型产生了如此变革性的影响?其实产生变革性影响的是人类反馈本身。使用人类反馈有很多方式,其中一些纯粹是监督学习,实际上并不是真正的强化学习。所以是HF,是HF这部分,嗯。然后呢,还有别的方式使用人类反馈,对吧?
91:00
bike might be wrong or whether the person you're fighting with MMA is going to
你可以让人类去给世界模型生成的多个答案打分,然后你训练一个目标函数来预测那个分数,接着你就可以用这个目标函数来预测某个答案好不好,再通过它反向传播梯度去找到新的系统,让系统只生成高分的答案。好了,这是一种方式。这在RL里就相当于训练一个所谓的奖励模型,对吧?就是一个小模型,用来估计一个答案有多好。这跟我之前讲规划时提到的目标函数非常相似,只不过现在它不是用来做规划,而是用来微调你的系统。我觉得如果把
91:06
do something and do something else um
它用在规划上会更高效,但嗯,目前它是用来微调系统的参数。现在呢,做这件事有好几种方法,有些是监督式的,就是直接问一个人,比如“这个问题的好答案是什么?”然后你就把答案打出来。嗯,方法很多。
91:09
so there uh so there's two ways you can be wrong either your your objective
没有产生预期的结果,我们就在那种情况下用RL来调整世界模型或critic。对
91:14
function does not reflect the actual objective function you want to optimize or your world
,你提到了RLHF,也就是reinforcement learning wi
91:20
model is inaccurate right so you didn't you the prediction you were making about what
th human feedback,为什么你还是讨厌reinforcement
91:25
was going to happen in the world is inaccurate so if you want to adjust
learning?我并不讨厌reinforcement learning,
91:31
your world model while you are operating the
我觉得它不应该被完全抛弃,但我认为它的……
91:34
world or your objective function that is basically in the realm of RL this is
你的世界模型或者说目标函数,本质上属于RL的范畴,这就是RL处理的问题,对吧,某种程度上。所以调整你的世界模型,甚至提前调整它的方法,就是去探索那些你知道世界模型不准确的空间区域,这基本上就叫
91:39
what RL deal deals with uh to some extent right so adjust your world model
好奇心,或者说是玩耍。当你玩耍的时候,就需要根据当前情况去调整它,这时候就要用RL。你觉得为什么RLHF效果这么好?就是基于人类反馈的强化学习,为什么它对大语言模型产生了如此变革性的影响?之前
91:44
and the way to adjust your world model even in advance uh is to explore
产生变革性影响的是人类反馈本身。使用人类反馈有很多方式,有些其实纯粹是监督学习,并不是真正的强化学习。所以关键是HF,就是HF部分。然后还有各种使用人类反馈的方法。你可以让人类对世界模型生成的
91:49
parts of the space where your world model where you know that your world model
多个答案进行评分,然后你训练一个目标函数来预测这个评分,接着用这个目标函数来判断一个答案好不好,再用来fine-tuning你的系统。我觉得用它来做planning会更高效,但目前它是用来fin
91:54
is inaccurate that's called curiosity basically or play right when you play
e-tune系统的参数。现在做这件事有好几种方式,有些是监督式的,就是直接问一个人,比如“这个问题的好答案是什么?”,然后你就直接把答案写进去。方法很多。
91:58
you kind of explore part of the St space that um you know you don't
你会在ST空间里探索一些你不想在现实中做的事情,因为那可能很危险,但你可以调整你的world model,基本上不会把自己搞死。所以这就是你想用RL的地方——当需要学习某个具体任务时,你已经有了所有好的表征,已经有了你的world model,但你需要针对当前的情况做调整,这时候就用RL。你觉得为什么RLHF效果这么好?这种reinforcement
92:03
want to do in for real because it might be dangerous but but you can
learning with human feedback,为什么它对大型语言模型产生了如此变革性的影响?其实真正产生变革性影响的是human feedback。有很多种使用方式,有些纯粹是supervised,实际上根本不是reinforcement learning。所以关键是HF,就是HF这部分。然后还有别的方式使用human feedback,比
92:08
adjust your world model uh without killing yourself basically um so that's what you want
如你可以让人对world model生成的多个答案进行评分,然后训练一个objective function来预测这个评分,再用这个objective function来判断一个答案好不好,然后fine-tuning你的系统。我觉得用它来做planning会更高效,但目前它是用来fine-tuning系统的参数。现在有好几种方法来做这个,有些是supe
92:13
to use ourl for when when when it comes time to learning a particular task
rvised的,就是直接问一个人,比如“这个问题的好答案是什么?”,然后你就把答案打出来。还有很多其他方式。历史上有很多图片,当然这些图片在中国受到政府的高度审查,所以大家就开始问:这些LLM的设计过程是怎样的?审查在这些模型里扮演什么角色?等等。所以你在Twitter上评论说,开放训练数据的分布,这些数据反映了社会中的偏见,可能对某些人来说有冒犯性,
92:18
you already have all the good representations you already have your world model but you
或者没有,而一些去偏见的技巧又可能对某些人造成冒犯,因为历史不正确性之类的原因。所以你可以问两个问题:第一个问题是,能不能造出一个没有偏见的AI系统?答案绝对是“不能”,而且这不只是因为技术上的挑战——虽然技术挑战确实存在——而是因为偏见在观察者眼里。不同的人对什么是偏见可能有不同的看法。很多事情,有些事实是无可争议的,但也有很多观点或表达方式可以不同
92:23
want you
,所以你不可能有一个无偏见的系统,那根本不可能。
92:24
need to adjust it for the situation at hand that's when you use RL why
它并没有产生预期的预测结果,在这种情况下我们就用RL来调整世界模型或critic。对,所以你提到了RL
92:29
do you think rhf works so well this reinforcement learning with human feedback why did
HF,基于人类反馈的强化学习,那你为什么还是讨厌强化学习呢?我并不讨厌强化学习,我认为它不应该被完全抛弃
92:35
it have such a transformational effect on large language models that before what's had the
,但我认为它需要根据具体情况进行调整,这时候才用RL。那你觉得RLHF为什么效果这么好?这种基于人类反馈
92:41
transformational effect is human feedback there's many ways to use it and some of it
的强化学习,为什么对大型语言模型产生了如此变革性的影响?其实产生变革性影响的是人类反馈本身。有很多方式
92:46
is just purely supervised actually it's not really reinforcement
可以利用它,其中一些纯粹是监督学习,实际上并不是真正的强化学习。
92:50
rning so it's the the HF it's the HF yeah uh and then there is
历史上那些图像,当然这些图像被中国政府高度审查,于是大家就开始问:设计这些LLM的过程是怎样的?审查在这些模型里扮演什么角色?等
92:55
ways to use human feedback right so you can you can ask humans to rate
等这类问题。所以你在Twitter上评论说,开放训练数据的分布,这些数据反映了社会中的偏见,可能对某些人具有冒犯性,也可能没有。而
92:59
answers multiple answers that are produced by World model and uh and and then what
一些去偏见的技巧,又可能因为历史上的不正确性之类的原因,冒犯到另一些人。所以你可以问两个问题:第一个问题是,有没有可能制造出一个没
93:04
you do is you train an objective function to predict that rating and then you
有偏见的AI系统?答案是绝对不可能。这不仅仅是因为技术上的挑战,虽然技术挑战确实存在,而是因为偏见是因人而异的。不同的人对什么构成
93:09
can use that objective function to predict you know whether an answer is good and
偏见可能有不同的看法。很多事情,有些事实是无可争议的,但也有很多观点,或者可以用不同方式表达的东西。所以你不可能有一个没有偏见的系
93:14
you can
统,这根本不可能。
93:15
back propagate gradient through this to find you new system so that it only produces
通过这个反向传播梯度来找到新的系统,让它只生成高评分的答案。好,这是一种方法,这在我们领域里
93:19
High highly rated answers okay so that's one way so that's like in ourl that
意味着训练一个所谓的reward model,对吧?就是一个小的模型,用来估计一个答案有多好。
93:24
means uh training what's called a reward model right so something that you know basically
这跟我之前讲planning时的目标非常相似,只不过现在它不是用来做planning,而是用
93:29
a small on that that estimates to what extent an answer is good right it's
来做fine-tuning你的系统。我觉得用它来做planning会更高效,但呃,目前它是用来
93:34
very similar to The Objective I was I was talking about talking about earlier for
fine-tune系统的参数。现在有好几种方法来做这个,有些是supervised的,就是直
93:39
planning except now it's not used for planning it's it's used for
接找个人问,比如这个问题的好答案是什么?然后你就把答案打进去。呃,方法很多。
93:43
fine-tuning your system I think it would be much more efficient to use it for
所以有两种方式你会出错:要么你的目标函数
93:48
planning but um but but uh currently it's used to fine tune the parameters the
没有反映你真正想优化的目标函数,要么你的世
93:53
system now there there's several ways to do this um you know some of some
界模型不准确。也就是说,你对世界会发生什么
93:57
of them are supervised you just you know ask a human person like what is
的预测是不准确的。所以如果你想在运行过程中
94:02
a good answer for this right then you just type the answer um uh I
调整你的世界模型,让它适应当前的情况,那就
94:07
mean there's there's lots of ways
是你用RL的时候。
94:09
that those systems are are being adjusted now a lot of people have been very
这些系统现在正在被调整。很多人对最近发布的Google Gemini 1.5非常不满,用我的话来说,它可以说是超级“觉醒”,而且是贬义的那种。它做了一些几乎荒谬可笑的事情,比如篡改历史,生成黑人乔治·华盛顿的图像,或者更严重一点,你在Twitter上评论过的那件事——它拒绝评论或生成关于天安门广场或“坦克人”的图像,甚至文字描述都不行,那可是
94:16
critical of the recently released Google's Gemini 1.5 for essentially in my words I could
历史上最具标志性的抗议画面之一。当然,这些图像被中国政府高度审查,所以大家开始质疑:设计这些LLM的过程是什么?审查在这些东西里扮演什么角色?等等。然后你在Twitter上评论说,开源才是答案。对,差不多就是这样。你能解释一下吗?我其实在我能用的每个社交网络上都发了那条评论,而且我在多个场合反复强调过这个观点。我的看法是这样的:人们可以抱怨AI
94:23
say super woke woke in the negative connotation of that word uh there is some
系统有偏见,它们通常确实有偏见,因为训练数据的分布反映了社会中的偏见,这可能会冒犯一些人,也可能不会。而一些去偏见的技巧又可能冒犯另一些人,因为历史不准确之类的问题。所以你可以问两个问题:第一个问题是,能不能造出一个没有偏见的AI系统?答案是绝对不可能,这不只是因为技术挑战——虽然确实有技术挑战——而是因为偏见是因人而异的。不同的人对什么是偏
94:30
almost hilariously absurd things that it does like it modifies history uh
见可能有不同的看法。很多事情,有些事实是无可争议的,但也有很多观点或表达方式可以不同。所以你不可能有一个没有偏见的系统,那根本不可能。那么答案是什么?答案和自由民主制度中关于媒体的答案一样:媒体要自由且多元化。我们之所以有言论自由,是有充分理由的,因为我们不希望所有声音都统一。
94:36
like generating images of a u black George Washington or um perhaps more seriously something
比如生成一张黑人乔治·华盛顿的图片,或者更严肃一点,你在Twitter上评论过的那种情况——拒绝评论或生成关
94:44
that you commented on Twitter which is refusing to comment on or generate images of
于天安门广场或Tank Man的图片,这是历史上最具传奇色彩的抗议图片之一。当然这些图片被中国政府高度审查,所
94:51
U or even descriptions of tianan square or the the Tank Man one of the
以大家就开始问,设计这些LLM的过程是怎样的?审查在这些东西里扮演什么角色?等等。然后你在Twitter上评
94:59
most sort of legendary protest
论说,open source就是答案。
95:02
images in history and of course these images are highly censored by the Chinese government
历史上那些图片,当然这些图片在中国政府那里是被高度审查的,所以大家就开始问,设计这些LLM的过程到底是什么样的,审查在其中扮演什么角色,诸如此类的问题。于是你在Twitter上评论说,Open跟法国政府谈了很多次,法国政府不会接受他们所有公民的数字生活被美国西海岸的三家公司控制,这完全不可接受,不管这些公司初衷多好,这对民主都是威胁。而且这也是一个广泛的社区,有成千上万的企业在用这个技术构建应用。所以Me
95:08
and therefore everybody start asking questions of what is the process of uh designing these
ta从这个技术中获取收入的能力,并不会因为开源基础模型的发布而受损。Gemini现在受到的根本批评是,效果不好之类的问题,而且那些糟糕的文本、图片、视频的公开传播,实际上又把那些例子放进了下一版本的训练数据里。所以这正好凸显了这件事有多难,因为各种人都不满意,就像你说的,你不可能做出一个让所有人都满意的系统。是的,所以如果你打算自己做fine-tune并且保持闭源,那问题就在于……在那之前,30年前我们还在
95:15
llms what is what is what is the role of censorship in these all that
搞35、搞com Nets、搞早期神经网络的时候,我就超级兴奋,因为我看到了一条通往人类级别智能的道路,系统可以理解世界、记忆、规划、推理。我管它叫“圆珠笔”,然后Twitter上就炸了,说“天哪,人们可以用它写可怕的东西,比如 misinformation、propaganda、hate speech,赶紧禁掉它”,然后那些写末日论的人就来了,就像AI末日论者一样,想象如果每个人都能拿到一支圆珠笔,那社会
95:21
kind of stuff so you uh commented on Twitter saying that open
就毁了,应该立法禁止用圆珠笔写仇恨言论,现在就要监管圆珠笔。还有一点,如果我们真的想要观点的多样性,在这个未来里我们都会通过AI系统互动,那这些系统本身就需要是多样化的,以保护思想、信仰、政治观点等等的多样性,以及保护哲学、理性主义、逃离宗教教条、民主、科学。当然,没有这些就不会有美国革命、法国革命,我们现在可能还在被统治着。
95:26
source is the answer yeah essentially so um can you explain I I actually made
对,基本上是这样。你能解释一下吗?我其实在我能用的每个社交网络上都发了那条评论,而且我在各种论坛里多次提过这个观点。呃,我的看法是这样的:人们可以抱
95:32
that comment on just about every Social Network I can and I've I I've uh
怨AI系统有偏见,它们通常确实有偏见,因为训练数据的分布反映了社会中的偏见,这可能会冒犯一些人,也可能不会。而一些去偏见的技巧又可能冒犯另一些人,因
95:38
I've made that point multiple times in in various forums um uh here's my my
为历史不正确性之类的原因。所以你可以问两个问题:第一个问题是,能不能造出一个没有偏见的AI系统?答案是绝对不可能,这不仅仅是因为技术挑战——虽然技术
95:44
point of view on this uh people can complain that AI systems are biased and
挑战确实存在——而是因为偏见在观察者眼里。不同的人对什么是偏见可能有不同的看法。很多事吧,有些事实是无可争议的,但也有很多观点或者可以用不同方式表达
95:50
they generally are biased by
的东西。所以你不可能有一个无偏见的系统,这根本就是
95:52
the distribution of the training data that they've been trained on um that reflects biases
为什么你觉得RLHF效果这么好?就是reinforcement learning with human feedba
96:02
in society um and that is potentially offensive to some people or potentially not and
ck,为什么它对大语言模型产生了如此变革性的影响?产生变革性影响的是human feedback。有很多方式可以用它,
96:12
and some techniques to debias then become offensive to some
有些其实纯粹是supervised的,并不是真正的reinforcement。
96:18
people um because of you know historical uh incorrectness and things like that um and
历史上的图像,当然这些图像被中国政府高度审查,所以大家开始问问题:设计这些LLM的过程是
96:24
so you can ask the question you can ask two questions the first question is
什么样的?审查在这些过程中扮演什么角色?诸如此类。所以你在Twitter上评论说,Ope
96:30
is it possible to produce an AI system that is not biased and the answer
nAI的人,因为你知道,历史上的不准确之类的原因,所以你可以问两个问题:第一个问题是,能
96:36
is absolutely not and it's not because of technological uh challenges although there are technological
不能造出一个没有偏见的AI系统?答案是绝对不可能,这不仅仅是因为技术上的挑战,虽然技术挑
96:41
challenges to
战确实存在。
96:42
that it's because bias is in the eye of the beholder um different people may
因为偏见是因人而异的,不同的人对什么是偏
96:48
have different ideas about what constitutes bias um you know for a lot of uh
见可能有不同的看法。你知道,很多事情,有些
96:54
a lot of things I mean there are facts that are you know indisputable but
事实是无可争议的,但也有很多观点,或者可
96:59
there are a lot of opinions or or things that can be expressed in different
以用不同方式表达的东西。所以,你不可能有一
97:05
ways U and so you cannot have an unbiased system that's just an
个完全没有偏见的系统,这根本做不到。
97:10
impossibility um and so what's the what's the answer to this and the the answer
不可能性,嗯,那这个问题的答案是什么?答案和我们之
97:17
is the same answer that we found in Liberal democracy about the press the Press
前在自由民主制度中关于媒体的结论一样——媒体必须自由
97:24
to be free and uh diverse we have free speech for a good reason is
且多元化。我们之所以有言论自由,是有充分理由的,因
97:31
because uh we don't want all
为我们不希望所有...
97:34
of our information to be uh to come from a unique Source um because that's
我们获取的信息不应该只来自单一来源,因为这完全违背了民主的理念,也违背了思想进步甚至科学发展的本质——在科学领域,人们需要争论不同的观点,科学正是在分歧中取得进步,最终达成共识,形成一致意见。全世界所有民主国家都是如此。所以,有一个未来已经在发生:我们与数字世界的每一次互动都将由AI系统、AI助手来中介。我们会戴上智能眼镜,现在你
97:39
opposite to the whole idea of democracy and uh you know progress of ideas and
就能从MAA买到,就是那个叫MAA的,你可以跟它对话,它连接着一个LLM,能回答你任何问题;或者你看着一座纪念碑,眼镜系统里有摄像头,你可以问它“能告诉我这栋建筑或这座纪念碑的信息吗”;你看着一份外语菜单,它能帮你翻译;如果我们说不同的语言,它还能做实时翻译。所以在不久的将来,我们与数字世界的很多互动都将由这些系统来中介。而且,我
97:45
even science right in in science people have to argue for different opinions and and
们使用的搜索引擎将不再是搜索引擎,而是对话系统——你直接问一个问题,它就会回答,然后可能给你指向合适的参考资料。但问题是,我们不能让这些系统只来自美国西海岸的少数几家公司,因为这些系统将构成全人类知识的存储库,我们不能让它们被一小部分人控制。出于同样的原因,媒体必须多元化,AI助手也必须多元化。那么,我们如何获得多样化的AI助手呢
97:50
science makes progress when people disagree and they come up with an answer and you
?目前训练一个基础模型——一个基础LLM——非常昂贵且困难,未来可能会有所不同,但现在就是LLM,所以只有少数公司能做好这件事。如果其中一些顶级系统是开源的,任何人都可以使用,任何人都可以对它们进行fine-tune。如果我们建立一些机制,允许任何群体——无论是普通公民、公民团体、政府机构、非营利组织还是公司——拿这些开源的AI系
97:55
know a consensus forms right and it's true in all democracies around the
统,用自己的数据针对自己的目的进行fine-tune,那么我们就会拥有大量多样化的、专门针对各种需求的AI系统。所以,我告诉你,我跟法国政府谈过不少次,法国政府绝不会接受所有公民的数字生活被美国西海岸的三家公司控制,这完全不可接受,无论这些公司多么善意,这对民主都是威胁。而且,这也是一种
97:59
world so there is a future which is already happening where every single one of
所以,有一个未来已经在发生——我们与数字世界
98:05
our interaction with the digital world will be mediated by ai ai systems AI assistants
的每一次互动都将由AI系统、AI助手来中介。我
98:11
right we're going to have smart glasses you can already buy them from MAA the
们会戴上智能眼镜,你现在就能从MAA、从ran
98:17
ran MAA where um you know you can talk to them and they are connected
MAA买到,嗯,你可以跟它们说话,它们背后连
98:23
with an llm and
着一个LLM。
98:24
you can get answers on any question you have or you can be looking at
所以有一个已经在发生的未来:我们与数字世界的每一次互动,都将由AI、
98:28
a monument and there is a camera in the in the system that in in
AI系统、AI助手来中介。对吧,我们会戴上智能眼镜,你现在就能从MAA
98:32
the glasses you can ask it like can what can you tell me about this
买到,就是那个MAA,嗯,你可以跟它说话,它连着一个LLM,然后你问
98:36
uh building or this Monument you can be looking at a menu in a foreign
任何问题都能得到答案。或者你看着一座纪念碑,眼镜系统里有摄像头,你可以
98:40
language and I thing will translate it for you or we can do real time
问它:“关于这栋建筑或这个纪念碑,你能告诉我什么?”或者你看着一份外
98:44
translation if we speak different languages so a lot of our interactions with the digital
语菜单,它就能帮你翻译。如果我们说不同的语言,还能做实时翻译。所以在不
98:48
world are going to be mediated by those systems in the near
久的将来,我们与数字世界的很多互动都将由这些系统来中介。
98:51
future um you know increasingly the search engines that we're going to use are not
未来,嗯,你懂的,我们使用的搜索引擎会越
98:57
going to be search engines they're going to be uh dialog systems that would just
来越少,取而代之的会是对话系统——你直接
99:03
ask a question and it will answer and then point you to perhaps appropriate reference
问一个问题,它就会回答,然后可能给你指向
99:09
for it but here is the thing we cannot afford those systems to come from
相关的参考资料。但问题是,我们不能让这些
99:15
a handful of companies on the west coast of the
系统只来自西海岸那几家公司。
99:19
US because those systems will constitute the repository of all human knowledge and we cannot
嗯,你懂的,我们将来用的搜索引擎,将不再是传统意义上的搜索引擎,而是对话系统——你直接问
99:24
have that be controlled by a small number of people right it has to be
一个问题,它就会回答,然后可能给你指向合适的参考资料。但问题是,我们不能让这些系统只来自
99:29
diverse for the same reason the Press has to be diverse so how do we
美国西海岸的那几家公司,因为这些系统将构成全人类知识的存储库,我们不能让少数人控制它,对
99:34
get a diverse set of AI assistant um it's very expensive and difficult to train
吧?它必须多元化,理由和媒体必须多元化一样。那么,我们如何获得多样化的AI助手呢?目前训练
99:39
a based model right a based llm at the moment you know in the future
一个基础模型,一个基础LLM,非常昂贵且困难。嗯,未来可能不一样,但眼下就是LLM,所以
99:44
it might be something different
只有少数公司能真正做好这件事。
99:46
but at the moment that's an llm uh so only a few companies can do
未来,嗯,我们使用的搜索引擎将不再是搜索引擎,而
99:53
this properly and if some of those Tob systems are open source anybody can use
是对话系统——你直接问一个问题,它就会回答,然后可
100:00
them anybody can fine-tune them um if we put in place some systems that allows
能给你指向合适的参考来源。但问题是,我们不能让这些
100:07
any group of people whether they are um individual
系统只来自西海岸的那几家公司。
100:11
citizens groups of citizens government organizations NOS uh companies whatever to take those open source
如果其中一些顶尖系统是开源的,任何人都能用,任何人都能对它进行fine-tune。如果我们建立一些机制,允许任何群体——无论是
100:18
uh systems AI systems and fine-tune them for their own purpose on their own data
个人、公民团体、政府组织、非政府组织、公司等等——拿这些开源的AI系统,用自己的数据针对自己的目的进行fine-tune,那么
100:24
then we're going to have a very large diversity of uh different AI systems that
我们就会拥有非常多样化的、专门针对各种用途的AI系统。对吧,所以我告诉你,我跟法国政府谈过不少,法国政府不会接受他们所有公民的
100:31
are specialized for all of those things right so I tell you I
数字生活被美国西海岸的三家公司控制,这完全不可接受。无论那些公司多么善意,这对民主都是威胁,嗯,而且这也...
100:37
talked to the French government quite a bit and the French government will not accept
跟法国政府谈了很多次,法国政府不会接受他们所有公民的数字生活被美国西海岸的三家公司控制,这完全不可接受,不管这些公司初衷多好,这对民主都是威胁。嗯,而且这也是
100:44
that the digital diet of all their citizen be controlled by three companies on the
一个……一个广泛的社区,实际上有成千上万的企业在用这个技术构建应用。所以,我们——Meta——从这个技术中获取收入的能力,并不会因为开源基础模型的分发而受损。G
100:50
west coast of the US that's just not acceptable it's a danger to democracy regardless
emini 现在受到的根本批评是……效果啊之类的问题,而五起不良文本、图片、视频的公关事件,实际上又把那些例子放进了下一版本的训练数据里。所以这恰恰说明了这件
100:57
of how well-intentioned those companies are right um and so uh and it's also a
事有多难——各种人都不满意,就像你说的,你不可能造出一个让所有人都满意的系统。对,嗯,所以如果你打算自己做fine-tuning,并且保持闭源,那问题就在于……
101:03
danger to local culture to values to language right I was talking with um uh
对当地文化、价值观和语言构成威胁。对,我之前和印度Infosys的创始人聊过,他正在资助一个项目,对Meta开源的Llama 2模型进行fine-tuning,让Llama 2能说印度全部22种官方语言。这对印度人来说非常重要。我还和我以前的一位同事Mustafa聊过,他曾在FAIR做科学家,后来回到非洲,为Google在非洲创建了一个研究实验室,现在他创立了一家新公司叫Kara。他做的事情就是让LLM能说塞
101:09
the uh founder of infosis in India um he's funding a project to fine tune
内加尔的本地语言,这样人们就能获取医疗信息,因为那里医生很少,人均医生数量极低。在塞内加尔,如果没有开源平台,这一切都不可能实现。有了开源平台,你就能拥有AI系统,这些系统不仅在政治观点等方面多样化,而且在语言、文化、价值观、政治观点、各领域技术能力上也都多样化。你还能形成一个产业和生态系统,让公司针对工业垂直应用去fine-tuning这些开源系统。比如,一家出版社有几千本书,他们想建一个系统,让顾客能直接提
101:15
Lama 2 the open source model produced by by meta so that Lama 2 speak
问任何一本书的内容,那就需要用他们的专有数据来训练。Meta内部也有一个叫Metam的系统,它就是一个LLM,能回答关于公司内部任何问题,非常有用。很多公司都想要这个,不仅给员工用,也给客户用,用来服务客户。所以,要拥有一个AI产业,要拥有不带有单一偏见的AI系统,唯一的方式就是有开源平台,任何群体都能在上面构建专门的系统。历史不可避免的方向是,绝大多数AI系统都将建立在开源平台之上。这是一个美好的愿景。也就是
101:22
all 22 official languages in India it's very important for people in India I was
说,像Meta或Google这样的公司,在构建基础pre-trained model之后,应该只做最少的fine-tuning步骤,越少越好。但Meta能承担得起吗?不行。我不知道你知不知道,公司总得赚钱吧,而开源基本上就是白送。马克·扎克伯格发过一个视频,一个非常性感的视频,讲的是35万块Nvidia H100。光GPU的数学账就是1000亿美元,再加上训练所需的基础设施。我不是搞商业的,但靠这个怎么赚钱呢?
101:28
talking to a former colleague of mine Mustafa used to be a scientist at fair
和我以前的一位同事聊过,Mustafa 以前是 FAIR 的科学家,后来搬回了非洲,在非洲为 Google 创建了一个研究实验室,现在他创立了一家新公司叫 Kara。他做的事情基本上就是让 LLM 能说当地的非洲语言,这样人们就能获取医疗信息,因为他们那里医生很少,人均医生数量非常低。在塞拉利昂,我的意思是,没有开源平台这一切都无从谈起
101:33
and then moved back to Africa I created a research lab for Google in Africa
。有了开源平台,你就能拥有 AI 系统,这些系统不仅在政治观点之类的事情上多样化,而且在语言、文化、价值体系、政治观点、以及各个领域的技术能力上也是多样化的。然后你就可以有一个产业,一个由公司组成的生态系统,这些公司针对行业里的垂直应用对开源系统进行 fine-tuning,对吧?比如,一个出版商有几千本书,他们想建一个系统,让客户可以
101:37
and now is as a new startup Kara and what he's trying to do is
直接问任何一本书里的内容,那就需要用他们的专有数据来训练,对吧?还有一家公司,Meta 内部就有一个,叫 Metam,它基本上是一个 LLM,能回答任何关于公司内部事务的问题,非常有用。很多公司都想要这个,对吧?很多公司不仅想让员工用,还想让客户用,用来服务客户。所以,要想拥有一个 AI 产业,要想拥有没有单一偏见的 AI 系统,唯一的办
101:42
basically have llm that speak the local languages in Sagal so that people can have
法就是有开源平台,任何群体都可以在上面构建专门的系统。所以历史不可避免的方向是,绝大多数 AI 系统都将建立在开源平台之上。这是个美好的愿景。所以意思是,像 Meta 或 Google 这样的公司,在构建好基础的 pre-trained 模型之后,应该只做最少的 fine-tuning 步骤,步骤越少越好。那 Meta 能负担得起这样做
101:47
access to medical information because they don't have access to doctors it's a very small
吗?不能。我不知道你知不知道,但公司总得想办法赚钱,而开源基本上就是白送。我不知道,Mark 发过一个视频,Mark Zuckerberg,一个非常性感的视频,讲的是 35 万块 Nvidia H100。光 GPU 的账就算一下,那就是 1000 亿美元,再加上训练所需的所有基础设施。我不是搞商业的,但你怎么靠这个赚钱呢?举个例子,如果
101:52
number of doctors per per capita in the
你有一个 LLM,能帮一个夫妻店披萨店,通过 WhatsApp 和顾客聊天,顾客可以直接订披萨,系统会问他们想要什么配料、什么配菜之类的,那这家店就会为这个服务付费。这就是一种模式。
101:55
in syal um I mean you can't have any of this unless you have open
但目前,这还是个LLM的事,嗯,所以只有少数
102:01
source platforms so with open source platforms you can have ai systems that are not
公司能做好。如果这些对话系统是开源的,任何人都
102:07
only diverse in terms of political opinions or things of that type but in terms
能用,任何人都能fine-tune它们。如果我
102:14
of uh uh language culture value systems political opinions um technical abilities in various
们建立一些机制,让任何群体——无论是个人——
102:20
domains and you can have an industry an ecosystem of companies that fine-tune those open
我跟法国政府聊过不少,他们绝
102:24
source systems for vertical applications in Industry right you you have I don't know a
不会接受自己国家公民的数字生
102:29
publisher has thousands of books and they want to build a system that allows a
活被美国西海岸的三家公司控制
102:34
customer to just just ask a question about any about the content of any of
,这完全不可接受。不管这些公司
102:39
their books you need to train on their proprietary data right um You have a
初衷多好,这对民主都是威胁。
102:43
company we have one within meta it's called metam and it's
嗯,而且这也涉及到——
102:47
basically an llm that can answer any question about internal uh stuff about about the
我跟法国政府聊过不少,法国政府
102:52
the company U very useful a lot of companies want this right a lot of
不会接受他们所有公民的数字生活
102:56
companies want this not just for their employees but also for their customers to take
被美国西海岸的三家公司控制,这
103:01
care of the customers so the only way you're going to have an AI industry
完全不可接受。无论这些公司初衷
103:06
the only way you're going to have ai systems that are not uniquely biased is
多好,这对民主都是威胁。嗯,而
103:11
if you have open source
且这也——
103:13
platforms on top of which uh any group Can U build specialized systems so the
我们和法国政府谈过很多次,法国政府不会接受他们所有公民的
103:20
the direction of of inevitable direction of history is that the vast majority of AI
数字生活被美国西海岸的三家公司控制,这是完全不可接受的。无
103:28
systems will be built on top of Open Source platforms so that's a beautiful Vision
论这些公司多么善意,这对民主来说都是一种威胁。而且,这也是
103:35
so meaning
一个……
103:36
like a company like meta or Google or so on should take only minimal fine-tuning
像Meta或Google这样的公司,在构建完基础预训练模型之后,应该只做最少的fine-tuning步骤,能少就少。但Meta能这么做吗?不能。我不知道你知不知道,公司总得想办法赚钱吧,而开源基本上就是白送。我不确定,但Mark发了个视频,Mark Zuckerberg,一个非常性感的视频,讲的是35万块Nvidia H100。算一下,光是GPU就要1000亿美元,再加上训练所需的基础设施。我不是搞商
103:42
steps after the building the foundation pre-trained model as few steps as possible basically can
业的,但靠这个怎么赚钱呢?举个例子,如果你有一个LLM,能帮一家夫妻店披萨店,通过WhatsApp跟顾客聊天,顾客可以直接订披萨,系统会问他们要什么配料、配菜之类的,然后商家为此付费。这是个商业模式。但如果你把开源模型放出去,别人也能做同样的事,跟你竞争,基本上就是为商家提供fine-tune模型。这其实是Meta在下的赌注——顺便说一句,我非常看好这一切——但Meta的赌注是“我们做得更好”?不,其
103:49
meta afford to do that no so I don't know if you you know this
实赌的是:我们已经有了庞大的用户群和客户群。啊,对,对。所以无论我们提供什么,对他们都有用,而且有办法从中获得收入。同时,我们把那个系统,或者说基础模型,也就是foundation model,以开源形式提供给其他人,让他们在上面构建应用,这也没什么坏处。如果那些应用对我们的客户没用,我们可以直接从他们那里买下来。实际上,他们可能会改进这个平台,我们已经看到了。我是说,Llama 2已经被下载了数百万
103:55
but companies are supposed to make money somehow and uh open source is is is
次,成千上万的人提供了改进建议。所以很明显,把系统开放给一个广泛的社区,能加速进步。而且有成千上万的企业在用这个模型构建应用。所以Meta从这项技术中获取收入的能力,并不会因为开源发布基础模型而受损。Gemini受到的根本批评是——就像你在西海岸指出的那样——澄清一下,我们现在在东海岸,我猜Meta AI总部会在那边。所以,嗯,对西海岸有些尖锐的话。但我觉得,问题在于,公平地说,大多数科技人都有左翼的
104:01
like
政治倾向,他们偏左,所以……
104:01
giving away I don't know Mark made a video Mark Zuckerberg um very sexy video
各个领域,你可以有一个行业、一个生态系统,里面有很多公司对这些开源系统
104:09
talking about 350,000 Nvidia h100s yeah the the math of that is just for the
做微调,用于垂直行业的应用。比如,一家出版社有几千本书,他们想建一个系
104:16
gpus that's 100 billion um plus the infrastructure for training everything so I'm no business
统,让客户可以直接问任何一本书的内容,那就需要用他们的专有数据来训练。嗯
104:24
guy but how do you make money on that so
,Meta内部就有这样一个例子,叫Metam——
104:29
the division you payt is a really powerful one but how is it possible to
你提到的这个划分确实很有道理,但怎么才能赚钱呢?好,其实有好几种商业模式,对吧。MAA 构建的商业模式是“服务即产品”,而这项服务的资金来源要么是广告,要么是企业客户。举个例子,如果你有一个 LLM,能帮一家夫妻店披萨店通过 WhatsApp 和顾客聊天,顾客可以直接下单披萨,系统会问他们想要什么配料、配菜之类的,那么这家店就会为此付费,这是一种模式。另外,如果系统偏向更传统的服务类型,那就可以靠广告支持,或者还有其他几种模式。但关键是
104:36
make money okay so you have several business models right the business model that uh
,如果你有足够大的潜在客户群,而且你本来就需要为他们搭建这个系统,那么把它以开源形式发布出去对你也没什么坏处。我不是搞商业的,但如果你发布开源模型,其他人也能做类似的任务并参与竞争,基本上就是为企业提供微调模型。这其实是 Meta 在下的赌注。顺便说一句,我特别看好这个方向,但 Meta 的赌注是不是说“我们会做得更好”?不,其实赌注更偏向于:我们已经拥有庞大的用户群和客户群。啊,对,没错。所以无论我们提供什么,对他们来说都会有用,而且有
104:43
MAA is built around is um youer a service and the the financing of that
办法从中获得收入。而且,我们把这个系统或者基础模型——也就是 foundation model——以开源形式提供出来,让别人在上面构建应用,这也没什么坏处。如果那些应用对我们的客户没用,我们也可以直接从他们那里买过来。事实上,他们可能会改进这个平台,我们已经看到这种情况了。我的意思是,Llama 2 已经被下载了数百万次,成千上万的人提供了改进建议。所以,把系统开放给一个广泛的社区,显然会加速进步。而且,确实有成千上万的企业在用这个模型
104:50
service is uh either through ads or through business customers so for
构建应用。所以,我们——也就是 Meta——从这项技术中获取收入的能力,并不会因为把基础模型以开源形式分发出去而受到损害。Gemini 受到的根本批评是,就像你在西海岸指出的那样——澄清一下,我们现在在东海岸,我猜 Meta AI 总部应该在那里,所以对西海岸有些狠话——但我觉得问题在于,公平地说,大多数科技人士在政治上都倾向于左翼,他们偏左,所以……
104:55
example if you have an llm that uh you know can help a mom and
我跟法国政府谈过不少次,法国政府不会接受他们所有公民的数字饮食被美国西
105:00
pop pizza place um by you know talking to their customers through WhatsApp and so
海岸的三家公司控制,这完全不可接受。无论这些公司初衷有多好,这对民主都是
105:05
the customers can just order a pizza and the system will just you know ask
个威胁。所以,还有一个例子,如果你有一个LLM,能帮一家夫妻店披萨店通
105:10
them like what topping do you want or what sites blah blah blah um the
过WhatsApp跟顾客聊天,顾客可以直接点披萨,系统会问你要什么配料、
105:15
business will pay for that okay that's a model
配菜之类的,然后商家为此付费,这是一种模式。
105:18
um and otherwise you know if it's a system that is on the more kind
嗯,还有,如果是一个更偏向传统服务的系统,那可能靠广告支持,或者你知道,有好几种模式。但关键是,如果你有足够大的潜在客户群,而且你本来就需要为他们搭建那个系统,那把它开源发布对你也没什么坏处。我不是搞商业的,但如果你发布开源模型,其他人就能做同样的任务,跟你竞争,基本上就是为企业提供微调模型。这其实是Meta在打的赌——顺便说一句,我
105:24
of classical Services it can be uh ad supported or you know there's several models
特别看好这一切——但Meta打的赌是,我们会做得更好。不,其实赌的是,我们已经有了庞大的用户群和客户群。啊,对,对。所以无论我们提供什么,对他们都有用,而且有办法从中获得收入。而且,我们把那个系统或基础模型、那个foundation model开源,让别人在上面构建应用,也没什么坏处。如果那些应用对我们的客户没用,我们也可以直接从他们那
105:30
but the point is uh if you have a big enough um potential customer base
里买下来。嗯,他们甚至可能改进这个平台——事实上我们已经看到了。我的意思是,Llama 2有数百万次下载,成千上万的人提出了改进建议。所以,让这个系统面向广泛的社区开放,显然加速了进步。而且有成千上万的企业在用这个模型构建应用。所以,Meta从这项技术中获取收入的能力,并不会因为开源发布基础模型而受损。Gemini受到的根本批评是——
105:37
and you need to build that system anyway for them it doesn't hurt you to
就像你在西海岸指出的那样——澄清一下,我们现在在东海岸,我猜Meta AI总部应该在那里。所以,对西海岸有些狠话。但我觉得,公平地说,大多数科技人士在政治上都偏左翼。所以,问题不在于设计这些系统的人的政治倾向,而在于他们客户群体的接受度或政治倾向。对吧,一个大公司不能得罪太多人,所以他们要确保推出的任何产品都是安全的——不管那意味着什么
105:43
actually distribute it in open source again I'm
。而且很容易做得过火,也不可能让每个人都满意。所以就像我之前说的,你不可能有一个系统是每个人都觉得没有偏见的。你往一边推,一群人会觉得有偏见;你往另一边推,另一群人又会觉得有偏见。
105:46
no business guy but if you release the open source model then other people can
在这些平台之上,任何团队
105:51
do the same kind of task and compete on it basically provide fine-tune models for
都可以构建专门的系统。所以
105:56
businesses the is the bet that meta is making by the way I'm a huge
历史的大方向是,绝大多数A
106:01
fan of all this but is is the bet that meta is making is like
I系统都会建立在开源平台之
106:06
we'll do a better job of it well no the the bet is is more
上。这是个很美好的愿景,
106:11
we have we already have a
也就是说——
106:13
huge user base and customer base ah right right so it's going to be useful
平台,任何群体都可以在上面
106:18
to them whatever we offer them is going to be useful and there is a
构建专门的系统。所以,历史
106:23
way to derive revenue from this uh and it doesn't hurt that you know we
的必然方向是,绝大多数AI
106:28
provide that system or the Bas the base model right the foundation model uh in
系统将建立在开源平台之上。
106:32
open source for others to build applications on top of it too if those applications
这是一个美好的愿景,意思是
106:37
are not
……
106:38
to be useful for our customers we can just buy it from them um uh
为了对我们的客户有用,我们直接向他们购买就行,嗯呃,他们可能会改进这个平台,实际上我们已经看到了这一点。嗯,我是说,你知道,Llama 2 有 literally 数百万次下载,成千上万的人提供了改进建议。嗯,所以,这 clearly 加速了进展,让系统对广泛的社区可用,而且有 literally 数千家企业在用它构建应用。所以,嗯,Meta 从这项技术中获
106:43
it could be that they will improve the platform in fact we see this already
取收入的能力,并不会因为开源基础模型的发布而受损。Gemini 受到的根本批评是,正如你提到的西海岸——先澄清一下,我们现在在东海岸,我猜 Meta AI 总部会在那边——所以,嗯,对西海岸有些激烈的言辞,但我觉得问题在于,公平地说,大多数 tech 人士在政治上倾向于左翼,他们偏左,所以,不管那意味着什么,很容易做得过火,而且也不可能让每个人都满意。所以就像
106:49
um I mean there is you know literally millions of downloads of Lama 2 and
我之前说的,你不可能搞出一个被所有人视为无偏见的系统。你往一边推,一群人会觉得有偏见;你往另一边推,另一群人又会觉得有偏见。他指出的 big Tech 的问题是,他问 big Tech 到底能不能推出生成式 AI 产品,面对一、来自内部活动家、员工暴民、疯狂高管、破碎董事会、压力团体、极端监管机构、政府机构、媒体、所谓的专家等等不断升级的要求,污染输出;二、随
106:54
thousands of people who have you know provided ideas about how to make it better
时可能生成糟糕回答的风险;三、有效性和诸如此类的东西;四、糟糕的文本、图片、视频被公开,实际上这些例子又进入了下一版本的训练数据。所以他只是强调了这有多难,因为各种人都不满意,就像你说的,你没法创造一个让所有人都开心的系统。是的。嗯,所以如果你打算自己做 fine-tune 并保持 closed source,问题在于,国会调查是其中之一,法律责任,嗯,你知道
107:00
um so you know this this clearly accelerates progress to make the system available to
,搞出让人伤害自己或他人的东西,嗯,大公司对此非常谨慎,首先他们不想伤害任何人,其次他们想保住自己的业务。所以,嗯,对于这类系统来说,基本上不可能避免,因为它们 inevitably 会形成政治观点,以及关于各种可能政治或不政治的事情的观点,但人们可能会在道德问题、宗教问题、文化问题上产生分歧,不同社区的人一开始就会不同意。嗯,所以只有相对少数的情况是可行的。
107:05
a sort of a a wide community of people and and there is literally thousands
你做得足够多,就能下意识地完成,对吧,不用思考。如果你是个有
107:11
of businesses who are building applications with it so um so our ability to meta's
经验的司机,你可以不假思索地开车,同时还能跟人聊天或者听广播
107:18
ability to derive revenue from this technology is not impaired uh by the distribution of
,对吧?如果你是个非常熟练的棋手,你可以跟一个新手对弈,也不用
107:24
it of based models in open source the fundamental criticism that Gemini is getting is
怎么思考,你只是识别出模式然后下棋,嗯,那就是系统一。所以所
107:30
that
有……
107:31
as you point out on the west coast just to just to clarify we're currently
举个例子,如果你有一个LLM,可以帮助
107:36
in the east coast where I would suppose meta AI headquarters would be so there
一家夫妻店披萨店,比如通过WhatsA
107:42
uh strong words about the West Coast but uh I guess the issue that happens
pp和顾客聊天,顾客可以直接点披萨,系统
107:48
is I think it's fair to say that most tech people have a political affiliation
会问他们想要什么配料、配菜等等。那么这
107:53
with the left wing they're they lean left and so the
家店会为此付费,这是一种模式。
107:57
problem that people are criticizing Gemini with is that there's in that debiasing process that
大家批评Gemini的一个问题是,在你提到的那个去偏见过程中,他们的意识形态倾向变得很明显。你觉得这是可以避免的吗?你是在说开源是唯一的办法?你见过这种让工程变得困难的意识形态倾向吗?不,我不认为这跟设计这些系统的人的政治倾向有关,而是跟他们的客户群体或受众的可接受度或政治倾向有关。对吧,大公司不能得罪太多人,所以他们要确保推出的任何产品都是“安全”的,不管这个词到底是什么意思。而且很容易做过头,也不可能让每个人都满意。你不可能让所有人都满意。所以就像我之
108:05
you mentioned that their ideological lean becomes obvious uh is this something that could be
前说的,你不可能有一个被所有人视为无偏见的系统。你往一个方向推,一群人会觉得有偏见;你往另一个方向推,另一群人又会觉得有偏见。除此之外,如果你把系统往一个方向推得太远,它就会变得不准确,对吧?就会出现那种黑人士兵是纳粹的情况——我们应该提一下生成黑人士兵纳粹图像这件事,这在事实上是不准确的,对吧?而且对某些人来说也可能很冒犯。所以,要制造出对所有人都无偏见的系统是不可能的。我看到的唯一解决方案就是多样性,而且是这个词的完整含义——全方位的多样性。嗯,Mark
108:12
escaped you're saying open source is the only way have have you witnessed this kind
Andreessen今天发了条推文,我总结一下:结论是只有初创公司和开源才能避免他指出的那个大科技公司的问题。他在问:大科技公司真的能推出生成式AI产品吗?面对一,来自内部活动家、员工暴民、疯狂高管、瘫痪董事会、压力团体、极端监管机构、政府机构、媒体所谓的“专家”等等不断升级的要求,这些都在腐蚀输出;二,随时可能生成糟糕答案、画出糟糕图片、渲染糟糕视频的风险,谁知道它会说什么或做什么;三,法律风险,产品责任、诽谤、选举法等等,任何会让国会生气的事;四,持
108:19
of ideological lean that makes engineering difficult no I don't think
续收紧对不可接受输出的控制,实际上是在降低模型的质量,比如它的可用性、使用体验、有效性等等;五,糟糕的文本、图像、视频被公开,实际上这些例子又成了下一版训练数据的一部分。所以他只是强调了这有多难,因为各方都不满意,就像你说的,你不可能创造一个让所有人都满意的系统。是的。那么,如果你自己做fine-tuning并保持闭源,问题本质上就是……
108:25
it has to do I don't think the issue has to do with the political
这个问题不在于设计系统的人的政治倾向,而在于他们客户群体的接受度或政治倾向。大公司不能得罪太多人,所
108:31
leaning of the people designing those systems it has to do with the uh acceptability
以他们会确保推出的任何产品都是安全的——不管“安全”具体指什么。而且很容易做得过火,也不可能让每个人都
108:37
or political leanings of the their customer based audience right so a big company cannot
满意。就像我之前说的,你不可能造出一个系统,让所有人都觉得它没有偏见。你往一边推,一群人觉得有偏见;你
108:43
afford to offend too many people so they're going to make sure that whatever product
往另一边推,另一群人又觉得有偏见。所以,要造出对所有人都没有偏见的系统是不可能的。我看到的唯一解决方案
108:50
they put out is safe
就是多样性,而且是全方位的多样性。
108:52
whatever that means and and it's very possible to overdo it and it's also very
不管那意味着什么,而且很容易做得过火,
108:56
possible to it's impossible to do it properly for everyone you're not going to satisfy
也很可能根本没法让所有人都满意——你不可
109:01
everyone so that's what I said before you cannot have a system that is unbiased
能让每个人都满意,所以我之前就说过,你没
109:05
is perceived as unbiased by everyone it's going to be you know you push it
法搞出一个对所有人来说都毫无偏见的系统。
109:10
in one way one set of people are going to see it as biased and
你往一边推,一群人会觉得有偏见;你往另一
109:15
then you push it the other way and another set of
边推,另一群人又会觉得有偏见。
109:18
people is going to see it that's biased and then in addition to this there's
大家看到就会觉得有偏见,而且除此之外还有个问题:如果你把系统往某个方向推得太远,它就会变得不真实,对吧?就会出现那种,你知道,黑皮肤的纳粹士兵——我们应该提一下图像生成中黑皮肤纳粹士兵的例子,这在事实上是不准确的,对吧?而且对某些人来说也可能冒犯,对吧?所以,嗯,你知道,要做出对所有人都没有偏见的系统是不可能的。所以我看到的唯一解决方案就是多样性,而
109:23
the issue of if you push the system perhaps a little too far in One
且这个词要取其完整含义——全方位的多样性。对。嗯,Mark Andre今天刚发了条推文,我帮你总结一下:结论是只有初创公司和开源才能避免他指出的那些大科技公司的问题。他问的是:大科技公司真的能推出生成式AI产品吗?第一,来自内部活动家、员工暴民、疯狂高管、失控董事会、施压团体、极端监管机构、政府机关、媒体、所谓的“专家”等等一切力量不断升级的要求,都在
109:28
Direction it's going to be non-factual right you're going to have you know you know
扭曲输出。第二,随时可能生成糟糕的回答、画出糟糕的图片、渲染出糟糕的视频——谁知道它下一秒会说什么或做什么。第三,法律风险:产品责任、诽谤、选举法,还有很多其他问题,总之任何会让国会生气的事。第四,持续收紧对不可接受输出的控制,实际上是在降低模型的质量——比如它原本有多好用、多让人愉快、多有效等等。第五,糟糕的文本、图像、视频被公开后,这些例子反而会进
109:33
black Nazi uh soldiers in the we should we should mention image generation of of
入下一版本的训练数据,如此循环。所以他只是强调了这件事有多难——正如你所说,来自各种人的不满,你不可能做出让所有人都满意的系统。对。嗯,所以如果你要自己做微调并保持闭源,那问题就在于:国会调查是其中之一,法律责任也是,嗯,你知道,做出那些会让人伤害自己或伤害他人的东西——比如,大公司真的很小心,不会生产这类东西,嗯,因为他们首先不想伤害任何人,其次想
109:39
uh black Nazi soldiers which is not factually accurate right and can be offensive for
保住自己的生意。所以,嗯,对于这类系统来说,基本上是不可能的,因为它们不可避免地会形成政治观点,以及关于各种可能政治或不政治的事情的观点,但人们可能会在道德问题、宗教问题、文化问题上产生分歧——不同社群的人本来就会在这些事情上意见不一,对吧?嗯,所以只有相对少数的事情是人们基本能达成一致的,比如基本原则。但除此之外,如果你想让这些系统有用,它们就必然会
109:44
some people as well right so
在某种程度上冒犯一些人。所以开源就是更好的选择,多样性也是更好的选择,对吧?而开源恰恰能实现多样性。没错,开源能实现多样性,这一点非常有意思。
109:46
uh so you know it's going to be impossible to kind of produce systems that
对了,Marc Andreessen今天发了条推文,我帮你总结一下。结论是:只有初创公司和开源才能避免他指出的那些大科技公司的问题。他在问:大科技公司真的能推出生成式AI产品吗?第一
109:51
are unbiased for everyone so the only solution that I see is diversity and diversity
,内部活动家、员工暴民、疯狂的 executives、破碎的董事会、压力团体、极端监管机构、政府机构、媒体和所谓的“专家”等等,不断升级的要求,全都在腐蚀输出。第二,随时有风险生成糟糕
109:57
in the full meaning of that word diversity of in every possible way yeah uh
的答案、画糟糕的图、渲染糟糕的视频——谁知道它下一秒会说什么或做什么。第三,法律风险:产品责任、诽谤、选举法,还有很多其他东西,总之任何让国会生气的事。第四,不断试图收紧对不可接受输
110:03
Mark Andre just tweeted today let me do a tldr the conclusion is only startups
出的控制,反而降低了模型的质量——比如它实际用起来好不好、顺不顺手、有没有效等等。第五,糟糕的文本、图像、视频被公开后,这些例子又会进入下一版本的训练数据。所以他只是强调了这有多难,就
110:09
and open source can avoid
像你说的,各方都不满意,你不可能造出一个让所有人都开心的系统。
110:11
the issue that he's highlighting with big Tech he's asking can big Tech actually field
他指出的问题是,大型科技公司真的能推出生成式AI产品吗?一方面,来自内部活动人士、员工群体、疯狂的 executives、混乱的董事会、压力团体、极端监
110:18
generative AI products one ever escalating demands from internal Act activists employee mobs crazed Executives
管机构、政府机构、媒体和所谓的“专家”等不断升级的需求,全都在腐蚀输出内容;另一方面,还要持续承担生成错误答案的风险,以及各种有效性之类的问题。第五点是,
110:26
broken boards pressure groups extremist Regulators government agencies the press in quotes experts and everything
那些糟糕的文本、图片、视频一旦公开,实际上会被纳入下一版本的训练数据中。所以他只是强调了这有多难,因为各种人都会不满意——就像你说的,你不可能创造一个让
110:33
uh corrupting the output two constant risk of generating a bad answer
所有人都满意的系统。是的,所以如果你打算自己做 fine-tuning 并保持 closed source,那问题就在于……
110:39
or drawing a bad picture or rendering a bad video who knows what is going
是的。所以如果你自己做 fine-tune 并且保持闭源,那问题就在于:国会调查是其中之一,法律责任
110:45
to say or do at any moment three legal exposure product liability slander election law
也是。还有,做出让人伤害自己或伤害他人的东西——大公司非常小心不生产这类东西,首先他们不想伤害任何人
110:51
many other things and so on anything that makes Congress mad four continuous attempts to
,其次他们想保住自己的生意。所以,对于这类系统来说,基本上是不可能的,因为它们不可避免地会形成政治观
110:57
tighten grip un acceptable output degrade the model like how good it actually is uh
点,以及各种可能涉及政治或不涉及政治、但人们会争论的道德问题、宗教问题、文化问题——不同社区的人本来
111:03
in terms of usable and pleasant to use and
就会在这些问题上产生分歧。所以,实际上只有相对很少的一部分……
111:07
effective and all that kind of stuff and five publicity of bad text images video
在那之前,嗯,30年前,我们在做Com Nets和早期神经网络的时候。嗯,我超级兴奋,因为我看到了一条通往
111:12
actually puts those examples into the training data for the next version so on so
人类级别智能的路径,你知道,就是那些能理解世界、记忆、规划、推理的系统。嗯,有一些……我管它叫“圆珠笔”,然
111:17
he just highlights how difficult this is from all kinds of people being unhappy as
后Twitter上就炸了:“天哪,人们可以用它写可怕的东西,比如 misinformation、propa
111:23
you said you can't create a system that makes everybody happy yes uh so if
ganda、hate speech,赶紧禁掉!”然后那些写“doomers”的人就来了,就像AI doome
111:28
you're going to do the fine tun yourself and keep a close Source essentially the
rs一样:想象一下如果每个人都能拿到一支圆珠笔,这可能会毁灭社会,应该立法禁止用圆珠笔写仇恨言论,现在就要监
111:33
problem there is
管圆珠笔!然后还有……
111:34
then trying to minimize the number of people who are going to be unhappy y
然后尽量让不开心的人越少越好,嗯,你其实是在说,这事儿几乎不可能做到完美,而开源是更好的方式。对,我的意思是,Mark 说的有几件事确实没错,那些事情确实让大公司很害怕,嗯,你懂的,比如国会调查就是其中之一,还有法律责任,嗯,还有做出那些会让人伤害自己或伤害别人的东西。大公司在这方面真的很小心,不想做出这类东西,因为首先他们不想伤害任何人,其次他们想保住自己的生意。所以基本上,像这样的系统几乎不可能避免会形成政治观点,还
111:40
um and you're saying like the only that that almost impossible to do right and
有对各种事情的看法,这些事可能涉及政治也可能不涉及,但人们可能会在道德问题上有分歧,嗯,还有像宗教问题之类的,或者文化问题,不同社群的人本来就会有不同的看法。所以其实只有相对少数的事情是人们能达成共识的,比如一些基本原则。但除此之外,如果你想让这些系统有用,它们就必然会冒犯到一些人,这是不可避免的。所以开源就是更好的选择,多样性也是更好的,对吧?开源能带来多样性,没错,开源能带来多样性。这会是一个很精彩的世界,如果开源世界
111:46
it's the better ways to do open source phasically yeah I mean he's Mark is
真的由 Meta 带头,创造出这种开源基础模型的世界,那么各国政府就会需要找到新的模型,嗯,然后可能投票给左派和右派的人也会有自己的模型和偏好,可以自己选择,这可能会让我们更加分裂。但那是我们人类自己的事,我们得自己去解决。基本上,技术让人类能更有效地做人类该做的事,而人类提出的所有棘手的伦理问题,最终还是要靠我们自己来解决。对,我是说,有些东西是有底线的,就像言论自由也有底线一样,这些系统被允许生成的内容也必须有一定的
111:51
right about uh a number number of things that you list that indeed scare um
限制,嗯,需要一些 guard rails。所以这也是我一直感兴趣的一点,就是我们之前讨论的那种架构,系统的输出是通过 inference 来满足某个目标,而这个目标可以包含 guard rails。我们可以在开源系统里加入 guard rails,如果我们最终按照这个 blueprint 来构建系统,我们可以在这些系统里加入 guard rails,确保有一套最低限度的 guard rails,让系统不危险、不 toxi
111:57
large companies uh you know certainly
c 等等,就是那些大家都能达成共识的基本东西。嗯,然后人们后续加的 fine-tuning,或者额外加的 guard rails,就会根据他们各自的社群来定制,就是这样。
111:59
Congressional investigations is one of them legal liability uh you know uh making things that
国会调查是其中之一,还有法律责任,以及做出那些可能导致人们伤害自己或他人的东西。你知道,大
112:05
get people to you know hurt themselves or hurt others like you know um big
公司在这方面非常谨慎,不想产出这类内容,首先是因为他们不想伤害任何人,其次是为了保住自己的业
112:11
companies are really careful about not um producing things of this type and um because
务。所以,对于这类系统来说,基本上是不可能避免的,因为它们不可避免地会形成政治观点,以及对各
112:17
they have you know they want to hurt anyone first of all and then second
种可能政治或不政治的事情的看法,但人们可能会在道德问题、宗教问题、文化问题上产生分歧,不同社
112:23
they want to preserve their business so um it's essentially
群的人从一开始就会不同意。所以,实际上只有相对较小的一部分……
112:27
impossible for systems like this that can inevitably formulate political opinions and you know opinions
有效啊,还有那些乱七八糟的。五条关于糟糕文本、图片、视
112:32
about various things that may be political or not but that people may disagree about
频的曝光,实际上是把这些例子放进了下一版本的训练数据里,
112:37
about you know moral issues and you know um things about like questions about religion
以此类推。所以他只是强调了这件事有多难,因为各种人都不满
112:42
and things like that right or or cultural issues that people from different communities would
意,就像你说的,你没法创造一个让所有人都开心的系统。对,
112:47
disagree with in the first place um so there's only kind of a relatively small
嗯,所以如果你打算自己做微调,并且保持闭源,那问题就在于
112:53
number
——
112:53
of things that people will uh sort of agree on you know basic principles but
大家会认同的基本原则其实就那么些,但如果你想
112:59
beyond that if you if you want those systems to be useful they will necessarily
让这些系统真正有用,它们就必然会冒犯到不少人,
113:05
have to uh offend a number of people inevitably and so open source is just
这是不可避免的。所以开源更好,多样性也更好,
113:11
better and then diversity is better right and open source enables diversity that's right open
对吧?开源能带来多样性,没错,开源确实能带来多
113:17
source enables diversity that this can be fascinating
样性,这一点非常有意思。
113:20
world where if it's true that the open source world if meta leads the way
如果开源世界真的由Meta引领,创造
113:26
and creates this kind of Open Source Foundation model world there's going to be like
出这种开源基础模型的世界,那么政府会找
113:31
governments will have a find new model and yeah and and then potentially uh you
到新的模式,然后,嗯,可能投票左派和右
113:37
know people that vote left and right will have their own model and preference and
派的人都会有自己的模型和偏好,能够自由
113:43
be able to choose and it will potentially divide us even more
选择,而这可能会让我们更加分裂。
113:47
but that's on us humans we get to figure out basically the technology enables humans
但这取决于我们人类自己,我们得想清楚——技术
113:52
to human more effectively and all the difficult ethical questions that humans raise will just
本质上让人类更高效地“做人”,而人类提出的所有
113:57
it'll um leave it up to us to figure it out yeah I mean there
棘手的伦理问题,最终还是要靠我们自己解决。是啊
114:01
are some limits to what you know the same way there are limits to free
,我觉得有些东西是有底线的,就像言论自由也有底
114:06
speech there has to be some limit to the kind of stuff that those systems
线一样,这些系统能生成什么内容也必须有一定的限
114:11
might
制。
114:11
be authorized to um to produce um you know some guard rails so I mean
嗯,要有一些护栏。所以这也是我一直感兴趣
114:17
that's one thing I've been interested in which is uh in the type of architecture
的一点——在我们之前讨论的那种架构里,系统
114:23
that we were discussing before where the output of a system is the result of
的输出是为了满足某个目标而进行推理的结果,
114:29
an inference to satisfy an objective that objective can include guard rails and uh we
那个目标可以包含护栏。而且我们可以在开源系
114:35
can put guard rails in open source systems I mean if we
统里加入护栏。我的意思是,如果我们
114:39
eventually have systems that are built with this blueprint uh we can put guard rails
最终按照这个蓝图构建系统,我们可以在这些系统
114:44
uh in those systems that guarantee that there is sort of a minimum set of
里设置护栏,确保有一套最低限度的防护措施,让系
114:49
guardrails that make the system non- dangerous and nontoxic Etc you know basic things that
统不会变得危险、不会产生有害内容等等。就是那些
114:54
everybody would agree on um and and then you know the the fine tuning that
大家都能达成共识的基本东西。嗯,然后,人们后续
114:59
people will add or the additional guardwell that people will add will kind of cater
添加的微调或者额外的护栏,会更多地服务于他们
115:04
to their um Community whatever it is and yeah the
自己的社区,不管是什么社区。对。
115:07
fine doing will be more about the gray areas of what is hate speech what
精细化的处理更多会涉及那些灰色地带,比如什么是仇恨言论、什么是危险内容之类的。你知道,不同的价值体系嘛。但即便如此,就拿如何制造生物武器来说吧,我记得你评论过,或者至少有一篇论文提到,一群研究人员正在试图理解这些LLM的社会影响。其中一个门槛挺有意思的:就是LLM是否比搜索引擎(比如Google搜索)更容易让人做到这些事?现在越来
115:12
is dangerous and all that kind of stuff I mean you've different value systems value
越多的研究似乎表明,它并没有帮助。也就是说,如果你已经能访问搜索引擎和图书馆,LLM并不会帮你设计或制造生物武器或化学武器。你获得的信息量增加,或者获取信息的便捷性提升,实际上并不会真正帮到你。这是第一点。第二点是,有一份如何制造化学武器或生物武器的指令清单是一回事,真正动手去造又是另一回事,这比你想的要难得多,LLM在这方面帮不了
115:17
systems I mean like uh but still even with the objectives of how to build
你。事实上,世界上没有人——甚至包括那些国家——会使用生物武器,因为大多数时候他们根本不知道如何保护自己的人口免受其害。所以这东西太危险了,根本没法用,而且实际上已经被国际条约禁止了。化学武器则不同,它也被条约禁止,但问题是类似的:很难在不反噬使用者的情况下使用。不过我们可以问问你,马斯克,我可以给你一份非常精确的火箭发动机制造指
115:22
a bioweapon for example I think something you've commented on or at least there's a
令清单,但即使你有一个由50名经验丰富的工程师组成的团队,你仍然得炸掉十几个原型才能让它正常工作。你知道,化学武器、生物武器之类的东西也是一样。它们需要现实世界中的专业知识,而LLM在这方面帮不了你。它甚至需要我们在讨论的那种常识性专业知识——如何把基于语言的指令在物理世界中实现出来,这需要大量不在指令里的知识。没错,很多生物学家其
115:27
paper where a collection of researchers is trying to understand the social impacts of these
实已经对此发表过看法,回应那些说法时表示:你们知道实际做实验有多难吗?这可不是闹着玩的。是的,Hans Marik又一次被提到了。再聊聊Llama吧,你知道Mark宣布Llama 3最终会发布,但我觉得还没有具体的发布日期。首先,Llama 2已经出来了,你对未来的Llama 3、4、5、6、10,或者说Meta开源模型的未来,最期
115:32
llms and I guess one threshold is nice
待的是什么?嗯,有好几点。首先会有各种版本的Llama,都是对之前版本的改进,更大、更好、支持多模态之类的。然后在未来的版本中,会有能够规划、真正理解世界运作方式的系统。也许吧。
115:35
is like does the llm make it any easier than a than a search would
问题是,LLM比搜索引擎(比如Google搜索)更容
115:41
like a Google search would right so the increasing uh number of studies on this
易做到这些吗?越来越多的研究似乎表明,它并没有帮助。所
115:48
seems to point to the fact that it doesn't help so having an llm doesn't
以,如果你已经能访问搜索引擎和图书馆,LLM并不会帮
115:54
help you design or build a bioweapon or a chemical weapon
你更容易地设计或制造生物武器或化学武器。
115:59
if you already have access to uh you know a search engine and a library
你获得的信息量增加或者获取信息的便
116:03
uh and and so the the S of increased information you get or the ease
利性,实际上并不会真正帮到你。这是第
116:08
with which you get it doesn't really help you um that's the first thing the
一点。第二点是,拿到一份如何制造化
116:12
second thing is it's one thing to have a list of instructions of how to
学武器或生物武器的指令清单是一回事,
116:17
make a chemical weapon for example or bioweapon it's another thing to actually build it
真正造出来是另一回事,而且这比你想象
116:21
and it's much harder than you might think and then LM will not help you
的要难得多。LLM在这方面也帮不了
116:26
with
你。
116:26
that um in fact you know nobody in the world not even like you know
事实上,世界上没有人——甚至包括那些国家——会使用生物
116:31
countries use bioweapons because most of the time they have no idea how to protect
武器,因为大多数时候他们根本不知道如何保护自己的人口免
116:36
their own populations against it so um so it's too dangerous actually to kind of
受其害。所以,嗯,这东西太危险了,根本没法用。而且它实际
116:42
ever use um and it's in fact banned by uh International treaties um chemical weapons
上被国际条约禁止了。化学武器不一样,它也被条约禁止了,
116:47
is different it's also banned by treaties U but um uh but it's the same
但嗯,它面临同样的问题:很难在不反噬使用者的情况下使用。
116:52
problem it's difficult to use in situations that doesn't turn against the perpetrators but we
不过我们可以问问马斯克——我可以给
116:57
could ask you on musk like I can I can give you a very precise
你一份非常精确的火箭发动机制造说明
117:01
list of instructions of how you build a rocket engine M and even if you
,即使你有一个由50名经验丰富的工
117:06
have a team of 50 Engineers that of re experienc building it you're still going
程师组成的团队,在造出能用的东西之
117:11
to have to blow up a dozen of them before you get when that works
前,你还是得炸掉十几个。嗯,这和前
117:15
um and you know it's the same with
面说的道理是一样的。
117:18
uh you know chemical weapons or biow weapons or things like this it requires expertise
嗯,你知道化学武器或生物武器这类东西,它需要专业知识,在现实世界里,光靠语言模型是帮不了你的。它甚至需要我们在讨论的那种常识性专业知识——也就是如何把基于语言的指令,在物理世界中具体实现出来——这需要很多指令里没有的知识。对,没错,很多生物学家其实已经发帖回应过这些说法了,他们说,你知道真正做实验有多难吗?这可不是简单的事。对,这就又提到了H
117:23
you know in the in the real world that n is not going to help
ans Marik的观点。再聊聊Llama吧,Mark宣布Llama 3最终会发布,但我不知道具体发布日期。你首先最期待什么?毕竟Llama 2已经出来了,还有未来的Llama 3、4、5、6、10,Meta旗下开源模型的未来。嗯,有好几点吧。首先会有各种版本的Llama,都是对之前版本的改进,更大、更好、支持多模态之类的。然后未来的系统会具
117:28
you with and it requires even the common sense expertise that we've been talking about
备规划能力,真正理解世界是如何运作的。也许我们的进展是因为我们公开了研究成果。比如上周我们发表了Via的工作,这是从视频训练系统的第一步。下一步就是基于这种思路的世界模型,从视频中训练。DeepMind那边也有类似的工作,还有UC Berkeley也在做从视频构建世界模型。很多人都在做这个,我觉得很多好想法正在涌现。我猜这些系统会像JEP那样
117:34
which is how to take uh language-based instructions and materialize them in the physical world
,不会是生成式模型。未来会告诉我们答案。有个叫Danar Hafner的人,不是DeepMind的,他做过这类模型,学习表征然后用它们来规划或通过强化学习完成任务。Berkeley那边也有Peter Iil S leine和其他人的很多工作。我其实通过一些项目和我的NYU身份在合作,也通过Meta有合作,因为Berkeley的实验室某种程度上
117:39
requires a lot of knowledge that's not in the instructions yeah exactly a lot
和Meta有关联,和FAIR有关。所以我觉得这非常令人兴奋。我真的很兴奋,我很久没有对机器学习和AI的方向这么兴奋过了,大概从十年前FAIR刚开始的时候,再往前就是三十年前我们做卷积神经网络和神经网络早期的时候。所以我超级兴奋,因为我看到了一条通往潜在人类级别智能的道路,系统能够理解世界、记忆、规划、推理。确实有一些——
117:44
of biologists have posted on this actually in response to those things saying like do
不少生物学家其实已经针对这些观点发帖回应了,
117:48
you realize how hard it is to actually do the the lab work I you
说你们知道做实验室工作有多难吗?这可不是闹着玩
117:53
know this is not trivial yeah and that's Hans Marik comes comes to light once
的。对,这时候Hans Marik又冒出来了
117:58
again uh just the Linger on llama you know Mark announced that llama 3 is
。再说回Llama,Mark宣布Llama 3
118:03
coming out eventually I don't think there's a release date but what what are you
迟早会出,我觉得还没定发布日期,但首先Lla
118:08
most excited about first of all llama 2 that's already out
ma 2已经发布了,你最期待什么?
118:11
there and maybe the future llama 3 4 5 6 10 just the the future
很多生物学家其实已经针对那些事情发帖回应了,说你们知不知道
118:17
of the open source under meta well a number of things so uh there's going
真正做实验有多难?这可不是小事。对,Hans Marik
118:22
to be like various versions of of Lama that are uh you know improvements of
又一次被提出来了。再聊聊 Llama 吧,Mark 宣布
118:28
previous llamas bigger better multimodal things like that and then in future Generations systems that
Llama 3 最终会出来,我觉得还没有具体的发布日期,但
118:33
are capable of planning that really understand how the world Works um maybe
你最期待的是什么?首先 Llama 2 已经发布了。
118:38
are trained from video so they have some World model maybe you know capable of
它们是从视频中训练的,所以可能具备某种世界模型,能进行我之前提到的推理和规划。这大概要多久?那些朝这个方向的研究,什么时候会反馈到产品线上?比如L的产品?我不知道,没法告诉你。我们还需要经历几个关键突破才能达到那个目标,但你可以通过我
118:42
the type of reasoning and planning I was talking about earlier like how long is
们发布的研究来跟踪进展——毕竟我们公开研究成果。比如上周我们发布了Via项目,这是从视频训练系统的第一步;下一步就是基于这种理念的世界模型,从视频中训练。DeepMind那边也有类似工作,还有UC Berkeley也在做视频世界模型。很
118:47
that going to take like when is the research that is doing going in that
多人都在研究这个,我觉得很多好想法正在涌现。我打赌这些系统会像JEP一样,不会是生成式模型。未来会告诉我们答案。有个叫Danar Hafner的学者,不是DeepMind的,他研究这类模型——学习表征然后用它们做规划或通过强化学习完成任
118:51
direction going to sort of feed into the product line if you want of L
务。还有Berkeley的Peter、Iil、Sleine等很多人也在做。我其实在跟其中一些人合作,通过我NYU的身份参与一些项目,也通过Meta合作——因为Berkeley的实验室跟Meta有关联,属于FAIR。我觉得这非常令人兴奋。
118:56
I don't know I can't tell you and there's you know a few breakthroughs that
我对机器学习和AI的发展方向,上一次这么兴奋还是十年前FAIR刚起步的时候,再往前是三十年前我们研究卷积网络和神经网络早期的时候。我超级兴奋,因为我看到了一条通往人类级别智能的路径——系统能理解世界、记忆、规划、推理。现在有了一些可能奏
119:00
we have to basically uh go through before we can get there but you'll be
效的思路,我真的很激动。我喜欢的是,我们终于走上了正确的方向,也许在我脑子变成白浆之前,或者在我退休之前,就能成功。是啊,你也很兴奋吧?光是看那些GPU的数量,整个训练过程用这么多算力,退一步看地球和人类——我们建造了这些计算设备,训练
119:05
able to monitor
出这一个大脑,然后开源它,就像在孕育一个新生命。
119:06
our progress because we publish our research right so you know if last week we
我们的进步是因为我们公开研究,对吧?比如上
119:11
published the Via work which is sort of a first step towards Training Systems from
周我们发布了 Via 的工作,这是从视频训练
119:17
video um and then the next step is going to be World models based on
系统的第一步,下一步将是基于这种理念的世界
119:23
on kind of this type of idea training training from video there similar work at
模型——从视频中训练。DeepMind 那边
119:28
at Deep Mind also and um the taking
也有类似的工作,而且……
119:31
place people and also at UC brookley on uh World models from video a lot
还有未来的 Llama 3、4、5、6、10,就是
119:36
of people are working on this I think a lot of good ideas are coming
Meta 旗下开源模型的未来。嗯,有好几个方面吧。会
119:42
are appearing my bet is that those systems are going to be Jep alike they're
有各种版本的 Llama,都是之前版本的改进,更大、
119:47
not going to be gener generative models um and uh we'll see what the future
更好、多模态之类的。然后在未来的几代系统里,会有能够
119:52
will tell um there's really good work at uh um a gentleman
规划、真正理解世界运作方式的系统,也许吧。
119:56
called danar Hafner who is not Deep Mind who who's worked on kind of models
我们的进步在于我们会公开研究成果,对吧。比如上周我们发表了Via的工作,算是从视频训练系统的
120:01
of this type that learn representations and then use them for planning or learning tasks
第一步,下一步就是基于这种思路的世界模型,从视频训练。DeepMind那边也有类似的工作,还
120:07
by reinforcement running um and a lot of work at brookley by Peter iil S
有一位叫Danar Hafner的,他不是DeepMind的人,研究的是那种学习表征然后用它
120:12
leine bunch of other people of that type uh I'm collaborating with actually in the
们做规划或通过强化学习完成任务的模型。Brookley那边Peter Iil S和一堆人也在
120:18
context of some grants with my NYU hat um
做这类工作,我其实也在合作,以NYU的身份参与一些项目。
120:21
and then collaborations also through meta because the the lab at brookley is associated with
我们的进步是因为我们公开研究成果,对吧?比如上周我们发布了 Via 的工作,算是从视频训练系统的第一步。下一步就是基于这种思路的世界模型,从视频训练。De
120:26
meta in some way so with fair so I I think uh it's very exciting
epMind 那边也有类似的工作,还有 UC Berkeley 的人也在做视频世界模型。很多人都在研究这个,我觉得很多好想法正在涌现。我的猜测是,这些系统会
120:32
you know I I think I'm super excited about I I haven't been that excited
像 JEP 那样,不会是生成式模型。至于未来会怎样,我们等着看吧。有一位叫 Danar Hafner 的先生做了很好的工作,他不是 DeepMind 的,
120:37
about like the direction of machine learning and AI you know since uh you know
他研究这类模型,学习表征然后用它们来做规划或者通过强化学习完成任务。还有 Berkeley 的 Peter iil S leine 和一堆其他人也在做类似的
120:43
10 years ago when Fairway started
工作。我其实正在和他们合作,以我 NYU 的身份参与一些项目。
120:45
and before that um 30 years ago when we working on 35 on on com
嗯,30年前,我们在搞35、搞通信网络,还有神
120:51
Nets and and and the early days of neural net so um I'm super excited
经网络早期那会儿。所以我现在特别激动,因为我看
120:58
because I see a path towards potentially human level intelligence uh with you know systems
到了一条路,有可能通向人类级别的智能——就是那
121:04
that can understand the world remember plan reason um there there is some some
种能理解世界、能记忆、能规划、能推理的系统。
121:10
set of ideas to make progress there that might have a chance of working and
有一整套思路可以推动进展,而且很可能行得通
121:16
I'm really excited about this what I like is that you know it uh somewhere
,我对此非常兴奋。我喜欢的是,我们终于走上
121:22
we we get onto like a good direction and perhaps succeed before my brain turns
了一个正确的方向,也许能在我的大脑变成白酱之
121:28
to white sauce or or before I need to retire yeah yeah uh you're also
前,或者在我退休之前成功。是啊是啊,你也很
121:34
excited by are
兴奋对吧?
121:35
you is it beautiful to you just the amount of gpus involved sort of the
你觉得光是看那些GPU的数量、整个训练过程消耗
121:42
the the whole training process on this much compute it's just zooming out just looking
的算力,是不是很美?把视角拉远,看地球和人类一
121:49
at Earth and humans together have built these Computing devices and are able to train
起造出了这些计算设备,然后能训练出这一个大脑,
121:55
this one brain then then we then open source like giving birth to this
接着我们又把它开源——就像在孕育一个新生命。
122:02
open-source brain trained on this gigantic compute system there's just the details of how to
open-source brain trained o
122:06
train on that how to build the infrastructure and the the hardware the cooling all
n this gigantic compute syst
122:11
of this kind of stuff U or you just still the most of your excitement
em,剩下的就是训练细节、怎么搭基础设施、硬件、冷却这
122:16
is in the the theory aspect of it the meaning like the software well I
些乱七八糟的东西。或者你大部分兴奋点还是在理论层面,就是
122:21
used to be a hardware guy many years ago yes yes that's decades ago Hardware
软件那部分?我很多年前是搞硬件的,对对对,几十年前了。
122:26
has improved a little bit
硬件确实进步了一点。
122:28
changed a little bit yeah I mean certainly scale is necessary but not sufficient absolutely
变化不大吧。我是说,规模当然必
122:33
so we certainly need competition I mean we're still far in terms of compute power
要,但光有规模不够。绝对需要竞
122:38
uh from you know what we would need to match the compute power of the
争。在算力上,离匹配人脑还差得
122:43
human brain um you know this may occur in the next couple decades but um
远。可能未来几十年能做到,但还
122:48
but we're still some ways away and certainly in terms of power efficiency were really
有距离。而且能效方面,差得更远。
122:53
far um so a lot of progress to make in uh in in in hardware
硬件还有很多进步空间。现在进步不
122:59
and you know right now a lot of progress is is is not I mean
全是来自硅技术,很多来自架构创新
123:04
there's a bit coming from Silicon technology but a lot of it coming from architectural
,还有更高效的方式来实现那些流行
123:10
Innovation and quite a bit coming from uh like more efficient ways of you know
的架构,基本上就是Transfor
123:15
implementing the architectures that have become popular basically combination of Transformers
mer和CETs的组合。
123:19
and cets right and uh so you know there's still some ways to go until
所以离饱和还有一段路要走。我们
123:26
uh we're going to saturate we're going to have to come up with like new
得想出新的原理、新的制造技术、新
123:32
new principles new fabrication technology new uh basic components um perhaps you know based on
的基础元件,可能基于跟经典数字半
123:38
sort of different principles than those classical digital semas interesting so
导体不同的原理。有意思。
123:43
you think in order to build Ami M me we need we potentially might need
你觉得要做出AGI,可能也需要硬件创新?如果想
123:50
some Hardware Innovation too well if you want to make it um ubiquitous yeah certainly
让它普及,当然。因为得降低功耗。现在一个GPU是
123:57
because we're going to have to reduce the you know comput the power consumption a
半千瓦到一千瓦,人脑才25瓦。GPU远低于人脑的
124:04
GPU today right is half a kilowatt to a kilowatt human brain is about 25
算力,需要十万甚至百万倍才能匹配。所以差距巨大。
124:12
wats uh and the GPU is way below the power of human brain you need
你常说AGI不会很快到来,不是今年也不是未来几年,可能
124:18
you know something like a 100,000 or million to match it so uh so you
更远。你直觉上为什么这么想?首先,它不会是一个事件。科幻
124:24
know we're off by huge Factor here you often say that AGI is not coming
和好莱坞喜欢渲染那种“有人发现了AGI或人类级AI的秘密
124:29
soon meaning like not this year not the next few years potentially farther away what's
,然后一开机就有了AGI”的场景,但那不会发生。它不是事
124:35
your
件。
124:36
basic intuition behind that so first of all it's not going to be an event
会是渐进式的进步。我们会有能从视频
124:41
right the idea somehow which you know is popularized by science fiction and Hollywood that
中学习世界如何运作、学到好的世界表示
124:45
you know somehow somebody is going to discover the secret the secret to a gii
的系统吗?有。但要达到人类那样的规模
124:50
or human level AI or Ami whatever you want to call it and then you
和性能,还需要很长时间,不会一天就实
124:55
know turn on a machine and then we have a gii that's just not going
现。我们会有能拥有大量关联记忆、能记
124:59
to happen it's not going to be an
住东西的系统吗?有。
125:02
event it's going to be gradual progress are we going to have systems that can
这会是渐进式的进步。我们会不会有系统能通
125:07
learn from video how the world works and learn good World presentations yeah before we
过视频学习世界如何运作,并掌握好的世界表
125:11
get them to the scale and performance that we observe in humans it's going to
征?会有的。但要达到人类那样的规模和性能
125:16
take quite a while it's not going to happen in one day um uh are
,还需要相当长的时间,不可能一天就实现。
125:21
we going to get systems that can uh have large amount of associative memory so
嗯,我们会不会有系统能拥有大量联想记忆,
125:25
they can they can remember stuff yeah
从而记住东西?会的。
125:28
but same it's not going to happen tomorrow I mean there is some basic techniques
但同样,这不会明天就发生。我是说,有些基础技术还需要开发,我们已经有很
125:31
that need to be developed we have a lot of them but like you know
多了,但要让这些东西和整个系统协同工作,又是另一回事。我们怎么才能有一个
125:35
to get this to work together with full system is another story how we going
能推理和规划的系统,也许沿着我之前描述的那种目标驱动型AI架构的思路?对
125:39
to have system that can reason and plan perhaps along the lines of the objective
,但在我们让这一切正常运作之前,还需要一段时间。而且,在让所有这些组件协
125:43
driven AI architectures that I I described before yeah but like before we get this
同工作之后,还要在上面构建能学习的系统,比如分层规划、分层表征,能够像
125:47
to work you know properly it's going to take a while so and before we
人脑一样针对各种不同场景进行配置的系统。嗯,所有这些至少需要十年,很可能
125:50
get all those things to work together and then on top of this have systems
更久,因为还有很多我们现在没看到、没遇到的问题,我们不知道在这个框架内是
125:54
that can learn like
否有简单的解决方案。
125:55
hierarchical planning hierarchical representations systems that can be configured for a lot of different situation
层次化规划、层次化表征,以及能像人脑一
126:01
at hands the way the human brain can um you know all of this is
样针对多种不同情境进行配置的系统——所有
126:06
going to take you know at least a decade and probably much more because there
这些至少需要十年,很可能更久,因为我们
126:12
are a lot of problems that we're not seeing right now we have not encountered
现在还没看到很多问题,还没遇到,所以也不
126:17
and so we don't know if there is a easy solution within this framework
知道在这个框架内是否有简单的解决方案。
126:22
um so you know it's it's not just around the corner I mean I've I've
嗯,所以这绝不是近在眼前的事。我是说,过去12到15年里,我一直听到有人声称AGI就快来了,而且他们一直错得离谱。他们这么说的时候我就知道他
126:27
been hearing people for the last 12 15 years claiming that you know edgi is
们是错的。我问他们为什么。首先,从人工智能这个词诞生开始,就一直有一种永恒的乐观主义,这也许和其他技术不一样。这是莫拉维克悖论吗?它解释了为
126:31
just around the corner and being systematically wrong and I knew they were wrong when
什么人们对AGI如此乐观?我不认为这只是莫拉维克悖论。莫拉维克悖论是意识到世界并不像我们想的那么简单之后的结果。首先,智力不是一种可以用单一
126:36
they were saying it I call their why do you think people have been calling
标量、单一数字来衡量的线性东西。嗯,你能说人类比老鼠更聪明吗?在某些方面是的,但在很多领域,老鼠比人类更聪明,比如让它们在森林里生存的能力。
126:40
first of all I mean from the beginning of from the birth of the term
所以智商是一个非常有限的智力衡量标准。智力比智商测量的东西要大得多。嗯,智商大概能衡量人类的一些东西,因为人类在某种程度上形态相对统一,对吧?
126:45
artificial intelligence there has been a Eternal
但它只衡量一种能力,这种能力可能对某些测试有用,但对其他测试没用。
126:47
optimism that's perhaps unlike other Technologies is it a Maric Paradox is the explanation for
这种乐观情绪可能与其他技术不同。是Moravec悖论
126:53
why people are so optimistic about AGI I don't think it's just Marx Paradox Marx
在解释为什么人们对AGI如此乐观吗?我不认为只是Mor
126:59
Paradox is a consequence of realizing that the world is not as easy as we
avec悖论。Moravec悖论是意识到世界并不像我们
127:06
think so first of all um intelligence is not a linear thing that you can
想的那么简单之后的结果。首先,智能不是线性的东西,不
127:12
measure with a scalar
能用标量来衡量。
127:13
with a single number um you know can you say that humans are smarter than
嗯,但如果你谈论的是其他智能实体,它们觉得容易的基本事情和我们非常不同,那智商就毫
127:19
WR tongs in some ways yes but in some waysons are smarter than humans in
无意义了。所以智力是一系列技能的集合,以及高效获取新技能的能力。嗯,对。而一个特定
127:24
a lot of domains that allows them to survive in the forest for example so
智能实体所拥有或能快速学会的技能集合,和另一个实体的技能集合是不同的。因为这是一个多
127:30
IQ is a very limited measure of intelligence T intelligence is bigger than what IQ
维的东西,技能集是一个高维空间,你无法衡量,也无法比较两个东西哪个更聪明。它是多维
127:35
for example measures well
的,所以你要反驳这一点。
127:37
IQ can measure you know approximately something for humans MH but um because humans kind
IQ可以大致衡量人类的一些能力,嗯,因为人类
127:44
of you know come in relatively kind of uniform form right right uh but it
在某种程度上形态相对统一,对吧?但它只衡量某
127:51
only measures one type of uh ability that you know may be relevant for some
一种能力,这种能力可能对某些测试有用,但对其
127:58
test but not others
他测试没用。
128:00
and uh but then if you talking about other intelligent entities for which the you
但如果你谈论其他智能实体,它们觉得容易的基本事情和我们非常不同,那这个衡量就毫无意义了。所
128:07
know the the basic things that are easy to them is very different then it
以智能是一组技能的集合,以及高效获取新技能的能力,对吧?而一个特定智能实体所拥有或能快速学
128:15
doesn't mean anything so intelligence is a collection of skills and an ability to acquire
习的技能集合,与另一个实体的技能集合是不同的。因为这是一个多维的东西,技能集是高维空间,你
128:22
new skills efficiently mhm right and the collection of skills that
无法测量,也无法比较两个实体谁更智能——它是多维的,所以你会反驳。
128:27
an need intelligent particular intelligent entity possess or is capable of learning quickly is different
一个智能体是否需要具备某种特定的智能,或者能否快速学习
128:34
from the collection skills of another one and because it's a multi-dimensional thing the set
,这与另一个智能体的技能集合是不同的。因为这是一个多维度
128:40
of skills is high dimensional space you can't measure you can compare you cannot compare
的东西,技能集处在一个高维空间里,你无法测量,只能比较。
128:47
two things as to whether one is more intelligent than the other it's multi-dimensional so
你无法比较两个东西哪个更智能,因为它是多维度的,所以你会
128:54
you push back
反驳这一点。
128:55
against what are called AI doomers a lot uh can you explain their perspective and
针对所谓的“AI末日论者”,很多人呃,你能解释一下他们的观点以及你为什么觉得他们错了吗?好,AI末日论者想象各种灾难场景,比如AI如何逃脱或控制,基本上把我们全杀了,呃,而这基于一大堆假设,大部分是错的。第一个假设是,超级智能的出现会是一个事件,在某个时刻我们会找到秘密,然后打开一台超级智能机器,因为我们从没做过,它会接管世界并杀死我们所有人。这是错的,它不会是一个事件。我们会拥有像猫一样聪明的系统,具备人类级智能
129:01
why you think they're wrong okay so a I doomers imagine all kinds of catastrophe
的所有特征,但它们的智能水平可能像猫或鹦鹉之类的,呃,然后我们会逐步提升,让这些东西更智能。随着我们让它们更智能,我们也会加入一些护栏,学会如何设置护栏,让它们行为得当。而且我们不会只做一个,不会只有一个努力方向,会有很多人同时在做,其中一些人会成功造出可控、安全且有正确护栏的智能系统。如果某个系统出问题,我们可以用好系统去对抗那些 rogue 系统,呃,所以就像我的智能AI警察对抗你的 rogue AI,呃,所以不
129:07
scenarios of how AI could Escape or control and basically kill us all uh and
会像我们暴露在一个单一的 rogue AI 下被全杀,这根本不会发生。现在还有另一个谬误,就是认为因为系统智能,它就必然想接管一切,嗯,有几个论点让人害怕这个,我觉得也完全是错的,呃。其中一个就是,嗯,在自然界中,似乎更智能的物种往往支配其他物种,甚至消灭它们,有时是有意,有时只是无意,呃,所以就有一种想法说,如果AI系统比我们更智能,它们肯定会消灭我们,如果不是有意,只是因为不在乎我们。这很荒谬,原因有几个,嗯。
129:13
that relies on a whole bunch of assumptions that are mostly false so the first
第一个原因是,它们不会成为一个物种,不会成为与我们竞争的物种,不会有支配的欲望,因为支配欲望必须硬编码进智能系统,嗯,它在人类中是硬编码的,在狒狒、黑猩猩、狼中也是硬编码的,但不是在所有物种中。这种支配、服从或以其他方式获取地位的欲望,是特定于社会性物种的。非社会性物种,比如猩猩,就没有这个,对吧?它们几乎和我们一样聪明,对吧?而且人类没有显著动机把这个编码进AI系统,就算有人做了,也会有AI系统呃,惩罚它们,与它们
129:18
assumption is that
竞争。嗯,有各种动机让AI系统服从人类,对吧?我的意思是,这就是我们构建它们的方式。
129:20
the emergence of super intelligence is going to be an event that at some point
超级智能的出现不会是一个突然的事件,我们不会在某天突然破解秘密,然后打开一台超级智能的机器,因为它我们从来没做过,它就接管世界、杀死所有人——这种说法是错的。它不会是一个事件。我们会先有像猫一样聪明的系统,具备人类级别智能的所有特征,但它们的智能水平大概就像猫或者鹦鹉那样。然后我们会一步步往上走,让这些东
129:24
we're going to have we're going to figure out the secret and we'll turn on
西变得更聪明。在让它们更聪明的过程中,我们也会给它们装上一些guard rails,学着怎么加这些guard rails,让它们行为得当。而且我们不会只靠一个团队来做这件事,不会是单一的努力,会有很多不同的人在做,其中一些人会成功造出可控、安全、有正确guard rails的智能系统。如果某个系统出了问题,
129:28
a machine that is super intelligent and because we've never done it before it's going
我们还可以用好的系统去对抗那些Rogue的系统。所以就像是我有聪明的AI警察,来对付你的Rogue AI。所以我们不会暴露在单个Rogue AI面前,被它杀光所有人——这根本不会发生。还有一个谬论是,因为系统很智能,它就一定想要接管一切。有几个论点让人们害怕这个,但我认为它们完全是错的。其中一个论点是,在自
129:33
to take over the world and kill us all that is false it's not going
然界里,似乎更聪明的物种最终会支配其他物种,甚至有时会有意或无意地消灭其他物种。所以有人就会想,如果AI系统比我们更聪明,它们肯定会消灭我们,就算不是故意,也是因为它们不在乎我们。这完全是荒谬的,原因有好几个。第一个原因是,它们不会成为一个物种,不会成为和我们竞争的物种,它们不会有支配的欲望。因为支配的欲望
129:37
to be an event we're going to have systems that are like as smart as
必须被硬编码进一个智能系统里——它在人类身上是硬编码的,在狒狒、黑猩猩、狼身上也是硬编码的,但不是在所有物种里。这种支配、服从或以其他方式获取地位的欲望,是特定于社会性物种的。非社会性物种,比如猩猩,就没有这种欲望,而它们几乎和我们一样聪明。而且人类没有显著的动机去把这种欲望编码进AI系统。就算有人这么做了
129:42
a cat has all the have all the characteristics of you know human level intelligence
,也会有别的AI系统去惩罚它们、和它们竞争。实际上,有各种各样的动机让AI系统对人类保持顺从。没错,我的意思是,这就是我们建造它们的方式。用AI去伤害其他人——这话我好像在哪听过,不记得了。也许是在一本书里吧。但说到那本书,会不会也有 unintended consequences?当然会。所以这不是一个简
129:46
but their level of
单的问题。设计那些guard rails让系统行为得当,不是一件有银弹就能解决的事。
129:47
intelligence would be like a cat or a pirrot maybe or something um and then
再往前30年,我们在搞35年前的Com Net
129:51
we're going to work our way up to kind of make those things more intelligent
s和神经网络早期阶段,所以我特别兴奋,因为我看到
129:55
and as we make them more intelligent we're also going to put some guard rails
了一条通往人类级别智能的路,系统能理解世界、记忆
129:59
in them and learn how to kind of put some guard rails so they behave
、规划、推理。还有一些智能可能像猫或鹦鹉那样,
130:02
properly and we're not going to do this with just one it's not going to
然后我们逐步提升它们的智能,同时加一些护栏,学会
130:06
be one effort there's going to be lots of different people doing this and some
怎么加护栏让它们行为得当。这不是一次性的事,会有
130:10
of them are going to succeed at making intelligent systems that are uh controllable and
很多人各自尝试,有些人会成功造出可控且安全的智
130:14
safe and
能系统。
130:14
have the right guard rails and if some other goes wrog then we can use
智能可能像猫或者鹦鹉之类的,然后我们会逐步提升,让这些东西变得更聪明。随着它们变得
130:19
the the good ones to go against the Rogue ones uh so it's going to
更聪明,我们也会给它们加上一些护栏,学习如何设置这些护栏,让它们行为得当。而且我们不
130:23
be my you know smart AI police against your Rogue AI um so it's not
会只靠一个项目来做这件事,不会只有一个努力方向,会有很多不同的人在做这件事。其中一
130:27
going to be like you know we're going to be exposed to like a single
些人会成功制造出可控、安全且有正确护栏的智能系统。如果有些系统出了问题,我们就可以用
130:32
Rogue AI that's going to kill us all that's just not not happening now there
好的系统去对抗那些 rogue 系统。所以,这就像是我这边的聪明 AI 警察去对付
130:36
is another fallacy which is the fact that because the system is intelligent it necessarily
你那边 rogue AI。因此,我们不会暴露在一个单一的 rogue AI 面前,被
130:40
wants to take over MH um
它全部消灭,这种情况根本不会发生。
130:42
and there is several arguments that make people scare of this which I think are
在那之前,30年前,我们还在研究 Conv
130:48
completely false uh as well so one of them is um you know in nature
Nets 和神经网络的早期阶段。所以我非常兴
130:55
it seems to be that the more intelligent species otherwi that end up dominating the
奋,因为我看到了一条通往人类级别智能的道路
131:01
other and uh and even you know extinguishing the others uh sometimes by Design sometimes
,系统可以理解世界、记忆、规划、推理。还有一
131:07
just by
些……
131:08
mistake and and so you know there is sort of uh Thinking by which you
现在还有另一个谬论,就是认为因为系统是智能的,它就必然想要接管一切。嗯,有一些论点
131:13
say well if AI systems are more intelligent than us surely they're going to eliminate
让人们对此感到恐惧,但我认为这些论点完全是错误的。其中一个论点是,在自然界中,似乎更
131:18
us if not by Design simply because they don't care about us and that's just
聪明的物种最终会支配其他物种,甚至有时是有意地、有时是无意地消灭其他物种。所以就有一
131:23
Preposterous for for a number of reasons um first reason is they're not going to
种想法认为,如果 AI 系统比我们更聪明,它们肯定会消灭我们,如果不是有意为之,那
131:28
be a species they're not going to be a species that competes with us they're
也只是因为它们不在乎我们。这完全是荒谬的,原因有很多。首先,它们不会成为一个物种,不
131:34
not going to have the desire to dominate
会成为与我们竞争的物种,它们不会有支配的欲望。
131:36
because the desire to dominate is something that has to be hardwired into an intelligent
有几个让人害怕的论点,我觉得完全是错的。其中一个是在自然界里,似乎越聪
131:44
system uh it is hardwired in humans it is hardwired in baboons in chimpanzees in
明的物种最终会支配其他物种,甚至让它们灭绝,有时是故意的,有时只是无意
131:51
wolves not in a wrong Tes the species in which this desire to dominate or
。因为支配欲必须被硬编码进智能系统,人类有,狒狒有,黑猩猩有,狼也有,但
131:58
submit or or attain status in other ways is is specific to social species
某些物种没有。这种支配或服从或追求地位的欲望,是社会性物种特有的。
132:05
non-social species like our tongs don't have it right and they are as smart as
因为支配的欲望必须被硬编码进一个智能系统里。这种欲望在人类身上是硬编码的,在狒狒、黑猩猩
132:10
we are almost right and to you there's not significant incentive for humans to encode
、狼身上也是硬编码的,但不是在所有物种身上。这种支配、服从或以其他方式获取地位的欲望,是特
132:15
that into the AI systems and to the degree they do there'll be AIS that
定于社会性物种的。非社会性物种,比如猩猩,就没有这种欲望,而它们几乎和我们一样聪明。而且
132:19
um sort of punish them for it I'll compete them over well there's all kinds
,人类没有显著的动机去把这种欲望编码进 AI 系统。即使有人这么做了,也会有其他 AI 系
132:24
of incentive to make AI system submissive to humans right right I mean this is
统为此惩罚它们,与它们竞争。实际上,有各种各样的动机让 AI 系统对人类保持顺从,对吧?我
132:29
the way we're going to build
的意思是,这就是我们构建它们的方式。
132:31
them right um and so so then people say oh but look at llms LMS
嗯,然后呢,人们就会说,你看LLM,LLM是不可控的,他们说得对,LLM确实不可控。但目标驱动型AI,也就是通过优化某个目标来得出答案的系统,它们必须优化这个目标,而这个目标可以包含护栏。一个护栏是“服从人类”,另一个护栏是“如果会伤害其他人类,就不要服从人类”。这话我好像在哪听过,记不清了。可能是在一本书里吧。嗯,说到那本书,这一切会不会也有意想不到的后果?当然会。所以这不是一个简单的问题。我是说,设
132:37
are not controllable and they're right LMS are not controllable but objective driven AI so
计那些护栏,让系统行为得当,这可不是一个简单的问题,不是什么有银弹、有数学证明能保证系统安全的事。这会是一个非常渐进、迭代的设计过程,我们以某种方式设置这些护栏,让系统行为得当。有时候它们会做出一些意料之外的事,因为护栏没设对,然后我们就去修正,让它们做对。那种“我们一点都不能搞错,因为搞错一点点我们全都会死”的想法,太荒谬了。我们就是一步步来。我经常用的一个类比是涡轮喷气发动机的设计。我们是怎么让涡轮
132:43
systems that derive their Answers by optimization of an objective means they have to optimize
喷气发动机变得如此不可思议地可靠的?那可是极其复杂的硬件,在极高温度下运行,有时候一次就是20个小时。然后我们可以坐着一架双引擎喷气客机,以接近音速的速度飞半个地球,这多不可思议?简直难以置信,对吧。我们做到这一点,是因为发明了一个让涡轮喷气发动机安全的一般性原则吗?不是。我们花了数十年时间,才慢慢微调这些系统的设计,让它们变得安全。那通用电气或者斯奈克玛内部,有没有一个专门负责涡轮喷气发动机安全的独立小
132:49
this objective and that objective can include guard rails one guardrail is uh obey humans
组?没有。设计本身就是关于安全的,因为更好的涡轮喷气发动机,同时也是更安全的涡轮喷气发动机,更可靠的那种。AI也是一样。你需要专门做点什么来让AI安全吗?不需要。你需要做出更好的AI系统,它们自然就会安全,因为它们的设计初衷就是更有用、更可控。那么,想象一个AI系统,它能够极其有说服力,能说服你任何事情。我至少能想象出这样一个系统,而且我能看到这样的系统像武器一样,因为它能控制人的思想。我们人类很容易上
132:54
another guardrail is don't obey humans if it's
当,我们愿意相信某些东西。你可以有一个控制它的AI系统,你也能看到政府把它当作武器来用。所以,如果你想象这样一个系统,你觉得它和核武器有什么相似之处吗?没有。那为什么,为什么这种技术不一样?你是说它会是一个渐进的过程。
132:57
hurting other humans with I've heard that before somewhere I don't remember yes maybe in
对于这类系统来说,这几乎是不可能的,因为它
133:02
a book yeah uh but speaking of that book what is could there be unintended
们不可避免地会形成政治观点,还有对各种事情的
133:07
consequences also from all of this no of course uh so this is not a
看法——可能涉及政治,也可能不涉及,但人们会
133:12
simple problem right I mean uh designing those guard rail so that the system behaves
因此产生分歧,比如道德问题、宗教相关的问题,
133:16
properly is not going to be a a simple uh issue that for which there
或者不同社群对文化议题的看法本来就不一致。嗯
133:21
is a silver bullet for which you have a
,所以其实只有相对少数的——
133:24
mathematical proof that the system can be safe it's going to be very Progressive iterative
数学上证明系统是安全的,这将会是一个非常渐进式的迭代设计过程,我们以这样的方
133:29
design system where we put those guard rails in such a way that the system
式设置那些guard rails,让系统行为正常。有时候它们会做出一些意料之外
133:33
behave properly and sometimes they're going to do something that you know was unexpected because
的事,因为guard rail没设对,然后我们会纠正它们,让它们做对。那种“
133:38
the guardare wasn't right and we're going to correct them so that they do it
我们稍微搞错一点就会全完蛋”的想法,简直荒谬。我们就是一步一步来,我经常用的一
133:42
right uh the idea somehow that we can't get it slightly wrong because if we
个类比是涡轮喷气发动机的设计。我们是怎么让涡轮喷气发动机变得如此不可思议地可靠
133:47
get it slightly wrong we all die is is ridiculous um we we're just going
的?我的意思是,那些东西是极其复杂的硬件,在极高的温度下运行,有时候一次就是
133:51
to go
20个小时。
133:52
progressively and it's it's just going to be the the analogy I've used many times
逐步地,我经常用的一个类比是涡轮喷气发
133:58
is um is uh turbojet design um how how did we figure out how to
动机的设计——我们是怎么让涡轮喷气发动机
134:04
make turbojet so unbelievably reliable right uh I mean those are like you know incredibly
变得如此不可思议地可靠的?那些东西可是极
134:11
complex uh pieces of Hardware that run at really high temperatures for you know 20
其复杂的硬件,在极高温度下运行,有时一次
134:17
20 hours at a time sometimes
就是20个小时。
134:20
and we can you know fly halfway around the world with a on a two
然后我们还能靠双引擎喷气客机以接近音速飞半个地球,这多厉害?简直难以置信,对吧?我们做到这一点,是因
134:25
engine uh jetliner at near the speed of sound like how incredible is this it's
为发明了一个让涡轮喷气发动机安全的通用原理吗?不是的,我们花了几十年才逐步fine-tune这些系统
134:31
just unbelievable right and did we do this because we invented like a general principle
的设计,让它们变得安全。通用电气或者斯奈克玛内部有没有一个专门负责涡轮喷气发动机安全的独立小组?没有
134:36
of how to make Turbo Jet safe no we it took decades to kind of
,设计本身就是围绕安全的,因为更好的涡轮喷气发动机同时也是更安全的涡轮喷气发动机,更可靠。AI也是一样
134:42
fine-tune the design of those
,你需要专门的规定来让AI安全吗?
134:43
systems so that they they were safe is there a separate uh group Within General
确保系统安全,是不是通用电气或者斯奈克玛之类的公
134:49
Electric or snma or whatever that is specialized in turo jet safety no it's the
司里,会有一个专门负责涡轮喷气发动机安全的团队?不
134:55
design is all about safety because a better Turbo Jet is also a safer Turbo
,设计本身就是关于安全的,因为更好的涡轮喷气发动机
135:00
Jet so um a more reliable one it's the same for AI like do you
也是更安全的涡轮喷气发动机,更可靠的那种。AI也是
135:06
do you need you know specific Provisions to make AI safe
一样,你需要专门的规定来让AI安全吗?
135:10
no you need to make better AI systems and they will be safe because they
不,你需要做出更好的AI系统,它们自然就会安全,因为它们被设计得更有用、更可控。那么想象一个AI系统
135:16
are designed to be more useful uh and more controllable so let's imagine a system
,它极其有说服力,能说服你任何事情——我至少能想象出这样的系统,而且我能看到这样的系统像武器一样,因
135:22
AI system that's able to be incredibly convincing and can convince you of anything I
为它能控制人的思想。我们人类相当容易轻信,我们愿意相信某些东西。你可以有一个控制它的AI系统,你也能看
135:28
I can at least imagine such a system and I can see such a system
到政府把它当作武器。所以你觉得,如果想象这样一个系统,它和核武器有什么相似之处吗?没有。那为什么这种
135:34
be weapon-like because it can control
技术不一样呢?你是说会有一个渐进的过程……
135:37
people's minds we're pretty gullible we we want to believe a thing you can have
为了让这些系统安全,通用电气或SNECMA内部有
135:43
an A system that controls it and you could see governments using that as a
没有专门负责涡轮喷气发动机安全的团队?不,设计本
135:48
weapon so do you think if you imagine such a system there's any parallel to
身就是为了安全,因为更好的涡轮喷气发动机同时也是更
135:54
something like nuclear weapons no so is why why why is that technology different so
安全的涡轮喷气发动机,更可靠的那种。AI也一样,
136:00
you're saying there's going to be gradual
你需要专门的安全措施吗?
136:03
development yeah there's going to be I mean it might be rapid but they'll be
发展嘛,肯定会有,可能很快,但会是迭代式的,然后我们也能做出应对,等等。所以那个由Vladimir Putin或者他的手下设计的AI系统,它会试图跟每个美国人对话,说服他们投票给让Putin满意的人,或者挑拨人们互相争斗——就像他们一直试图做的那样。但它们不会直接跟你对话,它们会跟你的AI助手对话,嗯哼,而你的AI助手跟它们的一样聪明。对,没错。因为就像我说的,未来你与数字世界的每一次互动,都会通过你的
136:08
iterative and then we'll be able to kind of respond and and so on so
AI助手来中介。所以你第一反应会是:这是诈骗吗?这东西在跟我说真话吗?它甚至根本到不了你这里,因为它只会跟你的AI助手对话,而你的AI助手根本不会——它就像个垃圾邮件过滤器,对吧?你连垃圾邮件都看不到,它自动被放进一个你永远不会打开的文件夹里。同样的事情也会发生:那个试图说服你什么的AI系统,会跟你的助手对话,而你的助手至少跟它一样聪明,然后会说:这是垃圾信息。它甚至不会引起你的注意。所以对你来说,任何一
136:14
that AI system designed by Vladimir Putin or whatever or his uh minions uh you
个AI系统想要领先一大步,达到能说服其他AI系统的程度,都非常困难。所以永远会有这种竞赛,没人能遥遥领先。这就是世界的历史——世界的历史就是,每当某个地方出现一种进步,就会有相应的反制措施,就像一场猫鼠游戏。嗯,大部分情况是这样,但这也是为什么核武器如此有趣——因为那是一种威力巨大的武器,谁先得到它至关重要。你可以想象,如果希特勒、斯大林、毛泽东先得到了核武器,那对世界的影响会跟美国先得到完全不同。但说
136:19
know is going to be uh like talking to trying to talk to every American
到核武器,你不会想象AI领域会有一个突破性发现,然后像曼哈顿计划那样全力以赴。不,就像我说的,这不会是一个事件,而是持续的进步。每当一个突破出现,它都会非常迅速地广泛传播。是的,可能首先在行业内。我的意思是,这不是一个政府或军事组织特别有创新性的领域,事实上它们远远落后。所以这会来自行业,而这种信息传播得极快。过去几年我们已经看到了这一点,对吧?比如你有一个新东西,就拿AlphaGo来说,它在三个月内就被
136:25
to uh convince
复现了,甚至没有特别详细的信息。对,这是一个不擅长保密的行业。但即使有保密措施——
136:26
them to vote for you know whoever whoever pleases Putin sure uh or whatever or
让他们投票给你想让他们投的人,比如讨好普京之类的,或者挑拨人们互相敌对——就像他们一直在试图做的那样。他
136:31
or you know or R people up against each other um as they've been trying
们不会直接跟你对话,他们会跟你的AI助手对话,而你的AI助手会和他们的AI一样聪明。没错,因为正如我说的,
136:37
to do they're not going to be talking to you they're going to be talking
未来你与数字世界的每一次互动,都将由你的AI助手来中介。所以你要问的第一件事就是:这是骗局吗?这东西跟我说
136:42
to your AI assistant mhm which is going to be as smart as theirs MH
的是真话吗?它甚至根本到不了你面前,因为它只会跟你的AI助手对话,而你的AI助手甚至不会——它就像个垃圾
136:48
right right that AI because as I said in the future every
邮件过滤器,对吧?你根本看不到那些垃圾邮件,它们自动被放进一个你永远不会看的文件夹里。
136:52
single one of your interaction with the digital world will be mediated by your AI
你每一次与数字世界的互动,都会通过你的AI
136:56
assistant so the first thing you're going to ask is is this a scam like
助手来中介。所以你第一件会问的事就是:这是
137:00
is this thing like telling me the truth like it's not even going to be
不是骗人的?这东西说的是真话吗?它甚至根本不
137:04
able to get to you because it's only going to talk to your AI assistant
会传到你这儿,因为它只会跟你的AI助手对话
137:08
and your AI assistant is not not even going to it's going to be like
,而你的AI助手就像个垃圾邮件过滤器一样,
137:12
a spam filter right you're not even seeing the email the spam email right it's
对吧?你连那封垃圾邮件都看不到,它自动被放进
137:16
automatically put in a folder that you never see um it's
一个你永远不会打开的文件夹里了。
137:19
going to be same thing that AI system that tries to convince you of something
就跟那个情况一样——一个AI系统想说服你什么,它得先跟另
137:24
is going to be talking to a assistant which is going to be at least
一个至少跟它一样聪明的助手对话,然后那个助手会说“这是垃圾
137:28
as smart as it and is going to say this is spam you know U
信息”,甚至都不会让你注意到。所以对你来说,任何一个AI系
137:32
it's not even going to bring it to your attention so to you it's very
统想领先到能说服其他AI系统,都极其困难。所以这就像一场
137:37
difficult for any one AI system to take such a big leap ahead to where
永远没人能遥遥领先的竞赛。这就是世界的历史,世界的历史就是
137:41
it can convince even the other AI systems so like it there's always going to
,每当某个地方出现一种进步,就会有相应的对策,这就是一场猫
137:46
be this
鼠游戏。
137:46
kind of race where nobody's way ahead that's the history of the world history of
人类的大脑很容易轻信,我们愿意相信某些东西。
137:51
the world is you know whenever there is a prog at some someplace there is
你可以有一个AI系统来控制它,政府可能会把它当
137:57
a countermeasure and and you know it's a it's a Katan mous game well this
作武器。所以你觉得,如果想象这样一个系统,它
138:02
is why mostly yes but this is why nuclear weapons are so interesting because that
和核武器有什么相似之处吗?不,为什么那个技术不
138:07
was such a powerful weapon that it mattered who got it
一样?你是说这会是一个渐进的过程?
138:11
first that you know you could imagine Hitler Stalin ma getting the weapon first and
嗯,大部分情况确实如此,但这也正是核武器那么有意思的原因。因为那是一种威力巨大的武
138:17
that that having a different kind of impact on the world than than the United
器,谁先拿到它至关重要。你可以想象,如果希特勒、斯大林先拿到这种武器,对世界的影响
138:24
States getting the weapon first but to you nuclear weapons is is like you you
会跟美国先拿到完全不同。但核武器这种东西,你不会觉得会有一个突破性发现,然后像曼哈
138:31
don't imagine a uh breakthrough
顿计划那样全力以赴去搞AI。
138:33
Discovery and then Manhattan Project like effort for AI no as I said it's not
你与数字世界的每一次互动,都将由你的AI助手来中介。所以你要问
138:39
going to be an event it's going to be you know continuous progress and and
的第一件事就是:这是不是个骗局?这东西跟我说的是真话吗?它甚至根
138:45
whenever you know one breakthrough occurs it's going to be widely disseminated really quickly yeah
本到不了你这里,因为它只会跟你的AI助手对话,而你的AI助手就
138:51
probably first within industry I mean this is not a domain where you know government
像个垃圾邮件过滤器一样,对吧?你根本看不到那些垃圾邮件,它们会自
138:57
or military organizations are particularly Innovative and they're in
动被放进一个你永远不会打开的文件夹里。
139:01
fact way behind um and so this is going to come from industry and and
不,就像我说的,这不会是一个单一事件,而是一个持续进步的过程。每当一个突破
139:06
this kind of information disseminates extremely quickly we've seen this over the last few years
出现,它很快就会广泛传播。对,很可能先在行业内传播。我的意思是,这不是一个政
139:10
right where you have a new like you know even take alphao this was reproduced
府或军事组织特别有创新性的领域,事实上它们远远落后。所以这东西会来自行业,而
139:15
within three months even without like particularly detailed information right yeah this is an industry
这种信息传播得极快。过去几年我们已经看到了,对吧?就拿AlphaGo来说,即
139:20
that's not good at secrecy no but even even if there is just the
使没有特别详细的信息,三个月内就被复现了。对,这是一个不擅长保密的行业。
139:24
fact that you know that something is possible yeah uh makes you like realize that
知道某件事是可行的,就会让你意识到值得花时间去做。你可能是第二个做这件事的人,但你知道你会去做。同样,对于所有的创新,比如self-supervised learning、transformer、decoder-only架构、LLM,你不需要确切知道它们的工作原理,就能知道它们是可行的,因为已
139:29
it's worth investing the time to actually do it you you may be the second
经部署了,然后被复现,接着那些公司的人跳槽,从一家公司到另一家,信息就传播开了。美国科技产业,尤其是硅谷的成功,恰恰就在于信息流通得非常非常快,传播得很快,所以整个地区因为这种信息流通而领先。所以,也许我想在AI末日论者的心理上多停留一下。你用经典的Yann LeCun方式举了一个很好的例子:当
139:34
person to do it but you know you'll you'll do it uh and you know
一项新技术出现时,你说工程师说“我发明了这个新东西,我叫它圆珠笔”,然后推特圈回应“天哪,人们可以用它写可怕的东西,比如 misinformation、propaganda、仇恨言论,现在禁止它”。然后写作末日论者登场,类似于AI末日论者,想象如果每个人都能拿到圆珠笔,这可能会摧毁社会,应该立法
139:39
same for you know all the Innovations you know self supervisor Transformers decoder only architectures
禁止用圆珠笔写仇恨言论,现在就要监管圆珠笔。接着铅笔行业大亨说,圆珠笔非常危险,不像铅笔写字可以擦掉,圆珠笔写下的东西永远留存,政府应该要求笔制造商获得许可证。这似乎确实是人类面对新技术时心理的一部分。那么,关于这一点,你有什么深刻的见解?嗯,人们对新技术及其对社会的影响有一种自然的恐惧,人们
139:44
llms I mean those things you don't need to know exactly the details of how
对于他们熟悉的世界受到重大变革的威胁有一种本能的反应,这些变革要么是文化现象,要么是技术革命。他们担心自己的文化、自己的工作、自己孩子的未来,以及自己的生活方式。所以任何变化都会让人害怕。纵观历史,任何技术革命或文化现象总是伴随着媒体上的某些群体或反应,基本上把当前社会的所有问题都归咎于那个特定
139:50
they work to know
的变化。比如,电曾经被认为会杀死所有人,火车曾经被认为是一件可怕的事情,因为……
139:51
that you know it's possible um because it's deployed and then it's getting reproduced and
你与数字世界的每一次互动都将由你的AI助手来中介。所以你要问的
139:57
then you know people who work for those companies move they go from one company
第一件事就是:这是不是个骗局?它是不是在跟我说真话?它甚至根本
140:03
to another and you know the information disseminates what makes the success of the the
不会传到你这儿,因为它只会跟你的AI助手对话,而你的AI助手就像
140:09
US tech industry and Silicon Valley in particular is exactly that is because information circulates
个垃圾邮件过滤器一样。你连垃圾邮件都看不到,对吧?它自动被放进
140:15
really really quickly and this you know
一个你永远不会打开的文件夹里。
140:18
disseminates very quickly and so you know the the whole region sort of is ahead
但即使有保密措施,因为这东西被部署了,然后被复现,
140:23
because of that circulation of information so maybe I just to linger on the psychology
再加上那些公司的人跳槽,从一家公司跑到另一家,信息就
140:29
of AI doomers you give uh in the classic Yan laon way a pretty good
传播开了。美国科技产业,尤其是硅谷的成功,恰恰就在于
140:35
example of just when a a new technology comes to be you say uh engineer
信息流通得非常非常快,而且传播得极快。所以整个地区因
140:41
says I invented this new thing
为这种信息流通而领先。
140:43
I call it a ballpen and then the Twitter sphere responds OMG people could write
另外,如果我们真的想要,嗯,观点的多样性,嗯,AI系统——你知道,在未来我
140:49
horrible things with it like misinformation propaganda Hast speech ban it now then writing doomers
们都会通过AI系统互动——我们需要这些系统是多样化的,以保护,嗯,思想的多
140:55
come in akin to the AI doomers imagine if everyone can get a ballpen this
样性,还有信仰、政治观点,以及各种东西,以及保护……嗯,哲学、理性主义,嗯,
141:02
could destroy Society there should be a law against using ballpen to write hate speech
逃离宗教教条,嗯,民主、科学。而且当然,没有这些,就不会有美国独立战争、法
141:08
regulate ballpens now and then
国大革命,我们还会活在……
141:10
the pencil industry Mogul says yeah ballpens are very dangerous unlike pencil writing which is
铅笔业大亨说,是啊,圆珠笔很危险,不像铅笔写字可以擦掉,圆珠笔写上去就永远留着,政府应该要求生产笔的人拿执照。我的意思是,这看起来确实是人类心理的一部分,当面对新技术的时候。那么,关于这个,你能谈谈什么深刻的见解吗?嗯,人们对新技术以及它可能对社会产生的影响有一种自然的恐惧,人们有一种本能的反应,就是他们熟悉的世界受到重大变革的威胁,这些变革要么是文化现象,要么是技术革命,他们担心自己的文化、自己的工作、自己
141:18
erasable ballpen writing stays forever government should require a license for a pen manufacturer I
孩子的未来,还有自己的生活方式,对吧。所以任何变化都会让人害怕,你可以在历史上看到,任何技术革命或文化现象总是伴随着媒体上的群体或反应,基本上把当前社会的所有问题都归咎于那个特定的变化。比如电曾经被认为会杀死所有人,火车曾经被说成是可怕的东西,还有爵士乐或漫画书被指责导致失业,或者年轻人不想工作之类的,对吧。这种情况已经存在了几个世纪,都是些下意识的反应。问题在于,我们是拥抱变化还是抗拒它,以及真正的危险是什么
141:26
mean this does seem to be part of um human psychology when when it comes
,而不是想象出来的危险。所以人们担心,我觉得有一件事他们担心的是大科技,我们一直在反复讨论但值得再提一下,他们担心AI会变得多强大,以及它落入一个中央集权或少数几个控制中心的手中。这就是对大科技的怀疑,这些公司可以赚大钱并控制这项技术,从而利用和欺负社会上的小人物。嗯,这正是我们需要开源平台的原因。对,我只是想更加强调这一点。那么让我问你,就像我说的,你在网上确实有点风趣,yos shabbach发了条推文,
141:33
up against new
你看到后笑了,提到H 9000,引用说“我理解你的论点,也完全明白你的沮丧,但是否”
141:35
technology so what what deep insights can you speak to about this well there is
技术方面,那你能谈谈什么深刻的见解?嗯,人们对新技术以及它可能对社会产生的影响有一种自然的恐惧,人们对自己熟悉的世界受到重大变革的威胁会有一种本能的反应,这些变革要么是文化现象,要么是技术革命,他们担心自己的文化、工作、孩子的未来,还有他们的生活方式,对吧。所以任何变化都会让人害怕,你会在历史上看到,任何技术革命或文化现象总是伴随着媒体上的某些群体
141:42
a a natural fear of uh new technology and the impact it can have in
或反应,基本上把当时社会的所有问题都归咎于那个特定的变化,对吧。比如电力一度被认为会害死所有人,火车也被说成是可怕的东西,还有爵士乐或漫画书被指责导致失业,或者年轻人不想工作之类的,这种情况已经存在了几个世纪,就是那种膝跳反应。问题是,我们是拥抱变化还是抗拒它,真正的危险是什么,而不是那些想象出来的危险。所以人们担心,我觉得他们担心的一点是,关于大
141:49
society and people have kind of instinctive reaction to um you know the world they
科技公司我们一直在反复讨论但值得再提的,就是他们担心AI会有多强大,以及它掌握在少数集中权力的手中,比如只有一小撮中央控制者。这就是对大科技公司的怀疑,这些公司可以赚大钱并控制这项技术,从而利用和欺负社会中的小人物。嗯,这正是我们需要开源平台的原因。是的,我只是想更加强调这一点。那么让我问你,就像我说的,你在网上确实有点风趣,yos shabbach
141:56
know being threatened by Major Transformations um that are either
发了一条推文你笑了,提到了H 9000,引用说“我欣赏你的论点,也完全理解你的沮丧,但……”这是你能评论一下的吗?比如在大公司工作,如何避免过度恐惧,或者说谨慎反而造成伤害。嗯,再说一次,答案就是开源平台,让各种各样的人都能参与进来。
142:01
cultural phenomena or technological um revolutions and they fear for their culture they feel for
发现,然后像曼哈顿计划那样全力搞AI?不,就像我说的,这
142:07
their job they feel for they fear for their you know the future of their
不会是一个单一事件,而是持续的进步。每当有突破发生,它都
142:12
children um and uh their way of life right so so any change um is
会迅速传播开来。是的,可能先在行业内传播。我的意思是,这
142:18
feared and and you see this you know long history like any technological Revolution or
不是一个政府或军事组织特别有创新精神的领域,它们往往……
142:24
cultural phenomenon was always accompanied by uh you know groups or reaction in the media
我把它叫做圆珠笔,然后Twitter上的人就炸了,说“天哪,有人能用它写可怕的东
142:31
uh that that basically attributed the all the problems the current problems of society to
西,比如 misinformation、propaganda、仇恨言论,赶紧禁止
142:38
that particular change right electricity was going to kill everyone at some point you know
它”。然后那些“写作末日论者”就出来了,跟AI末日论者一样,想象如果每个人都能拿
142:45
you uh the train was going to be a horrible thing because you know you
到一支圆珠笔,这能摧毁社会,应该立法禁止用圆珠笔写仇恨言论,现在就要监管圆珠笔。
142:52
can't breathe past 50 kilm an hour um and so there's a wonderful website called
喘口气都喘不过,时速超过50公里就不行了。有个很棒的网站叫pessimist archive,里面全是那些报纸剪报,记录了人们想象中会因为技术革新或者文化现象而出现的各种可怕事情。你知道,有很多精彩的例子,比如爵士乐或者漫画书被指责导致失业,或者年轻人不想工作了之类的。这种现象已经存在了几个世纪,都是些下意识的反应。问题在于,我们是拥抱变化,还是抗拒它?真正的危险是什么,而不是那些想象出来的危险。所以人们担心——我觉
142:59
a pessimist archive right which has all those newspaper clips of all the horrible things
得他们担心big Tech的一个问题,我们一直在反复讨论,但值得再提一次——他们担心AI会变得多强大,担心它落入少数几个中央集权势力的手中。这就是对big Tech的怀疑:这些公司可以赚大钱,控制这项技术,然后借此占便宜、欺负社会上的小人物。这正是我们需要open source平台的原因。没错,我就是想反复强调这一点。那么让我问你,就像我说的,你在网上偶尔会有点风趣。Yos Shabbach发了条推文,你看了都笑了,引
143:06
people imagine would would arrive because of uh either technological uh Innovation or uh a
用H 9000的话:“我理解你的论点,也完全明白你的挫败感,但pod bay doors是该开还是该关,这是一个复杂而微妙的问题。”你现在领导着Meta AI,你知道,这真的让我担心:我们的AI霸主会用这种企业腔调跟我们说话。而你用你的行事方式在抵制这一点。你能谈谈在大公司工作,如何避免过度谨慎反而造成伤害吗?是的,再次强调,答案还是open source平台,让各种各样的人都能构建代表全球文化、观点、语言和价值观多样
143:13
cultural phenomenon um you
性的AI助手,这样你就不会被单一的AI实体束缚在某种特定的思维方式里。所以我认为这对社会来说是一个非常重要的问题。
143:15
know the this is wonderful examples of uh uh you know jazz or comic books
我管它叫圆珠笔,然后推特上就炸了:“天哪,人们可以用它写可怕的东西!比如 mis
143:22
being blamed for uh unemployment or or you know young people not wanting to work
information、propaganda、仇恨言论,赶紧禁掉!”然后那些“写作
143:28
anymore and things like that right and and that has existed for for centuries um
末日论者”就来了,就像AI末日论者一样:“想象一下,如果每个人都能拿到一支圆珠笔
143:35
and it's you know knee-jerk reactions um the question is you know do
,这可能会摧毁社会!应该立法禁止用圆珠笔写仇恨言论!现在就监管圆珠笔!”
143:41
we Embrace change uh or do we resist it and what are the real dangers
在那之前,30年前,我们在搞35、搞Com N
143:47
as opposed to the imagined uh imagined ones so people worry about I think one
ets,还有神经网络的早期阶段。嗯,我超级兴奋,
143:54
thing they worry about with big Tech something we've been talking about over and over
因为我看到了一条通往人类级别智能的路,用那些能理
144:01
but I think worth mentioning again they worry about how powerful AI will be and
解世界、能记忆、能规划、能推理的系统。嗯,还有
144:08
they worry
一些——
144:09
about it being in the hands of one centralized power of just a handful of
关于它掌握在少数 centralized power 手中,就
144:14
central control and so that's the skepticism with big Tech you can make these companies
是那么一小撮 central control,所以这就是对 b
144:20
can make a huge amount of money and control this technology and by so doing
ig tech 的怀疑——你可以让这些公司赚大钱,控制这项技术,
144:26
you know take advantage uh abuse the little guy in society well that's exactly why
然后借此欺负社会上的小人物。嗯,这正是我们需要 open so
144:32
we need open source
urce 的原因。
144:34
platforms yeah I just wanted to nail the point home more and more yes um
文化现象总是伴随着,呃,你知道的,媒体
144:40
so let me ask you on your like I said you do get a little
里的一些群体和反应,基本上把当前社会所有
144:45
bit uh um you know flavorful on the internet uh yos shabbach tweeted something that
问题都归咎于那个特定的变化,对吧?电曾经
144:51
you loled at uh in reference to H 9000 quote I appreciate your argument and
被认为会害死所有人,火车也被说成是可怕的
144:57
I fully understand your frustration but whether
东西,因为你知道……
145:00
the pod bay doors should be opened or closed is a complex and nuanced issue
pod bay doors应该打开还是关闭,这是一个复杂且微妙的问题。你现在是Meta AI的负责人,嗯,你知道,这真的让我很担心——我们的AI霸主会用这种企业腔调跟我们说话,而你用你的方式在抵抗这一点。嗯,你能谈谈在大公司工作,如何避免过度谨慎反而造成伤害吗?我觉得答案还是open source platforms,让广泛多元的人群去构建AI assista
145:06
so you're at the head of meta AI um you know this is something that
nce,代表全球文化、观点、语言和价值体系的多样性。这样你就不会被单一AI实体束缚,被某种思维方式洗脑。所以我认为这对社会来说是一个非常重要的问题。如果我们真的想要观点多样性,想要AI系统——在未来我们都会通过AI系统互动——那么这些系统必须多样化,以保护思想、信仰、政治观点等等的多样性,以及保护democracy。而与此相悖的是那些认为出于安全原因,应该把AI
145:13
really worries me that AI our AI overlords will speak down to us with corporate
系统锁起来的人,因为他们觉得把技术交到所有人手里太危险了,可能会被恐怖分子利用。那会导致一个非常糟糕的未来,我们的信息摄入被少数公司通过proprietary systems控制。你相信人类能用这项技术构建出总体上对 humanity 有益的系统吗?这不就是democracy和free speech的意义吗?我认为是的。你相信机构会做正确的事吗?你相信人们会做正
145:19
speak um of this nature and you sort of resist that with your way of
确的事吗?确实有坏人会做坏事,但他们不会拥有比好人更先进的技术。所以,那就是我的好AI对抗你的坏AI,对吧?我们刚才聊到的例子,比如某个rogue country可能会构建一个AI系统,试图说服所有人发动内战,或者选出一个有利的统治者。但那时他们必须先突破我们的AI系统,对吧?一个带着浓重俄罗斯口音的AI系统,句子还不加冠词,试图说服别人——嗯,那至少会很明显。
145:25
being um is this something you can just comment on of working at a big
对,我就是想把这个观点反复强调清楚。好,那让我问你,就像我说的,你在网
145:33
company how you can avoid the over fearing I suppose the through caution create harm
上偶尔会有点情绪化,yos shabbach 发了一条推文,你看了 lo
145:41
yeah again I think the answer to this is open source platforms and then en
l 了,提到 H 9000,引用说“我理解你的论点,也完全明白你的 fr
145:49
enabling a widely diverse set of people
ustration,但问题是……”
145:52
to build AI assistance that represent the diversity of uh cultures opinions languages and value
构建能够代表全球文化、观点、语言和价值体系多样性的AI助手,这样你就不会因为单一的AI实体而被某种特定思维方式洗脑。所以我认为这对社会来说是一个非常非常重要的问题。如果我们真的想要观点多样性的AI系统,在这个未来里我们都会通过AI系统互动,那么这些系统必须多样化,以保护思想、信仰、政治观点等等的多样性,以及保护民主。而与此相悖的是那些认为出于安全原因应该把AI系统锁起来的人
145:59
systems across the world um so that you're not bound to just uh you know
,因为他们觉得把AI交到每个人手里太危险了,可能会被恐怖分子利用之类的。那会导致一个非常糟糕的未来,我们所有的信息摄入都被少数公司通过专有系统控制。你相信人类能用这项技术构建出总体上对 humanity 有益的系统吗?这不就是民主和言论自由的意义所在吗?我觉得是的。你相信机构会做正确的事吗?你相信人们会做正确的事吗?确实会有坏人做坏事,但他们不会拥有比好人更先进的技术,所以
146:05
be brainwashed by a particular way of thinking because of single AI entity um so
到时候就是我的好AI对抗你的坏AI,对吧?我们刚才讨论的例子,比如某个 rogue 国家可能会构建一个AI系统,试图说服所有人发动内战,或者选出一个对他们有利的统治者。但那样的话,他们必须先突破我们的AI系统,对吧?一个带着浓重俄罗斯口音的AI系统会试图说服我们,而且句子还不加冠词。不过在人形机器人方面的进展,我认为确实重振了整个行业。波士顿动力在这个领域已经领先了非常久。
146:12
I mean I I think it's really really important question for society
现在有各种各样的公司,Figure AI 当然,波士顿动力,还有 un tree,但真的很多公司,这很棒,我很喜欢。不过他们还是造不出家用机器人,对吧?而且我们距离完全自主的L5级驾驶还有一段距离。当然,我们离拥有一个能像17岁少年那样通过驾驶20小时来自我训练的L5级自动驾驶AI系统,也还非常遥远。
146:17
and the problem I'm seeing is um is that um which is why I've been
我看到的問題是——這也是為什麼我一直這麼直言不諱,甚至有時帶點諷刺,永遠別停,Yan,我們愛你——因為我看到了這種權力集中透過專有AI系統帶來的危險,比所有其他問題都大得多。如果我們真的想要意見多樣性,想要在未來我們都透過AI系統互動的世界裡,這些系統必須是多樣的,才能保護思想、信仰、政治觀點等等的多樣性,以及保護民主。而與此對立的是那些認為出於安全考量,應該把AI系統鎖起來、不讓所有人接觸的人,因為太危
146:24
so vocal and sometimes a little sardonic about it never stop never stop Yan we
險了,怕被恐怖分子利用之類的。這會導致一個非常糟糕的未來,我們所有的資訊攝取都被少數公司透過專有系統控制。你相信人類能用這項技術建立整體上對人類有益的系統嗎?這不就是民主和言論自由的核心嗎?我認為是的。你相信機構會做正確的事嗎?你相信人們會做正確的事嗎?確實有壞人會做壞事,但他們不會擁有比好人更先進的技術,所以到頭來就是我的好AI對抗你的壞AI,對吧?我的意思是,我們剛才談到的例子,比如某個流氓國家可能會建
146:31
love it is because I see the danger of this concentration of power through through
立一個AI系統,試圖說服所有人發動內戰,或選出一個對他們有利的統治者,但他們必須突破我們的AI系統,對吧?一個帶有濃重俄羅斯口音的AI系統,句子裡連冠詞都不加,試圖說服我們——這至少會荒謬到好笑的地步。好了,既然我們聊到了物理現實,我想問你對未來的願景,關於機器人在這個物理世界中的角色。你提到的許多智能類型,都能讓機器人成為我們人類更有效的合作夥伴。既然Tesla的Optimus團隊展示了一些人形機器人的進
146:38
proprietary AI systems has a much bigger danger than everything
展,我認為這真的重新點燃了整個行業,而這個行業長期以來一直由Boston Dynamics領先。現在有各種各樣的公司,Figure AI當然是,Boston Dynamics,UniTree,UniTree,還有很多,這很棒,真的很棒,我喜歡。所以你認為未來會有數百萬台嗎?
146:43
else that if we really want you know uh diversity of opinion uh AI systems
没有产生预期的结果,呃,我们在那种情况下用 RL 来调整世界模型或 critic。对,所
146:49
that you know in in this future that where we'll all be interacting through AI
以,呃,你提到了 RLHF,reinforcement learning with hu
146:55
systems we need those to be diverse for the preservation of uh uh diversity of
man feedback,为什么你还是讨厌 reinforcement learning
147:01
ideas and you know Creeds and political opinions and and and whatever uh and the
?我并不讨厌 reinforcement learning,我觉得它不应该被完全抛弃,但我
147:07
preservation of
认为它的……
147:08
democracy and what works against this is people who think that for reasons of security
嗯,这是你能评论一下的吗?在大公司工作,怎么避免过
147:16
we should keep AI systems under lock and key because it's too dangerous to put
度谨慎反而造成 harm?对,我觉得答案还是 op
147:24
it in the hands of of everybody um because it could be used by terrorists
en source platforms,然后让各种
147:31
or something um that would lead
各样的人都能参与进来。
147:34
to uh you know potentially a uh a very bad future in which all of
我们是拥抱变化,还是抗拒它?真正的危险是什么,而不是那
147:40
our information diet is controlled by a small number of uh companies through proprietary systems
些想象出来的危险?所以人们担心,我觉得他们担心的一件事
147:46
do you trust humans with this technology to uh to build systems that are on
是关于大科技公司——我们一直在反复讨论,但我觉得值得再
147:52
the whole good for Humanity isn't that what democracy and free speech is
提一次——他们担心AI会变得多强大,他们担心……
147:58
all about I think so do you trust institutions to do the right thing do
否则,如果我们真的想要 diversity of opinion,想要 AI syste
148:02
you trust people to do the right thing and and yeah there's bad people who
ms 在未来我们都会通过 AI systems 互动的世界里保持多样性,那这些系统就必须
148:06
are going to do bad things but they're not going to have Superior technology to
是多元的,才能保护 ideas、creeds、political opinions 等等
148:10
the good people so then it's going to be my good AI against your bad
的多样性,以及保护 democracy。而与之相反的是那些认为出于安全考虑,应该把 AI
148:15
AI right I mean there the examples that we were just talking about of you
systems 锁起来的人,因为他们觉得太危险了,不能交给所有人,怕被恐怖分子利用之类的
148:19
you know maybe uh some Rogue country will build you know some AI system that's
。那会导致一个非常糟糕的未来,我们所有的信息摄入都被少数公司通过 proprietary
148:23
going to try to
systems 控制。
148:24
convince everybody to go into a civil war or something or or or elect a
说服大家去打一场内战,或者选一个合意的统治者之
148:31
favorable U ruler and um but then they will have to go past our AI
类的,但他们得先过我们AI系统这一关——一个带
148:37
systems right an AI system with a strong Russian accent will be trying to conv
着浓重俄罗斯口音的AI系统,会试图说服我们,而
148:43
our and doesn't put any uh articles in their sentences um well it'll be at
且它句子里面一个冠词都不加。嗯,这至少会出现在
148:50
the very
最前沿。
148:50
least absurdly comedic okay uh so I uh since we talked about sort of the
最不荒诞的喜剧感,好吧,呃,所以,既然我们聊到了物理现实,我很想听听你对未来机器人的愿景,在这个物理现实中
148:56
uh physical reality I'd love to ask your vision of the future with with robots
。你提到的很多种智能,都能让机器人成为我们人类更有效的协作者。嗯,自从特斯拉的Optimus团队展示了一些人
149:02
in in this physical reality so many of the kinds of intelligence you've been speaking
形机器人的进展,我觉得这真的重新激活了整个行业,而这个行业我认为波士顿动力已经引领了非常非常久。所以现在有各
149:07
about would Empower robots to be more effective collaborators with us humans so um since
种各样的公司,Figure AI显然,波士顿动力当然,还有UniTree,UniTree,呃,但还有很多很多
149:13
uh Tesla's Optimus uh team has been showing off some
,这很棒,真的很棒,我很喜欢。那么,你觉得未来会有数百万个这样的系统吗?
149:17
progress on humanoid robots I think it really reinvigorated the whole industry that's that I
人形机器人方面的进展,我觉得真的重新激活了整个行业。波士顿动
149:22
think Boston Dynamics has been leading for a very very long time so now there's
力在这个领域已经领先了非常非常久,所以现在有各种各样的公司,
149:28
all kinds of companies figure AI obviously Boston Dynamics um un tree un tree uh
比如Figure AI,当然还有波士顿动力,嗯,还有UniTr
149:34
but there's like a lot of them it's great it's great I mean I love
ee,嗯,UniTree,呃,但还有很多其他的,很棒,真的很
149:39
it uh so do you think there'll be uh millions of
棒,我是说我很喜欢。那么你觉得会有数百万台吗?
149:44
humanoid robots walking around soon not soon but it's going to it's going to happen
人形机器人到处走,不是马上,但肯定会发生,我觉得未来十年机器人领域会非常有意思。机器人产业的崛起已经等了十年二十年了,一直没真正起来,除了那种预设好的行为之类的东西。主要问题还是Moravec悖论,就是怎么让这些系统理解世界是怎么运作的,然后规划行动。我们可以在非常专门的任务上做到这一点。波
149:48
like the next decade I think is going to be really interesting in robots like
士顿动力的做法基本上是靠大量手工构建的动力学模型和事先精心规划,这是非常经典的机器人学,加上很多创新和一点点感知能力。但这并不意味着他们能造出一个家用机器人,对吧。我们离完全自动驾驶的L5级别还有一段距离,嗯,而且离那种能像17岁少年一样自己开20小时车来训练自己的L5自动驾驶系统还差得很远。
149:53
the the emergence of the robotics industry has been in the waiting for you know
所以,除非我们有了世界模型,那种能自我训练理解世界运作的系统,否则机器人的重大进展不会到来。现在很多做机器人硬件的人都在押注AI能在这方面取得足够进展,他们也希望能从中发现一个产品。嗯,在真正强大的世界模型出现之前,会有一个接近强大的世界模型,人们正试图在一个笨拙的机器人身上找到产品,比如不是
149:58
10 20 years without really emerging other than for like you know kind of preprogram
那种完美高效的机器人。所以在工厂环境里,人形机器人可以帮助自动化一些方面,我觉得这因为安全要求之类的东西,是个极其困难的任务。我觉得家用场景更有趣,但然后你就会开始想,我记得你提到过装洗碗机对吧?嗯,我觉得那是你正在研究的主要问题之一。我的意思是,还有打扫房子、饭后收拾桌子、洗碗,所有这些任务
150:03
behavior and stuff like that um and uh and the main issue is again the
,还有做饭,原则上都可以自动化,但实际上非常复杂、非常困难。但即使只是在充满不确定性的非结构化空间里基本导航,现在某种程度上也能做到,导航没问题。但以对人类来说有吸引力的方式导航,那是另一回事。嗯,不一定,我们其实有演示,因为Fair有个所谓的具身AI小组,他们没有自己造机器人,而是用商用机器
150:08
Maric Paradox like you know how do
人。你可以告诉一个机器狗去冰箱那里,它真的能打开冰箱,可能还能从冰箱里拿一罐饮料之类的东西,然后带给你。所以它能导航,能抓取物体,只要——
150:10
we get those system to understand how the world works and and kind of you
我们让这些系统理解世界是如何运作的,并且,嗯,规划行动,这样我们就能完成非常专门的任务。嗯,波士顿动力的做法,基本上是用大量手工打造的动力学模型和精心的事先规划,这是非常经典的机器人学,加上很多创新和一点点感知能力。但即便如此,他们还是造不出家用机器人,对吧?嗯,我们距离完全自主的
150:15
know plan actions and so we can do it for really specialized tasks um and
L5级自动驾驶还有一段距离,嗯,我们当然也离那种像任何17岁青少年一样,通过开20小时车就能自我训练的L5级自动驾驶系统非常遥远。嗯,所以,除非我们再次拥有世界模型——那些能自我训练、理解世界运作的系统——否则我们在机器人领域不会有重大进展。所以,目前很多从事机器人硬件的人,都在押
150:21
uh the way Boston Dynamics goes about it is you know basically with a lot
注或指望AI能在这方面取得足够进展,他们也希望从中找到一个产品。嗯,是的,在你拥有一个真正强大的世界模型之前,会有一个近乎强大的世界模型,嗯,人们正试图在一个笨拙的机器人中找到产品,我想,不是一个完美高效的机器人。所以,在工厂环境中,人形机器人可以帮助自动化一些方面,我认为这因为所有
150:26
of um handcrafted dynamical models and careful planning uh in advance which is very classical
安全要求之类的东西,是一项极其困难的任务。我觉得在家用领域更有趣,但然后你会开始想,我记得你提到过装洗碗机,对吧?嗯,我想那是你正在研究的主要问题之一。我的意思是,还有,嗯,打扫卫生、收拾屋子、饭后清理桌子、洗碗,所有这些任务,嗯,还有做饭,我的意思是,所有那些原则上可以自动化、但
150:32
robotics with a lot of innovation a little bit of perception um
实际上极其复杂精妙的任务。但即使是基本导航,在一个充满不确定性的非结构化空间中,某种程度上是可行的,比如你现在可以做到导航没问题。嗯,但以对我们人类有吸引力的方式进行导航,那是另一回事。嗯,这不一定是,我的意思是,我们已经有……
150:36
but it's still not like they can't build a domestic robot right um and you
但他们也不是造不出家用机器人,对吧?嗯,距离完全自动驾驶的L5级别还有一段距离,嗯,我们离
150:43
know we're still some distance away from completely autonomous level five driving mhm uh and
那种能像17岁少年一样自己开20小时车、满足所有安全要求的L5自动驾驶系统还很远。我觉得家用
150:50
we're certainly very far away from having uh you know level five autonomous driving bi
场景更有意思,但接着你会想——你刚才提到洗碗机对吧?对,我猜那是你们正在攻克的主要问题之一。
150:57
A system that can train Itself by driving 20 hours like any 17y
我的意思是,还有打扫房间、饭后收拾桌子、洗碗,所有这些任务,还有做饭,全都算上。
151:03
old uh so until we have uh again World models systems that can train themselves
嗯,直到我们有了世界模型系统,能够自己训练去理解世界是如何运作的,我们才能在机器人领域取得重大进展。所以现在很多做机器人硬件的人,都在押注AI能在这方面取得足够进展,也希望从中找到产品。在你拥有一个真正强大的世界模型之前,会先有一个接近强大的世界模型,而人们正试图在一个笨拙的机器人身上找到产品,大概不是那种完美高效的机器人。比如在工厂环境中,人形机器人可以
151:09
to understand how the world Works uh we're not going to we're not going to
帮助自动化某些环节。我觉得那是个极其困难的任务,因为涉及所有安全要求之类的东西。我认为家庭场景更有趣,但接着你会开始想——我记得你提到过装洗碗机对吧?对,我想那是你正在研究的主要问题之一。还有打扫屋子、收拾餐桌、洗碗,所有这些任务,理论上可以自动化,但实际上非常复杂、非常困难。就连在充满不确定性的非结构化空间里基本导航,现在勉强能做到,导航还行。但以对人类
151:15
have significant progress in robotic so a lot of the people working on robotic Hardware
来说有吸引力的方式导航,那是另一回事。不一定,我们直接与AI系统在物理空间互动,这样能让我们从哲学和心理层面探索与机器人的关系,会非常非常有趣。所以我希望你在那个japa项目上尽快取得进展。嗯,我希望事情能按计划进行。我们研究从视频中自监督学习这个想法已经十年了,真正取得显著进展只是最近两三年。你提到过,有很多有趣的突破不需要大量算力就能实现。所以如果你对
151:21
at the moment are are betting or banking on the fact that AI is going
做这类方向的PhD感兴趣,还有很多可能性去做创新工作。那么,你会给一个想读研读博的本科生什么建议?基本上我已经列过了:如何通过观察训练世界模型,不一定非得用海量数据集——当然,要像语言模型那样涌现新特性,可能确实需要大数据集——但我觉得有很多好想法不需要一味扩大规模就能实现。还有,如何用学到的世界模型做规划,如果系统演化的环境不是物理世界,而是比如互联网的
151:27
to make sufficient progress
世界,或者某种世界,其中动作包括在搜索引擎里搜索、查询数据库、运行模拟、调用计算器或解微分方程。
151:29
towards that and they're hoping to discover a product in it too is uh yeah
但问题在于,它们还是造不出家用机器
151:34
before you have a really strong World model there'll be an almost Strong World model
人,对吧?嗯,而且我们距离完全自动驾
151:39
and um people are trying to find a product in a clumsy robot I suppose
驶的L5级别还有一段距离,嗯,我们当
151:44
like not a perfectly efficient robot so there's the fact factory setting where uh humanoid
然也离那种能像17岁司机一样开20
151:49
robots can help automate some aspects of the factory I think that's a crazy difficult
小时车来自我训练的L5自动驾驶系统非
151:54
task because of
常遥远。
151:55
all the safety required and all this kind of stuff I think in the home
直接与物理空间中的AI系统交互,这样一来,我
152:00
is more interesting but then you start to think I think you mentioned loading the
们就能从哲学和心理层面探索人与机器人的关系,这
152:06
dishwasher right yeah like I suppose that's one of the main problems you're working on
真的会非常非常有趣。所以我希望你在那个JAPA
152:11
I mean there's you know uh cleaning up cleaning the house uh clear clearing up
项目上能尽快取得进展。嗯,我也希望事情能按计
152:16
the table after meal washing the dishes you know all those tasks you know cooking
划推进。我的意思是,我们一直在研究这个自监督学
152:22
I mean all
习的想法。
152:23
the tasks that you know in principle could be automated but are actually incredibly sophisticated
朝着那个方向,他们希望也能从中发现一个产品,嗯,是的,在你拥有
152:28
really complicated but even just basic navigation around an un Space full of uncertainty that's
一个真正强大的世界模型之前,会有一个接近强大的世界模型,嗯,人们
152:33
sort of works like you can sort of do this now navigation is fine well
正试图在一个笨拙的机器人中找到产品,我想,不是一个完美高效的机
152:38
navigation in a way that's compelling to us humans is is is a different thing
器人。所以有工厂环境,人形机器人可以帮助自动化工厂的某些方面。我
152:43
yeah it's not going to be you know necessarily I mean we have
觉得那是一个极其困难的任务,因为所有安全要求之类的东西。
152:47
demos actually because you know there is a So-Cal embodied AI group at at fair
demos 实际上,你知道在 Fair 那边有一个 So-Cal embodied AI group,他们并没有自己造机器人,而是用商业机器人。你可以让一只机器狗去冰箱那边,它真的能打开冰箱,可能还能从冰箱里拿一罐饮料之类的东西,然后带给你。所以它能导航、能抓取物体,只要它能直接跟 AI 系统在物理空间里互动。这样一来,从哲学和心理学的角度去探索我们
152:52
and uh you know they've been not building their own robots but using commercial robots
跟机器人的关系,会变得非常非常有趣。所以我希望你在整个 japa 项目上能尽快有进展。嗯,我也希望事情能按计划顺利推进。我们其实一直在研究这个从视频中做 self-supervised learning 的想法,已经做了十年了,但真正有显著进展也就是最近两三年。而且你刚才也提到,有很多有趣的突破其实不需要大量 compute 就能实现。所以如果你对做这
152:57
um and you can you can tell a robot dog like you know go to
类方向的 PhD 感兴趣,还是有很多可能性去做创新工作的。那么,你会给一个想读 grad school 和 PhD 的本科生什么建议呢?基本上我刚才已经列了一些了:比如怎么通过观察来训练一个 world model,而且不一定非得用巨大的数据集来训练——当然,要出现像 LMS 那样的涌现特性,可能还是需要大规模数据——但我认为有很多好想法其实不需要
153:01
the fridge and they can actually open the fridge and they can probably pick up
scaling up 就能做。然后还有怎么用学到的 world model 来做 planning,如果系统演化的环境不是物理世界,而是比如互联网的世界,或者某种 action 是搜索搜索引擎、查询数据库、运行模拟、调用计算器、解微分方程的世界,那怎么让系统真正规划出一系列 action 来解决问题?所以 planning 的问题不只是规划物理动作,也
153:06
a can in the fridge and stuff like that and and bring it to you
可以是规划用工具来对话或者用于任何智能系统。这方面有一些工作,但不算特别多。Fair 那边有一个叫 tool former 的工作,是几年前做的,还有一些更近期的 planning 研究,但我觉得我们还没有一个很好的解决方案。然后还有 hierarchical planning 的问题。比如我提过的从纽约到巴黎的旅行规划,那是分层的,但实际上我们几乎
153:11
you know so it can navigate can grab objects as long as
每一个动作都涉及某种意义上的 hierarchical planning,而我们完全不知道怎么做——AI 领域几乎没有 hierarchical planning 的成功示范,尤其是那些必要的不同层级的表示是学出来的。我们最多能做到两层的 hierarchical planning。
153:14
it's been trying to recognize them which you know Vision systems work pretty well nowadays
它一直在尝试识别它们,你知道,现在的视觉系统已经做得相当不错了,但还不是那种完全通用的机器人,能够复杂到做像收拾餐桌这样的事情。对我来说,未来让人形机器人——或者说机器人整体——越来越多地出现,是一件令人兴奋的事,因为这能让人类真正在物理空间中与AI系统直接互动,从而在哲学和心理层面上探索我们与机器人的关系,这真的会非常非常有趣。所以我希望你在整个“jap
153:20
um but but it's not like a completely you know General robot that would be
a”项目上能尽快取得进展。嗯,我希望事情能按计划进行。我们其实一直在研究这个从视频中进行自监督学习的想法,已经做了十年了,但真正取得显著进展也就是最近两三年的事。而且你之前也提到过,有很多有趣的突破可以在不需要大量算力的情况下实现。所以如果你对做这类研究的博士感兴趣,现在仍然有很多可能性去做创新性的工作。那么,你会给一个想读研究生、攻读博士的本科生什么建议呢
153:26
you know sophisticated enough to do things like clearing up the dinner table Yeah to
?基本上我已经列出来了:如何通过观察来训练一个世界模型,而且不一定非得用巨大的数据集来训练——当然,要出现像语言模型那样的涌现特性,可能还是需要大数据集——但我认为有很多好想法可以在不依赖规模扩展的情况下实现。然后还有一点,就是如何用学到的世界模型进行规划,如果系统演化的环境不是物理世界,而是比如互联网的世界,或者某种行动包括在搜索引擎里搜索、查询数据库、运
153:31
me that's an exciting future of getting humanoid robots robots in general in the whole
行模拟、调用计算器、解微分方程的世界,那么如何让系统实际规划出一系列行动来解决问题呢?所以规划的问题不仅仅是规划物理行动,也可以是规划使用工具的行动,用于对话系统或任何智能系统。这方面有一些研究,但不算很多——Fair那边几年前有个叫Tool Former的工作,还有一些近期的规划研究,但我不认为我们有任何好的解决方案。然后还有层级规划的问题。我提到过规划从
153:37
more and more because that gets uh humans to really
纽约到巴黎的旅行,这本身就是层级性的,但实际上我们几乎每一个行动都在某种程度上涉及层级规划,而我们完全不知道怎么做——在AI中,几乎没有层级规划的示范案例,尤其是那些必要的不同层级的表征都是学出来的。我们可以做两层的层级规划,但再往上就……
153:41
directly interact with AI systems in the physical space and in so doing it allows
打算去读研读博,所以我基本上已经列出来
153:46
us to philosophically psychologically explore our relationships with robots can be really really really interesting
了——这个想法就是如何通过观察训练一个世
153:51
so I hope you make progress on the whole uh japa thing soon well I
界模型,而不必非得用海量数据集训练。当然
153:57
I mean I hope I hope things kind of you know work as uh as
,要像LLM那样涌现出特性,可能还是需
154:02
planned um I mean again we've been kind of working on this idea of self
要大规模数据,但我认为有很多好点子可以尝
154:07
supervised
试。
154:08
learning of from video for for 10 years and and you know only made significant
我认为在家用场景更有趣,但然后你开始想,我记得你提到过装洗碗机,对吧?嗯,我
154:12
progress in the last two or three and actually you've you've mentioned that there's a
想那是你正在研究的主要问题之一。我是说,还有打扫、清理房子、饭后收拾桌子、洗碗
154:16
lot of interesting breakthroughs that can happen without having access to a lot of compute
,所有这些任务,还有做饭,我是说,所有原则上可以自动化的任务,但实际上非常复杂
154:21
yeah so if you're interested in doing a PhD and this kind of stuff there's
、极其困难。但即使只是在充满不确定性的非结构化空间中基本导航,这现在某种程度
154:25
a lot of possibilities still yeah to do Innovative work so like what advice would
上可行,你可以做到。导航没问题,但以对人类来说有吸引力的方式导航,那是另一回事
154:29
you give to a undergrad that's
。嗯,它不一定会,我是说,我们有
154:31
looking to uh go to grad school and do a PhD so basically I've listed
那些能协助我们完成所有日常任务——无论是
154:36
them already uh this idea of how do you train a world model by observation
职业还是个人生活——的智能体,那将是非常美
154:40
and you don't have to train necessarily on gigantic data sets or I mean you
妙的事情,因为智能是最稀缺的商品。人类犯
154:45
could out to be necessary to actually train on large data sets to have emerging
下的所有错误,归根结底都是因为缺乏智能或知
154:50
properties like like we have with LMS but I think there's a lot of good
识,这两者是相关的。所以让人们变得更聪明,
154:55
ideas that can be done
只会带来好处。
154:56
without necessarily scaling up then there is how you do planning with a learn World
不一定非要靠规模扩展,那接下来就是如何用学到的
155:02
model if the world the system evolves in is not the physical world but it's
世界模型来做规划——如果系统演化的世界不是物理
155:07
the world of let's say the internet or you know some sort of uh world
世界,而是比如互联网的世界,或者某种世界里,一
155:12
of where an action consists in doing a search in a search engine or interrogating
个动作就是去搜索引擎里搜一下、或者查询数据库、
155:18
a data database or running a simulation or calling a calculator or solving a differential
或者跑个模拟、或者调用计算器、或者解个微分方程。
155:23
equation how do you get a system to actually plan a sequence of actions to
方程是,如何让一个系统真正规划出一系列动作,来给出问题的解决方案。规划的问题不仅仅是规划物理动作,也可以是规划使用工具的动作,比如对话系统,或者任何智能系统。这方面有一些研究,但不算特别多,只有一些工作。Fair 那边有一个叫 tool former 的,是几年前的工作,还有一些更近
155:29
you know give the solution to a problem um and so the question of planning
期的关于 planning 的研究。但我不认为我们有任何好的解决方案。接下来是层级规划的问题。比如我之前提到的,规划一次从纽约到巴黎的旅行,这就是层级式的。但实际上,几乎我们做的每一个动作都涉及某种意义上的层级规划。而我们完全不知道怎么做,AI 领域里几乎没有层级规划的示范,那些必要
155:34
is not just a question of planning physical actions could be you know planning actions
的不同层级的表征都是学出来的。我们只能做到两层级的层级规划。如果它能辅助我们所有的任务,无论是职业还是日常生活,那将是非常棒的事情。因为智能是最稀缺的商品,人类犯的所有错误,本质上都是因为缺乏智能,或者缺乏知识,这两者是相关的。所以让人们变得更聪明,只会是好事。让人变得更聪明,我一直在
155:40
to use tools for a dialog system or for any kind of intelligent system and
用的一个类比是,journalized 可能带来的影响,相当于人类历史上的印刷术的发明。它让每个人都变得更聪明了。人们能够接触到书籍,书比以前便宜得多,所以更多人有了动力去学习阅读,这在以前是不可能的。人们变得更聪明了,它催生了启蒙运动。没有印刷术就不会有启蒙运动。它带来了哲学、理性
155:45
um there's some work on this but not like not a huge amount some work
主义、摆脱宗教教条、民主、科学。当然,没有这些,就不会有美国革命、法国革命,我们可能还在封建制度下。所以它彻底改变了世界,因为人们变得更聪明了,开始学习各种东西。当然,它也带来了欧洲两百年的宗教冲突,因为人们读到的第一本书就是《圣经》,然后意识到也许对《圣经》的解释和神父们说的不一样。
155:51
at Fair um one called tool former which was couple years ago and some more
在Fair,有一个叫Tool Former的项目,是几年前做的,还有一些更近期的关于规
155:56
recent work on planning U but um but I don't think we have like a
划的工作,但我不觉得我们在这些方面有什么好的解决方案。然后还有分层规划的问题——比如我刚
156:01
good solution for any of that then there is the question of hierarchical planning so
才提到的从纽约到巴黎的旅行规划,这就是分层的,但几乎我们做的每一个动作在某种意义上都涉
156:07
the example I I mentioned of you know planning a trip from New York to
及分层规划。而我们真的完全不知道怎么做这个,AI里几乎没有分层规划的示范,尤其是那些必要
156:12
Paris that's hierarchical but almost every action that we take involves
的不同层级的表征都是学出来的。我们最多能做到两层层次的分层规划。
156:16
hierarchical planning in some in some sense and we really have absolutely no idea how
那些在我们所有任务中帮助我们、无论是职业还是个人日常生活中的东西,我觉得那会是非常棒的事情,因为智能是最紧缺的商品,真的,这就是我的意思——人类犯的所有错误都是因为缺乏智能,或者缺乏知识,这两者是相关的。所以让人们变得更聪明,只会是好事,我的意思是,为
156:23
to do this like this's zero demonstration of hierarchical planning uh in AI where the
了让人类更聪明,我一直在用的一个类比是,也许在人类历史上,与通用人工智能可能带来的影响相当的事件,就是印刷术的发明。它让每个人都变得更聪明了,因为人们能够接触到书籍,书比以前便宜得多,所以更多人有了动力去学习阅读,这在以前是没有的。人们变得更聪明了,它
156:31
various levels of representations that are necessary have been learned we can do like two
促成了启蒙运动,对吧?没有印刷术就不会有启蒙运动。它促成了哲学、理性主义、摆脱宗教教条、民主、科学,当然没有这些就不会有美国革命、法国革命,我们可能还在封建制度下。所以它彻底改变了世界,因为人们变得更聪明了,开始学习了解各种事物。当然,它也带来了欧洲两百
156:38
level hierarchy hierarchical planning when we
年的宗教冲突,对吧?因为人们最先读到的就是《圣经》,然后意识到也许对《圣经》的理解和神父们说的不一样。
156:41
design the two the two levels so for example you have like a a dog
设计这两个层级,比如你有一个机器狗,对吧
156:46
lag robot right you want it to go from the living room to the kitchen
?你想让它从客厅走到厨房,你可以规划一条避
156:51
you can plan a path that avoids the obstacle and then um you can send
开障碍物的路径,然后把这个路径发给一个更低
156:56
this to a lower lower level planner that figures out how to move the legs
层级的规划器,由它来决定如何移动腿部来跟随
157:02
to kind of Follow that trajectories right so that works but that twole planning is
那条轨迹。这样是可行的,但这种双层规划是人
157:07
designed by hand
工设计的。
157:08
right um we specify what the proper levels of abstraction the representation that each level
对,我们手动指定了合适的抽象层级,以及每个层级
157:13
of attraction has have to be how do you learn this how do you learn
应该用什么样的表示方式。那怎么学习这个呢?怎么学
157:18
that hierarchical representation of action plans right we you know with cight and deep learning
习这种层级化的行动计划表示?你看,用CNN和深度
157:23
we we can train the system to learn hierarchical representations of percepts mhm what is
学习,我们可以训练系统学习层级化的感知表示。嗯,
157:29
the equivalent when what you're trying to represent our
那如果你想表示的是行动计划呢?
157:32
action plans for action plans yeah so you want you want basically a robot dog
对,行动计划。所以你想要的,基本上就是一个机器
157:37
or humanoid robot that turns on and travels from New York to Paris all by
狗或者人形机器人,能自己开机,然后从纽约跑到巴
157:42
itself for example all right they might have some uh trouble at the at the
黎,全程自主完成。好吧,可能在TSA那里会遇到
157:47
TSA but yeah no but even doing something fairly simple like a household task sure
点麻烦,但就算做一件像家务活这样简单的事,比如
157:52
like you know uh cooking or something yeah that there's a lot involved it's a
做饭什么的,也涉及很多环节,是非常复杂的任务。
157:57
super complex task we take and once
我们往往觉得理所当然。
157:59
again we take it for granted what hope do you have for um the future
你对人类的未来有什么希望?我们聊了这么多激动人心的技术
158:05
of humanity we're talking about so many exciting Technologies so many exciting possibilities what gives
,这么多令人兴奋的可能性。当你展望未来10年、20年、
158:10
you hope when you look out over the next 10 20 50 100 years if
50年、100年的时候,是什么给了你希望?你看社交媒体
158:16
you look at social media media there's a lot of there's there's Wars going on
,有很多战争、分裂、仇恨,这些也是人性的一部分。但在这
158:21
there's division uh there's hatred all this kind
一切之中,是什么让你抱有希望?
158:24
of stuff that's also part of humanity but amidst all that what gives you hope
我没有那个问题。我们可以用AI让人类变得更聪明。我是说,AI本质上会放大人类的智能。就好像我们每
158:33
I don't have that question uh we can make Humanity Smarter with AI okay I
个人都会拥有一支由智能AI助手组成的团队,它们可能比我们更聪明,会执行我们的指令,甚至以远超我们
158:42
mean AI basically will amplify human intelligence it's as if if every one of
能力的方式完成任务,因为它们比我们更聪明。所以每个人都会成为一群超级聪明的虚拟员工的老板。
158:51
us will have a staff of smart AI assistants they might be smarter than us
我们不应该因此感到威胁,就像我们不应该因为管理一群比自己更聪明的人而感到威胁一样。我当然有很多这样的经验,
158:58
they'll do our bidding perhaps execute a task in ways that are much better than
和比我更聪明的人一起工作,这其实是一件很棒的事。所以,拥有比我们更聪明的机器,来协助我们完成所有任务,无论
159:05
we could do ourselves because they'll be smarter than us and so it's like everyone
是职业还是个人生活,我觉得这绝对是一件美妙的事情。因为智能是最稀缺的资源,人类犯下的所有错误,归根结底都是因
159:11
would be the the boss of a staff of super smart virtual
为缺乏智能,或者说缺乏知识,这两者是相关的。所以让人们变得更聪明,只会带来更好的结果。
159:17
people so we shouldn't feel threatened by by this any more than we should feel
所以我们不应该对此感到威胁,就像我们不会因为管理一群比自己更聪明的人而感到威胁一样。我在这方面确实有很多经验,比如和比我更聪明的人一起工作——这其实是一件很棒的事。所以,拥有比我们更聪明的机器,来协助我们完成所有任务、日常
159:22
threatened by being the manager of a group of people some of whom are more
生活中的方方面面,无论是职业还是个人事务,我认为这绝对是一件美妙的事情。因为智能是最稀缺、最被渴求的资源——人类犯下的所有错误,本质上都是因为缺乏智能,或者说缺乏相关的知识。所以,让人们变得更聪明只会带来更好的结果。为了让
159:28
intelligent than us I certainly have a lot of experience with this of uh you
人类更聪明,我一直在用的一个类比是:通用人工智能可能带来的影响,相当于人类历史上印刷术的发明。印刷术让每个人都变得更聪明了——人们能够接触到书籍,书籍变得比以前便宜得多,于是更多人有了动力去学习阅读,这在以前是不可能的。人
159:33
know having people working with me who are smarter than me um that's actually a
们变得更聪明,这催生了启蒙运动,对吧?没有印刷术就不会有启蒙运动。它推动了哲学、理性主义、摆脱宗教教条、民主、科学……当然,没有它就不会有美国革命、法国革命,我们可能还生活在封建制度下。所以它彻底改变了世界,因为人们变得更
159:39
wonderful thing so uh having machines that are smarter than
聪明了,开始学习了解各种事物。当然,它也带来了欧洲大约200年的宗教冲突——因为人们最初读到的就是《圣经》,然后发现也许存在与神父们不同的解读方式。
159:43
us that assist us in our all of our tasks our daily lives whether it's
为了让人类更聪明,我一直在用的一个类比是:通用人工智能可能带来的历史性事件
159:47
professional or personal I think would be absolutely wonderful thing because intelligence is the most
,或许相当于印刷术的发明。它让每个人都变得更聪明——人们能接触到书籍,书比
159:52
um is the commodity that is most in demand that that's really what I mean
以前便宜得多,所以更多人有了学阅读的动力,这在之前是不可能的。人们变得更聪明
159:57
all the mistakes that Humanity makes is because of lack of intelligence really or lack
,它催生了启蒙运动,对吧?没有印刷术就不会有启蒙运动。它推动了哲学、理性主
160:02
of knowledge which is you know related so um making people smarter would just can
义,让人摆脱宗教教条,还有民主、科学。当然,没有它就不会有美国独立战争和法国
160:07
only be better I mean for
大革命,我们可能还在被统治着。
160:09
the same reason that you know public education is a good thing and books are
你知道,公共教育是好事,书本是好事,互联网本质上也是好事,甚至社交网络,只要运营得当,也是好事——这很难,但你知道,嗯,它有助于信息和知识的传播,以及知识的传递。
160:15
a good thing and the internet is also a good thing intrinsically and even social
所以AI会让人变得更聪明。我一直在用的一个类比是,也许人类历史上与AI带来的影响相当的事件,就是印刷术的发明。它让每个人都更聪明了。人们能接触到书,书比以前便宜得多
160:21
networks are a good thing if you run them properly it's difficult but you know
,所以更多人有了动力去学习阅读,这在以前是没有的。然后人们变得更聪明了,它促成了启蒙运动,对吧?没有印刷术就不会有启蒙运动。它推动了哲学、理性主义、摆脱宗教教条、
160:27
you can um uh because you know it it's helps the communication of information and
民主、科学——当然,没有这些就不会有美国革命、法国革命,可能现在还处在封建制度下。所以它彻底改变了世界,因为人们变得更聪明了,开始学习了解各种事物。不过它也造成了欧
160:32
knowledge and the transmission of knowledge so AI is going
洲200多年的宗教冲突,对吧?因为人们最先读到的是《圣经》,然后发现也许对《圣经》的理解,和神父们说的不一样。
160:36
to make Humanity smarter and the analogy I've been using is the fact that perhaps
让人变得更聪明,我一直在用的一个类比是,也许在人类历史上,与journaliz
160:44
an equivalent event in a history of humanity to what might be provided by journalized
ed可能带来的影响相当的事件,就是印刷术的发明。它让每个人都变得更聪明了——人们
160:52
is the invention of the printing the printing press it made everybody smarter the fact
能够接触到哲学、理性主义,摆脱宗教教条,还有民主、科学等等。当然,没有它就不会有
161:00
that people could uh have access to um
美国革命、法国革命,我们可能还活在君主制下。
161:04
two books books were a lot cheaper than they were before and so a lot
为了让人类变得更聪明,我一直在用的一个类比是:也许人
161:12
more people had an incentive to learn to read which wasn't the case before um
类历史上能与journalized带来的变化相提并论
161:19
and people became smarter it it enabled the enlightenment right there wouldn't be an Enlightenment
的事件,就是印刷术的发明。它让所有人都变得更聪明了,
161:26
without the printing press it enabled
因为人们可以接触到……
161:29
uh philosophy rationalism U escape from religious Doctrine um democracy science uh and certainly without
在那之前,呃,30年前,当我们研究com Nets和神经网络早期阶段的时候,呃,我非常
161:40
this it wouldn't be there wouldn't have been theeran American Revolution the French Revolution and
兴奋,因为我看到了一条通往人类级别智能的道路,呃,你知道,就是那些能够理解世界、记忆、规
161:50
so was still be under
划、推理的系统。呃,有一些……
161:54
feudal regimes perhaps um and so it completely transformed the the world because people became
封建制度,嗯,然后它彻底改变了世界,因为人们变得
161:55
smarter and kind of learn learn about things now it also created 200 years of
更聪明了,开始学习各种东西。不过这也导致了欧洲大约
161:56
essentially religious conflicts in Europe right because the first thing that people read was the
两百年的宗教冲突,对吧?因为人们最先读到的是《圣经
161:57
Bible and uh realized that perhaps was a different interpretation of the Bible than what
》,然后发现,嗯,也许对《圣经》的理解和神父们讲的
161:58
the priests were
不太一样。