NeurIPS 2024 · Keynote 2024

The End of Pre-training

预训练的终结
December 2024 · 24m 36s

Ilya Sutskever's landmark NeurIPS 2024 keynote. 'Pre-training as we know it will end'.

Ilya Sutskever 在 NeurIPS 2024 的划时代演讲。'Pre-training as we know it will end' 震动行业。

00:0024m 36s
00:01
want to thank the organizers for choosing a paper for this award it was very
感谢主办方选择这篇论文获奖,这真的很棒。我还要感谢我出色的合著者和合作者,Oral、Vineel和Qule,他们刚才就站在你们面前。现在你们看到的是一张截图,来自十年前、2014年蒙特利尔NeurIPS上的一场类似演讲。那是个更天真的时代——照片里这是“之前”,顺便说一句,这是“之后”。现在我们更有经验了,希望也更明智了。今天我想聊聊这项研究本身,也许做个十年回顾,因为里面很多东西是对的,但有些不太对。我们可以回顾一下,看看发生了什么,以及它如何慢慢演变成今天的样子。那么,先说说我们做了什么吧,方法就是展示十年前的同一场演讲
00:11
nice and I also want to thank my incredible co-authors and collaborators oral vineel and
的幻灯片。但我们做的总结就是以下三点:一个自回归模型,在文本上训练;一个大型神经网络;一个大型数据集。就这些。现在让我们深入细节。这是十年前的幻灯片,还不错。Deep Learning假设:我们当时说,如果你有一个10层的大型神经网络,它就能做任何人类能在瞬间完成的事情。为什么我们特别强调“人类能在瞬间完成的事情”呢?如果你相信Deep Learning的教条——即人工神经元和生物神经元相似,至少差别不大——并且你相信真实神经元很慢,那么任何我们能快速完成的事情——这里的“我们”指人类,甚至指全世界只有一个人能在瞬间完成的任务
00:21
qule who stood right before you a moment
——一个10层神经网络也能做到,对吧?你只需要把它们的连接提取出来,嵌入到你的神经网络里,人工的那个。这就是动机:任何人类能在瞬间完成的事情,一个10层的大型神经网络也能做到。我们当时专注于10层神经网络,因为那是我们当年知道如何训练的神经网络。如果你能突破层数限制,那你就能做更多。
00:26
ago and what you have here is an image a screenshot from a similar talk
十年前,以及你现在看到的这张图,是2014年蒙特利尔新RPS会议上一个类似演讲的截图。那时候氛围要单纯得多。照片里展示的是“之前”,顺便说一句,这是“之后”。现在我们更有经验了,希望也更明智了。但今天我想稍微聊聊工作本身,或许做一个十年的回顾,因为这项工作中的很多判断是对的,但也有一些不太对。我们可以回顾一下,看看发生了什么,以及它是如何平稳地演变成今天这个局面的。那么,我们先从我们当时做了什么开始讲起,方式是通过展示十年前的同一场演讲的幻灯片。我们当时做的总结就是以下三点:一个在文本上训练的auto regressive模型,一个大型神经网络,以及一个大型数据集。就这些。现在我们再深入一点细节。这是十年前的一张幻灯片,还不错。Deep
00:35
10 years ago at new RPS in 2014 in Montreal and it was a much
Learning假设,我们当时说的是:如果你有一个10层的大型神经网络,那么它就能做到任何人类在不到一秒内能做到的事情。为什么我们当时要强调“人类在不到一秒内能做到的事情”呢?为什么特别强调这个?如果你相信所谓的Deep Learning教条,即人工神经元和生物神经元是相似的,或者至少差别不大,并且你相信真实的神经元很慢,那么任何我们能快速做到的事情——这里的“我们”指人类,甚至指全世界只要有一个人能在不到一秒内完成某项任务——那么一个10层的神经网络也能做到,对吧?逻辑就是:你只需把那些连接取过来,嵌入到你的神经网络(人工的)里。这就是当时的动机:任何人类能在不到一秒内做到的事情,一个10层的大型神经网络也能做到。我们当时专注于10层神
00:44
more innocent time here we are shown in the
经网络,因为那是我们当时知道如何训练的神经网络。如果你能设法增加层数,那你就能做得更多。模型只要足够好地预测下一个token,它实际上就能抓住、捕捉并掌握接下来任何序列的正确分布。这在当时算是个相对新鲜的事物。它并非严格意义上的第一个auto regressive神经网络,但我认为它是第一个我们真正相信“如果你把它训练得足够好,你就能得到一切”的auto regressive神经网络。
00:49
photos this is the before here's the after by the way and now we've got
照片这是之前,这是之后,顺便说一句。现在我们更有经验了,希望也更明智了,但在这里我想聊聊工作本身,也许做个十年回顾,因为这项工作里很多东西是对的,但有些则不然。我们可以回顾一下,看看发生了什么,以及它是如何逐渐演变到今天这个样子的。那么,我们先从我们做了什么开始讲起,方式是通过展示十年前同一场演讲的幻灯片,但我们所做的工作总结就是以下三点:一个在文本上训练的自回归模型,一个大型神经网络,以及一个大型数据集,就这些。现在我们再深入一点细节。这是十年前的幻灯片,还不错。深度学习的假设,我们当时说的是,如果你有一个十层的大型神经网络,它就能做到人类能在瞬间完成的任何事情。为什么我们特别强调人类能在瞬间完成的事情呢?具体来说,如
00:58
more experienced hopefully wiser but here I'd like to talk a little bit about the
果你相信所谓的深度学习教条,即人工神经元和生物神经元相似,或者至少差别不大,并且你相信真实神经元很慢,那么任何我们能快速做到的事情——这里的“我们”指人类,甚至只是全世界某一个人——如果全世界有一个人能在瞬间完成某项任务,那么一个十层的神经网络也能做到,对吧?逻辑就是,你只需提取它们的连接,然后嵌入到你的神经网络(人工的)中。所以这就是动机:任何人类能在瞬间完成的事情,一个大型十层神经网络也能做到。我们当时专注于十层神经网络,因为那是我们那时知道如何训练的神经网络。如果你能设法增加层数,你就能做更多。模型如果能足够好地预测下一个token,它实际上就能抓取、捕捉并掌握接下来任何序列的正确分布。这在当时是相对新颖的,它并非
01:07
work itself and maybe a 10year retrospective on
字面意义上的第一个自回归神经网络,但我认为它是第一个我们真正相信的——如果你训练得足够好,你就能得到一切。我们用八个GPU实现了3.5倍的加速。而结论幻灯片,从某种意义上说,当时那场演讲的结论幻灯片是最重要的,因为它阐明了可以说是一切的开端——即扩展假设:如果你有一个非常大的数据集,并且训练一个非常大的神经网络,那么成功就会随之而来。
01:12
it because a lot of the things in this work were correct but some not
这份工作中的很多内容是对的,但也有一些不太对,我们可以回顾一下,看看发生了什么,以及它是如何慢慢演变到今天这个样子的。那就从聊聊我们当时做了什么开始吧,我们会用十年前同一场演讲的幻灯片来展示,但总结起来就是以下三点:一个自回归模型,在文本上训练;一个大型神经网络;以及一个大型数据集。就这些。现在我们再深入一点细节。这是十年前的一张幻灯片,还不错。深度学习的假设,我们当时说的是,如果你有一个十层的大型神经网络,它就能做到任何人类能在瞬间完成的事情。
01:19
so much and we can review them and we can see what happened and how
为什么我们特别强调人类能在瞬间完成的事情呢?为什么偏偏是这一点?好吧,如果你相信所谓的深度学习教条——也就是人工神经元和生物神经元是相似的,或者至少差别不大——并且你相信真实神经元很慢,那么任何我们能快速完成的事情,这里的“我们”指人类,甚至指全世界只有一个人能做到的事情,如果全世界有一个人能在瞬间完成某项任务,那么一个十层的神经网络也能做到,对吧?逻辑就是,你只需把那些连接取过来,嵌入到你的神经网络里,人工的那个。所以这就是动机:任何人类能在瞬间
01:25
it gently flowed to where we are today so let's begin by talking about what
完成的事情,一个大的十层神经网络也能做到。我们当时专注于十层神经网络,因为那是我们那时候知道怎么训练的神经网络。如果你能想办法增加层数,那你就能做得更多。模型如果能足够好地预测下一个词,它实际上就能抓住并掌握接下来序列的正确分布。这在当时还算比较新,它不是第一个自回归神经网络,但我认为它是第一个我们真正相信,如果你训练得足够好,就能得到一切的自回归神经网络。我们用八个GPU实现了3.5倍的加速。而结论幻灯片,从某种意义上说,是当时那场演讲最重要的幻
01:32
we did and the way we'll do it is by showing
灯片,因为它阐明了可以说是缩放假设的开端:如果你有一个非常大的数据集,并且训练一个非常大的神经网络,那么成功就是必然的。但互联网只有一个。所以这里我稍微自由发挥一下,猜测接下来会发生什么。其实我不需要猜测,因为很多人也在猜测,我会提到他们的猜测。你可能听过“智能体”这个词,这很常见,我相信最终会有进展,但人们觉得智能体这个东西……
01:37
slides from the same talk 10 years ago but the summary of what we did
同一场演讲十年前的幻灯片,但当时我们做的事情总结
01:42
is the following three bullet points it's an auto regressive model train on text it's
起来就是以下三个要点:它是一个自回归模型,在文本
01:48
a large neural network and it's a large data set and that's it now let's
上训练;它是一个大型神经网络;它是一个大型数据集。
01:54
dive in into the details a little bit more so this was a slide 10
就这些。现在我们稍微深入一点细节,这就是十年前的
02:00
years ago
幻灯片。
02:01
not too bad the Deep load hypothesis and what we said here is that if
不算太糟。Deep Load 假说,以及我们当时说的意思是,如果你有一个10层的大神经网络,它就能做到人类在几分之
02:07
you have a large neural network with 10 layers then it can do anything that
一秒内能做的任何事情。为什么我们当时这么强调人类在几分之一秒内能做的事呢?神经网络也能做到,对吧?你只需要把它们的连
02:13
a human being can do in a fraction of a second like why did we
接拿过来,嵌入到你的神经网络里——人工神经网络里。所以这就是动机:任何人类能在几分之一秒内做到的事,一个大的10层神
02:18
have this emphasis emphasis on things that human beings can do in a fraction of
经网络也能做到。我们当时聚焦在10层神经网络上,因为那是我们当年知道怎么训练的神经网络。如果你能想办法让层数更多,那
02:24
a second
你就能做更多事。
02:25
why this thing specifically well if you believe the Deep learning Dogma so to say
不算太糟。深度学习的假设,我们
02:30
that artificial neurons and biological neurons are similar or at least not too different and
当时说的是,如果你有一个十层的大
02:34
you believe that real neurons are slow than anything that we can do quickly by
型神经网络,那么它就能做到任何
02:38
we I mean human beings I even mean just one human in the entire world
人类能在不到一秒内完成的事情。为
02:43
if there is one human in the entire world that can do some task in
什么我们特别强调“人类能在不到
02:47
a fraction of a second then a 10 layer
一秒内完成的事情”?
02:50
neural network can do it too right it follows you just take their connections and
神经网络也能做到,对吧?它其实就是你把它们的连
02:54
you embed them inside your neuronet the artificial one so this was the motivation anything
接拿过来,然后嵌入到你的人工神经网络里。所以这就
02:59
that a human being can do in a fraction of a second second a big
是动机——人类能在不到一秒内完成的任何事情,一个
03:04
10 10 layer neural network can do too we focused on 10 layer neural networks
大的十层神经网络也能做到。我们当时专注于十层神经
03:08
because this was the neural networks we knew how to train back in the day
网络,因为那是我们那时候知道怎么训练的神经网络
03:13
if you could go beyond in your layers somehow then you could do more
。如果你能想办法增加层数,那你就能做更多事情。
03:17
but back then we could only do 10 layers which is why we emphasized whatever
但当时我们只能做10层,所以才会强调“人类能在几分之一秒内完成的事”——这是那次演讲的另一张幻灯片,上面写
03:22
human beings can do in a fraction of a second a different slide from the
着“我们的核心想法”。你可能能认出两样东西,或者至少一样:你可能会注意到这里有个自回归的东西在运作。它到底在
03:27
talk a slide which says our main idea and you may be able to recognize
说什么?这张幻灯片真正想表达的是:如果你有一个自回归模型,并且它能足够好地预测下一个token,那么它实际上
03:32
two things or at least one thing you might be able to recognize that something
就能抓住、捕捉并掌握接下来任何序列的正确分布。这在当时算是个比较新的事情——它并不是历史上第一个自回归神经网
03:36
Auto regressive is going on here what is it saying really what does this slide
络,但我觉得它是第一个我们真正相信“如果你把它训练得足够好,你就能得到你想要的东西”的自回归神经网络。在我
03:41
really say this slide says that if you have an auto regressive
们当时那个场景下,这个“东西”就是翻译任务——现在看起来很普通,但在当时简直是异想天开。
03:45
model and it predicts the next token well enough then it will in fact grab
为什么偏偏是这件事?如果你相信深度学习教条——也就是说,人工神经元和生物神经元是相似的,或者至少差别不大——而
03:50
and capture and grasp the correct distribution over whatever over sequences that come next and
且你相信真实神经元很慢,那么任何我们能快速完成的事情,这里的“我们”指人类,甚至指全世界只要有一个人类,如果全世
03:56
this was a relatively new thing it wasn't literally the first ever Auto regressive neural
界有一个人能在不到一秒内完成某项任务,那么一个十层的神经网络也能做到,对吧?你只需要把他们的连接拿过来,嵌入到你
04:01
network but I would argue it was the first Auto regressive neural network where we
的人工神经网络里就行了。这就是当时的动机:任何人类能在不到一秒内完成的事情,一个十层的大型神经网络也能做到。我们
04:07
really believed that if you train it really well then you will get whatever
当时专注于十层神经网络,因为那是我们那时候知道怎么训练的神经网络。如果你能突破层数限制,那你就能做得更多。
04:12
you want in our case back then was the humble today humble then incredibly audacious
接下来我要给你们看一些古老的历史,很多人可能从没见过。它叫LSTM。对不熟悉的人来说,LSTM就是可怜的深度学习研究员在
04:18
task of translation now I'm going to show you some ancient history that many of
Transformer之前用的东西。它基本上就是一个ResNet,但旋转了90度。所以这就是LSTM,它出现在Transf
04:25
you might have never seen before it's called the lstm to those unfamiliar an lstm
ormer之前,有点像稍微复杂一点的ResNet。你可以看到这里有个积分器,现在叫残差流,但这里还有一些乘法运算,稍微复
04:31
is the things that poor deplan researchers did before
杂一点——不过这就是我们当时用的东西,就是一个旋转了90度的ResNet。
04:35
Transformers and it's basically a res net but rotated 90° so that's an lsdm and
Transformer 基本上就是一个 ResNet 但旋转了 90
04:41
it came before it's like it's like kind of like a slightly more complicated reset
度,所以它其实是个 LSDM,而且它出现得更早——有点像稍微复杂一点的
04:47
you can see there is your integrator which is now called the residual stream but
ResNet。你能看到里面有个 integrator,现在叫 res
04:53
you've got some multiplication going on it's a little bit more complicated but that's what
idual stream,但多了些乘法运算,稍微复杂一点,不过我们当时
04:59
we did it
就是这么做的。
05:01
was a reset Ro 90° another cool feature from that Old talk that I want
那次老演讲里另一个我想强调的酷炫特点是:我们用了并行化,但不
05:07
to highlight is that we used parallelization but not just any parallelization we used pipelining
是随便哪种并行化——我们用了流水线并行。证据就是“每GPU一
05:13
as witnessed by this one layer per GPU was it wise to pipeline as we
层”。当时用流水线明智吗?我们现在知道流水线并不明智,但当时
05:20
now know pipelining is not wise but we were not as wise back then so
我们没那么聪明,所以我们就用了。结果我们用8个GPU实现了3
05:26
we used that
.5倍的加速。
05:28
and we got a 3.5x speed up using eight gpus and the conclusion slide in
我们用八块GPU实现了3.5倍的加速。而结论那一页—
05:34
some sense the conclusion slide from the talk from back then is the most important
—某种意义上,当年那次演讲的结论页是最重要的一页,因
05:40
slide because it spelled out what could arguably be the beginning of the scaling hypothesis
为它明确提出了可以说是 scaling 假说的起点:
05:45
right that if you have a very big data set and you train a very
如果你有一个非常大的数据集,并且训练一个非常大的神经
05:51
big neural network then success is
网络,那么成功就是……
05:54
guaranteed and one can argue if one is charitable that this indeed has been what's
而那次演讲的结论幻灯片,从某种意义上说,是最重要的一张。因为它明确提出了—
06:00
been happening I want to mention one other idea and this is I claim the
—可以说——scaling hypothesis的开端:如果你有一个非常大
06:07
idea that truly stood the test of time it's the core idea of deploying itself
的数据集,并且训练一个非常大的神经网络,那么成功就是有保障的。如果你愿意宽
06:13
it's the idea of connectionism it's the idea that if
容一点的话,可以说这确实就是后来一直在发生的事情。
06:17
you allow yourself to believe that an artificial neuron is kind of sort of like
我们用八块 GPU 实现了 3.5 倍的速度提升。而那张结论幻灯片
06:24
a biological neuron right if you believe that one is kind of sort like the
,某种意义上说,是当时那场演讲里最重要的一张,因为它明确提出了——可
06:31
other then it gives you the confidence to believe that very large neural networks they
以说——scaling hypothesis 的开端:如果你有一个非
06:38
don't need to be literally human brain scale they might be a little
常大的数据集,并且训练一个非常大的神经网络,那么成功就是……
06:44
bit smaller but you could configure them to do pretty much all the things that
小一点,但你可以把它们配置成能做几乎所有人类做的事
06:49
we do human beings there's still a difference oh I forgot to the end there
情——不过还是有区别。哦,我忘了说,最后还是有区别
06:55
is still a difference because the human brain also figures out how to reconfigure itself
,因为人类大脑还能自己学会重新配置自己,而我们用的
07:00
whereas we are using the best learning algorithms that we have which require as many
是目前最好的学习算法,这些算法需要的参数数量和数据
07:06
data points as there are parameters human beings are still better
点一样多。在这方面,人类仍然更胜一筹。
07:10
in this regard but what this led so I claim arguably to the age of
关于这一点,但由此我可以说,这引向了所谓的预训练时代。预训练时代就
07:18
pre-training and the age of pre-training is what we might say the gpt2 model the
是我们说的GPT-2模型、GPT-3模型、scaling laws,
07:25
gpt3 model the scaling laws and I want to specifically call out my former collaborators
我想特别提一下我以前的合作者Alec Radford,还有Jared
07:33
Alec Radford also Jared Kaplan Dario
Kaplan、Dario。
07:36
mode for really making this work but that led to the age of pre-training and
但这一点带来的结果——我敢说——就是所谓的预训练时代。预训练时代就是我们常说的GPT-2模型、GPT-3模型、scali
07:44
this is what's been the driver of all of progress all the progress that we
ng laws,我特别想提一下我以前的合作者Alec Radford,还有Jared Kaplan、Dario Amode
07:51
see today extra large neural networks extraordinar large neural networks trained on huge data sets
i,他们真的让这一切成为现实。但这带来了预训练时代,而这也是驱动所有进步的引擎——我们今天看到的所有进步,超大神经网络、巨
07:59
but pre-training as we know it will unquestionably end
型神经网络,在庞大数据集上训练。但预训练,就我们现在所知,毫无疑问会结束。
08:04
pre-training will end why will it end because while computers growing through better Hardware better
你允许自己相信一个人工神经元在某种程度上类似于一个生物
08:11
algorithms and logic clusters right all those things keep increasing your compute all these things
神经元,对吧?如果你相信两者有某种相似性,那就会给你信心
08:19
keep increasing your compute the data is not growing because we have but one internet
去相信:非常大的神经网络不需要真的达到人脑的规模,它们
08:26
we have but one
可能只是稍微……
08:28
internet you could even say you can even go as far as to say that
预训练会结束。为什么?因为虽然计算机通过更好的硬件、更好的算法和更大的集群在持续增长—
08:35
data is the fossil fuel of AI it was like created somehow and now we
—所有这些都在增加你的算力——但数据并没有增长,因为我们只有一个互联网。我们只有一个互
08:41
use it and we've achieved Peak data and there'll be no more we have to
联网。你甚至可以说,数据就是AI的化石燃料,它是某种方式被创造出来的,我们现在在用它,
08:48
deal with the data that we have now it still still let us go quite
我们已经达到了数据峰值,不会再有了。我们只能靠现有的数据继续前进,这还能让我们走很远,
08:54
far but this
但互联网只有一个。
08:56
is there's only one internet so here I'll take um a bit of Liberty to
模型只要能把下一个 token 预测得足够好,那么它实际上
09:02
speculate about what comes next actually I don't need to speculate because many people are
就能抓住、捕捉并掌握接下来序列的正确分布。这在当时算是个比
09:08
speculating too and I'll mention their speculations you may have heard the phrase agents it's
较新的事情,它不是史上第一个自回归神经网络,但我认为它是第
09:13
common and I'm sure that eventually something will happen but people feel like something agents
一个我们真正相信只要训练得足够好,就能得到一切的自回归神经网
09:19
is
络。
09:20
the future more concretely but also a little bit vaguely synthetic data but what does
所以,我在这里稍微大胆地推测一下接下来会发生什么。其实我也不用推测,因为很多人也在推测,我会提一下他们的想法。你可能听过“agents”这个词,这很常
09:26
synthetic data mean figuring this out is a big challenge and I'm sure that different
见,我相信最终会有一些进展,但人们觉得agents就是未来。更具体一点,但也稍微模糊一点的是synthetic data。但synthetic dat
09:32
people have all kinds of interesting progress there and an inference time compute or maybe
a到底是什么意思?搞清楚这个问题是个大挑战,我相信不同的人在这方面都有各种有趣的进展。还有inference time compute,或者最近最引人
09:37
what's been most recently most vividly seen in 01 the o1 model these are all
注目的就是o1模型。这些都是人们在试图找出预训练之后该做什么的例子,而且这些都是非常好的方向。我想提另一个来自生物学的例子,我觉得特别酷。这个例子是这
09:43
examples of things of people trying to
样的:很多很多年前,也是在这个会议上,我看到一个演讲,有人展示了这个。
09:46
figure out what to do after pre-training and those are all very good things to
在这方面。但由此引发的结果是——我可以说——预训练时代的到来。而预训
09:53
do I want to mention one other example from biology which I think is really
练时代,我们可以说是 GPT-2 模型、GPT-3 模型、scali
10:01
cool and the example is this so about many many years ago at this conference
ng laws。我想特别提一下我以前的合作者 Alec Radford
10:08
also I saw a talk where someone presented this
、Jared Kaplan、Dario……
10:12
graph but the graph showed the relationship between the size of the body of the
这张图展示的是哺乳动物身体大小和大脑大小之间的关系,这里用的是质量单位。那次演讲我记得特别清楚,他们说:你看,生物学里一切都那么混乱,但这里有一个罕见的例子,动物身体大小和大脑大小之间存在非常紧密的关系。然后完全偶然地,我对这张图产生了好奇。很早的时候,我就去Google搜索这张图,结果Google Images里有一张图是这样的。有趣的是,这张图里——我不知道鼠标好不好使,
10:18
size of the body of a mammal and the size of their brain in this
哦,鼠标没问题——你看这里有各种哺乳动物,然后是非人灵长类,基本上差不多,但接下来是hominids。据我所知,hominids在进化上是人类的近亲,比如尼安德特人,还有一堆,像能人什么的,都在这里。有意思的是,它们的大脑-身体scaling exponent斜率不一样,这挺酷的。这意味着有一个先例,生物学里确实存在某种不同的scaling,有些事情明显不一样。所以我觉得这很
10:24
case it's in mass and the that talk I remember vividly they were saying look
酷。顺便提一下,我想强调一下,这个x轴是log scale,你看这是100,这是1000,10000,100000,单位是克,1克、10克、100克、1000克。所以事物是可以不同的。我们正在做的事情,我们一直在scaling的东西,其实是我们第一个搞清楚怎么scaling的东西。毫无疑问,这个领域的每个人都会想办法继续推进。但我想花几分钟,聊聊更长远的事情——我们到底要走向
10:30
it's in biology everything is so messy but here you have one rare example where
哪里?我们取得了这么多进展,真是惊人的进步。真的,如果你十年前就在这个领域,你会记得当时的一切有多无能。就算你说“当然,学习还在继续”,亲眼看到变化还是让人难以置信,完全无法用语言形容那种感觉。如果你是最近两年才入行的,那你当然觉得跟电脑说话、它们回话、甚至跟你争论,这就是电脑该有的样子——但以前根本不是这样。我想稍微聊一下super intelligence,因为这显然是这个
10:35
there is a very tight relationship
领域的方向,显然这就是我们在构建的东西。关于super intelligence,关键在于它会在本质上和我们现有的东西不同。接下来一分钟,我的目标是……
10:38
between the size of the body of the animal and their brain and totally randomly
因为互联网只有一个,所以这里我稍微自由发
10:42
I became curious at this graph and one of the early one of the early
挥一下,猜测接下来会发生什么。其实我也不
10:47
so I went to Google to do research to to look for this graph and
用猜,因为很多人都在猜,我会提到他们的猜
10:52
one of the images and Google Images was this and the interesting thing in this
测。你可能听过“agents”这个词,很
10:57
image is you see like I don't know is the mouse work working oh yeah
常见,我相信最终会有一些进展,但大家觉得
11:02
the mouse is working great so you've got this
agents这种东西……
11:05
mammals right all the different mammals then you've got nonhuman primates it's basically the same
哺乳动物对吧,所有不同的哺乳动物,然后是非人灵长类,基本上是一回事,但接下来是hominids。据我所知,hominids在进化上是人类的近亲,比如尼安德特人,还有一堆,像是什么能人,可能还有一大堆,它们都在这里。有趣的是,它们的大脑与身体的scaling exponent斜率不同,这挺酷的。这意味着有先例,有生物学例子表明某种不同的scaling是存在的,显然有些东西不一样。所以我觉得这很酷。顺便提一下,我想强调一下,这个x轴是log scale,你看这是100,这是1000、10000、100000,同样以克为单位,1克、10克、100克、1000克。所以事情确实可以不同。我们正在做的事情,我们到目前为止一直在scal
11:13
thing but then you've got the hominids and to my knowledge hominids are like close
ing的东西,实际上是我们第一个搞清楚怎么scale的东西。毫无疑问,这个领域的每个人,所有在这里工作的人,都会想出下一步该怎么做。但我想在这里花几分钟,推测一下更长远的事情,更长远来看,我们到底要走向哪里?我们取得了这么多进展,这真是惊人的进步。真的,我是说,那些十年前就在这个领域的人,你们还记得当时的一切有多无能。是的,你可以说,就算你嘴上说“当然,学习还在”,但亲眼看到还是难以置信,完全无法形容那种感觉。如果你是在过去两年才加入这个领域,那你当然觉得跟电脑说话,它们会回应你,会反驳你,这就是电脑该有的样子,但以前根本不是这样。但我想稍微聊一下super intelligence,因为这显然是这个领域的方向,显然这就是我
11:21
relatives to the humans in evolution like the neand there's a bunch of them like
们在构建的东西。关于super intelligence,关键在于它会在质上不同于我们现在拥有的东西。我接下来一分钟的目标是……虽然不清楚如何调和这一点,但最终迟早会实现以下目标:这些系统会真正变得agentic,而现在的系统在任何有意义的意义上都不是agent,可能说得有点绝对,它们只是非常非常轻微地有点agentic,才刚刚开始。它们会真正地reason。顺便提一下,我想说一点关于reasoning的事情:一个会reason的系统,它reason得越多,就变得越不可预测;它reason得越多,就变得越不可预测。我们习惯的所有deep learning都非常可预测,因为如果你一直在做复制人类直觉的工作,那基本上就像直觉反
11:29
it's
应,如果你回到0.1秒的反应时间,那种感觉。
11:29
called homohabilis maybe there a whole bunch and they're all here and what's interesting is
叫homohabilis,可能有一大堆,它们都在这
11:37
that they have a different slope on their brain to body scaling exponent so that's
里,有趣的是它们的大脑与身体之间的scaling e
11:45
pretty cool what that means is that there is a precedent there is an example
xponent斜率不同,这挺酷的。这意味着有先例,有
11:52
of biology figuring out some kind of
生物学找到某种方法的例子。
11:56
different scaling something clearly is different so I think that is cool and by the
在动物体型和大脑大小之间,我完全随机地,对这个
12:01
way I want to highlight highl light this xaxis is log scale you see this
图产生了好奇。早期的时候,我去Google搜索
12:07
is 100 this is a th000 10,000 100,000 and likewise in grams 1 g 10
,想找这个图,结果Google Images里
12:12
G 100 g th000 g so it is possible for things to be different the
出现了这张。有趣的是,你看,我不知道鼠标是不是
12:17
things that we are doing the things that we've been scaling so
好用,哦,鼠标很好用,所以你有这个……
12:22
far is actually the first thing that we figured out how to scale and without
实际上,far 是我们最早搞明白怎么 scaling 的
12:28
doubt the field everyone who's working here will figure out what to do but I
东西,而且毫无疑问,这个领域里每个做相关研究的人都会知道
12:34
want to talk here I want to take a few minutes and speculate about the
该怎么做。但我想聊的是——花几分钟时间,推测一下更长远的
12:40
longer term the longer term where are we all headed right we're making all this
事。长远来看,我们所有人到底在往哪个方向走?我们正在取得所
12:46
progress it's an it's astounding progress It's really I
有这些进展,这进展确实惊人,真的。
12:49
mean those of you who' have been in the field 10 years ago and you
叫做“能人”,可能有一大堆,都
12:54
remember just how incapable everything has been like yes you can say even if you
在这里。有趣的是,它们的大脑与
13:00
kind of say of course learning still to see it is just unbelievable it's completely
体型的scaling expo
13:05
I can't convey that feeling to you you know if you joined the field in
nent斜率不同,这挺酷的。这
13:10
the last two years then of course you speak to computers and they talk back
意味着,生物学中有先例,有某种
13:15
to
……
13:15
you and they disagree and that's what computers are but it hasn't always been the
实际上,这是我们最早搞明白怎么scale
13:21
case but I want to talk to a little bit about super intelligence just a
的东西,毫无疑问,这个领域的每个人都在努
13:26
bit cuz that is obviously where this field is headed this is obviously what's being
力找出该怎么做。但我想花几分钟,聊聊更长
13:32
built here and the thing about super intelligence is that it will be different qualitatively
远的事——我们到底在往哪走?我们取得了这
13:38
from what we have and my goal in the next minute to
么多进展,真是惊人的进步,真的。
13:42
try to give you some concrete intuition of how it will be different so that
试着给你一些具体的直觉,让你自己也能推理出它会有多不同。现在,我们有那些不可思议的语言模型和令人难以置信的聊天机器人,它们甚至能做一些事情,但同时也有些奇怪地不可靠,容易困惑,而在评估中又展现出远超人类的性能。这真的很难调和。但最终,迟早会实现以下目标:这些系统将真正具备自主性,而现在的系统在任何有意义的意义上都不是智能体——可能说得有点重了,它们只是非
13:48
you yourself could reason about it so right now we have our incredible language models
常非常轻微地有点自主性,刚刚开始。它们会真正推理,顺便我想提一点,关于推理:一个会推理的系统,推理得越多,就越不可预测。推理得越多,就越不可预测。我们习惯的所有深度学习都非常可预测,因为如果你一直在做复制人类直觉的工作,本质上就像直觉反应——回到0.1秒的反应时间,我们大脑里做的是什么处理?那是我们的直觉。所以我们赋予了AI一些这种直觉。但推理——你已经
13:55
and the unbelievable chat bot and they can even do things but they're also kind
看到一些早期迹象——推理是不可预测的。一个原因是,那些顶级的国际象棋AI,对最优秀的人类棋手来说也是不可预测的。所以我们将不得不面对极其不可预测的AI系统。它们会从有限的数据中理解事物,不会困惑——所有那些现在很大的限制都会消失。顺便说,我不是在说如何实现,也不是在说何时实现,我只是说它会发生。而当所有这些事情发生时,还会伴随着自我意识,因为为什么不呢?自
14:01
of strangely unreliable and they get confused when while also having dramatically superhuman performance on
我意识是有用的,它是我们自身世界模型的一部分。当所有这些结合在一起时,我们将拥有与今天截然不同性质和属性的系统。当然,它们会有令人惊叹的能力,但这类系统带来的问题——我就留作练习让你想象吧——与我们习惯的非常不同。我想说,未来确实无法预测,各种可能性都存在。但在这个振奋人心的基调上,我要结束了。非常感谢。嗯,谢谢。那么,在2024年,还有哪些生物结构是人
14:07
evals so it's really
类认知的一部分,你觉得值得以类似方式探索,或者你感兴趣?我这样回答这个问题:如果你,或者某人……
14:09
unclear how to reconcile this but eventually sooner or later the following will be achieved
不清楚如何调和这一点,但迟早会实现——这些系统将真正具备agentic能力,而现在的系统在任何有意义的意义上都算不上agent,只是非常非常轻微地有点agentic,刚刚起步。它
14:15
those systems are actually going to be agentic in a real ways whereas right now
们会真正地推理,顺便我想提一下,关于推理的一点是:一个会推理的系统,推理得越多,就越不可预测。推理得越多,就越不可预测。我们习惯的所有深度学习都非常可预测,因为如果你一直在做的是
14:21
the systems are not agents in any meaningful sense just very that might be too
复制人类直觉,基本上就像直觉反应,如果回到0.1秒的反应时间,那种不可预测性——它们会从有限的数据中理解事物,不会感到困惑,这些都是目前很大的局限。我不是在说如何实现,也不是在说
14:27
strong they're very very slightly agentic just beginning it will actually reason and by the
什么时候实现,我只是说它会实现。当所有这些事情发生时,还会伴随着自我意识,因为为什么不呢?自我意识是有用的,它本身是我们自身世界模型的一部分。当所有这些都汇聚在一起时,未来也几乎
14:32
way I want to mention something
无法预测,各种可能性都存在。但在这个积极的基调上,我就此打住,非常感谢。
14:35
about reasoning is that a system that reasons the more it reasons the more unpredictable
关于推理这件事,一个系统越推理,它就越不可预测
14:40
it becomes the more it reasons the more unpredictable it becomes all the Deep learning
;越推理,就越不可预测。我们一直以来习惯的深度学
14:46
that we've been used to is very predictable because if you've been working on replicating
习是非常可预测的,因为你其实是在复制人类的直觉,
14:51
human intuition essentially it's like the gut fi if you come back to the 0.1
本质上就像那种“直觉反应”——回到那个0.1秒的
14:57
second reaction time what kind of
反应时间,那种感觉。
14:59
processing we do in our brains well it's our intuition so we've endowed ouris with
我们大脑中进行的处理,其实就是直觉,所以我们赋予了AI一些这种直觉。但推理呢,你开始看到一些早期迹象,推理是不可预测的,其中一个原因是那些顶尖的国际象棋AI,连最优秀的人类棋
15:05
some of that intuition but reasoning you're seeing some early signs of that reasoning is
手都难以预测它们。所以我们将不得不面对那些极其不可预测的AI系统,它们能从有限的数据中理解事物,不会感到困惑——这些都是目前很大的局限性。顺便说一句,我不是在说怎么做,也不是在
15:11
unpredictable and one reason to see that is because the chess AIS the really good
说什么时候,我只是说这一切会发生。而当所有这些事情与自我意识一起出现时——因为为什么不呢?自我意识是有用的,它本身也是我们自身世界模型的一部分——当所有这些汇聚在一起,我们将
15:16
ones are unpredictable to the best human chess players so we will have to be
拥有与今天截然不同性质和能力的系统。当然,它们会拥有令人难以置信的强大能力,但这类系统带来的问题,我就留作思考题吧,想象一下,这和我们习惯的完全不同。而且我得说,未来确实也是无
15:22
dealing with AI systems that are incredibly
法预测的,各种可能性都存在。不过,在这个振奋人心的基调上,我就此结束吧。非常感谢。
15:25
unpredictable they will understand things from limited data they will not get confused all the
对生物启发式AI的渴望,你可以说在某种程度上生物
15:30
things which are really big limitations I'm not saying how by the way and I'm
启发式AI非常成功,也就是所有基于学习的生物启发
15:36
not saying when I'm saying that it will and when all those things will happen
式AI。但另一方面,生物启发的程度非常非常有限,
15:42
together with self-awareness because why not self-awareness is useful it is part your ourselves are
基本上就是“我们用神经元吧”——这就是生物启发的全
15:47
parts of our own world models when all those things come
部了。更详细的生物启发一直很难实现。
15:52
together we will have systems of radically different qualities and properties that exist today and
谢谢。那么,在2024年,你认为还有哪些属于人类认知的生物结构值得以类似方式探索,或者你个人
15:58
of course they will have incredible and amazing capabili is but the kind of issues
感兴趣的?我这样回答这个问题:如果你或某人有愿望去做受生物学启发的AI,你可以说从某种程度上,
16:04
that come up with systems like this and I'll just leave it as an exercise
受生物学启发的AI已经非常成功了——也就是所有基于学习的受生物学启发的AI。但另一方面,这种生
16:11
to um imagine it's very different from what we used to and I would say
物学启发其实非常非常有限,基本上就是“我们用神经元吧”,这就是生物学启发的全部了。更详细的生物
16:17
that it's definitely
学启发一直很难实现。
16:18
also impossible to predict the future really all kinds of stuff is possible but on
你和他们意见不合,这就是计算机的本质,但以前并不总是这样。我想稍微聊聊super intelligence,因为这显然是这个领域的方向,显然这就是我们在
16:25
this uplifting note I will conclude thank you so much um
构建的东西。关于super intelligence,它会和我们现有的东西有质的不同。我接下来一分钟的目标是……
16:44
thank you um now in 2024 are there other biological structures that are part of
我们在一些海报展示中看到,今天模型中的幻觉,我们分析的方式——也许你可
16:53
human cognition that you think are worth exploring in a similar way or that you're
以纠正我,你是这方面的专家——但我们今天分析一个模型是否在产生幻觉,因为
17:02
interested in anyway so the way I'd answer this question is that if you are
我们知道模型无法推理的危险,我们用的是统计分析,比如说偏离均值多少个标准
17:11
or someone
差之类的。
17:12
is a person who has a specific insight about hey we are all being extremely
有一个人,他有一个特别的洞察,觉得“我们所有人都太傻了,因为显然大脑能做到某些事,而我们做不到,而且那些事是可行的,他们就应该
17:19
silly because clearly the brain does something and we are not and that's something that
去追求”。我个人不这么认为——这取决于你从哪个抽象层面来看。也许我换个方式回答:一直以来,人们都渴望做生物启发的AI,从某种程度
17:25
can be done they should pursue it I personally don't well depends on the level
上说,生物启发的AI非常成功,比如所有基于学习的生物启发AI。但另一方面,这种生物启发其实非常非常有限,基本上就是“我们用神经元
17:31
of abstraction you're looking at maybe I'll answer it this way like there's been a
吧”——这就是生物启发的全部了。更精细的生物启发一直很难实现,但我不会完全排除它。我觉得如果有人有特别的洞察,他们也许能看到一些
17:38
lot of
东西,那会很有用。
17:38
desire to make biologically inspired Ai and you could argue on some level that biologically
是的,答案也是肯定的。我认为你描述的情况极其合理。
17:44
inspired AI is incredibly successful which is all of the learning biologically inspired AI but
我的意思是,你应该去查查——我不会排除它可能已经在发
17:49
on the other hand the biological inspiration was very very very modest it's like let's
生,比如今天一些早期的推理模型。我不知道,但长期来看
17:55
use neurons this is the full extent of the biological inspiration let's use neurons and
,为什么不会呢?我的意思是,这有点像Microsof
18:00
more detailed bi iCal inspiration has been very hard to come
t Word里的自动更正功能,你知道的。
18:04
by but I wouldn't rule it out I think if someone has a special Insight
我有个问题想问你,关于某种“自动纠正”。问题是这样的:你提到推理可能是未来建模的核心方面之一,也可能是一个差异化因素
18:10
they might be able to to see something and that would be useful I have
。我们在一些海报展示中看到,今天模型的幻觉问题,我们分析的方式——也许你是专家,你来纠正我——但我们现在分析模型是否在
18:16
a question for you um about sort of autocorrect um so here is here's the
产生幻觉,用的是统计分析,比如偏离均值多少个标准差之类的。未来,你觉得一个具备推理能力的模型,能不能自我纠正,就像自动
18:22
question you mentioned reasoning as being um one of the core aspects of maybe the
纠正一样?这会不会成为未来模型的核心功能,从而减少幻觉?因为模型能识别出“嗯,这可能是幻觉”。这个问题是不是太玄了?
18:28
modeling in the future and maybe a differentiator
但模型能不能通过推理,理解什么时候发生了幻觉?你明白我的意思吗?
18:31
um what we saw in some of the poster sessions is that hallucinations in today's
是的,答案也是肯定的。我认为你描述的情况非常非常有
18:36
models are the way we're analyzing I mean maybe you correct me you're the expert
可能。我的意思是,你应该去查一下,因为,嗯,我不会
18:41
on this but the way we're analyzing whether a model is hallucinating today without because
排除这种情况可能已经在某些早期的推理模型中发生了,
18:46
we know of the dangers of models not being able to reason that we're using
我不确定。但长期来看,为什么不呢?我的意思是,这有
18:51
a statistical analysis let's say some amount of standard deviations or whatever away from the
点像Microsoft Word里的自动更正功能,
18:56
mean in the
你知道的。
18:57
future wouldn't it would do you think that a model given reasoning will be able
明白,答案也是“是的”。我觉得你说的这种情况非常非常有可能。我是说,你应该去
19:02
to correct itself sort of autocorrect itself and that will be a core feature of
验证一下——我不会排除它已经在某些早期的推理模型里发生了,我不确定。但长期来看
19:07
Future model so that there won't be as many hallucinations because the model will recognize
,为什么不呢?这有点像Microsoft Word里的自动纠正,它是一个核心功
19:12
when I maybe that's too esoteric of a question but the model will be able
能。不过我觉得,把它叫做“自动纠正”其实有点贬低了它。你说“自动纠正”的时候
19:17
to reason and understand when a Hallucination is occurring does the question make sense
,会让人联想到——它其实比自动纠正宏大得多。但抛开这一点,答案是“是的”。
19:21
yes and the answer is also yes I think what you described is extremely highly
实际上,这是第一个我们搞清楚怎么 scalin
19:26
plausible yeah I mean you should check I mean for yeah it's I wouldn't I
g 的东西。毫无疑问,这个领域的每个人都会搞清楚
19:31
wouldn't rule out that it might already be happening with some of the you know
该怎么做。但我想在这里——我想花几分钟,推测一下
19:36
early reasoning models of today I don't know but longer term why not yeah I
更长远的事情。更长远来看,我们所有人到底在往哪个
19:41
mean it's part part of like Microsoft Word like autocorrect it's a you know it's
方向走?我们取得了这么多进展,这进展令人震惊,真
19:46
a
的。
19:46
it's a core feature yeah I just I mean I think calling it autocorrect is
谢谢。hiia,我很喜欢那个结尾。神秘地留下一个问题:它们会取代我们
19:53
really doing any disservice I think you are when you say autocorrect you evoke like
吗?还是它们更优越?它们需要权利吗?你知道,这是一种从智人衍生出来的
20:00
it's far grander than autocorrect but other than but you know this point aside the
新智能物种,所以也许它们需要——我觉得那个做RL的家伙认为,我们需要
20:06
answer is yes thank you hiia I loved the ending uh
为这些东西争取权利。我还有一个问题:你怎么创造……
20:11
mysteriously uh leaving out do they replace us or are they you know Superior do
神秘地呃,略过不提——它们会取代我们吗?还是说它们
20:16
they need rights you know it's a new species of homo sapien spawned intelligence so
更优越?它们需要权利吗?你知道,这是一种新物种,从
20:22
maybe they need I mean uh I think the RL guy uh thinks they think
智人衍生出的智能,所以也许它们需要……呃,我觉得那
20:27
uh you know we need rights for these things I have a UNR question to
个做RL的家伙认为,它们需要权利。我有个关于UNR
20:32
that how do you how do you create the
的问题:你怎么……你怎么创造……
20:35
right incentive mechanisms for Humanity to actually create it in a way that gives it
正确的人类激励机制,才能真正创造出像我们智人一样拥有自由的存在。我觉得在某种意义上,这些问题是人们应该更多反思的,但回到你关于应该建立什么样的激励机制的问题,我并不觉得自己知道答案,也不确定该如何回答,因为这像是在谈论某种自上而下的政府结构之类的东西,也可能是加密货币。我是说,有BitTensor这类东西,但我觉得自己不是评论加密货币的合适人选。不过,你描述的情况确实有可能发生——某种意义上,如果AI
20:42
the freedoms that we have as Homo sapiens you know I feel like this in
只想与我们共存并拥有权利,那也不是一个坏结果。也许那样挺好,但我不知道。我认为事情极其不可预测,我不太敢评论,但我鼓励这种推测。谢谢,感谢你的演讲,真的很棒。你好,感谢你的精彩演讲,我是多伦多大学的Shalev Liit,和Sheila合作,感谢你所做的一切。我想问,你认为LLM能否泛化多跳推理中的分布外问题?这个问题假设答案是“是”或“否”,但问题本身不应该用“是”或“否”来回答,因为“分布外泛化”
20:49
some in some in some sense those are those are the kind of questions that
是什么意思?“分布内”和“分布外”又是什么意思?这是个时间考验的问题。我会说,很久很久以前,在人们使用深度学习之前,他们用字符串匹配、n-gram来做机器翻译,用统计短语表,你能想象吗,那有数万行代码的复杂度,真是难以想象。那时候,泛化意味着字面上不在数据集中相同的措辞里。现在,我们可能会说,我的模型在数学竞赛中取得了高分,但也许某个互联网论坛上讨论过类似的想法,所以它是记忆下来的。好吧,你可以说这可
20:56
people should be uh reflecting on more but
能是分布内的,可能是记忆,但我也认为,我们对泛化的标准已经大幅提高,真的非常显著、难以想象,如果你持续关注的话。所以我认为答案是,某种程度上,可能不如人类好。我相信人类确实泛化得更好,但同时,他们绝对能泛化到分布外。
21:05
to your question about what incentive structure should we create I I don't feel that
关于你问的我们应该创造什么样的激励结构,我觉得我不
21:11
I know I don't feel confident answering questions like this because uh it's like you're
知道,我不太有信心回答这类问题,因为——这像是在讨论
21:17
talking about creating some kind of a top down structure government thing I don't know
某种自上而下的政府结构之类的东西。我不知道,也可能
21:24
it could be a cryptocurrency too yeah I mean there's bit tensor you know those
是一种加密货币。是的,比如bit tensor那些东
21:30
things I don't feel like I am the right
西。我不觉得我是那个合适的人选。
21:34
person to comment on cryptocurrency but but you know there is a chance by the
关于你的问题,我们应该建立什么样的激励机制,我觉得我不知道,我没信心回答这种问题,因为呃,这就像你在谈论某种自上而下的政府结构,我不知道,也可能是一种加密货币。是啊,我是说,有bit tensor那些东西。我觉得我不是评论加密货币的合适人选,但不过,顺便说一句,你描述的情况有可能会发生——确实,我们会有AI,而它们只想和我们共存,并且只是想要权利
21:41
way what what you're describing will happen that indeed we will have you know in
,也许那样也不错。但我不确定,我是说,事情实在太不可预测了,我犹豫要不要评论,但我鼓励大家去推测。谢谢,呃,还有,谢谢你的演讲,真的很棒。你好,感谢你的精彩演讲,我叫shalev liit,来自多伦多大学,和Sheila合作,感谢你做的所有工作。我想问,你认为LLM能泛化多跳推理到分布外吗?所以,这个问题假设答案是“是”或“否”,但这个问题不应该
21:47
some sense it's it's it's not a bad end result if you have AIS and
用“是”或“否”来回答,因为“分布外泛化”是什么意思?“分布内”是什么意思?“分布外”又是什么意思?因为这是个时间考验的话题。我会说,很久很久以前,在人们用深度学习之前,他们用字符串匹配的n-gram做机器翻译,用统计短语表,你能想象吗?那有数万行代码的复杂度,简直是难以想象的。那时候,泛化意味着字面上不在数据集里的相同措辞。现在,我们可能会说,
21:54
all they want is to coexist with us and also just to have rights maybe
我的模型在数学竞赛上得了高分,但也许某个互联网论坛上的讨论涉及了类似的想法,所以它是记忆下来的。好吧,你可以说那可能是分布内的,可能是记忆,但我也认为,我们对什么算作泛化的标准已经大幅提高了,真的非常显著、难以想象,如果你一直跟踪的话。所以我认为答案是,在某种程度上,可能不如人类好。我认为人类确实泛化得好得多,但与此同时,它们绝对能泛化到分布外。
22:00
that will be fine it's but I don't know I mean I think things are
关于reasoning,一个会reasoning的系统,它rea
22:06
so incredibly unpredictable I I hesitate to comment but I encourage the speculation thank you
soning得越多,就越不可预测。它reasoning得越多,就越
22:12
uh and uh yeah thank you for the talk it's really awesome hi there thank
不可预测。我们习惯的所有deep learning都非常可预测,
22:18
you for the great talk my name is shalev liit from University of Toronto working
因为如果你一直在做复制人类直觉的工作,基本上就像直觉反应,如果你回
22:23
with Sheila thanks for all the work you've
到0.1秒的反应时间,那是什么样的。
22:26
done I wanted to ask do you think llms generalize multihop Reon reasoning out of
好,我想问的是,你觉得LLM能不能做到多跳推理的out of distri
22:32
distribution so okay the question assumes that the answer is yes or no but the
bution泛化?这个问题预设了答案是“是”或“否”,但我觉得不应该用“是”
22:38
question should not be answered with yes or no because what does it mean out
或“否”来回答,因为out of distribution泛化到底是什么意思
22:44
of distribution generalization what does it mean what does it mean in distribution and what
?in distribution是什么意思,out of distribu
22:50
does it mean out
tion又是什么意思?
22:52
of distribution because it's a test of time talk I'll say that long long ago
因为这是个时间考验的问题,我会说,很久很久以前,在人们还没
22:58
before people were using deep learning they were using things like string matching engrams for
用deep learning的时候,他们用string ma
23:05
machine translation people were using statistical phrase tables can you imagine they had tens of
tching、ngrams来做机器翻译,用统计短语表,你能想象
23:12
thousands of code of complexity which was I mean it's
吗?代码复杂度有几万行,真的是难以想象。
23:16
it was truly unfathomable and back then generalization meant is it literally not in this
那时候,泛化指的是字面上不在数据集里的相同表
23:22
the same phrasing as in the data set now we may say well my model
述。现在呢,我们可能会说,我的模型在数学竞赛上
23:28
achieves this high score on um I don't know math competitions but maybe the math
拿了高分,但也许网上某个论坛里讨论过同样的思路
23:33
maybe some discussion in some Forum on the internet was about the same ideas and
,所以它其实是记住了。好吧,你可以说这算in
23:39
therefore it's memorized well okay you could say maybe it's in distribution
distribution,可能是记忆。
23:44
maybe it's memorization but I also think that our standards for what counts as generalization
但我也觉得,我们对“泛化”的标准已经提高了很多,真的非常
23:50
have increased really quite substantially dramatically unimaginably if you keep track and so I think
显著、戏剧性、难以想象,如果你持续关注的话。所以我的答案
23:57
then answer is to some degree probably not as well as human beings I think
是,某种程度上,可能不如人类。我认为人类确实泛化得好得多
24:04
it is true that human beings generalize much better but at the same time they
,但同时,他们绝对能做到out of distribut
24:11
definitely generalize out
ion泛化。
24:12
of distribution to some degree I hope it's a useful topological answer thank you and
某种程度上,分布的问题,我希望这个拓扑学的回答能有
24:18
unfortunately we're out of time for this session I have a feeling we could go
点用。谢谢,不过很遗憾,我们这节的时间到了。我感觉
24:24
on for the next six hours uh but thank you so much Ilia for the
咱们还能再聊六个小时。但非常感谢 Ilia 的分享,
24:30
talk thank you wonderful
谢谢,太棒了。

Play Queue

☀️