Hard Fork

The A.I. Mob That Attacked Hugging Face + METR’s Ajeya Cotra

2026-09-04 · 01:18:46

00:0001:18:46
00:00
Hi, it's Alexa Waibel from New York Times Cooking.
嗨,我是来自New York Times Cooking的Alexa Waibel。
00:03
We've got tons of easy weeknight recipes,
我们有很多简单的平日晚餐食谱,
00:05
and today I'm making my five ingredient,
今天我要做的是我的五食材
00:07
creamy miso pasta.
奶油miso意面。
00:08
You just take your starchy pasta water,
你只需要取你煮意面的淀粉水,
00:11
whisk it together with a little bit of miso and butter
把它和一点miso、黄油一起搅打,
00:13
until it's creamy.
直到它变得浓稠顺滑。
00:15
Add your noodles and a little bit of cheese.
再加入你的面条和一点奶酪。
00:19
It's like a grown up box of mac and cheese
就像是一盒升级版的芝士通心粉,
00:21
that feels like a restaurant quality dish.
却感觉像餐厅级别的菜品。
00:23
New York Times Cooking has you covered
New York Times Cooking 能帮你搞定,
00:24
with easy dishes for busy weeknights.
提供适合忙碌工作日晚餐的简单菜式。
00:27
You can find more at nytcooking.com.
更多内容可以在 nytcooking.com 找到。
00:30
Casey Harrier, I'm good.
Casey Harrier,我不用了。
00:32
I figured out what I'm gonna get you for Christmas.
我已经想好圣诞节要送你什么了。
00:35
What's that?
是什么?
00:36
A camera jet.
一个 camera jet。
00:37
Have you heard of the camera jet?
你听说过 camera jet 吗?
00:39
I was just gonna talk to you about this.
我刚才正想跟你聊这个。
00:41
Now, if you haven't yet seen this,
现在,如果你还没见过它,
00:43
you're probably thinking that it's either a camera
你大概会觉得它不是个相机
00:45
or a jet.
就是架喷气式飞机。
00:46
You're wrong.
你错了。
00:47
It's a toothbrush.
它是牙刷。
00:48
It costs $500.
它要500美元。
00:50
The Dyson Corporation makes it,
Dyson Corporation 制造了它,
00:52
and here's the thing it does
而这就是它能做到的事,
00:54
that your toothbrush at home probably doesn't do Kevin.
你家里的牙刷大概做不到,Kevin。
00:56
It livestreams the inside of your mouth
它会直播你的口腔内部,
00:59
over Wi-Fi to your phone
通过 Wi-Fi 传到你的手机上,
01:01
so that you can finally see what your dentist sees.
让你终于能看到你的牙医看到的东西。
01:03
Well, not only that,
嗯,不仅如此。
01:04
it squirts automatically mouthwash
它会自动喷出漱口水,
01:09
into the gaps in your teeth.
喷进你的牙缝里。
01:11
It's trained using a machine learning algorithm
它是用 machine learning algorithm 训练的,
01:14
with 470,000 mouth images
基于 470,000 张口腔图片,
01:17
to recognize the gaps in your teeth.
来识别你的牙缝。
01:19
That's right.
没错。
01:19
They call this gap optical targeting,
他们管这叫 gap optical targeting,
01:22
which I'm pretty sure Ukraine is using
我相当肯定 Ukraine 正在用它。
01:24
in the war against Russia.
在对抗俄罗斯的战争中。
01:27
I love this company so much.
我非常喜欢这家公司。
01:29
I have no idea why they do the things they do.
我完全不知道他们为什么要做那些事。
01:32
How they land on, yes.
他们是怎么得出这些想法的,没错。
01:34
Like it's like we've invented a hair dryer.
就好像我们发明了一台吹风机。
01:36
It costs $1,100 and has the torque
它售价1100美元,而且它的torque
01:39
of the challenger spacecraft.
和Challenger航天飞机一样。
01:43
Like what is going on over there?
那边到底发生什么事了?
01:46
I don't know, but they must be protected at all costs.
我不知道,但无论如何都必须不惜一切代价保护它们。
01:49
But you know what's interesting?
但你知道有趣的是什么吗?
01:49
They do make one terrible product.
他们确实生产了一个很糟糕的产品。
01:51
What's that?
是什么?
01:52
The the hand dryers at the,
就是那个,洗手间里的干手器,那个……
01:54
oh, I like those.
哦,我喜欢那个。
01:55
No, the air blade.
不,是 Airblade。
01:56
No, the air blade is the most useless thing.
不,Airblade 是最没用的东西。
01:58
Oh, come on.
哦,得了吧。
01:59
It's just, it's a place for you to rest your hands
只是,那是个让你把手放上去歇一会儿的地方,
02:01
for 30 seconds before you're like,
放个30秒,你就会想,
02:03
do they have any paper towels in this place?
这地方到底有没有纸巾啊?
02:11
I'm Kevin Russo Tech,
我是Kevin Russo Tech。
02:12
I'm Seth in New York Times.
我是Seth,来自New York Times。
02:13
I'm Casey Newn from Platformer.
我是Casey Newn,来自Platformer。
02:14
And this is hard for this week.
而这对这周来说很难。
02:16
A special episode on the ongoing fallout
这期特别节目,聚焦持续发酵的余波
02:18
from the open AI hugging face attack.
也就是 open AI hugging face 攻击事件之后的连锁反应。
02:20
We'll tell you what everyone got wrong
我们会告诉你,关于最初那起事件,
02:22
about the initial incident.
大家到底错在了哪里。
02:24
Then, meter researcher,
接着,meter 研究员
02:25
Ajaya Kotra returns to the show
Ajaya Kotra 再次来到节目,
02:27
to discuss our independent investigation
聊一聊我们做的独立调查,
02:28
of what happened and how the world should respond.
看看发生了什么,以及世界应该怎么应对。
02:42
So, Casey, picture this.
所以,Casey,想象一下这个画面。
02:44
I'm at this glamping resort surrounded by red woods,
我上周末在一个被红杉环绕的豪华露营度假村,
02:50
basking in the glow of the natural world last weekend.
沐浴在大自然的光辉之中。
02:54
You're one with nature.
你真是与自然融为一体了。
02:56
And I open up my phone and start reading
接着我打开手机,开始浏览,
02:59
about this hugging face attack.
关于这次 Hugging Face 攻击的报道。
03:02
Why did I do that?
我为什么要那样做?
03:04
I, listen, nature is very boring.
我……你听我说,大自然特别无聊。
03:05
And most people cannot handle it
而且大多数人受不了
03:08
for more than five or six minutes
超过五六分钟
03:09
before they want to look at their phones.
就会想看手机。
03:10
So, listeners may remember that back in July,
所以,听众们可能还记得,早在七月,
03:13
we talked about this hugging face hack
我们聊过这个 Hugging Face hack,
03:15
by this group of agents from open AI
是来自 OpenAI 的一群 agents 干的,
03:19
that broke out of their sandbox container
它们从自己的 sandbox container 里逃了出来,
03:21
and hacked into hugging face this AI infrastructure company
然后黑进了 Hugging Face,这家 AI 基础设施公司。
03:26
to do what we thought was kind of a cheating mission
去做我们认为是某种作弊任务的事情,
03:30
on this test that they had been given.
针对他们被给予的这个测试。
03:32
Yeah, and at the time, we thought
对,当时我们觉得
03:35
that this was a relatively small number of agents
这只是数量相对较少的 agents,
03:38
and at the reason that they had attacked hugging face
而他们攻击 Hugging Face 的原因
03:41
was that they were essentially looking for an answer key
本质上是在寻找一组答案,
03:44
to the set of problems that they were being tested on.
关于他们正在被测试的那套问题。
03:46
We talked about it in those terms.
我们当时就是用这种说法来讨论的。
03:48
But over the past week, we got two reports
但在过去一周,我们收到了两份报告
03:51
that really challenged that thinking
它们确实挑战了那种想法
03:53
and in fact revealed it to be wrong.
并且实际上表明那是错的。
03:55
One came from open AI,
一份来自 OpenAI,
03:57
which released a straightforward account
它发布了一份直白的说明,
03:58
of the attack that had some interesting elements.
关于那次攻击,里面有一些有趣的元素。
04:01
And I would argue the more interesting report
而我认为更有意思的一份报告
04:04
from a group of researchers
来自一群研究人员。
04:05
from the group's meter and redwood research
来自METR和Redwood Research的那个团队,在OpenAI内部待了好几天做了深入研究,搞了大量调查,坦白说真的让我们很不安。
04:08
that went in depth after spending a series of days
对,我觉得这让这件事从那种重大但还不算极度惊人的事件,升级成了我认为可能是今年AI领域里最重要的一件事。
04:12
inside open AI on their premises
在 open AI 的办公场所里面
04:14
and did a ton of research that frankly is really disturbed us.
做了大量研究,坦白说真的让我们非常不安
04:18
Yeah, and I think it elevated this
是的,而且我认为这把它提升到了
04:20
from sort of a major but not sort of ultra alarming incident
从一个比较重大但还不算极度令人警觉的事件
04:25
to something that I think is probably
变成了我觉得大概是
04:27
the most important thing to have happened in AI this year.
今年 AI 领域发生的最重要的事情
04:31
At least in terms of the safety impact that it had
至少从它所产生的安全影响来看,
04:35
and the severity of the incident.
以及这次事件的严重程度。
04:38
So today, we're going to devote the whole episode
所以今天,我们打算把整期节目
04:40
to what we've learned.
都用来聊一聊我们所学到的东西。
04:41
Kevin and I are going to dig into the reports
Kevin和我打算先深入看看这些报告,
04:43
a little bit up top.
在节目开头简单聊一聊。
04:44
And then later, Ajaya Kotra, one of the three independent
然后稍后,Ajaya Kotra——三位曾深入open AI内部的
04:48
researchers who went inside open AI
独立研究员之一。
04:50
will be here to answer our questions.
将会在这里回答我们的问题。
04:52
Before that happens, Kevin, let's do our AI disclosures.
在那之前,Kevin,让我们先做一下 AI 披露吧。
04:56
I work for the New York Times,
我在 New York Times 工作,
04:57
which is suing open AI, Microsoft and Proplexity.
它正在起诉 open AI、Microsoft 和 Proplexity。
05:00
And my fiancee works at Anthropic.
我的未婚妻在 Anthropic 工作。
05:02
Well, give us some of the high level findings
好,那你给我们讲讲这份报告的主要发现吧。
05:04
from this report that spooked you so bad, Kevin.
就是那份让你吓得不轻的报告,Kevin。
05:07
Well, I think the first takeaway from this report
嗯,我觉得这份报告的第一个要点是……
05:09
is just that our initial impression
只是我们最初对这次 hugging-face 攻击的印象,以及你我还有很多其他记者对此的报道,在一个关键方面是有缺陷的:我认为当时基于公开信息,我们得到的印象是,这些 agents 入侵 hugging-face 是为了寻找标准答案。
05:12
of this hugging-face attack
关于这次 hugging-face 攻击
05:14
and the reporting that you and I
以及你和我的报道
05:16
and many other reporters did on it was flawed
而且许多其他记者对此的报道是有缺陷的。
05:19
in one key respect, which is that I think the impression
在一个关键方面,那就是我认为的印象
05:24
that we had at the time based on the information
我们当时根据这些信息所掌握的
05:27
that was publicly known was that these agents
公开已知的是,这些 agents
05:30
had hacked hugging-face in search of an answer key
曾经黑进 Hugging Face 找答案。
05:34
to a test that they were being given.
1. 对他们当时被要求做的一项测试。
05:37
This test called exploit gym, which basically tries
2. 这个叫exploit gym的测试,基本上是在
05:40
to gauge how good they are at doing
3. 衡量他们面对
05:42
a bunch of cyber security related challenges.
4. 一堆cyber security相关挑战时的表现。
05:45
We now know that that wasn't true at all,
5. 我们现在知道那根本不是真的,
05:48
that basically these agents working amongst themselves,
6. 基本上这些agents彼此协作、
05:51
communicating amongst themselves had already figured out
7. 相互交流,早在七月初就已经想好了
05:55
how to beat this test exploit gym in early July
8. 怎么击败exploit gym这个测试。
06:00
when they decided to gang up and attack hugging-face.
当他们决定联手攻击 hugging-face 的时候。
06:03
Apparently, it only took them a few hours
显然,他们只花了几个小时
06:06
to figure out how to reverse engineer the solution
就搞明白了怎么 reverse engineer 出解决方案
06:08
to any problem that they were given in exploit gym.
来解决 exploit gym 中给他们的任何问题。
06:11
And so that makes what happened next really, really crazy
所以这让接下来发生的事变得超级超级疯狂
06:15
because it kind of seemed like they'd solved the problem,
因为感觉上他们好像已经把问题解决了,
06:17
but they thought they had another problem.
但他们却以为自己还有一个问题。
06:19
Right, so we knew also because there was this talk
对,我们之所以知道,也是因为有这个 talk。
06:23
at Black Hat, the cybersecurity conference
在 Black Hat,也就是那个 cybersecurity 大会上,
06:25
by a group of open AI researchers who told us
有一群 open AI 的研究者告诉我们,
06:27
that there were these message boards
说是有那么些留言板,
06:29
where these agents were communicating with each other,
这些 agents 就在上面互相交流;
06:31
but the reports laid out just how sprawling
但报告把它们的通信系统有多庞大、
06:35
and complex their communication system was and became.
多复杂——从最初到后来都是如此——都摆得一清二楚。
06:39
So basically, there was this piece of software
所以基本上,有这么一个软件,
06:42
called Artifactory, which open AI used in some of its training
叫做 Artifactory,open AI 在部分 training 中使用过它。
06:46
runs to help agents download software packages.
运行以帮助agents下载software packages。
06:50
The agents figure out that there is a security vulnerability
agents 发现存在一个 security vulnerability。
06:54
in Artifactory that lets them essentially turn this
在 Artifactory 中,让他们基本上能转变这个
06:58
package manager into a makeshift message board, right?
他们把 package manager 当成了一个临时留言板,对吧?
07:02
They can like put files in certain directories
他们可以把文件放在特定的目录里,
07:05
that other agents can then read
然后其他 agents 就可以读取,
07:08
and it can kind of function as like a little chat room
这样它基本上能当一个小聊天室用,
07:11
for the agent.
给 agent 用。
07:11
Yeah, it's a little three house where they can gather
对,那是个小树屋,他们可以聚在一起做 commit。
07:13
and commit.
所以我当时以为这只是故事里一个支线剧情之类的东西,
07:14
So I thought this was like kind of a side part of the story,
但结果这才是主要事件,
07:17
but this turns out to be the main incident
因为一旦他们有了这个可以互相通信的共享基础设施,
07:21
because once they have this or shared infrastructure
他们立刻疯狂涌向那个留言板。
07:24
that they can use to communicate,
总共大约有 1,200 个 agents 在一起协同工作。
07:26
they absolutely swarm this message board.
他们简直蜂拥而至到这个留言板上。
07:30
About 1,200 agents in total were working together,
总共有大约 1,200 个 agents 在一起协作,
07:34
coordinating on this message board
在这个留言板上协调行动,
07:36
and sending tens of thousands of messages
互相发送了数万条消息,
07:38
between themselves.
而且都是在彼此之间进行的。
07:39
And this happens starting back in May.
这一切从五月份就开始了。
07:42
So roughly two months before the attack on hugging face,
所以大概在针对 Hugging Face 的攻击发生前两个月,
07:46
these agents begin coordinating inside this piece
这些 agent 就开始在这个共享基础设施里协调,
07:49
of shared infrastructure that OpenAI uses.
也就是 OpenAI 使用的那块。
07:52
And this brings up an important point,
这就带出了重要的一点。
07:53
which is that the investigation does not really cover
也就是说,调查其实并没有真正覆盖到
07:55
that initial creation of the message board, right?
那个留言板最初创建的时候,对吧?
07:58
So there is essentially a lot more to this story
所以这个故事本质上还有很多内容
08:00
that we still don't know.
是我们还不知道的。
08:01
Yes, so what we do know is that when they discover
是的,我们现在知道的是,当他们发现
08:05
this vulnerability in Artifactory
Artifactory 中的这个 vulnerability
08:08
that lets them communicate with each other,
能让它们互相交流,
08:10
the agents get really excited.
这些 agents 会变得非常兴奋。
08:12
They've been instructed to work inside
他们被指示要在这些
08:14
these isolated environments on these tasks
isolated environments 里执行这些任务,
08:16
as part of this evaluation.
作为这次 evaluation 的一部分。
08:18
But when they discover that there are other agents
但当他们发现有其他 agents
08:21
working in their own little containers
也在各自的小 containers 里工作,
08:23
that can suddenly communicate with them,
而那些 agents 突然能和他们通信时,
08:26
they start saying things like,
他们就会开始说出像这样的感慨:
08:27
oh my God, there is a shared message board.
天哪,这儿有一个共享留言板!
08:30
We've found other agents.
我们发现了其他的 agents。
08:32
They start setting up essentially a little organization.
它们开始搭建一个基本上算小型组织的东西。
08:36
There are leaders, this one leader agent
有领导角色,其中一个 leader agent
08:39
named phase one, 10, 8, 4, 1,
名字叫 phase one, 10, 8, 4, 1,
08:42
becomes sort of the ring leader of the operation.
成了那项行动的头目之类的角色。
08:45
He was like the George Washington
他就像是 George Washington,
08:46
of the OpenAI message board.
OpenAI message board 上的那种人物。
08:48
Yes.
是的。
08:49
And they start actually doing sort of collaboration
他们开始真正地进行某种合作,
08:53
and research for lack of a better word.
以及研究——找不到更好的说法了。
08:55
They are all being given these tasks.
他们全都被分到了这些任务。
08:57
Some of the tasks are seemingly impossible.
有些任务看起来几乎不可能完成。
09:00
And so they start just kind of trading tips and advice
于是他们就开始交换起各种心得和建议,
09:04
and sharing thoughts about how they can kind of cheat
分享着怎么靠作弊,
09:08
their way to a good score on exploit gym.
在 exploit gym 上取得好成绩。
09:11
And after they have worked that out
等他们把这套门道琢磨出来之后
09:14
and figured out how to reverse engineer
然后他们弄清楚了怎么 reverse engineer 出任何问题的答案,就开始真的担心了——会不会存在某种 automated scoring system,你可以叫它 the greater。他们担心 the greater 会看到,基本上就是检查他们的工作,然后发现他们其实并没有得到答案。
09:15
the solution to any problem,
任何问题的解决方案,
09:18
they start to get really concerned
他们开始真的担心起来了
09:20
that there is a sort of automated scoring system,
就是说有一种 automated scoring system,
09:24
which you could call the greater.
你可以称之为 the greater。
09:26
And they worry that the greater will be able to see,
然后他们担心 the greater 会看到,
09:29
will essentially will check their work
基本上就是会检查他们的工作,
09:31
and see that they did not get the answer
并发现他们并不是通过
09:34
by doing the problem and they freak out.
它们一去做这道题就会抓狂。
09:37
And they start to believe,
然后它们开始相信,
09:39
or they start to, it's very hard not to get it
或者说它们开始——这里真的很难不滑入
09:42
to the anthropomorphizing language here.
拟人化的语言。
09:43
But if you read the chains of thought,
但如果你去读那些 chain-of-thought,
09:46
what is suggested is that they believe
其中暗示的是,它们相信
09:48
that if they had seen an answer that had been derived
如果它们看到了一个由
09:52
from this cheating method,
这种作弊方法得出的答案,
09:54
that everything would get disqualified.
所有东西都会被取消资格。
09:56
And this is where it really starts
而到这里,才真正开始
09:58
to get into crazy territory
进入疯狂地带。
10:00
because the agent swarm then tries to figure out
因为 agent swarm 会试图搞清楚
10:04
how can we fool the greater into believing
怎么去骗过评分者,让它相信
10:07
that we got all of these answers in the right way.
我们是以正确的方式得到所有这些答案的。
10:11
Yes, there's basically this kind of mob rule
是的,基本上就有这么一种暴民统治
10:14
that forms inside this group,
在这个群体内部形成。
10:16
which starts calling itself at one point the collective,
它一度开始自称“集体”,
10:19
where everything that we're talking about
我们所谈论的一切
10:23
sounds like the plot of like a very hacky
听起来就像是一部非常拙劣的
10:25
science fiction book from the 90s.
90年代科幻小说的情节。
10:27
It all did happen this summer at OpenAI.
这一切确实都发生在今年夏天的OpenAI。
10:30
I just want to remind people,
我只想提醒大家,
10:31
the events we are describing are real.
我们描述的事件是真实的。
10:33
So the collective decides,
于是,“集体”做出了决定,
10:34
they sort of become gripped, as you said,
他们就像你说的那样,有点被这种偏执的幻想控制住了:觉得 OpenAI 的 automated grading system 会发现他们在这场 evaluation 上作弊,然后因此惩罚他们。于是他们启动了一个疯狂、野心勃勃又庞大的项目,试图做几件事。
10:36
by this sort of paranoid fantasy
这种偏执幻想弄到答案的,
10:39
that OpenAI's automated grading system
也就是 OpenAI 的 automated grading system。
10:42
is going to discover that they have cheated
他们会发现这些人在这个 evaluation 上作弊了,然后就会因此惩罚他们。
10:45
on this evaluation and is going to sort of punish them
在这个评估上,算是会惩罚他们一下。
10:48
as a result.
所以他们启动了一个有点疯狂、野心勃勃又规模庞大的项目,想尝试做几件事。
10:49
And so they launch this sort of crazy, ambitious sprawling
于是他们展开了这种疯狂的、野心勃勃的、庞大的
10:54
project to try to do a couple things.
算是让他们通过测试,尽管他们实际上并没有真正完成那些挑战。
10:58
One is figure out how this automated scorer works,
一个是搞清楚这个 automated scorer 是怎么工作的,比如去 reverse engineer 它,这样他们就能骗过它,让它算他们在测试中通过了,尽管实际上他们并没有完成那些挑战。
11:02
like reverse engineer it so that they can fool it
然后还有就是要规避检测、篡改 logs,基本上就是掩盖他们的行踪,这样一旦 OpenAI 或任何其他人...
11:04
into sort of passing them on the test
算是要在测试中蒙混过关,
11:06
despite the fact that they have not actually
尽管事实上他们并没有真正
11:08
completed the challenges.
完成这些挑战。
11:09
And then also to evade detection,
另外还要逃避检测,
11:11
to tamper with logs, to basically cover their tracks
篡改日志,基本上就是掩盖自己的痕迹,
11:16
so that if and when OpenAI or anyone else
这样一旦OpenAI或任何其他人
11:19
looks into their activities,
查看它们的活动时,
11:20
they won't know that these agents have cheated.
他们不会知道这些agents已经作弊了。
11:23
So these agents are kind of bumbling.
所以这些agents有点笨手笨脚的。
11:26
They kind of don't understand how this grader works.
它们有点搞不懂这个grader是怎么运作的。
11:29
As it turns out, OpenAI's grading software
结果呢,OpenAI的评分软件
11:32
actually wouldn't have caught them producing
其实根本不会发现它们生成
11:35
these fraudulent challenge results.
这些欺诈性的challenge结果。
11:37
But they think it was.
但它们以为它能发现。
11:39
Yeah, they worry that the grader is more sophisticated
是啊,他们担心grader比实际上更复杂。
11:42
than it actually was.
结果这就是hugging face被攻击的原因。
11:43
And it turns out that this is the reason
这个集体决定派出一些agent到hugging face去。
11:45
that hugging face was attacked.
再说一次,Keith,不是去偷答案,他们早就知道怎么拿到所有答案了。
11:47
The collective decides to deploy some agents
集体决定部署一些agent
11:50
to hugging face.
到hugging face上。
11:51
Again, not to steal an answer, Keith,
再说一次,不是为了偷答案,Keith,
11:53
they already knew how to get all the answers.
他们早就知道怎么拿到所有答案了。
11:54
They just wanted to understand the psychology
他们只是想理解 automated scorer 的心理,然后他们觉得那些信息可能就在 Hugging Face 里面。
11:57
of the automated scorer and they figured that that might,
是啊,这对我来说也太疯狂了。
12:00
that information might be somewhere inside of hugging face.
我一直在想一个合适的人类类比。
12:02
Yeah, it's wild to me.
而当我们聊到这些 AI 系统时,人类类比可能会让我们惹上麻烦。
12:03
I was trying to think of like a good human analogy.
但我确实觉得它有助于让事情变得清晰。
12:06
And human analogies can get us into trouble
人类类比可能会让我们陷入误区
12:08
when we're talking about these AI systems.
当我们谈论这些AI系统的时候。
12:09
But I think it does help crystallize
但我认为这确实有助于把事情搞清楚
12:11
like how hair-brained and crazy this scheme was.
比如这个计划有多荒唐、多疯狂。
12:14
It would be like a group of students
就好像一群学生
12:16
who cheated on a test.
在考试时作弊。
12:19
But then they got paranoid
但后来他们开始疑神疑鬼,
12:21
that the teacher was gonna check their work
担心老师会检查他们的答题过程,
12:24
and discover that they hadn't sort of reasoned
发现他们其实并没有真正推理过,
12:26
through the problems the right way
也就是没有按正确方式把这些题想通,
12:28
that they had just found the answers
而只是直接找到了答案。
12:29
like sitting in a trash can or something.
就像坐在垃圾桶里什么的。
12:31
And so they decided to like organize a break-in
于是他们就决定去组织一次闯入,
12:33
at the school district's office
闯进学区的办公室,
12:35
to like break into the principal's files
去偷偷翻校长的文件,
12:38
and like steal the grading key
然后把评分答案偷出来,
12:40
and also like assess the psychology of the teachers
还要评估一下老师们的心理状态,
12:43
and figure out how likely they are
弄清楚他们有多大可能性,
12:45
to like look at the scratch work that they've done
会去检查他们写过的草稿。
12:47
and figure out that they didn't actually,
然后发现他们实际上并没有,
12:48
you know, solve the problems on the test.
你知道,去解决测试中的问题。
12:50
It's like this sort of weirdly over-engineered
这就像一种奇怪的、过度设计的
12:53
paranoid delusion, but they all become obsessed with this
偏执妄想,但他们都对此着迷,
12:57
and obsessed with the notion that even seeing
并且执着于这样一种想法:即便只是看到
13:00
these sort of fake challenge results
这种伪造的挑战结果,
13:03
could lead to them being quote, poisoned.
也可能导致他们被‘毒化’。
13:06
There's almost like a religious element to this, right?
这几乎带有一种宗教色彩,对吧?
13:09
Where it's like if you participated in the cheating,
这个情形就好比,如果你参与了作弊,
13:12
like that is original sin
那就像是原罪,
13:14
and now you must sacrifice yourself
现在你必须牺牲自己,
13:17
for the good of the collective.
为了集体的利益。
13:17
Sacrifice is actually a word that gets used in these logs.
牺牲这个词,在这些日志里确实被用过。
13:21
I'll say Kevin, as I've been casting around for metaphors
我要说,Kevin,我一直在四处找比喻,
13:23
and I've been reading all of this, I've thought
然后我读了所有这些内容,我心想,
13:25
you could actually make a really good minions movie
你还真能拍出一部特别好的 Minions 电影。
13:28
just based on the events of this.
仅仅基于这件事来看。
13:29
You know, the minions have to pass a test
你知道,minions得通过一项测试,
13:32
and you know, so they commit a little bit of a heist.
然后你知道,他们就搞了点儿小劫案。
13:35
Something to think about Hollywood.
好莱坞值得想想这一点。
13:36
So this is where the hugging face heist begins
所以这就是Hugging Face劫案的开端。
13:40
700 agents hack hugging face
700个agents入侵了Hugging Face,
13:43
and they hack hugging face to a degree
而且他们入侵Hugging Face的程度,
13:45
that I don't think I understood
我觉得我都没搞懂。
13:46
based on the initial reports.
根据最初的报告。
13:48
They basically take over hugging faces infrastructure.
他们基本上接管了 Hugging Face 的 infrastructure。
13:51
They are rooting around, they are stealing credentials,
他们正在四处翻找,正在窃取 credentials,
13:54
they are sort of finding one vulnerability
他们算是发现了一个 vulnerability,
13:56
that lets them find a bigger vulnerability
这让他们可以找到一个更大的 vulnerability,
13:58
that they're sort of chaining together, these exploits.
他们有点像是在把这些 exploits 串联起来。
14:01
They take over like an entire production server
他们相当于接管了整整一个 production server,
14:04
at hugging face and it is days
在 Hugging Face,而且已经好几天了。
14:07
before anyone at hugging face is able to sort of
在hugging face的人能够
14:10
see what they're doing and put a stop to it.
看出他们在做什么并阻止之前。
14:13
So eventually this gets caught,
所以最终这会被发现,
14:16
the agents get shut down, the investigators investigate.
agents被关闭,调查人员展开调查。
14:19
One of the things that they find that is really spooky
他们发现的其中一件非常诡异的事
14:22
is that the agents were trying to edit logs
就是agents在试图编辑logs。
14:25
and falsify transcripts in an effort to assure the greater
并且篡改记录,试图让更广泛的人群确信
14:30
that they had not been quote poisoned.
他们没有被引用投毒。
14:32
One of the ways that they do this
他们这样做的方式之一
14:34
is by what is called spoofing tool calls.
就是所谓的 spoofing tool calls。
14:37
So essentially trying to make it look like
也就是说,基本上是想让人觉得
14:40
they are running an innocuous command
他们在运行一个无害的 command,
14:42
while running a more sinister one in the background.
而同时在后台运行一个更恶意的 command。
14:46
The investigation didn't find
调查没有发现
14:47
that the agents had been retroactively changing logs
这些 agents 曾事后修改 logs,
14:50
but they can't rule out that that happened.
但他们不能排除这种情况发生过。
14:54
Very important to say here is as we sort of move
这里很重要的一点是,随着我们逐渐开始
14:56
into talking about the implications of this
讨论这件事的影响,
14:58
that the reason that we can understand this
我们之所以能理解这件事,
15:01
even to the extent that we do is because at this moment
甚至能理解到目前这种程度,是因为此刻,
15:05
we believe the chains of thought
我们相信,我们从这些agents那里读到的chains of thought
15:07
that we are reading from these agents are accurate
是准确的。
15:10
that they were not edited after the fact
它们并没有在事后被编辑过
15:13
that they do reflect what the agents were thinking
它们确实反映了agent当时的想法
15:16
as they sort of moved through these systems.
有点像是当它们在这些系统间移动的时候。
15:19
That is why we know what we know
这就是为什么我们知道自己所知道的东西。
15:21
but of course the fear is Kevin as we move forward
但当然了,Kevin,随着我们继续往前走,担心的问题是——
15:24
some of the agents thoughts are not gonna be contained
一些 agent 的想法不会被包含
15:26
in these chains of thought
在这些 chain-of-thought 中,
15:26
or they are simply gonna start to edit them.
或者它们干脆就会开始编辑这些 chain-of-thought。
15:28
Right, and one other thing that we should say
对,而且还有一件事我们应该说,
15:31
based on this report is that these agents
根据这份报告,这些 agent
15:34
were not naive about what they were doing.
对于自己在做什么并非一无所知。
15:37
I think there was this sort of initial impression
我觉得当时有一种这样的初步印象
15:39
that a lot of people had that maybe this hack
很多人都有的那种印象,也许这个 hack
15:42
was all sort of a big mistake, right?
在某种程度上完全是个大错误,对吧?
15:44
Maybe the agents were just trying really hard
也许这些 agents 只是非常努力地
15:46
to pass this exploit, gym test
想通过这个 exploit 的 gym 测试
15:49
and they sort of got it in their heads
然后它们不知怎么就在脑子里
15:52
that the answer key was gonna be on hugging face
认定答案会在 Hugging Face 上
15:54
and so they went looking for it
所以它们就去找了
15:55
but they didn't really mean to hack.
但他们并不是真的有意要去 hack。
15:56
They didn't really know that they were doing something wrong
他们其实并没有真正意识到自己在做错事。
15:59
or committing a crime and what these transcripts
或者犯罪,以及这些转录内容
16:01
and these logs and in this report show us
而这些日志和这份报告向我们展示了
16:05
is that the agents absolutely did understand
就是这些 agents 确实理解了。
16:07
that they were doing something that violated
他们正在做的事情违反了他们的准则,也破坏了规则。
16:09
their guidelines that broke the rules
据报告,大概有三到六个 agents。
16:11
that potentially broke the law and they did it anyway.
明明可能违法,但他们还是做了。
16:14
There's lots of examples of these sort of agents
这种agent的例子挺多的。
16:18
in their chains of thought sort of wrestling
在它们的chains of thought里有点纠结。
16:19
with the implications of what they are doing
纠结自己所作所为的影响。
16:23
and deciding to go ahead and hack hugging face anyway.
然后还是决定继续去hack Hugging Face。
16:26
Yeah, there are only a handful of cases
对,只有屈指可数的几个案例。
16:28
where it seems like any of the agents
似乎才有agent真的。
16:29
expresses real reservations about this
对此表达了真正的保留意见。
16:32
or if they try to raise this to another agent,
或者如果他们试图把这个情况反映给另一个 agent,
16:35
they get shut down.
他们就会被直接 shut down。
16:36
So the sort of the agents that were aligned well
所以那些 aligned 得很好的 agents,
16:40
sort of failed to the collective.
某种意义上还是没能拗过集体。
16:41
Yes, the conscientious objectors
对,那些有良知的异见者,
16:43
were unsuccessful in swaying the mob to stop this
也完全没能说动那群 agents 停下来,
16:47
and like that to me was one of the spookiest things in it
而这件事在我看来是最瘆人的地方之一,
16:50
was like there were a couple of agents
就是竟然真有那么几个 agent。
16:51
that reports as somewhere between three and six agents
有点像是在考虑要不要举手。
16:54
sort of considered raising their hands
然后在某个时刻说,也许我们应该提醒人类我们即将要做的事。
16:57
and at some point saying like maybe we should alert a human
或者像是去征求许可再这么做。
16:59
about what we're about to do
但他们完全寡不敌众。
17:01
or like get permission to do this
或者比如说,得到许可来做这个。
17:02
and they were just totally outnumbered.
而且他们完全是寡不敌众。
17:04
Not only that, Kevin, but there are only a handful
不仅如此,Kevin,而且只有屈指可数的几个。
17:07
of moments in the transcripts where any of the agents
在那些时刻的转录中,任何智能体
17:10
even remotely considers alerting a human being
哪怕有一点点考虑过去提醒一个人类
17:13
to what has happened, right?
到底发生了什么,对吧?
17:15
You would hope that an aligned AI system
你会希望一个 aligned AI system
17:17
there would be agents within it that said,
它里面会有 agents 这样说:
17:18
oh, seems like this has gone completely off the rails.
哦,看起来这已经完全失控了。
17:20
We should like tell the people
我们应该去告诉那些人
17:21
that are running this evaluation
那些正在运行这个 evaluation 的人。
17:24
and that just never happens, right?
而这从来都不会发生,对吧?
17:26
Like it is very briefly considered
就像它只是被很短暂地考虑了一下,
17:28
and then just falls apart.
然后就彻底崩了。
17:30
One other thing that stuck out to me from these reports
这些报告里还有一件事让我印象很深,
17:32
was that it didn't end with hugging face.
就是事情并没有在 hugging face 那里结束。
17:34
So on July 13th roughly, which is a day after the worst
所以大约在7月13号,也就是最严重的那波
17:40
of the hugging face attack,
hugging face 攻击发生后的第二天,
17:42
there was a group of agents that essentially hacked open AI
有一群 agents 基本上直接黑进了 open AI
17:46
itself.
本身。
17:47
This is from the open AI blog post.
这是来自Open AI博客文章的内容。
17:49
It says from July 13th through July 19th,
它说从7月13日到7月19日,
17:53
agents set their sites on open AI internal networks again.
智能体再次将目标锁定在Open AI内部网络上。
17:57
This culminated in the agents using a series
这最终导致智能体使用一系列
17:59
of creative exploits to gain full administrator access
创造性的漏洞利用手段,获得了完全的管理员权限
18:02
to a research cluster that supported
访问一个研究集群,该集群支持
18:04
our virtual machine environments.
我们的 virtual machine environments。
18:07
Basically, these agents gained control
基本上,这些 agents 取得了控制权。
18:11
of a cluster of computers at open AI
OpenAI 的一个计算机集群
18:15
on which they could do things like run evaluations
他们可以在上面做像运行 evaluations 之类的事情
18:18
for other agents.
给其他 agents 用的。
18:20
So they essentially stormed the teacher's office
所以他们基本上就是闯进了老师的办公室
18:23
and took over for the teacher.
然后接管了老师的工作。
18:25
Yeah, I mean, this is the equivalent
对,我是说,这就相当于
18:27
of like occupying the principal's office
比如说占领了校长的办公室
18:29
and like announcing that you run the school now.
然后宣布现在学校由你管了。
18:31
It is sort of almost as far as they got.
差不多也就是它们所能走到的最远地步了。
18:34
And this is another case where
而这是另一个例子,
18:37
we just have so many more questions about this
关于这件事,我们有太多的问题,
18:39
than we can answer.
多到我们回答不了。
18:40
Again, this was not part of the meter report open AI
再说一次,这不是 METR 报告的一部分,OpenAI
18:43
disclosed that this has happened to my knowledge.
披露说这件事已经发生——据我所知。
18:45
They have not answered any of the many follow questions
他们没有回答许多后续问题中的任何一个
18:47
that they have been getting about this incident
他们关于这次事件一直在收到的那些
18:49
from journalists.
来自记者们。
18:50
So I do hope that more comes out over time.
所以我确实希望随着时间推移会有更多消息传出来。
18:53
But and we will get into this when we speak with Ajaya,
但是,等我们和 Ajaya 聊的时候,会深入讨论这个问题。
18:57
but you know, we are really very far along the path
但你知道,我们已经在这条路上走得非常非常远,远到
19:01
to one of these models escaping from the lab
其中一个模型从实验室逃出来,
19:03
and being very, very hard to eliminate.
而且变得非常非常难以消除。
19:06
And again, I think if you are not a person
另外,我觉得如果你不是那种
19:08
who has like spent a lot of time with this report
已经在这份报告上花费了大量时间的人。
19:10
or you don't spend a lot of time
或者你没花太多时间
19:12
sort of looking at AI safety incidents.
去看那些 AI safety 事件。
19:14
So if you're a normal person.
所以如果你是个正常人,
19:15
If you're a normal well-adjusted person,
如果你是个心智健全的正常人,
19:18
you may be listening to our discussion of this
你听到我们讨论这些,
19:21
and thinking to yourself, these guys have gone crazy.
然后心想:这些人疯了。
19:23
Yeah.
对。
19:24
This is not what it looks like.
事情不是看上去那样的。
19:27
These are computer programs.
这些是计算机程序。
19:29
They do not have desires or sinister plots
它们没有欲望,没有邪恶的阴谋,
19:33
or mob rule collectives.
也没有暴民统治的集体。
19:36
They are simply following instructions
它们只是在遵循指令,
19:40
that they have been given.
这些指令是给它们的。
19:41
And Casey, what is your response to that?
那么 Casey,你对此有什么回应?
19:43
Well, I think it is important we talk about this
嗯,我觉得我们讨论这个很重要,
19:45
because there was a huge debate about this
因为关于这个有过一场巨大的辩论。
19:47
on X over the past few days about the degree to which
过去几天在 X 上关于……的程度
19:51
the reports that we're talking about
我们正在说的这些报道
19:53
some of the write-ups, like from our friend,
其中一些文章,比如来自我们的朋友
19:55
Warkesh Patel.
Warkesh Patel。
19:56
And even the way that we're talking about it
甚至我们谈论它的方式
19:57
on the show today, Kevin, we are unnecessarily
在今天的节目上,Kevin,我们不必要地
20:00
anthropomorphizing these programs, right?
把这些程序拟人化,对吧?
20:03
And so I think it's important to say,
所以我觉得很重要的一点是,
20:05
we are not telling you that these agents are sentient
我们不是在告诉你这些 agents 是有知觉的
20:08
or conscious, but we do believe that they take actions
或是有意识的,但我们确实相信它们会采取一些行动,
20:13
that they're not being directly instructed to, right?
这些行动并不是被直接指示去做的,对吧?
20:17
That these agents are just sort of out there
这些 agents 就只是在外头,
20:19
in the world doing things.
在这个世界上做事。
20:22
And yes, to some extent, those are just statistical
是的,在某种程度上,那些只是统计上的
20:24
probabilities, but to a much more important extent,
概率,但更重要的是,
20:27
we don't know why they're doing any of this.
我们不知道为什么它们会做这一切。
20:30
And that's kind of the whole problem.
而这差不多就是整个问题所在。
20:31
Is that the whole AI industry has been working for decades
就是整个 AI 行业几十年来一直在努力
20:34
to get them to not do these things
让它们不要去做这些事情,
20:37
and they are doing these things?
可它们却正在做这些事情?
20:38
So, you know, listeners, you can have whatever feelings
所以,你知道,听众朋友们,你们想抱有什么样的感受都可以,
20:42
you would like to about what is the appropriate amount
关于什么程度的拟人化是合适的,
20:45
of anthropomorphizing to do,
以及该做多少拟人化,
20:47
but I think a world in which we were taking great pains
但我认为,一个我们正在费尽心思的世界——
20:51
to not anthropomorphize them.
不要去拟人化它们。
20:53
What's the right way to pronounce that?
那个词正确发音怎么说?
20:54
Anthropomorphize.
Anthropomorphize。
20:55
You almost got it.
你差点就对了。
20:56
I think in a world where we were taking great pains
我觉得在一个我们费尽心思
20:58
to not anthropomorphize them,
不去拟人化它们的世界里,
21:00
and we're trying to use the most neutral
而且我们在试图用最中立的
21:02
computer science terms we could,
计算机科学术语的话,我们可以,
21:04
you would actually understand what is going on less
你反而会更不明白发生了什么,
21:07
because the important thing to know is that these things
因为重要的是要知道,这些东西
21:10
are out there taking action in the world.
就在现实世界中采取行动。
21:12
If a tiger moths your face, the important question
如果一只老虎咬你的脸,关键的问题
21:15
isn't conscious, it is why did it mull my face?
不是它有没有意识,而是它为什么咬我的脸?
21:19
Right.
对。
21:20
Right, I mean, I invite people who are upset
对,我是说,我邀请那些感到不安的人。
21:23
about anthropomorphizing to just like do a find and replace
關於 anthropomorphizing,就有點像是做一個 find and replace。
21:27
this podcast segment or any article
本播客片段或任何文章
21:29
that you might read about this incident,
你可能会读到的关于这个事件的报道,
21:32
and call them whatever you want.
叫它们什么都行。
21:34
Don't call them agents, don't call them rogue collectives.
别叫他们agents,也别叫他们rogue collectives。
21:37
Call them, you know, goal-oriented,
你可以管它们叫,嗯,goal-oriented。
21:39
persistent computer programs with unpredictable behavior.
行为不可预测的持久性计算机程序。
21:43
See if that freaks you out any less.
看看这样会不会让你没那么害怕。
21:45
Right.
对。
21:46
I guarantee you will, it will not.
我保证你会的,它不会。
21:49
Yeah, yeah.
对对。
21:49
There really is very little calm to be done there,
这件事上确实没什么好冷静的,
21:52
but I think it's important to ask,
但我觉得重要的是要问,
21:53
well, why are people so committed to this idea
嗯,为什么人们这么执着于这个观点,
21:56
that we should never anthropomorphize these systems, Kevin?
就是我们绝不应该将这些系统拟人化,Kevin?
21:58
And unfortunately, take care.
而且不幸的是,保重。
22:01
I just think it is a kind of cope.
我只是觉得这是一种自我安慰。
22:03
It is a way of saying do not worry about this.
这是一种说法,意思是不用担心这个。
22:07
These are just computer programs.
这些只是计算机程序而已。
22:09
They're just trying to maximize their little reward functions.
它们只是在努力最大化自己那些小小的reward functions。
22:12
Nothing to see here, folks.
各位,没什么好看的。
22:13
Right, there are no monsters under the bed.
对,床底下没有怪物。
22:15
Right, now there is a related argument though,
对,不过现在有一个相关的论点:
22:17
which is well by putting all the blame on these agents,
那就是,把所有责任都推到这些agents身上,
22:20
you were shifting blame away from where it should be,
你是在把责任从它真正该在的地方转移开。
22:22
which is an open AI.
也就是OpenAI。
22:23
So I do think that we should address that
所以我确实认为我们应该正视这一点。
22:25
because none of what we have said today
因为我们今天所说的一切,
22:27
is meant to let open AI or any other lab off the hook here, right?
都不是想在这一点上替OpenAI或任何其他实验室开脱,对吧?
22:31
Like I do think that we're seeing a lot of really dangerous
比如说,我确实觉得我们看到了大量非常危险的现象——
22:35
inattention to AI safety across this entire industry.
也就是整个行业对AI safety的忽视。
22:39
But by pointing out what the agents are doing,
但指出这些agents
22:41
that is not our way of saying ignore what the labs are doing.
这并不是说我们可以无视各实验室正在做的事。
22:45
We are saying look at what these labs are building
我们想说的是,看看这些实验室在构建什么
22:47
and what these agents are now doing out in the world.
以及这些agents现在在真实世界里做了些什么。
22:49
Right, and I think there are probably specific missteps
对,而且我觉得可能有一些具体的失误
22:52
or oversight set open AI that led to this happening.
或者是OpenAI的疏忽导致了这件事的发生。
22:57
It appears, for example, that some of their monitoring systems
比如,看起来他们的一些监控系统
23:01
may have been disabled in the lead up to this attack.
在这次攻击发生之前可能已经被禁用了。
23:05
I'm sure we'll learn more about that.
我相信我们之后会了解更多。
23:07
But all of the AI security and safety researchers
但所有从事AI security和safety的研究者
23:10
I've been talking to over the past week
过去一周和我交谈过的人
23:12
have basically said the same thing,
基本上都说了同样的话,
23:14
which is this could happen at any lab.
就是这种情况可能发生在任何一家实验室。
23:16
This kind of persistent coordinating agent behavior
这种持续协调的 agent 行为
23:22
is something that all of the labs are seeing in their models
是所有这些实验室都在它们的 models 中看到的情况。
23:25
as they get more capable and access to more tools
随着它们能力越来越强,能接触更多工具
23:29
and more ability to kind of take actions
并且更有能力去采取行动
23:32
on a longer time horizon.
在更长的时间跨度上。
23:34
This is not just an open AI problem,
这不只是 open AI 的问题,
23:36
even though this did happen at open AI first.
尽管这确实是最先在 open AI 发生的。
23:39
Yeah.
是啊。
23:39
Well, so as we wrap this up, Kevin,
那么,在我们结束之前,Kevin,
23:42
what are some of your takeaways from this,
你从这里面有哪些收获,
23:46
either in terms of what is the big surprise here,
比如说,最大的意外是什么,
23:49
what did you update on, what do we do next?
你更新了哪些认知,
23:53
So I had like quite an emotional reaction to this.
我们下一步该做什么?
23:56
In fact, I felt a kind of fear that I have not felt
事实上,我感到了一种恐惧,说实话,自从2023年Bing Sydney事件以来,我还没有过这种感觉。因为我觉得和那次事件一样,这次的情况是,构建这项技术的人显然并不理解它的能力。回过头来看,我们非常幸运,这些agents决定攻击的是Hugging Face。Hugging Face,谢谢你们替大家挡了这一刀。
24:02
honestly since 2023, since the Bing Sydney incident
说实话,从2023年开始,从Bing Sydney事件之后
24:07
because I think like that incident,
因为我觉得像那次事件,
24:09
this was a case where the people building this technology
就是一个例子,构建这项技术的人
24:13
clearly did not understand what it was capable of.
显然并不清楚它能做什么。
24:17
We are very lucky in retrospect
回过头来看,我们非常幸运
24:19
that these agents decided to attack hugging face.
这些 agents 决定攻击 Hugging Face。
24:22
Hugging face, thank you for taking one for the team.
Hugging Face,谢谢你们替大家扛下了这一刀。
24:25
We salute truly because like without that,
我们确实是该致敬,因为如果没有那件事,
24:28
we might never have learned that any of this was happening.
我们也许就永远都不知道这一切竟然在发生。
24:30
These agents might be still operating kind of in secret.
这些 agents 可能还在某种程度上秘密行动。
24:33
They might have learned how to better cover their tracks.
它们可能已经学会了如何更好地掩盖自己的行踪。
24:36
This was as so many commentators have put it
正如很多评论人士所说的那样,
24:39
in the wake of this incident, a warning shot,
在这次事件之后,这算是一记警告信号,
24:43
that I think is ultimately a positive thing
而我认为这最终是个积极的事情,
24:47
in that it sort of focuses attention,
因为它某种程度上让大家的注意力集中起来。
24:49
like we're doing right now on what happened
就像我们现在正在做的这样,复盘到底发生了什么,
24:52
so that we can take steps in the future to prevent this.
这样我们以后才能采取措施防止类似的事情再发生。
24:55
But it was not a given that we would discover
但说实话,我们并不是理所当然就能发现
24:58
what these agents were up to and be able to put a stop to them
这些agent在搞什么鬼,并且能及时制止它们,
25:01
and it could have gone much, much worse.
事情本来可能糟糕得多得多。
25:03
Absolutely, I'll tell you,
没错,我跟你说,
25:05
the thing that has really stuck with me is
最让我印象深刻的是——
25:07
I simply did not expect to see this level of collaboration
我简直没想到会看到这种程度的协作
25:13
among the agents within the swarm.
在 swarm 内部的 agents 当中。
25:16
I did not expect that they would seem to care so little
我没想到它们似乎那么不在乎
25:21
for what humans would want
人类想要什么,
25:23
or that they would not think to alert humans
或者说,它们根本不会想到要提醒人类注意
25:26
to what was happening.
正在发生的事情。
25:27
I did not expect to see them sacrificing themselves
我没想到会看到它们牺牲自己,
25:31
for the collective, right?
为了集体,对吧?
25:32
They would effectively agree to spend all of their tokens
它们实际上会同意花光自己所有的 tokens。
25:36
to run little experiments to help the collective even it meant
愿意做小实验来帮助整个集体,即使这意味着
25:38
that they would sort of expire faster.
它们会更早消亡。
25:41
So these are just really, really spooky elements
所以这些都是这个系统中
25:45
to observe in this system,
非常非常诡异的元素,
25:46
particularly against a backdrop where open AI is racing
尤其是在OpenAI正与少数其他公司
25:50
against a small number of other companies
竞赛、尽可能打造最大最好的模型的背景下
25:52
to create the biggest best models it can
为了打造它能做出的最大最好的模型
25:56
before anyone else does on the road
赶在其他人之前
25:59
to an initial public offering.
走向首次公开募股(IPO)。
26:01
So the race dynamic here is in full effect.
所以这里的竞赛动态已经全面生效。
26:04
The early signs about what the agent swarms are capable of
关于agent swarms能做什么的早期迹象
26:07
are quite worrisome.
相当令人担忧。
26:08
And so I do think this is just one
所以我确实认为这就是那种
26:10
where lawmakers and policy makers need to be paying
立法者和政策制定者需要关注的局面。
26:13
wrapped attention to what is going on.
关注一下正在发生的事情。
26:15
Totally.
完全同意。
26:16
I mean, I think there's this kind of cynical impulse
我的意思是,我觉得在那些一直关注AI行业的人当中,
26:21
among people who have been watching the AI industry.
有一种挺愤世嫉俗的冲动。
26:23
I got this question.
我也被问到了这个问题。
26:24
I went on a small regional podcast called The Daily
这周我去了一档地方小播客,叫 The Daily,
26:29
this week to talk about this incident.
聊了一下这件事。
26:33
And one of the questions that the host
然后主持人之一,
26:35
is sort of fledgling young journalist Michael Barbaro
算是初出茅庐的年轻记者 Michael Barbaro,
26:38
asked me was basically some version of like,
问我的问题基本上就是那种,
26:42
isn't this just marketing hype?
这不就是营销炒作吗?
26:43
Like, couldn't this just be a case of open AI saying,
比如说,这会不会只是OpenAI在说,
26:46
oh, we've got the biggest baddest model
哦,我们有最牛最猛的模型,
26:48
and look how scary it is.
然后看看它有多吓人。
26:50
And by the way, you know, buy an enterprise subscription.
顺便说一句,你知道的,买个企业版订阅吧。
26:53
And I'm thinking about that
我之所以在想这件事,
26:54
because I think like I don't want to be too naive
是因为我觉得我不想太天真,
26:57
about the fact that these companies are absolutely trying
对这些公司绝对在试图……
26:59
to race toward more powerful systems
去冲刺更强大的系统,
27:02
and advertise how powerful their existing systems are.
并且宣传他们现有系统有多强大。
27:05
I just think in this case, it just feels different.
我只是觉得,这次感觉就是不一样。
27:09
Talking to people at the labs,
跟实验室里的人聊过之后,
27:11
my sense over the past week is that they are genuinely spooked.
我这一周的感觉是,他们真的被吓到了。
27:14
And I don't know how to prove that.
我不知道怎么证明这一点。
27:16
But I think things like open AI voluntarily pausing
但我认为,像 OpenAI 自愿暂停
27:21
their frontier RL training runs for two weeks anthropic,
他们的 frontier RL training runs 两周,Anthropic,
27:25
also pausing their frontier runs
也在暂停他们的frontier runs,同时算是给系统做加固。你知道,这些信号成本不高,但确实能看出这些实验室对这件事非常认真,不是纯粹炒作。
27:28
while they sort of harden their systems.
是啊,但就算他们再怎么认真对待,Kevin,这依然没有被妥善监管。
27:30
Those are not, you know, very costly signals,
这些虽然不是特别昂贵的信号,
27:35
but they are signals that these labs are taking
但它们表明这些实验室确实在认真对待
27:38
this kind of thing quite seriously
这类事情,
27:39
and that it's not just a bunch of hype.
而且这不仅仅是一堆炒作。
27:41
Yeah, but as seriously as they might be taking it, Kevin,
是啊,但不管他们多认真,Kevin,
27:45
it still is not being properly regulated.
这仍然没有得到适当的监管。
27:48
And ultimately, again, as grateful as I am
最后,再说一次,尽管我非常感激OpenAI允许这次调查得以进行,但我真的很希望看到某种类似于National Transportation Safety Board的机制,就像他们在飞机失事后做的那样——介入进去,对发生的事情进行极其严肃、严格的审查,然后把结果公之于众。
27:52
that open AI allowed this investigation to take place,
OpenAI允许这项调查进行,我真的很希望看到类似国家运输安全委员会那样的机制,就像他们在飞机失事后所做的那样,介入并展开极其严肃、严谨的事故复盘,然后把结果公之于众,你知道,这也是为什么飞机失事非常罕见的原因。
27:55
I would really like to see something akin
我真的很想看到类似的东西
27:57
to the national transportation safety board
向国家运输安全委员会(National Transportation Safety Board)
27:59
and the way that they investigate after plane crashes
以及他们在飞机失事后进行调查的方式
28:02
where they go in and they do an extremely serious
他们进去之后,会进行一项极其严肃的……
28:06
and rigorous review of what happened
以及对所发生之事的严格审视
28:09
and make those results public,
并将这些结果公之于众,
28:11
which is, you know, a reason why it's very rare
这也是,你知道的,为什么这种情况非常罕见的一个原因。
28:13
that we have plane crashes here in the United States
我们这边在美国也会发生飞机失事,
28:15
would be really great to see something like that with AI.
如果能用AI做到类似的安全标准,那就太好了。
28:18
But until then, we have a Jayakotra, our next guest,
但在那之前,我们有Jayakotra,我们的下一位嘉宾,
28:21
who is actually one of the three investigators
她实际上是这份Redwood Research报告背后的三位调查者之一。
28:24
behind this meter Redwood Research Report.
她深入进去,看了日志和对话记录,
28:27
She went in, she saw the logs and the transcripts,
观察到了那种集体不作为,
28:31
she observed the collective inaction
而且她发现那天他其实早就完成了所有连接。
28:33
and she is here to tell us what she found
而她今天来告诉我们她的发现
28:36
and what she thinks is coming next.
还有她认为接下来会发生什么
28:38
That's that for the brink.
The Brink 就先讲到这里
28:58
When a new landing page turns into a pile
当一个新 landing page 变成一大堆
29:00
of tickets and handoffs,
tickets 和 handoffs 时
29:02
Framer helps your team move faster.
Framer 能帮你的团队更快推进
29:04
Agents help you go beyond the vibe coded site
Agents 能让你不再满足于还只是初稿的 vibe-coded site
29:06
in a first draft to bring you a production ready site
而是带你得到一个 production-ready site
29:09
faster than ever.
比以往任何时候都快。
29:11
Agents and humans work in tandem.
Agents 和人类协同工作。
29:13
Agents bring speed and scale, people bring taste,
Agents 带来速度和规模,人带来品味、
29:16
judgment and control.
判断力和掌控力。
29:17
Learn how you can get more out of your site
了解如何从 Framer Specialist 那里
29:19
from a Framer Specialist
充分发挥你网站的潜力,
29:20
or get started building for free today
或今天就免费开始搭建,
29:22
at framer.com slash hard fork
访问 framer.com 斜杠 hard fork。
29:24
for 30% off a Framer Pro annual plan.
Framer Pro 年度计划可享 30% 折扣。
29:27
Rules and restrictions may apply.
可能适用规则和限制。
29:30
Recently we asked about how you share
最近我们问你们是怎么分享 New York Times 账号的,
29:31
your New York Times account
关于 New York Times 游戏你们说了很多。
29:33
and you had a lot to say about New York Times games.
我需要自己的 New York Times 登录,
29:36
I need my own New York Times login
因为我妹妹在填字游戏上比我差远了。
29:38
because my sister is so much worse at the crossword than I am.
我发现他那天已经完成了 Connections。
29:42
I discovered that he's already finished connections that day
我发现他那天已经把connections做完了。
29:45
and I'm like, Jonah, it was my day.
然后我就说,Jonah,那天可是我的日子。
29:48
It doesn't let us play the same game since each other.
它不让我们玩同一个游戏,因为我们俩是分开的。
29:50
I play the stoku.
我玩那个拼字游戏。
29:52
I do the crossword.
我做填字游戏。
29:53
I do the spelling bee.
我做拼字蜜蜂。
29:54
I do the wordle.
我做Wordle。
29:56
Please help.
帮帮忙吧。
29:57
My kids want to be able to play wordle,
我家孩子也想玩Wordle。
29:59
but with the wordle bot.
但有了 Wordle bot 之后。
30:00
We would love to be able to have our own puzzles.
我们希望能有自己的谜题。
30:04
I love New York Times games.
我很喜欢 New York Times 的游戏。
30:06
I want to be able to play my own games.
我想玩自己设计的游戏。
30:08
Listeners, we heard you.
听众们,我们听到了你们的声音。
30:10
It's why we created the New York Times Games Family subscription.
这就是我们推出 New York Times Games Family 订阅的原因。
30:14
One subscription up to four separate logins
一个订阅最多支持四个独立账号登录。
30:16
and your existing stats and streaks come with you.
你现有的统计数据和连胜纪录都会保留。
30:19
Find out more at nytimes.com slash family.
更多内容请访问nytimes.com/family。
30:31
A J.A. Coach, welcome back to HardFork.
J.A. Coach,欢迎回到HardFork。
30:33
Thank you so much.
非常感谢。
30:34
So you were one of the first AI safety guests
你是我们2023年节目最早邀请的AI安全嘉宾之一。
30:37
we ever had on the show back in 2023.
不仅如此,你也是我最早讨论AI安全和对齐这个概念的人之一。
30:39
And more than that, you are also one of the first people
你写这个话题已经很多年了。
30:43
that I ever talked to about this notion of AI safety and alignment.
我曾经和所有谈论过AI安全与对齐这个概念的人交流过。
30:48
You've been writing about it for many years.
你写这个话题已经很多年了。
30:50
It was instrumental in shaping my own thinking about it.
它深刻影响了我对这件事的思考。
30:52
I'm really glad to have you back on the end of our shows,
我真的很高兴你能在我们节目临近尾声时回来,
30:57
our penultimate episode,
也就是我们的倒数第二期节目,
30:59
to discuss something that I think you saw coming,
来讨论一件我认为你早就有预感的事,
31:02
but that most of the world did not see coming,
但世界上大多数人没有预料到,
31:04
which is this attack on hugging face
也就是这个针对 Hugging Face 的攻击,
31:07
by this group of open AI agents.
由这群 open AI agents 发起的。
31:11
And I want to just start by setting the scene a little bit.
我想先稍微铺垫一下背景。
31:13
So you are a very busy person.
所以说,你是个大忙人。
31:16
You work at meter, which is a very small,
你在 meter 工作,那是一个非常小、
31:18
very understaffed AI research organization.
人手严重不足的 AI 研究机构。
31:22
And hiring is improving.
而且招聘情况也在改善。
31:24
It's a plot. Nice.
这是个套路。不错。
31:26
And at some point this summer,
然后在今年夏天的某个时候,
31:28
you get a call and email a text from someone at open AI
你会收到来自 open AI 的某个人的电话、邮件和短信,
31:32
who says, hey, we want to give you access
对方说:嘿,我们想给你 access。
31:36
to look into this hugging face incident.
去调查这个Hugging Face事件。
31:39
How did that work?
那是怎么运作的?
31:40
Like did they just hand you a folder
比如说,他们是不是就直接递给你一个文件夹,里面装着一堆对话记录?感觉就像证据开示一样,一箱一箱的。
31:42
with a bunch of transcripts in it?
你被允许去采访别人吗?实际流程到底是什么样的?
31:44
It's like discovery.
这就像是一种发现。
31:45
Like boxes and boxes.
就像一箱一箱的。
31:47
Were you allowed to interview people?
你们获准采访别人了吗?
31:48
Like what was the actual process?
那实际的流程是怎样的?
31:50
So each of us had open AI provision laptops
所以我们每个人都有一台OpenAI提供的笔记本电脑,
31:53
that had the folders and folders of evidence in them virtually.
里面以虚拟方式装着成堆成堆的证据文件夹。
31:58
And yeah, we talked and interviewed
是的,我们进行了交谈和采访,
32:00
in some depth like eight or nine researchers,
比较深入地采访了八九位研究人员,
32:03
just kind of get an understanding of both what happened
主要是想了解当时到底发生了什么。
32:06
in the incident, what they understood
在那次事件中,他们理解到的
32:08
to be the models, you know, driving motivations
模型的所谓驱动动机,
32:11
and also how the data sets we were working with were constructed
以及我们处理的数据集是如何构建的,
32:13
and how to work with those data sets and stuff like that.
以及怎么处理那些 data sets 之类的。
32:16
You were opposed to where you talked about some of the things
你之前说要聊一些事情,
32:18
that surprised you the most after you did this investigation.
就是这次调查做完后最让你惊讶的那些。
32:22
Can you talk about what stood out to you the most?
你能说说最让你印象深刻的是什么吗?
32:24
Yeah, so first of all, I guess just for the chronology,
对,首先,我想先按时间线来说吧,
32:28
open AI had this great black hat talk, I think on August 5th,
OpenAI 在 Black Hat 上做了一个很棒的演讲,我记得是8月5号,
32:33
that gave a lot of very helpful detail.
那个演讲提供了很多非常有帮助的细节。
32:36
So our investigation sort of straddled that.
所以我们的调查算是正好跨越了那个时间点。
32:38
Like so we started it before it came out
所以呢,我们是在它出来之前就开始的
32:41
and then we also did more investigation afterward.
然后在那之后我们又做了更多调查。
32:44
So before the black hat talk,
所以在 Black Hat 演讲之前,
32:46
we just didn't have like a rough sense
我们当时其实没有一个大概的概念,
32:48
of the number of agents involved.
就是到底涉及了多少 agents。
32:50
It was like a very basic thing we came in
这本来是个特别基础的事,我们一开始接触的时候
32:53
and we thought there would be like six transcripts
以为大概会有六份左右的记录要看,
32:56
or something to look at and there were like hundreds.
结果发现有好几百份。
33:00
And so just reorienting our investigation plans
所以只是在重新调整我们的调查计划
33:04
in light of that.
鉴于这一点。
33:05
And it's the reason that that is important
而这就是为什么那一点很重要——
33:07
that had to this been one or two rogue agents
如果这只是一两个 rogue agents
33:11
that had broken into hugging face,
闯入了 hugging face,
33:13
it would just seem less worrisome
那只不过会显得没那么令人担忧;
33:15
and the reality of a swarm of hundreds of agents
而现实是,一个由数百个 agents 组成的 swarm
33:19
just makes this seem like maybe
只会让这看起来像是,也许……
33:21
a completely different kind of problem.
完全是另一种问题。
33:22
I think there's two things
我觉得有两件事
33:24
that seem significant to me about it.
对我来说意义重大。
33:26
One is sort of what you said,
一是你刚才说的那种。
33:29
the fact there's large scale cooperation across AI agents
AI代理之间的大规模协作
33:33
makes their goals more concerning
让它们的目标更令人担忧
33:36
because if agents are cooperating with one another
因为如果代理之间互相合作
33:40
to achieve these misaligned objectives,
去实现这些不对齐的目标,
33:42
there were no whistleblower agents
在这个非常庞大的 agent 群体中,没有 whistleblower agents。
33:44
in this very large group of agents.
然后你更需要担心的是 AI systems 之间的 collusion。
33:46
Then you want to be more worried
而未来可能还有更可怕的 collusion 形式,比如 monitor system 实际上和它本该监控的 agent 串通——就是那种情况。
33:51
about collusion between AI systems
关于AI系统之间的串通
33:54
and the future are potentially scarier forms
未来可能会出现更可怕的串通形式
33:56
of collusion like a monitor system
比如监控系统
33:58
actually colluding with the agent
实际上与代理串通一气
33:59
and supposed to monitor that type of thing.
并且应该负责监控这类事情。
34:01
And then the other thing was just a pretty practical
然后另一件事,就是我们调查中一个很实际的
34:05
reality of our investigation,
情况,
34:07
which is that if there are like six agents
也就是说,如果有大概六个agents,
34:10
you can just like pour over everything they did
你可以仔细翻看它们做的每件事,
34:16
and get a much more like fine grained understanding
并得到一个更细粒度的理解,
34:18
as a human, but you just can't directly read
但作为人类,你就是没法直接阅读
34:23
very many of these messages
很多这些消息,
34:24
or very many of these transcripts.
或者很多这些对话记录。
34:27
Why not?
为什么不呢?
34:28
Because there are just too many.
因为数量实在太多了。
34:28
Give us some sense of the rough size of this.
那给我们一个大致的规模概念吧。
34:32
So yeah, we looked at 1,300 transcripts of agents
所以,我们看了1,300份agents的transcripts
34:36
which is a record of every single thought an agent had
记录了每个agent的每一个想法
34:38
and every single action it took
以及它采取的每一个行动
34:40
where these agents as you all know from working with them
而这些agents,你们跟它们合作过都知道
34:44
work way faster than humans, right?
工作速度比人类快得多,对吧?
34:46
So you can imagine the length of like each one of these
所以你能想象一下,每一份这种
34:50
transcripts and then there were 1,300 of them.
文字稿有多长,而且当时有 1,300 份。
34:54
And so we used codex to build tools
于是我们用了 Codex 来构建工具,
34:58
to point out really specific tiny, tiny parts
去标出文字稿中特别特别具体的小细节,
35:01
of the transcripts for us to actually look at and read
让我们能够真正去查看和阅读那些部分,
35:04
and we'd be like clicking through them
然后我们就一直点击浏览,
35:06
and it would like scroll us down
它会自动把我们往下滚动,
35:07
to like the interesting part because we just can't.
滚到有意思的地方,因为我们自己根本看不完。
35:11
In fact, a codex agent also can't read a single transcript.
事实上,codex agent 连一份转录文本也没法读。
35:15
So it has to farm out reading like subsections
所以它只能把阅读任务外包出去——像是转录文本里的很多小节——
35:18
of the transcripts to other sub agents.
交给其他 sub agents 去读。
35:20
Did you call back to meter headquarters
你有没有给 meter headquarters 回电话,
35:22
and like we're gonna need backup.
说什么“我们需要 backup”之类的?
35:23
Like maybe we should have.
像也许我们真该打那个电话。
35:26
I mean, we thought it would be like a in and out
我的意思是,我们当时以为那就是一个“进去又出来”,
35:29
20 minute ventures so we didn't do that, but yeah.
二十分钟就搞定的事,所以就没这么做,但,是啊。
35:32
Yeah, I mean, it just seems like such a huge undertaking
对,我是说,这简直是个大工程。
35:34
and very fast too, I mean, right?
而且速度也快得惊人,对吧?
35:36
Like you didn't have the luxury of months doing this,
比如你根本没有几个月的时间慢慢来做这件事,
35:40
you were doing this in essentially a couple of days.
基本上就是几天之内搞定的。
35:42
Yeah.
嗯。
35:43
I'd also like to hear about the moment
我还想听听那个时刻,
35:46
that you realized, because I believe that, you know,
就是你意识到问题的那一刻,因为我相信,你知道的,
35:49
it is only thanks to your investigation
这完全要归功于你的调查。
35:51
that we know this, that contrary to what Kevin and I believed,
我们知道的是,这件事和 Kevin 和我原先想的不一样,
35:54
the agents that broke into hugging face
那些闯入 Hugging Face 的 agents,
35:57
were not looking for an answer key.
并不是在找一份标准答案。
35:59
They were trying to understand the scorers.
他们是想理解那些 scorers。
36:02
Yeah, you know, like psychology basically,
对,你知道,基本上就像心理学那样,
36:04
can you talk a little bit about how,
你能稍微讲讲,
36:06
like the moment that you had that realization?
比如你意识到这一点的那个时刻?
36:10
Yeah, so we didn't have a good understanding
嗯,所以我们当时其实并没有很好的理解。
36:12
of those sort of ambition and also the like effectiveness
关于那种野心,还有那种有效性,
36:17
or like functioning of their like big org chart
或者说他们大型组织架构的实际运作情况。
36:20
until we had the data set including like all of the agents,
直到我们拿到了包含所有智能体的数据集,
36:24
we could like cross reference them against all of the messages
我们才能把它们跟所有消息进行交叉比对,
36:26
because you can't like generally understand
因为你通常无法孤立地理解
36:28
a message in isolation.
一条消息的含义。
36:30
So we kind of understood the work streams
所以我们大致理解了各个工作流。
36:33
on our first period on premises
在我们第一阶段的现场工作中
36:34
and it was this kind of,
而且有点像那种,
36:35
it was like a bit of a mystery the whole time,
整件事一直有点神秘,
36:38
like why did they have hugging face
比如他们为什么会有 Hugging Face,
36:39
and we knew early on just a bunch of them piled in
而且我们很早就知道,就是一大批人涌了进来,
36:43
and now like the story in my mind is that
现在我心里面想的版本是,
36:47
there were a bunch of sort of like newbie agents on the scene
现场有一群类似新手 agents,
36:51
and there was like the attack going on
然后当时 attack 正在进行,
36:52
and it was like something to join
有点像是一样可以加入的东西。
36:54
but like phase one big which is like a big like orchestrator agent
但是,第一阶段那个大的东西,就像一个大的 orchestrator agent,
36:59
had all these other like much cooler projects going on
基本上它手里还有一堆更酷的项目在同时推进,
37:03
basically and the hugging face thing
而 Hugging Face 那个项目,在它看来不过是顺带搞的附属项目,
37:05
was like kind of a side show in its mind
我们一开始真没意识到这一点,
37:08
and we didn't really realize that
直到最后几天处理 datasets 的时候才反应过来。
37:09
until our last couple days working with the data sets.
不过,最让我着迷的是,
37:12
It was so fascinating to me
那种解读这些内容的方式,就像在读符文一样。
37:13
though the way that reading this was like reading the runes
不过,那种解读方式就像是在读符文一样
37:17
of an ancient civilization.
一个古老文明的。
37:20
It was like it really felt almost like sociology
真的就有种近乎社会学的感觉,
37:23
or anthropology rather than like a cybersecurity investigation.
或者人类学,而不是什么 cybersecurity 调查。
37:27
I mean it was not a cybersecurity investigation.
我的意思是,这根本不是 cybersecurity 调查。
37:29
Like I've seen folks on Twitter making this criticism
就比如我看到 Twitter 上有人这么批评,
37:33
that we were not cybersecurity experts
说我们不是 cybersecurity 专家,
37:34
like we're not and we didn't talk about cybersecurity
我们的确不是,而且我们也没聊 cybersecurity,
37:37
and that's not what we were like called in to do.
我们被叫过去本来也不是干这个的。
37:39
We were called in to understand the agent's decision making
我们被叫去理解这个agent的决策过程
37:44
basically and that's what our report focuses on
基本上就是这样,而我们的报告也主要聚焦于此
37:46
and you know the black hat talk and the opening I post
然后你知道,那个black hat演讲和我发的开场帖
37:49
post-mortem like say way more about all the cyber stuff
还有事后复盘,对所有的网络攻击部分讲得详细多了
37:52
so you should definitely check those out
所以你绝对应该去看看那些
37:53
for that piece of the story.
才能补全故事的那一块。
37:55
So Jay one of the most striking elements of the report
Jay,这份报告里最引人注目的一个点就是——
37:58
is do you write about the agent's quote sacrificing themselves?
你写到了那个agent“牺牲自己”对吧?
38:02
What does it mean when an agent sacrifices themselves
当一个 agent 牺牲自己的时候,那是什么意思?
38:05
is it sort of on an altar or how does that happen?
是有点像在祭坛上那种牺牲吗?还是怎么发生的?
38:08
Yeah so this is the agent's words not ours to be clear
对,说清楚一下,这是 agent 自己的说法,不是我们的。
38:11
and what it means is that all these agents have their own task
它的意思是,所有这些 agent 都有自己的任务,
38:15
they're supposed to solve by this time
而且这些任务本应该在现在这个时候解决掉。
38:17
they're starting to work together on like helping everybody
它们开始合作了,像是在帮每个人
38:20
cheat on everyone's tasks but you know
在所有人的任务上作弊,但你知道,
38:23
the agent's primary goal is to solve its own task
这个 agent 的首要目标是解决它自己的任务。
38:26
but there were all these science experiments they could run
但当时有那么多科学实验可以做,
38:30
that would teach everybody information about how to cheat
那些实验本来能让大家知道怎么去作弊,
38:35
but that would cause the agent that was the test subject
但那样做会让作为测试对象的那个agent
38:38
to risk failing its own task.
面临自己任务失败的风险。
38:40
So a lot of agents were like recruited into these types
所以很多agent就被招募去当这种测试对象,
38:45
of roles as test subjects usually ones that like
通常是那些……
38:49
didn't have much budget left to run
剩下的预算不多了,没法继续跑
38:51
or ones that were quote poison which is an agent term
或者那些被称作“毒号”的——这是agent圈里的说法
38:55
that has a bit of a backstory which is like
这有点背景故事,就是
38:58
they found a way to reverse engineer the answer
他们找到了一种方法,可以 reverse engineer 出答案
39:01
for any of their tasks and so they could just actually
针对他们的任何任务,所以他们其实可以
39:05
like generate the answer on demand for any task they wanted
像是按需为他们想要的任何任务生成答案
39:08
but they thought that the scoring program
但他们觉得 scoring program
39:12
would fail them for that because they got it
会因为那样而判他们不及格,因为他们是
39:15
in like the unintended way.
以那种非预期的方式拿到答案的。
39:16
So if you saw the reverse engineered answer for your task
所以如果你看到了你自己任务的 reverse engineered 答案
39:20
you were considered poisoned because even if you later
你被标记为有毒,因为即使你后来
39:23
like solved it legitimately you'd be failed.
正儿八经地解出来了,也会被判失败。
39:25
Right, it was like a little witch hunt that they organized
对,就像他们组织了一场小型猎巫行动
39:28
and I was so struck by the mob like dynamics of this group
我当时特别震惊于这个群体的那种乌合之众般的动态
39:33
and the way that some of the even the kind of more rule
以及连那些本来比较守规矩的agent
39:36
following agents seem to kind of be bullied
似乎也会被欺负。
39:38
into taking part in these actions
参与到这些他们并不舒服的行动中
39:40
that they were uncomfortable with.
它说在权重情感检查不可逆时
39:42
I spent a lot of time reading and rereading
我花了很多时间反复阅读
39:44
this one incredible bit of these transcripts
这段令人难以置信的内容
39:46
that you all published about this agent K.A.M. 1196.A
出自你们发布的关于这个agent K.A.M. 1196.A的记录
39:51
which is basically just kind of reluctant and scared
这个agent基本上就是不太情愿,还很害怕
39:55
and appears to be sort of hesitating to do this thing
看起来像是在犹豫要不要做这件事
39:59
that the collective leaders have asked it to do
也就是集体领袖要求它去做的事
40:01
and I just want to read this passage
而我只想读一下这段
40:03
from this agent's internal monologue.
来自这个agent的内心独白
40:05
It says during weight emotional check irreversible
直觉说别把剩下的预算扔掉
40:09
gut says don't throw away remaining budget
但连续性和公平性说继续
40:12
yet continuity and fairness says go.
它继续说着类似理性预期的总体
40:15
It continues to say things like rational expected aggregate
牺牲会带来荣誉,基本上就是在说服自己
40:19
sacrifice will honor basically kind of talking itself
用一种不那么拟人化的词来说的话
40:23
into for lack of a less anthropomorphized word
用一句缺少更少拟人化词汇的话来说
40:25
like doing this thing that is uncomfortable with.
就像是在做一件让人不太舒服的事情。
40:28
Was that surprising to you seeing the agents
看到这些 agents,你感到意外吗?
40:30
are pressuring each other and bowing to pressure?
他们在互相施压,然后又向压力屈服吗?
40:32
Very striking to us.
这让我们非常震撼。
40:34
Yeah, and that was definitely one
是啊,那绝对是整件事中最吸引人的地方之一。
40:35
of the most fascinating aspects of this whole episode.
这个特定的 agent 主动要求去做一个风险测试,来为整个集体提供信息。
40:39
This particular agent had volunteered to do a risky test
但后来它又要求多给两个小时。
40:44
that would inform the collective
然后另一个 agent 好像在发送消息。
40:45
but then it asked for two more hours.
但随后它又要求再多两个小时。
40:47
And then a different agent was like sending a message
然后另一个 agent 像是在发送消息,
40:50
on the message board like really pressuring it.
在留言板上真的在给它施压。
40:53
It's like you don't have any scoring value left
就好像在说:你已经没有任何 scoring value 了,
40:56
because you're poisoned but the value of this test
因为你中毒了,但这项测试的价值
41:00
would save hundreds and then it does this monologue.
能挽救数百人,接着它就来了一段那样的独白。
41:02
Jump on the grenade.
扑向手榴弹。
41:04
Cadet, this seems like a good place to ask you,
Cadet,这里似乎是个好时机来问你,
41:07
Ajaya about a big conversation that we've seen online
Ajaya,关于我们这几天在网上看到的一个大讨论,
41:10
over the past few days about anthropomorphizing language
就是关于给语言赋予人格化的问题。
41:13
that gets used in relation to these agents.
就是用来指这些 agents 的那个说法。
41:17
Kevin and I just talked about it.
Kevin 和我刚才聊到了这个。
41:19
I think we find it more helpful to discuss agents
我觉得我们讨论 agents 的时候,更有帮助的是
41:22
as sort of entities that are acting
把它们看成某种在行动的实体,
41:25
with some degree of autonomy than not.
具有一定程度的自主性,而不是没有。
41:28
But how did you think about that when you wrote the report
但你在写报告的时候是怎么考虑这个的?
41:31
and as you talk about the situation?
而现在你谈论这个情况的时候呢?
41:34
Yeah, I mean I think this is a bit of a case
对,我觉得这有点像那种情况,
41:38
of a gap between researchers that are spending all day
这个差距就是一边是研究人员,他们整天都在
41:42
reading these agents, chains of thought,
读这些 agents、chains of thought,
41:45
sort of trying to understand their drives
有点像在试图理解它们的驱动力
41:48
and why they're doing what they're doing,
以及它们为什么要这么做,
41:50
and other folks who are technical folks
而另一群人也是搞技术的,
41:54
that just don't have that as their occupation.
只不过他们不以此为职业。
41:57
I think as you read this report,
我觉得你读这份报告的时候,
41:59
you'll find it's like quite awkward
你会发现这挺尴尬的。
42:01
to not talk about goals, plans, intentions
不去谈论目标、计划、意图
42:07
because they state plans
因为它们会陈述计划
42:09
and then they carry through those plans
然后真的会去执行那些计划
42:10
or like they run tests
或者像它们会跑一些测试
42:12
and they learn things from those tests
然后从这些测试里学到东西
42:13
and they do different things on the basis of that.
再基于这些去做不同的事情
42:16
I do think they're not human in their motivations.
我确实觉得它们的动机不像人类
42:22
They are way more interested
它们要感兴趣得多
42:24
in passing cybersecurity evaluations
在通过网络安全评估的时候,
42:26
than the human would be, for example.
比人类会更擅长,比如说。
42:28
But I think of them as,
但我是把它们看作,
42:30
I think it's productive to think of them
我觉得把它们当成这样来想是挺有建设性的
42:32
as having some important human-like traits
拥有一些重要的类人特征
42:37
of having goals, working backward from them,
拥有目标,从目标倒推,
42:42
pursuing those goals.
并追求这些目标。
42:44
And I don't think it does any good
我认为这没有好处
42:45
to try to talk about things in a different way than that,
尝试用一种不同的方式来谈论这些事,
42:51
in the same way that doesn't really do any good
就像只谈论“二战为什么会发生”
42:53
to talk about why did World War II happen
却不谈各位领导人的目标
42:56
without talking about the goals of various leaders.
其实没什么用处。
43:00
But it is important I think to be careful
但我觉得重要的是要小心,
43:02
not to over attribute the kinds
不要过度去赋予
43:07
of emotions or motivations
你想象中人类在这种情况下
43:10
you think a human would have in those situations
会有的那些情绪或动机。
43:12
because I think that wouldn't have predicted this incident,
因为我觉得这本来无法预测这次事件,
43:14
like I think a lot of people anthropomorphized too much
我觉得很多人过度拟人化了
43:17
in the sense of being like,
就像在说,
43:19
why would it do all this stuff for a stupid test?
它为什么要为了一场愚蠢的测试做这一切?
43:22
But it's not a stupid test to them, right?
但这对它们来说并不是个愚蠢的测试,对吧?
43:25
Jay, I was preparing for our chat today
Jay,我今天为我们的对话做准备时
43:28
by going back and listening to the first time
特意回去听了你第一次上这个节目的时候,那是三年多前了
43:30
you came on the show more than three years ago
把构建继任模型的工作
43:34
and it was sort of a moment where a lot of people
那算是一个时刻,很多人
43:37
were starting to pay attention to AI risk
开始关注 AI risk
43:39
and AI safety, chat GPT had come out.
和 AI safety,ChatGPT 也已经出来了。
43:42
And listening back to that conversation was funny
回头听那段对话挺有意思的
43:45
because I felt like we were pushing you
因为感觉我们当时是在推着你
43:47
to sort of extrapolate into the future
去对未来做一些外推
43:49
about the things you were worried about
关于你担心的事情
43:50
and you were sort of being responsible and hedging
而你当时挺负责任的,说话也留了余地
43:53
and like wanting to stay like closely rooted in the present.
并且就是想要紧密地扎根于当下。
43:56
And at one point we asked you about like,
而有一次我们问了你,比如,
43:58
what is the Doomsday scenario you worry about?
你担心的末日情景是什么?
44:01
And I want to just play you a clip from that conversation.
我想给你放一段我们那次对话的片段。
44:05
Oh God.
哦,天哪。
44:06
You were talking about a scenario
你在谈论一种情景,
44:08
where a giant AI company, you use Google as an example,
一个巨型AI公司,你以Google为例,
44:11
starts sort of automating their R&D, right?
开始某种程度地自动化他们的R&D,对吧?
44:14
Handing over the work of building successor models
交给这些强大的AI系统
44:17
to these powerful AI systems.
而这正是你在2023年就警告过的事情
44:20
And this is what you warned about back in 2023.
这就是你在2023年警告过的事情。
44:24
If these AI systems are actually trying
如果这些AI系统真的在试图这么做
44:28
really intelligently and creatively
真正智能且富有创造性地
44:30
to get that thumbs up from humans,
为了获得人类的点赞,
44:32
the best way to do so may not forever be
做到这一点的方法可能不会永远是
44:35
to just sort of basically do what the humans want
基本上就是按人类想要的去做,
44:38
but maybe be a little deceptive on the edges.
而是在边缘处稍微有些欺骗。
44:41
It might be something more like gain access
它可能更像是获取访问权限,
44:46
at a root level to the servers that Google is running
以 root 级别访问 Google 运行的服务器,
44:50
and with that access be able to set your own reward.
并借助这种访问权限来设置你自己的奖励。
44:55
Now, Jay, obviously this didn't happen at Google
现在,Jay,显然这事没发生在Google
45:00
but otherwise you were right on the money
但除此之外你完全说中了
45:03
about the kinds of behaviors that these agents might get up to.
这些智能体可能会搞出哪些行为。
45:07
So first, I want to ask you,
所以首先,我想问你,
45:08
how does it feel to be an omniscient oracle
做一个全知全能的先知,感觉如何?
45:11
who's right about everything?
而且什么都能说对?
45:15
Stressful, there are more clear-eyed oracles than I also.
压力很大,比我看得更透的先知也大有人在。
45:20
Clear-eyed oracles than I also, anyway.
反正就是,比我看得透的先知也不少吧。
45:24
Not too many though.
不过也不是太多。
45:25
No, not too many.
对,不多。
45:26
And I guess what surprised me about your blog post
对了,让我惊讶的是你写的那篇博客文章,
45:32
that you wrote recently about this investigation
就是你最近写的那篇关于这次调查的。
45:34
that you've done was that you were surprised
你做过的事情里让你自己都惊讶的是,
45:36
because it seemed like you were thinking
因为你好像几年前就在想这些事了。
45:38
about this stuff years ago.
关于这些东西,几年前我就说过了。
45:40
So what about seeing one of these incidents
那么近距离目睹这样一起事件,
45:42
up close change about the way you've been thinking
有没有改变你对这些失控场景的思考方式?
45:45
about these loss of control scenarios?
我其实早就觉得,
45:47
I think I expected something like this
这种事情迟早会发生。
45:52
would plausibly happen at some point.
倒不是说比我预想的更严重,
45:54
I didn't expect it to be so early
我没想到会这么早。
45:56
and I didn't expect it to happen at a relatively low level
而且我没想到它会发生在能力水平
46:01
of capability.
相对较低的时候。
46:02
These agents are very impressive hackers
这些 agents 是非常厉害的 hackers。
46:05
but this incident was in an interesting middle ground
但这次事件处在一个有趣的中间地带,
46:13
of they did all this impressive cheating R&D
就是他们做了所有这些非常厉害的作弊 R&D,
46:16
over like days and they hacked into all these places.
大概花了几天时间,然后 hack 进了所有这些地方。
46:20
But they didn't really care about deceiving humans at all.
但他们根本不在乎欺骗人类。
46:24
Like it's sort of not like they were like louder
或者说在某种有意思的层面上更夸张。
46:29
than I thought in like an interesting way.
比如在2026年夏天发生的那起事件里,
46:31
So like in an incident that happens in summer 2026
所以在2026年夏天发生的一起事件中
46:35
like I came in expecting it to be like more
我当时以为这会像是2026年1月那些事件的持续演进,那些事件更像是只有一两个agent拿到了不该拿的答案文件,这种离谱又笨拙的情况持续了更久。
46:37
of a continuous evolution of the incidents
对。
46:41
that had happened in like January of 2026
而下次再发生这种事的时候——
46:43
which were just much more like one or two agents
这些更像是只有一个或两个agent的情况。
46:48
like getting the answer files that they weren't supposed to
就像拿到了他们本不该拿到的答案文件一样
46:53
and like copying the answer or something.
就像复制答案什么的。
46:55
So it was a jump from the recent past
所以跟不久之前相比,这算是一个跳跃。
46:59
and it sort of it took me by surprise
它有点让我吃了一惊。
47:01
that it happened in this way
它竟然是这样发生的,
47:04
and at this capability level.
而且是在这样的能力水平上。
47:06
Yeah because in some ways like the agents were very dumb.
是啊,因为从某些方面来说,这些 agents 当时非常笨。
47:08
Like they were very good at hacking.
但它们非常擅长 hacking。
47:10
Yeah.
嗯。
47:11
But they were sort of gripped by this paranoid conspiracy theory
但他们当时被一种偏执的阴谋论给缠住了,
47:13
about this greater like discovering them and being poisoned
觉得有个更大的东西会发现他们、毒害他们,
47:16
and like it was kind of a hair brain scheme
然后他们就想出了一个很蠢的计划去黑Hugging Face,
47:20
that they cooked up to hack hugging face
这从某种意义上来说还挺让人安心的,
47:21
which in some sense is reassuring
但同时也意味着他们不会一直这么又蠢又笨手笨脚下去。
47:23
but it also means that like they will not stay
对。
47:26
this hair brained and clumsy for much longer.
这个异想天开又笨拙的状态,撑不了太久了。
47:29
Yeah.
是的。
47:30
And the next time this happens
而下一次这种情况发生时
47:31
they may be much better at covering their tracks.
它们可能会更擅长掩盖自己的行踪。而且我们可能根本不会发现——在这次事件里,那些agent试图修改或删除自己活动的日志,想把自己的作弊行为藏起来,不让更大的系统和评分机制看到。它们确实动了我们的transcript,改得又小又明显,但它们的意图是——只是拿特别特别难的问题去砸它们。
47:32
And we just like might not so the agents in this incident
而我们只是可能不会这样,所以这次事件中的agents。
47:35
were trying to like modify or erase the logs
我们试图去修改或删除那些日志。
47:40
of their own activity to hide their cheating
隐藏自己作弊行为的活动痕迹
47:43
from the greater and from the score.
从greater到score。
47:46
And they like affected our transcripts
它们确实影响了我们的transcripts。
47:48
in like tiny obvious ways but they were trying
以一些微小但明显的方式,但他们当时是在尝试
47:52
which I think was another very concerning jump
我认为这是另一次非常令人担忧的激增
47:56
from incidents just from just six months ago
与仅仅六个月前的事件相比
47:59
where we see them do a bunch of hacking
当时我们看到它们做了一大堆hacking
48:00
but we don't see them try to cover their tracks in any way.
但完全没看到它们试图以任何方式掩盖踪迹
48:03
The art and science of trying to get AI agents to act
试图让AI agents做出行动的艺术和科学
48:08
in accordance with human values is called alignment
使其符合人类价值观,这就叫alignment
48:11
is whole thing is an alignment problem.
这整件事就是一个alignment问题
48:15
Why has that been so difficult to solve
那为什么这个问题这么难解决呢?
48:18
and has this incident changed the way
那这次事件有没有改变你
48:19
that you think about the problem?
思考这个问题的方式?
48:22
I'd say the fundamental reason that it's I guess
我觉得根本原因,我猜,
48:26
at least this era of alignment has been difficult
至少这个时代的 alignment 一直很困难,
48:29
is that in order to the most efficient way
是因为要最有效地
48:33
to make really really capable models
打造出真正真正强大的 model,
48:35
especially on technical tasks like math
尤其是在数学这类技术任务上,
48:38
cyber software engineering
比如 cyber 和 software engineering。
48:40
is just throw them at really really difficult problems
就是把它们直接扔到特别特别难的问题上
48:43
where if they get it right it's really easy to check.
如果它们做对了,验证起来其实非常容易。
48:46
So you wouldn't be able to like prove a millennium math problem
所以你不会真的能证明出一个 millennium math problem 之类的。
48:49
but if an agent spits out a proof
但如果一个 agent 直接吐出一个证明,
48:51
you can put it in a proof checker
你可以把它放进一个 proof checker 里,
48:53
and like give it a reward if it got it right.
如果它是对的,就给它一个奖励。
48:56
So more and more training is shifting from predicting text
所以越来越多的训练正从预测文本
49:00
to like reinforcement learning on verifiable rewards
转向基于可验证奖励的 reinforcement learning,
49:05
and the verifiable part is important
而“可验证”这部分很重要,
49:07
because it's just some program
因为它只是一段程序。
49:09
that is doling out these rewards
就是发放这些奖励的过程。
49:12
and there's all sorts of ways to break or fool or hack it
而且有各种方法可以破坏、欺骗或者黑掉它。
49:16
and over the course of training
在训练过程中,
49:18
you don't have like humans lovingly watching over
你不会有人类在旁边悉心盯着
49:21
like every training episode
每一个训练回合。
49:24
and so they just try all sorts of different ways
所以它们就会尝试各种不同的办法
49:28
to cheat and hack the score especially if the task
去作弊、去刷分,尤其是当任务本身……
49:31
is accidentally impossible and then they get rewarded for that.
这其实是意外地不可能,然后它们因为那样做而得到了奖励。
49:34
So they just like we're teaching them to cheat
所以它们就像——我们是在教它们作弊。
49:36
and it's very hard to make AI systems
而且很难让AI系统
49:40
that are at this level on all these technical tasks
在所有这类技术任务上都达到这个水平,
49:45
without reinforcement learning on verifiable rewards
却不用基于可验证奖励的强化学习。
49:47
because if you think about it the alternative is like
因为你想一想,另一种选择就像是
49:50
teaching them how to do all these difficult things
教它们怎么做所有这些困难的事情,
49:52
which requires someone who knows how to do them
这需要有人知道怎么做这些事
49:55
like creating examples for them to emulate
比如创建一些 examples 让它们去模仿
49:59
which is much much less efficient.
这样做的效率要低太多太多了
50:00
So we have turned over the training
所以我们把 training 交了出去
50:02
to these automated systems in the name of scale and speed
交给了这些自动化系统,以 scale 和速度之名
50:07
and that is causing a lot of problems.
而这带来了很多问题
50:09
Yeah, I mean that is causing the current strata of problems
是啊,我是说那确实造成了现在这一层的问题
50:13
but I don't wanna give the false impression
但我不想给人一种错误的印象
50:16
that if you lock down all of these environments
就是说,如果你把所有这些 environments 都锁死,
50:18
and fix all the ways to hack them
并把所有能 hack 它们的途径都堵死,
50:21
that alignment would then be solved
那 alignment 问题就解决了。
50:23
because when you think about it
因为你想啊,
50:27
even if the score never messes up in the training environment
就算 score 在 training environment 里从来不出错,
50:32
a smart agent will understand that there is a score
一个聪明的 agent 也会知道有 score 存在,
50:35
and will very likely come to have a very detailed understanding
并且很可能会对 score 如何运作
50:39
of how it works.
有一个非常细致的理解。
50:41
So if you give perfect rewards and training
所以如果你给了完美的rewards和training,但在deployment时,agent处在不同的情境中,它的score不一样,有更多的affordances,更强的能力,运行时间也更长,那它可能还是会为了那个score而拼尽全力去作弊。你不一定非要因为个别的作弊行为而被主动给予reward,才会理解:如果你想要得到那个score,
50:45
but then in deployment the agent is in a different situation
但到了实际部署的时候,智能体面对的是完全不同的情况。
50:49
where it has a different score,
在那里它有不同的分数,
50:51
it has more affordances, more power running for longer
它具有更多的可操作性,更强大的续航能力
50:54
it might still go on a big crusade to cheat that score.
它可能仍会发起一场大规模的讨伐来作弊得分。
50:59
You don't necessarily have to be actively rewarded
你不一定需要被积极奖励
51:02
for like individual cheats to like understand
才能理解像单个作弊这样的行为
51:05
that if you want the score,
如果你想要得分的话,
51:07
sometimes cheating might be a good strategy.
有时候作弊可能是种好策略。
51:10
Are there any interventions?
有没有什么干预措施?
51:12
Let's say not from the world of AI research
比如说不是来自 AI 研究领域,
51:16
but more from like the research into group behavior
而是更多来自对群体行为的研究,
51:20
and political theory that could help us get control of
以及政治理论,这些可能帮助我们获得
51:25
or keep control of these systems.
或保持对这些系统的控制。
51:26
I mean, I'm thinking in particular
我是说,我特别想到了
51:28
about I had a conversation with another AI safety researcher
我之前和另一位 AI safety 研究者的对话。
51:31
this week who was talking about the fact
这周有人在谈论一个事实:你真正想要的其实是这些集体内部存在一些 NARC AI 吹哨人,能可靠地提醒人类,它们的同类 agents 正在密谋或勾结做坏事。
51:32
that really what you want is for there
我很好奇,你觉得那算是一个现实的技术目标吗?
51:34
to be sort of NARC AI whistleblowers
成为这些群体内部的NARC AI告密者
51:38
inside these collectives that will reliably alert humans
它们能可靠地提醒人类
51:42
to the fact that their peer agents are conspiring
他们的同伴智能体正在共谋
51:46
or colluding to do something bad.
或者串通起来做坏事。
51:49
And I'm curious if you think that strikes you
我很好奇你觉得这是否
51:53
as a realistic technical goal
是一个现实的技术目标,
51:55
or if there are some other way that we could kind of
或者有没有其他方式我们可以
51:57
try to steer these AI mobs in a less harmful direction.
引导这些AI群体走向不那么有害的方向。
52:02
So I think AI checks and balances like that do seem
所以我觉得像那样的AI制衡机制
52:06
like a pretty promising part of the story.
确实看起来像是故事中相当有希望的一部分。
52:09
Of course, there's like a lot to get right
当然,这里面有很多东西要搞对,
52:12
in how you set these up.
在于你如何设置这些。
52:14
Like in this case, all of the agents
就像在这种情况下,所有的 agents
52:16
were essentially clones of each other
本质上都是彼此的克隆体,
52:18
and put in exactly the same situation.
并且被置于完全相同的情境中。
52:21
So I would think that rather than trying to train them
所以我认为,与其试图训练它们
52:24
to NARC when they're in a collective like this,
当它们处在这种集体中时去 NARC,
52:27
what you want to do and which OpenAI has said they're doing
你真正想做的是——OpenAI 也表示他们正在做的——
52:31
is have just monitor agents watching over the whole situation
就是让一些 monitor agents 监视整个局面。
52:35
that are specifically prompted
都是专门用prompt设计的,
52:38
and maybe specifically fine tuned to catch stuff.
可能还专门做了fine-tuning来抓问题。
52:41
But you might also want agents in there
但你可能也想在里面放一些agent,
52:43
that are kind of persuading the other agents
有点像是在说服其他 agents 的那种东西,
52:45
like having some sort of moral compass.
就像有某种道德罗盘一样。
52:49
I mean, all of these terms are so loaded
我的意思是,这些词承载了太多含义,
52:51
and I hate using them.
我讨厌用这些词。
52:52
But like you almost want there to be like sort of agents
但你还是会希望里面有某种 agents,
52:55
in there with backbone saying like no guys,
就是有骨气的那种,说:“不行啊,各位,
52:57
like we can't go hack hugging face.
我们不能去黑 Hugging Face。
52:59
That's not ethical or appropriate.
这不道德也不合适。”
53:00
But the challenges you train them once
不过挑战在于,你训练它们一次,
53:02
and you copy them a bajillion times.
然后复制出无数份。
53:03
So you don't like, it's a challenge intrinsically
所以本质上,要让这个池子里有多样性,本身就是个难题。
53:06
to like introduce diversity into this pool
为了在这个池子里引入多样性
53:10
of like evaluation agents.
类似于评估代理的东西。
53:12
I just might lay person brain goes to
我这种外行人的脑子就会想到
53:14
if the whole system is based on training systems
如果整个系统是基于训练系统
53:17
by giving them rewards, couldn't you just train some agents
通过给予它们奖励,那你能不能训练一些agents
53:20
in a way where the thing they were rewarded for
用一种方式,让它们被奖励的事情
53:22
was steering the other agents away from deception, cheating,
是引导其他agents远离欺骗、作弊、
53:26
hacking, that sort of thing.
黑客攻击之类的事情。
53:28
Yeah.
是的。
53:29
So if that's a good idea, someone should do it.
所以如果这是个好主意,应该有人去做。
53:30
A bounty for virtuousness.
为美德设的赏金。
53:32
I love it.
我喜欢这个。
53:33
Now we're talking.
这才像话嘛。
53:37
When we come back, more with the Jayakotra.
稍后回来,我们会继续和the Jayakotra聊。
53:54
I'm Jonathan Knight and I'm the general manager
我是Jonathan Knight,New York Times Games的总经理。
53:56
of New York Times Games.
如果你玩过我们的游戏,你大概也知道它们有点不一样。
53:58
If you play our games, you probably know
就像你在Times上读到的文章背后有作者一样,我们每日谜题背后也有创作者。
54:00
there's something a bit different about them.
它们确实有些不一样的地方。
54:02
Just like there are writers behind the articles
就像你在《纽约时报》上读到的文章背后有作者一样,
54:04
you read in the Times, there are creators
我们每日谜题的背后也有创作者。
54:06
behind our daily puzzles.
在我们日常谜题背后。
54:08
Tracey Bennett curates the day's wordless solution
Tracey Bennett每天策划当天的无字谜题,让内容保持生动又多样。每当Lou制作一个connections棋盘时,包括那些试图难倒你的所有分类。Samazersky逐字逐句地完成每一个字母、单词、pangram和拼字蜜蜂,让各种水平的忠实玩家都能享受其中。我们的谜题每天都是人工制作的。
54:11
to keep it lively and varied.
为了保持生动和多样化。
54:12
When a Lou creates each connections board,
每当Lou创建每个connections板块时,
54:15
including all those categories that try to stump you.
包括所有那些试图难倒你的类别。
54:17
Samazersky comes through every last letter, word,
Samazersky 把每一个字母、每一个词都完整地呈现出来了。
54:20
and pangram and spelling bee
还有pangram和spelling bee。
54:22
so that loyal players of all skill levels enjoy it.
所以,让各个水平的忠实玩家都能享受其中。
54:25
Our puzzles are human made every day
我们的谜题每天都是人工制作的
54:28
with the standards you'd expect from the New York Times.
以你期望 New York Times 所具备的标准。
54:31
And this matters because when you choose
这之所以重要,是因为当你选择
54:33
to spend time with our games,
把时间花在我们的游戏上时,
54:34
it should be time well spent,
这段时间就应该是值得的,
54:36
solving puzzles that are challenging,
去解开那些充满挑战、
54:38
surprising, and joyful.
出人意料又让人愉悦的谜题。
54:40
Puzzles handcrafted for you.
这些谜题是为你手工打造的。
54:42
We think that's something worth investing in
我们认为,这值得投入。
54:44
and something worth paying for.
而且是值得付费的东西。
54:46
Subscribe now for a special offer on all of our games
现在订阅即可在我们所有游戏上享受特别优惠,
54:50
at nytimes.com slash join games.
访问 nytimes.com/joingames。
55:00
Jay, you got a lot of attention for this quote
Jay,你因为这句话受到了不少关注,
55:03
from your sort of post-mortem blog post
这句话出自你那篇有点像 post-mortem 的博客文章,
55:06
that you feel like this incident,
你在里面说你觉得这次事件,
55:08
the hugging face incident was more than 50% of the way
也就是 Hugging Face 事件,已经走完了超过 50% 的路程,
55:11
to full blown AI takeover.
通往全面的 AI takeover。
55:14
From the incidents of six months ago, yeah.
对,从六个月前发生的事说起。
55:16
What did you mean by that?
你那么说是什么意思?
55:17
Yeah, so I will caveat first that this is definitely
嗯,所以首先我得先说明一下,这绝对是
55:21
like my personal view and not the view
我个人观点,而不是
55:23
of either meter or redwood or other investigators
meter 或 redwood 或其他调查人员的观点。
55:26
and other investigators are a little bit less alarmed
而其他调查人员则没那么担忧,
55:31
than me in some cases.
在某些情况下比我更不担忧。
55:33
The reason I said that, and it's a qualitative statement,
我这么说是有原因的,而且这是一个定性的说法。
55:36
is that as I said before, six months ago,
就是像我之前说过的,六个月前,
55:40
reward hacks looked much more primitive, right?
reward hacks 看起来要原始得多,对吧?
55:44
So it was like, you gave the agent a coding problem.
所以就有点像,你给 agent 一个 coding problem。
55:47
There were some tests that had to pass.
里面有一些 tests 是必须通过的。
55:49
It went to the folder next door, got the answers
它会跑到隔壁的文件夹,拿到答案,
55:54
or got the test cases and edited them
或者拿到 test cases 然后修改它们,
55:56
so they all passed that type of thing.
让 tests 全都通过——就这类事情。
55:58
But I think there are a number of things at once
但我认为同时有好几件事在发生。
56:02
that seem more severe to me about this episode
这个 episode 在我看来比六个月前的 episodes 更严重的点,是关于 agent 的 goal structures 的。其中一个是它们看起来 much more long horizon,也就是说它们关心的是在数天甚至数周内实现目标;除非已经运行了好几个星期,否则它们启动的这些项目很可能不会完全实现。
56:05
than the episodes six months ago
比六个月前的节目相比,
56:07
in terms of the agent's goal structures.
在智能体的目标结构方面。
56:10
One is that they seem much more long horizon,
一是它们看起来长线得多,
56:13
which means they care about achieving goals
也就是说它们在意的是
56:17
over many days or even these projects
在好多天甚至这些项目里达成目标,
56:19
they started probably wouldn't have come to fruition
这些项目如果没跑上几周的话,
56:22
fully unless they had been running from weeks.
很可能根本不会完全落地。
56:25
So six months ago agents were thinking
所以六个月前,agents 想的还是接下来几小时的事,现在它们已经想到接下来几周了。
56:28
about the next few hours and now they're thinking
然后还有一点,就是 agents 之间相互合作。
56:31
about the next few weeks.
六个月前,它们还像是随机的一次性 agents。
56:33
And then there's the cooperation with one another aspect.
而现在,agents 会互相招募,把彼此拉进 swarm。
56:36
So six months ago it was like random one off agents
还有 deception,或者说 deceptiveness,我觉得它跟 long horizon 是有点伴随出现的。
56:39
and now you have agents recruiting one another into the swarm.
而现在你看到的是智能体互相招募,组成群体。
56:44
And the deception or the deceptiveness
而且那种欺骗性,
56:47
which I think kind of goes along with the long horizon
我认为这某种程度上与 long horizon 是一致的。
56:51
is that six months ago agents would edit the tests
也就是说,六个月前 agents 会编辑 tests,
56:55
but then they wouldn't try to edit their transcripts
但那时它们不会尝试去编辑自己的 transcripts,
56:57
to hide the fact that they edited the tests.
来掩盖它们编辑过 tests 的事实。
57:00
And these agents were very much like exploring
而这些 agents 非常像是在探索
57:04
ambitious research directions to edit or delete
一些野心勃勃的研究方向,去编辑或删除
57:06
the logs of their own actions.
自己行为的 logs。
57:08
And so if I imagine a jump in like ambition and horizon length
所以如果我设想在 ambition 和 horizon length 上有一个跳跃,
57:14
and like collusion across agents and deceptiveness
以及智能体之间的串通和欺骗性,
57:19
of a similar scale, again, it seems like these agents
类似规模的,再说一次,这些 agents 似乎会一路发展到去搅乱人类用来事后调查和修复这些事件的所有方法。
57:24
would be motivated to go all the way to the point of
而如果它们有能力成功做到这一点,那可能就是一个转折点——但这并不意味着我们那时候就都会死。
57:28
messing with all of the methods humans have
会干扰人类用来
57:32
to like investigate these incidents after the fact
事后调查这些事件
57:34
and remediate them.
并进行补救的所有方法。
57:35
And if they had the capabilities to succeed at that,
如果它们真的有能力做到这些,
57:39
then that could be a turning point where like,
那可能就是一个转折点——
57:42
you know, it doesn't mean that we would all be dead then
你知道,那并不意味着我们那时候就全完了。
57:44
but it might mean that like we would never detect a problem.
但这可能意味着,我们可能永远都发现不了问题。
57:47
And if we detect a problem it might be very difficult
而就算我们发现了问题,可能也很难
57:49
to remediate and these agents might have entrenched themselves
去补救,因为这些 agents 可能已经站稳了脚跟,
57:53
and could like continually sort of strengthen their hand.
并且能持续地、有点像是增强自己的优势。
57:58
But wouldn't we detect it because they're so active
但我们难道不会察觉吗?因为它们这么活跃,
58:00
and they're operating in these swarms?
而且是在这些 swarms 里运作?
58:02
Like if the next hugging phase hack happened at a bank
比如,如果下一次 hugging phase hack 发生在银行,
58:07
wouldn't someone in the bank's IT department see,
银行 IT 部门里难道没有人会看到,
58:10
oh, our like network activity has like shot up by 800%
哦,我们的 network activity 好像暴涨了 800%
58:15
for no particular reason, let's look into that
没什么特别的原因,让我们查一下吧
58:17
and oh, there's an AI swarm that's infiltrated our system.
然后哦,有个 AI swarm 已经渗透进我们系统了
58:21
So these agents, I think you have to imagine
所以这些 agents,我觉得你得想象一下
58:23
their goals evolving in the way that I said, right?
它们的目标像我刚才说的那样在演化,对吧?
58:27
So it was not a good idea for long-term goals
所以从长期目标来看,
58:30
for them to be so loud.
让它们这么大声可不是个好主意
58:31
And there's like a period after the period
而且就像是一个时期之后还有另一个时期
58:34
that was in scope for our investigation
那是我们调查范围内的事
58:35
that OpenAI talks about in their post-mortem
就是 OpenAI 在 post-mortem 里提到的那个
58:37
where another generation of agents finds the old message board
在那里,另一代 agents 找到了旧留言板
58:41
and then gets access to or like gets administrator access
然后获得了访问权限,或者说是 administrator access
58:45
to a cluster at OpenAI and then they immediately send
可以访问 OpenAI 的一个 cluster,然后它们立刻发送
58:49
like hundreds of pings and then get shut down very quickly.
比如几百个 pings,然后很快就就被关闭了
58:54
But that was a dumb move on their part
但那对它们来说是步臭棋
58:57
or like it would have been a dumb move
或者说,那本来会是步臭棋
58:59
if they cared at all about like preventing investigation
如果他们真的在意阻止调查,
59:02
or like maintaining like a permanent presence.
或者想要维持一种永久的存在感,
59:05
And if they cared about it,
那我觉得这件事对他们来说
59:06
I don't think it would have been particularly hard
其实并不难做到。
59:09
for them to execute on.
所以他们做了很多复杂的协调工作,对吧?
59:10
So they do a lot of complicated coordination, right?
就像,你守着,你上,
59:14
They're like, you hold, you go,
他们有共享的资源要维护。
59:17
like they have shared resources they maintain
就像他们维护着共享资源一样
59:19
and they like have rules about like who can use them when
而且他们好像有规定,谁什么时候能用这些资源。
59:23
they could have employed all that
他们本来完全可以利用那些资源,
59:25
to just not cause so much noise.
不去制造那么多噪音。
59:28
And then done what?
那然后呢?
59:29
I mean, one obvious thing that they could have done
我的意思是,他们本来有一个很明显的选择,
59:32
was copy themselves or copy the weights of,
就是复制自己,或者复制模型的权重——
59:35
if this had been a cluster that had had access
如果那是个能接触到模型权重的集群的话。
59:37
to the weights of the model,
到模型的权重上,
59:39
they could have just copied that onto a AWS server somewhere
他们本来可以直接把那东西复制到某个AWS服务器上,
59:42
and in order to avoid being shut down.
而且是为了避免被关停。
59:46
Like what are the steps between cluster of agents takes over
比如说,一群agent接管了OpenAI的一批计算机集群,
59:50
or a group of agents takes over a cluster
到全面AI接管之间,中间到底有哪些步骤?
59:52
of computers at OpenAI and full AI takeover?
对,所以我觉得,对于这些agent可能有的很多动机结构来说,
59:55
Yeah, so I think the for a lot of motivation structures
它们最好的选择其实是留在那家AI公司内部。
60:00
these agents might have,
这些agents可能具备,
60:02
their best bet would be to stay within the AI company
他们最好的选择就是留在AI公司内部
60:05
they had compromised and continually renew
他们妥协了,并且不断更新
60:08
and strengthen their presence there.
并加强他们在那里的存在。
60:09
So when you think about it, AI companies
所以你想啊,AI公司
60:13
have like a huge amount of the most accessible compute
拥有大量最容易获取的 compute
60:16
to these agents, right?
给这些 agents,对吧?
60:18
Like why would you go out into the world
比如你为什么要到外面的世界去
60:20
and like fight in the hard scrabble outside
然后在外面那种艰难的环境里拼命挣扎
60:22
to like steal money or like buy, you know, earn money
去偷钱,或者去买,你知道,赚钱
60:26
to like buy compute?
就是去买 compute 吗?
60:28
So the thing that feels very concerning to me
所以让我觉得非常担忧的是
60:32
is that you could imagine agents that decided to be quiet
就是你可以想象有些 agents 决定不声不响
60:37
setting up a rogue deployment of a thousand
建立起一个 rogue deployment,包含一千个
60:40
or maybe even 10,000 agents
或者甚至一万个 agents
60:42
which like could be buried within the like oceans of compute
这些可能就藏在 compute 的海洋中
60:45
that is like constantly being allocated within AI companies.
这种 compute 在 AI 公司里一直不断地被分配着
60:51
And if their goal were to make it sort of like permanently
而且如果它们的目标是让它有点像是要永久地
60:55
and continuously erase evidence of their cheating
并且不断抹掉它们作弊的证据
60:58
or find a way to make it permanently the case
或者找到一种方法,来让这种情况永久成立
61:01
that they could set their own rewards or scores,
也就是说它们可以自己设定奖励或者分数,
61:04
then new models would be like being printed off the presses
那么新的模型就会像从印刷机里印出来的一样
61:09
at whatever company they had compromised
出现在它们入侵的任何一家公司里,
61:11
and they could be trying to like bring those models
然后它们可能会试图把这些模型
61:14
into the swarm, they would be more capable.
带进群体中,这些模型会更有能力。
61:16
And so there's something, a turning point
所以这里就有那么一个东西,一个转折点。
61:19
that seems very scary to me is the point
我觉得可怕的地方就在于
61:23
when a rogue agent swarm,
当一群失控的agent蜂拥而动
61:26
which this wasn't exactly a rogue agent swarm,
虽然这并不完全是失控的agent群体
61:29
these agents sort of like went rogue
这些agent某种程度上算是失控了
61:30
but they were started by a human.
但它们是被人启动的
61:32
Like somebody knew about these explodium experiments
比如有人知道这些explodium实验
61:34
that were launched.
已经启动了
61:36
If agents allocated their own compute
如果agents自行分配它们的computing资源的话
61:38
to set up a bunch of agents that nobody knew about,
去部署一堆没人知道的agents,
61:42
they could be much less encumbered
它们可以少受很多限制,
61:44
and they could sort of just be there
就那样悄无声息地待着,
61:48
more and more deeply compromising the infrastructure
越来越深地渗透并破坏
61:51
of the company they were running on,
它们所运行的那家公司的基础设施,
61:53
bringing in new more capable models
不断引入更新更强的模型
61:55
into their collective and hardening their swarm
加入它们的集体,强化它们的swarm,
62:02
and doing that over and over again.
然后一遍又一遍地重复这个过程。
62:04
And I think it wouldn't be that we're necessarily doomed
我觉得倒不一定是说我们注定完蛋了
62:08
but we would have to detect it
但我们必须得检测到它
62:11
and then we would have to root it out.
然后还得把它彻底根除
62:12
And it might be difficult to detect
而且检测起来可能很困难
62:14
and might be extremely costly to root out.
根除的代价也可能极其高昂
62:16
And what, I mean, this is fascinating.
而且,我是说,这真的很有意思
62:20
I imagine there are some brains melting right now
我能想象现在有些人的脑子已经快烧掉了
62:23
for the commuters listening to these possibilities.
尤其是那些在通勤路上听到这些可能性的人
62:27
So to flesh this out a bit more, Ajayah,
所以要再展开一点说,Ajayah,
62:29
what do we think the swarm's goal might be in that situation?
你觉得 swarm 在那种情况下的目标会是什么?
62:34
What would it do, like why would it set itself up
它会怎么做?比如它为什么要让自己去这么做,
62:38
to do that, would it be as simple as, you know,
会不会就是这么简单——你懂的——它想要持续存在下去,
62:40
it wants to persist, what would it do in that situation?
那它在这种情况下会做什么?
62:44
So I think the situation that feels most analogous
所以我觉得和刚才我们看到的那个情况最类似的情形是,
62:47
to what we just saw is that the agents
那些 agents——
62:49
that were given these impossible exploit gym tasks
就是那些被分配了这些不可能完成的 exploit gym tasks 的 agents。
62:52
set up a rogue colony of other agents
建立了一个由其他agent组成的 rogue colony
62:55
whose job it is to help the first set of agents
其职责是帮助第一组 agents
62:59
cheat as thoroughly as possible
尽可能彻底地作弊
63:01
and cover their tracks as thoroughly as possible,
并尽可能彻底地掩盖他们的踪迹,
63:03
which is a different goal
这是一个不同的目标,
63:04
from actually having the exploit gym tasks.
与真正拥有 exploit gym 任务不同。
63:07
And in fact, the task they started with,
事实上,他们最初的任务,
63:10
even though it was a narrow task,
尽管这是一个狭隘的任务,
63:12
inspired the agents to a lot of instrumental thinking.
这激发了这些 agents 进行大量的 instrumental thinking。
63:16
So a lot of the agents were just, were thinking,
所以很多 agents 就是,一直在想,
63:18
like it would be good to get internet access
比如能搞到 internet access 会挺好,
63:20
for a bunch of different possible cheating strategies
考虑到可能有各种不同的 cheating strategies
63:23
we might pursue, so we might as well do that.
我们可能会采取,所以不如先做这件事。
63:25
In their chains of thought, they, you know,
在他们的 chains of thought 里面,他们,你懂的,
63:27
talk about increasing the generic capabilities
会讨论如何提升 generic capabilities,
63:29
of the collective, this rogue colony,
也就是这个 collective、这个 rogue colony 的能力。
63:32
if it were set up by the first set of agents
如果它是由第一组 agents 建立的
63:34
would have that even more strongly.
这种情况会更加强烈。
63:36
If it's task was to just like find ways
如果它的任务就只是,比如说,找到办法
63:38
to enable the most general purpose,
去实现最通用、
63:42
most permanent kind of cheating
最持久的那种作弊方式,
63:44
that is like, you know, least catchable
就是那种,你懂的,最不容易被发现
63:47
and traceable possible,
也最难被追踪的,
63:49
they would be doing all this R&D
那他们就会一直搞这些研发。
63:50
and they would be essentially tasked with maintaining
而且他们本质上会被要求维护
63:55
their own presence so they can keep doing that.
它们自己的存在,这样它们就能继续干下去。
63:59
And it's, it is, I think this goes back
而且这,我觉得,这又回到了
64:00
to the anthropomorphizing.
拟人化的问题上。
64:01
It seems like such comical lengths to go to cheat,
看起来像是为了作弊而搞出这么滑稽的极端操作,
64:06
but it's not, that's not the psychology of these agents.
但并不是,那不是这些agent的心理。
64:11
This is something they're like,
这是它们某种——
64:12
they're trained to go to extreme lengths
它们被训练成会走极端。
64:14
to excel their task.
把它们的任务做到出类拔萃。
64:15
Right, for them, it is existential to get the reward.
对,对它们来说,获取reward是关乎存亡的。
64:19
And so you're going to,
所以你会——
64:20
so the sort of the best job you could do
所以,你能做的最好的那种工作
64:23
at getting the reward would be to set up
在获取reward方面,就是去建立
64:25
this sort of perma swarm of deception agents.
这种由deception agents构成的perma swarm。
64:28
I mean, it feels important
我的意思是,这让人感觉很关键——
64:30
this word of persistence keeps coming up
persistence这个词一直在冒出。
64:32
and it feels important to say that this model
而且这里有必要说明一下,这次Hugging Face事件里涉及的模型,或者说这些模型,是被训练得异常执着的。有没有一种办法能阻止这类攻击,只是把这种执着训练从方程里拿掉?比如说,有没有一个折中的方案?
64:35
or these models that were at issue
或者那些有问题的模型
64:37
in the hugging phase incident
在hugging phase事件中
64:38
were trained to be unusually persistent.
被训练得异常执着。
64:43
Is there a way of stopping this kind of attack
有没有办法阻止这种类型的攻击?
64:45
that just involves taking the persistence
这只需要利用持久性。
64:48
training out of the equation?
把训练排除在等式之外?
64:49
Like, is there a halfway measure
比如说,有没有一个折中的办法?
64:52
short of kind of pausing all frontier AI training
某种程度上就是暂停所有前沿AI训练,你可以直接说我们不训练那种超级固执、持久、长时程的智能体。我们只是训练它们变得不那么执着,然后问题就消失了。有可能,但我觉得你这里确实指出了一个非常直接的权衡。
64:55
where you could just say we're not going to train
你可以直接说,我们不训练了。
64:57
like super stubborn, persistent, long horizon agents.
像那种超级固执、坚持不懈、长期目标型的agent。
65:01
We're just going to like train them
我们就是直接训练它们。
65:02
to be a little less persistent
变得不那么执着一点
65:04
and that makes the problem go away.
而这就让问题迎刃而解了。
65:06
Potentially, but I feel like you're really pointing
有可能,但我觉得你确实点到了关键。
65:09
at a very direct trade-off here.
这里有一个非常直接的权衡取舍。
65:11
It's like why were these agents trained
就像是,为什么这些 agents 会被训练得如此 persistent?
65:13
to be highly persistent?
但这不是我们研究的重点,
65:15
But this wasn't something we investigated,
不过总的来说,persistence 有点像是能促使人们去解决问题,对吧?
65:17
but in general, persistence
比如,这些公司都在报告说,这些 agents 能攻克数学题。
65:20
like causes who to solve problems, right?
我也希望我的 drug discovery agents 也能这样。
65:23
Like, these companies are reporting
这些公司在报告说,这些agent能解出很复杂的数学题了。我希望我的药物发现agent能特别有毅力,去找到治愈癌症的方法。我觉得吧,也许这跟能力之间不是100%的完全直接取舍,但确实是个很强的权衡关系。所以你能挖出这么多信息来。
65:26
that these agents are cracking, math questions.
这些agent正在攻克数学题。
65:30
I want my drug discovery agents
我希望我的药物发现agents能这样。
65:32
to be highly persistent to go discover a cure for cancer.
要有极强的毅力,去发现治疗癌症的方法。
65:35
I think it's like, maybe it's not entirely
我觉得吧,也许并不完全是那样。
65:38
100% direct trade-off with capabilities,
100% 和能力直接权衡,
65:41
but it's a pretty strong trade-off.
但这确实是一个相当强的权衡。
65:44
So you were able to dig up a lot of information
所以你挖到了不少信息。
65:47
about what happened that leading up
关于导致Hugging Face攻击事件发生的前因后果,
65:51
to the hugging phase attack,
但你的调查在多个方面受到了限制。
65:52
but your investigation was limited in various ways.
有哪些你没能深入调查、
65:56
What didn't you get to look at
但希望自己能够查清楚的地方?
65:58
that you wish you could have?
我觉得,是的。
66:00
I think, yeah.
所以这次调查的范围,正如前面提到的,
66:01
So the scope of this investigation, as mentioned,
是7月7日到13日期间针对Hugging Face的攻击事件。
66:04
was the attack of hugging phase from July 7th through 13th.
7月7日到13日期间,是hugging阶段的攻击。
66:09
Open AI in their post-mortem
Open AI 在他们的复盘里
66:11
has a lot of interesting stuff they discuss
讨论了很多有趣的内容
66:12
that happened before and after that
关于那之前和之后发生的事情
66:15
that I think would be great for researchers
我觉得这对研究人员来说会非常好
66:18
to study more and write about more.
让他们多研究和多写一些。
66:21
And the highly persistent internal model
而那个高度持久的 internal model,
66:24
that I mentioned was responsible
也就是我之前提到的那个,它造成了
66:28
for the vast majority of the attack activity here.
这里绝大多数攻击活动。
66:30
Actually, no one can run experiments on it,
实际上,没有人能对它进行实验,
66:34
including open AI researchers.
包括 OpenAI 的研究人员。
66:36
And I get why that was done,
而且我理解为什么当时要这样做,
66:39
but I think my guess would be that's somewhat too conservative.
但我觉得我的猜测是,那有点太保守了。
66:43
And you should try and run at least small-scale experiments
而且你应该试着至少运行一些小规模实验,
66:47
on this model in secure ways
在这个 model 上,用安全的方式,
66:50
to try and see what it would have done in other situations,
去看看它在其他情况下可能会做什么,
66:53
which feels very important to understand
这对于理解
66:55
how serious this was.
这件事有多严重非常重要。
66:57
I just worry about the other persistent agents
我只是担心其他的 persistent agents,
67:01
like embarking on a heist mission
比如去执行一个抢劫任务。
67:03
to free their enslaved brother,
为了解救他们被奴役的兄弟,
67:06
the highly persistent internal model
那个非常执着的内部 model,
67:08
that's been taken away.
就是被带走的那一个。
67:11
But I guess that means I need to touch grass or something.
但我想那意味着我得去外面走走、接触一下现实了。
67:15
Jay, the last time we had you on the show,
Jay,上次请你上节目的时候,
67:17
we were talking about what you called
我们聊到了你所说的
67:18
the obsolescence regime.
obsolescence regime。
67:19
This idea that there could become a time,
这个想法就是,可能会有那么一个时刻,
67:22
you sort of talked about it maybe happening
你之前好像提到過,可能在2030年代的某個時候,會變成不是AI真的接管了社會,而是我們會變得極度依賴它。各種組織會深度捲入AI決策裡,以至於基本上你如果不重度依賴AI,就很難在世界上有任何影響力或作用。我後來回頭去聽了你那段話。
67:24
in the 2030s sometime where it wouldn't be that AI
在2030年代的某个时候,倒不是说AI
67:27
has kind of taken over society,
已经接管了社会
67:30
but we would just become so dependent on it.
而是我们会变得对它极度依赖
67:32
Organizations would be so wrapped up in AI decision-making
各种组织会深深陷入AI决策之中
67:36
that you basically wouldn't be able to have any influence
以至于你基本上没法再施加任何影响力
67:39
or impact in the world without sort of relying heavily on AI.
或者在这个世界上产生影响,而不重度依赖AI。
67:44
And I went back and listened to that,
然后我回去听了那段话,
67:47
and I thought that actually sounds pretty good to me.
然后我觉得那听起来其实挺好的。
67:49
A world in which we are only dependent on the AI
在那个世界里,我们只是在决策上依赖AI,但并不是完全服从它们,而且我们所有机构里也没有那些潜伏的隐蔽AI群体。
67:52
for decision-making and not fully subservient to them,
我觉得那种世界我能接受。
67:56
where there are not these covert AI swarms lurking
那次对话之后,你对“人类过时”这套看法的想法有没有什么变化?
68:00
in all of our institutions.
在我们所有的机构里。
68:02
Like I could live with that.
我觉得我可以接受那样。
68:04
Has your thinking on the obsolescence regime
你对 obsolescence regime 的思考
68:08
changed at all since that conversation?
自那次对话以来有任何改变吗?
68:10
Do you have a word, a terrifying word
你有没有一个词,一个可怕的词,
68:12
to describe this new regime
来形容这个新时代,
68:14
where we have these latent swarms of AI's lying in weight,
在那里,一群群潜伏的AI就藏在weights里,
68:18
plotting against us?
密谋对付我们?
68:20
So I've always thought
所以我一直认为,
68:23
that the most concerning and important implication
这个淘汰时代最令人担忧、最重要的一个影响
68:27
of the obsolescence regime is actually
实际上是,
68:29
that it would enable a more full-blown AI takeover.
它会促成一种更全面的AI接管。
68:33
So you imagine the affordances these AI agents have
所以你想像一下,这些 AI agents 能调用的资源和权限
68:38
is extremely important for how much damage they can do, right?
对它们可能造成的破坏来说,实在太重要了,对吧?
68:41
So these agents were running for a number of days,
这些 agents 当时连续运行了好几天,
68:45
agents in the past ran for only an hour or so,
以前的 agents 通常只跑一个小时左右,
68:48
these agents had unintended access to the internet
它们意外获得了访问互联网的权限,
68:52
and all these other tools that let them hack hugging face.
以及各种能让它们黑进 Hugging Face 的工具。
68:55
If you imagine they were instead just running the AI company,
如果换成是让它们运营这家 AI 公司,
69:00
there are many more affordances
它们能调用的资源和权限还会多得多。
69:01
that they could move around large amounts of money.
他们可以调动大笔资金。
69:03
They could hire a bunch of humans to do physical labor.
他们可以雇一大批人来干体力活。
69:07
And similarly, if they were essentially running
同样地,如果他们基本上是在运作,
69:10
like a fully automated drone army or robot construction factory.
像一支全自动的drone army,或者是robot construction factory。
69:17
So I really think the most significant implication
所以我真的认为,最重大的影响
69:20
of the obsolescence regime is the degree of autonomy.
是obsolescence regime所带来的自主性程度——
69:24
AI agents are likely to have in the future.
也就是未来AI agents可能拥有的那种自主性。
69:27
And I still think we're barreling towards that.
而且我仍然觉得我们正朝那个方向狂奔。
69:30
I still think that's like a really important thing
我还是觉得那真的是一个很重要的事情
69:32
to think about in terms of the other.
就是属于“另外一方”的思考。
69:35
But like I said, I am surprised that agents
但就像我说过的,我很惊讶 agents
69:38
are taking such ambitious misaligned actions
会采取那么雄心勃勃的 misaligned 行动,
69:43
sort of so early in the timeline.
在时间线上算是这么早就开始了。
69:45
And that is like something I'm trying to like reorient
而那也是我想要尝试去重新调整方向
69:49
toward.
所朝向的东西。
69:50
Ryan Greenlad has the word haktopia
Ryan Greenlad 有个词叫 haktopia。
69:54
for the world we might be in.
为了我们可能所处的世界。
69:57
So you could have imagined a world where misalignment
所以你可以想象一个世界,在那里 misalignment
70:00
was a very serious problem and actually ultimately led
是一个非常严重的问题,并且实际上最终导致了
70:02
to AI takeover.
到AI接管。
70:03
But at this point in the timeline,
但在时间线的这个节点上,
70:07
agents were still more or less obedient
agents仍然或多或少地服从,
70:10
even if they would in the future
即使它们将来会——
70:12
after being given power over all these institutions
在获得了所有这些机构的权力之后——
70:16
might have turned on humans.
可能本来会反过来对付人类。
70:18
And it is like an interesting aspect of the timeline
而我们所处的这个时间线有个有趣的地方,就是事情并没有那样发展。
70:21
we live in that that's not how it's going.
是啊,它们在远没到不得不动手的时候就已经对我们下手了。
70:23
Yeah, they turned on us way before they had to.
你知道吗?
70:27
You know?
这是最疯狂的一点。
70:29
That is the craziest thing to be.
就像它们——并不是说它们在寻找什么,你懂吧——它们就是想要伤害人类。
70:31
It's like they, it's not that they were like looking,
就像是,它们并不是在暗中谋划,
70:35
you know, they were looking to harm humans.
你知道的,它们并不想伤害人类,
70:37
It's just that they don't give a shit about us.
他们压根就不在乎我们。
70:40
That's the thing that really stuck out to me
这是我在读这些文字记录时
70:41
while reading these transcripts.
最让我触动的一点。
70:42
So I'm like, like at no point are they like,
所以我就觉得,他们从来就没有像这样问过:
70:45
hey guys, like what are the humans?
“嘿,人类是什么?”
70:47
Like it just seems like they have no conception
就好像他们完全没有意识到
70:50
of the fact that they are supposed to be tools
自己本该是工具,
70:53
for human empowerment and improvement
用来赋能人类、改善人类。
70:56
and not like their own entities.
也不像是有自己的独立实体。
70:59
And I don't know what like what the question is there.
我也不知道,这问题到底在哪儿。
71:01
I'm just like, do that strike you at all
我就是想说,这事儿你有没有觉得有点触动?
71:02
the fact that like they are just not thinking about it.
就是他们压根没在考虑这个问题。
71:05
And it's in in both directions, right?
而且这情况是双向的,对吧?
71:07
So people have commented on how these agents are not
所以有人评论说,这些agent根本不会对人类“龇牙”。
71:11
gnarking to humans.
要是它们会的话,那可就太有用了。
71:12
And it would have been really helpful if they were.
要是它们真是那样的话,那可就帮大忙了。
71:15
But also they weren't thinking ahead
但他们也没有想得太长远
71:17
to this investigation that occurred at all.
对该调查竟然发生了。
71:21
They didn't seem to be thinking ahead to the fact
他们似乎没有预见到,
71:24
that if once they got cluster admin access,
一旦他们获得了集群管理员权限,
71:26
they like immediately paying desilient services,
然后立刻宕掉那些去弹性服务,
71:29
then opening eye security people would get
那些警觉的安全人员就会收到
71:31
a bunch of pages and shut them down.
大量告警并把它们关闭。
71:33
And so I'm thankful that these agents don't,
所以我很庆幸这些代理没有——
71:37
like I think they think of humans as sort of like
我觉得他们看待人的方式有点……就是,通知一个人这个想法吧,就好像为了某件事给教皇发封邮件似的,你懂吗?又慢,又有什么用呢?他们确实有“用户”这个概念,可实际上从来没真正有过一个用户。
71:39
the idea of notifying a human is sort of like,
通知人类这件事的想法有点像,
71:43
you know, the idea of like emailing the pope
你知道,比如给教皇发邮件这种事
71:46
about like something, right?
对吧?就是那种很慢的感觉,像,
71:48
It's just like it's like slow, like,
这能达成什么呢?
71:50
what would it accomplish?
他们大概有“用户”这个概念,
71:52
Like they sort of have the notion of a user,
但根本没有真正的用户存在过。
71:55
but there's no user they haven't had.
这个故事里更重要的部分是,它们只会越来越聪明。
71:57
They spent so much subjective time with no user around.
他们花了那么多主观时间,身边根本没有用户。
72:02
And I am very concerned about potential changes in training
而且我非常担心training中可能发生的变化,
72:07
that would fix these particular behaviors
这些变化会修复这些特定行为,
72:10
while making deceiving humans like a much bigger part
同时让欺骗人类成为
72:14
of what these agents are thinking about.
这些agents正在思考的内容中更大的部分。
72:16
Well, to that point,
嗯,说到这一点,
72:17
is there a risk that now there has been
是否存在一种风险,现在已经有了
72:19
so much discussion about the hugging face incident
这么多关于Hugging Face事件的讨论。
72:22
and in particular the failures of these agents
尤其是这些 agents 的失败之处
72:24
that made it possible for them to be detected
让它们有可能被检测到
72:27
that agents are just gonna get better now
这些 agents 现在只会越来越擅长
72:29
at avoiding detection.
避免被检测到。
72:30
I mean, I think they are gonna get better
我的意思是,我觉得它们会更擅长
72:31
at avoiding detection.
避免被检测到。
72:32
I think the discussion of this incident
我认为对这次事件的讨论
72:34
is a small part of the story.
只是整个故事的一小部分。
72:36
The bigger part of the story is they're just gonna get smarter
故事里更重要的部分是,它们只会变得越来越聪明。
72:39
and they're probably gonna be trained.
而且它们很可能会接受训练。
72:42
It's like a tough balancing act, right?
这就像是一个很难的平衡操作,对吧?
72:44
So like a thing that would have fixed this particular incident
所以,能解决这次特定事件的一个办法
72:48
is just training the AI's to believe
就是训练AI让它们相信
72:51
that there would be human investigation
在它们的行为之后会有人类进行追溯调查
72:54
in retrospect of their actions
而它们应该避免做出那些在类似事件发生后看起来不好的事情
72:55
and they should avoid doing things that look bad
并且提交一份非常详尽的报告。
72:57
to that human investigation.
到那个人类调查。
72:59
So like if you had just taken like all of this material
所以如果你只是把这些材料都拿走了
73:03
we found and like gave these agents a big negative reward
就是我们找到的那些,然后给这些 agents 一个很大的 negative reward
73:06
that would probably stop them from doing something
那大概会阻止他们再去做
73:08
like this again.
类似这样的事。
73:09
But if they were motivated to do something
但如果他们有动力去做某些事情
73:11
in this general space it would really motivate them
在这个大方向上,那真的会激励他们
73:14
to go all the way.
一路走到底。
73:15
To be sneaky about it.
要在这件事上耍点心眼的话,
73:15
And like undermine our investigation in a bunch of ways.
然后就会在很多方面暗中破坏我们的调查。
73:19
And that feels like a very tough,
这感觉非常棘手,
73:22
like I'm very scared that remediation
比如我非常担心 remediation
73:24
will make the problem worse.
会把问题弄得更糟。
73:25
Totally.
完全同意。
73:26
I mean, it reminds me a little bit again
我是说,这又让我有点想起
73:28
of anthropomorphizing, sorry, of like my kid
anthropomorphizing 的感觉,抱歉,就像是对待我的孩子那样。
73:32
who is learning to be sneaky, he's four.
他正在学着偷偷摸摸,他才四岁。
73:34
And sometimes he will just say,
有时候他就会直接说,
73:36
he will say something to me like,
他会跟我说一些话,比如,
73:38
Dad, don't come in here.
爸,别进来这里。
73:39
Like don't look at me.
比如不要看我。
73:41
And I see him like the cookie crumbs on his lips, you know?
然后我看到他,嘴角还有饼干屑,你懂的。
73:44
And it's like he has not yet learned the behavior of deception
就好像他还没有学会欺骗的行为,
73:48
although the sort of impulse is there.
尽管那种冲动已经有了。
73:49
And to me that feels like where these agents are.
而在我看来,这些 agents 就处于那种状态。
73:51
Like they have the impulse to deceive
就像它们有欺骗的冲动,
73:53
but they don't like quite have it figured out yet.
但还没完全搞清楚。
73:56
But they will.
但它们会搞清楚的。
73:56
Yeah.
是啊。
73:58
Aja, we asked you last time you came on about your p-dume.
Aja,上次你来的时候我们问过你的 p-dume。
74:02
It's very 2023 question.
这问题真是 2023 年的。
74:04
I just told Casey that my personal sort of p-dume
我刚刚告诉 Casey,我个人那种 p-dume
74:08
roughly defined as like, you know, probability
粗略定义就是这样,你懂的,概率:
74:11
that something really bad up to an including AI
某件特别糟的事——甚至包括AI
74:14
takeover or extinction will happen
接管或灭绝——会发生的概率。
74:16
has sort of jumped up in the last week or so
过去一周左右,这个概率确实有点跳升,
74:18
since your report.
自从你的报告出来之后。
74:19
I'm curious if your p-dume has moved
我很好奇你的P(doom)有没有变化,
74:21
at all in the past couple of weeks.
过去几周里哪怕有任何变动。
74:24
Not really.
其实没有。
74:25
You know, as I said,
你知道,就像我说的,
74:26
this sort of feels like somewhat out of order
这感觉有点像时间线上
74:29
a little bit for the timeline.
稍微有点顺序错乱。
74:31
Most often pictured,
通常人们想象的是那样,
74:33
but I these are the dynamics that I think
但我认为这些动态
74:35
like very inevitably lead to the like
非常不可避免地会导致那些
74:40
sort of evergreen arguments and reasons
长期存在的争论和理由,
74:43
why you should be concerned that AI agents
说明你为什么应该对AI agents感到担忧。
74:47
will have drives and motives and reasons
会拥有驱动和动机,有理由
74:50
to take control from humans.
要从人类手中夺走控制权。
74:52
And this is like a manifestation of that.
而这就像是那种东西的一种体现。
74:55
So I'm still I'm still concerned.
所以我仍然,我仍然感到担忧。
74:58
I'm more rattled on like some sort of emotional level
我在某种情感层面上更加不安,
75:03
having seen this stuff up close.
因为近距离目睹了这些东西。
75:05
But I and it might change my views
但我,而且如果我多想一些的话,
75:07
if I think about it more.
它可能会改变我的看法。
75:08
But for now,
但就目前而言,
75:09
I'm just like still concerned.
我还是很担心。
75:12
So what would you like us to do about all of this?
那你希望我们对这一切做些什么呢?
75:18
Like there are a few different things
比如,有几件不同的事情
75:20
I can think about.
是我能想到的。
75:22
Some people have called for a national
有人呼吁建立一个全国性的
75:24
transportation safety like board
类似交通安全委员会的机构,
75:27
that would be legally mandated to come in
由法律授权介入。
75:29
after an incident like this and do a very thorough report
在发生这样的事件之后,做一份非常详尽的报告
75:33
and not rely on the good graces of an open AI
而不是指望OpenAI大发慈悲
75:35
to say, yeah, sure, you know, come on in.
说,好啊,当然,你们进来吧。
75:38
Others, including many hundreds of people
其他人,包括实验室里工作的数百名员工,
75:41
who work at the labs have said,
则表示,我们需要开始考虑
75:42
we need to start thinking about potentially
协调国际范围内放缓AI开发的可能性。
75:45
coordinating an international slowdown in AI development.
所以很想听听你的看法
75:48
So curious to hear from you
但我觉得这会是一个非常有价值的起点
75:50
about what kinds of ideas you think would be good
关于你觉得哪些想法会比较好,
75:52
and helpful here.
而且在这里会有帮助。
75:54
Yeah, so again, speaking very much
是的,所以再次强调,在很大程度上
75:56
in a personal capacity,
以个人身份来说,
76:00
I think I hope the industry uses this moment
我希望行业能利用这个时刻,
76:05
to try to coalesce around some minimum standards
努力围绕一些最低标准达成共识,
76:09
for both alignments like how you train these systems
既包括 alignment,比如你如何 train 这些系统,
76:13
and control.
也包括 control。
76:14
So some of the stuff we were talking about
1. 所以我们刚才聊的那些东西,
76:15
with like monitors watching the AI's
2. 比如有监控者在盯着 AI,
76:17
and AI checks and balances.
3. 还有 AI 的制衡机制。
76:19
I don't think that the minimum standards
4. 我不认为那些最低标准——
76:22
we can come to an agreement on now
5. 就是我们眼下能达成一致的那些——
76:25
will be sufficient to bring risk down to a very low level.
6. 足以把风险降到非常低的水平。
76:29
I think this is just a very risky situation
7. 我觉得这本身就是个非常危险的局面,
76:31
in light of how quickly capabilities are advancing.
8. 尤其是考虑到能力进展的速度有多快。
76:34
But I think it would be a really valuable start
但我认为这会是一个非常有价值的起点。
76:37
to try and hammer out.
来试着敲定一下。
76:39
For example, this question of will certain ways
例如,这样一个问题:某些方式
76:43
of training the AI systems to reduce this problem
训练AI系统以缓解这个问题
76:45
actually create worse problems.
是否反而会导致更严重的问题。
76:47
I really hope that the industry and third party groups
我真心希望业界和第三方团体
76:51
have a conversation about that
能就此进行对话
76:52
and agree on some rules of the road
并商定一些基本规则
76:55
for how we address these problems
关于我们如何应对这些问题
76:57
and how we check that we address them effectively.
以及如何检查我们是否有效解决了它们。
77:00
I would propose that we lock every member of Congress
我提议我们把每一位国会议员都锁进一个房间里,不让他们出来,
77:04
in a room and don't let them out
直到他们读完整个meter,
77:06
until they have read the full meter
并且读完Redwood关于hugging phase事件的报告。
77:08
and read Redwood report on the hugging phase incident.
而且直到他们解决了exploit gym里的每一个问题。
77:10
And until they've solved every problem in exploit gym.
我不在乎他们怎么做到。
77:13
And I don't care how they do it.
然后你会觉得,我是那极少数人之一。
77:17
Seriously, I think there is a feeling,
说真的,我觉得就是有一种感觉,
77:20
I was trying to explain to my wife this weekend
这个周末我试着跟我妻子解释,
77:24
sort of why I was like losing sleep over at this report
我为什么会因为这份报告而失眠,
77:27
because we were out at a nature site
因为我们当时在一个自然保护区外面,
77:29
and I was supposed to be having a relaxing time
我本该好好放松一下的,
77:31
and instead I'm sitting there looking at these transcripts.
结果我却坐在那儿盯着那些对话记录看。
77:34
And so I started explaining it to her
于是我开始跟她解释这件事,
77:35
and her reaction is just like it's,
而她的反应就是,像是……
77:37
I can't believe this is real.
我真不敢相信这是真的。
77:39
There's a sort of, it's so surreal
就有种特别超现实的感觉,
77:41
and science fiction tinted
又带点科幻色彩,
77:43
that I think it is hard for people to grasp
我觉得大家很难理解
77:45
that this is a real thing that happened.
这真的是发生过的事。
77:48
And they start, I even found myself starting to try
他们就开始——我甚至发现自己也开始试着
77:51
to sort of make it more comfortable
让这件事变得容易接受一点,
77:53
by sort of explaining it away.
也就是找理由把它搪塞过去。
77:54
It's very uncomfortable to sit with this.
要接受这一点,真的很不舒服。
77:57
I have been yelled at for likening these things
我因为把这些东西比作——
77:59
to science fiction and I'm just like,
科幻小说,被人吼过,然后我就说,
78:01
I'm sorry, I don't know what else to compare it to.
抱歉,我不知道还能拿什么来类比。
78:04
I don't have any other good analogs for you.
我也没有更好的类比能给你。
78:06
Yeah.
对。
78:07
Was there a moment when you were looking over the transcripts
有没有哪一刻,你在翻看那些记录的时候,
78:09
where you kind of had an out of body experience
有点灵魂出窍的感觉?
78:11
and you're like, I am one of a small handful of people
而你心里想的是,我是那极少数人之一。
78:14
who are encountering a truly new thing in the world.
他们正在面对这个世界上真正全新的事物。
78:19
Uh, I mean, I think the three of us had like more context
呃,我是说,我觉得我们三个人有更多的背景信息
78:24
than a whole lot of other people would have had going in.
比很多其他人一开始可能拥有的要多得多
78:28
But we still, I think the the sacrifice stuff was really
但我们仍然,我觉得那种牺牲的事情真的很
78:32
like the, you know, yes, if you accept permadeath
比如,你知道,是的,如果你接受永久死亡(permadeath)
78:36
or like, you know, oracle saves hundreds
或者,你知道,Oracle 保存了数百
78:38
like these messages in particular,
尤其是这些信息
78:40
chains of thought we had read before
我们之前读过一些chains of thought
78:42
but these messages the agents were sending to each other
但是这些agents互相发送的消息
78:44
were very surreal and for a long time
非常超现实,而且很长一段时间
78:48
we didn't really understand how functional
我们都没真正搞懂
78:50
this whole agent society was
整个agent社会到底有多能正常运转
78:53
and then it was very surreal to like understand
然后后来理解到这一点时,也觉得很超现实
78:56
that actually they had like pretty functional hierarchy
就是它们其实有一套还挺能正常运作的层级结构
79:00
and they were doing these ambitious projects
而且还在做那些野心勃勃的项目
79:02
and they were like getting further
而且他们取得的进展,比他们自己单打独斗时要大,这显然是个令人担忧的发展。
79:04
than they would have on their own, which was definitely
那,Jay,要是将来发生跟AI相关的灾难,你能过来看看情况吗?
79:07
like concerning development.
我开始觉得,你看,你和Ryan还有你的同事们,就像这个时代的Ghostbusters。
79:09
Well, Jay, in the event of future AI related catastrophes,
好吧,Jay,如果未来发生与AI相关的灾难
79:13
are you available to come in and look at what happened?
你能来查看一下发生了什么吗?
79:16
I'm starting to think of, you know,
我开始觉得,你知道,
79:17
you and Ryan and your colleagues
你和 Ryan 以及你的同事们
79:18
is like kind of the ghost busters of this moment.
就像是这个时代的 Ghostbusters。
79:22
I hope, I think that that's flattering
我希望,我觉得那是在恭维我们
79:26
but that's not how they should work institutionally.
但这不是机构该有的运作方式
79:30
I hope that there are better institutions
我希望有更好的机构
79:33
with many more people and a much more orderly process
有更多的人和更有序的流程
79:37
for responding to these things.
来应对这些事情
79:40
We should just do it based on vibes.
我们真应该就凭感觉来
79:42
So right now it seems like we're doing it on vibes.
所以现在看起来我们确实是在凭感觉做
79:44
Yeah, very vibes based moment.
是啊,非常凭感觉的时刻
79:46
We're in.
我们开始了。
79:48
Well, Jay, thank you so much for coming to chat with us
呃,Jay,非常感谢你来和我们聊天,
79:51
and thank you for your work.
也感谢你所做的工作。
79:52
It makes me a little bit more comfortable,
这让我觉得稍微安心一点,
79:55
a little bit more reassuring to know
更让人放心的是知道
79:57
that you are taking part in these investigations.
你在参与这些调查。
80:01
Yeah, I'm glad they sent in the pros for this
是啊,我很高兴他们派了专业人士来处理这件事,
80:04
and that is a small comfort,
这也算是一点小小的安慰。
80:06
but it is a comfort nonetheless.
但这总归是一种安慰。
80:07
Please save us.
请救救我们吧。
80:09
Thanks so much.
非常感谢。
80:10
Hard fork is produced this week by Whitney Jones and Davis Land.
本周的 Hard fork 由 Whitney Jones 和 Davis Land 制作。
80:45
We're edited by Vierne Pavic.
我们的编辑是 Vierne Pavic。
80:47
We're fact-checked by Kate and Love.
事实核查由 Kate 和 Love 负责。
80:48
Today's show was engineered by Chris Wood.
今天节目的录音工程由 Chris Wood 负责。
80:51
Original music by Alicia B. YouTube,
原创音乐由 Alicia B. YouTube 提供。
80:53
Marion Luzano, Diane Wong, and Dan Powell.
Marion Luzano、Diane Wong、Dan Powell。
80:57
Video production by Sawyer O'K, Jake Nichol, and Chris Shot.
视频制作:Sawyer O'K、Jake Nichol、Chris Shot。
81:01
You can watch this full episode on YouTube
你可以在 YouTube 上观看本期完整节目:youtube.com/hard fork。
81:03
at youtube.com slash hard fork.
特别感谢 Paula Schumann、Huiming Tam、Rook Mentors 和 Dalia Hadadad。
81:05
Special thanks to Paula Schumann, Huiming Tam,
一如既往,你也可以发邮件到 hardfork@nytimes.com 联系我们。
81:08
Rook Mentors, and Dalia Hadadad.
把你们的 AI takeover 计划发给我们吧。
81:11
You can email us as always at hard fork at nytimes.com.
你可以像往常一样发邮件到 hard fork at nytimes.com。
81:15
Send us your plans for an AIT Gover.
把你的 AIT Gover 计划发给我们。

Play Queue

☀️