TWIML AI
Why the Next AI Breakthrough May Come from Physics with Max Welling - #774
2026-08-25 · 3337
In this episode of TWIML AI, Max Welling discusses why physics may drive the next AI breakthrough. He covers geometric and gauge-equivariant neural networks, AI-driven materials discovery at CuspAI, and applications such as metal-organic frameworks for carbon capture and improved solar panels. Welling also describes generative and agentic workflows, machine-learned force fields, self-driving labs, and deeper links between machine learning, information theory, entropy, and thermodynamics.
本期 TWIML AI 的嘉宾是 Max Welling,他讨论了物理学为何可能成为下一轮 AI 突破的重要来源。Welling 结合自己在 geometric neural networks、gauge equivariant neural networks 和 machine learning force fields 方面的研究,介绍了 CuspAI 如何用 AI 发现新材料,例如用于 carbon capture 的 metal-organic frameworks 和更高效的 solar panels。他认为,生成式模型与 agentic workflow 可以加速分子搜索、模拟和实验闭环,并推动 self-driving labs 的发展。他还提出,machine learning、information theory、entropy 与 thermodynamics 之间存在深层联系,例如 Landauer's principle 所体现的信息与能量关系。
00:00
I want to send a big thanks to Blitzie for supporting the podcast and sponsoring this episode.
我想对 Blitzie 表示大大的感谢,感谢他们支持播客并赞助本期节目。
00:06
Once we accelerate software development velocity by 5X, you need Blitzie, which brings
一旦我们把 software development velocity 提升 5X,你就需要 Blitzie,它能
00:10
autonomous software engineering to your enterprise.
将 autonomous software engineering 带到你的企业。
00:14
Blitzie starts by reverse engineering your codebase, building a dynamic understanding of your
Blitzie 首先通过 reverse engineering 你的 codebase,建立对你
00:18
entire application ecosystem.
整个 application ecosystem 的动态理解。
00:20
Your engineer simply declares intent, and once approved, Blitzie autonomously executes
你的工程师只需声明 intent,一旦获批,Blitzie 就会自主执行
00:25
entire software epics, delivering validated end-to-end tested code, more than 80% of
整个 software epics,交付经过 end-to-end 测试并验证的代码,超过 80% 的
00:31
the work completed in a single run.
工作在 single run 中完成。
00:34
Blitzie is not just generating code, it's developing software at the speed of compute.
Blitzie 不只是生成 code,而是以 compute 的速度在开发 software。
00:39
Experience Blitzie firsthand at Blitzie.com-slash2mall.
访问 Blitzie.com-slash-2mall,亲身体验 Blitzie。
00:43
That's B-L-I-T-Z-Y.com-slash-T-W-I-M-L.
就是 B-L-I-T-Z-Y.com-slash-T-W-I-M-L。
00:49
Max is so great to be on the line with you again.
Max,能再次和你连线真是太好了。
00:52
It's been a while.
有一阵子了。
00:53
It's great to be back, Sam.
Sam,能回来真是太好了。
00:54
I'm looking forward to our discussion.
我很期待我们的讨论。
00:59
Our audience can look up our conversations from, I think, 2019 and 2020, where we covered
我们的听众可以查看我们的对话,大概是2019年和2020年的,当时我们聊了
01:07
what you were working on at the time, and I think still echoes into your work today,
你那时候在做的工作,而且我觉得那些内容至今仍然影响着你的研究,
01:12
geometric neural networks, and gauge, aquaverience, neural networks, and the like.
几何神经网络,还有 gauge、aquaverience、神经网络之类的。
01:20
But I'd love to have you kind of catch us up on what you've been up to since, it's been
但我想请你跟我们更新一下你从那之后的进展,已经过了
01:26
Yeah, actually, the aquaverience theme has definitely continued, so, in fact, I found out
嗯,其实 aquaverience 这个主题确实一直在延续。事实上,我发现
01:33
that aquaverience was used very fruitfully in chemistry and material science.
aquaverience 在化学和材料科学中得到了非常成功的应用。
01:39
So in chemistry and material science, people train neural network models to predict
所以在化学和材料科学中,人们训练神经网络模型来预测
01:46
the forces on atoms, because people want to evolve atoms forward in time in order to
原子所受的力,因为人们想要让原子在时间上向前演化,以便
01:51
compute their properties, which is called molecular dynamics.
计算它们的性质,这被称为molecular dynamics。
01:56
And typically, you need to use quantum mechanics to compute these forces, because a large
而通常,你需要用quantum mechanics来计算这些力,因为很大一部分
02:03
contribution comes from the electrons, and electrons are very light.
贡献来自电子,而电子非常轻。
02:08
And so you need to treat them with quantum mechanics.
所以你需要用quantum mechanics来处理它们。
02:10
But, you know, if you get ten electrons or more, at big as it becomes completely unfeasible
但是,你知道,如果电子达到十个或更多,当体系变得足够大时,就完全不可行
02:15
to solve the so-called Schrodinger equation, so people have come up with approximations
去求解所谓的Schrodinger equation,所以人们提出了近似方法
02:20
like density functional theory, known as DFT, and people, you know, the inventors of that
比如density functional theory,也就是DFT,然后人们,你知道,那些发明者
02:27
got the Nobel Prize for that.
就因为这个拿了 Nobel Prize。
02:29
But what now people do is they train surrogate, so they provide data, computing this, using
但现在人们做的是训练 surrogate,提供数据,用这种昂贵的 quantum mechanics 近似来计算。
02:35
this expensive approximation to quantum mechanics.
然后他们用 neural networks 来走捷径,预测那个计算的结果,但速度要快得多。
02:39
And then they use neural networks to shortcut the computation, so to predict the outcome of
所以换句话说,相对于这些 quantum mechanics 近似,加速了三到四个数量级,效率更高。
02:45
that computation, but at a much more accelerated pace.
而在那些模型中,因为世界是三维对称的,所以如果我旋转一个分子,
02:48
So in other words, three orders or four orders of magnitude, acceleration, more efficiently
所以换句话说,三到四个数量级的加速,更高效,
02:54
relative to these quantum mechanical approximations.
是相对于这些 quantum mechanical approximations 而言的。
02:58
And in those models, because the world is three-dimensional symmetric, so if I rotate a molecule,
而在那些模型里,因为世界是 three-dimensional symmetric 的,所以如果我旋转一个分子,
03:05
all the forces will rotate with it, and so we could now put the same ideas that we use
所有的力都会随之旋转,所以我们就能把用在图像上的那套想法,放到这些分子、这些预测力的模型上,而且我们也可以用 equivariants。所以这也是为什么我——也因为我背景是科学,我博士读的是理论物理——我当时想,好吧,这简直是我旧热情和新热情的完美结合,我可以把它们放在一起,也就是从那时起我开始对 AI for science 感兴趣。所以这也就很直接地引向了 Cusp AI 的创立。
03:10
for images, we could put them in these molecules, these models that predict the forces, and
比如,我们可以把它们放进这些模型里,这些预测 forces 的模型里,
03:15
we could use agrivarians.
我们可以用 equivariants。
03:18
And so that's why I kind of, also because my background is in science, I did my PhD in
所以这就是为什么我有点——也因为我的科学背景,我读 PhD 时做的是
03:23
theoretical physics.
theoretical physics。
03:24
I thought, okay, this is a perfect unification of my old sort of passion, and my new passion,
我想,好吧,这完美统一了我以前的热情和现在的热情。
03:33
I can put it together, and that's when I started to be interested in AI for science.
我能把这些整合起来,也正是从那时起,我开始对AI用于科学产生了兴趣。
03:36
So that leads it pretty directly to the founding of Cusp AI.
所以这几乎直接促成了Cusp AI的创立。
03:41
Actually, first, I spent two years at Microsoft Research as a VP, because they were building
其实,首先,我在Microsoft Research做了两年VP,因为他们在组建AI for science实验室,我回应了他们,然后帮忙推进了那件事。
03:47
their AI for science lab, and I answered them, and so I helped that along.
但两年后,我想自己搞一个startup。
03:52
But after two years, I wanted to start a startup.
我之前也做过一个startup,但那家后来被Qualcomm收购了,之后我在Qualcomm待了一段时间。
03:56
I already did a startup a while ago, but this was actually the startup I got acquired
我真的很喜欢startup,那种充满活力的环境,以及你能产生的影响力。
04:01
by Qualcomm, and then I spent some time at Qualcomm.
而且我想用Silicon Valley的方式来做这件事,和我的联合创始人Chad Edwards一起,所以我们在2024年创立了Cusp AI。
04:05
And I really like startups, the dynamical environment, and the impact you can make.
我真的很喜欢startups,那种充满活力的环境,以及你能产生的影响力。
04:10
And I wanted to do it the Silicon Valleyway, together with my co-founder, Chad Edwards,
我想以Silicon Valley的方式来做这件事,和我的联合创始人Chad Edwards一起,
04:16
and so we started Cusp AI in 2024.
所以我们在2024年创立了Cusp AI。
04:19
Talk a little bit about the progress that you've made since then.
跟我们聊聊你从那以后的进展吧。
04:23
What is the kind of shape of the company today?
公司现在大概是什么形态?
04:25
This is being a huge ride, actually, it's a roller coaster.
这其实是一段很疯狂的旅程,就像坐过山车一样。
04:29
So we started, I think about two years ago, so may spring 24, yeah, we started with
所以我们大概两年前开始的,也就是2024年春天,对,我们一开始拿到了
04:39
a good initial investment of about 30 million from which we could hire an excellent team.
一笔大约3000万的初始投资,让我们能组建一支很棒的团队。
04:45
So the team has grown to about 50 people right now across different geographies, so
现在团队已经发展到大约50人,分布在不同地区,
04:53
there's a half-quarter, both in Amsterdam and in Cambridge.
在阿姆斯特丹和剑桥都有总部。
04:58
Actually, the half-quarter officially is in Cambridge, but the two initial labs were
实际上,总部正式设在剑桥,但最初的两个实验室是在
05:02
Amsterdam and Cambridge because Chad is from Cambridge, you know, from Amsterdam.
阿姆斯特丹和剑桥,因为Chad是剑桥人——你知道,阿姆斯特丹来的。
05:07
We now also have labs in London and Berlin, and we're also expanding into Asia and North
我们现在在伦敦和柏林也有实验室,同时也在扩展到亚洲和北美。
05:15
Got it.
毫不意外,你的顾问名单有点像名人录,Jeff Hinton和Jan LeCoon排在最前面?
05:16
And no surprise, your list of advisors is a bit of a who's who, with Jeff Hinton and
是的,这些顾问确实非常棒。
05:23
Jan LeCoon at the top of the list?
所以我们有Jeff Hinton和Jan LeCoon。
05:25
Yes, the advisors are actually fantastic.
是的,我们的顾问们其实都非常棒。
05:29
So we have Jeff Hinton and Jan LeCoon.
我们有Jeff Hinton和Jan LeCoon。
05:33
We added to that also Martin Vandembrink and Lord Brown.
我们还加上了Martin Vandembrink和Lord Brown。
05:39
So Lord Brown is the former CEO of BP, and Martin Vandembrink is the former president and
所以Lord Brown是BP的前CEO,而Martin Vandembrink是前总裁,也是
05:46
CTO of ASML, they're both retired, and they like to spend their time with new startups
ASML的CTO,他们俩都退休了,喜欢把时间花在新的startups上
05:53
and help them along.
并帮助他们发展。
05:56
And then there's a variety of harding, she's working for the UK government and also deep
然后还有一个叫Harding的人,她在英国政府工作,同时也在DeepMind
06:01
And then, or maybe formerly a debind.
然后,或者说,也许以前是在DeepMind。
06:04
And then Kristen Person who has sort of initiated the materials project, and she's also advising
然后是Kristen Person,她算是发起了Materials Project,而且她也在做顾问。
06:13
So let's dig into the technology and the approach that you're taking to kind of apply
那我们就来深入探讨一下你采用的技术和方法,如何应用
06:21
your, this original set of work that you developed to materials.
你开发的这套原始工作到材料上。
06:28
How did you get started with that effort?
你是怎么开始这项工作的?
06:31
So I think the opportunity is, to say to me, there's a deep fascination with the fact
所以我认为机会在于,对我来说,有一个事实让我深感着迷,
06:37
that there isn't sheer infinite amount of possibilities in which you can put together
那就是,并不是有无穷无尽的可能性,让你能够组合
06:42
And the universe has only figured out so many of them, because you know, they stay, they
而宇宙目前也就演化出了这么多种,因为你知道,它们会留在宇宙中,自然形成,我猜是这样。
06:47
form naturally, I guess, in the universe.
但还有更多你可以自己设计的材料,带有各种奇异的性质。
06:50
But there's many more that you can design yourself with all sorts of exotic properties.
我们当初创办这家公司的时候,我和Chad都挺担心气候问题的,现在也依然如此。
06:55
When we started this company, both Chad and I were kind of concerned about the climate,
所以我们觉得,很有必要加速能源转型,转向更可持续的能源来源,同时也要试着,你知道,把大气中的二氧化碳给弄出来。
07:04
And so we felt there's a strong need to accelerate the energy transition to more sustainable
所以我们觉得很有必要加速 energy transition,转向更可持续的能源来源,同时也要尝试,你懂的,把大气里的 carbon dioxide 弄出来。
07:09
energy sources, as well as trying to, you know, take the carbon dioxide that's in
那是极其庞大的量。
07:16
the atmosphere out.
那大概是,我觉得,有 Lake of Geneva 那么大,装满液态 carbon dioxide,一个巨大的量。
07:17
So not many people know that by the time it's, but 2015, of course, we really like to
其实没多少人知道,到2015年的时候,我们当然真的很想做到完全carbon neutral。
07:23
be completely carbon neutral.
但在这之后,还有50到100年的时间,我们每年都得把目前排放量的大约一半再移除掉。
07:26
But after that, there is still 50 to 100 years where we have to take out every year, about
所以那就是每年20 gigatons。
07:32
a half of what we currently put in.
这个量极其巨大。
07:33
So that's 20 gigatons a year.
我觉得那差不多就是整个Lake of Geneva装满liquid carbon dioxide的规模——一个庞大的量。
07:35
That's an enormous amount.
所以如果还没有现成的合适方案,它实际上会完全从头设计它们。
07:36
That's the kind of the, I think the size of the lake of Geneva filled with, sort of,
那大概就是,我想,有Geneva湖那么大,装满了,怎么说呢,
07:41
liquid carbon dioxide, a gigantic amount.
liquid carbon dioxide,巨量。
07:44
And we don't have the technology for that because it's actually very hard to take it out.
而且我们还没有那个技术,因为实际上把它提取出来非常难。
07:49
Because it's so dilute in the atmosphere, it's very expensive too.
因为它在大气中太稀薄了,所以成本也非常高。
07:53
And, you know, the two, the two factors which are more expensive are energy and the material,
而且,你知道,两个更贵的因素是能源和材料,
08:00
the sort of material that you use in order to take the carbon dioxide out of the atmosphere.
就是你用来把 carbon dioxide 从大气中提取出来的那种材料。
08:06
And so we've been working the first project we've been doing was improving the materials
所以我们一直在做的第一个项目,就是改进这些材料,
08:11
that take out this carbon dioxide from the atmosphere.
也就是把 carbon dioxide 从大气中提取出来的材料。
08:14
And these are called metal organic frameworks.
这些被称为 metal organic frameworks。
08:17
And this is the material that this year won the Nobel Prize in Chemistry.
而且这种材料今年获得了 Nobel Prize in Chemistry。
08:23
And yeah, we, so we build a platform and I can go into much more detail, but we build
嗯,对,我们建了一个 platform,我可以更详细地讲,就是我们用 AI 来设计这些材料。
08:26
a platform that designs these material with the use of AI.
所以你可以大致把它想成一个 search engine,但它不是搜索现有的文档。
08:30
So you can sort of think of it as a search engine, but it's not searching over existing
它搜索的是已知和未知的材料。
08:34
documents.
所以如果没有现成的合适材料,它实际上会从零开始完全设计它们。
08:35
It's existing, it searches over known and unknown materials.
而且它有很多组件。
08:39
So it actually completely designs them from scratch if, if there's no suitable ones
所以它实际上会从头开始完全设计它们,如果没有合适的
08:43
already available.
现成可用的话。
08:45
And there's many components.
而且有很多组成部分。
08:46
It's a genetics.
这是genetics。
08:47
There is an agent sitting at the core of it who orchestrates a long computation.
它核心有一个agent,在编排一个很长的computation。
08:53
And it searches through existing databases.
它会搜索现有的databases。
08:55
It generates entirely new molecules.
它生成全新的molecules。
08:59
And it also evaluates all these molecules with all sorts of tools.
它也会用各种tools来评估所有这些molecules。
09:02
And in this generation and evaluation, you know, ingredients and all sorts of methods
而且在这个generation和evaluation中,你知道,我们在我的academic lab里这些年开发出的ingredients和各种methods,你知道,扮演了非常重要的角色。
09:08
that, you know, we have developed over the years in my academic lab are play a very important
这些,你知道,是我们在我的学术实验室里多年来开发的,起着非常重要的
09:14
To be clear, you mentioned a molecule that won the Nobel Prize, were you involved in
说清楚一下,你提到一个得了 Nobel Prize 的 molecule,你参与发现那个 particular molecule 了吗?
09:22
the discovery of that particular molecule?
没有没有,我倒希望是。
09:24
No, no, I wish.
那是化学家做的,是他们发现的。
09:28
Those were chemists who did that, who discovered it.
我想是 Professor Gita Gawa、Professor Yagi 和 Professor Robson。
09:32
I think Professor Gita Gawa, Professor Yagi, Professor Robson.
我想就是这三位,最后一位我不太确定。
09:37
I think those are the three I'm not quite sure of last one.
不过这类 molecule 是,你知道,在 metal node 里有一个 metal complex,有各种 atoms 沿着它,还有在 graph 的 vertices 上。
09:41
But this is a class of molecules where there is a metal complex in a metal node with all
但这是一类分子,其中一个metal complex在一个metal node里,带有所有...
09:47
sorts of, you know, atoms along it and at the vertices of a graph.
就是,你懂的,原子沿着它以及 graph 的 vertices 排列。
09:53
And then there is so called linkers, which are also organic complexes, which are connecting
然后还有所谓的 linkers,它们也是 organic complexes,用来连接 graph 上的这些 vertices。
09:59
these vertices on the graph.
而且它们 extremely porous。
10:02
And they're extremely porous.
所以中间有非常大的 holes,surface area 也巨大。
10:04
So they have a very large holes in the middle with an enormous surface area.
所以如果你让空气 atmosphere 吹过去,空气中的 molecules,也就是 water、nitrogen 和 carbon dioxide——carbon dioxide 只占很小一部分。
10:11
And so if you blow sort of air atmosphere through it, the molecules in the air, which
它们 tend to stick to the sides。
10:17
is water and nitrogen and carbon dioxide, the carbon dioxide is only a small fraction
是水、氮气和二氧化碳,二氧化碳只是
10:23
They tend to stick to the sides.
它们往往会粘在边上。
10:26
And so then what you need to make is a molecule, you need to design a molecule where actually
所以接下来你需要制造的是一个 molecule,你需要 design 一个 molecule,也就是说,实际上
10:33
only the carbon dioxide sticks inside these pores and the rest, you know, goes through.
只有 carbon dioxide 会吸附在这些 pores 里面,其他的,你懂的,就直接通过。
10:37
So that's a design question.
所以那是一个 design 问题。
10:39
And also when it's full, you want to shake it or heat it to get it out so that you can
而且当它满了之后,你要摇晃它或者加热它,把它弄出来,这样你才能
10:46
actually reuse that particular material.
真正重复使用那个特定的材料。
10:49
What we added to this is a way to fine tune or design a molecule for a particular purpose.
我们在此基础上加入的是一种方法,可以针对特定目的去 fine-tune 或者 design 一个 molecule。
10:58
So people have, you know, actually made maybe around 100,000 of these molecules and labs
所以人们,你知道,实际上可能已经在 labs 里制造了大概 100,000 个这样的 molecules,
11:06
and and verify their structure.
然后,然后验证了它们的 structure。
11:08
And so what we can do is we can now basically come up with an entirely new molecule for a
所以现在我们可以做的,基本上就是针对一个非常具体的任务,设计出一种全新的 molecule,然后在实验室里把它做出来,再用于那个特定的任务。
11:14
very specific task and then make it in the lab and then use that for that, for that particular
我要说的是,这项工作不仅仅是在 molecule 上。事实上,这只是我们合作的第一批 molecule。我们后来已经扩展到了 semiconductor 领域。我们做了很多关于 perovskite 的工作,perovskite 是用于太阳能面板的材料,可以改进太阳能面板。
11:20
I should say cost was not only working on months.
我要说的是,我们不只是研究分子。
11:23
In fact, this was just the first set of molecules that we worked with.
事实上,这只是我们接触的第一批分子。
11:27
We have hence expanded to semiconductors.
因此,我们已经扩展到半导体领域。
11:31
We're doing a lot of work on and perovskides, which is materials for solar panels improved
我们在perovskites上做了很多工作,这是一种用于改进太阳能电池板的材料。
11:39
improved solar panels.
改进的太阳能电池板。
11:41
We look at semiconductors for new chip materials.
我们研究 semiconductors,寻找新的 chip materials。
11:47
We look at, you know, battery materials, fuel cells.
我们也会看,你知道,battery materials、fuel cells。
11:53
We also look at removing PFAS from water and we in the current set of molecules we use
我们也研究从水中去除PFAS,而我们目前用的那组 molecules
11:59
for that is again, metal organic frameworks.
用于这个的,还是 metal organic frameworks。
12:02
Do you partner with other companies that have an interest in these particular molecules
你们会和那些对这些特定 molecules 感兴趣的其他公司合作吗?
12:09
or are you out exploring and then if you find something, you will find partners and maybe
还是你们自己去探索,然后如果发现了什么,再找合作伙伴,也许
12:15
What's the thinking around the business model?
关于商业模式,你们是怎么考虑的?
12:18
So we like to work with partners because there's a very broad class of materials and
所以我们喜欢和合作伙伴一起工作,因为材料类别非常广泛,
12:23
every class has its own super experts that focus on those particular areas.
每一类都有自己的超级专家,专注于那些特定领域。
12:27
And they're either in academia or they're in companies.
他们要么在学术界,要么在公司里。
12:32
And of course, we also like to partner on the actual synthesis of these materials.
当然,我们也喜欢在这些材料的实际 synthesis 上进行合作。
12:35
So we partner with academic labs, but also with the labs inside of these companies.
所以我们既与学术实验室合作,也和公司内部的实验室合作。
12:40
And so we build an ecosystem or a network where our engine can actually help in all of these
这样我们就建立了一个 ecosystem 或 network,让我们的 engine 能够在所有这些
12:49
different material classes design the materials and then we work with those companies to actually
不同的材料类别中帮助设计材料,然后我们与那些公司合作,实际
12:55
That's a partnership model, but there's also internal projects that we run.
那是一个合作模式,但我们也有内部开展的项目。
13:03
So for instance, the project on metal organic frameworks for carbon capture, we ran self-funded
比如,那个关于metal organic frameworks用于carbon capture的项目,我们就是自筹资金在内部做的。
13:08
internally.
然后我们现在在semiconductor方面也有另一个项目,也是我们在做。
13:11
And then we also have another project now in the semiconductor side where we run it.
所以如果我们发现了很棒的东西,因为是我们自己出资的,我们就拥有IP,然后可以为这个IP找客户。
13:16
And so if we discover something fantastic, we self-funded, so we got the IP and then
最重要的是,我觉得最重要的是我们发现了某些东西,能让人们对我们有信心,相信我们能做这件事。
13:20
we can find customers for that IP.
我们可以为这些IP找到客户。
13:24
Most important, I think the most important thing is that we discover something that gives
最重要的是,我认为最重要的是,我们发现了一些东西,它能带来
13:28
people certain confidence that we can do this.
人们对这件事有一定的信心。
13:31
We actually really own this process and we know how to do it.
我们其实完全掌控这个过程,而且我们知道怎么做。
13:35
And so then the customers to design materials for their specific needs.
然后客户就可以针对他们的特定需求来设计 materials。
13:40
You talked a little bit about the generative or agentic nature of the scanning scientific
你之前提到了一点,关于扫描 scientific literature 的 generative 或 agentic 特性,以及如何用它来识别潜在的 molecules。
13:45
literature and using that to identify potential molecules.
但我们也聊过你工作里的一些 geometric 含义。
13:51
But we've also talked about some of the geometric implications of your work.
你暗示过可能会用到 simulation。
13:58
You alluded to potentially the use of simulation.
你能再多讲讲识别这些 molecules 的 end-to-end 过程吗?
14:03
Can you talk a little bit more about the end-to-end process of identifying these molecules?
好的,我很乐意。
14:10
So I guess there's a sequence of things that happens, right?
所以我猜事情是有一个先后顺序的,对吧?
14:14
The first thing is like in a search engine, you actually type a request.
第一件事就像在 search engine 里,你实际上输入一个请求。
14:19
So you basically say, I want a material and these are all the properties that it should
所以你基本上会说,我想要一种材料,这些是它应该
14:24
And these are the things it should not have.
还有这些是它不应该有的东西。
14:26
These are the properties it should not have.
这些是它不应该有的性质。
14:29
And so you give this as a query.
然后你就把这个作为 query 提交。
14:32
And then you could also tell if you have prior knowledge about how you want this particular
然后你还可以告诉它,你是否拥有 prior knowledge,关于你想要这个特定的
14:36
search to happen.
为了让搜索发生,
14:37
You could tell the system, maybe use these tools and maybe sequence it in this way.
你可以告诉系统,也许使用这些 tools,也许以这种方式来安排顺序。
14:43
So you can also give it some instructions.
所以你也可以给它一些 instructions。
14:46
And then it goes through a process of steps.
然后它会经历一个多步骤的 process。
14:48
So the first step is it will look through its database.
所以第一步是它会在它的 database 里查找。
14:51
So we have a very large database of materials, which we have all ingested into this database.
所以我们有一个非常大的 materials database,我们把这些材料都 ingest 进了这个 database。
14:58
Lots of it is sort of exclusive licenses from the big publishing houses.
其中很多是从大型出版社获得的 exclusive licenses。
15:04
And then it will start to look through all of the literature, whether something exists
然后它会开始浏览所有的 literature,看看是否存在相关内容。
15:08
out there that has these properties, or which is close to having these properties.
外面已经有的,能具备这些性质的,或者接近具备这些性质的。
15:14
And so if it doesn't, then it will have to go into a new phase.
如果没有的话,那就得进入一个新阶段。
15:18
So you can hold the conversation with this agent and talk about it.
所以你可以和这个 agent 对话,跟它讨论这件事。
15:21
So that's already quite useful.
那这已经挺有用了。
15:23
But typically, then the next step is that you go to a generative model.
但通常来说,下一步就是转向 generative model。
15:27
So in this case, it's the same generative model that generates images or video.
所以在这种情况下,它就是同一个能生成图像或视频的 generative model。
15:33
And so but in this case, it will generate molecules for you.
但在这里,它会为你生成分子。
15:38
And so you tell it, I want molecules with these very specific properties, the conditioning
然后你告诉它,我想要具有这些非常特定性质的分子,这就是 conditioning。
15:44
And then it will start to generate these molecules, often, you know, hundreds of thousands
然后它就会开始生成这些 molecules,通常,你知道,会有成千上万个。
15:49
Because it's quite cheap in the computer.
然后就到了下一个阶段,就是从这些生成的 molecules 中。
15:52
And then comes the next phase, which is out of these generated molecules.
我们现在必须筛出那些看起来,看起来非常有前景的。
15:57
We now have to sieve out the ones which are look, look very promising.
而这可能是一个非常昂贵的步骤。
16:01
And this can be a very expensive step.
而这可能是非常昂贵的一步。
16:05
In some sense, you're built a multi-scale digital twin of the process that you really want
从某种意义上说,你构建了一个多尺度的 digital twin,模拟你真正想让这些分子在其中运作的那个过程。所以第一步基本上就是把这个分子 relax 到它的 ground state,确保它处于最佳能量状态。然后你会做一堆检查,比如它带不带电荷?如果我晃它一下,它会不会散掉?如果孔洞大小重要的话,孔洞有多大,你知道的,各种容易快速算出来的东西,用来筛掉一大堆看起来没什么希望的候选。
16:12
to these molecules to operate in.
让这些 molecule 可以在其中运作。
16:14
And so the first step is basically you relax the molecule to its ground state to make
所以第一步基本上就是把这个 molecule relax 到它的 ground state,确保它处于最佳的 energy state。
16:20
sure that it's the best energy state.
然后你会做一堆检查,比如它是不是 charged?如果我摇一摇它,会不会散架?
16:24
Then you do a bunch of checks, like is it charged?
然后你做一堆检查,比如它有没有带电?
16:29
If I shake it, will it fall apart?
如果我摇一摇它,会不会散架?
16:31
How big are the pores inside if that's important, you know, all sorts of things that are easy
如果这很重要的话,里面的pores有多大,你知道,有各种各样容易快速compute然后扔掉的东西,你懂的,一大堆看起来不太有前景的。
16:37
to compute fast to throw away, you know, a whole bunch of things that do not look promising.
然后我们把这个force field提炼成一个对这个特定类别的问题非常高效的force field。
16:43
And so then filtering it down.
然后就把它过滤下来。
16:45
Yes, definitely a filtering step, yes.
对,肯定有一步是 filtering。
16:47
And then go to the next step.
然后再进入下一步。
16:48
So we have a pipeline that fine tunes or distills machine learning force fields for that
所以我们有一个 pipeline,专门针对那种特定材料去 fine-tune 或 distill machine learning force fields。
16:57
particular material.
这个过程就是,我们收集关于那种材料的所有数据。
16:58
So this is a process by which we take all the data there is about that material.
我们有一个 foundation model,它在更广泛的数据上训练过。
17:03
We have a foundation model that's trained on a much wider range of data.
然后我们把这个 force field distill 成一个——对这种材料来说非常高效的 force field。
17:08
And then we distill this force field into this is a very efficient force field for this
然后我们在MD loop,通常是molecular dynamics loop里用它来模拟分子,分子会扭来扭去、动来动去,从中你通常能计算出那个分子的非常关键的properties。
17:14
particular class of problem.
这些properties之后通常会进入更高尺度上的partial differential equation。
17:16
And we use that in an MD loop, typically the molecular dynamics loop, to simulate the molecule
然后我们在 MD loop 里用到这个,通常是 molecular dynamics loop,来模拟这个分子
17:22
as it wiggles around and moves around from which you can often compute very key properties
当它扭来扭去、动来动去的时候,你通常能从中算出非常关键的性质。
17:27
of that particular molecule.
那个特定分子的。
17:30
Those properties then often go into partial differential equation at a higher scale.
这些属性随后通常会进入更高尺度的partial differential equation。
17:35
And or in a process that actually models the device in which you want this up this material
或者说,在一个实际模拟你要让这种材料运行的设备的流程中。
17:43
And so that's again, more expensive.
然后你得到一个candidate,我们目前,我会说,这个领域正在发生的革命是self-driving labs,你能做的实验数量要快得多得多。
17:45
And so again, you want to do this with fewer and fewer candidate materials.
所以再说一遍,你希望用越来越少的候选材料来做这件事。
17:50
And then at the very end, you go to, you know, to an experiment, right, you go so now
然后到最后关头,你会去,你知道,做实验,对吧,就是说现在真的去做那个实验。
17:55
actually do the experiment.
那更贵,而且更慢。
17:56
Now that's more expensive and even slower.
而且你应该只用大概10种材料,用老方法来做。
17:59
And that's you should do this only with order 10 materials post in the old way.
所以然后你会得到一个候选物,我们目前——当前这个领域正在发生的革命,我会说,就是self-driving labs,在那里你做实验要快得多,快得多。
18:05
And so then, so and then you get a candidate, what we are currently, the current, I would
所以对于这些来说,experimental loop也要快得多。
18:09
say, revolution that's happening in this space is self-driving labs where the amount of
这是一个非常有趣的发展,我们现在正把我们的platform和它整合。
18:15
experiments you can do is much, much faster.
你能做的实验要快得多得多。
18:17
So you could do maybe a hundred experiments a day.
所以你可能一天大概能做一百个 experiments。
18:21
And then the game is more like the agent figures out what the settings of the experiment
然后这个游戏更像是 agent 去搞清楚 experiment 的 settings 应该是什么。
18:26
should be.
Experiments 做完了。
18:27
The experiments are done.
Data 会传回来。
18:28
The data comes back in.
然后你就有来自 experiment 的 data 和来自 computations 的 data。
18:30
And then you have the data from the experiment and the data from your computations.
你把这些结合起来,去设定 experiments 的下一阶段。
18:35
You combine them to set the next stage for the experiments.
所以 experimental loop 对这些来说会快非常多。
18:39
And so the experimental loop is much, much faster for those.
所以对于这些来说,experimental loop要快得多得多。
18:44
And that's a very interesting development that we are now integrating our platform with.
这是一个非常有趣的进展,我们现在正在将我们的平台与之整合。
18:50
And so how many molecule classes and individual molecules have you kind of gone, you know,
那你们大概接触过多少种 molecule classes 和 individual molecules,你知道,
18:59
all the way through this cycle with?
完整走完这个 cycle 的?
19:01
Yeah, we are engaged in in a few of those, but you know, the question is a little bit,
是啊,我们确实参与了其中一些,但你知道,这个问题有点,
19:07
what do you mean by all the way?
你说的 all the way 是什么意思?
19:09
The different materials are at different stages of maturity.
不同的 materials 处于不同的成熟阶段。
19:12
So one of them we went all the way to, you know, actually doing, you know, doing the lab experiments.
所以其中一个我们真的走到了,你知道,真正去做 lab experiments 那一步。
19:18
Another way on semiconductors is on its way.
另一种关于 semiconductors 的方法也快出来了。
19:22
And probably in a few months, we'll start doing the experiments.
大概几个月后,我们就会开始做实验。
19:26
Met all the way to lab experiments.
一路做到了 lab experiments。
19:28
And I was curious if you have enough data to say that you're hit rate with this process
而且我很好奇,你们是否有足够的数据可以说你们在这个 process 上的 hit rate
19:35
is higher, you know, versus lower or, you know, there's anything you can say qualitatively
是更高,你知道,还是更低,或者,你知道,有没有什么能定性地说的,
19:42
or quantitatively about the candidates that you produce relative to, you know, the traditional
或者定量地,关于你们产生的 candidates,相对于,你知道,传统的
19:49
So I definitely think that, you know, there's definitely evidence that these things are
所以我确实认为,你知道,确实有证据表明这些东西是
19:55
much more efficient.
效率高太多了。
19:56
So some of our scientists, they have said things like, we can do now in a few days, what took
所以,我们有些科学家说过类似的话:以前一个博士要花很久才能做完的事,我们现在几天就能搞定,但这更多是在digital领域。
20:02
a PhD before, but that's more in the digital domain.
这些科学家其实更多是做simulation方面工作的。
20:07
These are actually scientists that are more simulation-based work.
我们有两个正在进行中的项目,里面有实验性的部分。但我们有一个——你知道——我们和亚洲一个大型国家实验室签了合同,目前还不能说太多细节,我们在那里规划了更多类似的实验。
20:12
We have two projects going, which have experimental pieces to it, but we have a, you know, we have
所以基本上,安排就是我们带一些客户过来,然后我们可以在那里做实验。
20:21
a contract with a big national lab in Asia, which cannot quite say the details of yet, where
与亚洲一家大型国家实验室的合同,我们还不能透露太多细节,在那里
20:28
we have many more of these experiments planned out.
我们已经规划了更多这样的实验。
20:31
So basically, the setup is that we bring some customers and we can do experiments in that
所以基本上,我们的安排是,我们带来一些客户,然后我们可以在那里做实验。
20:37
lab, but they're lab scientists and they can also bring their partners and they can
实验室,但他们是实验室科学家,他们也可以带上他们的合作伙伴,他们可以
20:42
use our platform.
使用我们的platform。
20:44
And then we can collect data that way.
然后我们就可以通过那种方式收集data。
20:46
And I think, oh, and then there is one other one that we're currently doing in Amsterdam
而且我想,哦,还有一个我们目前在Amsterdam做的
20:53
on perovskites, which is running right now.
关于perovskites的,现在正在进行中。
20:56
So that's also a lab engagement.
所以那也是一个实验室合作项目。
20:58
There is one in...
还有一个在……
20:59
Perovskites is one.
Perovskites就是其中一个。
21:02
Perovskites is a material class that's a semiconductor crystal structure that you would put
Perovskites 是一种材料类别,它是一种 semiconductor 晶体结构,通常你会把它放在……放在 silicon 上,也就是普通 solar cells 的上面。
21:08
on a... on silicon, typically, on top of the normal solar cells.
这可以帮助你过滤掉,或者说基本上把可见光中的大量能量转化为能量。
21:14
And that can help you filter out or basically convert a large amount of the energy in visible
对吧?
21:21
light to energy.
所以我们和 DTU,也就是 Danish Technical University,也有一些关于 catalysis 的合作。
21:24
So we also have something with catalysis with DTU, which is the Danish Technical University.
所以我们和DTU也有关于catalysis的合作,也就是Danish Technical University。
21:29
So that's on currently running on catalysis, and we have a project that's about to start
所以那目前是在做催化(catalysis),我们有一个项目即将开始
21:37
And quite a few engagement with labs that are either running or are starting to run,
并且有不少实验室的合作,有的已经在运行,有的即将开始运行,
21:44
but they haven't finished completely.
但它们还没有完全结束。
21:46
And is your lab, or does CUSP publish?
那你的实验室,或者CUSP,会发论文吗?
21:54
Are you still active in academic publishing around materials science now?
你现在还活跃在材料科学(materials science)的学术发表领域吗?
21:59
So myself, I have one day at the university, and of course I still work in that day with
我自己嘛,在大学里有一天的时间,当然那天我还是会和
22:03
students and there we publish everything.
学生们,而且我们在那里发布所有内容。
22:06
I'm very interested in all sorts of things we have to do with AI for science.
我对我们为科学所做的各种AI相关的事情非常感兴趣。
22:10
So definitely yes.
所以当然是肯定的。
22:12
But also CUSP actually publishes.
但CUSP实际上也有发布。
22:15
So we have interns that work with our scientists, where we publish the results.
所以我们有实习生和我们的科学家一起工作,我们会发布结果。
22:22
So we've, Francis recently had results on machine learning force fields with uncertainty,
所以,我们——Francis最近在带uncertainty的machine learning force fields方面有了结果,
22:26
prediction, property predictors.
prediction、property predictors。
22:30
And the most important thing, I think, that we've recently released open source and published
而且最重要的是,我认为,我们最近发布了open source并发表了。
22:35
a blog post about is our new molecular dynamics framework called Cups, confusingly.
有篇博客文章讲的是我们的新 molecular dynamics framework,叫 Cups,挺让人困惑的。
22:43
So that's universal, it's universal particle simulation simulator.
所以它是 universal 的,是一个 universal 的 particle simulation simulator。
22:49
And so the reason why we did this is actually together with Nvidia is that, you know, now
而我们做这个的原因,实际上是和 Nvidia 一起的,就是,你知道,现在
22:56
with this new development of these machine learning force fields that I talked about
随着我之前提到的这些 machine learning force fields 的新发展
23:00
is neural networks that replace quantum mechanical calculations for the forces.
也就是用 neural networks 来替代用于计算 forces 的 quantum mechanical calculations。
23:05
You cannot run them very efficiently in the current sort of MD simulators, the simulators
你不能在现有的 MD simulators 中很高效地运行它们,也就是那些 simulators
23:12
that evolve a material or, you know, a molecule forward in time.
那些让材料或者,你知道,molecule 随时间向前演化的 simulators。
23:17
Because you need to, you need to run these neural networks on GPUs.
因为你需要,你需要把这些 neural networks 跑在 GPUs 上。
23:22
And typically, you know, these, these simulators, they don't run GPUs or on CPUs.
而且通常,你知道,这些 simulators 并不运行在 GPUs 或 CPUs 上。
23:29
And you want to run things in parallel.
而且你想要并行运行。
23:32
And so what was built by our team is a framework where you can compile these force fields into
所以我们团队构建的是一个 framework,你可以将这些 force fields 编译成
23:41
jacks, which is Python based.
jacks,它是基于 Python 的。
23:44
And then it will actually very efficiently run these MD simulators.
然后它实际上会非常高效地运行这些 MD simulators。
23:50
And you can run them in parallel as well so that you use your GPUs, the utilization of
而且你也可以并行运行它们,这样你使用你的 GPUs,
23:54
your GPUs is high.
你的 GPUs 的利用率会很高。
23:55
And so we released this open source a few days ago, actually, where we could go at
所以我们实际上几天前发布了这个 open source,然后我们可以去
24:01
And yeah, we got a lot of excited responses to that are the representations of the molecules,
而且,关于分子的representations,我们收到了很多热烈回应。
24:09
like are there standard formats for representing these things that someone working in the
比如,在这个领域工作的人是不是已经有了一些standard formats来表示这些东西?
24:14
space would would already have.
然后你的machinery就直接在这些已有的格式上运行。
24:15
And then your machinery just works on those existing ones.
你可以把representations看作是一种化学领域的foundation model。
24:20
The representations, you can think of that as maybe a foundation model for chemistry.
所以你要做的是——这也是我们和Meta一起做的工作。
24:25
So what you do is, and this is also work that we did together with meta.
所以在那种情况下,你就自己训练一个这样的force field。
24:31
So what you do in that case is you train yourself one of these force fields.
所以在这种情况下,你要自己训练一个这样的force field。
24:37
So you take a molecular structure, so that is basically the position of the atoms and
所以你取一个 molecular structure,那基本上就是 atoms 的位置,
24:43
their and their and their class, like if it's oxygen or hydrogen.
以及它们的,它们的,它们的 class,比如是 oxygen 还是 hydrogen。
24:49
And of course, you want this to be aquivarian because rotations and translations don't
当然,你会希望这个是 equivariant 的,因为 rotations 和 translations 并不
24:54
And then you map this into a latent space where they get represented by some code.
然后你把它映射到一个 latent space,在那里它们用某种 code 来表示。
24:59
That is, that is meaningful.
也就是说,这是有意义的。
25:01
And if you do a machine learning force field, you will then actually predict the forces
如果你用 machine learning force field,你就能真正预测 forces
25:04
and the energies from that.
以及从那里得到的 energies。
25:06
But you can stop at that intermediate level.
但你可以就停在那个中间层级。
25:08
And then you have a, you know, a representation for that particular material.
然后你就有了,你懂的,针对那种特定材料的 representation。
25:14
And you can train this on a very wide range of materials and chemistry.
而且你可以在非常广泛的材料和化学上训练这个。
25:19
And from that representation, you can build property predictors.
然后从那个 representation,你就可以构建 property predictors。
25:23
You can use it to condition your generative models.
你可以用它来 condition 你的 generative models。
25:25
There's also some uses for that, you know, foundation model.
你懂的,那个 foundation model 也还有一些用途。
25:30
But it's also the starting point for distilling, let's say models for, for specific material
但它也是,比如说,针对特定材料类别来 distilling 模型的起点。
25:38
And to be clear, are these foundation models trained, you know, per molecule or per class
而且说清楚一点,这些foundation model是每个分子训练,还是按类别训练?
25:43
or are they, you know, very broad?
还是说它们非常宽泛?
25:48
So that's a great question.
这是个好问题。
25:49
A foundation model, almost by definition, is trained on a very broad class of materials.
一个foundation model,几乎从定义上讲,就是在很宽泛的材料类别上训练的。
25:57
The best data set for that is the materials project data set and, and omol from, from
最好的数据集是materials project数据集,还有Meta的omol。
26:03
meta.
所以他们生成了非常大量的DFT计算,这些量子...
26:04
So they have, they have generated very large number of dft calculations, these quantum
所以他们生成了非常大量的DFT计算,这些quantum……
26:08
mechanical calculations to create that data set.
用力学计算来生成那个数据集。
26:12
And those calculations are being used to train these force fields that I talked about
而这些计算正被用来训练我提到的那些 force fields。
26:16
or those representations, those foundation models, right?
或者那些 representations,那些 foundation models,对吧?
26:19
And then if, if you say, but I'm actually interested in this particular material, right?
然后如果你说,但我实际上对这个特定材料感兴趣,对吧?
26:24
What you can then do is start from that very broad representation, this foundation model
那你要做的就是从那个非常宽泛的 representation,也就是这个 foundation model 出发。
26:30
and then fine tuned model for that particular material class.
然后再对那个特定材料类别做 fine-tuned 的模型。
26:34
That way it is specialized for the material class, but that way it's also very fast because
这样一来它对这类材料是专精的,但同时也非常快,因为
26:38
you need these things to do very fast because in a molecular dynamics simulation, you have
你需要这些东西运行得非常快,因为在 molecular dynamics simulation 里面,你要面对的是……
26:43
to call them many, many times, right?
要 call 它们很多很多次,对吧?
26:45
Which is one, one little step in an MD simulator is a femtosecond.
在 MD simulator 中,一小步就是一 femtosecond。
26:49
It's a tiny step.
这是非常小的一步。
26:51
And so you have to call them many times to make any progress.
所以你必须 call 它们很多次才能有所进展。
26:54
And you mentioned distillation earlier, is that where distillation comes in?
你之前提到了 distillation,distillation 就是在这里派上用场的吗?
26:57
You're trying to get a smaller model that's more focused on the molecules that you care
你想得到一个更小的 model,让它更专注于你关心的那些分子。
27:05
That's, that's, that's what distillation is called.
这,这,这就是所谓的distillation。
27:06
You distill the big model into a small model.
你把大模型distill成一个小模型。
27:09
In this case, the use of fine tuning, is it, how analogous is it to the process of fine
在这种情况下,使用fine tuning,它跟LLM里的fine tuning过程有多相似?你知道,reinforcement fine tuning,用traces,就是language-based traces?
27:17
tuning in LLM, you know, reinforcement fine tuning with, you know, traces, you know,
比如,你的dataset是什么?你在这个过程里用来做fine tuning的dataset是什么?
27:23
language-based traces?
所以实际上它跟LLM很相关。
27:24
Like what is the, your dataset that you're fine tuning in your process that you're fine
所以有,有两种做法,通常你可以用graph neural network,它实际上看三维结构,或者你可以把它当作LLM,比如……
27:29
tuning with here?
用这里来 tuning 吗?
27:30
So actually it's quite related to an LLM.
所以实际上这和 LLM 关系很大。
27:34
In fact, the models we use are very related to an LLM, so you can basically think of the
事实上,我们用的模型和 LLM 非常相关,所以基本上你可以把现有的信息看作一个 sequence,然后把它 sequentialize。
27:39
sequence, the information that you have, you can sequentialize it.
然后从那里,你就可以,呃,把它映射到这个 latent space 里。
27:45
And then from there, you can actually, you know, map it into this latent space.
所以有,有两种做法,通常你可以用 graph neural network,它实际上看三维结构;或者你把它当 LLM 用,像 sequence 版本,比如分子经常被表示成一种叫 smile string 的字符串,这样它就变成了一维 string,你也可以用那个来表示分子。
27:49
So there's, there's either, and typically you can do like a graph neural network, which
所以有,有一种,通常你可以用 graph neural network,它
27:58
actually looks at the three-dimensional structure, or you can use it as an LLM, like
实际上查看三维结构,或者你可以把它用作 LLM,比如
28:01
a sequential version, like often molecules are represented as, you know, as a, as a
一种 sequential 的版本,就像分子经常被表示成,你知道,作为一个叫 SMILES string 的 string,然后就变成了一个 one-dimensional string,你也可以用那个来表示分子。
28:08
string called a smile string, so then it becomes a one-dimensional string, and you can
它甚至更进一步,因为你基本上还可以看 language、graphs 和 molecules 的组合。
28:12
use that to represent the molecule as well.
所以你可以某种程度上,从你知道的,比如说 literature 里学习,那里有很多 text,然后每当提到一个 molecule 的时候,你可以表示,你
28:15
So you can use both LLM style models as well as a graph neural network style models
所以你可以同时使用LLM风格的模型以及graph neural network风格的模型来做这个。
28:19
for this.
所以你的基础模型可能和LLM类似,但在表示molecules的时候,你有一个非常独特的tokenizer?
28:20
So your base model might be similar to an LLM, but you've got a very unique tokenizer
这更进一步,因为你基本上还可以同时考虑语言、graphs和molecules的组合。
28:25
in the case of representing molecules?
所以你可以有点像是,你可以从,你知道的,比如说文献中学习,那里有大量文本,然后你可以表示,每当提到一个molecule的时候,你...
28:29
It goes even further because you can basically also look at combinations of language and
我认为那是一个非常令人兴奋的 proof point,会帮助整个 industry 向前推进。
28:35
graphs and molecules.
图和分子。
28:36
So you can sort of have, you can learn from, you know, let's say the literature, where
所以你可以有点,你可以从,你懂的,比如说,文献中学习,那里
28:42
there's a lot of text, and then you can repress, every time a molecule is mentioned, you
有大量的文本,然后你可以表示,每次提到一个分子时,你
28:47
can actually use the molecular representation, which is then the little graph neural network
实际上可以用 molecular representation,然后它就是那个小型的 graph neural network
28:52
that represents, so that becomes the token.
它表示出来的东西就成了 token。
28:54
You turn this into a bunch of tokens using your graph neural network.
你用 graph neural network 把它转换成一大堆 token。
28:58
And so you can, you can really start to combine these things actually.
所以实际上你真的可以开始把这些东西组合起来。
29:01
What do you see the work that you're doing yet, cusp going in the future, like what are
你觉得你现在做的这项工作,未来会怎么发展?比如
29:07
the kind of near and midterm and longer term things that you're most excited about with
你对近期、中期和长期的哪些事情最感兴趣,尤其是
29:14
Okay, so you're constantly expanding the tools that we need for the different class, material
好的,所以你是在不断扩展我们为不同类别材料所需的工具。
29:22
So we know how to train these tools, but every new material class has to be retrained.
所以我们知道如何训练这些工具,但每一种新的材料类别都需要重新训练。
29:27
We also going up the stack, so we're going into more coarse, grained representations,
我们也在顺着stack往上走,进入更粗粒度的表示,
29:34
like larger scale, digital twins, all the way up to modeling the actual reactor or device,
比如更大规模的digital twins,一直到建模实际的反应堆或设备,
29:41
which is material sitting.
也就是材料系统。
29:45
The most, you know, the effort that's most important for us right now is connecting
目前对我们来说最重要的,你懂的,是连接
29:50
the platform to self-driving labs, so to really be very, you know, generate large amounts
平台和self-driving labs,从而真正地,你知道,生成大量的
29:56
of data from the self-driving lab.
来自self-driving lab的数据。
30:02
And I think ultimately it's very important to go through the entire process of predicting
而且我认为,最终最重要的是走完整个流程,从预测
30:08
the molecule, making it in the lab, scaling the material and then actually putting it in
molecule,在 lab 里把它做出来,对这个 material 做 scaling,然后真正把它放进
30:12
an actual device and getting a customer excited about that particular material and willing
一个实际的 device,让客户对那种特定的 material 感到兴奋,并且愿意
30:19
to pay money for it, right?
为它付钱,对吧?
30:21
But this basically means that, you know, this last part of scaling is something that
但这基本上意味着,你知道,最后的 scaling 这部分是
30:27
is still needs to be done, but it's, yeah, you know, getting there as quickly as possible
还需要去做的,但,是的,你知道,尽快做到那一步
30:32
as I think absolutely key.
我认为绝对是关键。
30:36
And so I'm really looking forward to discovering a material that is unique, it could not
所以我真的很期待能发现一种独特的材料——它不可能由AI单独完成——而且能真正做成semiconductor device或者类似的东西,或者是一种新的solar cell。
30:43
have been done by AI, and that makes it into an actual semiconductor device or something
然后我们可以说,你知道,AI确实帮了忙,确实在这种材料的设计中是非常关键的一环。
30:49
like this or a new solar cell.
我觉得那会是一个非常令人兴奋的证明点,能推动整个行业向前发展。
30:51
And we can say, you know, AI actually helped, you know, was a very important piece of the
你还在写一本书,书名特别有意思。
30:57
design of this particular material.
这本书讲的是generative AI和stochastic thermodynamics。
30:59
I think that's a very exciting proof point that will help the entire industry forward.
我觉得那是一个非常令人兴奋的 proof point,能推动整个行业向前发展。
31:04
You also have a book that you're working on that has a really interesting title.
你还有一本正在写的书,书名特别有意思。
31:10
It's talking about generative AI and stochastic thermodynamics.
它讲的是generative AI和stochastic thermodynamics。
31:18
What's the connection between generative AI and thermodynamics and tell us a little bit
generative AI和thermodynamics之间有什么联系,给我们稍微讲讲
31:23
about the direction you're taking with the book?
你写这本书的方向又是什么?
31:26
So maybe first a bit of a history on this.
那也许先简单讲讲这件事的背景。
31:28
So I started this almost two years ago, not even before CAUSEP.
我差不多两年前开始做这件事,甚至都没早于CAUSEP。
31:34
So I had a bit of a, sort of a law between, you know, when I stopped working for Microsoft
所以我有那么一点,算是,一个间歇期,你知道,就是在我从微软离职之后
31:39
and the start of CAUSEP, I started, so there was about six months between that.
到CAUSEP开始,中间大概有六个月吧。
31:45
And I started to, we were on, me and my wife were on vacation in Italy and see the, you
然后我开始,我们当时在,我和我妻子在意大利度假,然后你看,你
31:50
know, I cannot do absolutely nothing.
知道,
31:53
And so in the mornings, I mean, I'm just going to write a book.
我完全没法什么都不做。
31:56
Yeah.
所以到了早上,我是说,我就打算写一本书。
31:57
I had my cappuccino, you know, in the sun, and I would start a book.
对。
32:01
So wonderful, the best vacation you can have.
我端着我的卡布奇诺,你知道,在太阳底下,然后就开始写书。
32:04
And then the afternoon, we would hike through the mountains and enjoy life.
太棒了,你能拥有的最好的假期。
32:08
And so, and then I was teaching a course at, in a town called Mausenberg, which is
然后到了下午,我们就在山里徒步,享受生活。
32:17
a Dutch word, but it probably people said Musenburg, but it's close to Cape Town in South
一个荷兰词,不过人们可能说的是Musenburg,它在南非开普敦附近。
32:22
Africa.
那里有African Institute for Mathematical Sciences。
32:23
There is the African Institute for Mathematical Sciences.
所以我非常荣幸在那里待了几周,围绕我书里写的那些想法去教学。然后,然后我找了两个学生来帮我把它写完,因为你知道,开始容易,完成难,确实需要帮手。
32:27
And so I had the pleasure of spending a couple of weeks there teaching from these kind
所以我很高兴能在那里待上几周,教这些
32:32
of ideas that I was writing a book about.
我当时写书时涉及的那些想法。
32:36
And then, and then I recruited two students to help me out, actually, you know, finish it
然后,然后我找了两个学生来帮我,其实,你知道,就是完成它
32:40
because it's, you know, it's, you know, starting as easy finishing as hard, and you need some
因为,你知道,开始容易,完成难,所以你需要一些
32:45
help.
帮助。最后说到information theory,对吧,物理的核心是information。
32:46
And so it took two more years to actually finish it as a huge amount of work.
所以又花了两年时间才真正完成,工作量非常大。
32:50
But what is the book about the book is about actually the mathematics that describes modern
但这本书讲的是什么呢?这本书实际上讲的是描述现代
32:59
generative AI, including probabilistic models, including diffusion models and many other
generative AI,包括probabilistic models,包括diffusion models以及许多其他
33:06
Their mathematics turns out to be equivalent to the mathematics that describes modern
它们的数学实际上等同于描述现代
33:12
non-equilibrium, statistical mechanics or thermodynamics.
non-equilibrium statistical mechanics或thermodynamics。
33:18
And how specific is that statement, is that statement, you know, everything is a PDE
这个说法具体到什么程度呢?这个说法,你知道,是不是说一切归根结底都是PDE
33:26
at the end of the day, or is that statement, you know, more kind of specific and concrete?
还是说这个说法,你知道,更偏向具体和明确?
33:33
It's surprisingly specific, and I say surprisingly because I don't think, maybe many people think
它具体得惊人,我之所以说“惊人”,是因为我不觉得——也许很多人会想它是否有类比。
33:40
if it has an analogy.
所以我认为这是一个非常好的问题,因为我觉得这要深入得多。
33:41
So that's what I think is a really good question because I think this runs much deeper.
所以我觉得,在我看来,这不仅仅是一个类比。
33:45
So I think the, it's more than an analogy in my mind.
我认为它们之间有着非常非常深的联系,而这归根结底与 information theory 有关,对吧?物理学的核心就是 information theory。
33:49
I think there is a very, very deep connection and this has to do with the fact that we are
而 machine learning 的核心也是 information theory。
33:54
talking about information theory in the end, right, that the core of physics is information
最后在说 information theory,对吧,物理学的核心就是信息。
34:02
And in the core of machine learning is information theory.
而 machine learning 的核心是 information theory。
34:06
And both of these are described.
而这两者都是这样描述的。
34:09
So if you talk about information theory with loss of information, in other words, processes
所以,如果你谈论的是带 information loss 的 information theory,换句话说,就是那些随着时间演化而丢失信息的过程。
34:15
where you lose information as you are evolving over time.
所以有一个 observer,试图去描述一个过程,但有太多 degrees of freedom 需要追踪,你根本做不到。
34:19
So there's an observer that tries to, you know, describe a process, but there's so many
所以发生的情况就是,你在丢失 information,而你需要用 probabilities 来捕捉这一点。
34:25
degrees of freedom to keep track of, you can't.
而这门数学正是 thermodynamics 和 machine learning 两者的核心。
34:28
And so what happens is that you're losing information and you need to capture that by probabilities.
而我认为这就是核心观点。
34:36
And that mathematics is the core of both thermodynamics as well as machine learning.
而这一数学是 thermodynamics 和 machine learning 两者的核心。
34:43
And I think that's the core statement.
而我认为这就是核心陈述。
34:47
And this thing, I think, you know, there's this concept of entropy, which is maybe interesting
而且这个东西,我觉得,你知道,有一个叫 entropy 的概念,可能在两个理论里都挺有意思的。
34:52
in both of these theories.
所以也许我们值得稍微聊一下这个。
34:54
And so maybe it's good to talk a little bit about that.
所以 entropy 基本上就是你对某个 particular state 所感受到的 surprise。
34:56
So entropy is basically the surprise that you find for a particular state.
你知道,你清楚那个 precise state 吗?还是你只有一个非常模糊的 probability distribution 来描述那个 particular state?
35:03
You know, do you know the precise state or is you only have a very vague probability distribution
确实如此。
35:08
for that particular state?
但在物理学里,你会把 entropy 看作一个可以对 system 进行测量的东西。
35:14
But in physics, you think of entropy as maybe a thing you can measure of the system.
但在物理学中,你会把 entropy 视为系统可测量的某个东西。
35:23
Because you have to subtract entropy from energy actually to get free energy, which is the
因为实际上你得从能量中减去熵才能得到自由能,也就是你可以自由用来在这个世界上做有用功的那部分能量。
35:27
energy you're free to use to do useful work in the world.
所以它感觉是一个真实的东西,而不是某种我对世界一无所知的东西。
35:32
And so it feels like a real thing, not like something that I don't know about the world.
而在贝叶斯统计学里,情况正好就是这样,你谈论的是主观统计,因为它描述的是我对这个世界不了解的部分。
35:39
And in a Bayesian statistics, this is exactly what it's like, you talk about a subjective
我需要,你知道,对那些我不了解的事情去描述概率。
35:47
statistics because it describes what I do not know about the world.
但在我看来,很多物理学家也这么认为,尤其是E.T. James很出名,熵在物理学中也精确地描述了你所缺失的所有信息。
35:50
And I need to, you know, describe probabilities to the things that I do not know about the world.
而我需要,你知道,对我不知道的世界上的事物赋予 probabilities。
35:56
But in my view, and the number of physicists as well, in particular, E.T. James is a famous
但在我看来,也有不少物理学家,特别是 E.T. James,是一位著名的...
36:01
one, entropy and physics also precisely describes all the information you're missing about
第一,entropy 和物理学也精确地描述了所有你缺失的信息——关于
36:09
And so there's this very deep connection between these two fields.
所以這兩個領域之間有非常深的連結。
36:13
Maybe the last thing I want to say about this, there is this famous theorem by Landauer who
關於這個,我想說的最後一件事是,Landauer 有一個著名的定理,他
36:18
basically said that if you erase a bit from a device, you must radiate K L and T, which
基本上說,如果你從一個裝置抹除一個 bit,你必須輻射出 K L 和 T,而這
36:30
is just a number, amount of energy as heat to the environment.
只是一個數字,是以熱的形式釋放到環境中的能量。
36:36
So here you can see the direct relationship.
所以你在這裡可以看到直接關係。
36:38
I'm deleting a bit, which is a piece of Shannon information from a chip.
我正在刪除一個 bit,也就是從晶片刪掉一塊 Shannon 資訊。
36:45
And heat determined at that, when you do that, you must radiate heat to the environment.
而熱量就取決於此——當你這麼做時,你必須向環境輻射熱量。
36:50
And so you can see the directness.
所以你可以看到这种直接性。
36:51
That's just kind of a conservation of information along the lines of conservation of mass, which,
这就有点像 conservation of information,类似于 conservation of mass,而,
36:56
you know, is very foundational in physics.
你知道,这在物理学中是非常基础的东西。
37:00
In fact, the statement is the second law of thermodynamics says that the entropy must
事实上,这个说法就是 second law of thermodynamics,说的是 entropy 必须
37:06
always increase, which means, you know, you either keep the level of information the same,
总是增加,也就是说,你知道,你要么保持 information 的水平不变,
37:12
which is when the entropy stays the same, or you lose information in the process.
也就是 entropy 保持不变的时候,要么在这个过程中丢失 information。
37:17
And that's when the entropy goes up.
而那就是 entropy 增加的时候。
37:19
So of course, what happens is the information in the universe doesn't go away, but it's transferred
所以当然啦,实际情况就是,宇宙中的信息并没有消失,而是从一个你熟悉的系统转移到了heat bath里面,在那里你已经彻底失去了这些信息,完全不可恢复。
37:24
from something you know, the system to the heat bath, where you've completely lost this
所以,嗯,总之,这有点技术性,但核心思想是,information theory 是这两个理论共同的基础,基本上让它们的数学形式变得非常非常相似。
37:29
information, it unrecoverable.
而且,很多在一个领域里发展出来的工具,在另一个领域里都有完全对应的东西。
37:32
And so yeah, anyway, it's a bit technical, but the idea is that information theory is
因为很多时候,找到这些不同的——你知道——就像是两者之间的一本词典。
37:36
behind both of these theories, which basically makes the math very, very similar.
这两个理论背后的东西,这基本上让数学变得非常非常相似。
37:42
And many of the tools which have been developed in one field have their exact analog in the
而且,很多在一个领域里开发出来的工具,在另一个
37:48
Because a lot about finding these different, you know, it is kind of a dictionary between
因为很多关于寻找这些不同对应关系的事情,你知道,它有点像一本词典,在两者之间。
37:55
the two fields, right?
这两个领域,对吧?
37:57
There's something called stochastic normalizing flows in one in the machine learning and
在 machine learning 那边有一种叫 stochastic normalizing flows 的东西,
38:03
then there is escorted free energy estimation in the other field and there turn out to be
然后另一个领域里则有 escorted free energy estimation,结果它们
38:08
exactly the same methods.
其实是完全一样的方法。
38:11
And so ultimately, who is the book for and how do you see the book changing the way they
那么说到底,这本书是写给谁的?你觉得它会如何改变他们
38:17
view the world or the way they're able to do the things they do?
看待世界的方式,或者他们做事的方式?
38:22
For actually, you know, for both sides, it's written for the machine learner who is interested
实际上,你知道,对两边来说,这本书是写给那些有兴趣的 machine learner——
38:28
to learn a little bit about, you know, thermodynamics and non-equilibrium thermodynamics.
想稍微了解一下,嗯,thermodynamics 和 non-equilibrium thermodynamics 的人。
38:34
And I think, and in reverse, right, it's for also for the physicists who wants to get into
而且我觉得,反过来也一样,对吧,这也是对那些想进入 machine learning 的物理学家说的,他们可以基于自己已经知道的东西继续构建。
38:39
machine learning and build on the things they already know.
我只想指出一点:关于 fusion models 的第一篇论文,标题里其实就有 non-equilibrium thermodynamics 这个词。
38:42
And I just want to point out that the first paper that was written about the fusion models
所以那篇论文的作者其实早就知道这层联系。
38:46
actually had the word non-equilibrium thermodynamics in the title.
所以 diffusion model 本质上就是一个过程:你先有结构,然后摧毁它——这通常就是世界上发生的事情。
38:49
So the authors of that paper actually already knew about this connection.
熵不断增加,然后我们试图在时间上逆向反转这个过程,也就是从...
38:54
So it's a diffusion model really are a process by which you take structure and you destroy
所以 diffusion model 其实就是一个过程,你取出结构然后摧毁它,
39:00
it, which is typically what happens in the world.
这通常是世界上发生的事。
39:03
The entropy goes up and then we try to reverse that backward in time, which is to start with
entropy 会上升,然后我们试图在时间上把这个过程倒回去,也就是从
39:09
noise and create structure, which is our generative models.
从噪声中创造结构,这就是我们的 generative models。
39:14
And the way I think it will evolve into the future is, of course, there's, you know,
而我认为它未来会如何演变,当然,你知道的,
39:18
or the way it can be used fruitfully is, first of all, it can help people in physics to
或者说它能被有效利用的方式,首先是它可以帮助物理学领域的人
39:27
use these tools for machine learning.
使用这些工具去做 machine learning。
39:29
And let me give you one example.
我给你们举个例子。
39:31
So it's a very important problem to compute the free energy difference between, let's
所以,计算 free energy 之间的差异是个非常重要的问题,比如
39:37
say a unbound system, which is like a protein and a drug that you want to neutralize the
说一个 unbound system,就像一个 protein 和一个 drug,你想要 neutralize 这个
39:46
And so you want this drug to attach or bind to the particular protein.
所以你想要这个药物去attach或者bind到特定的protein上。
39:50
And so you want to figure out how much does it want to bind to this particular pocket
所以你要搞清楚它到底有多想bind到这个protein上的那个特定的pocket。
39:56
on the protein.
所以你需要计算unbound state和bound state之间的free energy difference,unbound state就是这两个东西分开的状态,bound state就是它们在一起的状态。
39:57
And so you need to compute the free energy difference between the unbound state where
这是一个非常重要的问题。
40:01
these two things are for a part and the bound state where they're together.
因为这是化学家通常都会想要做的事情。
40:04
That's a very important problem.
但现在你可以用现代的machine learning方法,比如diffusion models,来加速这个过程。
40:07
Because it's something that a chemist would typically want to do.
因为那是化学家通常会想做的事情。
40:10
But now you can use modern machine learning methods like diffusion models in order to accelerate
但现在你可以用现代 machine learning 方法,比如 diffusion models,来加速
40:15
those particular calculations.
那些特定的计算。
40:17
So that's a clear way in which you can start from a machine learning method and help
所以这是一个清晰的方式,你可以从一种 machine learning 方法出发,去帮助
40:22
the chemists do their calculations better.
化学家们把他们的计算做得更好。
40:25
But then the other direction is also true, right, because there is things which have been
但反过来也一样,对吧,因为有一些东西已经
40:29
developed in stochastic thermodynamics, like some like a concept called counter diabetic
在 stochastic thermodynamics 中发展出来,比如有一个概念叫 counterdiabatic
40:35
It's something that they have figured out on how to do very efficient control of particular
这是他们已经搞明白的,就是如何对特定的
40:41
physical systems.
物理系统进行非常高效的控制。
40:44
And those concepts can now be used in diffusion models in order to do a better job at generating
而这些概念现在可以用于 diffusion models,以便更好地生成图像,因为事实证明,这非常有助于降低 diffusion 过程中的统计误差或噪声。
40:51
images because it turns out that that's very helpful in bringing down sort of statistical
图像,因为事实证明这其实非常有助于降低某种 statistical
40:59
error or noise in the diffusion process.
所以这实际上可以帮助你构建更好的 diffusion models。
41:03
And so it can actually help you to build better diffusion models.
所以它确实可以帮助你构建更好的 diffusion models
41:06
So this is a cross fertilization between these two fields.
所以这是这两个领域之间的交叉融合。
41:10
And do you see on the machine learning side, the primary target of the analogy or relationship
那在 machine learning 方面,你认为这个类比或关系的主要目标是什么?
41:19
as being diffusion models, or is it more fundamental than that?
那么从 machine learning 的角度来看,你认为这种类比或关系的主要对象是 diffusion models,还是比这更根本?
41:25
It's more fundamental than that.
这比那更根本
41:27
In fact, the book goes into variation of autoencoders.
实际上,这本书深入讲解了autoencoders的各种变体。
41:32
It goes into MCMC methods, Markov chain Monte Carlo methods, which are used to sample,
它涵盖了MCMC方法,也就是Markov chain Monte Carlo方法,用来从特定分布中采样。
41:39
sort of sample from particular distributions.
它还涉及free energy estimation,以及很多不同的应用和交叉融合。
41:42
It goes into free energy estimation, it goes, there's many different applications and cross
所以我觉得作为一个领域,我们其实有点挣扎——我们在创造越来越复杂的foundation models,但对它们的理解仍然停留在表面,比如在它们的mechanics层面。
41:48
fertilizations that can happen.
可能发生的 fertilizations
41:50
So I think as a field, we struggle a little bit with, we're creating these ever more
所以我觉得作为一个领域,我们有点挣扎,我们在创造这些越来越复杂的 foundation models,但对它们的理解仍然停留在表面,比如在它们的 mechanics 方面。
41:57
complex foundation models, our understanding of them, still somewhat surface level, like
复杂 foundation models,我们对它们的理解仍然有些表面,比如
42:06
in terms of their mechanics.
就它们的机制而言。
42:08
We've thermodynamics as a field is much more mature.
Thermodynamics 作为一个领域要成熟得多。
42:14
And do you see us being able to use this relationship to better understand the models
你觉得我们能不能利用这种关系,从核心上更好地理解这些 models?
42:23
kind of at their core?
当然能。
42:25
Absolutely.
至少它提供了一个不同的视角来理解事物。
42:26
At least it gives a different perspective in trying to understand things.
所以我会说,stochastic thermodynamics 这个领域本身其实相当新,很有意思。
42:30
So I would say the field of stochastic thermodynamics actually itself quite new, interestingly.
所以它仍然是一个非常活跃的研究领域。
42:36
So there's still very actively resourge.
所以仍有一股非常活跃的复兴势头。
42:38
So thermodynamics is systems in equilibrium.
所以热力学是研究处于平衡态的系统。
42:41
Stochastic thermodynamics is systems out of equilibrium.
Stochastic thermodynamics 讲的是非平衡态的系统。
42:43
And that's actually quite new field.
这其实是个相当新的领域。
42:47
But it gives you a different way to try and understand how diffusion models work.
但它给了你一种不同的方式来试着理解 diffusion models 是如何工作的。
42:54
Because concepts like heat and work and entropy production, these things we don't use when
因为像 heat、work 和 entropy production 这些概念,我们在谈论 machine learning models 的时候
43:05
we talk about machine learning models.
是不会用到这些东西的。
43:08
But I think they give you a very new and interesting perspective on how to think about what's
但我认为它们提供了一个非常新颖有趣的视角,去思考正在发生什么,
43:13
going on and how to also improve them.
以及如何改进它们。
43:17
And so, yeah, I think in reverse, it's also true, right?
所以,是的,我觉得反过来也同样成立,对吧?
43:21
Machine learners have their own way of thinking about how to improve models.
做machine learning的人有他们自己关于如何改进模型的思考方式。
43:25
And that could also help the physicists to think about problems.
而这也可以帮助物理学家去思考问题。
43:28
And I've found both parties actually to be very interesting, interested in the other
而且我发现双方实际上都对彼此的工作非常感兴趣,也很有意思。
43:34
So that's good.
也许为了更进一步地抽象,你最近做了一个主题演讲,那个我很清楚。
43:35
Maybe to take the next step in getting even more abstract, you recently did a keynote
而且除了谈到你在CUSP的工作和那本书之外,你还提出了另一个类比。
43:44
And in addition to talking about your work at CUSP and the book, you put forth another analogy,
而且除了谈到你在CUSP的工作和这本书之外,你还提出了另一个类比。
43:54
let's call it, or opportunity to kind of learn from the physical world, and that is
我们姑且称之为,或者说是一种从物理世界学习的机会,那就是
43:59
in looking at waves and applying that to machine learning and stochastic systems.
观察waves,并将其应用到machine learning和stochastic systems上。
44:06
Talk a little bit about that work.
稍微聊聊这方面的工作吧。
44:08
Yeah, so this is, I think, very exciting in the sense that the thing that there's two
是的,我觉得这非常令人兴奋,因为有两个
44:15
reasons why we thought we needed to think more about waves in science.
原因让我们觉得需要在科学中更多地思考waves。
44:21
And that is because first of all, waves are actually seen in the brain now because we've
首先是因为,我们现在确实能在大脑中看到waves,因为我们已经
44:27
gone from single electrode measurements to maybe hundreds or thousands of electrode measurements.
从single electrode measurements发展到了可能有几百或几千个electrode measurements。
44:33
And of course, in a single electrode, you cannot see a moving wave, a traveling wave.
当然,在single electrode上,你看不到移动的wave,也就是traveling wave。
44:36
But if you put hundreds or thousands in them, you can actually see these traveling waves.
但如果你把几百上千个放进去,你确实能看到这些 traveling waves。
44:39
And people have now observed them everywhere.
现在人们已经在各处都观察到它们了。
44:44
And then the question becomes, is this function always just a side effect or something?
然后问题就变成了,这个 function 是不是总只是一个 side effect 之类的?
44:49
That's the first reason.
这是第一个原因。
44:50
The other reason is that in neural networks, you have this phenomenon called over smoothing,
另一个原因是,在 neural networks 里,有一种现象叫 over smoothing,
44:55
which means that you start with a piece of information and you have thousands of layers.
也就是说,你从一条信息开始,后面有成千上万层。
45:00
And this information basically is exponentially suppressed as you go through it.
而这条信息在穿过这些层时,基本上会被 exponentially suppressed。
45:06
So it's the input and the output of the neural network become independent of each other.
所以这个 neural network 的 input 和 output 就变得互相独立了。
45:12
Of course, it's something that you really do not want because you want the output to say
当然,这是你绝对不想要的,因为你希望output能说出关于input的某些信息,比如这个特定input的class是什么。
45:16
something about the input, like maybe what's the class of this particular input.
但是一直以来,要确保信息从input层经过成千上万个layers传到output层,总是有点困难。
45:19
But it's always been a little hard to make sure that the information travels all the
人们用了很多技巧来实现这一点。
45:25
way from the input to the output layers through thousands of layers.
但我觉得我们在deep learning中开发的很多技巧,恰恰就是为了做到这一点。
45:29
People have used many tricks to do that.
另一个应用则是reasoning。
45:32
But I feel many of the tricks we have developed in deep learning is exactly to try to do that.
所以如果你需要对问题进行reasoning,有时候就得在memory里保留一些东西。
45:38
And in the other application is in reasoning.
而另一个应用是在推理方面。
45:41
So if you have to do reason about the problem, you have to sometimes keep things in memory.
所以如果你要推理一个问题,有时候你得把一些东西记在memory里。
45:47
Of course, we kind of write it maybe to shorter memory.
当然,我们可能会把它写到更短的 memory 里。
45:51
And then we read and write from this kind of memory.
然后我们从这种 memory 里读写。
45:53
It's a quite a more stable piece where things don't get mixed up so fast.
这是一个更稳定的部分,东西不会那么快就搞混。
45:57
So you need some kind of memory.
所以你需要某种 memory。
46:02
And basically the question becomes how do you communicate very deep into time or reason
基本上问题就变成了,你怎么和很深的时间交流,或者推理
46:11
very deep into time or how do you communicate with things that are very far away in your brain
很深的时间,或者你怎么和你大脑里离得很远的东西交流
46:16
or take information from a material that actually has very long range interactions.
或者从一种实际上有 very long-range interactions 的材料中获取信息。
46:25
And also in materials, materials use waves for long range interaction, which are called
而且在材料中,材料用波来进行 long-range interaction,这些波叫做
46:31
Phonons are lattice vibrations and they carry information over long distances in a material.
Phonons 是 lattice vibrations,它们能在材料中长距离传递信息。
46:40
In fact, if you think about the universe, the only reason we can see very far and deep
事实上,如果你想想宇宙,我们能看得很远很深、深入宇宙的唯一原因,
46:46
into the universe is because light waves hit our instruments in our eyes.
就是 light waves 击中我们眼睛里的仪器。
46:53
And that is, again, waves that are doing the communication over these long distances.
而这又是 waves 在长距离中完成信息传递。
47:00
But it's now interesting is that in neural networks, we don't use this tool at all.
但有趣的是,在 neural networks 里,我们完全不使用这个工具。
47:03
So we create maps and the maps, you know, map numbers to new numbers.
所以我们创建 maps,而这些 maps,你知道的,把数字映射到新的数字。
47:11
And so the question becomes can we actually start using this phenomenon, this wave phenomenon
所以问题就变成了,我们能不能真正开始利用这个现象,这个 wave phenomenon。
47:16
that we see in the brain, can we also start to use it in neural networks.
我们在大脑里看到的那些 waves,能不能也开始用在 neural networks 里呢?
47:19
And so this has been the work that we've been developing and especially for, so we've
所以呢,这就是我们一直在做的工作,尤其是,我们...
47:24
basically designed neural networks that naturally create these waves, it's moved this information
基本上就是设计了能自然产生这些 waves 的 neural networks,它把这些信息移动
47:30
around in what we call channels or memory channels, or you can call them capsules where
到我们叫做 channels 或 memory channels 的地方,你也可以叫它们 capsules,在那里
47:36
this information gets sort of moved around and stays stable.
这些信息会被移动来移动去,但保持稳定。
47:41
And we've been very successful in, for instance, task that requires memory.
而且我们非常成功地处理了一些需要 memory 的任务,比如
47:46
So a task where a neural network or an RNN, basically, you know, the task is you take
就是让 neural network 或 RNN 去做,基本上,你知道,任务是拿一个
47:52
a number, you hold it in memory, and then as a random other point, another number arrives,
数字,把它记在 memory 里,然后在某个随机的时间点,另一个数字就来了。
47:56
you hold it in memory, and then you add them, and then another random point you're asked
你在记忆里存着它,然后你把它们加起来,接着又有一个随机点需要你
48:01
to produce the sum.
给出总和。
48:03
And that point could be very far into the future.
而那个点可能出现在很远的未来。
48:05
So you have to do the computation and you have to hold things in memory.
所以你必须做计算,而且要在记忆里保持这些东西。
48:09
And these models that naturally work with these waves, they can actually, they can do
而那些天然处理这些波的模型,它们实际上能做
48:13
these tasks much, much better than RNNs.
这些任务比RNN好得多,好太多了。
48:16
And I hear RNNs and long-running memory, I think, about challenges like exploding gradients
而我一听到RNN和长期记忆,就想到像是exploding gradients这样的挑战
48:24
and then I think about kind of the oscillatory nature of waves and that that has some inherent
然后我又想到波的振荡性质,那里面有某种固有的
48:32
dampening properties that allow you to kind of access memory further into the future
这种 dampening 特性让你可以某种程度上在更远的未来 access 到 memory
48:37
without, you know, these exploding gradient challenges, is that part of the mechanism
而且你知道,没有 exploding gradient 这些挑战,这是不是 mechanism
48:42
at work here?
在这里起作用的一部分?
48:43
That's a very good point.
这个观点非常好。
48:45
So if you think of this a neural network as a dynamical system, so that's basically,
所以如果你把 neural network 看作一个 dynamical system,基本上就是,
48:50
you think of the layers as time, and you're trying to get a signal from the beginning and
你把 layers 看作 time,然后你试图从开头获取一个 signal,并且
48:54
you're trying to propagate it through the layers over time.
你试图让 signal 随着 time 通过各个 layers 进行 propagate。
48:57
So a couple of things can happen, just different types of dynamical system.
所以可能会有几种情况,也就是不同类型的 dynamical system。
49:03
The first one has a stable system where basically, if you take two different inputs, they
第一个有一个稳定的系统,基本上,如果你拿两个不同的输入,它们
49:10
map to this, you know, after a while the collapse onto a same point and then from there
映射到这里,你知道,过一段时间后,它们会collapse到同一个点,然后从那里
49:14
on that one point moves forward.
开始,那个点继续向前移动。
49:16
So basically every converges onto what's called a point, a table point, and then from there
所以基本上每个都会收敛到一个所谓的点,一个stable point,然后从那里
49:23
on basically that's what propagates.
开始,基本上那就是传播下去的东西。
49:26
That's not what you want because, you know, information gets lost because everything gets
这不是你想要的,因为,你知道,信息会丢失,因为所有东西都
49:31
mapped to a point.
被映射到一个点。
49:32
Now that's an imploding gradient if you wish, that's something where the information gets
现在,如果你愿意的话,这就是一个imploding gradient,这是一种信息会
49:38
There's another extreme where you take two input points, I say two images that are similar,
还有另一种极端情况,你取两个 input points,比如说两张相似的图像,
49:44
and then they start to sort of wildly move around, and that's called chaos.
然后它们开始有点疯狂地移动,这就叫做 chaos。
49:49
So if it's a non-linear system, so this starts to wildly move around, and they also, you
所以如果是一个 non-linear system,它就会开始疯狂地移动,而且它们也,你
49:55
also lose the information, but for a different reason.
也会丢失信息,但原因不同。
49:59
The information is still there, but you cannot track it numerically.
信息仍然存在,但你不能 numerically 追踪它。
50:03
Your numerical precision is too small, and that's actually what we mean when we say entropy
你的 numerical precision 太小了,这实际上就是我们说 entropy
50:09
That means it basically becomes a random, it's a random process, and we all know that
那基本上就变成随机了,它是一个 random process,而且我们都知道,
50:15
a mark of chain is a random process that depends on the previous step, it will lose information
Markov chain 是一个依赖于前一步的 random process,它会丢失与输入相关的信息,
50:21
about the input because we can track it.
因为我们可以追踪到它。
50:24
That's also not what you want.
那也不是你想要的。
50:26
So what you really want is sitting somewhere in the middle, somewhere that's not too stable
所以你真正想要的是处在中间的某个位置,一个不是太稳定,
50:31
and not too unstable, and that's called edge of chaos.
又不过于不稳定,这就是所谓的 edge of chaos。
50:36
And people have found that, in fact, neural networks that perform best operate at this edge
而且人们发现,事实上,表现最好的 neural networks 就是在 edge of chaos 上运行的。
50:43
of chaos, and people have also found that, unsurprisingly, the brain operates at the edge
人们还发现,毫不意外的是,大脑也运行在这个 edge 上。
50:50
And we found that if you include these waves in your neural network, you, very naturally
而且我们发现,如果你把这些waves纳入你的神经网络,你非常自然地
50:59
operate in this regime edge of chaos, and you don't have to find you in the system to
在这个edge of chaos的regime里运作,而且你不需要在系统中找到自己才能在那里。
51:05
be there.
要进入这个edge of chaos regime非常容易。
51:06
It's very easy to get to this edge of chaos regime.
然后,如果时间允许,我们可以谈谈我们是怎么做到这一点的——这是通过spontaneous symmetry breaking实现的,但我觉得这是相当技术性的讨论。
51:12
And then maybe if we have time, we can talk about how we do this, and this is done by
那么什么是spontaneous symmetry breaking呢?
51:16
spontaneous symmetry breaking, but that's a rather technical discussion, I think.
spontaneous symmetry breaking就是一个具有特定对称性的系统——比如完美的translation symmetry,就像一种气体具有完美的translation symmetry那样。
51:22
So what is spontaneous symmetry breaking?
那么什么是 spontaneous symmetry breaking?
51:24
Okay, this is a very interesting phenomenon in physics.
好的,这是物理学中一个非常有趣的现象。
51:28
So here's another example of something quite deep from physics that we can start to use
那么,这里还有一个来自物理学的相当深奥的例子,我们可以开始用它来理解 neural networks,这也是你之前问过的问题之一。
51:33
to understand neural networks, which is one of your previous questions.
所以 spontaneous symmetry breaking 就是:一个系统具有某种特定的对称性,比如完美的 translation symmetry。你可以想象气体就有完美的 translation symmetry,或者液体更合适。然后液体在温度降低时,实际上会凝结成固体,就像水变成冰一样,对吧?
51:38
And so spontaneous symmetry breaking is where a system that has a particular symmetry,
所以 spontaneous symmetry breaking 就是,一个系统在具有特定对称性时,
51:44
let's say a perfect translation symmetry, just think of a gas as perfect translation symmetry
比如说完美的 translation symmetry,就把气体想象成完美的 translation symmetry。
51:50
or a liquid better.
或者液体更好。
51:52
And then the liquid at lower temperature, the liquid will then actually condense it into
然后温度更低的那个液体,就会真的把它凝结成
52:00
as a solid, you can go from water to ice, right?
固态,你可以从水变成冰,对吧?
52:07
And if you look at ice, you know, it has a crystal structure, so it has less symmetry
你看冰,你知道,它有crystal structure,所以它的symmetry更低。
52:12
now, because it's now the continuous translation group has now been turned into a discrete translation
现在,因为原来的continuous translation group现在已经变成了一个discrete translation
52:19
group where you can only translate over the lattice spacing.
group,你只能沿着lattice spacing来平移。
52:23
And so you've actually broken the symmetry into something smaller.
所以你实际上是把symmetry破缺成了更小的东西。
52:27
And there is a very deep theory from physics that says when that happens, when a continuous
物理学里有一个非常深刻的理论,它说当这种情况发生,也就是当一个continuous
52:32
symmetry breaks into a discrete symmetry or breaks into a smaller symmetry, then you'll
symmetry破缺成discrete symmetry,或者破缺成更小的symmetry时,那么你就会有
52:39
have new wave-like modes, which can propagate without using any information or with very,
新的wave-like modes,它们传播时不依赖任何信息,或者只需要非常,
52:50
very little information.
非常少的信息。
52:51
And these are these phonons or light, or these are the waves that can travel very long distances.
而这些就是这些 phonons 或光,或者说这些是能传播非常远距离的波。
52:59
They don't decay, they don't disperse, they hold the shape, you start with a certain
它们不会 decay,不会 disperse,能够保持形状,你从一个特定的
53:04
shape, then the shape can just translate and just move.
形状开始,然后这个形状就可以直接 translate 和移动。
53:09
And so we thought that's a really good candidate for these traveling waves.
所以我们觉得那真的是 traveling waves 的很好的候选者。
53:13
And so what we did is we created the neural network with the large symmetry, we just
所以我们做的就是,我们创建了一个具有 large symmetry 的 neural network,我们只是
53:17
bake it in, it's not like rotation or translation that we typically use, we just give every neuron
把它 bake in,这不像我们通常用的 rotation 或 translation,我们只是给每个 neuron
53:25
some extra dimensions and we say there is some symmetry.
一些额外的 dimensions,然后我们说存在一些 symmetry。
53:30
And then we do something, we generate random weights, we generate random activations,
然后我们做一件事,我们生成 random weights,我们生成 random activations,
53:38
and if you make the weights large enough, the distribution from which you sample the
如果你把 weights 设得足够大,你从中 sample weights 的 distribution 也足够大,原本的 symmetry 就会被打破,这个解释起来很复杂,但总之 symmetry 被打破了。
53:44
weights large enough, the original symmetry gets broken, and that's complex to explain,
现在你就处于这个 broken symmetry phase 里,有这些 Goldstone modes,或者说这些 waves,它们可以在不消耗任何能量的情况下传播,而且确实如此。
53:50
but the symmetry gets broken.
所以当你创造出那种情况时,你就会发现它们,你会得到这些 waves,你能非常自然地看到东西在周围 oscillate。
53:52
And now you're in this broken symmetry phase where you have these goldstone modes or
然后这个想法是,你不能 train neural network 去利用这些 oscillations。
53:56
these waves which can travel without any using any energy, and they do.
这些波可以在不消耗任何能量的情况下传播,而且确实会这样。
54:01
So you find them when you create that, you get these waves where you can just see things
所以当你创造那种条件时,你得到这些波,你能看到事物
54:05
oscillate around, very naturally.
非常自然地振荡来振荡去。
54:09
And then the idea is, you cannot train the neural network to make use of these oscillations
然后这个想法是,你没法训练 neural network 去利用这些振荡,
54:15
which are basically for free, they're running there without costing any energy, and they're
基本上是免费的,它们就在那里运行,不耗费任何能量,而且非常稳定。
54:19
very stable.
所以它们也从 neural network 的开头一直加到 neural network 的末尾。
54:20
So they also add all the way from the beginning of the neural network to the end of the
而且你无法设计系统让它拥有这些 waves,并且拥有这种 edge of chaos 的行为。
54:23
neural network.
所以信息从 neural network 的 input 开始,通过所产生的这些 wave-like patterns,一路传播到 neural network 的 output。
54:25
And you cannot design the system so that it has these waves and it has this edge of chaos
你也没法设计系统,让它具有这些波和这种 edge of chaos。
54:31
So information from the beginning, the input of the neural network propagates all the way
所以从一开始,neural network 的 input 就一路传播
54:36
to the output of the neural network through these wave-like patterns that are created
到 neural network 的 output,通过这些产生的 wave-like patterns
54:42
by this phenomenon called spontaneous symmetry breaking.
通过这个叫做 spontaneous symmetry breaking 的现象。
54:45
And so here we use something a deep result from physics as a design principle for neural
所以我们在这里使用了物理学中一个很深刻的结果,作为 neural networks 的设计原则。
54:50
networks.
嗯,我很喜欢你能玩转所有这些物理工具,把它们引入 machine learning,试图找到 correlations,以及我们能利用它们的方法。
54:51
Well, I love that you're getting to play with all of these tools from physics and bringing
我觉得后面还有很多。
54:56
them into machine learning and trying to find the correlations and ways that we can take
数学和物理学中有一个非常丰富的领域,都可以被 machine learning 利用。
55:03
There's many more to come, I think.
我觉得还会有更多。
55:04
There's a very rich field of mathematics and physics that can all be leveraged by machine
数学和物理学中有一个非常丰富的领域,都可以被机器利用。
55:10
But it's beautiful because the cross-fertilization goes in two directions, right?
但这很美好,因为这种交叉融合是双向的,对吧?
55:14
Because with our cost platform, we are using AI tools to help the scientists discover new
因为通过我们的 cost platform,我们用 AI tools 来帮助科学家发现新的
55:23
materials and accelerate their simulation tools by using AI.
材料,并通过使用 AI 来加速他们的 simulation tools。
55:30
And so it's really a healthy cross-fertilization between these two communities, which are
所以这真的是这两个社区之间一种健康的交叉融合,而这些社区非常
55:36
Well, Max, it's been great catching up with you.
嗯,Max,和你聊聊近况真是太好了。
55:40
You have a lot going on, you're working on a lot of different angles.
你有很多事情在进行,在从很多不同的角度开展工作。
55:44
It's been enjoyable to learn about them, looking forward to keeping in touch.
了解它们很有趣,期待保持联系。
55:49
Well, thank you, Sam.
嗯,谢谢你,Sam。
55:50
It was a pleasure again to talk to you.
再次和你交谈很愉快。