Latent Space
🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
2026-08-26 ·
Anima Anandkumar, Bren Professor of Computing, joins Latent Space to discuss why today’s foundation models are largely for language rather than physics. She explains how neural operators, especially Fourier neural operators, can learn mappings between function spaces across resolutions and combine data with physical constraints. The conversation covers AI weather and climate modeling, including ForecastNet and the use of Lean/TorchLean for verifying neural-network behavior in control and simulation.
本集 Latent Space 的嘉宾是 Anima Anandkumar,Bren Professor of Computing。她讨论的主题是为物理系统构建 foundation models,以及为何当前的语言 foundation models 不能直接迁移到 physics。她主张用 neural operators(尤其是 Fourier neural operator)学习函数空间之间的映射,从而支持多尺度、任意分辨率和数据与物理约束结合。她还谈到 AI 在天气与气候建模中的应用,例如 ForecastNet 能以较小数据量和计算成本接近传统天气预报,并延伸到 climate modeling。她还介绍了用 Lean/TorchLean 对神经网络进行形式化验证,以用于 control loops 等需要可靠性保证的场景。
00:00
So we, you know, set out looking for interesting examples and one of them was like weather modeling
于是我们,你知道,开始去寻找一些有趣的例子,其中一个就是像 weather modeling 这样的。
00:06
because the weather data is open source. And so given that the data was there, we were like,
因为天气数据是 open source 的。所以既然数据就在那,我们就想,
00:11
okay, let's just go try it, right? And that's the beauty of it whenever data is available,
好吧,直接去试试嘛,对吧?这就是它的美妙之处,只要数据可得,
00:15
it's really good news. But a lot of weather scientists did caution us back then. This was back in
那真的是好消息。但当时很多气象学家确实提醒过我们。那是在
00:20
21 and they said, no, no, no, this is so difficult. You know, there have been decades of development
21 年,他们说,不不不,这太难了。你知道,传统的 weather forecasting 已经发展了几十年了,
00:26
traditional weather forecasting, and that's very careful. Bottom-up physics-based modeling,
而且非常谨慎,是 bottom-up 的 physics-based modeling,
00:31
right? So assuming, oh, this is the fluid dynamics, can you go predict the weather, the next day,
对吧?所以假设,哦,这是 fluid dynamics,你能预测下一天的天气,
00:38
and so on? And so that's how a lot of the thinking was that AI is just not going to be able to beat
等等?所以当时很多人的想法就是,AI 根本赢不了。
00:45
the decades of work in weather modeling. But to our surprise, we just went ahead, we trained them,
weather modeling 领域几十年的工作。但出乎我们意料的是,我们直接就去训练了它们,
00:51
we used neural operators to be able to effectively capture the phenomena. And then we found that
我们用 neural operators 来有效地捕捉这些现象。然后我们发现,
00:59
it's not only accurate, it's almost as close to what the traditional weather models can do accurately,
它不仅准确,几乎接近传统 weather models 的准确度,
01:06
but also tens of thousands of times faster. So what would take a big supercomputer to run,
而且还快上数万倍。所以以前需要大型 supercomputer 才能跑的东西,
01:11
can now be run? And we only needed a consumer grade like GPU, like, you know, it was a small model,
现在也能跑了?而且我们只需要一个消费级的 GPU,就是,你知道,模型很小,
01:18
it fit very well, it's very fast, and it's accurate. And I think that just changed everybody's
它非常合适,速度很快,而且很准确。我觉得这真的改变了每个人的
01:24
thinking. Welcome to LeanSpace. This is the AI for Science section of LeanSpace. I'm Brandon. I work
想法。欢迎来到 LeanSpace。这里是 LeanSpace 的 AI for Science 板块。我是 Brandon,我在
01:33
on RNA therapeutics using AI at Atomic AI. I'm joined by my co-host, RJ Honakki, who develops
Atomic AI 用 AI 做 RNA therapeutics。和我一起的是我的搭档 RJ Honakki,他开发
01:40
spatial transcriptomics in the CTO and founder of Mirroromics. Today we're excited to be joined
Mirroromics 的 CTO 和创始人,谈 spatial transcriptomics。今天我们很高兴邀请到
01:45
by Anima Anankamar, a deep brain professor of mathematics and computer science at Caltech.
Anima Anankamar,一位在 Caltech 任教的数学和计算机科学教授,思想深邃。
01:53
Anima has done all sorts of really cool work combining AI with basically models of the physical
Anima 做过各种超酷的研究,把 AI 和基本上物理
01:59
world and has a really diverse background. I don't think I could even remotely cover it.
世界的模型结合起来,而且她的背景非常多元。我觉得我完全没法全面概括。
02:04
But anyway, Anima, introduce yourselves. Thank you for coming in the show.
不过总之,Anima,介绍一下你自己。谢谢你来到节目。
02:08
Yeah, thank you, Brandon and RJ. It's a pleasure to be there. And I really like the term
是的,谢谢 Brandon 和 RJ。能来很高兴。而且我真的很喜欢这个术语
02:13
latent space because that very much figures in a lot of my work because it's really, you know,
latent space,因为它在我的很多工作中都非常重要,因为,说真的,你知道,
02:20
the world is latent. But yeah, just as a brief introduction, you know, I've been working in AI
世界是 latent 的。不过嘛,简单介绍一下,你知道,我一直在做 AI。
02:26
for more than two decades in a way before even deep learning when a lot of the theoretical
在超过二十年的时间里,甚至在 deep learning 出现之前,当时 probabilistic models 的很多理论基础都必须被建立起来。我做过这些。然后随着 deep learning 开始起飞,我也一直有只脚踩在工业界,直到最近。所以我曾在 NVIDIA 领导那边的 AI 研究。再之前,在 Amazon Web Services,我创立了 cloud AI 团队,并在差不多十年前构建了第一个面向 AI 产品的 cloud。所以,你知道,像这样一只脚在工业界、一只脚在学术界,我觉得给了我很多有趣的视角,知道如何把 theory 和 practice 结合起来,去思考大规模的 AI,同时也是有理论依据的 AI。你的很多工作都与使用某些...对物理系统进行建模有关。
02:33
foundations had to be built for probabilistic models. I worked on them. And then as deep learning
得先为probabilistic models打好基础。我参与过这些工作。然后随着deep learning
02:39
started taking off, I also had a foot in industry until recently. So I was at NVIDIA, I led AI
开始腾飞,我也一直一只脚踩在产业界,直到最近。所以我在NVIDIA,领导AI
02:44
research there. And before that, at Amazon Web Services, I found the cloud AI team and
研究。在那之前,在Amazon Web Services,我创立了cloud AI团队并
02:50
built the first cloud for AI products back almost a decade ago. So, you know, like kind of having
在差不多十年前打造了首个面向AI产品的cloud。所以,你知道,就好像同时有
02:58
this one foot in industry in academia, I think has given me a lot of interesting perspective of how
一只脚在产业界,一只脚在学术界,我觉得这给了我很多有趣的视角,关于如何
03:05
to bring theory and practice together and think of AI at large scale, but also AI that is
将理论和实践结合起来,从大规模角度思考AI,同时也要让AI是
03:11
principled. A lot of your work has been related to the modeling of physical systems using certain
有原则的。你的很多工作都涉及使用某些
03:18
types of physical systems, which you model with differential equations and you help model them
你用 differential equations 建模的那类物理系统,然后用 machine learning 来辅助建模。
03:23
with using machine learning. So maybe first let's go in and talk a little bit about that as a high
所以也许我们先从 high level 聊一下这个,不过我们也会讲到 single law freighters 的细节,以及后面的一些应用,比如天气预测。
03:30
level, but we'll get to kind of the details about the single law freighters and some of the
但首先,我其实特别好奇想听听 Torchling 和你最近在做的这项工作是怎么和更大的 research program 连接起来的。
03:34
applications like whether later. But first, I'm actually really curious to hear about
对我来说,粗略来说,我的 thesis 就是把 AI 和科学结合在一起,对吧?
03:38
Torchling and how this recent work even doing connects with that larger research program.
所以,大概十年前我开始在 Caltech 的时候,你知道,我的热情一直是科学,也就是物理。但我在做 AI,所以怎么...
03:46
To me, broadly, like, you know, my thesis is AI and science, how we bring that together,
对我来说,宽泛地说,就像,你知道,我的论题是AI和科学,我们如何将它们结合起来,
03:52
right? So, you know, when I started at Caltech almost a decade ago, that's when
对吧?所以,你知道,差不多十年前我刚在Caltech开始的时候,那就是
03:58
you know, my passion was always science was physics. But you know, I was doing AI, so how to
你知道吗,我的热情一向是在科学,也就是物理。但你知道吗,我当时在做AI,所以怎么
04:04
bring that together, I was where, you know, the first kind of foundations got laid there.
把这些结合起来,我当时就在那里,你知道,最早的那些基础就是在那里奠定的。
04:10
And to me, like, you know, there are several aspects to that. One is people have been
而对我而言,你知道,这有几个方面。一是人们一直在
04:16
thinking how to use language models for science. Yes, you can do a lot of hypothesis generation,
思考如何用 language models 来做科学。是的,你可以做很多 hypothesis generation,
04:22
you can have ideas, but ideas are not enough, right? So you can have a lot of ideas.
你可以有想法,但想法不够,对吧?所以你可以有很多想法。
04:29
The bottleneck is going, testing and verifying that they work in the real world.
瓶颈在于去测试和验证它们在真实世界里是否有效。
04:34
And so this aspect is where a lot of my recent focus has been on how do we ensure that we can
所以这方面是我最近关注很多的地方:我们如何确保能够
04:41
build AI that has guarantees that it will work in the physical world or any aspects in scientific
构建有保证的 AI,让它能在物理世界或科学领域的任何方面
04:50
domains. And one way to think about it is, you know, can we model the physical world and keep
都起作用。一个思考方式是,你知道,我们能否对物理世界建模并保持
04:57
the physics correct? And that's where neural operators come in. The other aspect is can we verify
物理上正确吗?这就是 neural operators 发挥作用的地方。另一个方面是,我们能否以符号化的方式验证
05:05
symbolically certain aspects, so for instance, you know, if we claim that the theorem is correct,
某些方面,比如说,你知道,如果我们声称 theorem 是正确的,
05:12
we have to go verify that, you know, that's where lean as a formal language can be useful for
我们就得去验证它,你知道,这就是 Lean 作为一个 formal language 可以用于
05:18
verification. So how do we bring that together with language models is where a lot of mathematical
verification。那么如何把它和 language models 结合起来,这正是很多 mathematical
05:26
reasoning has been at the forefront. And so torch lean kind of is in that realm where we say,
reasoning 一直处于前沿的地方。所以 TorchLean 有点属于这个领域,我们说,
05:32
you know, not only that you want to verify mathematical statements, you may want to verify
你知道,不仅你想要验证 mathematical statements,你可能还想验证
05:39
what neural networks themselves claim to deliver. You know, for instance, if you're now using a neural
neural networks 本身声称要交付的东西。你知道,比如说,如果你现在用一个 neural
05:46
network and you want to ask whether it's going to be robust, say you want to use a neural network
network,你想问它是否 robust,比如说你想用个 neural network
05:52
in a control loop, you want to control, you know, whether it's a drone, whether it's a nuclear
在一个 control loop 里,你想控制,你知道的,不管是一个 drone 还是一个 nuclear reactor。
05:59
reactor. So all of this ultimately when we build AI systems with deep learning into control loops,
所以所有这些,最终当我们用 deep learning 把 AI 系统构建到 control loops 里的时候,我们想要的是 robustness。
06:06
we want robustness. And so now torch lean can help us do those verification seamlessly. So we can
而现在 torch lean 可以让我们无缝地做这些 verification。
06:14
now have neural networks be part of the verification loop and have confidence that we can use them
所以我们现在可以把 neural networks 放进 verification loop 里,并且有信心能恰当地使用它们。
06:21
appropriately. We have already discussed on the podcast lean and everyone should probably be
我们之前在播客里已经聊过 Lean,大家应该也都熟悉 neural networks。
06:27
familiar with neural networks. But neural networks seem very unconstrained. What kinds of proofs
但 neural networks 看起来非常不受约束。
06:34
are you talking about? Are you talking about bounds on the outputs inputs? What can you prove with
你说的 proofs 是指哪种?
06:40
the torch thing? Yeah. So torch lean is a no-world framework, right? So what it really enables is
你是说对 outputs 和 inputs 的 bounds 吗?
06:47
that you can now write neural networks essentially in lean. So instead of writing in like PyTorch,
现在你基本上可以用 Lean 来写 neural networks 了。所以不是像在 PyTorch 里写那样,
06:53
it's like a PyTorch like abstraction, but you can like kind of write it in lean and so it can be
它是一个类似 PyTorch 的 abstraction,但你可以用 Lean 来写它,这样它就可以被
07:00
fully formalized in lean. And then there are several implementations. You know, we have algorithms
在 Lean 里完全 formalized。然后有好几个 implementations。你知道,我们有一些算法,
07:05
for certified robustness like crown. You know, those are implemented under this framework. So
用于 certified robustness,比如 CROWN。这些算法都是在这个 framework 下实现的。所以
07:11
sorry, was it like what? Like crown? Crown. Crown is one of the. So there are different ways to
抱歉,你刚说的是什么?CROWN 吗?对,CROWN。CROWN 是其中一种... 所以说有很多不同的方式可以
07:18
bound, you know, for certified robustness. You know, how tight those bounds can be? It depends
对 certified robustness 做 bound。你知道,这些 bounds 能有多 tight,取决于
07:23
on the relaxation techniques. And sort of without going into those, there's many such algorithms.
relaxation techniques。先不展开讲那些吧,其实有很多类似的算法。
07:30
But you know, we're kind of like implementing them and enabling them in lean. So we can seamlessly
但你知道,我们基本上是在实现这些算法,并让它们在 Lean 里可用。这样我们就能无缝地...
07:36
run both, you know, we can both first kind of write down, not torch like framework, neural networks
两者都做,你知道,我们可以先把 neural networks 写下来,不是用像 torch 那样的 framework
07:44
very simply, right? And then we can also make statements about them formally and verify them. So
非常简单,对吧?然后我们还可以对它们做形式化的陈述,并验证它们。所以
07:50
all of that can be brought together in one framework. So what's an example of a bound that you could
所有这些都可以整合到一个 framework 中。那么,你能提出的 bound 的例子是什么?
07:57
claim or like so we're operating a nuclear reactor. We don't want it to melt down. What are the
或者这样,假设我们在运行一个核反应堆。我们不希望它熔毁。那么
08:02
sort of guarantees that you could provide to the outputs and outputs that would help that
你能为 inputs 和 outputs 提供什么样的保证,来帮助它
08:08
not melt down? I mean, the natural one is the certified robustness that I mentioned. So
不熔毁?我的意思是,最自然的就是我之前提到的 certified robustness。所以
08:14
saying that if your inputs are, you know, put up by a certain amount, how much is the output
也就是说,如果你的 inputs 你知道,被扰动了一定程度,output 会
08:20
going to be put up, right? The sensitivity analysis is another term. And so having those kinds
变化多少,对吧?sensitivity analysis 是另一个术语。所以拥有这些类型的
08:26
of bounds for different neural architecture, so you kind of automatically get those bounds can
对于不同 neural architecture 的 bounds,所以你会自动得到这些 bounds,然后它们能帮助我们,你知道,不仅训练 neural network 在 control loop 中表现良好,还要担心 safety、robustness 和 stability。
08:32
then help us, you know, not only train neural networks to do well in a control loop, but also
这些都是人们担心的 control system 的一部分。
08:38
worry about safety and robustness was stability. These are all part of control systems that people
所以那是一个 application 的例子。
08:44
worry about. So that's one example of an application. So it's really more broadly. The idea is you
所以它其实是更广义的。
08:50
need verification in lots of scenarios that in more neural networks, so control loops are one.
这个想法是,你需要在很多 scenarios 中进行 verification,在更多 neural networks 中,所以 control loop 只是其中之一。
08:58
Another example is, you know, we used physics inform neural networks to say solve partial
另一个例子是,你知道,我们用 physics-informed neural network 来求解 partial differential equation,或者提出能保证满足某些 physical law 的系统。
09:05
differential equations or come up with systems that guarantee to satisfy certain physical laws.
但我们也要 verify,比如说,我们的 neural network 只在 finite
09:12
But we also want to verify, for instance, that our neural network is only trained in finite
但我们也想验证,比如说,我们的neural network只在finite analysis课上训练过
09:18
precision, right? So can we overcome those requirements and what happens when we are, what are the
precision,对吧?
09:26
shortcomings because we are using this finite precision? Can we also bound those? So those are
那么我们能克服这些要求吗?当我们使用这种 finite precision 时,会发生什么,有哪些不足?
09:31
other kinds of bounds that work in torch lane. So all aspects of like, you know, the effect of
我们能不能也对那些进行 bound?
09:37
precision, the effect of perturbation, all of these we can, you know, we can have algorithms that
所以那些就是 torch lane 中起作用的其他类型的 bounds。
09:44
are implemented in lean that can be seamlessly now part of the verification loop.
所以所有这些方面,比如,你懂的,precision 的影响,perturbation 的影响,所有这些我们都可以,你知道的,我们可以在 lean 中实现算法,这些算法现在可以无缝地成为 verification loop 的一部分。
09:50
So and do is the descriptive power of the torch lane is that sufficient to describe basically any
那么,torch lane 的描述能力是否足以描述基本上任何 neural network?还是说在这方面还有其他限制?
09:57
neural network or is there other constraints on that? Yeah. So it's essentially a, you know,
是的。所以它本质上是一个,你懂的,类似 pie torch 的 framework,对吧?
10:02
pie torch like, you know, framework, right? So you can just kind of nicely define neural
所以你可以很方便地定义 neural...
10:09
layers in the same way. But the backend having like lean helps us formalize and prove it.
layers 也是同样的方式。
10:15
And for like transformer architecture, for example, is it reasonable to prove these kinds of bounds
但 backend 如果用 Lean 这种东西,能帮我们 formalize 和 prove 它。
10:21
on a very large neural network? So the, you know, there is the aspect of one is like kind of having
比如说像 transformer architecture,在一个非常大的 neural network 上去 prove 这些 bounds,合理吗?
10:27
the framework, right? The other is scalability low. So leads still has a lot of shortcomings there.
所以你知道,一个方面就是 framework 的问题,对吧?
10:34
It's CPU based and, you know, it's not like, like, getting that onto the GPU has a lot of nuances
另一个是 scalability 还很低。
10:42
there. So, you know, a lot of work needs to be done. So what we've started with is a framework,
所以 Lean 在这方面还有很多不足。
10:47
you know, making that more efficient, especially at a very large scale requires still a lot of work
它是基于 CPU 的,而且你知道,把它弄到 GPU 上有很多微妙的地方。
10:53
to be done. But that's true broadly for lean as well. And just trying to understand how
所以你知道,还有很多工作要做。
11:00
I imagine this. So if I were to say numerical analysis classes, you know, graduate level numerical
我想象一下。比如说numerical analysis课程,你知道,研究生级别的numerical analysis课,你有一个differential equation,你有一些discretization error之类的东西,然后你像这样bound——给定这些properties,我可以bound这个solution,对吧。所以用这些physics-based或AI-based的方法来解differential equations,我觉得历史上这有点像wild west。所以我想你提到了physics-inspired neural networks,这真是个很酷的想法,聊聊这个会很有趣。但我知道有时候它们很特别,而且人们并不——它们并不总是有效,而且我觉得人们并不总是知道它们什么时候会有效或无效。我的意思是,我不是专家,但我就是想知道这是不是你的经验。
11:05
analysis class, you have a differential equation. You have some discretization error or something
在analysis class里,你有一个differential equation,然后有一些discretization error之类的东西
11:09
and you've bound like given these properties, I can bound the solution. Right. So solving some of
然后你像这样界定了,基于这些性质,我可以把解给界住。对吧。所以解决这些
11:16
these physics based or AI based solutions to differential equations, I think historically it's been
基于物理或者基于AI的differential equation求解方法,我觉得历史上一直是
11:23
kind of the wireless. So I think you mentioned physics inspired neural networks. Really cool idea.
有点Wild West那种感觉。所以我觉得你提到的physics inspired neural networks,真的是个很棒的想法。
11:28
It'd be fun to talk about that a little bit. But I know that sometimes they are particular and
聊聊这个会很有意思。但我知道有时候它们特别,而且
11:34
that people don't, they don't always work. And I think people don't always know when they will or
人们不会,它们不总是有效。而且我觉得人们并不总是知道它们什么时候会有效或
11:38
won't work. I mean, I'm not an expert, but I'm just wondering if that's of been your experience.
不会有效。我的意思是,我不是专家,但我只是想知道这是否也是你的经验。
11:44
And what I'm wondering is like, has this helped you understand like the domain of applicability
我在想的是,这些有没有帮助你理解 PINNs 的 domain of applicability?也就是说,目标是不是在于你能严格地说某个 solution 会 converge?还是说对于 neural networks,其实不一定有那种在 controlled way 下的 convergence 概念?
11:50
for pins or and is that sort of like the goal is like you can rigorously say like this solution
对于 PINNs 来说,还是说,目标就是你可以严谨地说,这个解决方案
11:56
will converge or is there not necessarily the same concepts of convergence in the controlled way
是的。你知道,physics-informed neural nets 本质上是说,你写下一个 PDE(partial differential equations),然后希望 optimization 能成功,然后得到答案,对吧。当然,如果 optimization 不是一个难题,那这就通用了,什么都能解,大家都会很开心,但事实并非如此。所以 optimization 最终通常非常困难,尤其是对于 time-dependent 的问题。
12:02
for neural networks? Yeah. So you know, like physics and form neural nets are about like saying that,
对于神经网络?是的。所以你知道,像 physics-informed neural nets 就是说,
12:08
you know, I write down like a PDE partial differential equations and hopefully the optimization
你知道,我写下类似 PDE 偏微分方程,然后希望优化
12:15
succeeds. And I get the answer, right. And of course, if like optimization was not a tall
能成功。然后我得到了答案,对吧。当然,如果优化不是个难事
12:21
issue, this would be universal. You solve everything, you know, we're all happy, but that's not the case.
那这就是通用的。你解决所有问题,你知道,我们都开心,但事实并非如此。
12:28
And so optimization ends up being usually very difficult, especially for problems that are time
所以优化通常最终会非常困难,尤其是对于时间相关的
12:34
dependent, meaning it's not just stationary. You also have time and the time component in many
它是dependent的,意思就是不是静止的。你还有时间,而且在很多情况下,time component可能是turbulent的,比如在fluid dynamics里,你知道,如果你跑得足够久,就可能变得chaotic。所以你真的,你知道,非常小的fine scale effects会起作用。所以在那些情况下,想要在所有时间点上求解partial differential equation,基本上就是hopeless。这就像一个optimization landscape,我觉得我们根本没办法handle。而那种说从零开始就能用neural net来解这些方程的想法,是不可能的。所以pins不是哪儿都有用的。我们提出neural operators,就是为了克服这个问题。
12:41
cases could be turbulent, like in the case of fluid dynamics, you know, you kind of like if you
有些情况可能是 turbulent 的,比如 fluid dynamics 那种,你知道,就有点像如果你
12:48
run it long enough, you can become chaotic. So you really, you know, have like very small
让它跑得足够久,就可能变得 chaotic。所以你真的,你知道,会有非常小的
12:55
fine scale effects matter. And so in those cases, just trying to solve a partial differential
fine scale effects 变得很重要。所以在那些情况下,光想求解一个 partial differential
13:02
equation at all times is just hopeless. Like, you know, this is not an optimization landscape that,
equation 在所有时刻根本就是 hopeless。就像,你知道,这不是一个 optimization landscape,
13:10
you know, I think we'll, you know, we can have any handle on. And this is where the idea that
你知道,我觉得我们,嗯,能 handle 得了的 landscape。而这就引出了那个想法:
13:15
from scratch, we would be able to solve these equations using a neural net is not possible.
就是从头开始用 neural net 去解这些方程是不可能的。
13:23
So pins don't work everywhere. And our idea of neural operators came as a way to overcome this,
所以 PINNs 不是到处都适用。而我们提出 neural operators 就是为了克服这个,
13:29
right. So say, you know, we can't rely just on physics constraints alone to come up with answers.
对,所以你知道,我们不能光靠 physics constraints 来得出答案。
13:36
We have lots of data available, you know, I'll talk about the weather example where we when
我们有很多 data 可用,你知道,我会讲一个天气的例子,就是当我们
13:42
collect data, right? So we don't just solve equations and have synthetic data, but we also have
收集 data 的时候,对吧?所以我们不只是解 equations、用 synthetic data,我们还有
13:47
real data by observing the weather as one example. So why not make use of all of the data available?
通过观察天气得到的 real data 作为一个例子。那为什么不利用所有可用的 data 呢?
13:54
So we don't just rely on trying to solve partial differential equations and other physical problems
所以我们不只是依赖尝试从零开始解 partial differential equations 和其他物理问题,
14:00
from scratch because it's really the data driven approach that makes it possible to get quick
因为真正能让我们快速得到
14:06
answers. And so with neural operators, we can bring both of them together. We can have all
答案的,是 data driven approach。所以有了 neural operators,我们可以把两者结合起来。我们可以拥有所有
14:12
the data that's available. We can utilize it. We can add physical constraints. And then that
可用的 data。我们可以利用它。我们可以加上 physical constraints。然后那样
14:18
overcomes the limitations that pins face. Can you give a little bit more intuition on the
克服 PINNs 面临的这些限制。你能再给一点关于这个的 intuition 吗?
14:24
difference there and why that is possible? So I heard you mention, you know, in with pins,
这其中的差别是什么,为什么能做到?所以我听你提到,你知道,在 PIN 里面,
14:30
you're basically just baking the physics constraints into the neural network, but that this becomes
你基本上就是把物理约束直接烘进神经网络里,但这样做会随时间或其他变量变得不稳定,
14:36
unstable over time or other variables, whereas if you add a little bit of data, like I can kind of
而如果你加一点点数据,我大概能
14:43
intuitively understand why that might help, but can you give a little intuition for what's going
凭直觉理解那为什么会帮上忙,但你能不能给点直觉上的解释,说说这到底是怎么回事?
14:47
on? What's the difference here? So with the pin, like, you know, every instance of an equation,
这里的区别是什么?所以用 PIN 的时候,你知道,每一个方程实例,
14:54
you solve from scratch, right? At least in the classical sense. So you start, you take the
你都是从零开始求解,对吧?至少在经典意义上是这样。所以你开始的时候,你拿到的是
15:00
specification of what equation you want to solve when you hope that the optimization labs
你想求解的方程的规格说明,你希望这些优化规律
15:05
succeeds, which many cases it doesn't. Whereas with the neural operators, what we do is we have
能成功,但很多时候并不会。而使用 neural operator 时,我们的做法是,我们有
15:12
lots of great data. So we have a training phase. We teach it how to come up with solution for
大量优质数据。所以我们有一个 training phase。我们教它如何针对
15:18
different instances of equation. And so just as another supervised learning at test time, you can
不同的 equation 实例生成 solution。所以就像普通的 supervised learning,在 test time,你可以
15:23
now ask, you know, can you come up with an answer? And you can still have physics constraints as a
现在问,你知道,你能给出答案吗?而且你仍然可以把 physics constraints 作为
15:28
way to guide that. So, you know, it can be both data driven and physics informed together,
一种引导方式。所以,你知道,它可以同时是 data driven 和 physics informed 的,
15:34
but the benefit is because we have data, you know, it's like you're not stuck in an optimization
但好处是,因为我们有数据,你知道,就像你不会被困在一个 optimization
15:40
landscape, right? So you know what the answers are during training. So you're now at a better chance
landscape 里,对吧?所以你在 training 期间知道答案。所以你现在有更好的机会
15:47
to come up with the right answers even at test time. My understanding is that a neural operator
在 test time 也能得出正确答案。我的理解是,一个 neural operator
15:54
is a function fit to data or a neural network, you know, learns to fit functions to data. Is that
是一个 function 拟合到 data 上,还是一个 neural network——你知道,学习把 function 拟合到 data 上。这是
16:04
a good intuition here? Yeah. So, you know, neural operators are in that sense similar to, you know,
一个好的直觉吗?是的。所以,你知道,neural operators 在这个意义上和那个很像,你知道,
16:10
it's the same as neural networks, right? You're learning on data. But the difference is neural operators
和 neural network 是一样的,对吧?你都是基于 data 学习。但区别在于,neural operators
16:16
are you can think of it as a generalization of neural networks. So with standard neural networks,
你可以把它看成是 neural network 的一种 generalization。对标准的 neural network 来说,
16:22
the inputs and outputs are a fixed size. So in language, we have fixed vocabulary, we fix what
input 和 output 的大小是固定的。所以在 language 中,我们有固定的 vocabulary,我们固定了
16:27
the input and output are. And same with images in computer vision in videos, we assume a fixed
input 和 output 是什么。同样,在 computer vision 的 image 和 video 中,我们假设一个固定的
16:34
resolution. And we always, you know, our inputs and outputs are always at that fixed resolution.
resolution。而且我们总是——你知道,我们的 input 和 output 始终是在那个固定的 resolution 下。
16:39
We can't change it post-all. Whereas with a lot of this physical data, the idea is our world is
我们事后无法改变。而对于很多 physical data,关键在于,我们的 world 是
16:46
inherently multi-scale. So you should not be like deciding beforehand what the resolution is.
本质上是 multi-scale 的。所以你不应该事先就决定 resolution 是什么。
16:52
You know, maybe you have like whether data available only at course resolution,
你知道,也许你手上的数据只有 coarse resolution 可用,
16:57
but really the actual phenomena is happening at a finer scale, right? And maybe you want to
但实际上真正的现象是在更细的尺度上发生的,对吧?而你可能想要
17:02
after that incorporate either additional data, finer resolution or add in physical constraints
在那之后加入额外的数据、更细的 resolution,或者加入物理约束
17:09
at finer resolution. So we should be having that flexibility and we should really think of the
在更细的 resolution 下。所以我们应该有这种灵活性,我们应该真正把
17:14
world not at these fixed resolution, but one that's happening infinitely, you know, that one,
世界看作不是固定 resolution 的,而是无限发生着的那种,你知道,就是那个
17:22
the real world happens at that infinite resolution. And that's what neural operators enable
真实世界发生在那种无限的 resolution 下。而这正是 neural operators 所实现的
17:27
because they model inputs and outputs as continuous functions that can be infinitely resolved,
因为它们把输入和输出建模为 continuous functions,可以被无限地解析。
17:34
that can have infinite discretization. And now we can have, you know, at inference time,
这可以有无限的 discretization。然后现在我们可以,你知道,在 inference 的时候,
17:40
you can give it now inputs and ask for outputs at any resolution. So you're not just limited to
你可以给它新的 inputs,然后要求任何 resolution 的 outputs。所以你不再局限于
17:47
the resolution of training that we see in standard neural networks. And that's what neural operators
training 的 resolution,就像我们在 standard neural networks 中看到的那样。而这就是 neural operators 所
17:53
enable. So neural operators enable us to zoom in and out as we like.
实现的。所以 neural operators 让我们能随意 zoom in 和 zoom out。
17:59
So obviously that has stated that any function that's under constraint, right? You could have
所以显然,这也就是说任何在 constraint 下的 function,对吧?你可能会有
18:07
many, many functions that fit the data. It'll be easier to overface. So how do you regularize that?
很多很多 functions 都能 fit the data。那样就更容易 overfit。所以你如何 regularize 呢?
18:12
Yeah, certainly, like, you know, if you're asking about making predictions at a higher resolution
对,当然,比如说,你懂的,如果你要求在更高的 resolution 下做 predictions,
18:19
than what's seen, like what we call zero short super resolution, you're kind of making some guesses,
比之前见过的更高,就像我们所谓的 zero-shot super-resolution,你其实是在做一些猜测,
18:25
right? And that's what these models are doing. They're trying to regularize and kind of smoothly
对吧?
18:30
extent to higher resolution. But of course, if you now give it the model additional information
这些模型就是在做这件事。
18:37
in terms of, let's say, a physical loss. So you could give it partial differential equation,
它们试图 regularize,并平滑地扩展到更高分辨率。
18:42
constraints, conservation loss. And you can now enforce them at a finer resolution than the data
但当然,如果你现在给模型额外的信息,比如说一个 physical loss。
18:48
you have, then there's more guidance in a way. So that way it can now come up with the right answers
所以你可以给它 partial differential equation 约束、conservation loss 之类的。
18:56
even at higher resolution because you're, you know, giving it constraints at higher resolution.
然后你可以在比现有数据更精细的分辨率上去施加这些约束,这样某种程度上就有更多指导。
19:02
And so that's how we can ensure that these physics and form neural operators can work at higher
所以这样一来,即使是在更高分辨率下,它也能得出正确答案,因为你是在更高分辨率下给它约束的。
19:08
fidelity and higher resolution than even the training data that was available.
这样我们就能确保这些 physics-informed neural operators 能够在比训练数据更高的保真度和更高的分辨率下工作。
19:13
My understanding is a lot of your work uses a particular kind of an operator, a Fourier operator.
我的理解是,你的很多工作都用到了一种特定的 operator,叫 Fourier operator。
19:21
So Fourier is a dual domain. It is extended across the entire domain of the inputs. That's a
所以 Fourier 是一个 dual domain,它延伸到 input 的整个 domain 上。
19:30
lot of jargon, maybe. Can you give some intuition for why, why is that important? How does that help?
这可能有点太术语化了。你能给点 intuition 吗?为什么这个很重要?它到底有什么帮助?
19:38
What I mentioned neural operators as a class of models that allow us to have any resolution input
我刚才提到 neural operators 是一类模型,可以允许任意 resolution 的 input,
19:46
and any resolution output, right, and learns the mapping between them. You know, that's really
以及任意 resolution 的 output,对吧,然后学习这两者之间的 mapping。
19:51
called an operator. So the mapping between function spaces. So that's the reasoning behind the name
这其实就是所谓的 operator,也就是 function spaces 之间的 mapping。所以这就是 neural operator 这个名字的由来。
19:57
neural operator. And in a Fourier neural operator was one of the early setups we, or architectures,
而 Fourier neural operator 是我们早期提出的 setup 之一,或者说 architecture 之一。
20:06
we came up with. And the reason why that's been so successful is because it kind of strikes a
它之所以这么成功,是因为它有点像是……
20:12
nice trade-off between efficiency and expressivity, right? So why is the Fourier space a good one?
efficiency 和 expressivity 之间很好的 trade-off,对吧?那为什么 Fourier space 是个好的选择呢?
20:21
The Fourier space allows us to, you know, it's a dual space like you mentioned, but it really allows
Fourier space 让我们能,你知道,它像你提到的,是一个 dual space,但它确实让我们能
20:26
us to capture non-local phenomena, right? So beating something that's like a non-local in the
去捕捉 non-local 现象,对吧?所以像 non-local 这样的东西,在
20:34
Fourier domain could be even efficiently captured. And a lot of phenomena like we see in nature whether
Fourier domain 里可以被高效地捕捉到。而且我们在大自然中看到的很多现象,无论是
20:41
it's fluid dynamics, material deformation, quantum chemistry, it's all, you know, there's a lot
fluid dynamics、material deformation、quantum chemistry,都是,你知道,有很多
20:47
of them are non-local, you know, the differential equation like the derivative is local, but the
现象都是 non-local 的。你知道,比如 differential equation,derivative 是 local 的,但
20:52
inverse of it is you're kind of doing essentially integration, it's non-local, right? So the solutions
它的逆运算本质上就是在做 integration,它是 non-local 的,对吧?所以 solutions
20:58
are non-local. And these models are able to capture that, but at the same time doing Fourier
都是 non-local 的。而这些模型能够捕捉到这一点,但同时做 Fourier
21:05
transform is efficient. And it kind of like nicely captures a lot of inductive bias we see in many
Transform 是高效的。而且它有点像很好地捕捉到了我们在很多自然现象中看到的 inductive bias。但这并不意味着我们完全是在用 Fourier basis 来捕捉这个世界,对吧?它并不是 Fourier basis 里的线性表示——那是经典 numerical methods 的做法。我们在 Fourier layers 之间加入非线性,就像 transformer 和其他 neural nets 那样。我们还加了 residual connections。所以所有这些从其他 neural nets 里借鉴来的 architecture 方面的东西,那些在其他 neural nets 里效果不错的点,把它们整合在一起,确实能帮我们更好地理解世界。所以你可以这么想,如果我们用 transformer 并且要求非常高的 resolution,那么由于 quadratic complexity,就会变得不可行。
21:12
of these natural phenomena. But this doesn't mean that we are capturing the world entirely in the
能够描述这些自然现象。但这并不意味着我们就把整个世界完全捕捉进了……
21:20
Fourier basis, right? It's not a linear representation in the Fourier basis, which is what classical numerical
Fourier basis,对吧?它不是在 Fourier basis 下的线性表示,那是经典数值方法做的事。我们在 Fourier layers 之间像 transformer 和其他神经网络一样加入非线性,还加了 residual connections。所以所有这些受其他神经网络启发的架构设计——那些在别的神经网络里效果很好的东西——把它们结合在一起,真的能帮我们取得两全其美的效果。所以你可以这么想:如果我们用 transformer,又需要非常高的分辨率,那会由于 quadratic complexity 而变得不可行。但我只是在 Fourier domain 里操作呢,还是说还需要其他方面才能把它做好?
21:28
methods do. We add non-linearity just as a transformer and other neural nets in between four-year
方法确实如此。
21:35
layers. And we also add residual collections. So all of these architectural aspects that are
我们在 Fourier layers 之间添加 non-linearity,就像 transformer 和其他 neural nets 那样。
21:42
inspired by other neural nets that work well in other neural nets, bringing that together really
我们也添加 residual connections。
21:48
kind of helps us get best to put the world. So you can think of like if we were to use transformers
所以所有这些受其他运作良好的 neural nets 启发的 architectural aspects,把它们整合在一起,真的能帮助我们获得两全其美的效果。
21:54
and we require a very high resolution, it would become untenable because of the quadratic complexity
所以你可以想象,如果我们使用 transformers 并且需要 very high resolution,由于 quadratic complexity,这会变得不可行。
22:01
and all-to-all connections. On the other hand, if you did that with Fourier transforms, we have like
还有 all-to-all connections。
22:08
quasi-linear complexity and still we have global connections in a way we can model these non-local
另一方面,如果用的是 Fourier transforms,我们会有 quasi-linear complexity,而且仍然有 global connections,可以用这种方式建模这些 non-local phenomena。
22:16
phenomena. And so that's why it's a nice middle ground. So that allows you to learn from what is
所以这就是为什么它是个很好的中间点。
22:23
happening on the like it's talking whether like what's happening in Chicago may have some impact on
这样可以让你从正在发生的事情中学习,就是比如说芝加哥发生的事可能对旧金山发生的事有影响。
22:29
what's happening in San Francisco. Well, maybe not, but that's the idea. Yeah, so that's the idea and
嗯,也许不会,但就是这个想法。
22:34
time like kind of yes at this part maybe local, but eventually they have an impact in the other
对,就是这个想法,时间上也是,比如这部分可能只是 local,但最终它们会对其他位置产生影响。
22:40
locations. And yes, so both in space and time we want to capture that dependence. Yeah, so what
对,所以无论在空间还是时间上,我们都想捕捉这种 dependence。
22:47
happens today in Chicago will happen will have an impact in a month in San Francisco or something like
对,所以今天芝加哥发生的事,一个月后会对旧金山产生影响,或者类似这样的。
22:53
yeah, so you know, so there is like both the short term and the long term effect. So in a short
对,所以你看,就是有短期和长期两种效应。短期
22:58
term like we think about predictable weather, but longer term we're talking about climate, right?
来看的话,我们说的是可预测的天气,但长期来说,我们在说气候,对吧?
23:04
So what happens you'll be not be able to say precisely you know what happens in Chicago,
所以结果是,你没法精确地说出芝加哥会发生什么,
23:12
what will happen in San Francisco, that's like the butterfly effect. On the other hand, we can kind
旧金山会发生什么,这就是蝴蝶效应。另一方面,我们可以大致
23:17
of give averages. You know, if there's heat wave in this kind of overall region, you know, we kind
给出平均值。你知道,如果整个区域出现热浪,我们大致
23:23
of have one idea that it's going to be higher than average temperatures. So those are the aspects
会有个概念,就是气温会高于平均值。所以这些方面
23:29
we can capture together from an architectural standpoint for all the AI engineers here. Are we
我们可以从 architecture 的角度一起捕捉到,对在座所有 AI 工程师来说。我们是不是
23:34
just talking about doing all the work in the 4.8 domain, but it's basically the same neural network,
只是在 4.8 domain 里做所有工作,但基本上还是同一个 neural network?
23:40
but I'm just operating in the 4-year domain or is there other other aspects that are that are
但我只是在 Fourier domain 中操作,还是有其他方面需要正确做到这一点?
23:46
required in order to do this proper? So think of it, I guess the maybe the easiest way to think
所以想想看,我觉得也许最简单的思考方式,如果做我们所谓的 ensembles,也就是说你有一个关于什么的 probabilistic estimate……
23:51
about it is, you know, you can if you think of a transformer architecture instead of like the
关于这个问题,你知道,你可以这么想,transformer architecture,而不是说那种
23:58
you know, attention map, you know, have the 4-year, but you still have other nonlinearities,
你知道,attention map 那种,用 Fourier,但你还是会有其他的 nonlinearities
24:03
you have like, you know, the residual, you have, you know, many other parts of the architecture
比如说,residual,还有,你知道,architecture 里的很多其他部分
24:08
still there that give it like expressivity, and we have lifting to higher dimension like,
仍然在,给它那种 expressivity,而且我们还有提升到更高 dimension 的操作,比如说
24:14
you know, in a channel space to give it more expressivity. So all of those kind of best principles are
你知道,在 channel space 里给它更多 expressivity。所以所有这些最好的原则都
24:21
still available, but the 4-year helps us capture that all to all, you know, for dependence
仍然可用,但 Fourier 帮我们捕捉那种 all-to-all,你知道,对于 dependence,
24:29
without requiring very huge complexity that AC transformers. The other advantage is that it gives
不需要像AC transformers所要求的那种非常巨大的复杂度。
24:35
you the natural multi-scale, what's sort of an implicit cutoff is, you know, the sort of if you
另一个优势是它给你天然的多尺度,有点像是隐含的cutoff,你知道吗,如果你有signals背景或者physics背景,你可能会问,在linear的情况下,如果你完全linear地处理,会有一个maximum frequency,而高于那个你就无法表示任何东西。
24:42
have a signals background or physics background, you might ask, you know, in linear, if you're doing
但是,添加这些其他的architectural changes在nonlinear domain中,实际上会怎样影响你对frequency bounds的选择呢?
24:47
everything linearly, there's a maximum frequency and, you know, above that, you can't represent anything.
对对,不,这是个很好的问题,而这正是expressivity发挥作用的地方,对吧?
24:53
But how does that, how does adding these other architectural changes in a nonlinear domain actually
否则,如果你只是对一个signal做Fourier transform,然后试图表示它,你知道吗,这正是numerical methods也尝试过的做法。
24:59
affect, you know, your choices of frequency bounds? Yeah, yeah, no, that's a great question,
会影响你对 frequency bounds 的选择吗?对对,不,这是个很好的问题,
25:05
and that's where the expressivity comes in, right? Otherwise, if you're just taking a 4-year transform
而这就是 expressivity 发挥作用的地方,对吧?否则,如果你只是做一个 Fourier transform
25:10
of a signal and trying to represent it, you know, that's what numerical methods have also attempted
对于一个 signal,然后试图 represent 它,你知道,numerical methods 也一直在尝试这样做。
25:15
to do, and that requires very fine discretization, and that's why it's very expensive to do the
要做这个,就需要非常精细的 discretization,所以用 classical 方式做 simulations 会非常昂贵。
25:21
simulations in a classical way. And instead, if you want to move away from that and say we want to
而如果你想摆脱这个,说我们要 learn features——这就是 deep learning 的核心——那我们就不能只把它限制在 Fourier domain 里。
25:27
learn the features, which is what deep learning is all about, then we cannot force it to be only in
我们得给它 nonlinearity,让它自己去搞明白正确的 basis,你懂的,表示 signals 的最佳 basis 是什么,所以这就成了一种很好的 combination。
25:33
the 4-year domain. We have to give it nonlinearity to figure out what the right basis for, you know,
也就是说,所有这些 nonlinearities 会帮它,嗯,找到正确的 latent space。
25:39
the best basis to represent the signals are, and so that's the kind of like nice kind of combination.
不是故意玩谐音梗。
25:47
We have that it's like all these nonlinearities will help it kind of, you know, find the right
然后,你知道,如果你在那个 latent space 里做 Fourier,那可能会是一种更高效的 representation 方式。
25:53
latent space. No fun intended. And so, you know, if you do 4-year in that latent space, you know,
所以,这是一种思考方式,因为,你知道……
26:04
that may be a more efficient way to represent. So, that's one way of thinking because, you know,
那可能是一种更高效的 represent 方式。所以,这是一种思考方式,因为,你知道,
26:09
first of all, we're lifting the signal to more dimensions, even if the signal is two or three
首先,我们在把 signal 提升到更高的 dimensions,即使 signal 本身只是 two or three dimensions,
26:15
dimensions, we are now lifting it to much higher dimension. So, in that space, the idea is it's easier
我们现在把它提升到 much higher dimension。所以,在那个 space 里,想法是更容易
26:21
to learn, and we're doing it as a nonlinear lifting, right? So, there's already a latent space there,
去 learn,而且我们做的是 nonlinear lifting,对吧?所以,那里已经有一个 latent space,
26:26
and then we're doing further nonlinear transformations in between our 4-year transforms. So,
然后我们在 Fourier transforms 之间做进一步的 nonlinear transformations。所以,
26:32
that means we're saying yes, you know, maybe with this limited number of frequency modes,
这意味着我们在说,是的,你知道,也许用这么有限数量的 frequency modes,
26:38
it's not expressive enough, but when I add nonlinearities, I can, you know, I can kind of nice,
它不够 expressive,但当我加上 nonlinearities,我可以,你知道,我可以有点 nice,
26:45
more nicely capture that. So, you started your career back before neural networks where I
更好地捕捉到这一点。
26:52
guess taken off, right? So, I think back then, people really did think a lot about, you know,
所以我猜,你的职业生涯开始于 neural networks 还没起飞的时候,对吧?
26:56
appropriate basis sets and, you know, function expansions and orthogonal polynomials or whatever.
合适的 basis sets,还有,你懂的,function expansions、orthogonal polynomials 之类的。
27:03
How does that evolution from, you know, your research standpoint, like, as the community has evolved
从你的研究角度来看,这种演变是怎么发生的?就是随着整个领域逐渐发展,从原来那套
27:09
from that to, oh, just screw it, throw it all in, it seems like you still believe in at least some
到“算了,管他呢,全扔进去”。看起来你仍然相信至少有某些
27:14
of those concepts is being guiding principles. Do you think that that there is actually so less
概念是可以作为指导原则的。你觉得,实际上从经典数学里能借鉴的东西是不是真的那么少,
27:22
to be taken from, you know, classical mathematical, like reverse mathematical techniques that you
比如 reverse mathematical techniques 这类技术,你可以用它们来真正改善对现实世界的建模,
27:28
can use those techniques to actually help improve modeling of the real world, even if you so are
即使你其实就是在把所有东西都一股脑扔进去?
27:33
just throwing the kitchen sink at things? No, I think it's a nice, I think there's a trade-off
一股脑全扔进去?不,我觉得这挺好的,我觉得这里面有一个 trade-off。
27:39
involved. I mean, it's funny, my undergraduate thesis more than two decades ago now was on
我是说,有意思的是,我二十多年前的本科毕业论文就是关于
27:45
fractional four-year transform, right? So, and so yes, I mean, by themselves, like, you know,
fractional four-year transform,对吧?所以,是的,我的意思是,单独这些技术,你懂的,
27:50
that wasn't enough to do computer vision, but I was curious, okay, what are these techniques and
那还不足以做 computer vision,但我很好奇,好吧,这些技术到底是什么,
27:55
how well do they were? And so, you know, I'm completely with you that we cannot just force ourselves to
它们表现怎么样?所以,你懂的,我完全同意你的看法,我们不能强迫自己
28:02
use stone age techniques or classical techniques, right? I mean, so we have to have feature learning,
去用 stone age 的技术或者 classical techniques,对吧?我的意思是,我们必须要有 feature learning,
28:08
we have to have flexibility, expressivity, you know, they have to be easily optimized. So, all of these
我们必须要有 flexibility、expressivity,你懂的,它们还得容易优化。所以,所有这些
28:15
aspects are very important with deep learning, but when it comes to the physical world and physical
方面对 deep learning 来说都很重要,但涉及到物理世界和物理
28:22
data, it's never going to be as plentiful as we see with language models because we are, you know,
数据时,数据量永远不会像我们在 language models 里看到的那么丰富,因为我们,你知道,
28:29
our weather model, like, had about like 50,000 samples, right? 50,000 samples of fairly high
我们的 weather model,大概有 50,000 个 samples,对吧?50,000 个相当高
28:35
resolution, like, were global weather maps, but it's nothing like what we see with language and
分辨率,比如全球天气图,但它跟我们在语言和其他领域看到的完全不一样,
28:41
in other domains, it's even less because it's so expensive to simulate, and the real data may just
在其他领域,情况更差,因为模拟太贵了,而且真实数据可能根本
28:47
not be available. And so, here we have to think about the inductive biases more, we have to
拿不到。所以这时候我们得更认真地思考 inductive biases,必须
28:53
add in the physics constraints, it cannot be just reliant on data. And that's where, I think,
加入 physics constraints,不能只靠数据。而我认为,这正是
28:59
a little bit more thinking of the architectural design comes up. The other aspect is computational
更需要在 architectural design 上多花心思的地方。另一方面是 computational
29:06
complexity. So, think about language, it's just one dimension, and even there, the context lens,
complexity。你看语言只是一维的,即使如此,context length——
29:12
you know, we're getting to millions and we are struggling, right? I mean, on the other hand,
你知道,我们做到上百万已经很吃力了,对吧?我的意思是,反过来,
29:18
now we're thinking about not just 2D, 3D, even 40, you know, 3D and time, and each of the
现在我们考虑的不仅是 2D、3D,甚至是 4D——就是 3D 加时间,而且每一个维度
29:26
dimension is even a few hundred grid points, which is where, you know, industrial scale starts at,
维度哪怕只是几百个grid points,也就是,你懂的,industrial scale开始的地方,
29:32
like, like, a thousand grid points in each dimension, we're talking like hundreds of billions
比如说,比如说每个维度一千个grid points,我们说的是几千亿
29:37
to even a trillion context length, right? So, forget ever having a transformer for anything
甚至到一万亿的context length,对吧?所以,别指望transformer能处理任何事
29:43
of this scale, all of the world's computer will not be enough. And first of all, the how will I
在这个规模上,全世界的计算机都不够。而且首先,我怎么
29:48
have to be co-located to be able to ever do this? So, that's where we need other architectures.
必须得co-located才能有可能做到这件事?所以,这就是我们需要其他architectures的地方。
29:54
But I would push back a little bit, right? We have the vision and video language models, right? And
但我要稍微反驳一下,对吧?我们有vision和video language models,对吧?而且
29:59
they use, they basically weren't a mapping. But the resolution is very low. That's the key. Like,
它们用——它们基本上不是一种mapping。但resolution非常低。这是关键。比如,
30:05
for the physical world, the resolution, what do I require? I mentioned like, a thousand by
对于物理世界,resolution,我需要多少?我提到过,比如,一千乘
30:10
thousand by thousand. So, there, you know, if you count that, that's like already in hundreds of
所以我觉得,那时候大家确实会想很多,你知道,一千乘一千。
30:18
millions. Yeah. So, we are not doing that high resolution when we think about images and videos
所以,你知道,如果算一下的话,那已经是好几亿了。
30:24
currently. I have a friend here. And the video is also like autoregressive. So, it's essentially
对。
30:30
only like, you only need to do the next step. Right. Yeah. But you're learning, I mean, like,
所以,我们现在处理图像和视频的时候,并不会用到那么高的 resolution。
30:35
generally you're learning a code book, right? So, you have, you're kind of learning the bias
我这边有个朋友。
30:40
of the latent space or the real world to the latent space. And so, if there is a compression
而且视频也像是 autoregressive 的。
30:49
that you can do from the physical world into the latent space, then, you know, these are regressive
你可以从物理世界做到 latent space 里的那些事,然后,你知道,这些是 regressive
30:55
techniques have been successful. Yeah. But the idea is, you know, a lot of these autoregressive and,
techniques 已经成功了。对。但关键是,你知道,很多这种 autoregressive 和,
31:00
you know, techniques for vision and video models are for mostly like, you know, looking good.
你知道,vision 和 video models 的技术大多是为了看起来好看。对,所以它们并不是用于非常精确的 simulation,而且拥有更高的 resolution 和 details 确实非常关键。所以我们至少需要接收那种高 resolution 的数据,对吧。我们需要能够处理这些数据并对其进行 reasoning。这就是很多 bottleneck 的所在,因为我们不能直接抛弃一切,然后说,哦,我们就在每个维度上取 100 个 grid points 或 50 个 grid points 吧,因为根本没有足够的 detail 来正确建模 fluid dynamics、plasma、材料如何变形这些现象。所有这些都需要 high fidelity。而为此,我们需要 high...
31:08
Right. So, they are like not for very precise simulations and they're, you know, having that
对。所以,它们不太适合做非常精确的 simulations,而且它们,你知道,有那种
31:14
higher resolution and details is really parted. And so, we need to at least take in the data of
更高分辨率和细节真的很重要。所以,我们至少需要接收这些数据,
31:20
that high resolution, right. So, we need to be able to process that and reason over them. And so,
也就是高分辨率的数据,对吧。所以,我们需要能够处理它们并在此基础上推理。然后,
31:26
this is where a lot of the bottleneck is because we, you know, cannot afford to just throw away
这就是很多瓶颈所在,因为我们,你知道,不能承受就这样丢掉
31:32
everything and say, oh, let's just like have 100 grid points in each dimension or 50 grid points
所有东西,然后说,哦,咱们就在每个维度用 100 个 grid points 或者 50 个 grid points 吧,
31:38
because there just isn't enough detail to correctly model phenomena like fluid dynamics,
因为根本没有足够的细节来正确建模像 fluid dynamics 这样的现象。
31:45
plasma, how materials deforms. All of this requires high fidelity. And for that, we need high
plasma,材料如何变形。所有这些都需要 high fidelity。为此,我们需要
31:52
resolution. My understanding you have a thesis that AI needs, you know, to incorporate the physical
resolution. 我的理解是,你有一个论点,AI需要——你知道——把物理世界整合进来,才能在今后scale并保持准确。很多人都有这个论点。
31:59
world into it in order to scale and be accurate going forward. Many people have a thesis.
你有点独特,因为你有几个例子是以这种方式把我们的operators应用到物理世界中的。
32:07
You are somewhat unique and that you have several examples of applying our operators to
然后看起来你正在围绕你的经验构建一个论点。所以,你能和我们分享一些你用过neural operators和其他技术做的非常有趣和令人兴奋的事情吗?
32:14
the physical world in this way. And then it seems like you're constructing a thesis around your
是的,我的意思是,你知道,对我们来说,当我们开始的时候,比如用neural operators处理partial differential equations,但更广泛地说,你甚至不需要假设它们是partial differential equations。
32:20
experience here. So, can you share with us some of the really interesting and exciting looking
丰富的经验。那么,你能不能和我们分享一些非常有趣且看起来令人兴奋的
32:28
things that you've done using neural operators and other techniques?
你使用 neural operators 和其他技术做过的事情?
32:32
Yeah, I mean, you know, for us, when we started with like neural operators for partial differential
是的,我的意思是,你知道,对我们来说,当我们开始时,比如用 neural operators 处理 partial differential
32:38
equations, but also more broadly, you don't even need to assume their partial differential equations.
equations,但更广泛地说,你甚至不需要假设它们是 partial differential equations。
32:43
So, it could be any spatiotemporal or data at multiple scales. So, we, you know, set out looking
所以,它可以是任何 spatiotemporal 或者多尺度的 data。于是,我们,你懂的,开始找一些有趣的例子。其中一个就是 weather modeling,因为 weather data 是 open source 的。数据可以从 ECMWF(European Agency for Global Weather Modeling)拿到,叫做 data file。所以,既然 data 就在那儿,我们就说,好吧,那就去试试吧,对吧?这就是它美妙的地方。只要有 data 可用,就是天大的好消息。但当时很多气象科学家确实提醒过我们。那是 21 年的时候。他们说,不不不,这太难了。你懂的,传统 weather forecasting 已经发展了几十年,都是那种非常谨慎的、自底向上的、基于 physics 的 modeling,对吧?所以,假设,哦,这是 fluid。
32:49
for interesting examples. And one of them was like weather modeling because the weather data is
举些有趣的例子吧。其中一个就是 weather modeling,因为天气数据是
32:55
open source. It's available called data file from the ECMWF, the European Agency for Global Weather
开源的。有一个叫做 data file 的数据,来自 ECMWF,即 European Agency for Global Weather
33:04
Modeling. And so, given that the data was there, we were like, okay, let's just go try it, right?
Modeling。所以,既然数据就在那里,我们就想,好吧,我们直接试一下吧,对吧?
33:08
And that's the beauty of it. Whenever data is available, it's really good news. But a lot of weather
这就是它的美妙之处。
33:14
scientists did caution us back then. This was back in 21. And they said, no, no, no, this is so
只要有 data 可用,就真的是好消息。
33:19
difficult. You know, there have been like decades of like development in traditional weather forecasting
但当时很多气象科学家确实警告过我们。
33:25
and that's very careful bottom up physics based modeling, right? So, assuming, oh, this is the fluid
那是在21年。
33:31
dynamics. Can you go predict the weather, the next day and so on. And so, that's how a lot of the
dynamics。你能去预测天气、第二天的情况等等。所以当时很多人的想法是,
33:38
thinking was that AI is just not going to be able to beat the, you know, decades of work in weather
AI 就是没法击败,你懂的,几十年积累下来的 weather modeling 工作。
33:45
modeling. But to our surprise, we just went ahead, we trained them, we used neural operators to
但让我们惊讶的是,我们还是直接去做了,训练了它们,用 neural operators 来有效捕捉这些现象。
33:52
be able to effectively capture the phenomena. And then now we, to our surprise, we found that
然后现在,让我们惊讶的是,我们发现,
33:59
it's not only, you know, accurate, it's almost as close to the, what the traditional weather
它不仅准确,甚至几乎达到了传统 weather models 能做到的准确度,
34:06
models can do accurately, but also tens of thousands of times faster. So, what would take a big
而且还要快上几万倍。所以以前需要一台大型 supercomputer 才能跑的东西,现在也能跑了。
34:12
super computer to run can now be run. And we only needed a consumer grade like GPU, like, you know,
而且我们只需要消费级 GPU 就够了。你知道,那是个很小的模型,它非常契合。
34:19
it was a small model. It fit very well. It's very fast and it's accurate. And I think that just
它又快又准。我觉得这恰恰——
34:25
changed everybody's thinking. And after that deep mind, Huawei, many others followed us a year later,
改变了所有人的想法。
34:32
released their own models. We were the first to actually open source our weather model for
之后,DeepMind、华为,还有很多其他公司,一年后也跟进发布了他们自己的模型。
34:36
Casnet and do it permissively. So, that's what allowed companies, weather agencies, everybody to
我们是最早真正把我们的 Casnet 天气模型开源的,而且是以 permissive 许可的方式。
34:44
build on us. And so, you know, it's been a really interesting revolution to see that the weather
所以,这才让公司、气象机构、所有人都能在我们的基础上进行开发。
34:50
models are now out there. Weather agencies are adopting them. And it allows us to now have small
所以,你知道,看到天气模型现在到处都是,这真的是一场很有意思的革命。
34:58
weather agencies in the Global South, for instance, have the same kind of fidelity that very
气象机构都在采用它们。
35:05
big agencies were in the past only able to do. Right. So, it's democratizing weather modeling.
而且这让全球南方的小型气象机构,比如说,也能拥有那种过去只有非常大的机构才能达到的精度。
35:11
And so, that's just one example of where there's been very quick rapid progress and a paradigm shift
对。
35:17
in terms of saying that, oh, no, we can have AI as a reliable way to do weather modeling.
就是说,哦,不,我们可以让 AI 成为做 weather modeling 的可靠方式。
35:24
I saw that, you know, the models that you mentioned, you made an insight that nobody else had.
我看到,嗯,你提到的那些模型,你提出了一个别人都没有的洞见。
35:30
And that people were able to devise other mechanisms to kind of follow behind you. But there was
而且人们能够想出其他机制来某种程度上跟在你后面。但这其中需要
35:37
some sort of shift in thinking that was required here. And was it simply we believe that there's
某种思维上的转变。那么是否仅仅是因为我们相信
35:46
enough structure in this data to learn and that people were just doing it wrong and people found
这些数据中有足够的结构可以去学习,而人们之前只是搞错了,然后人们找到了
35:51
other ways to learn the structure, but that your method was very...
其他方式来学习这种结构,但你的方法非常……
35:56
So, let me clarify, right. So, there is, you know, first of all, the very first work was to just
所以,让我澄清一下,对吧。嗯,首先,最早的工作就是只是
36:01
say that, you know, look traditionally this has been done with trying to solve partial differential
说,你知道,传统上这件事是靠尝试求解 partial differential equations 来做的。
36:08
equations each time doing it again and again, whereas AI learns from data, learns patterns and can be
每次都反复求解方程,而AI则是从数据中学习,学习模式,并且可以
36:16
just as accurate but fast. And then the next iterations was to say, you know, how do we make it even
同样准确但速度更快。然后下一步的迭代就是,你知道,我们怎么让它更
36:22
more accurate? And there's the aspect that, you know, there is the short term weather like what is
准确?还有一个方面就是,你知道,短期天气预报,比如
36:28
predictable for the next two weeks. And then there's the long term, you know, going to
未来两周内可预测的部分。然后还有长期,你知道,从
36:33
subsisinal, to ultimately climate modeling. And traditionally, what people did was to have
subseasonal,再到最终的气候建模。传统上,人们的做法是
36:39
different models for these different scenarios. So, there's a different kind of system that works for
针对不同场景使用不同模型。所以,有一种系统适用于
36:45
short term and other system works for long term. But to me, there's only one earth. You know,
短期,另一种系统适用于长期。但对我来说,地球只有一个。你知道,
36:50
if you want a foundation model, if the claim is that, it should be able to do both very short term,
如果你想要一个foundation model,如果其主张是,它应该能同时处理非常短期的,
36:55
as well as very long term together. And that's where in forecast net three, the latest iteration
以及非常长期的预测。而在 forecast net three 里,也就是这个模型的最新迭代中,我们能够两者兼顾。这是因为你还把地球的 spherical geometry 也整合了进来。所以很多其他架构在短期天气预报上已经能拿到相当不错的准确率,但当你把它们跑更长期,比如几个月甚至一年,甚至在那之前,它很快就会爆掉,对吧,因为它假设这个世界是个矩形,但实际并不是。所以把所有这些几何信息整合进 neural operators 里,意味着我们可以让同一个模型忠实地运行更长期,并把它变成气候模型。这就是 Alenea Institute 所
37:01
of the model we're able to do both. And that's because you also, you know, incorporate the spherical
他们说,不不不,这太难了。
37:07
geometry of the earth. And so with a lot of the other architectures that have been getting fairly
你知道,传统天气预报已经发展了几十年,
37:14
good accuracies for the short term weather, when you run them for longer term, when you run them
而且那是非常谨慎的 bottom-up physics-based modeling,对吧?
37:20
for like several months to even year, even before that, it just very quickly blows up, right,
所以,假设,哦,这就是模型的 fluid,我们两者都能做。
37:26
because it assumes the world is a rectangle, which it isn't. And so incorporating all of the
因为它假设世界是一个矩形,但事实并非如此。
37:32
geometry and that information into neural operators means that we can faithfully run the same model
所以把所有的几何信息和那些数据都整合进 neural operators,就意味着我们可以让同一个模型更长期地忠实运行,并把它变成一个气候模型。
37:40
also longer term and make this into a climate model. This is where the Alenea Institute has
这就是 Alenea Institute 的做法——比如说输入风况、湿度等等,以 autoregressive 的方式预测接下来会发生什么,每六小时一次。
37:46
now built climate models based on our neural operator architecture. And that's the only one that
现在我们构建的 climate model 是基于我们的 neural operator architecture 的。
37:52
works as an AI emulator, right? None of the other architectures work for climate because
而且那是唯一一个能作为 AI emulator 工作的 architecture,对吧?
38:00
climate requires us to assume the world is a globe. And that if you are repeatedly rolling out,
其他 architecture 都不适用于 climate,因为 climate 要求我们假设世界是一个球体。
38:07
you kind of keep that information. Whereas if it's like a narrow surrogate, that's what I consider
而且如果你反复 rollout,你就能保留那个信息。
38:13
a weather model. You just narrowly look at a few metrics. Many different architectures should do
相比之下,如果它只是一个窄的 surrogate,那就是我所说的 weather model。
38:19
the job, right? But if you are asking one architecture do a range of different tasks like a foundation
你只是狭隘地盯着几个 metrics。
38:26
model, that's where incorporating the geometry of the earth, which is that it's a sphere,
很多不同的 architecture 都能胜任,对吧?
38:32
and using neural operators an efficient way to do that enables us to accomplish that.
但如果你想让一个 architecture 像 foundation model 一样执行一系列不同任务,那就需要结合地球的 geometry(也就是它是一个 sphere),并通过 neural operator 高效地实现这一点,这样我们才能做到。
38:39
Yeah, so forecast net, you said trained on 50,000 data points. Can we just talk about what does
对,所以forecast net,你说是在50,000个data points上training的。我们能不能聊聊它
38:44
this look like? What does the data input look like? What are you actually trying to predict from here?
看起来是什么样?data input长什么样?你实际上想从这里predict什么?
38:50
And then what is the large scale? You said you're going from weather to climate. What does it look
然后什么是large scale?你说你要从weather到climate。那是什么样
38:56
like to do that generalization? Because if you have 50,000 data points, these are some high resolution
去做那种generalization?因为如果你有50,000个data points,这些是某种high resolution
39:03
in North America, then they might even depend on the local geography. Like if you are always
在North America,那它们可能还取决于当地地理。就像如果你总是
39:10
modeling Kansas, this is going to transfer to, let's say, the Swiss Alps or something,
在modeling Kansas,那这会transfer到,比如说,Swiss Alps或者什么地方,
39:16
and then does that transfer to, you know, to Hanoi's?
然后那能不能transfer到,你知道,到Hanoi's?
39:20
First of all, to clarify, we are training it on the global weather model. So we have all the
首先,澄清一下,我们是在global weather model上training的。所以我们有所有的
39:26
information around the earth, and then we are asking it to predict, like given the current weather,
关于地球的信息,然后我们让它去预测,比如给定当前的天气,比如说风况、湿度等等,以 autoregressive 的方式会发生什么,而且是每六个小时一次。所以接下来六个小时会发生什么,然后你 rollout,然后训练模型去预测。所以呢,你知道,你有可能让同一个模型永远预测下去,对吧?但可预测的窗口就像天气一样。如果你想超越,如果要去做我们所谓的 ensembles,意思是你有一个 probabilistic estimate,关于几个月、两年后会发生什么,那样你就得到一个 climate model。好,所以你有,实际上你有一组这些 local predictors,然后你用
39:33
like, say, wind conditions, humidity and so on, what happens in an autoregressive way,
所以接下来六小时会发生什么,以此类推,然后你不断 rollout,训练模型去预测。
39:40
and it's every six hours. So what happens in the next six hours and so on, and you roll out,
而且,你知道,你有可能让同一个模型永远预测下去,对吧?
39:44
and you train the model to predict. And so, you know, you can have potentially the same model
但可预测性窗口就像天气一样。
39:50
predict forever, right? But the predictability window is like the weather. And if you want to go
如果你想更进一步,去做我们所说的 ensembles,也就是说你有一个 probabilistic estimate,关于……
39:56
beyond, if to do what we call ensembles, meaning you have like a probabilistic estimate of what
另外,如果做我们所说的 ensembles,意思就是你有一个 probabilistic estimate,关于什么
40:03
happens in several months, two years, and that's how you get a climate model.
发生在几个月、两年内,然后你就得到了一个climate model。
40:09
Okay. So you have, you actually have an ensemble of these local predictors, and then you use
好的。所以你实际上有这些local predictors的ensemble,然后你在ensemble上使用某种统计方法得到climate prediction。
40:15
some sort of statistics on the ensemble to get a climate prediction.
是的。所以你有点像有多个rollouts。本质上,你有多个rollouts的trajectories,然后你得到——所以麻烦在这里实际上很关键。
40:20
Yeah. So you kind of like have several rollouts. Essentially, you have several trajectories
是的。这就是传统climate modeling最大的bottleneck,因为做起来太贵了。即使是一次single run,你也必须做很长的trajectories,还要非常高的resolution。而且,你知道,这就是为什么我们没有很多非常高的resolution。
40:24
of rollouts, and then you get, and so you get troubles to actually pretty key here.
关于 rollouts,然后你会得到——所以你会遇到问题,但这里其实很关键。
40:29
Yes. And that's why that's the biggest bottleneck with traditional climate modeling,
是的。这就是为什么那是传统 climate modeling 最大的瓶颈,
40:34
that it's so expensive to do. Even one single run, you have to do long trajectories,
就是因为它太昂贵了。哪怕只跑一次,你也得做很长的 trajectories,
40:38
where a very high resolution. And, you know, that's why we don't have a lot of very high resolution
需要非常高的分辨率。而且,你知道,这就是为什么我们没有很多 very high resolution 的数据。
40:46
ability to do climate change predictions, for instance. So how do you validate climate?
比如说,做 climate change predictions 的能力。那你怎么 validate climate 呢?
40:52
So you do lots of rollouts and you look at sort of how things evolve kind of in aggregate.
所以你会做很多的 rollouts,然后看事情在 aggregate 层面是怎么演变的。
40:57
I mean, Nesey said, butterfly effect, locally, I think we can believe even a perfect climate model,
我是说,Nesey说过,butterfly effect,局部来看,我觉得哪怕是完美的climate model,
41:02
or perfect weather model would only give you maybe two weeks before it sort of,
或者完美的weather model,大概也就只能给你两周的时间,然后就会有点,
41:07
it becomes non-computable. So you do lots of rollouts, and there's some chaos, you average all
变得non-computable。所以你要做大量的rollouts,里面会有一些chaos,然后把所有这些
41:15
these things. How do you actually validate that this works over long enough times? Yeah, timescales?
东西求平均。那你实际上怎么验证这在足够长的时间尺度上有效?对,timescales?
41:21
Yeah, and it's a tricky question, right? So for instance, you have to kind of ensure that you
对,这是个棘手的问题,对吧?比如说,你某种程度上得确保你
41:27
satisfy all of the physical constraints. And if you just do a standard rollout, that is, you know,
满足所有的physical constraints。如果你只是做一个标准的rollout,那,你知道,
41:34
likely not going to happen. And so some of the ongoing research we're doing is how do you kind of
很可能不会满足。所以我们正在进行的一些研究就是,你怎么某种程度上去
41:40
enforce the right physical constraints as you do the rollouts? Like you don't want to, you know,
你在做rollouts的时候去强制正确的physical constraints?就像你不会想,你知道,
41:45
completely wash out the fine details, because then it's not accurate. But on the other hand,
把精细细节完全抹掉,因为那样就不准确了。但另一方面,
41:50
if you keep them, you may be physically, they're invalid. So this is still an open problem.
如果保留它们,结果可能在物理上就不成立。所以这仍然是一个未解问题。
41:55
And that's what makes it difficult, that you want AI to be fast and you want to be able to do
这就是难点所在:你希望AI速度快,又希望能进行
42:02
these long climate simulations, and at the same time be able to have full confidence in them.
这些长期的气候模拟,同时还要对它们有充分的信心。
42:09
But these are things we are working on now. I don't know if this falls into weather or climate,
但这些正是我们现在在努力的事情。我不知道这算天气还是气候的范畴,
42:13
but we recently just had one of the most extreme heat waves in the history of modern climate
不过最近我们刚刚经历了现代气候
42:20
data in the just right here in the kind of Southwest United States. I'm wondering,
数据记录史上最极端的熱浪之一,就在美国西南部这一带。我在想,
42:26
where did you, I don't know if you were involved in this regularly, but do you know if you or anyone
你当时——我不知道你是否经常参与这方面的工作——但你知道你或其他人是否
42:33
actually modeled that or predicted that correctly? Yeah. So we have, I know, I don't have information
实际上有没有正确地 modeled 或 predicted 出来?是的。所以我们有,我知道,我没有信息。
42:37
on this specific one, but we've tested in our latest forecast in a three model extreme weather
关于这个具体的,但我们在最新的 forecast 中,在 three model extreme weather 事件里测试了。
42:44
events of all kinds, right? And that's the key, like, you know, that's where you need probabilistic
各种事件,对吧?这就是关键,比如说,你知道,那就是你需要 probabilistic。
42:50
answers. So having just one deterministic output and saying that this is the weather is not
答案。所以只有一个 deterministic output,然后说这就是天气,并不。
42:55
enough when we are looking at extreme events. So we need careful probabilistic calibration,
足够当我们看 extreme events 的时候。所以我们需要仔细的 probabilistic calibration,
43:00
and we show that we are able to capture those well. And I think that was the surprise even in our
然后我们展示了我们能够很好地 capture 这些。而且我认为那就是惊讶之处,甚至在我们的。
43:06
very first attempt that we visualized certain hurricanes and storms. And it was able to do well,
这是我们第一次尝试把某些飓风和风暴可视化,而且它做得很好,
43:14
which is very surprising because you would think that rare events are not something AI would do
这非常令人惊讶,因为你会觉得罕见事件不是AI能做到的,
43:19
well, right? It would do a well on typical events. But I think this is where more broadly,
嗯,对吧?
43:24
the lesson is the physical world may be more forgiving because, you know, where there are extreme
在典型事件上它会做得很好。
43:31
events like hurricanes that have very specific physical signature, right? So it's like extreme,
但我觉得这里更广泛的教训是,物理世界可能更宽容,因为你知道,像飓风这样的极端事件有非常特定的physical signature,对吧?
43:37
but in a very specific way. So maybe you don't need as many samples because the physical world has
所以它像是极端的,但以一种非常特定的方式。
43:42
a lot of structure. And that's something we see this again, and again, that there is a lot of
所以也许你不需要那么多samples,因为物理世界有很多结构。
43:47
structure in so many other examples, you know, talk about plasma and fusion reactor. You know,
而这正是我们一次又一次看到的,在很多其他例子里都有很多结构,你知道,比如plasma和fusion reactor。
43:53
we barely have a few thousand samples, but we are able to accurately predict
你知道,我们只有几千个samples,但我们能够非常准确地预测像disruption这样的事件。
43:59
events like disruption very well. And we are able to do that a million times faster than what
而且我们能以比……快一百万倍的速度做到
44:05
traditional simulations were able to do. To me, yes, all these sound very surprising, but it's
传统模拟能够做到的。
44:11
because I think the nature helps us a lot. It has a lot of latent space structure.
对我来说,是的,所有这些听起来都很令人惊讶,但这是因为我觉得大自然帮了我们很多。
44:20
That I don't think traditional numerical methods are able to uncover because they are
它有大量的 latent space 结构。
44:26
focusing more on correctness. That in any scenario, you should be able to solve these equations.
我不认为传统 numerical methods 能够揭示这些,因为它们更注重正确性。
44:31
On the other hand, with AI, it's learning from data, it's uncovering the structure, it's uncovering
也就是说,在任何场景下,你都应该能求解这些方程。
44:37
how easy or kind of tractable these problems are, and that's what we are seeing in many cases.
另一方面,AI 从数据中学习,它在揭示结构,揭示这些问题到底有多容易、或者说有多 tractable,而这正是我们在很多情况下看到的。
44:44
Yeah, I've had a similar analogy. So from my own domain is probably closer to computational biology,
是的,我也有过类似的类比。
44:50
but you know, alpha fold is the obvious, like really exciting development in the community.
所以从我的领域来说,可能更接近 computational biology,但你知道,AlphaFold 显然是社区里一个非常令人兴奋的发展。
44:56
So the solving protein structure prediction, and of course, all the caveats of what was
所以解决 protein structure prediction,当然,关于到底解决了什么,所有那些需要注意的地方,我想我们之前在跟 Boltz 团队的那一期里讨论过,如果听众想了解更多,我鼓励大家去听那一期。但我觉得关于 protein structure 的关键点之一是,它确实受到物理的约束,所以从某种意义上说,这是生物学领域少数几个重大胜利之一,而这个领域其他方面则非常复杂,而且总的来说,我们一直很难取得大量成功。看起来,那些由 differential equations 解决,或者说能被 differential equations 很好建模的问题,在把这类技术作为一种通用形式整合进来时,有更大的空间。我觉得这里有一个问题。
45:01
actually solved, or I think we discussed this in a previous episode with the Boltz team,
实际上已经解决了,或者我记得我们在之前一期节目里和Boltz团队讨论过这个问题,
45:05
I encourage listeners to listen to that if they want more, but one of the, I think key points
如果大家想了解更多,我鼓励听众去听那期节目。但我觉得其中一个关键点是,
45:10
about protein structure is that it really is constrained by physics, and that's why it was in
关于 protein structure,它确实受到物理规律的约束,这也是为什么它在
45:15
some sense one of the few big wins in the field of biology, which is otherwise very complex, and
某种意义上成为生物学领域为数不多的大突破之一,而生物学本身非常复杂,并且
45:22
we've had trouble generally speaking, making a lot of success, and it seems like problems which
总的来说我们一直很难取得很多成功。而且看起来那些
45:28
are solved by differential equations, or say modeled well by differential equations have a lot more
能用differential equations解决,或者说能被differential equations很好建模的问题,就有更多
45:34
room for also integrating these techniques as a general form. I think there's a question there.
也有空间把这些技术整合成一种通用形式。我觉得这里有个问题。
45:43
So getting back to my question, we have climate or weather and climate, we have plasma, we have
所以回到我的问题,我们有气候或者天气和气候,我们有等离子体,我们有
45:52
biology, and I know that you prepared for us a few visualizations. Can you just share with us
生物学,而且我知道你为我们准备了一些visualizations。你能跟我们分享一下
46:00
what does that look like? So the visualizations, and then what is the thread that runs through here,
它们是什么样的吗?就是这些visualizations,然后贯穿其中的线索是什么,
46:07
and I think maybe the listeners will already have a hint about that, but I'd be really excited to see that.
我觉得也许听众们已经有一些猜测了,但我真的很期待看到。
46:14
Yeah, I can certainly share some of them. I mean, this one is just kind of showing that we have
是的,我当然可以分享一些。我的意思是,这个visualization主要是展示我们有
46:23
world of different scales, right? And these are examples of phenomena happening at different scales
不同尺度的世界,对吧?这些就是发生在不同尺度上的现象的例子,
46:29
from atomic to protein to even planetary scales, like the weather we talked about, and we need
从原子到蛋白质,甚至到行星尺度,比如我们之前谈到的天气,我们需要
46:37
to capture all of that. That's what neural operators are designed to do, and to feed in data
去捕捉所有这些。这正是neural operators的设计目的,以及输入数据。
46:45
of these different scales. And that's really like the aspect that makes a lot of the physical world
这些不同的scales。而正是这种对fine scale的需求,让很多物理世界的问题变得困难。
46:52
problems hard, that need for fine scale. We talked about how a lot of traditional computer vision
我们之前聊过,很多传统的computer vision video models只是为了视觉效果好看而设计的,这需要
47:00
video models are just designed to make things look visually good, and that requires
足够低的resolution,tractable,auto regressive,只要能work out就行,视频也足够短,但很多physical simulation并不是这样,
47:06
low enough resolution, it's tractable, auto regressive, it's enough that it works out,
industrial scale、high fidelity的工作。你真的需要高resolution。大气就是一个
47:11
it's short enough videos, but that's not how a lot of the physical simulation for
例子,比如说取决于你观察的resolution,能捕捉到不同的phenomena,如果没有高resolution,就会错过这些。
47:18
industrial scale, high fidelity work. You really need high resolution. That atmosphere is one
而现在的问题是,有了AI之后,
47:26
example, like depending on the resolution you observe, different phenomena can be captured,
比如,根据你观察的resolution不同,能捕捉到的phenomena也不同,
47:32
so you just miss that out if you don't have that high resolution. And now the question is with AI,
所以如果没有那么高的resolution,你就会错过这些。而现在的关键是,对于AI,
47:38
can we do this much faster than what we could with traditional simulation? And this one with
我们做这个能比传统 simulation 快得多吗?而这个用
47:43
neural operators is showing that if you used a standard neural network and you had a fixed number
neural operators 表明,如果你用一个标准的 neural network,而且你有固定数量的
47:51
of pixels like you're seeing on this side, and you zoom in, it gets blurry, that's the end of it,
pixels,就像你在这边看到的,然后你放大,它会变模糊,就这样了,
47:57
there's nothing beyond those fixed resolution that you can capture. But the idea is with neural
超过那个 fixed resolution,你就捕捉不到任何东西了。但关键在于用 neural
48:04
operators, because it's a function space representation, meaning you can keep zooming in, you can
operators,因为它是一种 function space representation,意思是你可以不断放大,你可以
48:10
add the relevant details, either by giving it data at higher resolution, or physical constraints
添加相关细节,要么给它更高 resolution 的数据,要么用 physical constraints
48:16
at higher resolution, you can kind of bring that multi-scale phenomena together.
在更高 resolution 下,你可以某种程度上把 multi-scale phenomena 整合到一起。
48:22
So where are the physical constraints? I mean, I see them physical constraints here are like
那么 physical constraints 在哪里?我是说,我看到这些 physical constraints 在这里就像
48:26
local simulating fluid equations or some sort of dynamic. Yeah, it could be, right? So it could be
局部模拟流体方程或者某种动态。对,有可能是吧?所以它可能是
48:33
of any nature. The idea is now you can add like conservation laws, for instance in an incompressible
任何性质的都行。现在的想法是,你可以加上像 conservation laws 这样的东西,比如在一个 incompressible
48:40
fluid, you can add like material deformation, like how things stretch. So all it can be a full
fluid 中,你可以加上 material deformation,比如物体怎么拉伸。所以它可以是完整的
48:47
partial differential equation. So that's also an interesting question we've been researching,
partial differential equation。所以这也是我们一直在研究的一个有意思的问题,
48:52
how is the curriculum of different physics? Some physics may be very hard to impose or
不同 physics 的 curriculum 是怎样的?有些 physics 可能很难施加,或者
49:01
add as a loss function, others may be easier, so you also need to kind of think about what to impose.
作为 loss function 加进去,其他的可能更容易一些,所以你也需要想一想到底该施加什么。
49:09
And the intuition is here is that when I'm adding a physical constraint, I'm adding it to the
这里的直觉是,当我加一个 physical constraint 的时候,我是把它加到
49:13
loss function, is that more or less what's happening? Yeah, because that's what is tractable,
loss function 里,大致就是这样吧?对,因为那是 tractable 的。
49:17
you know, making it a hard constraint is not tractable, whereas adding it as a loss function. And
你知道,把它做成 hard constraint 是不 tractable 的,但作为 loss function 加进去就可以了。
49:23
of course, there's still the balancing of that loss with the data we have, so we have to, you know,
当然,仍然要平衡这个 loss 和我们手上的 data,所以必须用合适的方式来做。
49:29
do that in the appropriate way. Yeah, so as I mentioned, this is the example of the weather
对,就像我刚才说的,这是 weather model 的例子,我们在这里展示的是怎么捕捉 atmospheric rivers,也就是我们在加州看到的这种现象,你知道,正好是在 La Estonia。
49:35
model, where here we are showing how we are able to capture like atmospheric rivers, which is
所以我觉得这周晚些时候会有一个预计出现。
49:43
the phenomena we see here in California, you know, exactly in La Estonia. So I think we have one
那怎么做呢?
49:49
expected later this week. So how to do that? So the idea of like, you know, why I showed this is
所以我展示这个的原因是,这是一种 global phenomena,对吧?
49:59
this kind of global phenomena, right? These are like thousands of miles wide. So you really need
这些大概有数千英里宽。
50:05
non-local models that capture these very large span phenomena and do that accurately, and that's
所以你确实需要 non-local models 来准确捕捉这些跨度非常大的 phenomena,而这就是...
50:14
what our neural operators are able to do. So, so the training data for this is you were talking
我们的 neural operators 能做什么。
50:19
about this a bit before, but I'm still like wondering what is this, this is weather satellites,
所以,所以这个训练数据就是——你之前稍微提到过,但我还是好奇,这到底是什么?是气象卫星吗?
50:23
or is there ground-based data? Is it some hybrid of the two? It's it's kind of a combination of
还是说地面数据?还是两者的 hybrid?
50:30
different sources. So it's what we call re-announced this data. So this is historical weather data
它是,它是不同来源的一种组合。
50:35
that is in a way re-analyzed, meaning that the satellite data is combined with essentially what
所以这就是我们所说的 reanalysis 数据。
50:43
the physics solvers tell you together assimilate it. And so this is a way made available by the weather
所以这是历史天气数据,某种程度上是重新分析过的,也就是把卫星数据和 physics solvers 告诉你的信息结合起来,一起做 data assimilation。
50:50
agency. So we can train on them. So you're saying that they take a low resolution data set,
这些数据由气象机构提供。
50:57
which is compiling all of the world's data set, you know, all of our meteorological data we have
所以我们可以在上面训练。
51:01
across the world. And then they do short time simulations using physics-based, you know,
在全世界范围内。
51:08
classical techniques to to fill in the details. And you can do that over short time spans,
然后他们用 physics-based 的,你懂的,classical techniques 做短时间的 simulation,来填补细节。
51:16
but as you go farther, it breaks down very quickly. So you're amortizing that across all the
你可以把它用在短时间跨度上,但时间一拉长,它很快就崩了。
51:22
everyone would have to do that. And so somebody does it, and then you are able to take advantage.
所以你在把它分摊到所有地方,每个人都得这么做。
51:26
Yeah, I mean, this is data, right? That's already prepared. But that year's already this data
所以有人做了,然后你就能利用起来。
51:32
simulation with physics kind of makes our model physics informed implicitly. So it's able to kind
是啊,我是说,这就是 data,对吧?
51:38
of, you know, keep that information. And that's why maybe that's one reason, maybe it does well
那是已经准备好的。
51:45
on even extreme weather events. Yeah, so this is just showing that our model is available in
但那一年的 data 已经,这种带 physics 的 simulation 隐性地让我们的 model 具备 physics-informed 的能力。
51:52
ECMWF, which is the weather agency, like, you know, the European weather agency. And so this was
ECMWF,就是那个气象机构,你知道,欧洲的气象机构。所以这个
51:59
launched like more than two years ago, but you know, I think fall 2023. So, you know, I think ECMWF
是在两年多前推出的,不过你知道,我觉得是2023年秋天。所以,我觉得ECMWF
52:08
making these AI based weather models available to the public, to me was a very big step because
把这些基于AI的天气预报模型开放给公众,对我来说是很大的一步,因为
52:15
that's where, you know, everybody could see what's happening. There were several hurricanes,
在那里,你知道,大家都能看到发生了什么。当时有几个飓风,
52:21
like, for instance, there was hurricane Lee. And that's where the public could see what are these
比如说,有飓风Lee。而且那时候公众就能看到这些
52:26
weather models doing? For instance, our forecast net was able to correctly predict that the
天气预报模型在做什么?比如,我们的forecast net能够准确预测到
52:32
hurricane making the landfall several days earlier compared to the standard weather forecasting models.
飓风登陆的时间,比标准的天气预报模型要提前好几天。
52:41
And so the idea that these models could be very good for extreme weather events and do early
所以这个想法就是,这些模型可能对极端天气事件非常有用,而且能提前
52:48
prediction, you know, both for human lives, for economic costs is a very big deal. And so that's
prediction,你知道,对人的生命、经济成本来说都是很大的事。所以那就是
52:54
when the public kind of got a lot more, I think, buy in from weather scientists because of how well
当公众,我想,对气象科学家更加买账的时候,因为它在这些事件中表现得太好了。
53:01
it was doing in these events. And this is what I was talking about in an ensemble prediction,
而这就是我之前说的ensemble prediction——
53:07
both for extreme weather, or if you're thinking about climate, it's not just about looking at
不管是extreme weather,还是你考虑climate,都不是只看一条trajectory,对吧?
53:13
one trajectory, right? Because, you know, unless you're somebody with a sharpie somehow saying,
因为,你知道,除非你是那种拿着sharpie的人,莫名其妙地说一句,
53:22
fun intended. But the word you really want is the probabilistic prediction, meaning, you know,
fun intended。但你真正想要的是probabilistic prediction,意思是,你知道,
53:29
I'm going to try different adding noise levels to my initial condition. When the hurricane
我要尝试在我的initial condition上加不同的noise levels。当hurricane
53:36
is forming in the Caribbean, I'm going to add some noise because anyway, it's noisy. I don't know
在Caribbean形成的时候,我会加一些noise,因为反正它本来就是noisy的。我不知道
53:42
truly what the measurement there is. And then I'm going to look at what happens to the possible
那里的measurement到底是什么。
53:49
hurricane tracks. And then I can come up with the probability of landfall in different regions.
然后我要看看可能的hurricane tracks会怎么样。
53:54
And that's how I can do risk assessment. And so this is where it gets even more expensive for
然后我可以得出不同地区的landfall概率。
54:00
traditional weather models, because if to run all of these ensembles and now AI weather models being
这就是我做risk assessment的方式。
54:06
so fast, tens of thousands of times faster means we can now do very large ensembles. And this
所以对于传统weather models来说,这就更贵了,因为要运行所有这些ensembles,而现在的AI weather models快得非常多,快数万倍,意味着我们现在可以做非常大的ensembles。
54:12
is a very big improvement in terms of what we can do for risk assessment. Have you gone through and
这在risk assessment方面是一个非常大的改进。
54:19
done, let's say, looked over the historical hurricane maps and then tried to do ensemble
你有没有去,比如,回顾历史的hurricane maps,然后试着做ensemble predictions,并且calibrate你的predictions有多像core?
54:25
predictions and calibrated how often your predictions are like a core. Yeah, yeah. So in forecast
是的,是的。
54:30
three paper that are, you know, we have metrics of like extreme weather events and ensemble
有三篇论文,你知道,我们有像extreme weather events和ensemble predictions这样的metrics。
54:37
predictions. And in fact, we've trained the model to do go down ensemble prediction. And so this
事实上,我们训练了模型去做ensemble prediction。
54:42
is where the calibration matters for these kind of events. What was the sort of key insights
所以这就是为什么calibration对这些类型的事件很重要。
54:49
or developments in forecast net three in versus two versus the first version? Yeah, so the first
那么forecast net三相对于二和第一个版本,有哪些关键insights或发展?
54:53
version was kind of the, you know, the using like the Fourier neural operators, but we didn't incorporate
是的,第一个版本算是,你知道,使用了Fourier neural operators,但我们没有加入spherical geometry,对吧?
55:00
the spherical geometry, right? In this next version, we said, I think, you know, it's important that
在这个新版本中,我们认为,你知道,世界是球体很重要,因为首先,否则就会有扭曲。
55:05
the world is a sphere because first of all, otherwise it's distorted. So you kind of are not predicting
所以你并不是在那种情况下进行预测,但如果它不是球形的,你会怎么做?像Mercator projection那样。
55:12
it in that question, but if it wasn't spherical, what did you do like a marketer projection?
但在这个问题里,如果它不是spherical,你会怎么做?像Mercator projection那样?
55:17
Yeah, the standard, like kind of the, you know, like as you all and all the other weather models
嗯,标准的,就是那种,你懂的,就像你们和所有其他weather models一样。
55:22
do the same, right? So they just kind of have the standard projection and then, you know, predict the
都是这么做的,对吧?所以它们就是有个标准的projection,然后预测天气。
55:28
weather. And which is okay for short term prediction, but when we in our goal was to have the same
这短期预测还行,但我们的目标是让同一个模型也能做更长远的预测。
55:34
model also do longer term. And that's when incorporating the spherical geometry added this
模型也能做更长期。那时候加入spherical geometry就增加了这种
55:39
additional stability we could do longer rollouts. And then in forecast net three, the idea was
额外的稳定性,让我们能做更长的rollouts。然后在forecast net three中,想法是
55:46
it's not just about deterministic prediction. We want to get ensemble predictions, right? So we have
不只是关于deterministic prediction。我们想要得到ensemble predictions,对吧?所以我们必须
55:52
to train them based on this objective that we get the probabilistic predictions correct as well.
基于这个目标来训练它们,让probabilistic predictions也准确。
55:57
How long are you predicting out? And how many rollouts are you doing?
你们预测多长时间?做多少个rollouts?
56:01
Yeah, so the, you know, rollout is how long you predict, right? So each step is six hours and then
对,所以这个,你知道,rollout 就是你要预测多长,对吧?每一步是六个小时,然后
56:08
you predict for how we have a long you want. You know, you just have to roll out.
你预测你想要多长的时间。你知道,你只需要 roll out。
56:13
Sorry, how many examples in the ensemble TF?
抱歉,ensemble TF 里有多少个 examples?
56:18
So and again, that's our choice. We can have like ensembles of different levels. So we,
所以再说一次,那是我们的选择。我们可以有不同级别的 ensembles。所以我们
56:24
I think it's like a few tens or something like it's what we are currently, you know,
我觉得大概是几十个或者差不多的,就是我们目前,你知道,
56:30
shown, but you can do much larger too. And that's adequate to get out how far my intuition is
展示的,但你可以做得更大。而这足以得出我的直觉是
56:36
the longer you want to predict the more. So not necessarily looks really like about, again,
你想预测得越长,就需要越多。所以不一定是那样,真的更像是,再次,
56:41
calibrating the ensembles and ensuring that they have the right spread rather than
calibrating ensembles 并确保它们有正确的 spread,而不是
56:47
you know, so, okay, so you have tens of these models or examples in your ensemble and that
你知道,所以,好吧,你的 ensemble 里有几十个这样的模型或例子,然后那个
56:55
even with a very, very long rollout, that's adequate. So again, like, you know, there's as I said,
即使 rollout 非常非常长,那也是足够的。然后再说,就像,你知道,正如我所说,
57:01
a lot of still outstanding questions to do very, very long rollouts, right? Because you do need
对于非常非常长的 rollout,还有很多悬而未决的问题,对吧?因为你确实需要
57:07
to incorporate the physical constraints in a way to ensure that that's something that we are
以某种方式整合 physical constraints,以确保那正是我们正在
57:12
actively researching now. But these models that we have are able to do the longest rollouts
积极研究的事情。但我们现有的这些模型能够做最长的 rollout
57:19
compared to any of the other weather models that completely ignore spherical assumption of a
与任何其他完全忽略 spherical assumption 以及其他一系列东西的 weather models 相比。
57:25
range of other things. So when you say incorporate physical laws for climate over long times,
所以当你说要在长时间尺度上为 climate 引入 physical laws 时,
57:32
I mean, what does that look like? Because there's a lot of the local conservation, which may be just
我的意思是,那会是什么样子?因为有很多 local conservation,可能只是
57:38
broken if you take an ensemble average, even though any given snapshot respects that.
如果你取 ensemble average 就会破坏掉,即使任何给定的 snapshot 都满足那个。
57:43
No, the idea is to make sure you look at each ensemble member and it's respecting the physical.
不,关键是确保你去看每一个 ensemble member,并且它本身是符合 physical 的。
57:48
Okay, okay. You're not assuming the ensemble, you're not deriving a coarse-grained equivalent of
好吧好吧。你不是在假设 ensemble,也不是在推导一个 probability 或什么的 coarse-grained 等价物。不,因为那样你会失去那个 resolution。这让我有点困惑。所以它不是 average,那你是怎么组合 ensemble 的?
57:55
a probability or something. No, because then you would lose that resolution. That's a little
不,你是在做 average,但你是分别预测每一个。哦,你是分别预测每一个的。
58:02
confusing to me. So it's not an average, how are you combining the ensemble?
好吧好吧。所以每一个都独立地满足这些 constraints,但 ensemble 不满足,对。所以那就是你确保 physical biology 的方式。
58:07
No, you are doing the average, but you're predicting each one. Oh, you're predicting each one separately.
不,你是在做average,但你是对每一个进行预测。哦,你是分别预测每一个。
58:12
Okay, okay. So each one is independently satisfies these constraints, but the ensemble
好的,好的。所以每一个都independently满足这些constraints,但ensemble并不满足,对。这样就确保了physical biology。
58:17
does not, yeah. And so that's how you ensure physical biology.
所以当你处理sphere上的问题时,你基本上会使用spherical harmonics或者某种类似的东西。
58:22
So when you go on a sphere, you operate and use basically spherical harmonics or some sort of,
所以当你在一个球体上操作时,你基本上用的是 spherical harmonics 或者某种
58:27
okay, yeah, you have a spherical basis for it, which is actually very natural with for a
好,对,你有一个 spherical basis,这实际上非常自然,对于一个
58:33
probably much harder if you're doing other. Yeah, exactly. So that's where the, you know,
可能更难如果你做其他的。对,没错。所以这就是,你知道,
58:39
like the 40-year safe session can incorporate these geometry as well. And so I think be very
像 40-year safe session 也可以融入这些 geometry。所以我认为非常
58:45
faithful to, you know, what the globe is. Yeah, are you, it's much more natural than a,
忠实于,你知道,地球本身。对,是吧,它比一个
58:51
didn't like a marketer projection or whatever other. Yeah, which is, you know, is like Greenland
不像 marketer projection 之类的。对,那就是,你知道,就像 Greenland
58:56
becomes huge. So that is a different story. But, but yeah, but I think that this is where I think
变得巨大。所以那是另一回事。但是,但是对,但是我认为这就是
59:03
the aspect of, you know, more broadly incorporating more of geometry and information about the domain
这个方面,你知道,更广泛地融入更多的 geometry 和关于 domain 的信息
59:10
becomes a lot more important for the physical world, right? So this is me again, emphasizing that
对于物理世界来说,这就变得更加重要了,对吧?所以我再次强调,
59:16
we need to incorporate more of the structures, because one is the data is limited. And the other is
我们需要融入更多的结构,一方面是因为数据有限,另一方面是
59:22
a lot of what we are asking is extrapolation, you know, to go beyond then what that's trained on.
我们要求的很多东西都是 extrapolation,你知道,要超越训练所覆盖的内容。
59:28
You know, we're just training it to predict the next six hours and maybe do a little bit of
你知道,我们只是训练它预测接下来的六个小时,也许做一点
59:32
multi-step fine tuning for autoregressive rollouts, right? So we're not training it to do very long
multi-step fine-tuning 用于 autoregressive rollouts,对吧?所以我们并没有训练它做非常长期的
59:38
like a climate simulation, because that's just too expensive. But we hope magically it works
比如气候模拟,因为那太昂贵了。但我们希望它奇迹般地有效
59:44
well. And it cannot, if you just say I'm just going to put a standard transformer or whatever else
然而,如果你只是说,我就放一个标准的 transformer 或者其他什么,
59:52
there, and it won't work out. So we added more of the domain constraints like spherical geometry,
那是不行的。所以我们就加入了更多的 domain constraints,比如 spherical geometry,
59:59
we added maybe more of the physics in certain ways. And that's where it becomes more interesting
我们在某些方面可能加入了更多的物理。
60:05
algorithmically as well. You know, there's more involved design here. So the time scale that you train
这也是算法上更有意思的地方。
60:12
on, how long is that? To predict for the next six hours. Oh, so only six hours. Yeah. And a little
你知道,这里的设计更复杂。
60:18
bit of multi-step fine tuning like I said, yeah, I said, okay, understood. So which is very
那你训练的时间尺度是多长?
60:24
surprising. Yeah, that is very surprising. I would have expected it was weeks or months. Yeah, no,
预测接下来六个小时。
60:28
I've done it kind of just works well, even for like now we are showing for several months that it's
哦,所以只有六个小时。
60:34
able to do that. The number of steps is on hundreds or yeah. And your foyer basis, it's many times
对。
60:46
the sort of base harmonic. I mean, this is like that's in space, right? So we're talking rollout
然后还有一点multi-step fine-tuning,就像我说的。
60:51
as autoregressive at time. Nobody likes it in time. Maybe I misunderstood here. Because it's a
它在时间上是 autoregressive。没人喜欢它在时间上。也许我在这里理解错了。因为这是一个
61:00
in 4.8 domain. No, no, in time, it's not. That's what I'm saying. It's autoregressive. Oh,
在 spatial domain 里。不,不,在时间上不是。我就是这个意思。它是 autoregressive。哦,
61:05
oh, I understand. Okay. So it's space. Yeah, interesting. Yeah. What's the angular resolution?
哦,我明白了。好吧,所以是空间上的。对,有意思。对。angular resolution 是多少?
61:11
In this scenario, in other cases, we also have in time is also represented in the 4-year domain.
在这个场景中,在其他情况下,我们也把时间维度表示在 spatial domain 里。
61:17
And that's a question as well. Can we do that? But in this example, yes, autoregressive. Got it.
这也是一个问题。我们可以那样做吗?但在这个例子里,是的,autoregressive。明白了。
61:24
What's the angular resolution that you use for this in, you know, in this spherical version?
你在这个,你知道,这个 spherical 版本里用的 angular resolution 是多少?
61:30
Yeah. So it's all of the data that's available is like, I think a quarter, like 0.25 degrees.
对。所以所有可用的数据大概是,我觉得四分之一,也就是 0.25 degrees。
61:38
So in terms of cell, maybe, or I mean, in terms of like a spherical harmonic
所以在 cell 方面,也许,或者我是说,在 spherical harmonic 方面。
61:46
Oh, you mean like how many modes we utilize? I think we, so for that resolution,
哦,你是说我们用多少种 modes?我想我们……对于那个 resolution,我们基本上利用了大部分,只有少数几个会活下来。细节我忘了,但我就是好奇,你们实际解析到地球上的 angular resolution 是多少,或者可能只是 physical resolution?对,我的意思是,那就是因为我们拿到的数据就是这样,对吧?所以现在我们拿到的数据大概是四分之一,就是 0.25 degrees。哦,好。像是 0.25 solid angle。我觉得算下来大概是,你知道,700 乘以几千这样的 resolution。所以,但这已经是标准的,就是处理过的了。对对,我只是想理解一下。
61:52
and we essentially utilize, I think most of it, only a few of them will live out. I forget the
而我们基本上利用,我觉得大部分,只有少数会存活下来。我忘了那个
61:57
details, but I'm just curious, like, what is the actual angular resolution on the globe that you
细节,但我就是好奇,比如你在地球上解析到的实际 angular resolution 到底是啥,你
62:02
are resolving to or maybe just the at the physical resolution? Yeah. I mean, that's what that
解析到的是什么,还是说就只是在 physical resolution 上?对,我是说,那就是因为那是我们拿到的 data,对吧?
62:07
because that's the data we get, right? So right now, the data we get is like a quarter,
因为那是我们拿到的 data,对吧?所以现在我们拿到的 data 大概是四分之一,
62:12
like 0.25 degrees. Oh, okay. Like 0.25 solid angle.
就是 0.25 degrees。哦,好的。就是 0.25 solid angle。
62:17
I think kind of comes out to like, you know, 700 by a few thousand like resolution. So,
我觉得大概算下来,就是,你知道吧,700 乘以几千这样的 resolution。所以,
62:25
but this is already standard, like kind of processed. Yeah, yeah. And just trying to understand
但这已经是标准的了,就是处理过的。对对,然后只是想理解一下
62:31
like how large is the basis do you need to represent? Yeah, I mean, that's really depends on
比如你需要多大的 basis 来表示?是啊,我是说,这真的取决于
62:35
the resolution. And the idea is, you know, right now, or whether data is just limited by this
resolution。然后想法是,你知道,现在,或者说数据是否只是被这个
62:40
resolution, but if you could, you know, you could like kind of do synthetic climate simulations
resolution 限制,但如果你能,你知道,你可以有点像做 synthetic climate simulations
62:45
of even higher resolution, right? And that's kind of the next thing on how to combine these together.
甚至更高 resolution 的,对吧?这就是接下来要做的事,怎么把这些结合起来。
62:51
Do you think you can predict with super resolution and basically resolution
你觉得你能用 super resolution 来预测,而且基本上是 resolution
62:56
lower than the data provided? Again, like yes, we can always predict them with the neural operators.
比提供的数据更低吗?另外,是的,我们总是可以用 neural operators 来预测它们。
63:03
But, you know, you do want to incorporate more of the physical constraints to ensure that they are
但是,你知道,你确实想加入更多的 physical constraints 来确保它们是
63:08
valid. Okay, so can we talk about some of the other? Yes, yes, I know, it's a lot. So this is just
有效的。好,那我们能聊聊其他的事情吗?是的是的,我知道,内容很多。所以这只是
63:17
showing like how, you know, the one I described that on the left where the world is being assumed,
就是展示一下,你懂的,我之前描述的那个左边的情况,假设世界是个矩形,它很快就爆掉了。而右边,因为我们假设世界是个球体,它就一直滚动下去,一直保持稳定。所以那里也有一点 singularity,对吧?而且它还是像,你懂的,所以想法是,没错,因为这是一个非常长的 rollout,而且我们没有任何物理上的 guardrails。我们并不是,你懂的,在把它投射到正确的物理上,对吧?这是完全的 extrapolation。但关键是,球体假设会在更大程度上稳定它,跟左边相比,好太多了。对。但如果你在南极,你仍然不会精确地到达那个极点,先生,难点就在这。所以所以,这就是一个例子。
63:24
it's a rectangle, it blows up very quickly. And on the right, because we assume the world was a
它是一个 rectangle,会非常快地放大。而在右边,因为我们假设世界是一个
63:29
sphere, it kept rolling it out and it kept being stable. So it's also a little bit of a singularity
sphere,它就一直在滚动展开,而且保持稳定。所以它也有点像一个 singularity。
63:36
there, right? And it's still like, you know, so the idea is yes, because it's a very long rollout
那里,对吧?而且它还是,你知道,所以想法是——对,因为这是一个非常长的 rollout
63:42
and we have no guardrails of physics. We are not, you know, kind of projecting it to the right
而且我们没有物理上的 guardrails。我们并不是,你知道,在把它投影到正确的物理上,对吧?这是完全的 extrapolation。但关键在于 sphere assumption 能把它稳定住,稳定得多。跟左边相比好太多了。对。但如果你在南极,你还是没法精确到达那个极点,老兄,难就难在这。所以所以这是一个事先就知道赢家的例子,对吧?我喜欢用不同的方法,你知道,看看 AI 能不能加速所有这些。然后我们就可以,你知道,不会过早地排除一种方法而选择另一种。所以这就是 AI 让我们更能承担风险、去——
63:47
physics, right? This is full extrapolation. But the idea is the sphere assumption stabilizes it
物理,对吧?这完全是 extrapolation。但关键是,sphere assumption 会让它稳定
63:55
to a much greater extent. Compared to the left, it's much better. Yeah. But if you're in the South
到更大的程度。跟左边比,好得多。是的。但如果你在南极
63:59
Pole, you're still not going to get exactly that pole, sir, the hard part. So so this is an example
附近,你还是没法精确得到那个极点,先生,这就是难点。所以,所以这是一个例子。
64:04
of the fusion reactor. So this is a tokamak. And we are able to model the complex plasma evolution
关于 fusion reactor 的。所以这是个 tokamak。我们能够模拟复杂的 plasma 演化
64:14
and do this a million times faster than what we could do with traditional simulations. And
并且这比传统 simulations 要快一百万倍。而且
64:20
this was in a way, we're creating a digital twin of the plasma, right? And then we can, you know,
这某种程度上是在创建 plasma 的 digital twin,对吗?然后我们可以,你知道,
64:27
do further things like right now, we are as a next step looking at like control. But with a full
做进一步的事情,比如现在,作为下一步我们在考虑 control。但要基于完整的
64:32
valid physics, like being able to prevent disruptions ideally and make fusion sustainable.
有效 physics,比如理想情况下能够预防 disruptions,让 fusion 可持续。
64:39
So are you simulating MHD equations here? Or sorry, I mean, you know, hydrodynamic equations.
所以你们是在模拟 MHD equations 吗?抱歉,我的意思是,hydrodynamic equations。
64:44
Yes. Okay. And then so for context disruption in this case is this phenomenon, which plagues,
是的。好的。那么为了说明背景,这里的 disruption 是一种现象,困扰着,
64:50
which plagues plasma physicists where at some point your entire plasma collects in a little
困扰着 plasma 物理学家,就是某个时刻你的整个 plasma 会聚集成一个很小的
64:55
tiny beam and then shoots a strong, you know, right to your containment vessel. And it can
微小的 beam,然后射出一道很强的,你懂的,直接打到你的 containment vessel 上。它可能会
65:00
damage the reactor. That's the that's a big bottleneck because then you have to kind of shut it
损坏 reactor。这就是个很大的 bottleneck,因为在那之前你不得不把它关掉
65:05
down before that happens. And then plus it's no longer possible to have a sustainable fusion.
以防止那种情况发生。而且再说了,可持续的 fusion 已经不可能了。
65:11
So there's a lot of open challenges here. But the idea is, you know, it's very expensive to go
所以这里有很多开放的挑战。但想法是,你知道,做物理实验非常昂贵。
65:17
to physical experiments. The more you can capture that in the digital twin, but ensure physical
你越能在 digital twin 中捕捉这些,同时确保物理
65:23
validity, the more you can even do design and other considerations in the digital realm.
有效性,你就越能在 digital 领域里做设计和其他考量。
65:31
You know, we can hopefully make advances. And these are the first steps towards that.
你知道,我们有希望取得进展。而这些是迈向那个目标的第一步。
65:36
The goal is that if you have one of these events that you can somehow adjust the magnetic field
目标是,如果发生这些事件之一,你能以某种方式调整 magnetic field。
65:45
so that it contains that and stabilizes. Yes. And that's the next step we are doing now.
所以它能包含那个并稳定下来。是的。这就是我们现在正在做的下一步。
65:50
We are looking at like designing both the control and the simulation together.
我们在考虑把 control 和 simulation 一起设计。
65:56
Are you working with the specific lab? I'm just curious. So this one was with the UK Atomic Energy
你们是和具体的实验室合作吗?我只是好奇。所以这个项目是和 UK Atomic Energy
66:01
Agency. And now we're also working with a few others here in the US as well. So we are, you know,
Agency。现在我们也在美国这边和其他几个合作。所以我们,你知道,
66:08
kind of getting the information from many different approaches of fusion itself. So this is the
从 fusion 本身的多种不同方法中获取信息。所以这就是
66:14
talkum out. We're also working with Stellarators. We are working with different. Stellarators are
tokamak。我们也在和 Stellarators 合作。我们在做不同的。Stellarators 很
66:20
tricky. But the idea is ideally, you know, like our goal is to be able to design them in the digital
棘手。但理想情况下,你知道,我们的目标是能在 digital
66:27
twin. So can we come up with good designs that would make it maybe more practical? And so that's
twin 里设计它们。所以我们能不能想出好的设计让它们更实用一些?所以那就是
66:34
I think also good thing as an AI person and much more like, you know, agnostic and not picking a
我觉得作为AI领域的人,还有一个好处就是可以更加agnostic,不会事先就选定哪个赢家,对吧?我喜欢尝试不同的方法,然后看看AI能不能加速所有这些方法。这样我们就可以,嗯,不会过早地把某一种方法排除掉。所以这就是AI让我们能够更愿意去冒险、探索不同的方法,而不是像在物理世界里,你要去建造这些东西的话,就不得不砍掉很多有风险的方向,然后说我只做这个,因为这是最有可能成功的。而且注意看你的职业生涯,你一开始花了很多时间在非常理论的基础上,就是machine learning的数学。然后也许……
66:41
winner beforehand, right? Like I like to work with different approaches, you know, and see
事先就有赢家,对吧?比如我喜欢尝试不同的方法,你知道的,看看 AI 是否能加速所有这些方法。然后我们可以,怎么说呢,不会过早地排除一个方法而偏向另一个。所以这就是 AI 让我们更能承担风险,以及各种其他原因,你知道的,然后你还得建立理论基础,然后... 还有,你知道的,我们能不能模拟它们几十年的迁移?所以这个,我们能做到比 traditional simulations 能做的快得多。我的意思是,另一个方面是能够处理各种 geometric shapes,比如,你知道的,在汽车、飞机等上面建模 aerodynamics。所以同样地,这是一个很好的 latent space 的例子,因为...
66:48
whether AI can accelerate all of them. And then we can kind of, you know, not prematurely rule
AI 是否能加速所有这些。然后我们可以有点,你知道,不要过早地排除
66:54
out one approach over the other. So that's what AI enables us to be more kind of taking risks and
掉一个方法而不是另一个。所以这就是 AI 让我们能更愿意冒险,而且
67:01
exploring different approaches as opposed to in the physical world, trying to build any of these.
探索不同的方法,而不是在物理世界里尝试构建这些东西。
67:07
You kind of have to cut a lot of the risky ones and say I'm only going to do this because
你多少得砍掉很多风险高的方向,然后说我只做这个,因为
67:12
this is the most likely to work. And notice over your career, you started out spending a lot of time
这是最有可能成功的。而且你注意到,在你的职业生涯里,你一开始花了很多时间
67:17
on, you know, really theoretical foundations, the mathematics of machine learning. And maybe,
在那种非常理论的基础上,你知道,就是机器学习的数学。然后可能,
67:23
I don't know, some like six, eight years ago, you started working really, working a lot on
我不知道,大概六到八年前吧,你开始非常投入地、大量地做 applications,扩展到各种各样的问题上。是什么促使你转换了关注的方向?我是说,你其实也还在做非常难的 math problems,比如 torch lane 那个工作。但 applications 确实扩展了很多,我很好奇是什么原因?还有,你从那时起学到了哪些 lessons?
67:29
applications and branching out in a diverse set of problems. What sort of prompted that shift in
应用以及向各种不同问题的拓展。是什么促使了这种
67:35
your approach of what you're looking at? I mean, so you're still working on very hard math
当然。我是说,对我来说,就是,你知道,我觉得我是和 AI 一起成长的,对吧?所以当 AI 还处于——你知道——neural nets 不 work 的阶段,因为 data 不够,还有各种其他原因,你知道,那时候你就得去构建 theoretical foundations,然后...
67:40
problems as well, like, for example, the torch lane work. But the applications have really grown
问题方向上的转变,比如 torch lane 的工作?但应用确实已经发展起来了,
67:48
and I'm wondering what prompted that? And like, where are some of the lessons you've learned since then?
我很想知道是什么推动了这一变化?还有,你从那时起学到了一些什么教训呢?
67:53
Sure. I mean, to me, it's like, you know, I feel like I've grown along with AI, right? So when AI
当然。我的意思是,对我来说,你知道,我觉得自己是和 AI 一起成长的,对吧?所以当 AI
68:00
was, you know, in this where neural nets were not working because there wasn't enough data and
就是,你知道,在那个阶段 neural nets 就是不 work,因为 data 不够,还有各种其他原因,你知道,那时候你就得去建立 theoretical foundations。然后,你知道,我们能不能 model 它们几十年间是怎么迁移的?所以这个我们能做到的比传统 simulations 快得多。我的意思是,另一个方面是能够做各种 geometric shapes,比如,你知道,能够 modeling 汽车、飞机等等的 aerodynamics。所以这又是一个很好的 latent space 的例子,因为你可以把一个 car 或者任何其他 shape 变换成一个 donut,然后在 donut 上 modeling,然后再把 donut 变回 car。所以你把它变成了 coffee cup。这就是那个经典笑话吗?
68:06
all kinds of other reasons, you know, then you kind of have to build the theoretical foundations and
各种其他的原因,你知道,然后你就得建立理论基础,而且
68:11
try to hope that that leads you to a place where, you know, you get algorithms to work, right?
试着希望那能带你到一个,你知道的,让 algorithms 能正常运转的地方,对吧?
68:18
And, and, and, you know, back then, extensive methods was with that idea that, you know,
而且,而且,而且,你知道,那时候,extensive methods 就是基于那个想法,你知道,
68:24
pre-deep learning, we still want a structure. We have probabilistic models, like late and
在 deep learning 之前,我们仍然想要一种结构。我们有 probabilistic models,比如 Latent
68:28
Dirichlet allocation for topic modeling and solving those were hard. But now,
Dirichlet allocation 用于 topic modeling,解决那些问题很难。但现在,
68:34
tens of methods gave us a way to be very practical. It's parallel and can be done at large scale,
几十种 methods 给了我们一个非常实用的途径。它是 parallel 的,而且可以在 large scale 上做,
68:40
but still has nice theoretical basis. So that was where, you know, we're starting off and then as
但仍然有很好的理论基础。所以那就是,你知道,我们开始的地方,然后随着
68:46
deep learning started taking off and we could see that it works well in practice. And yes,
deep learning 开始起飞,我们能看到它在实践中效果很好。而且,没错,
68:51
there is little bit of maybe theoretical understanding, but not a whole lot because of the way
也许有一点点理论上的理解,但不是很多,因为这种方式
68:56
how complex it is. To me, theory should not be a constraint, right? It should be an enabler. And so
它有多复杂。对我来说,理论不应该是约束,对吧?它应该是赋能者。所以
69:03
that's where a lot of like the exploration was out to make this work well in practice and over
这就是很多探索的所在——让它在实践中运行良好,并一直到
69:10
until Amazon Web Services, then Nvidia. So really like making things work at scale and really kind of
Amazon Web Services,然后是 Nvidia。所以真的是让事情在规模上运作,而且真的有点
69:17
getting hands dirty, right? It was kind of like where a lot of the development is. And now I see
动手实践,对吧?那里才是大量开发工作发生的地方。而现在我看到
69:23
a full circle because a lot of purely data driven approaches in a way seeing saturation, right? So
一个完整的循环,因为很多纯粹的 data-driven approaches 某种意义上正在趋于饱和,对吧?所以
69:30
now we want to ask, okay, either make them more hardware efficient, right? There's a lot of now room
现在我们要问,好吧,要么让它们的硬件效率更高,对吧?现在有好多空间
69:37
to kind of say, can we now, you know, make them much more energy efficient or hardware efficient?
可以说,我们现在能不能,你知道,让它们更节能,或者硬件效率更高?
69:43
So that's one aspect. But the other is areas like this where in the physical world, we don't have
所以这是一个方面。但另一个是像这样的领域,在物理世界里,我们没有
69:49
enough data. We are asking for hard extrapolation. You know, we want to think of doing discovery by
数据足够了。
69:55
nature. It's about extrapolation. So we will never have data about a new discovery, right? That's
我们要求的是 hard extrapolation。
70:00
my definition. And so there we need to again go back to thinking in principle ways in
你知道,我们想的是,discovery 本质上应该是什么样的。
70:07
whether it's architecture design, algorithm design, the right loss functions. So we need to be
它关乎 extrapolation。
70:14
much more mindful. So I see that coming a full circle because all of the things that work with
所以我们永远不会有一个新 discovery 的数据,对吧?
70:20
deep learning, let's take them but make them a bit more principled. There's several other
这是我的定义。
70:25
applications which seem very natural. I'm wondering if you've worked on these or did I just miss
所以在那方面,我们得再次回到更原则性的思考上,不管是 architecture design、algorithm design 还是正确的 loss functions。
70:30
some papers if I did? I'm sorry. So some examples are design of like electromagnetic circuits.
所以我们需要更留心。
70:36
I think it's a big one or maybe not a big one, but I think we'll be coming up in the near future.
我觉得这可能是个大事,或者也许不是,但我觉得在不久的将来我们会遇到。
70:42
Design of, let's say materials, design of let's say dissipation in heat sinks or any sort of
比如说materials的设计,比如说heat sinks里的dissipation设计,或者任何类型的
70:50
like fluid flow. I guess I'm going through what differential equations. So I know electromagnetic
比如fluid flow。我想我是在过differential equations。所以我了解electromagnetic
70:55
system, I know diffusion equations, you know, MHD. Yeah. I'm wondering some of the other domains
system,我知道diffusion equations,你知道,MHD。对,我在想其他的一些领域
71:04
Yeah, I mean, to me, there is just endless possibility, right? So there, you know, as like you
对,我的意思是,对我来说,这就是无限的可能性,对吧?所以,你知道,就像你
71:10
can just have this work on any data and we have several other examples. So this was like,
可以让它在任何data上工作,而且我们还有几个其他例子。所以这就像,
71:16
you know, being able to ask, can we sequester carbon dioxide underground and model
你知道,能够问,我们能不能把carbon dioxide封存在地下,然后model
71:23
how carbon dioxide expands or, you know, what is the pressure build up in these reservoirs?
carbon dioxide如何膨胀,或者,你知道,这些reservoirs里面的pressure build up是怎样的?
71:30
And, you know, can we kind of model how they migrate over several decades? And so this one,
还有,你知道,我们能不能 model 它们几十年的迁移?所以这个案例,
71:37
we were able to do much faster than what traditional simulations could do. I mean, the other aspect
我们能做到比传统 simulations 快得多。我的意思是,另一方面
71:42
is being able to do all kinds of geometric shapes like, you know, being able to model aerodynamics
是能够处理各种几何形状,比如,你知道,能够 model aerodynamics
71:48
in cars, planes, and so on. And so again, this is a nice example of a latent space because
在汽车、飞机等等上面。所以再说一次,这是一个很好的 latent space 例子,因为
71:55
you can transform a car or any other shape to a donut and then model on the donut and then
你可以把一辆汽车或任何其他形状变形成一个甜甜圈,然后在甜甜圈上 model,然后再
72:01
transform the donut back to the car. So you turn it into a coffee cup. Is that the classic joke?
把甜甜圈变回汽车。所以你就把它变成咖啡杯。这是那个经典笑话吗?
72:07
You're done in a coffee cup actually. So my idea of like a latent space to handle all kinds
你的 domain 其实只是一个咖啡杯。所以我的想法是,用像 latent space 这样的东西来处理各种不同的 geometries,并且能够在 latent space 里很好地捕获物理信息,这意味着我们现在可以有一个 model,能够 generalize 到很多不同的 geometries。而我的理解是,这里更大的愿景可能是,你可以 train 一个 foundation model,用同一个 model 去建模很多不同的物理现象。所以你可能 fine-tune,或者给它某种 prompt,让它理解特定的 geometry,但你在所有这些不同的物理问题上 train,然后你就有了你自己的那一个,然后你就能够...
72:12
of different geometries and and be able to capture the physics there in the latent space well
关于不同的 geometries,并且能够在 latent space 中很好地捕捉物理
72:18
means we can now have a model that generalizes across a lot of different geometries.
意味着我们现在可以有一个 model,能泛化到很多不同的 geometries。
72:24
And my understanding that the maybe the larger vision here is that you can train a
而我的理解是,也许更大的愿景是你可以训练一个
72:33
foundation model in the sense of being able to model many different physical phenomena
foundation model,也就是那种能够 model 很多不同物理现象的 model
72:38
with the same model. And so you may fine tune or there may be some kind of prompt that you give it
用同一个 model。所以你可以 fine tune,或者给它某种 prompt
72:43
to have you understand the particular geometry but that you sort of on all these different
让你理解特定的 geometry,但你有点像是在所有这些不同的
72:50
physical problems you train and then and then you have your particular one and you're able to
对于物理问题,你进行训练,然后得到你的特定模型,然后你能够
72:54
model that very effectively. Yeah, I mean, that's really the future, right? Because we have
非常有效地建模。是的,我的意思是,这真的是未来,对吧?因为我们有
72:58
foundation models for language, maybe vision but not for physics. So, you know, the idea is
用于语言的 foundation models,也许还有视觉,但没有针对物理的。所以,你知道,现在的想法是
73:05
instead of like right now what we've seen are narrow surrogates and we're trying to broaden
不像现在这样,我们看到的都是 narrow surrogates,我们正试图不断拓宽
73:10
their scope more and more but ideally we have much broader models that can work on a range of
它们的范围,但理想情况下,我们会有更广泛的模型,能够处理一系列
73:16
phenomena but also multi-physics so not just have like one single physics but coupled physics
现象,还要支持 multi-physics,所以不只是单一的物理,而是耦合的物理。
73:22
the real world has all of the physics kind of coming together in coupled ways so can we bring
现实世界中的所有物理现象都以耦合的方式结合在一起,所以我们能不能把
73:29
all that together. So that's one aspect like, you know, have foundation models that can do
这一切都整合起来?所以这是一个方面,比如,让 foundation models 能够做到...
73:34
design that can do simulation. But the other aspect that's really interesting is the inverse problem,
能够进行 simulation 的设计。但另一个真正有趣的方面是 inverse problem,
73:40
right? So can I now not just simulate but ask what is the best design and then these kinds of
对吧?所以我现在能不能不只是做 simulation,而是直接问什么是最好的设计?
73:47
like models can like do simulation but you can even do that implicitly and come up with the best
这类模型确实可以做 simulation,但你甚至能隐式地做到这一点,然后得出最好的 standard model。
73:54
design rather than in the earlier era of it was humans trying to come up with design then you go
设计,而不是在早期那种人类试图提出设计、然后你去
74:00
and try to simulate or go to the wind tunnel whatever physical testing and validate that but now you
做模拟或者去风洞做物理测试来验证,但现在你
74:06
have AI come up with optimized designs but you have the guardrails of physics so you have models
让AI提出优化设计,但你有物理的护栏,所以你有模型
74:12
that are accurate in physics you have the confidence they work well so you're kind of able to do
在物理上是准确的,你有信心它们能正常工作,所以你差不多能够做到
74:19
that as well in the same model. Is there reason have you seen any evidence that you talked about
在同一个模型中也做到这一点。有没有理由,你有没有见过任何证据,你谈到的
74:25
these like sort of multi-physics being able to transfer or that you may be able to generalize
这些multi-physics之类的能够transfer,或者你可能能够generalize
74:34
to sort of unseen physics? So I mean, you know, like the physics by nature, if it's completely
到某种unseen physics?所以我的意思是,你知道,物理本质上,如果它完全是
74:41
unseen it's not possible to transfer, right? I mean, if we are saying that we're going beyond the
unseen的,那就不可能transfer,对吧?我的意思是,如果我们说要超越这个
74:47
standard model there's absolutely no data that's not possible. But if you're asking about like,
完全没有数据的话,那是不可能的。
74:53
you know, for instance like, you know, there is the like say I've like, you know, shown it examples
但如果你问的是,比如说,你知道,比如我给它看过一些例子,像热量怎么传播,还有别的例子是材料怎么拉伸。
75:02
of like just how the heat propagates and the other examples of how the material like stretches
现在有了 coupling,因为热也会导致拉伸,也就是某种联合现象。
75:08
and now there's coupling like because of heat there's also stretching or kind of the joint phenomena.
那现在你可以希望用少得多的样本做 fine-tuning,因为它已经分别掌握了这些现象,然后再把它们组合起来。
75:14
You could like now hope to fine tune with much fewer samples because it kind of individually
也许它没法从零开始做到,因为那是……
75:20
knows this phenomena then combining them together maybe it can't do it from scratch because that's
知道这些现象,然后把它们组合在一起,也许它不能从零开始做到这一点,因为那是
75:26
too much to ask. It's highly non-linear and coupled but it can do it with fewer examples and
这要求太高了。它是高度 non-linear 且 coupled 的,但可以用更少的例子做到,而且
75:32
we've seen evidence of that in a lot of our papers that you're able to kind of essentially build up
我们在很多论文里都看到了证据,你基本上能够构建起
75:38
a curriculum and that's what we see again and again in many of these examples that, you know,
一个 curriculum,我们在很多例子里反复看到的就是,你知道,
75:43
the real world we can kind of control a lot of curriculum and say, you know, let's kind of build
在现实世界里,我们其实能在很大程度上控制 curriculum,然后说,你知道,我们来构建
75:50
in like modules and put them together and that's what it now allows us to do in a systematic way here.
像 modules 这样的东西,再把它们拼起来,而这就是它现在让我们能在这里系统化地做的事情。
75:58
I guess the design aspect I don't know if we wanted to show very quickly so this one was like,
我想说设计这方面,我不知道我们是不是想很快展示一下,所以这个就像是,
76:04
you know, looking at like designing the mask for inversely photography meaning now this
你知道,比如为 inverse photography 设计 mask,也就是说这现在变成了
76:10
an inverse design problem and we are also able to do that for designing gates and quantum dots.
一个 inverse design 问题。我们也可以用它来设计 gates 和 quantum dots。
76:18
This is like non-linear photonics and all of this what is common is the idea that,
这就像是 non-linear photonics,所有这些的共同点是,
76:24
you know, there's a forward model that is simulating the physics but now what we want is the
你知道,有一个 forward model 在模拟物理,但现在我们想要的是那个……
76:31
inverse design like the problem that we can optimize the best design and humans are usually not good
inverse design 就像是我们能优化出最佳设计,但人类通常不擅长这个,对吧?
76:39
at this, right? We are not good at like looking at highly non-linear phenomena and say,
我们不擅长观察高度 non-linear 的现象,然后说,哦,也许所有这些 gates 组合在一起,能帮助在 quantum gate 中把电子聚集起来。
76:44
oh, somehow maybe this combination of all these gates coming together helps pool the electrons
所以我们的 collaborators 之前手动做这件事非常吃力,而有了 AI,我们现在能提出非常高效的 designs,而且我们也知道这些 designs 确实有效,因为我们已经把 simulation 放到 loop 里面,它会说这些设计效果很好,所以我觉得
76:50
together in a quantum gate. And so our collaborators were struggling to do that manually and with AI
这些例子让我们看到,这不仅仅是 simulation 的问题,而是真正新颖的 designs 和发现,能够推动 innovation 本身向前发展。
76:56
we are now able to come up with very efficient designs but also those we know actually work
我们现在能提出非常高效的设计,而且我们也知道这些设计确实管用
77:04
because we have already the simulation as part of the loop saying that they work well so I think
因为我们已经把 simulation 作为 loop 的一部分,来表明它们效果很好,所以我觉得
77:10
these are examples where we see that it's not just about simulation, it's about really novel designs
这些例子说明,这不只是 simulation 的问题,而是真正新颖的设计
77:17
and novel discoveries that enable us to move the needle of innovation itself.
以及新颖的发现,让我们能够真正推动 innovation 本身的发展
77:23
Each one of these examples takes a lot of domain knowledge. How could somebody take your basic
这些例子每一个都需要大量的领域知识。如果一位领域专家想基于你的基础研究,并且快速开始把 neural operators 和你开发的其他 frameworks 应用到他们的问题上,该怎么做呢?对——在 neural operators 或者说 open source library 里,它已经被广泛采用了。它是 PyTorch ecosystem 的一部分。你知道,它不仅被很多研究人员使用,也被很多公司使用。我们那里有大量的 documentation,所以我鼓励大家去看。我们有你看,很多不同的 architectures、examples、recipes。所以我觉得那是一个很好的起点。
77:31
research if a domain expert and quickly get started applying neural operators and the other frameworks
如果你是领域专家,研究一下就能快速上手应用 neural operators 和其他框架
77:39
that you've developed to their problem? Yeah. In a neural operators or an open source library,
你最近加入了 UN Scientific Advisory Board。我知道我们时间不多了,
77:44
it's extensively already adopted. It's part of the PyTorch ecosystem. It's, you know,
它已经被广泛采用了。它是 PyTorch ecosystem 的一部分。你知道,
77:50
used by a number of not only researchers but also in companies. We have a lot of documentation
它不仅被很多研究者使用,也被公司使用。我们这里有大量的文档
77:56
there so I encourage people to go there. We have like, you know, many different architectures,
所以我鼓励大家去看看。我们有,嗯,很多不同的 architectures、
78:03
examples, recipes. So I think that's a great place to get started.
examples、recipes 等等。所以我觉得那是个很好的入门地方。
78:07
You recently joined the UN Scientific Advisory Board. I know we're running out of time,
你最近加入了 UN Scientific Advisory Board。我知道我们时间不多了,
78:11
but maybe just can you quickly give a bit of the story behind this and what you hope to accomplish?
但能不能快速讲一下这背后的故事,以及你希望达成什么?
78:17
Yeah. I'm really honored to be part of that advisory board for the UN and in these tricky times
嗯,我真的很荣幸能成为 UN 那个咨询委员会的一员,而且在这个棘手的时期,
78:24
with a lot of geopolitics there, you know, which again, I'm not the expert on that, but when it comes
地缘政治很复杂,你知道,我也不是这方面的专家,但说到
78:29
to, you know, aspects, especially related to AI, having scientists in the room is something that,
嗯,特别是与 AI 相关的方面,让科学家参与其中是一件……
78:35
you know, I think is very important. I hope I can have an unbiased view and try to provide scientific
你知道,我认为这非常重要。我希望我能有一个客观的视角,并尝试提供科学的
78:42
evidence for any aspect, right? We want to think about how AI impacts globally, like, you know,
证据,对吧?我们想思考 AI 如何在全球范围内产生影响,比如,你知道,
78:48
how do we ensure the benefits of AI reach everybody? How do we democratize access to AI?
我们如何确保 AI 的好处能惠及每个人?我们如何让 AI 的获取民主化?
78:55
How do we ensure the unintended consequences and harmful impacts can be controlled? I think these are
我们如何确保意外后果和有害影响能得到控制?我认为这些是
79:02
just the beginning aspects. Of course, the other side when it comes to weather models, I'm already
这些还只是开始的一些方面。当然,另一方面,说到 weather models,我已经很兴奋了,你知道,现在有一股推动力,想看看我们怎么才能把 weather 和 climate modeling 做得更好,然后用到食物上,比如用天气来改善农业。所以所有这些方面,UN 在全球也有很多专门的机构和一线人员。所以我非常期待能参与进去,成为其中的一部分。
79:08
excited like, you know, there are, there's a push to seeing how we can have better weather
令人兴奋的,比如,你知道,有人正在推动探索我们如何能拥有更好的天气
79:13
climate modeling, so then our food, you know, like using weather for better agriculture. So all
你知道,看看你的职业经历,听你说话的方式,还有你做的事情,感觉你是一个非常喜欢解决具体问题的人。你不喜欢对事情做哲学式的空想,我觉得这样你反而会,可能更乐观一点,比……
79:19
these aspects are also where UN has a lot of dedicated agencies and people on the ground across
这些方面也是联合国拥有很多专门机构和遍布各地的实地人员的地方
79:26
the world. So I'm looking forward to contributing and being part of this.
这个世界。所以我期待能为此做出贡献,成为其中一员。
79:30
You know, looking, you know, looking at your career and how you, you know, talk and what you
你知道,看着,你知道,看着你的职业生涯,以及你,你知道,说话的方式和你所
79:35
work on, it seems like you very much are a person who likes to solve concrete problems. You don't like
做的事情,感觉你很像是一个喜欢解决具体问题的人。你不喜欢
79:40
to philosophize about things, which then you're also seeing, I think, maybe more optimistic than
去空谈哲理,而且你也会发现,我觉得,可能比一些人更乐观。
79:45
a lot of people in the AI space. You have a very hopeful view of the world, I think, not honestly
AI 领域里有很多人。你对世界抱有一种非常乐观的看法,我觉得,说实话,这并不总是真的。你能以哪些独特的方式把这种观点带到董事会,而不是像某些人那样,嗯,你知道的,
79:52
always true. What are the ways that you can uniquely bring that viewpoint to the board versus maybe
谢谢。我,呃,对我来说,我想,正如我所说,我努力保持公正。作为一名科学家,我认为 AI 有很多有益的方面,当我们只考虑有害影响时,这些方面有时会被忽略,对吧?而且尤其是在 AI for science 方面,因为很多监管框架把 AI 等同于 language models。是的,language models 可以,呃,操纵人们,能产生各种我们应该考虑控制的有害影响。
80:00
some, you know, you know, thank you. I, you know, to me, I think, as I said, I tried to be
一些人,你知道,你知道,谢谢。我,你知道,对我来说,我觉得,如我所说,我尽量保持
80:06
unbiased and as a scientist and as a scientist, I think that there's a lot of beneficial
客观,而且作为一名科学家,作为一名科学家,我觉得AI有很多有益的
80:13
aspects of AI that are sometimes missed when we think of only the harmful impacts, right? And
方面,这些方面有时被忽略了,当我们只想到有害影响的时候,对吧?而且
80:19
and especially that is with respect to AI for science, because a lot of regulatory frameworks
尤其是关于AI for Science,因为很多监管框架
80:24
equate AI with language models and yes, language models can, you know, have manipulate people,
把AI等同于语言模型的话,是的,语言模型确实能操纵人,可能带来各种有害影响,这些是我们需要考虑去控制的。
80:32
can have all these kinds of harmful impacts that we should think about controlling.
可能会产生所有这些类型的有害影响,我们应该考虑如何控制。
80:36
But AI for science is different. So I think this one size fits all is where a lot of problems come
但是 AI 用于科学就不一样了。
80:43
up. So we have to be mindful that there is, you know, AI that can change the world with new
所以我觉得“一刀切”的做法会引发很多问题。
80:50
discoveries and we should enable people around the world to not only benefit from them but also
所以我们得意识到,有些 AI 能通过新发现改变世界,
80:56
be able to do research, you know, have access to AI that they can go in a way and use them in
我们应该让世界各地的人不仅能从中受益,还能自己做研究,能够接触到 AI,并以有趣的方式去使用它。
81:04
interesting ways. One question that we have been trying to ask every guest is if you could
我们一直尝试问每位嘉宾一个问题:如果你能神奇地消除你所在领域的一个瓶颈,那会是什么?为什么?
81:11
pick a bottleneck in your domain that you could magically remove, what would that be? And why?
更多的 compute。
81:18
More compute. I know that's an easy one or maybe a lazy one, right? Because, you know,
我知道这是个很简单的问题,或者说有点偷懒,对吧?
81:25
and you know, of course, our compute that we have is growing so much more than even a few years ago,
因为,你知道的,当然,我们现在拥有的 compute 比几年前增长太多了。
81:33
thanks to NVIDIA, thanks to others. So no, they're coming so bad. But what I mean by that is also like
感谢NVIDIA,感谢其他所有人。所以不,他们来得太糟糕了。但我的意思是,这也像是
81:41
for research enabling more and more compute, you know, it's very important. I know there are
为了研究,让越来越多的compute成为可能,你知道,这非常重要。我知道有
81:47
national labs building more super computers, you know, hoping that we can have more compute for
国家实验室在建造更多的超级计算机,你知道,希望我们能有更多的compute用于
81:54
research. But I think, you know, without that, we cannot experiment, we cannot innovate. I think
研究。但我认为,你知道,没有那个,我们就无法实验,无法创新。我觉得
81:59
this is a part that I push a lot and, you know, I think I cannot emphasize that it's so critical.
这是我一直大力推动的部分,而且,你知道,我觉得我无法强调它有多关键。
82:07
If you had a call to action or something that you would like people to do or think about,
如果你有一个行动号召,或者你想让人们去做或思考的事情,
82:15
or learn about what would that be? Yeah, so, you know, it can go to neural operator libraries,
或者去了解什么,那会是什么?是的,所以,你知道,可以去看看neural operator libraries,
82:20
so you can kind of hands-on play with different architectures, recipes, you know, look at use cases.
这样你就可以亲手试试不同的architectures、recipes,你知道,看看use cases。
82:28
But also think about like, you know, AI for science is not just language models and agents. Yes,
但也要想想,你知道,AI for science 不仅仅是 language models 和 agents。
82:34
that's one aspect of it. But ultimately, you know, those are still like external rappers,
是的,这只是其中一个方面。
82:40
in a way, right? Until we have AI that fully understands the physical world, not just as symbols,
但归根结底,你知道,那些在某种程度上还是像外部的包装,对吧?
82:46
but as one that can simulate and design and control based on that, you know, there's a big piece
除非我们有了能真正理解物理世界的 AI,不只是把它当作符号,而是能基于这种理解去模拟、设计和控制,你知道,这还缺一大块,所以这就是另一个方面。
82:54
missing, so that's the other aspect. But I think the people should really think about AI for the
但我觉得大家应该真的用这种方式来思考 AI for the physical world。
83:00
physical world in this way. Anima, this has been so fascinating. I'm excited to check out
Anima,这真的太迷人了。
83:07
neural operators myself. I have some ideas in my head already. I really appreciate you taking the
我自己也很期待去了解一下 neural operators。
83:12
time to sit down with us. Yeah, thank you, Arches. Thank you, Brandon. I really enjoyed it. We
我脑子里已经有些想法了。
83:17
really dug deep into a number of things, so I appreciate you doing that. Thank you. Thank you.
你真的深入探讨了很多事情,所以我很感谢你这么做。谢谢,谢谢。