TWIML AI

World Models and the Future of Spatial AI with Justin Johnson - #775

2026-09-01 · 3962

In this episode of TWIML AI, Justin Johnson of the University of Michigan and World Labs discusses world models and the future of spatial AI. Johnson clarifies competing definitions of world models, linking them to MDPs, agents, observations, simulation, and rendering, and contrasts explicit 3D representations such as Gaussian splats with neural, end-to-end approaches. He also discusses World Labs’ Marble system, which generates navigable 3D Gaussian splat worlds from text or images and raises open questions about consistency, data, and model design.

本期 TWIML AI 邀请 University of Michigan 的 Justin Johnson(同时与 World Labs 相关)讨论 world models 与 spatial AI。Johnson 指出,world model 的定义在文献和业界并不统一,并借助 MDP 框架解释 agent、state、observation、action、simulator 与 renderer 之间的关系。他对比了显式 3D 表示(如 Gaussian splatting)与端到端 neural 表示,讨论两者在一致性、训练数据和生成质量上的取舍。他还介绍了 World Labs 的 Marble:可从 text prompt 或单张图像生成可导航的 3D Gaussian splat 世界。最后,他谈到缺乏大规模显式 3D 数据,以及如何用大量真实和想象数据训练可扩展的 spatial AI 系统。

00:003962
00:00
I want to send a big thanks to Blitzie for supporting the podcast and sponsoring this episode.
我想特别感谢 Blitzie 对播客的支持,并赞助了本期节目。
00:06
Want to accelerate software development velocity by 5X?
想要把软件开发 velocity 提升 5 倍吗?
00:09
You need Blitzie, which brings autonomous software engineering to your enterprise.
你需要 Blitzie,它能将 autonomous software engineering 带入你的企业。
00:14
Blitzie starts by reverse engineering your code base, building a dynamic understanding
Blitzie 首先会 reverse engineering 你的 code base,建立对你整个 application ecosystem 的动态理解。
00:18
of your entire application ecosystem.
你的工程师只需要 declare intent,一旦获得批准,Blitzie 就会 autonomously 执行整个 software epics,交付经过验证的、end-to-end tested code。
00:20
Your engineers simply declare intent, and once approved, Blitzie autonomously executes
并且,80% 的工作能在 single run 中完成。
00:25
entire software epics, delivering validated end-to-end tested code.
整个软件史诗级的项目,交付经过验证且端到端测试的代码。
00:30
With an 80% of the work completed in a single run.
其中80%的工作在一次运行中就完成了。
00:34
Blitzie is not just generating code, it's developing software at the speed of compute.
Blitzie 不只是生成 code,而是以 compute 的速度开发软件。
00:39
Experience Blitzie first hand at blitzie.com-slash-twimble, that's b-l-i-t-z-y.com-slash-t-w-i-m-l.
亲身体验 Blitzie,请访问 blitzie.com-slash-twimble,也就是 b-l-i-t-z-y.com-slash-t-w-i-m-l。
00:50
The race to build more capable AI isn't just about making language models bigger.
构建更强大 AI 的竞赛不只是让 language models 变得更大。
00:54
Increasingly, it's about world models, systems that understand space, predict how environments
现在越来越关乎 world models,这些系统理解空间、预测环境如何
00:59
change and act in the world around them.
变化,并在周围的世界中行动。
01:03
Justin Johnson is helping shape this shift.
Justin Johnson 正在帮助塑造这一转变。
01:05
He co-founded world labs with Faith Ali and is an associate professor of computer science
他与 Faith Ali 共同创立了 world labs,并且是 computer science 的副教授
01:10
at the University of Michigan.
任职于 University of Michigan。
01:12
I asked him why so many researchers think world models are AI's next frontier.
我问他,为什么这么多研究者认为 world models 是 AI 的下一个前沿。
01:16
There's a shared low-level belief among many researchers in the field that there's
在这个领域里,很多研究者其实有一个共同的底层信念,就是
01:21
something that language models aren't doing, but that there's other kind of models
语言模型有些东西是做不到的,但还有另一种模型
01:23
that we should be building they do other kinds of things.
是我们应该去构建的,它们能做别的事情。
01:26
Something around understanding the world, generating worlds, simulating worlds, reconstructing
跟理解世界有关,生成世界、模拟世界、重建
01:31
worlds, planning actions through worlds.
世界、通过世界来规划行动。
01:34
These are all capabilities that we want to build models to have.
这些都是我们希望模型能拥有的能力。
01:37
Why do we care about this?
我们为什么在意这个?
01:38
Because we want to build systems that are not just stuck in a terminal or stuck as a virtual
因为我们想要构建的系统,不只是被困在 terminal 里,或者作为 virtual agent 存在。
01:43
agent.
你想要的 AI 系统愿景,是它们会成为 robots,在现实世界中行动,或者也许你想构建 virtual worlds,生活在其中,并在那里模拟有趣的事情。
01:44
You want to have visions of AI systems that are going to be robots that are out in the world
这些能力,真的感觉不像是会自然而然地从 language modeling 范式中产生出来的。
01:47
acting in the world or maybe want to build virtual worlds and live in those and simulate
我是 Sam Charrington,这里是 Twimble AI podcast。
01:52
interesting things there.
那里有些有趣的东西。
01:54
All of these are capabilities that really don't feel like they're falling naturally
所有这些能力,感觉真的不像会自然地从语言建模的范式里冒出来。
01:57
out of the language modeling paradigm.
我是Sam Charrington,这里是Twimble AI播客。
01:59
I'm Sam Charrington and this is the Twimble AI podcast.
最近,你知道,关于世界模型的讨论,我觉得已经转向谈论那个了,也许还有其他类型的模型能产生某些种类的输入和输出。
02:02
For over a decade, I've been exploring the ideas and innovations shaping the future of
十多年来,我一直在探索那些塑造
02:06
AI through conversations like this one that help you understand what's real, what's next,
AI 未来的思想和创新,通过像这样的对话,帮助你理解什么是真实的、什么即将到来,
02:12
and what matters.
以及什么才是重要的。
02:13
Let's jump in.
我们开始吧。
02:23
Over the past few years, one idea I think that has been espouses, you know, essentially
过去几年里,有一个想法,我觉得它一直被倡导,你知道,本质上就是
02:31
the idea that a world model is like an emergent property of, you know, even language models
world model 就像是,你知道,甚至 language models 的 emergent property,
02:39
or diffusion.
或者 diffusion。
02:40
Like if you, you know, diffusion can, you know, illustrate some physical properties
比如,你知道,diffusion 可以,呃,展示一些物理性质
02:48
of the world or, you know, language models can tell stories that, you know, kind of seem
关于世界的,或者说,你知道,language models 能讲故事,那些故事,你知道,似乎
02:53
like they know something about the world.
好像它们对世界有所了解。
02:56
And I guess there's a couple of questions emerging in my, you know, my question here.
我猜我这儿冒出来好几个问题,你知道,我的问题就在这里。
03:02
One is like, maybe it's asking for a concrete definition of a world model because I think,
一个是,也许它是在要求给 world model 一个具体的定义,因为我觉得,
03:10
you know, early on in those conversations, world model was really talking about, you know,
你知道,早期那些对话里,world model 其实说的是,你知道,
03:19
these models having like foundational knowledge of the real world, like the world that we live
这些模型具备对真实世界的那种基础性知识,就像我们生活的这个世界。
03:25
in.
而最近呢,你知道,关于 world model 的讨论,我觉得已经转向谈论
03:29
More recently, you know, the world model conversation, I feel like it shifted to talking about
最近嘛,你知道,关于world model的讨论,我感觉已经转向聊
03:34
being able to create artificial worlds, but have them be self consistent and navigable
能够创造人工世界,但要让它们自洽且可导航
03:41
and, you know, properties like that.
以及,你知道,诸如此类的特性。
03:42
So, you know, a kind of, I'd love to hear you, you know, elaborate on the relationship.
所以,你知道,有点像,我很想听听你详细讲讲这其中的关系。
03:50
And if you see kind of the same shift in the terminology, but also this idea of like,
而且你是否看到了术语上的类似转变,以及这种想法,比如,
03:57
you know, emergent and, you know, do we need new things or like, if we throw enough, you
你知道,emergent,还有,你懂的,我们是否需要新的东西,或者,如果我们投入足够的
04:01
know, data, compute, et cetera, at the models that we have, you know, does that get us there
data、compute 等等,给现有的 models,你知道,那会让我们达到目标吗
04:06
or more likely, why won't that get us there?
或者更可能的是,为什么不能达到?
04:08
I think there's a lot of interesting questions to unpack there.
我觉得这里有很多有趣的问题值得展开探讨。
04:11
One, I mean, the biggest one is just like, let's get it out of the way.
第一,我是说,最大的一个就是,咱们先把它说清楚吧。
04:15
There isn't a clear definition of world models that everyone in the field agrees on.
对于 world models,整个领域并没有一个大家都认同的清晰定义。
04:19
And I think that's causing part of the confusion, right?
我觉得这在一定程度上造成了困惑,对吧?
04:21
Like, there is not like a thing where we can say a model that has X property or produces
就像,并没有一个明确的东西,能让我们说一个模型具有 X 属性,或者产生
04:26
X kind of input and produces X kind of output, like, is a world model definitionally.
X 种输入并产生 X 种输出,从定义上讲就是 world model。
04:31
I think we don't have that crisp definition as a field of what we mean, which leads to
我觉得作为这个领域,我们并没有一个清晰的、关于我们到底在说什么的定义,这导致了
04:35
the confusion.
这样的困惑。
04:37
But to your point, I think there's a couple of variants of this that feel like they're,
但你说到点子上了,我觉得这里有几种变体,感觉它们像是……
04:40
like, I think there is some notion of like implicit world knowledge that you mentioned
我觉得你提到的 implicit world knowledge 这个概念是有点道理的。可能还有其他类型的 model 会产生特定类型的 inputs 和 outputs。比如一个生成 text 的 model,如果它生成了正确的 text,那么它唯一能生成这种 text 答案的方式,就是因为它了解真实世界,或者因为它在其 neural network weights 内部建模了某种 implicit world。video model 也类似,对吧?比如,如果我能生成一个超 photorealistic 的 video,有细节,有 physics,水以非常特定的方式流动,那么也许它也在 implicitly 地建模一些东西。
04:44
that, may there are other kinds of models that produce certain kinds of inputs and outputs.
也许还有其他类型的模型,能产生特定种类的输入和输出。
04:48
Maybe if a model that is producing text, if it produces the right kind of text, the
也许如果一个模型在生成文本,如果它生成了正确类型的文本,那么
04:52
only way it could have produced this kind of text answer is because it knows something
它能够生成这种文本答案的唯一方式就是因为它了解一些
04:56
about the real world, or because it's modeling some kind of implicit world internally in
关于真实世界的东西,或者因为它在内部建模了某种隐含的世界,在其
05:01
its neural network weights, or similarly for video models, right?
神经网络权重中,或者对于视频模型也是如此,对吧?
05:04
Like, if I am able to generate a video that is super photorealistic and has detailed
比如,如果我能够生成一个超级逼真且细节丰富的视频
05:09
and like has physics and has water running in very particular ways, like, maybe implicitly
并且具有物理效果,水以非常特定的方式流动,那么可能隐含地
05:14
the model must have been modeling something about a real world in order to generate an output
这个 model 一定是在 modeling 真实世界的某些东西,才能生成那种输出。
05:18
of that kind.
所以我觉得这就是那种隐式 world model 的概念,就是说,你可能是做任何数量的任务,但某些种类的回答,你会需要隐式地 modeling 一些关于世界的东西。
05:20
So I think that's kind of this notion of an implicit world model that you did this other,
但我觉得那边还有另外两条线很有意思。
05:25
like, maybe you could be doing any number of tasks, but there's certain kinds of answers
我觉得 world model 这个词本身其实可以追溯到 reinforcement learning 的文献。
05:30
that you could give that would require implicitly modeling something about the world.
而在那里有一个具体的技术定义,和 POMDPs 相关,也许我们接下来会讲到。
05:35
But then I think there's like two other threads there that are really interesting.
但我觉得还有另外两条线索非常有趣。
05:38
I think the term itself world model actually goes back to reinforcement learning literature.
我认为“世界模型”这个词本身实际上源自强化学习的文献。
05:42
And there there's a specific technical definition related to poem DPs that maybe we'll get
而且这里有一个跟poem DPs相关的特定技术定义,也许我们后面会讲到
05:46
into later, but there's a specific technical addition term of world modeling that goes
到后面,但有一个具体的附加技术术语叫 world modeling,它在 reinforcement learning 中已经有相当长一段时间了。
05:51
back to reinforcement learning for quite some time.
然后还有另一条线索,我觉得,最后就变成,你知道,过去几年我们一直在聊 generative models,image model 是生成图像的模型,video model 是生成视频的模型,然后或许 world model 就应该是生成世界的模型。
05:54
And then there's another thread, I think, that ended up is like, you know, we've been
那么,产生或生成一个世界又意味着什么?
05:58
talking a lot about generative models the last couple of years, and an image model is
所以我觉得我们有大概三种不同的线索,不同的人在
06:02
a model that produces images, a video model is a model that produces videos, and then
一个生成图像的模型,视频模型就是生成视频的模型,然后
06:06
maybe a world model should be a model that produces worlds.
也许世界模型应该是一个生成世界的模型。
06:09
And then what does it mean to produce or generate a world?
那生成或创造一个世界又是什么意思呢?
06:12
So I think we got these like three kind of different threads that different people in
所以我觉得我们有三条不同的线索,不同的人
06:15
the community use is around like some notion of implicit world modeling that I can answer
社区里用的概念大概是某种 implicit world modeling,就是我能回答特别难的问题。
06:19
really hard problems.
但我必须对世界的某些方面建模才能答对,或者还有那种特定的,你懂的,RL formulation of world model。
06:20
But I must be modeling something about a world to answer the right answer, or like there's
然后还有一种,就是像 generative model 那样创建或生成世界。
06:25
the specific, you know, RL formulation of world model.
你觉得这种 implicit world model 是,你知道,只是另一种定义,还是像一种错误的信念集合?当你听到这个说法时,你,怎么说呢,你觉得那些算是 world models 吗?
06:29
And then there's the there's like a generative model that creates or generates worlds.
然后还有一种,就是那种生成式模型,它创造或生成世界。
06:33
Do you see this implicit world model as, you know, just another definition or like an
你把这个隐式世界模型看作,你知道,只是另一个定义,还是说
06:40
incorrect set of beliefs, like, when you hear that, do you, you know, do you feel like
一套错误的信念?当你听到这个说法的时候,你,你知道,你会不会觉得……
06:45
those are world models?
那些是世界模型吗?
06:46
You think that's a valid way of thinking about world models, or do you think, you know,
你觉得那是思考 world models 的有效方式,还是你觉得,你知道,
06:50
world models have properties that aren't really characteristic of, you know, the current
world models 有一些属性,并不真正具备,你知道,当前的
06:58
models that we are talking about, you know, vision, language, etc.
我们正在讨论的 models 的特征,你知道,vision、language 等等。
07:02
I almost think like, I think they're all valid.
我几乎觉得,我觉得它们都是有效的。
07:04
I mean, the hard part is that I think all three of the things I just said are very
我的意思是,难点在于,我觉得我刚才说的三样东西都非常
07:08
interesting systems.
有意思的 systems。
07:09
Like they're very interesting models, like they all have, like we should probably build
就像它们都是非常有意思的 models,就像它们都有,我们也许应该
07:13
all of them as a community.
作为一个社区,把所有这些都构建出来。
07:15
But we should probably come up with better terms so that we don't confuse each other by
但我们应该想出更好的术语,免得我们用同一个词指代不同的系统,造成混淆。
07:18
calling different systems by the same term.
我觉得这就是问题所在。
07:21
And I think that's the, that's the thing.
所以我想,过去一年的学术文献里,world model 这个词更多凝聚成了一种特定的——呃,或者说 real-time interactive video model 那种东西。
07:23
So I think like maybe in like the academic literature in the last year, like the world
而最近一年左右的文献里,人们通常管这个叫 world model。
07:28
model has more coalesced into a particular flavor of a, of a, of a, or like real time interactive
模型已经更多地融合成了一种特定的风格,一种,一种,或者说类似实时交互式的
07:33
video model.
视频模型。
07:34
And that's like been more commonly what people call world models in the literature of the
而这在过去一两年左右的文献里,更像是人们通常所说的世界模型。
07:37
last couple last year or so.
我觉得围绕AI的“圣杯”式问题,有一部分是,怎么去——
07:39
But I think all of the other properties that we just talked about are really interesting
但我觉得我们刚才讨论的所有其他特性都非常有趣
07:43
and useful.
而且很有用。
07:44
And like especially the implicit world model notion, I think that that sort of can be applied
尤其是 implicit world model 这个概念,我觉得它在某种程度上可以应用
07:48
to anything.
到任何事情上。
07:49
And like no matter what kind of beta you're processing, no matter what kind of system
就像无论你处理的是什么样的 beta,无论什么类型的 system
07:53
you're building, I think if it gets to some level of interesting complexity, then it's
你正在构建的 system,我认为如果它达到了某种有趣的复杂度,那么它就会
07:57
going to end up with some notion of implicit world model somewhere in the system.
最终在 system 的某个地方产生某种 implicit world model 的概念。
08:00
And that's, that's really interesting.
而这一点,真的非常有趣。
08:02
One of the ways that, that occurs to me is like thinking about it in the context of,
我想起来的一个方式是,就是在这个背景下去思考它,比如,
08:08
you know, the question that we grappled with around, or even so grapple with around,
你知道,我们曾经纠结过的问题,或者说甚至现在还在纠结的,
08:13
large language models is like do, do they really understand language?
关于large language models,就是它们真的理解语言吗?
08:16
Do they, you know, is next token prediction?
它们,你知道,是不是只是next token prediction?
08:19
Is it just faking an understanding of language or is it an understanding of language?
它只是在假装理解语言,还是真的理解语言?
08:24
And I think, you know, for the most part, maybe we've moved off from that question and said
而我觉得,你知道,在很大程度上,可能我们已经不再纠结这个问题了,而是说,
08:29
these things are so amazing, like does it really matter?
这些东西太厉害了,所以真的重要吗?
08:33
And so from that perspective, like if we apply that to the world model question, like
所以从那个角度看,如果我们把这种思维应用到world model的问题上,就像——
08:38
if, you know, those types of models get so good at, you know, generating consistent to
如果说,你知道,那些类型的模型变得非常擅长,嗯,生成与
08:46
some underlying world results, maybe it doesn't really matter, but it, you know, I still find
某些底层世界结果保持一致,也许这并不重要,但你知道,我仍然觉得
08:53
it an interesting, you know, if, if nothing else, like philosophical question.
这挺有意思的,你知道,就算没有别的,也像一个哲学问题。
08:59
I agree.
我同意。
09:00
There's an interesting philosophical question in there, but at some point, it's sort of
这里面有一个有趣的哲学问题,但到了某个时候,这有点
09:03
unanswerable, right?
无法回答,对吧?
09:04
Like as a scientist, you want to be like, at least for me, I like to think about what
就像作为科学家,你会想要,至少对我来说,我喜欢思考
09:07
are things that I can measure about this system?
关于这个系统,我能测量哪些东西?
09:09
What are like concrete questions that I can ask or falsify, ideally, about a system?
比如说,理想情况下,关于一个系统,有哪些具体的问题我可以问,或者可以证伪呢?
09:14
So like, but I think there, but I think the, another thing you're getting at is there
所以,就是,但我觉得……呃,但我觉得,你想说的另一件事是
09:18
is like another notion that I think some people sometimes have when talking about world
另一个类似的概念,我觉得有些人在谈论 world models 的时候有时候会有
09:22
models, which is pretty different.
这非常不一样。
09:25
And I don't think we know how to get there.
而且我觉得我们并不知道怎么达到那一步。
09:27
And that's more like world model as theory builder, right?
而那更像是 world model 作为理论构建者,对吧?
09:30
Because we have this notion that as humans, we're kind of like traversing this really
因为我们有这样一个想法:作为人类,我们有点像在穿越一个非常
09:34
complicated world all around us and having these experiences, but we're not just like
复杂的世界,就在我们周围,并且经历着这些体验,但我们不仅仅像是
09:38
letting the, letting the photons fall on our retinas and like letting it happen.
就是让,让光子落在我们的视网膜上,然后就像,就让它自然发生。
09:42
Like we're constantly building theories in our mind for what's happening, right?
就像我们脑子里一直在为正在发生的事构建理论,对吧?
09:46
We come up with these great theories about how does gravity work and how does physics work
我们会想出这些很棒的理论,关于引力是怎么运作的,物理是怎么运作的,
09:49
and how does fluid dynamics work?
还有流体动力学是怎么运作的?
09:51
And it's not just that we observe these things.
而且我们不只是观察这些现象。
09:53
We actually end up with these very compact and very powerful theories about that explain
我们实际上最终得到了一些非常简洁又非常强大的理论,它们解释了外面正在发生的一切背后的机制。
09:58
all the mechanisms behind what's what's happening out there.
而且我觉得,部分的,部分的,就像围绕 AI 的那个“圣杯”式的问题,就是如何——
10:01
And I think part of the, part of the like the Holy Grail question around AI is how do
而且我不觉得LLM是个糟糕的例子,对吧?
10:05
we get machines to do that kind of a thing, too?
我们也让机器去做那种事情吗?
10:08
And maybe there's this notion of like, well, a world model should not just be directly
而且可能有一种观念,比如说,嗯,world model 不应该只是直接
10:12
thinking about observations, but should be building deep explanatory theories about
思考 observations,而是应该构建关于
10:16
the world.
这个世界的深层解释性理论。
10:18
And I think that one's really, really hard.
而且我觉得那真的、真的很难。
10:20
And I don't know that anyone has a great angle on how to get there.
而且我不确定有没有人对怎么做到那一步有什么好的角度。
10:24
But I think that's kind of a like a slightly different notion of the question that sometimes
但我觉得这有点像是对这个问题的一种稍微不同的理解,有时候
10:29
is mixed in there.
会被混在里面。
10:30
Yeah, you could argue that even before worlds, it would be great if the, you know, if we
是啊,你可能会争辩说,甚至在这些 worlds 之前,如果能,你知道,如果我们
10:34
can get, for example, a model that can build a deep theory about language, you know,
能拿到,比如说,一个能构建关于语言的深层理论的模型,你知道,
10:39
for example, or, you know, any other domain that is currently handled by the types of
比如说,或者,你知道,任何其他目前由我们当前处理的这些类型的
10:45
models we deal with currently, you know, language, you know, graphic arts, yeah, interesting,
模型所覆盖的领域,你知道,语言,你知道,图形艺术,是的,有趣,
10:53
interesting.
有趣。
10:54
But there's something like, you know, what if I had this like perfect, I mean, an
但有一种东西是,你知道,如果我有一个这样完美的,我是说,一个
10:58
L.M.
L.M.
10:59
kind of like a perfect L.M.
有点像是一个完美的 L.M.
11:00
It would maybe L.M.
这也许就是 L.M.
11:01
is the wrong example, but like, suppose you had a perfect video model that could generate
是个错误的例子,但比如说,假设你有一个完美的 video model,它可以生成
11:04
any video you asked for.
任何你要求的视频。
11:06
But like maybe what I wanted wasn't actually a video.
但可能我真正想要的其实并不是一个视频。
11:09
What I wanted as a human as a scientist was I wanted to learn something about the world.
作为一个人,作为一个科学家,我想要的是了解世界。
11:15
And like even if I can generate video of any kind or like of any, of any structure, like
而且即使我能生成任何种类、任何结构的视频,
11:19
I didn't learn what I wanted about the world, then maybe the video I want is like Einstein
如果我没能学到我想知道的世界,那我想要的视频可能就是像 Einstein
11:24
giving a lecture explaining like the new theory of quantum gravity or something like that.
在做一个讲座,解释像 quantum gravity 这种新理论,或者类似的东西。
11:28
And then it's like not actually the pixels themselves, not the capability to generate
然后其实我真正想要的不是pixels本身,也不是生成video的能力。
11:31
video that was what I really wanted.
我是想以某种方式从这个model中获得一些对世界的新理解。
11:33
I wanted to gain some new understanding of the world from from this model somehow.
这确实是个很难的问题。
11:38
And that's that's a really hard one.
对。
11:39
Yeah.
而且我觉得LLM并不是个坏的例子,对吧?
11:40
And I don't think L.L.M.
而且我觉得LLM并不是……
11:41
is a bad example, right?
这是个反面例子,对吧?
11:42
If an L.L.M.
如果是一个 L.L.M.
11:43
had the ability to generate theories, you could argue that it wouldn't hallucinate because
如果它有生成理论的能力,你可以说它不会 hallucinate,因为
11:48
it would think more deeply about the relationship between the things that it's generating and
它会更深入地思考它生成的内容之间的关系,并且
11:54
would kind of self correct or could self correct.
会某种程度上自我修正,或者说能够 self-correct。
11:59
So we talked about kind of implicit world models.
所以我们聊到了某种 implicit world models。
12:05
We talked about generative world models.
我们聊了 generative world models。
12:09
There's this state machine interpretation that you mentioned, Palm D.P. partially observable
你提到的那个 state machine 解释,Palm D.P. partially observable
12:15
Markov decision processes.
Markov decision processes。
12:17
Talk a little bit about that and how that history kind of plays into the way you and others
稍微聊聊那个,以及那段历史如何影响了你和其他人
12:25
are thinking about world models.
在思考 world models。
12:26
So there's this there's a abstraction called partially observed Markov decision processes
所以有一种叫做 partially observed Markov decision processes 的抽象概念,
12:30
or poem D.P.s that goes back quite a long time that is a really nice mathematical formalism
或者叫 POMDPs,它的历史相当悠久,是一个非常棒的数学形式化框架,
12:35
for thinking about how agents can interact with worlds.
用来思考 agent 如何与世界互动。
12:39
So then there's basically like two, you imagine like you basically decompose your system
那么基本上就是说,你可以想象你基本上把自己的系统分解开,
12:43
into two parts.
分成两个部分。
12:44
One is the world and that's like everything that happens around you.
一部分是 world,也就是你周围发生的一切。
12:48
And then there's an agent and an agent is something that can take actions in the world.
然后还有一个 agent,agent 是能够在世界中采取行动的东西。
12:52
They can move around.
它们可以四处移动。
12:53
They can maybe pick up objects.
它们也许能拿起物体。
12:54
They can do things in the world or to the world.
它们可以在 world 里做事,或者对 world 做事。
12:58
So you've sort of partitioned the whole universe into agent which moves around it does stuff
所以你就有点像是把整个宇宙分成了 agent,它会四处移动、做一些事情。
13:02
and then world which has stuff done to it or by the agent and also maybe evolves in time.
然后还有 world,它会被施加作用,或者由 agent 施加作用,而且也可能随时间演化。
13:07
So you've got the world and the agent.
所以你就有了 world 和 agent。
13:09
And then there's a you talk about this formalism of how do the two interact with each other.
然后呢,你讲到一种 formalism,就是这两者如何相互作用。
13:13
Then the agent is going to take actions and the different situations where your vocabulary
然后 agent 要采取行动,以及不同的情境,你的 vocabulary...
13:18
of actions might be different, right?
actions 的种类可能会不同,对吧?
13:20
Like maybe you're a robot and I can actuate my motors in a certain way.
比如你可能是一个机器人,而我可以以某种方式驱动我的马达。
13:24
Maybe I'm in a video game and I can push buttons on the controller.
或者我在一个视频游戏里,我可以按控制器上的按钮。
13:28
So in different situations, the agent might have available to them different kinds of actions.
所以在不同的情境下,agent 可用的 actions 种类会不同。
13:32
But whatever the situation is, when an agent makes actions on the world, the world will
但无论情况如何,当 agent 对世界执行 actions 时,世界就会
13:37
be changed in some way.
以某种方式被改变。
13:40
And the way we to note that is we say that the world has a state internal to it.
我们用来解释这一点的方式是:我们说世界内部有一个 state。
13:44
And the state kind of defines everything that makes the world what it is.
而这个 state 在某种程度上定义了构成世界本身的一切。
13:48
And the state might be very large and very complex.
而且这个 state 可能非常大,非常复杂。
13:50
It might not be understandable, it might be very very large and high dimensional, but
它可能无法被理解,可能非常非常大,而且是 high-dimensional,但它可以说是世界上正在发生的一切的完整解释。
13:54
it kind of is the full explanation of what's happening in the world.
所以 agent 会对世界采取行动。
13:58
So then the agent makes actions on the world.
然后因为 action 改变了世界,这意味着 action 会导致 state 在世界内部以某种方式发生变化或转移。
14:01
And then because the action changes the world, that means the action will cause the state
但现在这个 state 太大了,就像是非常复杂的东西。
14:05
to change or transition sometimes somehow inside the world.
也许它是那种,你知道的,你无法直接观察的东西。
14:09
But now the state is so big like it's something very complicated.
但现在状态空间非常大,像是非常复杂的东西。
14:12
Maybe it's, you know, you can't directly observe it.
也许,你知道,你无法直接观察它。
14:14
So then what the agent gets back are observations.
所以呢,agent 拿回来的就是 observations。
14:16
And then observations are somehow some kind of low dimensional projection of the
然后 observations 在某种程度上就是 full world state 的一种 low dimensional projection。
14:20
full world state.
而且,这在不同的 contexts 里含义也不同。
14:22
And again, what that means is different in different contexts.
也许作为人类,我们得到的 observations 就是眼睛看到的图像、耳朵听到的声音、以及身体感受到的触觉。
14:24
Maybe as a human, the observations we get are the images that we see in our eyes, the
这些都是我们收到的 sensory signals。
14:29
sounds that come into our ears, the feelings of touch that we feel on our bodies.
而这些 sensory signals 告诉我们一些关于世界的信息,但它们只告诉我们一些...
14:33
Those are all the sensory signals that we get.
那些都是我们获得的感觉信号。
14:36
And those sensory signals tell us something about the world, but they only tell us something
这些感觉信号告诉我们一些关于世界的东西,但它们只告诉我们一些事情
14:39
very sparse and local about everything that's happening in the world around us.
关于我们周围世界发生的一切,信息都非常稀疏且局部。
14:43
So then the poem D.P. loop is that, you know, you have an agent does actions to the world
所以这个 POMDP loop 就是,你知道,有个 agent 对世界采取行动,
14:48
that causes the state to transition.
这会导致 state 发生 transition。
14:50
Then based on the state, the agent gets an observation that tells it something about
然后基于这个 state,agent 会得到一个 observation,这个 observation 告诉它关于
14:53
the world.
世界的信息。
14:54
And then the agent is going to, this is going to happen in a loop over and over again as
然后 agent 会——这个会在一个 loop 里一遍又一遍地发生,随着
14:57
the agent tries to do things in the world.
agent 试图在世界中做事。
14:59
That sounds a lot like the setting for reinforcement learning, the agent's operating in the world.
这听起来很像 reinforcement learning 的设置,就是 agent 在世界中运作。
15:06
It's making observations.
它在进行观察。
15:07
There's some reward associated with its actions, et cetera.
它的行为会关联某种 reward,等等。
15:13
Is the implication then that you need a reinforcement learning type setup in order to have a robust
那么,这是否意味着你需要一个 reinforcement learning 类型的设置才能拥有一个稳健的 world model?还是说,P.O. Palm D.P.'s 和扭曲模型之间的具体关系是什么?
15:22
world model or what's like the concrete relationship between P.O. Palm D.P.'s and warped models?
你问到我了。
15:27
You got me there.
所以,那个重要的技术点——你懂的——我在最初的定义里漏掉的就是 reward,对吧?
15:28
So like the important, you know, technical piece that I left out of the initial definition
所以就像是在更……这个……所有这些 formalism 都是在……的背景下发展起来的。
15:31
was the reward, right?
那就是奖励,对吧?
15:33
So like in sort of the more, like this, those all formalism was developed in the context
所以就像在某种程度上,比如这个,所有这些形式化方法都是在……背景下发展起来的
15:37
of reinforcement learning.
关于 reinforcement learning 的。
15:39
And there, it's like, well, the agent has some goal that they're trying to achieve.
在那里,就是说,agent 有某个它试图达成的目标。
15:42
And how do they get signal about whether they're achieving the goal or not?
那它们怎么得到 signal,来判断自己有没有达成目标呢?
15:45
Then they get some reward signal.
然后它们会得到一些 reward signal。
15:47
So then the idea is like, well, the agent wants to try to take actions that will maximize
所以接下来,想法就是,agent 想要尝试采取一些 actions 来最大化
15:51
its reward.
它的 reward。
15:52
And once you go to that level, like that's exactly the reinforcement learning setup.
而且一旦你到了那个层面,就会发现这完全就是 reinforcement learning 的框架。
15:56
And you know, that's where this formalism comes from.
而且你知道,这个 formalism 就是从这儿来的。
15:59
But the reason I chose not to talk about the reward a moment ago is because I think
但刚才我选择不谈论reward的原因是,我认为
16:03
we can take that original abstraction of the P.O.M. D.P. and then pull it out and use
我们可以把P.O.M. D.P.的那个原始抽象提取出来,然后用在
16:06
in other contexts.
其他场景中。
16:08
So I think this notion about thinking about agents and states and observations, that becomes
所以我认为,这种关于agent、state和observation的思考方式,变得
16:12
applicable and useful in a lot of other contexts, even outside of reinforcement learning specifically.
在很多其他场景中都非常适用和有用,即使是超出reinforcement learning的范围。
16:17
And reinforcement learning is maybe one setting in which this formalism was originally developed
而reinforcement learning也许只是这个formalism最初被开发出来的一个场景。
16:22
and is super useful, but we can apply it elsewhere as well nowadays.
它在那个场景中非常有用,但如今我们也可以把它应用到其他地方。
16:25
Give us an example of how it's applied out of that context.
那就举个例子,说说它在那个场景之外是怎么应用的。
16:30
So one example, one like pretty concrete example might be behavior cloning in robotics,
所以有一个例子,一个非常具体的例子,就是 robotics 里的 behavior cloning,
16:34
right?
对吧?
16:35
So you're building a robot and your goal of end of the day is I want to build this robot
所以你在造一个机器人,归根结底,你的目标就是我想造出这样一个机器人
16:38
that's going to go around the world and do stuff and maybe make my bed for me or clean
它能在现实世界里到处转,做各种事情,也许帮我铺床,或者打扫
16:42
the kitchen or whatever it is.
厨房什么的。
16:44
Then once that robot is out there, it's an agent, it's interacting in a world, the world
然后一旦这个机器人部署出去,它就是一个 agent,它在和世界互动,世界
16:49
has state, the robot is getting observations about the world.
是有 state 的,机器人会不断得到关于世界的 observations。
16:52
So it's kind of operating on that P.O.M. D.P. loop once it's trained.
所以训练好之后,它基本上就是在跑那个 POMDP loop。
16:55
But there's a question of like, what was the training signal?
但有个问题是,training signal 到底是什么?
16:59
And you could have had that thing trained via reinforcement learning, where like every
你也可以用 reinforcement learning 来训练它,比如每当
17:03
time it got a reward, like every time it took an action, it got a reward, or you could
它获得 reward 的时候,也就是每次采取 action 得到 reward,或者你也可以
17:07
train it via behavior cloning, right?
用 behavior cloning 来训练,对吧?
17:09
Like maybe I had a large supervised data set of like in this data set, when you receive
比如说,我有一大堆 supervised data set,里面有这样的数据:当你收到
17:13
this observation, you should take this action.
这个 observation,就应该采取这个 action。
17:17
And then you could train a supervised learning model that was be very different.
然后你可以训练一个 supervised learning model,那就会非常不同。
17:20
That wouldn't have an explicit reward signal.
那样就不会有明确的 reward signal。
17:22
You would use like gradient descent and some supervised loss.
你会用像 gradient descent 和某个 supervised loss 这样的东西。
17:25
So then even though you end up with a system at the end of the day that kind of operates
所以即使到头来你得到的系统是在这种 POMDP 类似的循环里运作的,你仍然可以有一个 training objective,是 peer supervised learning,不需要 reinforcement learning。
17:28
in this P.O.M. D.P. like loop, you could have a training objective that was peer supervised
所以总结就是,POMDP 有两个部分。一个部分描述了 agent 和世界之间的关系,也就是那个想法:agent 可以观察到反映某种抽象状态的东西,但它永远无法真正知道那个抽象状态。
17:32
learning that didn't require reinforcement learning.
而它需要基于它的 observations 和 reward 来运作。
17:34
So the summary is that there's two parts of this, two parts of P.O.M. D.P. one kind of captures
所以总结是,这里有两部分,P.O.M.D.P. 的两个部分,一种捕捉……
17:42
the relationship between an agent and the world and the idea that the agent can observe
智能体与世界之间的关系,以及智能体能观察
17:49
things that reflect some abstract state, but it can never really know that abstract state.
反映某种抽象状态的事物,但它永远无法真正知道那种抽象状态。
17:55
And it needs to operate on its observations and the reward.
它需要基于自身的观察和奖励来运作。
18:06
And training is all about how it translates those observations to actions, but that doesn't
而 training 讲的就是怎么把那些 observations 转化成 actions,但这并不一定要围绕 reward maximization 本身。
18:12
necessarily need to be about reward maximization per se.
它也可以是关于别的事情,确实。
18:15
It could be about other things, exactly.
而这之所以和 world modeling 有关,是因为这就是这个词最初的来源,对吧?
18:17
And then the reason why this is connected to world modeling is because this is where
比如说,world model 最初的技术定义大概是,如果我们处于 P.O.M. D.P. 这个设定里,那么 world model 就是输入 world state,输入 agent 采取的 action,然后预测下一个 world state 是什么。
18:21
the term originally comes from, right?
这个术语最初是从哪里来的,对吧?
18:23
Like the kind of original technical definition of a world model is then if we're in this
就像世界模型的原始技术定义是,如果我们处在
18:27
P.O.M. D.P. setting, then a world model is something that inputs the world state,
部分可观测马尔可夫决策过程(POMDP)环境中,那么世界模型就是一个输入世界状态、
18:32
inputs the action, the agent takes, and then predicts what's the next world state.
输入智能体采取的行动,然后预测下一个世界状态的模型。
18:37
So then that's the kind of the original technical setting where the term world model was
所以,那其实就是最初使用 "world model" 这个术语的技术背景。
18:41
used.
而且,带着这个想法,我也不太确定这算不算一个更广的背景,或者只是有历史意义,而不一定关乎我们今天 world models 的发展方向。
18:42
And I don't know how with that in mind that this is in a sense broader context or of
不过,我问的问题并不是关于 state 的。
18:51
historical interest, and not necessarily about where we're going with world models today.
很多时候,我们听到 state 被说成是一种表示,几乎就像 observation 之于 state 那样,而且在 P.O.M.D.P. 这种形式化框架里,state 往往是某个 ground truth 的低维表示。
18:55
But hey, I didn't have a question about state.
但嘿,我没有关于状态的问题。
18:58
A lot of times we hear state talked about as a representation of almost like the observation
很多时候我们听到state被讨论成一种几乎像是observation对state的关系,然后就是P.O.M. D.P.那套形式化框架,通常state是某个其他东西——也就是ground truth——的低维表示。
19:06
is to the state and kind of this formalism of P.O.M. D.P. often state is some lower
关于这个有什么问题吗?
19:17
dimensional representation of some other thing that is the ground truth.
比如说坐在红绿灯前面,observation可以被简化成一维的红绿黄,或者二维的,随便怎样。
19:26
Is that a distinction that is useful for us to explore, do you think?
你觉得这个区分值得我们探讨吗?
19:29
Maybe a little bit.
可能有一点吧。
19:31
I think there's actually maybe two different notions of state that people talk about
我觉得实际上人们说的“state”可能有两种不同的概念,
19:34
sometimes that are getting conflated, so that might be useful to try to unpack.
有时候会被混为一谈,所以试着拆解一下可能会有用。
19:38
One is the notion, like I think in my mind when I say state, I'm thinking about sort
一个概念是,比如我脑海里说state的时候,我想到的是
19:41
of the ground truth state of the world.
世界的ground truth状态。
19:43
It's like the complete description of everything in the world that would be required to answer
就是关于世界中一切事物的完整描述,要回答关于它的任何问题都需要这个。
19:47
any question about it.
而state,按照P.O.M. D.P.的定义,是这个世界上的每一个原子。
19:48
This is another way to come at it or a question that popped up from me earlier that I'll surface
这是另一个切入角度,或者说是之前我脑子里冒出来的一个问题,现在拿出来说一下。
19:53
now.
你知道,刚才你在讲 loop 里的 state 的时候,我在想 state 和 observation 之间关系的一个极端例子:比如你坐在红绿灯前,observation 可以简化成红黄绿这种一维的,或者二维的,随便啦。
19:54
You know, when you were talking about state earlier in the context of the loop, I was
而 state,按照 POMDP 的定义,差不多就是世界上的每一个原子。
19:59
thinking like an extreme example of the relationship between state and observation is like you're
我的问题其实是,当你说 world models 的时候,state 是否必然
20:09
sitting at a traffic light, the observation can be boiled down into like a one-dimensional
我的问题其实是,当你们在说world models的时候,state是不是必然
20:14
red-green yellow or two-dimensional, whatever.
红绿黄,或者二维的,随便吧。
20:25
Whereas the state and by the definition of the P.O.M. D.P. is like every atom in the world.
而根据P.O.M. D.P.的定义,国家就像世界上的每一个原子。
20:35
My question is really when you're talking about world models, does the state necessarily
我的问题是,当你在谈论世界模型的时候,这个状态是否必然
20:43
consist of every atom in the world, or is it just a representation?
它是由世界上每一个 atom 组成的,还是仅仅是一个 representation?
20:51
Maybe that's another way of coming at this question.
也许那是理解这个问题的另一种方式。
20:53
I think so.
我觉得是这样。
20:54
I think it's all about the question of what's the abstraction that makes sense for the
我认为关键在于,什么样的 abstraction 对当前的问题是有意义的,对吧?
20:57
problem at hand, right?
即使是物理学家,也会根据你问的问题使用不同的 representation。
20:58
Even physicists will use different representations depending on the question you're asking.
对于某些问题,我想要这个 quantum system 的 wave function,这算是 state 的最佳版本。
21:03
For some problems, I want the wave function of this quantum system, and that's kind of
对于某些问题,我想要这个量子系统的波函数,那差不多就是状态的最佳版本了。
21:07
the best version of a state.
或者你是在说一个化学系统,那可能就只是这些离子在某个浓度下的溶液,又或者你在说一个热力学系统,你觉得它是个理想气体,那也足够描述系统的状态了。
21:09
For other systems, maybe it's macroscopic, and quantum effects don't come into play, and
对于其他 system,也许它是 macroscopic 的,quantum effects 不会介入,然后你可以把它看作一个 Newtonian system,state 就是这些 particles 的集合,看它们的 masses 是什么,positions 和 momentums 是什么。
21:13
then you could think of it as a Newtonian system, and the state is now this collection of
或者你是在说 chemical system,那可能就是一些 solutions,里面有这些 ions 在某种 concentration 下,或者你在说 thermodynamical system,你把它想成 idealized gas,这样就足以描述这个 system 的 state 了。
21:16
particles, what are their masses, and what are their positions and momentums.
但它从来不是世界上的每一个 atom。
21:21
Or maybe you're talking about like a chemical system, and then maybe it's just like these
但它从来不是世界上的每一个原子。
21:24
solutions with these ions in this concentration, or maybe you're talking about thermodynamical
去学会拿观测结果来预测某种神经向量,你懂的。
21:28
system, and you think it's an idealized gas, and that's enough to describe the state
系统,而你认为它是理想气体,这就足以描述其状态了。
21:32
of the system.
系统的。
21:33
But it's never every atom in the world.
但并不是世界上的每一个原子都这样。
21:36
No, no, I think it's always, I think even if we're talking about like ground truth state,
不,不,我觉得它总是,我是说即使我们在谈 ground truth state,
21:40
it's always like relative to the abstraction, like the abstraction that's necessary for
它总是相对于 abstraction 而言的,就像那种为了
21:43
solving the problems you care about.
解决你关心的问题所必需的 abstraction。
21:46
But I think there is another notion, which is a learned state, and that's another interesting
但我觉得还有另一个概念,就是 learned state,这也是另一个有意思的
21:50
thing that people are now talking about, is that the ground truth state, like we're probably
人们现在在讨论的事情,即 ground truth state,我们可能
21:54
not going to have direct access to in the physical world, right?
在物理世界中无法直接接触到,对吧?
21:57
Like even if we think about like every atom or like an idealized gas, like in a lot of situations,
就像即使我们考虑每一个原子,或者一种 idealized gas,在很多情况下,
22:02
we're just not going to observe that in the world.
我们在现实世界中就是观察不到那个东西。
22:04
So instead, there's another question of could I have a model which I learn, which is going
所以换个角度,还有一个问题是,我能不能有一个我学出来的 model,它会学着怎么拿 observation 去预测某种 neural vector。这个学到的 vector,虽然是 model 自己学出来的,但它的行为就好像是一个 ground truth state——我可以基于它去预测 action,预测从这个 state 会产生的 observation。我可以想象这个 state 在 action 下会怎么 transition,但它其实是 neural network 学出来的内部 vector representation,并不是物理学上那种 ground truth state,只是行为上很接近 ground truth state。
22:08
to learn to take observations and predict some kind of, you know, neural vector, and that
去学习获取观察结果,然后预测某种神经向量,你懂的,就是这样。
22:13
learned vector, even though it's something learned by the model, it behaves as if it were
学习到的向量,尽管它是模型学到的东西,但它的行为就好像它是
22:17
though, as if it were a ground truth state, where I can, you know, predict action, predict
不过,就像它是一个真实状态,在这里我可以,你知道,预测动作,预测
22:21
observations that would result from that state.
从该状态产生的观测结果。
22:24
I can imagine how that state would transition in response to actions, but it's kind of a learned
我可以想象这个状态如何根据动作进行转移,但这是一种学习到的
22:30
internal, you know, vector representation from a neural network, and it's not like a
内部表示,你知道,来自神经网络的向量表示,它并不像是一个
22:33
ground truth state from physics, but it kind of behaves like that ground truth state.
来自物理学的真实状态,但它的行为有点像那个真实状态。
22:37
Is that analogous to the idea of model based and model free RL?
这是不是和 model based 和 model free RL 的想法类似?
22:41
A little bit.
有一点。
22:42
Yeah.
对。
22:43
So if you're doing model based RL, then you have some explicit model of the state, and
所以如果你在做 model based RL,你有一个关于 state 的显式 model,然后
22:47
that's where you have a world model, right?
那就是你有 world model 的地方,对吧?
22:49
So that's, then I have an explicit part of my RL system that in a model based approach,
所以那就是,在 model based 方法中,我的 RL 系统里有一个显式的部分,
22:54
I have a component that tries to predict the next state based on the current state
我有一个组件,试图根据当前 state 预测下一个 state
22:57
in the action, or in a model free approach, then I don't have that explicit notion, maybe
和 action,或者在 model free 方法中,那我就没有那个显式的概念,也许吧。
23:01
I'm just getting the observations, feeding the observations directly to the model, and
我只是把 observations 直接喂给 model,然后它输出 actions,并没有对 state 做显式的 modeling,但是,你知道,回到我们之前聊过的 implicit world models,也许那个 neural network,如果它足够大、足够复杂,其实是在自己的 neural network weights 里面做某种 state modeling,如果它能在正确的情境下预测出特别好的 actions。
23:04
it outputs actions, and it's, there's no explicit modeling of state, but, you know, maybe
它输出动作,而且,并没有对状态进行显式建模,但是,你知道,也许
23:09
then going back to implicit world models, we talked about before, you know, maybe that
我觉得我们在这里想说明的是 state 这个词本身的模糊性。
23:13
neural network, if it's big and complicated enough, is doing something like state modeling
神经网络,如果它足够大和复杂,确实在做类似状态建模的事情。
23:16
inside its neural network weights, if it's able to predict really good actions in the right
你可以是,嗯,世界的任何一种 representation 或 abstraction。
23:20
circumstances.
情况就是这样。
23:21
I think what we're doing here is illustrating the ambiguity of the term state.
我觉得我们在这里想说明的是“状态”这个词本身的模糊性。
23:26
You can either be, you know, whatever the representation or abstraction of the world
你可以是,你知道的,无论是对世界的某种表征还是抽象。
23:32
is, or in some settings, it could be a surrogate for that that the agent creates and predicts
或者在某些設定下,它可能是 agent 在選擇行動時所創建並用來預測的替代物。
23:43
against in choosing actions.
而我覺得,到底哪一種最終會是正確的做法,現在還沒有定論。
23:46
And I think the jury's still out on which of these is going to be the right approach,
但這很令人興奮,對吧?
23:49
but that's exciting, right?
而這就是為什麼,我覺得現在很多人會被這個領域吸引的原因之一,就是因為 language modeling 讓我們感覺好像已經有了一套相當有效的 best practices,而且你知道,還有東西可以探索,但我們確實有了一套行之有效的 recipe。
23:51
And that's why, that's one of the reasons why I think a lot of people are attracted to
这就是为什么,这也是为什么我觉得现在很多人被吸引到
23:53
this area right now, is that language modeling feels like we kind of have a set of best
这个领域的原因之一,就是语言建模感觉我们好像有一套最佳
23:58
practices that work pretty well, and you know, there's things to be explored, but there's
实践,效果还不错,而且你知道,还有很多东西值得探索,但
24:03
a recipe that we have that works pretty well.
我们已经有一个配方,运行得挺好。
24:05
But if you want to talk about like how do we build agent, how do we model environments,
但如果你想谈论比如说我们如何构建 agent,如何对环境建模,
24:08
how do we model worlds, how do we model agents that interact with worlds, there's a lot of
我们如何对世界建模,如何对与世界交互的 agent 建模,有很多
24:13
kind of basic, you know, technical questions or like basic engineering approaches that, you
基础的、你懂的技术性问题,或者说基础的 engineering approaches,你
24:18
know, you could do it this way, you could do it that way, and we don't really know what's
知道,你可以这样做,也可以那样做,而且我们并不真正知道什么
24:20
going to be the best way.
会是最好的方式。
24:21
And that's a really exciting place to be working in and doing research.
而这确实是一个令人兴奋的工作和研究领域。
24:24
So let's switch gears a little bit and talk a little bit about the graphical side of things.
那么让我们稍微换个话题,聊聊图形方面的事情。
24:29
You know, as you mentioned, when we talk about world models, a lot of that conversation
你知道,正如你提到的,当我们谈论 world models 时,很多这类对话
24:36
is focused on the generation of graphical worlds, and one of the kind of foundational technologies
专注于图形世界的生成,而其中一项基础技术
24:46
in that space is this idea of a Gaussian splat.
在这个领域里,就是 Gaussian splat 这个概念。
24:50
So talk a little bit about, you know, maybe, you know, broadly the way folks approach,
所以请稍微谈谈,嗯,也许,嗯,大致上大家是怎么处理
24:56
you know, this world generation and, you know, the role of Gaussian splats, how that has
嗯,这个世界生成,以及,嗯,Gaussian splats 的作用,它是如何
25:02
evolved over the past few years as well.
在过去几年中演变的。
25:04
I think there's actually a top level fork there that we should call out even before we
我觉得实际上有一个顶层的分叉,我们应该指出来,甚至在我们
25:08
get down that road.
深入之前。
25:09
And that's like, you know, are you doing explicit 3D or are you doing implicit 3D?
那就是,嗯,你是做 explicit 3D 还是 implicit 3D?
25:16
And the Gaussian splats are an example of an explicit 3D approach where I'm going to
而 Gaussian splats 是 explicit 3D approach 的一个例子,我要去
25:20
generate, when I generate my world or model my world, I'm going to have some explicit 3D
生成——当我生成我的世界或者建模我的世界时,我会有一些 explicit 3D
25:24
representation of that world.
对这个世界的 representation。
25:26
And then I'm going to do things with that explicit 3D representation of the world.
然后我会用这个 explicit 3D representation 来做一些事情。
25:29
And one of those things might be render it down to pixels that you see.
其中一件事就是把它 render 成你看到的 pixels。
25:32
There's a different approach, which is sort of implicit.
还有另一种方法,有点像是 implicit 的。
25:35
I'm going to have a model that's directly generating pixels.
我会有一个模型直接生成 pixels。
25:38
Like it's a video model, it generates pixels, it generates observations directly, and
就像一个 video model,它生成 pixels,直接生成 observations,然后
25:42
it never goes through this intermediate explicit 3D representational bottleneck.
它永远不会经过这个中间显式的3D表征瓶颈。
25:47
And I think both angles have developed a lot over the last couple of years.
而且我觉得这两个角度在过去几年里都发展了很多。
25:52
So then, you know, that's an interesting fork in the road even talking at that point.
所以,你知道,即使当时聊到那个点,那也是个有趣的分叉路口。
25:56
It is.
确实。
25:57
And do you, so world labs, for example, take the, well, is it even fair to say that you
那么你,比如World Labs,采用——嗯,甚至可以说你们采用Gaussian splat方法吗?就像在marble以及你们发表和展示的一些工作中那样,你们用了那种方法,但你觉得那是基础性的,还是说那只是你们目前展示出来的东西,而你们对其他的方法持开放态度?
26:04
take the Gaussian splat approach like in marble and some of the work that you, you know,
像在marble里那样采用Gaussian splat的方法,还有你做的那些工作,你知道的。
26:10
publish and demonstrate, you use that approach, but do you feel that that is foundational
发布和展示的时候,你是用那种方法,但你觉得那是根本性的东西,
26:17
or is it just kind of what you have, you know, shown thus far, but you're open to other
还是说那只是你目前展示出来的成果,但你对其他方向也持开放态度?
26:22
ideas?
想法呢?
26:23
Well, I think we've actually done both already, right?
嗯,我觉得我们其实已经两者都做了,对吧?
26:25
So it's true.
所以这是真的。
26:26
We put out a product called marble, which takes the Gaussian splatting approach.
我们推出了一款叫 marble 的产品,它采用的是 Gaussian splatting 的方法。
26:31
And there, like what marble does is, let's use our input and image, or a sequence of images,
而且,marble 做的事情是,比如我们的输入是一张图像,或者一系列图像,
26:36
or a text prompt, and then generates a world where that world is represented as a set
或者一个 text prompt,然后生成一个世界,这个世界被表示为一组
26:40
of Gaussian splats, which is an explicit 3D representation.
Gaussian splats,这是一种 explicit 3D representation。
26:43
And then once you have that, once you have that, that Gaussian splat representation, then
然后一旦你有了那个,一旦你有了那个 Gaussian splat representation,然后...
26:47
you can render it to view nice images of the world, or you could imagine compositing objects
你可以 render 它来查看世界的漂亮 images,或者你可以想象把物体 compositing 进去,或者 export 到 graphic engine 之类的。
26:52
into it or export it to a graphic engine or something like this.
但后来我们有个,我们有个,我们在去年年末发了一篇研究博客叫做 RTFM,它采用了一种非常不同的方法,叫做 real-time frame model。
26:56
But then we had a, we had a, we put out a research blog post late last year called RTFM,
而这和我们的 marble approach 形成对比,因为它更倾向于这种 straight to video、pixels only 的方向。
27:02
which takes a very different approach called the real-time frame model.
没有对 3D 世界的显式 representation。
27:05
And this contrasts with the, with our marble approach in that it kind of goes more this
这个 model,你有一个 real-time 运行的 model,real-time 生成 images 来响应。
27:09
straight to video, pixels only direction.
直接生成视频,纯像素路线。
27:12
There is no explicit representation of the 3D world.
没有对3D世界的显式表征。
27:16
The model, you have a model running in real-time that generates images in real-time in response
这个模型,你有一个实时运行的模型,能实时生成图像来响应——
27:20
to user inputs.
对用户输入。
27:21
So the user, you know, asks to move around, and if the user sees an image of the world
所以用户,你知道,会要求移动视角;如果用户看到世界在移动的画面,比如你,你,你说,你,你说 “move left”,你看到自己 move left 的视频,而这一切都是 real-time 发生的。
27:25
moving, like you, you, you, you say, you, you say move left, you see a video of yourself
但在底层,并没有显式的 3D representation。
27:30
moving left, and this all happens in real-time.
它只是 model 实时生成的一帧帧画面。
27:33
But under the hood, there was no explicit 3D representation.
所以 World Labs 实际上一直在同时推进这两个方向。
27:36
It's just frames being generated from a model in real-time.
我们有 marble,它是一种 explicit Gaussian spot-based world model。
27:39
So World Labs actually has been approaching both directions.
所以World Labs实际上一直在两条方向上都推进。
27:42
We have marble, which is our explicit Gaussian spot-based world model.
我们有Marble,那是我们基于显式高斯泼溅的世界模型。
27:45
And RTFM, we have as a more implicit real-time frame generating world model.
And RTFM,我们视之为一个更隐式的实时帧生成世界模型。
27:50
You know, this raises a question for me that goes back to kind of our, you know, very
你知道,这让我想到一个问题,得追溯回我们之前关于什么是世界模型的那种非常基础的对话。
27:57
foundational conversation about what is a world model.
我在想比如 Meta 最初的 Gaussian spot 工作和演示,以及那些我认为让 Gaussian spots 流行起来的工具,之后它们才成为这个世界模型理念的基础。
28:03
I'm thinking about like Meta's initial Gaussian spot, you know, work and demos, and
你知道,它们让你拍一堆照片,然后它们会把那些照片拼在一起——这里用“拼接”这个词非常不严谨——来创建一个本质上...
28:11
the tools that I think have popularized Gaussian spots before, you know, they became kind
那些工具,我觉得之前就是它们让高斯泼溅流行起来的,你知道,在它们变得有点……
28:17
of foundational to this world model idea.
这是基础性世界观想法的一部分。你知道,他们让你,嗯,拍一堆照片,然后他们会把这些照片“拼”在一起——我用“拼”这个词很随意——来创造一个本质上……
28:19
You know, they allowed you to, you know, take a bunch of pictures, and they would like
但有趣的问题是,那个Gaussian splat点云是从哪来的?
28:23
stitch those pictures together, using the term stitch very loosely to create a essentially
就像是,我有1000张图片。
28:31
a navigable 3D model.
一个可导航的3D模型。
28:33
And even, you know, prior to, I don't think it used Gaussian spots, but like Apple's
而且,你知道,在那之前,我不觉得它用的是 Gaussian spots,但像 Apple 的
28:39
AR kit and those kind of things produced like these modestly navigable 3D, I'm trying
AR kit 这类东西,确实产生了一些适度可导航的 3D,我尽量
28:51
not to call them worlds because the question is like, how do you draw the line between,
不把它们叫作“世界”,因为问题是,你该怎么划清
28:56
you know, that thing in a world?
你知道,那个东西和世界之间的界限?
29:00
I want to say like static versus dynamic, but that doesn't seem right.
我想说类似于静态对动态,但那样好像不对。
29:04
No, I know what you mean.
不,我懂你的意思。
29:06
There is a really interesting distinction here that I think a lot of people get confused
这里有一个非常有趣的区别,我觉得很多人都会搞混。
29:09
about with Gaussian splatting, with Gaussian splatting in particular, right?
就是关于 Gaussian splatting, 尤其是 Gaussian splatting, 对吧?
29:12
Because a Gaussian splat is actually just a particular, it's just a representation.
因为一个 Gaussian splat 实际上就是一种特定的东西, 它就是一种 representation。
29:17
Like a Gaussian splat is just like, it's basically a 3D point cloud.
比如一个 Gaussian splat 就像是, 它基本上就是一个 3D point cloud。
29:20
I have a bunch of points in space and each of those points has a position and a color
我有一堆在空间中的点, 每个点都有一个位置和一个颜色。
29:23
and they get a couple other properties attached to it.
然后它们还会附带一些其他的 properties。
29:26
And if your point cloud is large enough, you can move around in it from the perspective
而且如果你的 point cloud 足够大, 你就可以从一个点的视角在里面移动。
29:31
of a point.
但有趣的问题就像是, 这个 Gaussian splat point cloud 是从哪儿来
29:33
But the interesting question is like, where did that Gaussian splat point cloud come
整个宇宙就像这1000张图片,我在拟合一个Gaussian splat来匹配这些图片。
29:36
from?
从哪里?
29:37
And there, there's like two big divides that people get confused because where we're
而且,那里有两个大的分歧,人们会搞混,因为我们——
29:41
Gaussian splatting originated as a technology is around reconstruction.
Gaussian splatting作为一种技术,其起源是重建。
29:46
So there, the idea is like, I'm going to take a lot of views of a space, like hundreds
所以,这个想法是,我要对一个空间取很多视角,比如几百
29:50
or even thousands of views that cover this one space in extraordinary detail.
甚至上千个视角,以极其详细的方式覆盖这个空间。
29:55
Then I'm going to fit a Gaussian, like a 3D point cloud Gaussian splat representation
然后我要拟合一个Gaussian,类似于3D point cloud Gaussian splat representation,
29:59
that explains the images that I saw.
来解释我所看到的图像。
30:01
Well, put like that, the distinction is obvious.
嗯,这么一说,区别就很明显了。
30:04
Like it's, it's not just world, it's world model.
就是,它不只是世界,而是 world model。
30:06
Exactly.
对。
30:07
There's no world model in there, right?
这里面没有 world model,对吧?
30:08
Like it was a, there was no big powerful model in here that, like, learned on a ton
就像,它没有……这里没有一个强大的模型,像那种学习了海量数据的。这些 optimization-based reconstruction 的方法来做 Gaussian splatting,就像是我有 1000 张图片。
30:12
of data, like these kind of optimization-based reconstruction approaches to Gaussian splatting
整个宇宙就相当于这 1000 张图片,然后我拟合一个 Gaussian splat 去匹配这些图片。
30:17
are like, I've got 1,000 images.
而那告诉你的是,当你从不同角度看它时,它看起来是什么样、是什么颜色。
30:19
The whole universe is like these 1,000 images and I'm fitting a Gaussian splat to match
整个宇宙就像这一千张图像,而我在拟合一个高斯泼溅来匹配它。
30:22
these images.
这些图像。
30:23
There's no generalizable knowledge here.
这里没有什么可泛化的知识。
30:25
That's actually very different from what we're doing in marble.
这其实和我们在 marble 里做的很不一样。
30:28
In marble, we have trained a large, powerful model that's been trained on a lot of data
在 marble 里,我们训练了一个 large, powerful model,它接受了大量各种类型数据的训练,
30:31
of various kinds, and that model happens to output Gaussian splats.
而这个 model 恰好输出 Gaussian splats。
30:36
So that, the marble world model is the model that knows how to model worlds.
所以,marble 的 world model 就是那个知道怎么对世界建模的 model。
30:42
It inputs images, it inputs tax, and it outputs Gaussian splats.
它输入图像,输入 text,然后输出 Gaussian splats。
30:45
That's very different from this situation where I'm just kind of dumbly fit, you know,
这跟那种情况完全不同——我只是在傻傻地拟合,你懂的,
30:49
a point cloud to the 1,000 images, and there's no notion of a model that learned a ton
把一个 point cloud 拟合到那 1,000 张图像上,根本没有“model 学到了大量知识”这回事。
30:53
of data.
的数据。
30:54
We touched on this before we started rolling, but I always thought of Gaussian splatts
我们在开始录制之前提到过这个,但我一直以为 Gaussian splats
30:59
kind of, you know, the broad idea or process, you know, kind of the technical and mathematical
就是,你知道,那种宽泛的概念或过程,你懂的,就是那种技术性和数学上的
31:05
process of creating these rendered point clouds, but it sounds like that has evolved quite
创建这些 rendered point clouds 的过程,但听起来这已经发展得相当
31:15
a bit.
多了。
31:16
And now, like, there are standard file formats and other things, you know, talk a little
而现在,像是有了标准的 file formats 和其他东西,你知道,稍微讲一讲
31:19
bit about kind of that ecosystem and the tooling that is used around Gaussian splats.
围绕 Gaussian splats 的那个 ecosystem 和 tooling。
31:28
Yeah.
对。
31:29
So their Gaussian splats are a pretty particular thing in most of the time.
所以他们的 Gaussian splats 在多数情况下是一种挺特别的东西。
31:32
It's a collection of points.
它就是一堆点的集合。
31:34
Each point has a position in 3D space, which is 3 coordinates XYZ.
每个点在 3D 空间里都有一个位置,也就是 XYZ 三个坐标。
31:38
It has an opacity, which is a number between 0 and 1 that tells you how big it is or
它有一个 opacity,是 0 到 1 之间的一个数,表示它有多大或者有多透明。
31:44
how transparent it is.
通常还会有一个颜色,就是 RGB 值,同样也是三个数字。
31:46
You've got usually a color, which is like a RGB value, which is, again, 3 numbers.
然后你经常会接触到所谓的 spherical harmonics 这个概念。
31:50
And then you'll often have some notion of what are called spherical harmonics.
这个就是告诉你,从不同角度看它的时候,它看起来是什么样、颜色是什么。
31:55
And that tells you, what does it look, what color is it when you look at it from different
这告诉你,从不同角度看它是什么样子的,是什么颜色的。
31:58
positions?
位置?
31:59
Right?
对吧?
32:00
Because you want to model this notion that maybe this point is one color if I look at it
因为你想要建模这样一个概念:同一个点,从下面看是一种颜色,从上面看又是另一种颜色。
32:03
from the bottom and a different color if I look at it from the top.
这就有助于你对反射进行建模,或者,或者,对吧,因为如果我有一个闪亮的表面,
32:07
And that helps you model reflections or, or, right, because then if I have a shiny surface,
比如镜子或光泽表面,那么如果我从一个角度看,我会看到来自顶部反射光线的高光。
32:11
like a mirror or a glossy surface, then if I look at it from one angle, I kind of see
如果我换一个角度看,就会看到不同的颜色。
32:15
a highlight from a light that bounces from the top.
顶部反射的光会产生一个高光。
32:17
If I look at it from a different angle, I see a different color.
如果我从不同角度看,会看到不同的颜色。
32:21
And though these spherical harmonics are a particular concrete way to capture this
而且虽然这些 spherical harmonics 是一种很具体的方式,用来捕捉这种 view-dependent color 的概念。
32:24
notion of view-dependent color.
所以呢,你知道,一个 Gaussian splat 就是一组点。
32:27
So then, you know, a Gaussian splat is this set of points.
每个点都有一个 XYZ position、一个 opacity、一个 color,还有这些 spherical harmonics,用来告诉你从不同角度看的时候颜色怎么变。
32:30
Each one has XYZ position, has an opacity, has a color, and also has these spherical harmonics
然后呢,把这些数据打包成文件里的 bytes 有各种不同的方法,对吧?
32:34
that tell you how the color varies as you look at it from different angles.
那么 PLY 是 Gaussian splats 一个挺流行的文件格式,读写起来相对容易,但是效率很低,然后还有 compressed representations。
32:37
And then there's different ways to pack that data into bytes in a file, right?
然后还有不同的方式把这些数据打包成文件里的字节,对吧?
32:42
So then PLY is a pretty popular file format for Gaussian splats, which is, you know, relatively
所以PLY是Gaussian splats比较流行的文件格式,你知道,读写起来相对容易,但效率挺低的,然后还有压缩表示法来处理其中一部分。
32:47
easy to read and write, but pretty inefficient, and then there's compressed representations
它们能让你存下看起来挺相似的东西,你知道,看起来效果不错,但文件体积小很多。
32:51
like SPZ that, you know, compress some of that data and use lower precision for some
比如SPZ那种,你知道,就是compress一些data,对其中一部分用更低的precision。它们能让你存一些看起来非常相似的东西,你知道,看起来挺好的,但是file size会小很多。然后你应该这么想,你知道,PLY有点像,呃,有点像GIF之于PNG,就是那种非常uncompressed的东西,而SPZ更有点像JPEG,你会compress一部分data,会丢掉一部分data,但我们觉得那些部分都是,你,
32:56
parts of it.
其中的一部分。
32:57
They'll let you store something that looks pretty similar, you know, looks pretty good,
他们会让你存储看起来非常相似的东西,你知道的,看起来相当不错。
33:01
but it takes a lot smaller file size.
但它的文件体积要小得多。
33:04
Then you should think that, you know, like a PLY is kind of like a, like a, like a GIF
那你应该这么想,PLY 有点像 GIF 之于 PNG,就是那种几乎不压缩的格式,而 SPZ 更像 JPEG,你会压缩一部分数据,也会丢掉一部分数据,但我们觉得丢掉的都是那些……
33:08
for a PNG, where it's something that's like pretty uncompressed, and an SPZ is a little
Gaussian splats 不是基于 grid 的,那它们相对于我们表示的空间维度来说,整体上是不是比较 sparse?
33:12
bit more like a JPEG, where you're going to compress some parts of the data, you're
对,我的意思是,它们不是基于 grid 的。
33:15
going to throw away some parts of the data, but we think those are parts that are, you,
这些点可以存在于空间中的任何位置,但通常来说,它们……
33:19
that aren't going to, you're not going to notice as a human that you throw away some
那些不会的,你作为人类不会注意到你丢掉了数据里的某些部分。
33:22
of these parts of the data.
相对于我们说,像传统2D和3D图像那种思考方式,Gaussian splats不是基于grid的,而且相对于我们表示的空间的dimensionality来说,它们通常是sparse的吗?
33:24
Relative to kind of the way we think about like traditional 2D and 3D images, the
是的,我是说,它们不是基于grid的。这些点中的任何一个都可以在空间里任意位置,但通常,我是说,那些……哦,哦,对。
33:30
Gaussian splats are not based on a grid, and are they sparse in general, relative to
高斯splat不是基于网格的,而且一般来说,相对于
33:37
the, you know, the dimensionality of the space we're representing?
你知道,我们正在表征的这个空间的维度,对吧?
33:42
Yeah, I mean, they're not based on a grid.
是的,我的意思是,它们并不是基于网格的。
33:44
Any of these points can live anywhere in space, but they typically, I mean, what are the,
这些点可以分布在空间中的任何位置,但通常来说,我的意思是,什么是——
33:48
oh, oh, right.
哦,哦,对。
33:49
The other important thing I forgot, like the silly me for not having my notes in front
还有一件重要的事我忘了,都怪我没把笔记放在面前,傻了我。但是 Gaussian splat 也是有大小的,对吧?它不是一个无穷小的点。
33:52
of me, but like a Gaussian splat also has a size, right, it's not an infinitesimal point.
Gaussian splat 是有大小的,它不是,它不只是一个空间里的点。
33:56
A Gaussian splat has a size, it's not a, it's not just a point in space.
它有一个,有一个类似 radius 的东西,然后你知道,还有一个 covariance matrix,因为它可能不是球体,可能是个椭球。然后 radius、covariance matrix 以及 radius 它们,嗯,其实真的就只是 covariance matrix 在告诉你它有多大,以及它在不同维度上被拉伸或压扁了多少。
33:59
It has, it has kind of a radius, and you know, also a covariance matrix, because it might
所以呢,你知道,Gaussian splats 最初卖点的一部分可能就是……
34:05
not be a sphere, it might be an ellipse away, and the radius, the covariance matrix and
它不是个球体,可能是个椭圆的形状,半径、协方差矩阵,还有那个半径,其实协方差矩阵就是告诉你它有多大,以及在不同维度上是被拉伸还是被压扁了。
34:09
the radius kind of, well, really just a covariance matrix kind of tells you like how big is
所以呢,Gaussian splats 最初的核心卖点之一,基本上就是——
34:13
it and how like stretched or squashed is it along different dimensions?
嗯,所有东西都是由小三角形组成的,基本上任何游戏都是这样。
34:17
So then like part of the, you know, part of the original pitch of Gaussian splats is maybe
所以现在我在图像里会看到一个剧烈的变化,一个突然不连续的跳变,而这个我是能看到的。
34:21
they could be pretty sparse because, or maybe, you know, maybe they could be fairly sparse
它们可能会非常稀疏,因为,或者你知道,也许它们可以相当稀疏。
34:25
and you end up with like really big Gaussian splats to cover like a big part of the wall.
然后你最终会得到很大的Gaussian splats来覆盖墙壁的很大一部分。
34:29
But in practice, that's usually ends up in pretty low quality.
但实际上,这通常最终会导致质量很差。
34:32
So usually you want the splats to be fairly dense over the geometry that you want to cover.
所以通常你希望splats在你想要覆盖的geometry上相当密集。
34:36
I mean, that ends up in looking, that ends up generating nicer images in general.
我的意思是,这最终会生成更漂亮的图像。
34:41
That being the case, like, is a part of the, the technology or what makes Gaussian splats
既然如此,这是技术的一部分吗,或者是什么让Gaussian splats
34:50
work like that the, the renderer is able to like focus on, you know, what's in kind
能够这样工作——renderer能够专注于,你知道,那些
34:57
of the, the user viewpoint versus things that are extraneous to it or, you know, maybe
用户视角里的东西,而不是与之无关的东西,或者,你知道,也许。
35:03
the broader question is like, um, you know, what makes Gaussian splats so interesting
更广泛的问题就是,嗯,你知道,是什么让 Gaussian splats 这么有趣。
35:10
like a, you know, quicker fresher on that relative to the way we've approached this before.
就是,你知道,相对于我们之前的方式,快速回顾一下。
35:14
I think the, the contrast with Gaussian splats you should be thinking about is triangle
我觉得你该考虑的那个对比,和 Gaussian splats 相对的,就是 triangle meshes。
35:17
meshes.
嗯,所以 triangle meshes 差不多是 computer graphics 里的标准 representation。
35:18
Um, so triangle meshes are kind of your standard representation in computer graphics.
也就是说,基本上我们要把整个世界都表示成一个个小三角形。
35:21
And that's saying that we're going to represent the whole world as like little triangles
嗯,所有东西都是由小三角形组成的,还有,几乎任何游戏...
35:24
basically.
基本上是这样。
35:25
Um, and everything is like made up of little triangles and, and like pretty much any game
嗯,然后所有东西都是由小三角形构成的,而且,基本上任何游戏都是这样。
35:29
you've ever played, like any VFX shot you've ever seen, any computer graph or any computer
你玩过的,比如你看过的任何VFX镜头,任何computer graph或任何computer
35:33
generated image or you've ever seen, um, they pretty much model the world as lots of little
生成的图像,或者你见过的,嗯,他们几乎把世界建模成很多小
35:37
triangles.
三角形。
35:38
Um, and that works.
嗯,那很管用。
35:39
That's works amazing for computer graphics for decades.
这在computer graphics领域几十年来都表现出色。
35:41
Um, but the problem is that triangles don't fit with neural networks very well.
嗯,但问题是,三角形和neural networks不太契合。
35:45
Um, because if you, if you, the important part is differentiability.
嗯,因为如果你,如果你,关键在于differentiability。
35:49
You want to be able to differentiate through this, uh, through this representation past
你希望能够通过这个,呃,通过这个representation进行微分。
35:52
gradients.
gradients.
35:53
And then particular, that means that you want your representation to have the property
然后具体来说,这意味着你希望你的 representation 具有这样一个特性:如果我改变 input 一点点,output 也会改变一点点。
35:56
that if I change the input a little bit, the output also changes a little bit.
嗯,但三角形就不是这样了,因为如果我这里有一个三角形,然后我稍微移动它一下,突然之前看不见的东西现在变得看得见了。
36:00
Um, and that's not the case with a triangle because if I've got a triangle here, and I move
所以现在我会有急剧的 discontinuous 变化,在我将要看到的图像里,而这个图像是 parameters 的 function。
36:03
it a little bit, all of a sudden something that became that was invisible now becomes visible.
嗯,所以,你知道,Gaussian splats 没有这个特性,因为一切都是 smooth 的。
36:08
So now I have a sharp change, a sharp discontinuous change in the image that I'm going to see as
所以现在图像里会出现一个剧烈的、不连续的变化,我接下来会看到这个变化。
36:13
a function of the, of the parameters.
是参数的函数。
36:15
Um, so, you know, Gaussian splats don't have that property because everything is smooth.
嗯,所以,你知道,高斯泼溅没有那个性质,因为一切都是平滑的。
36:20
Everything is partially transparent.
一切都是部分透明的。
36:22
So at like the image that you see is a continued, continuously changes as you vary any of the
所以就像你看到的那个图像,它是一个持续的、连续变化的东西,当你以无穷小的幅度
36:26
parameters of the Gaussian splats infinitesimally.
改变 Gaussian splats 的参数。
36:29
So what that means is they, they integrate with neural networks really well.
所以这意味着它们,它们与 neural networks 集成得非常好。
36:32
So neural networks are all, all our gradient based learners, right?
所以 neural networks 都是,都是我们的 gradient based learners,对吧?
36:34
I'm going to have an objective function.
我要有一个 objective function。
36:36
I'm going to minimize that objective function via gradient descent.
我要通过 gradient descent 来最小化那个 objective function。
36:39
In order to do that, I need to be able to pass gradient signal through whatever representation
为了做到这一点,我需要能够通过任意 representation 传递 gradient signal。
36:42
I'm using and Gaussian splats like past gradients really, really well.
我在用 Gaussian splats,它们真的非常非常像过去的 gradients。
36:46
So that means they can be plugged into gradient base optimizers and either like directly
所以这意味着它们可以插入到 gradient base optimizers 里,要么直接针对一组图像进行优化,也就是 reconstruction 的情况,要么插入到 neural network 的输出中,这更像是我们在 marble 案例里做的。
36:50
optimize against a set of images, which happens in the reconstruction case or be plugged
所以那里的设置是:我有一个 neural network。
36:55
into the output of a neural network, which is more what we do in the marble case.
它会输出 Gaussians,这些 Gaussians 会连接到某个 loss function 上。
36:59
So there the setup is like, I've got a neural network.
然后我可以把 loss 一路 back propagate 到 neural network 的参数里。
37:01
It spits out gouchions, those gouchions get attached to some loss function.
所以这大概就是最大的 delta,嗯,这大概就是 Gaussian splats 的核心创新。
37:05
Then I can back propagate my loss all the way into the parameters of the neural network.
然后我就可以把损失一路反向传播到神经网络的参数中。
37:08
So that's kind of the biggest delta, um, that's kind of the big innovation of Gaussian splats
所以这就是最大的差异,嗯,这是高斯泼溅的大创新。
37:12
and why people got excited about them over the past few years is because it's a, it's a
以及为什么过去几年人们对它们感到兴奋,是因为它是一种,一种
37:16
graphics representation that integrates with neural networks really cleanly.
与 neural networks 集成得非常干净的 graphics representation。
37:19
Kind of tying that to some of the core properties of world models, namely consistency, like
把它和 world models 的一些核心特性联系起来,也就是 consistency,比如
37:28
is it that ability to back propagate that, or how do we get that consistency?
是 back propagate 的能力吗,还是我们如何获得这种 consistency?
37:36
Are we, you know, is that coming from modeling?
我们,你知道,这是来自 modeling 吗?
37:40
Is it coming from, you know, scale, like, where does that, is it an architecture thing?
还是来自,你知道,scale,比如,这到底是从哪里来的,是 architecture 的问题吗?
37:48
What does that come from?
那它来自哪里?
37:49
I think it can come from many places and that's that's actually really interesting jumping
我认为它可以来自很多地方,而这实际上真的很有趣,跳跃
37:52
off point because there's this property of the world around us that it's consistent,
有点跑题了,因为我们周围的世界有一个特性,就是它是一致的,对吧?如果我看着你,你看起来每时每刻都差不多,或者我,你知道,走出去,走到另一个房间再回来,你还在这里,嗯,这就是一致性的概念。而且有不同的,Gaussian splats 在结构上就是一致的,对吧?因为我对世界有一个显式的 3D 表示。所以如果我看它一眼,然后移开视线,再回头看,它就在那里,而且它是 3D 的,所以
37:56
right?
对吧?
37:57
If I, if I look at you, you kind of look pretty similar from moment to moment, or if
如果我看着你,你每一时刻看起来都挺相似的,或者如果
38:00
I, you know, walk in, walk to a Duffer room and come back, like, you're still going
我,你知道,走进来,走到另一个房间再回来,你还是
38:03
to be here, um, and that's this notion of consistency.
会在这里,嗯,这就是一致性的概念。
38:07
And there's different, and gouchions splats are kind of consistent by construction, right?
而且还有不同的——高斯溅射在构造上就有一致性,对吧?
38:10
Because I've got this explicit 3D representation of the world.
因为我有了这个显式的3D世界表示。
38:13
So if I look at it, then look away, then look back, like, it's, it's all there and 3D, so
所以如果我看着它,然后移开视线,再看回来,它就在那儿,而且是3D的,所以
38:17
it's going to look the same.
它看起来会是一样的。
38:19
But you could also get consistency via large scale data and large scale training and large
但你也可以通过大规模数据、大规模训练和大规模 compute
38:22
scale compute, and that kind of leans into the more implicit, you know, representations
来获得一致性,这就比较倾向于那种更隐式的,你懂的,representations。
38:27
that we've done in RTFM and in other places.
就是我们在 RTFM 和其他地方做过的那些。
38:29
So there, the idea is, what if I'm going to have a model, there is no gouchions splats,
所以那里的想法是,如果我有一个模型,没有 Gaussian splats,
38:33
there are no 3D, there are no explicit points, I just have a model that's spitting out RGB
没有 3D,没有显式的点,只有一个模型在输出世界的 RGB
38:38
pixel values of the world.
像素值。
38:40
But if that model is really, really smart and really powerful and having been trained
但如果那个模型真的非常非常聪明、非常强大,而且经过训练
38:43
on a lot of data, and maybe with the right expressive architecture, even though it's
在大量数据上,而且也许有了合适的表达架构,即使它
38:47
not mathematically guaranteed, there's nothing mathematically guaranteeing it to be consistent,
并非数学上得到保证,没有什么在数学上保证它的一致性,
38:52
you know, it still ends up being consistent.
你知道,它最终还是一致的。
38:54
And I think that's, I think it's actually not that one is better than the other, there
而且我认为那是,我认为实际上并不是说一种比另一种好,
38:57
are just different points on the technology curve, right?
只是处于技术曲线的不同点上,对吧?
39:00
If you have a relatively, you know, low compute budget and you want things to run embedded,
如果你的计算预算相对较低,你知道,而且你想要在嵌入式设备上运行,
39:05
like you don't want to train giant models, then gouchions splats are appealing because
比如你不想训练巨型模型,那么 gouchions splats 就很有吸引力,因为
39:09
they're consistent by construction.
它们从构造上就是一致的。
39:11
But if you can scale up and use a lot of data and use a lot of compute, then I actually
但如果你能scale up,用大量的data和大量的compute,那我其实认为implicit 3D这条路就是能scale到无穷的那个东西。
39:14
think that the implicit 3D route is going to be the thing that scales up to infinity.
所以对我来说,这更像是一个,嗯,engineering question——我现在面临的问题的设计约束是什么——而不是一个哲学上的分歧,对吧?
39:19
So it's more a question of like, it's more of an engineering question of what are the
如果你想要便宜且构造上一致的,gouchions splats非常有吸引力。
39:22
design constraints of the problem facing me right now and less a philosophical divide
如果你想要,你知道,某种能scale到无穷的东西,而且,你知道,可以依赖无限的data和无限的compute,但你愿意为那个无限的computed inference买单。
39:26
for me, right?
对我来说,对吧?
39:27
If you want cheap and consistent by construction, gouchions splats are very appealing.
如果你想要便宜且构造上一致的方案,高斯溅射非常吸引人。
39:31
If you want, you know, something that scales to infinity and has, you know, can rely on
如果你想要,你知道,某种能扩展到无限并且有——你知道——可以依赖的
39:35
infinite data, infinite compute, but you're willing to pay that infinite computed inference
无限的数据、无限的计算,但愿意付出那种无限计算的推理代价
39:39
time.
时间。
39:40
And you're willing to pay the big training costs.
而且你愿意支付高昂的 training 成本。
39:42
And I think the implicit pixels only approach is very appealing.
而且我认为只使用 implicit pixels 的方法非常有吸引力。
39:46
And that's exactly why we've done both at World Labs.
这正是我们在 World Labs 两者都做的原因。
39:48
I think the direction I was trying to go with that question was focusing more on the model
我觉得我当时那个问题想引导的方向,是更关注 model
39:54
and the generation perspective and the idea that, you know, we're describing a world,
以及 generation 的视角,还有那个想法,你知道,我们在描述一个世界,
40:02
generating on the fly, the user can navigate through this world, turn away, come back
实时生成,用户可以在这个世界中导航,转身离开,再回来
40:08
and we're generating consistent, well, frames in the case of RTFM or splats in the case
然后我们在生成一致的——嗯——在 RTFM 的情况下是 frames,或者在 splats 的情况下
40:14
of marble.
大理石。
40:18
And the question is like, I think part of what you're saying is that consistency is kind
然后问题是,我觉得你所说的部分内容是,consistency 有点
40:24
of orthogonal to whether we're talking about pixel generation or splat generation.
正交于我们谈论的是 pixel generation 还是 splat generation。
40:30
And so the next part of that question is, you know, that consistency seems like a really
所以这个问题的下一个部分是,你懂的,consistency 似乎是这一切中一个非常
40:35
important part of all this, like, where does it come from?
重要的部分,就是说,它从哪里来呢?
40:38
I mean, I think it basically comes from your data ultimately, right?
我是说,我觉得它最终基本上来自你的 data,对吧?
40:40
Like no matter what representation you're using, if it's if it's raw pixels or it's
比如说,不管你在使用什么 representation,如果它是,如果它是 raw pixels 或者它是
40:45
Gaussian splats, like at the end of the day, you have to have a neural network in the
Gaussian splats,归根结底,你必须有一个 neural network 在
40:49
system that's been taught to create consistent data.
一个被训练来生成一致数据的系统。
40:53
And like Gaussian splats kind of make it easier for a neural network to produce consistent
像 Gaussian splats 这样的东西,会让 neural network 更容易生成一致的
40:57
outputs because they're more consistent by construction.
输出,因为它们在构造上就更一致。
41:00
But ultimately, it has to come from data of like views of the world that are consistent
但归根结底,它必须来自数据,比如那些关于世界的视角,而这些视角是一致的
41:04
in the way you want, you want your model to learn.
以你想要的方式,也就是你希望模型去学习的那种方式。
41:08
I think then the question is or the opportunity is to like dig into marble and talk a little
我觉得问题在于,或者说机会在于,深入去挖一挖 marble,聊一聊
41:13
bit about kind of the recipe and, you know, how that creates a world model or how that
那种方法,你知道的,它怎么创造出 world model,或者怎么
41:18
creates the marble world models.
创造出 marble 的 world models。
41:20
I mean, we have actually haven't talked explicitly about the model, the marble architecture.
我是说,我们其实还没明确聊过那个 model,就是 Marble architecture。
41:25
But at a high level, it lets you input as a user different kinds of things.
但简单来说,它让用户能 input 不同类型的东西。
41:30
You can input a text prompt.
你可以 input 一个 text prompt。
41:31
You can input an image.
你可以 input 一张 image。
41:32
You can input multiple images.
你可以 input 多张 images。
41:34
You can input a video.
你可以 input 一段 video。
41:35
Then from that, we generate a 3D Gaussian splat world.
然后从那些,我们生成一个 3D Gaussian splat world。
41:40
And one intermediate step in that is generating a 360 panorama image.
而其中一个中间步骤就是生成一个 360 panorama image。
41:45
And this is where like given those inputs, you kind of have one model, one part of a model
那么这里就是,比如说给定这些输入,你基本上有一个模型,模型的一个部分。
41:50
that generates a 360 panorama view of the world you're about to generate.
它生成一个360 panorama view,也就是你即将生成的那个世界的全景。
41:55
And that's a that's a big powerful generative model that needs to take whatever that user
那是一个,那是一个很强大的 generative model,它需要把用户输入的任何内容先映射到一个360 panel里面。
42:00
input was and map it first into a 360 panel.
然后从那里,再把它提升成一个完整的3D Gaussian splat 世界。
42:04
And then from there, lift it up into a full 3D Gaussian splat world.
那么这是否意味着,用户的输入要么是单个点,要么是单张图像?
42:07
Is the implication then that the user's input is either a single point or single image?
或者我在想,比如说,如果用户提供的是,嗯,一组空间上分散的图像,那是不是只是球体变得更大?
42:16
Or I guess I'm thinking about like if you if the user's providing, you know, a spatially
或者我在想,比如,如果用户提供了一个——你知道——空间上的
42:26
diverse set of images, is it just that the sphere is much bigger?
多种多样的图像,是不是只是球体更大了?
42:31
Oh, no.
哦,不。
42:32
So basically what we're doing is like the kind of marquee use case is almost single image.
所以基本上我们在做的事情,那种招牌用例几乎就是 single image。
42:36
And that's probably what works best in marble today.
而且这大概是目前 modeling 中效果最好的。
42:39
So there it's like, I'm going to input a single image.
所以在这个场景里,就像我要 input 一张 single image。
42:41
And then for the stuff in the world that I can see in that image, my generated world
然后对于我在那张 image 里能看到的世界里的东西,我生成的世界应该和我 input image 里看到的相匹配。
42:45
should match what I see in my input image.
然后 model 会尝试去补全它,像是为世界里那些在 input image 中看不到的其他东西找一个合理的 completion,对吧?
42:47
And then the model will try to complete what it, like try to get a plausible completion for
然后模型会尝试补全它,像是为输入图像中看不见的世界其余部分生成一个合理的补全,对吧?
42:53
everything else in the world that's not visible in the input image, right?
三个你可以导航的。
42:56
So if I'm like, maybe taking a picture of a blackboard in the front of a classroom,
所以如果比如说,我在教室前面拍一张黑板的照片,
43:01
then the model should know that behind the blackboard is going to be all these chairs
那这个 model 应该知道,黑板后面会是一排排椅子,就是学生们坐的那种椅子。
43:05
where the people sit where the students sit.
然后这个 model 应该能根据一张黑板的照片,补全并生成黑板后面的椅子,然后把整个场景提升到可以导航的 3D。
43:07
And then the model should be able to take a picture of a blackboard and then complete
那用户也会提供 text prompts 吗?
43:09
that and like generate the chairs behind the blackboard and then lift all that into
他们可以。
43:12
three that you can navigate.
用户也会提供文本提示吗?
43:15
And is the user also providing text prompts?
他们,他们可以。
43:17
They they can.
是的,我的意思是,这就是将世界模型纳入其中的美妙之处,你有一个强大的大型生成模型,它在大规模数据上训练过,既有
43:18
Yeah.
嗯。
43:19
They can prompt this directly from text, that's kind of optional.
他们可以直接从文本中进行 prompt,那是可选的。
43:23
If you, you know, if you want to generate your world purely from text, we can do that.
如果你,你知道,如果你想纯粹从文本生成你的 world,我们能做得到。
43:27
If you want to provide text as an auxiliary input to give some extra guidance to what's
如果你想提供文本作为辅助输入,来给正在发生的事情一些额外指导,关于
43:30
happening in your input image, that can work too.
你的输入图像中发生了什么,这也可以。
43:33
But we kind of wanted to take this approach with marble of maximum flexibility for users
但我们有点想通过 marble 采取这种给用户最大灵活性的方法,
43:37
that no matter what kind of signal you got, no matter what kind of edit you do, no
不管你有什么样的信号,不管你做哪种编辑,不管
43:41
matter where you want to take it after it's generated, we want to give you a lot of pathways
你想在生成后把它带到哪里,我们都想给你很多途径。
43:44
to use this thing and not try to like, you know, guide everyone down one rigid pathway
使用这个东西,而不是试图,呃,你知道的,去引导每个人走一条死板的路径。
43:48
for how to generate things or what to use it after, after you generate it.
关于怎么生成东西,或者生成之后用它来干嘛。
43:52
I think where my questions around or my assumption of multi image was coming from like,
我觉得我对 multi image 的疑问或者说假设,是这么来的,比如说,
43:58
I've seen some marble worlds that were kind of your classic, you know, room and you
我见过一些 marble worlds,就是那种很经典的,你知道,一个房间,然后你
44:07
spin around the room and there's like a ton of detail, super impressive.
在房间里转来转去,细节特别多,超级惊艳。
44:11
But then I've seen other ones where they're more like kind of these video game immersive
但我也见过另外一些,更像是那种沉浸式电子游戏
44:15
worlds where you can like navigate, you know, through, you know, fantasy kind of village
世界,你可以在里面到处走,你知道,穿过那种奇幻风格的村庄
44:22
kind of situation.
那种情况。
44:24
And I think are those all kind of generated from a single image generally and potentially
然后我在想,这些是不是通常都是从一个 single image 生成的,可能再加上一些 text conditioning?
44:31
some text conditioning?
是的,我的意思是,这就是让 world model 参与到 loop 里的好处——你有一个非常强大的单独的 large generative model,它是在海量 data 上训练过的,既有 real world data,也有 fantastical data,还有 photorealistic 和 non photorealistic 的内容。所以,在后端你有这么一个强大的 model,无论你带来什么样的 image prompt,不管那个 image 是你家里的一个房间,还是一个奇幻的游戏环境那种,你都会有一个 model 可以...
44:33
Yeah, I mean, that's the beauty of having a world model in the loop there is that you
对,我觉得这就是把world model放进整个流程里的妙处——你有一个非常强大的大型生成模型,在大量数据上训练过,既有真实世界的数据,也有幻想风格的数据,既有photorealistic的内容,也有非photorealistic的内容。所以你在后端有这么一个强大的模型,不管你给它什么样的image prompt,不管是你自己家里的房间,还是那种奇幻的、像电子游戏一样的环境,这个模型都懂所有这些不同类型的世界,并且能在需要的时候按任务要求把它们生成出来。
44:37
have a single powerful large generative model that's been trained on a ton of data, both
拥有一个强大的大型生成模型,它已经在大量数据上完成了训练,包括
44:41
real world data and fantastical data and photorealistic stuff and non photorealistic stuff.
真实世界的数据、奇幻的数据、照片级逼真的内容,以及非照片级逼真的内容。
44:46
So you've got this big powerful model in the back end that no matter what kind of image
所以你在后端有一个非常强大的大模型,不管是什么样的图像
44:50
you're prompt you bring, whether that image is of, you know, you're a room in your house
你带来的提示词,不管是那种画面,比如你家里的某个房间。
44:54
or, you know, a fantastical like video game kind of environment, you've got a model that
或者,你知道的,一个像奇幻电子游戏那样的环境,你有一个模型
44:59
knows all of these different kinds of worlds and can generate them as it's, as it's necessary
了解所有这些不同类型的“世界”,并且能在需要时实时生成它们。
45:03
for the task at hand.
用于当前这个任务。
45:04
And that's, that's, and that's kind of the major difference between, you know, having
这就是,这就是,这就是那种主要的区别,你知道,就是这些生成结果背后有一个世界模型,一个大的世界模型,而不是那种经典的gouching、splatting,只是拟合到我恰好拥有的那一千张图像上。
45:07
these generations backed by a world model, by a big world model versus, you know, classical
关于用来创建这些模型的数据集,还有训练方法和配方,你能说点什么吗?
45:11
gouching, splatting, just fitting to these thousand images that I happen to have.
嗯,我是说,你需要训练大量的数据。
45:15
What can you say about the data sets that are used to create these models and like the
外面有大量的视频,而图像和视频都是,你知道,3D世界的2D投影。
45:20
training approach and recipe?
训练方法和配方?
45:21
Yeah, I mean, you need to train a lot of data.
是啊,我的意思是,你需要训练大量的数据。
45:24
And one thing that's really important is being able to train on a variety of different
而且非常重要的一点是,能够在各种不同类型的 data 上进行训练,对吧?
45:27
kinds of data, right?
因为说到底,世界本身是 3D 的,而且有很多,你懂的,3D 结构。
45:29
Because ultimately the world itself is 3D and like has a lot of, you know, 3D structure.
但是,你懂的,没有太多明确的 3D data 可以让你去学习。
45:35
But there's not a lot of, you know, explicit 3D data for you to learn on.
那是相当少见的一种 data 形式。
45:38
That's the pretty rare form of data.
但是世界上有大量的图片。
45:40
But there's a lot of images out there.
世界上也有很多视频,而图片和视频都是,你懂的,3D 世界的 2D 投影。
45:42
There's a lot of videos out there and images and videos are both, you know, 2D projections
现在网上有很多视频,图像和视频本质上都是,你懂的,二维投影。
45:46
of a 3D world.
关于一个3D世界。
45:47
So even if you want a model that at the end of the day is going to produce 3D, it's still
所以即使你最终想要的是一个能生成3D的模型,它仍然
45:51
very powerful for it to learn on large quantities of image and video data.
在大量图像和视频数据上进行学习非常有用。
45:55
Because those you can get in those you can get in in very large quantities.
因为那些数据你可以大量获取。
45:59
And then when you've got 3D explicit 3D data, that's very powerful to learn from.
而当你有了显式3D数据,那也是非常强大的学习来源。
46:03
But you don't want to be bottlenecked only on learning from 3D data.
但你不希望只被3D数据的学习所限制。
46:06
I'm assuming that you're trying to train on whatever data you have natively as opposed
我假设你是想用你天然就有的任何数据进行训练,而不是
46:12
to like you take an image and project it into 3D or something that sounds super expensive
比如你拿一张图像,把它投影成3D,或者某种听起来超级昂贵
46:17
and noisy.
而且噪声很大的做法。
46:18
Yeah.
是的。
46:19
I mean, one of the lessons we learned from deep learning over the past decade is that
我是说,我们从过去十年的深度学习中学到的一个教训就是,
46:23
you want, you want big models trained on a lot of data that are trained end to end.
你想,你想要的是在海量数据上训练、并且端到端训练的大模型。
46:26
So if you've got a, like what is your task, is your task to like generate Gaussian
所以如果你有一个——比如你的任务是什么?是现在生成高斯泼溅世界,还是接下来生成真正强大的3D一致帧?
46:30
splat worlds today, or is it to generate really powerful 3D consistent frames the next
就像,想一想你要解决的任务到底是什么。
46:34
day?
以及我如何组织海量数据,让模型学会解决
46:35
Like just think about what is the task that you want to solve?
或者来自这些世界的视频。
46:37
And how can I marshal very large quantities of data to let a model learn how to solve
我们也可以将网格表示拟合到这个世界。
46:41
that task in a very general way?
那个任务是以一种非常通用的方式处理的吗?
46:43
Is the models output directly splats or is there some, I think this is kind of what
这个model的输出是直接就是splats,还是说有一些——我觉得这正是
46:50
you were just saying, like is it, is it producing some kind of normalized 3D representation
你刚才说的,就是它是不是在生成某种normalized 3D representation
46:56
or whatever that is?
或者不管那是什么?
46:57
And you can convert that to pixels or splats or you just convert your, it's just spitting
然后你可以把它转换成pixels或者splats,或者你就直接转换,它就是在吐出
47:02
out splats.
splats。
47:03
Yeah, the kind of rawest output from this model is splats in some way.
对,这个model最原始的输出在某种程度上就是splats。
47:07
And then once you've got splats, that's kind of our lowest common denominator 3D format.
然后一旦你有了splats,那就是我们最基础的通用3D format。
47:11
So once you've got splats, you can render them to an image and that can give you images
所以一旦你有了 splats,你就可以把它们 render 成图像,从这些世界得到图像或视频。我们也可以把 mesh representation 拟合到世界上。然后在某些上下文中,比如你想把它导入到 game engine 或 VFX engine 中,有时候这些引擎目前对 splats 的支持不是很好,所以有一个更经典的世界 3D triangle mesh 会很有用。所以 splat 差不多是我们从模型得到的最原始的 output format,然后你可以从它生成图像、视频或 meshes。
47:15
or videos from these worlds.
或者来自这些世界的视频。
47:18
We can also fit a mesh representation to the world.
我们也可以把网格表示拟合成这个世界。
47:23
And then in some contexts, like you want to import this into a game engine or a VFX engine,
然后在某些场景下,比如你想把这个导入到游戏引擎或者VFX引擎里,
47:29
sometimes those don't work so well with splats today, and it's useful to have a more
有时候这些引擎目前对splat的支持不太好,所以有一个更
47:33
classical 3D triangle mesh of the world.
经典的三维三角形网格来表示这个世界会更有用。
47:35
So the splat is kind of our most rawest output format from the model models and then you
所以splat基本上是我们从模型里输出的最原始的格式,然后你再
47:39
can go from that to images or videos or meshes.
可以从那里生成图像、视频或者网格。
47:44
But again, that's in pretty stark contrast to the RTFM model where there's no splats
但再说一次,这与 RTFM 模型形成鲜明对比,那里没有 splats
47:48
that just directly spits up pixels.
它直接吐出像素。
47:50
When you alluded earlier to a blog post that you recently wrote or your team recently published
当你之前提到你最近写的或团队最近发表的一篇博客文章时
48:00
and I thought what was really interesting about that was like, you know, as kind of an
我觉得那特别有意思的是,你知道,作为一个
48:06
analyst, anytime I see a taxonomy that tries to kind of pick apart the distinctions in
分析师,每当我看到一种试图区分其中差别的 taxonomy
48:12
the space and talk about those, I'm interested.
并讨论这些区别时,我就会感兴趣。
48:17
And that is what you try to do, at least in a particular dimension of world models.
而这就是你试图做的,至少在 world models 的某个特定维度上。
48:24
I think we've introduced like three or four other taxonomies and this conversations
我觉得我们已经介绍了三四个其他的 taxonomy,而这次对话...
48:27
so far, but talk a little bit about the way you kind of divided up that particular
目前就这样,不过聊聊你是如何划分世界模型的那个特定维度吧。
48:35
dimension of world models.
所以,像我说过的,我们聊过几次了,我觉得有很多不同种类的世界模型,大家都在训练,也都叫世界模型,但外表看起来差别很大。
48:38
So there, like I said, like we talked about a couple of times, I think there's a lot
而当我们非常认真地去想这件事时,意识到其实有一种框架,能让它们都变成同一事物的不同视角。
48:41
of different flavors of things that people are training and calling world models that
然后我们意识到,我们可以把这个追溯回 POMDP 形式体系。
48:45
look pretty different from the outside.
从外面看起来差别很大。
48:47
And when we thought about this really hard and realized that there is a framing where
当我们认真思考这个问题,意识到其实有一种框架,
48:50
they all actually are kind of like different, different views onto the same thing.
它们本质上都是同一个事物的不同视角。
48:56
And there we realized we could have grounded this back in this poem DP formalism that
然后我们意识到,可以把它重新落脚到那个DP形式体系上,
48:59
we talked about before.
我们之前聊过。
49:01
Because what in this poem DP, remember there are three things that are moving around
因为,在这个 POMDP 里你要记住,有三个东西在系统里运转。
49:05
the system.
你有一个在世界里的 agent,然后 agent 产生 actions,然后 world 进行 state 转移,再然后 agent 接收 observations。
49:06
You've got the agent in the world, then you've got the agent producing actions, you've
然后我们发现,如果你从这个角度看,现在大家训练的那些所谓的 world models,基本上都是在这个循环里输出三种东西之一。
49:10
got the world transitioning states, and then you've got the agent receiving observations.
让世界在状态之间转换,然后智能体接收观测。
49:15
And we realized, if you think about it that way, pretty much everything that people are
我们意识到,如果你这么想的话,现在人们训练的那些所谓的世界模型,基本上输出的都是三类东西之一。
49:19
training today that are called world models are usually outputting one of three things in
今天所谓的world models训练,通常输出的东西不外乎以下三种之一。
49:24
that loop.
那个循环。
49:25
Either you're building a model that outputs the outputs actions, you're building a model
要么你是在做一个输出 actions 的模型,要么是输出 states 的模型,再要么是输出 observations 的模型。
49:30
that outputs states, or you're building a model that outputs observations.
然后我们一旦意识到这点,就有种恍然大悟的感觉:哦,原来这些人不是在搞完全不同的东西,他们只是专注于这个基础的 POMDP loop 的不同部分。
49:34
And once we realized that, it was kind of an aha moment that, oh, it's not that these
而这些部分都是连在一起的,大家都在试图建模世界,理解世界如何随时间响应、演化和变化。
49:38
people are all building totally different things, they're just focusing on different parts
但不同的人因为不同的应用场景,会专注于这个 loop 的不同部分。
49:42
of this fundamental poem DP loop.
这个基础诗篇DP循环的核心。
49:44
And they're all connected together trying to model the world and understand how worlds
它们全都连接在一起,试图建模世界,理解世界如何
49:48
can respond and evolve and change over time.
随时间响应、演化、变化。
49:51
But different people are focusing on different parts of that loop for different applications.
但是不同的人关注的是这个循环的不同部分,用于不同的应用。
49:54
Got it.
明白了。
49:55
So something like a genie that's producing a stream of pixels that would be observations
所以,比如一个Genie,它输出的是像素流,这些像素就是模型中的observation;还有一个机器人系统,输出的是机器人在世界中要采取的动作,那种更专注于动作输出。
50:03
in this model and something like a robotic system that's producing actions for the robot
没错。
50:10
to take in the world, that's more focused on that as an output.
所以,然后我们就根据它们的输出,把这些大致分成三类world models。
50:15
Exactly.
就像你说的,比如如果输出的是observation,像Genie或者RTFM,那么...
50:16
So then, you know, then we kind of like taxonomize these into like these three different categories
所以呢,你知道,我们有点把这些分类成三种不同的类别,
50:22
of world models than based on what they're outputting.
就是根据这些世界模型的输出是什么来划分。
50:24
So like you said, like if you're outputting the observation like genie or like RTFM, then
就像你说的,比如如果你的输出是观察,像Genie或者RTFM,那么
50:29
we're calling that a rendering world model or just a renderer, right?
我们管那叫 rendering world model,或者就叫 renderer,对吧?
50:33
Because it's producing a final observation that can be consumed by a person or maybe
因为它生成的是一个最终的 observation,可以被人类,或者也许
50:37
by an agent.
被一个 agent 来使用。
50:38
So those are kind of what we, what we're terming a rendering or renderer as a world model.
所以这些就是我们所说的 rendering 或 renderer 作为 world model 的意思。
50:42
Then the other, the other easy, the other, the other cool one are all these robot people
然后另一个,另一个简单的,另一个,另一个很酷的是所有这些搞 robot 的人
50:46
training robotics policies, right?
在训练 robotics policies,对吧?
50:48
Like people are training models that, you know, the model operates a robot body and then
比如说人们在训练模型,你懂的,模型操作一个 robot body,然后
50:52
the robot body does something cool in the world.
这个 robot body 就在现实世界里做一些很酷的事情。
50:54
But then what is that model doing?
但那个模型到底在做什么呢?
50:56
That model is kind of on the other side of the, the poem DP loop, right?
那个模型有点像是处于 POMDP 循环的另一端,对吧?
50:59
It's receiving observations from the real world and now the model needs to take actions
它从现实世界接收 observations,现在这个模型需要采取 actions,
51:03
to, you know, try to make a change in the world.
去,你知道,试着改变世界。
51:05
And then they've got a world model that is inputting real world observations and outputting
然后他们有一个 world model,它输入现实世界的 observations,并输出
51:10
actions to be made in the real world.
要在现实世界中执行的 actions。
51:12
Whereas though that's that we're calling a planner world model as planner because it's
不过,那就是我们所说的 planner——也就是 world model as planner——因为它
51:16
planning out a sequence of actions to take in the world.
正在规划出一系列要在现实世界中采取的 actions。
51:19
And that's almost like exactly dual to the world model as renderer because the renderer
那这几乎和作为 renderer 的 world model 完全是 dual 的,因为 renderer
51:23
is sort of receiving actions from a, probably from a human user and then outputting observations
大概是在接收来自,可能是人类用户的 actions,然后输出 observations
51:28
of what a world might look like under those sequence of actions.
关于世界在这些 actions 序列下可能是什么样子。
51:32
So the renderer and the planner are kind of like almost perfect duels to each other.
所以 renderer 和 planner 几乎是彼此的 perfect duals。
51:37
And then the third one is what about the state?
然后第三个问题是,那 state 呢?
51:38
Like that's a tricky one that we keep coming back to.
这是个很棘手的问题,我们总是会绕回来。
51:42
So that we're calling world model as simulator, right?
所以我们称之为 world model as simulator,对吧?
51:44
Because another important property of world models is that maybe they should do some kind
因为 world models 的另一个重要特性是,也许它们应该做某种
51:48
of simulation of the state.
对 state 的 simulation。
51:51
In some contexts, you only care about the action of the observation, but in other contexts,
在某些 contexts 里,你只关心 observation 的 action,但在另一些 contexts 下,
51:55
you might want to know something about the state of the world under consideration and
你可能会想知道所考虑的 state of the world,并且
51:59
maybe understand or simulate how that state might evolve in response to actions in some
也许去理解或 simulate 那个 state 如何随着 actions 以 explicit 或 semi explicit 的方式演化。
52:03
explicit or semi explicit way.
所以这就是 world model 作为 simulator 的意思。
52:05
So that's the world model as simulator.
然后我们有点意识到,planner、simulator、renderer——当你去细想、拆解成这些 terms——基本上所有人们在训练的 world models
52:07
And then we kind of realized that planner simulator, renderer, when you kind of think
然后我们有点意识到,Planner、Simulator、Renderer,当你仔细
52:12
it, break it down to these terms, pretty much all the world models that people are training
去拆解成这些术语时,几乎所有人们训练的世界模型
52:15
these days actually can be bucket into one of these three camps, usually.
如今其实通常可以归入这三个阵营之一。
52:19
I thought the simulator was interesting in that I typically think of simulation as
我觉得simulator很有意思,因为我通常把simulation看作
52:27
like this external tool that we're using to develop models as opposed to this framing
一种外部工具,用来开发model,而不是像这种框架
52:34
of the model itself, you know, being fundamentally a simulator.
把model本身看作本质上就是一个simulator,你知道。
52:39
And there that like there, I think like another interesting thing is these start to blend
而且我觉得另一个有趣的点是,这些东西开始融合
52:42
together.
到一起。
52:43
And I think marble is actually an example of something that is somewhat straddling the boundary
然后我觉得marble其实就是一个例子,某种程度上它跨越了
52:47
between world model as renderer and world model as simulator, right?
world model as renderer和world model as simulator之间的边界,对吧?
52:50
Because because marble ends like when you when you when you as a user interact with
因为因为 marble 最终就是,当你作为用户与 marble 交互的时候,对吧?
52:55
marble, right?
你看到的是屏幕上的像素,对吧?
52:56
You're seeing pixels on the screen, right?
而那些像素,你知道,在某种意义上,你看到的是一个 observation。
52:58
And those pixels, you know, sort of in that sense, you're seeing an observation.
但那个 observation 并不是直接来自 neural network,而是来自一组 Gaussian splats。
53:02
But that observation did not come directly out of the neural network though that observation
而 Gaussian splats 有点像是一种 state representation,一种我们对世界的 explicit state representation。
53:05
came out of a set of Gaussian splats.
都来自一组高斯溅射。
53:07
And the Gaussian splats are kind of this state representation, this explicit state representation
而高斯溅射有点像这种状态表示,一种显式状态表示。
53:11
that we have at the world.
我们面对世界所拥有的。
53:13
And then once you have that explicit state representation, you can do other things with
而一旦你有了那个 explicit state representation,你就可以用它做除 rendering 之外的其他事情,对吧?
53:17
it other than rendering, right?
因为它是一个 explicit 3D representation,你可以测量两点之间的距离,或者我可以插入——我可以明确地操纵这个 state,比如拉进另一个 object asset 放到那个世界里。
53:19
Because it's an explicit 3D representation, you can measure the distance between two points
所以至少就拿我们现在的 marble 来说,作为用户的 end-to-end 体验,你在看 observations,看到屏幕上的 pixels,所以它有点像 renderer,但模型本身输出的是这种 explicit 或 semi explicit 的东西。
53:23
or I can insert I can manipulate the state explicitly by pulling in another object
或者我可以插入——我可以显式地操纵状态,把另一个对象
53:27
asset and putting it into that world.
资产拉进来,放进那个世界里。
53:29
So like at least the version of marble we have today, like the kind of the end-to-end experience
所以至少像我们今天拥有的Marble版本,那种端到端的体验。
53:34
as a user as you're seeing observations and you're seeing pixels on the screen so it's
作为用户,你看到的是观察结果,看到的是屏幕上的像素,所以它有点像renderer,但模型本身输出的是这种显式或半显式的东西。
53:37
kind of a renderer, but the model itself is outputting this explicit or semi explicit
我们最终想要的是一个统一的系统,能把所有这些事都做了。
53:42
world state that you can then do for other things.
world state,然后你可以拿它做其他事情。
53:45
That feels like a much more nuanced distinction than renderer and planner.
这感觉比 renderer 和 planner 之间的区别要微妙得多。
53:49
Like if I think about a renderer that as opposed to spitting out 2D frames was spitting
比如我在想一个 renderer,它不是输出 2D frames,而是直接输出
53:57
out somehow 3D directly, then you can do distances between points in that it's a representation,
某种 3D 的东西,那你就可以计算点之间的距离,因为它是一个 representation,
54:09
but it's also the observation.
但它同时也是 observation。
54:11
Exactly.
对。
54:12
And that's the other kind of point that we wanted to make here is that while there is
而且这就是我们在这里想说的另一点:虽然
54:15
this taxonomy, it's not very rigid.
有这样一个 taxonomy,它其实并不是很 rigid。
54:17
And I think sticking too rigidly to it would do a disservice to all of us.
而且我觉得,过于死板地坚持这个框架,对大家都没好处。
54:22
So the real world is messy and all of our taxonomies break down, but it's a useful framework
现实世界是很混乱的,我们所有的taxonomy都会失效,但它仍然是一个有用的思考框架。
54:26
to think about.
但现实中,我认为那些能再次带我们走到那一天的world models,实际上会把所有这些方面融合在一起。
54:27
But in reality, I think the world models that are going to carry us again to the day are
我们最终想要的,是一个能够同时做所有这些事情的统一系统。
54:31
actually going to blend all of these aspects together.
对,但我觉得我想多听一些关于simulator的详细阐述,除了大理石那个例子之外,还有没有其他例子能体现modelist simulator这个想法?
54:34
We ultimately want to have one combined system that could do all of these things.
对,但我觉得我想问得更细一点,关于simulator这个概念——除了大理石那个例子之外,还有没有其他例子能体现这种modelist simulator的想法?
54:38
Yeah, but I think I'm also asking for more elaboration on simulator and beyond the
你可以有网络来预测状态如何随动作演化,预测状态如何随时间变化,然后某种程度上有一个neural network的模拟,对应那种显式的world state。
54:49
marble example, are there other examples that capture this idea of modelist simulator?
大理石这个例子之外,还有没有其他例子能体现这种“模型即模拟器”的想法?
54:57
Yeah.
嗯。
54:58
I think there's a couple examples.
我觉得有几个例子。
54:59
I actually don't think anyone's nailed that one right now, but I think there's a
实际上我觉得现在还没人真正搞定那个,但我觉得有
55:03
couple of flavors of future systems I can imagine here.
几种我能想象的未来系统风格。
55:07
One is maybe like the marble explicit state future future version, not to say that this
一种是也许像 marble explicit state 的未来未来版本,这并不是说
55:12
is actually what we're going to do, but you one could do is also not to say that we're
我们真的会去做,但你也可以说,也并不是说我们
55:17
not doing it either, right?
不会这么做,对吧?
55:20
But you could imagine a version of this that treats something like Gaussian splats as
但你可以想象一个版本,把类似 Gaussian splats 的东西当作
55:25
a world state, but actually has a model evolve that state over time.
一个 world state,但实际上有一个模型让这个 state 随时间演化。
55:29
So you could have a model that inputs an image, outputs a Gaussian splat world, and then
所以你可以有一个模型,输入一张图像,输出一个 Gaussian splat world,然后
55:34
a user is going to take some action like pick up the bottle or move something around, and
用户会采取一些 action,比如拿起瓶子或者移动东西,然后
55:38
then you'll have a model that will then go in and update the Gaussian splat world with
然后你会有一个模型,它会去更新这个 Gaussian splat world,用
55:41
a powerful neural network model.
一个强大的 neural network model。
55:44
So that would be then a model that is working with this explicit or semi explicit world state
所以那就是一个模型,它处理的是这种 explicit 或 semi-explicit 的 world state
55:48
and then actually being able to evolve that world state over time in response to actions.
然后实际上能够根据 actions 让这个 world state 随时间演化。
55:53
And I don't think anyone's really built a system like that, but what they could.
而且我不认为有人真正构建过这样的系统,但他们可以做到。
55:56
And the other would be kind of the implicit world state.
另一个就是那种 implicit world state。
55:59
And maybe we can build systems that maybe they're not working on Gaussian splats directly,
也许我们可以构建一些系统,它们可能不是直接处理 Gaussian splats,
56:03
but they have some kind of vector implicit vector representation of a world state.
而是有某种 vector implicit vector representation 来表示 world state。
56:07
And it behaves in the way that an explicit state would where you can render observations
它的行为方式就像 explicit state 那样,你可以从中 render observations。
56:11
from it.
你可以让 networks 预测这个 state 如何根据 actions 演变,预测它如何随时间演变,从而有点像给那个 explicit world state 一个 neural network 的模拟。
56:12
You can have networks that predict how that state evolves in response to actions, predict
你可以有那种网络,用来预测状态如何随着动作而变化,也就是做预测。
56:17
how that state is going to evolve in time, and sort of have a neural network analog of
这个状态会如何随时间演化,并且某种程度上拥有一个神经网络的类比。
56:21
that explicit world state.
那个显式的世界状态。
56:23
And I think people are working on that, but that feels like it's still a pretty open
我觉得有人在研究这个,但感觉这仍然是个挺开放的研究问题——到底什么才是让它奏效的确切方法。
56:26
research question as to what's the exact right recipe to get that to work.
对,另一个让我在思考 simulator 时跳出来的点是,也许 simulator 的一个完美体现就是我们聊过的那个 theory builder。
56:30
Yeah, the other thing that jumps out of me in thinking about the simulator is that like,
就像它把这些抽象的想法,某种程度上提炼成——你知道嘛——在这种情况下,state 就是一组关于世界的 theories。
56:37
maybe a perfect expression of a simulator is this theory builder that we talked about.
对,我是说,那你就得聊 states 和 meta states,还有 theories 和 meta-steered theories,然后把它往 abstraction 的层级上再推一层,对吧?
56:44
Like it takes these abstract ideas and kind of boils them down into, you know, in this
就像它把这些抽象的概念提炼出来,你知道,在这种情况下,状态就是关于世界的一套理论。
56:49
case, the state is a set of theories about the world.
对,我的意思是,那你就得聊状态和元状态、理论和元引导理论,然后把抽象层次再往上推一层,对吧?
56:52
Yeah, I mean, then you got to talk about like states and meta states and theories and
但那种东西,我觉得没人知道该怎么做。
56:56
meta-steered theories and push it up a level of abstraction, right?
是啊。
56:59
Like, you could say like maybe the theory builder is the one who's writing the laws of
比如说,你可以说也许 theory builder 就是那个在 states 里写下物理定律的人。
57:03
physics in the states.
所以 theory 有点像是一种封装,囊括了可能存在的各种世界,以及这些世界如何被允许演化;而 state 则是 theory 的一个特定 instantiation,告诉你一个特定的世界。
57:04
So like the theory kind of encapsulates the kinds of worlds that may exist and how those
但那就……我觉得没人知道该怎么做。
57:08
kinds of worlds are allowed to evolve, and then the state is like one particular instantiation
对。
57:12
of the theory that tells you a particular world.
对。
57:14
But that's like that's that I don't think anyone has any clue how to do.
是啊。
57:17
Yeah.
我觉得我们这个领域最终会达到的,是能同时处理所有这些事情的模型。
57:18
Yeah.
是的。
57:19
Okay.
好的。
57:20
We're getting a bit too abstract there.
我们刚才讨论得有点太抽象了。
57:21
You talk a little bit about this idea of a unified world model.
你稍微讲了一下这个统一世界模型的概念。
57:25
I think we capture that a little bit.
我觉得我们算是捕捉到了一点。
57:28
It's just the idea that it's kind of a leaky abstraction and you're not expecting, you're
其实就是说,它有点像一种漏水的抽象,而且你并不指望——你指望的是在真实世界产品中出现某种交叉融合。
57:36
expecting kind of crossover in real world products.
没错。
57:40
Exactly.
我觉得我们作为这个领域,最终会走向的是能够同时处理所有这些事情的模型。
57:41
What I think we're going to get to as a field is models that can do all of these jointly,
我认为,作为一个领域,我们最终会发展到能够同时完成所有这些任务的模型。
57:47
and they're all going to benefit from each other.
而且它们都会互相受益。
57:48
And why is that?
那为什么呢?
57:49
It's because they're all kind of solve asking similar questions like they want to understand
因为它们都在解决类似的问题,就像它们想要理解
57:53
what are the kinds of worlds that could exist?
可能存在哪些种类的世界?
57:55
How could those worlds respond to action?
这些世界如何对行动做出反应?
57:57
What do those worlds look like?
这些世界看起来是什么样的?
57:59
What kind of observations arise from those worlds?
从这些世界中会产生什么样的观察?
58:01
How do they evolve in time?
它们如何随时间演化?
58:02
What can you do with them?
你能拿它们做什么?
58:04
These are all kind of fundamental questions that all connect to each other, right?
这些都是互相关联的根本性问题,对吧?
58:08
If I get better at understanding how the world state evolves, probably I'll also get
如果我更擅长理解 world state 如何演变,可能作为 renderer 我也能更好地预判从不同角度看起来会是什么样。
58:11
better at anticipating what it's going to look like from different angles as a renderer.
而且如果我很擅长想象 implicit world state 是怎么演变的,那对 planning 应该也很有帮助。
58:15
And if I'm really good at imagining how the implicit world state has evolved, probably
如果我知道我能隐式地演化自己的 world state,那我大概也能针对那个状态做 planning,知道该采取什么行动去影响世界。
58:19
that's pretty good for planning.
这对规划来说已经相当不错了。
58:21
If I know if I can implicitly evolve my world state, probably I can also plan against
如果我知道我能隐式地演化我的世界状态,大概我也能针对那个状态做规划,并知道怎么采取行动去影响世界。
58:25
that state and know how to take actions to affect the world.
但更多的问题是,今天我想不想开一个机器人?
58:28
So I don't think we're there yet, but I think over the next couple of years we'll start
所以我不觉得我们到了那一步,但我觉得接下来几年里,我们会开始看到 models 把越来越多的这些能力整合进一个强大的 unified model。
58:32
to see models that combine more and more of these capabilities into one powerful unified
而到那时候,你在某个时间点想要什么输出,与其说取决于专门的 models,不如说取决于我们的 renders、simulators 或 planners。
58:37
model.
但更多是——今天我是想 drive 一个 robot 吗?
58:38
And then what output you want at one moment in time is less a function of having specialized
所以它就会以 planner mode 运行。
58:42
models than our renders or simulators or planners.
而明天我想 drive 一个 virtual video game,所以它就会以 render mode 运行。
58:45
But more about, is it today I want to drive a robot?
生成模型,我觉得它很可能会继续演化,而且已经——就像我们以前那样——
58:48
So it's going to operate in planner mode.
所以它将以规划者模式运行。
58:50
And tomorrow I want to drive a virtual video game, so it's operating in render mode.
然后明天我想开一个虚拟电子游戏,所以它是以渲染模式运行的。
58:54
Or the next day I want to simulate possible counterfactuals in a world.
或者第二天,我想在一个世界里模拟各种可能的 counterfactuals。
58:58
And now I want it to operate in simulator mode.
而现在我想让它运行在 simulator mode 下。
59:00
So I think that's where the field is going to get the next couple of years.
所以我觉得这就是这个领域接下来几年要发展的方向。
59:04
That makes me think of this idea of the output layer in a neural network is it could be classification
这让我想到一个想法——neural network 的 output layer,可以是 classification,也可以是 regression 之类的,但 core representation 基本是一样的。
59:11
or it could be regression or something else, but like the core representation is for
我们只是用不同的方式去使用它。
59:16
the most part the same.
对,就是这样。
59:18
And we're just kind of using it in different ways.
我们只是用不同的方式使用它。
59:23
Yeah, exactly.
是的,没错。
59:24
I think that's what we're going to get.
我觉得这就是我们会得到的东西。
59:25
We're going to have these giant unified world models that have maybe different input
我们会有这些巨大的统一world models,可能有不同的input heads、不同的output heads,知道怎么输入和输出不同类型的东西。
59:28
heads, different output heads that know how to input and output different kinds of things.
但最终,所有的compute、模型所有的parameters都会集中在一个共享的trunk上,这个trunk就是那个world model,知道怎么模拟任何类型的事物。
59:32
But ultimately, like all the compute, all the parameters of this model are going to be
而且这些会隐含在它的weights里面。
59:35
this shared trunk that is this world model that knows how to simulate any kind of a thing.
然后它就可以按应用需要,把这些世界知识以actions、states或observations的形式呈现出来。
59:40
And it's going to be implicitly in its weights.
而且它会隐含在它的权重里。
59:41
Then it can surface that world knowledge as actions or as states or as observations as needed
然后它可以按需将该世界知识呈现为行动、状态或观察结果。
59:47
for the application.
生成式模型,我认为它很可能会不断进化,而且它已经像我们过去那样了。
59:48
I think the next question I wanted to get at was, and really kind of a closing question,
我觉得我接下来想问的问题,真的是一个收尾的问题,
59:52
is like do current architectures get us there like, and you just said, you know, kind
就是说,当前的 architectures 能让我们达到那个目标吗?你刚才说,你知道,有点
60:00
of know like the architecture evolves, but it's kind of consistent with the way we think
像是 architecture 会演变,但和我们想的方式挺一致的,
60:04
about, you know, today's models, transformers, et cetera.
就是,你知道,今天的 models,transformers 等等。
60:09
Do you foresee kind of an architectural step function being required to, you know, fulfill
你觉得会不会需要一个 architecture 上的 step function,才能,你知道,实现
60:18
the potential of world models?
world models 的潜力?
60:20
Maybe yes, but not as big a one, probably not too big.
也许需要,但不需要那么大的一个,可能不会太大。
60:22
I think transformers are really powerful.
我觉得 transformers 真的很强大。
60:25
Yeah, I get this question a lot, like does that mean transformers are dead?
对啊,我经常被问到这个问题,像是“那是不是说 transformers 就完蛋了?”
60:28
Do we need to architectures?
我们是不是需要新的 architectures?
60:29
Like no, transformers are great.
不是啦,transformers 很棒。
60:31
Transformers are super powerful.
transformers 超级强大。
60:32
They scale up really well.
它们 scale up 得特别好。
60:33
They can work on all the different kinds of data, all different kinds of data.
它们能处理各种不同类型的数据,各种不同类型的数据。
60:37
So transformers are very powerful.
所以 transformers 非常强大。
60:39
They could use them for all the different kinds of problems.
你可以用它们来解决各种不同的问题。
60:41
I think there is a, like a loss function question about what is the right loss function
我觉得有一个,像是损失函数的问题,就是训练这类生成模型时,什么才是正确的损失函数?
60:45
for training these kinds of generative models?
而这更是个悬而未决的问题。
60:47
And that's more up in the air.
比如说,我们到底想不想把这些东西当作生成模型来训练?
60:49
Like do we want to train these things, you know, as a generative model?
如果它是生成模型,你是用 diffusion 还是 rectified flow 来训练?
60:52
If it's a generative model, do you train it via diffusion or rectified flow?
你是把它当作 discrete order regression 来训练,还是用别的方式?
60:55
Do you train it as a discrete order regression or something else?
所以有一个损失函数的问题,它确实适用于任何一种
60:59
So there is a kind of a loss function question that's really applicable to any kind of
生成模型,而且我认为这个问题很可能会继续演变,而且已经像我们以前那样
61:02
generative model that I think is probably going to evolve and has already like we used to
至少从某个特定角度来看,如果我们谈的是空间信息,那它可能有用武之地。
61:05
do gans, now it's diffusion, like that's a kind of a change in loss function more than
比如 GANs,现在是 diffusion,这更像是 loss function 的改变,而不是
61:09
it is a change in architecture.
architecture 的改变。
61:12
But you know, another one area where I do think we need to see some evolution architecturally
但你知道,另一个我觉得我们需要在 architecture 上看到一些演进的领域
61:17
is how do we deal with really, really long contacts and really, really, lots of tokens?
就是我们如何处理特别特别长的 contexts 和特别特别多的 tokens?
61:22
So you know, this is something that comes up in LLM is already like LLMs are transformers,
所以你知道,这是 LLM 中已经出现的情况,就像 LLM 都是 transformers,
61:26
they work on tokens.
它们处理 tokens。
61:27
But you can do a lot of problems, you know, even maybe, maybe this is less true now
但你可以解决很多问题,你知道,甚至也许,也许这在 agentic 时代不那么成立了,
61:30
in the agentic era, but a couple of years ago, like it was hard to imagine situations
但在几年前,真的很难想象那些情况。
61:34
for that LLM needs to operate on million or 10 million tokens, because that's like
要做到那一点,LLM需要在百万或千万个tokens上运行,因为那就像是整本书一样。
61:39
whole books.
对,原则上,有些问题你需要把整本书,甚至很多本书塞进你的context里。
61:40
Like, yeah, in principle, there's problems you need to fit whole book, cram whole books
但你知道,也许那些不是必须的,你可以在不需要那种context的情况下用语言做很多事情。
61:43
into your contacts.
但现在,如果你想谈论世界,比如,你知道,我需要生成,你知道,对很多很多high-dimensional图像或很多很多3D空间进行建模,就像
61:44
But, you know, maybe those are not, you could do a lot with language without needing that
但是,你知道,也许那些并不是必需的,没有上下文这种东西,你也能用语言做很多事。
61:47
thing of a context.
但现在,如果你想谈论世界,比如,你知道,我需要生成,你知道,对很多很多高维图像或者很多很多三维空间进行建模,就像
61:49
But now, but if you want to talk about worlds, like, you know, I need to generate, you
上下文长度会限制世界的规模吗,还是说,你知道,我们就像
61:53
know, model maybe lots and lots of high-dimensional images or lots and lots of space in 3D, like
而且至少从某个视角来看,如果谈到空间信息,那或许确实有它的作用。
61:59
it's very quick, very easy to get situations where you want hundreds of thousands or millions
很快就会很容易遇到这种情况:你想要几十万,或者几百万
62:03
or tens of millions of tokens of context for different world modeling problems.
甚至几千万 tokens 的 context 来解决不同的 world modeling 问题。
62:07
So I do think we need to see some evolution in, you know, how do we adapt transformers
所以我确实认为,我们需要看到一些演进,你懂的,怎么调整 transformers
62:10
to work at really, really long context lengths to, because that becomes, that becomes kind
让它们能在非常非常长的 context lengths 下工作,因为那会变成一种
62:17
of a nice to have in LLMs, that becomes the everyday problem for any, any skilled up world
在 LLMs 里的 nice to have,但对任何、任何训练有素的 world model 来说,
62:22
model.
这就会变成日常问题。
62:23
That begs the question, how should we think about both tokens and context in world models,
这就引出了一个问题:我们应该怎么看待 world models 里的 tokens 和 context?比如 context length 会不会限制世界的规模,还是说,你懂的,我们是不是有点像...
62:29
like does the context length limit the size of the world, or, you know, are we like
也许吧。
62:34
paging out sections of the world, and so it's not a hard constraint, and like does token
把世界的各个部分paging out,所以它不是一个hard constraint,然后就像token是否
62:39
correspond to a splat or some other construct, like how do these things relate?
对应到一个splat或者别的什么construct,比如这些东西之间是怎么关联的?
62:44
I think that's where there's a lot of, you know, different things happening, and that's
我觉得那里有很多,你知道的,不同的事情在发生,然后那就是
62:46
where I do see a lot of evolution.
我确实看到很多演化的地方。
62:48
So that's like in some architectures, a token will be maybe a little chunk of a video,
所以就像在一些architecture里,一个token可能是一小块视频,
62:53
you know, maybe a 16 by 16 spatial 16 by 16 pixels and like four frames in time.
你知道的,可能是一个16 by 16 spatial,16 by 16 pixels,然后时间上大概四帧。
63:00
So some architectures, a token is literally like a little patch of a video.
所以有些architecture里,一个token literally就是视频的一小块patch。
63:04
In some architectures, that token might be a little bundle of splats somewhere in the
在一些architecture里,那个token可能是splats的一小捆,在某个...
63:08
world.
世界。
63:09
In some architectures, that token might be, might be a little chunk of 3D space.
在某些架构里,那个 token 可能是,可能是一个小 chunk of 3D space。
63:14
Maybe I've carved up my 3D space into voxels, and now each token corresponds to some chunk
也许我把我的 3D space 切分成了 voxels,现在每个 token 对应着某个 chunk
63:18
of 3D space.
of 3D space。
63:20
Or maybe those tokens are something abstract and latent.
或者那些 tokens 是某种抽象且 latent 的东西。
63:23
Like maybe I've got a latent world state that is just a num, like maybe I've just allocated,
就比如我有一个 latent world state,它就是一个数字,就像我刚刚分配了,
63:27
you know, my world state has a thousand tokens.
你懂的,我的 world state 有一千个 tokens。
63:29
What do they mean?
它们是什么意思?
63:30
It's opaque.
它是 opaque 的。
63:31
Model figured out.
Model 搞明白了。
63:32
So I think that's where we're seeing a lot of evolution and a lot of different approaches
所以我觉得这就是我们正在看到很多演变和很多不同方法的地方,
63:35
taking, doing different things architecturally.
它们在架构上采取并尝试不同的做法。
63:37
Also in this architectural thread, do you see a role for kind of these ideas around geometric
另外,在 architectural 这条线上,你觉得这些关于 geometric
63:43
deep learning, or, you know, different models that incorporate symmetry?
deep learning 的想法,或者,你知道,那些整合了 symmetry 的不同模型,有作用吗?
63:48
You know, some, you know, it's been a variety of work around trying to incorporate geometry
你知道,有些,你知道,有很多工作试图将 geometry 纳入
63:55
into deep learning.
deep learning 中。
63:59
And at least from a particular lens, it seems like if we're talking about spatial information,
也许吧。
64:04
there may be a role for that.
这可能有它的用武之地。
64:05
Maybe.
也许吧。
64:06
But I think the lesson we've learned over and over again in deep learning is that you want
但我觉得我们在深度学习中反复学到的教训是,你想要
64:10
simple representations and then scale up the model behind the simple representations.
简单的表示,然后在简单表示的基础上去扩展模型。
64:14
So you know, when you, and then also like choose a representation that's adapted for
所以你知道,当你——而且还要选择一种适合
64:18
the task at hand, right?
手头任务的表示,对吧?
64:19
Like if you want to, if the thing you actually want out is video frames or images like just
比如如果你想要——如果你真正想得到的是视频帧或图像,那就直接
64:24
do that, you don't need to bottleneck yourself through an explicit 3D representation.
这么做,你不需要通过显
64:28
If you do want a 3D representation, be it a splat or a mesh, find a simple way to interface
如果你确实想要一个3D表示,不管是splat还是mesh,那就找个简单的方式把那个架构和那个3D表示跟神经网络对接起来。
64:32
that architecture, that, that 3D representation with the neural network.
你越是往里塞对称性和各种复杂、繁琐的表示,往往就越难优化,而且你做的假设也越多。
64:36
And the more you put in symmetries and complex and complicated representations, often the
所以它就越难在scale上work,对吧?
64:42
harder it could, it'll be to optimize and the heart and the more assumptions you're
比如你加入某种对称性的概念,嗯,人通常是对称的,我们有两条腿两条胳膊,但并不是每个人都有两条腿两条胳膊。
64:45
making.
making。
64:46
So the less it's going to work at scale, right?
所以它在规模化之后就越不管用,对吧?
64:48
Like you put in some notion of, you know, symmetry, well, people are usually kind of symmetric.
就像你加入某种对称性的概念,你知道,人通常大体上是对称的。
64:53
We have two legs and two arms, but not everybody has two legs and two arms.
我们有两条腿和两条胳膊,但不是每个人都有两条腿和两条胳膊。
64:56
So those, those like hard assumptions of symmetry, maybe get you, maybe 80% or 90% of
所以那些,那些关于对称性的硬性假设,也许能帮你解决80%或90%的问题,但最终还是会失效的。
65:01
the way there, but they're going to break down eventually.
你今天就可以注册去玩玩看。
65:03
And then you'd like to be in a position where your architecture or your model is expressive
然后你会希望自己的架构或模型有足够的表达能力,去应对那些捷径开始失效的情况。
65:08
enough to handle the cases where those shortcuts are going to break down.
大家应该去哪里了解更多这方面的内容呢?
65:13
Where should folks go to learn more about this?
有没有什么经典资源,是你想推荐给刚进入这个领域、想更深入探索的人?
65:16
Are there any like canonical resources that you like to point folks to who are new to
有的,你绝对可以试试Marble,那是我们的产品,在marble.worldlabs.ai上。
65:20
the space and want to dig in more deeply?
今天就可以注册并试用一下。
65:23
Yeah, so you can definitely try out Marble, that's our product at marble.worldlabs.ai.
我真希望能推荐一个更好的关于世界模型的系列讲座或书籍之类的,
65:27
You can sign up and play around with that today.
我倒是希望能推荐一套更好的关于世界模型的讲座系列或者书什么的,但说实话,我觉得目前还没人写出特别牛的东西,这也是我们写那篇博客文章想部分解决的问题,但我觉得其实有人可以挖得更深,把这些想法拆解得更好。
65:30
I wish I could recommend a better lecture series or book or something on world models,
所以,如果我有遗漏什么的话,我很乐意——很乐意听你说说。
65:35
but I just, I don't think anyone's written down anything super awesome yet, which is partially
但我只是觉得,目前还没人写出什么特别厉害的东西,这部分是因为——
65:40
what we tried to solve with our blog post, but I think I think someone could go a lot deeper
我们写那篇博客想解决的问题就是这个,但我觉得,其实有人可以做得更深。
65:44
and unpack a lot of these ideas in a better way.
并且用更好的方式把这些想法展开来讲。
65:46
So I, you know, if someone, if I'm missing something, I'd love to, I'd love to, for your
所以,我,你知道的,如果有人——如果我有遗漏的地方,我很乐意,很乐意听你的。
65:50
viewers or listeners to let me know.
观众或听众们,记得告诉我你们的想法。
65:52
Justin, thank you so much for taking the time to share with us a bit about what you've
Justin,非常感谢你抽出时间跟我们分享你最近在做的事情。
65:56
been up to.
你和world labs的团队,做的那些东西真的很酷。
65:58
You and the team at world labs, very cool stuff.
你和world labs团队做的那些东西,真的非常酷。
66:00
Yeah, thanks so much for having me.
感谢你邀请我,真的非常荣幸。
66:01
This was a lot of fun.
这次聊得非常开心。

Play Queue

☀️