TWIML AI
Do AI Tokenomics Matter More Than Model Benchmarks? with Chris Potts - #776
2026-09-09 · 3569
TWIML AI episode #776 features Stanford professor and Big Spin co-founder Chris Potts in a discussion about whether AI tokenomics—how tokens are consumed, priced, and converted into value—matters more than conventional model benchmarks. The conversation spans Potts's linguistics background, the nature of LLMs as distributional learners, architecture and efficiency research, DSPy, and the challenge of measuring ROI from AI systems, including ideas like tokenflation and a token-based market basket.
本期 TWIML AI 播客 #776 邀请 Stanford 教授、Big Spin 联合创始人 Chris Potts,讨论 AI tokenomics 是否比 model benchmarks 更重要。他从语言学背景出发,谈到 LLM 作为 distributional learners 的特点,以及 transformer 架构、效率研究和 DSPy 对 AI 系统设计的影响。核心论点是,随着 AI 被嵌入更多产品和工作流,token 消耗、成本和推理效率正在成为衡量模型价值的关键维度,而传统 benchmark 可能不足以反映真实回报。Potts 还提出 tokenflation 等概念,并讨论如何像 CPI 一篮子商品那样,用更贴近任务和产出的指标来评估 token 花费是否带来实际价值。DSPy、Big Spin、token efficiency 和模型成本等话题被反复提及。
00:00
This episode is brought to you by Blitzie, the Autonomous Software Development Platform
本期节目由 Blitzie 赞助,它是 Autonomous Software Development Platform
00:05
built for enterprise scale.
专为 enterprise scale 打造。
00:07
Today's code bases have grown beyond human comprehension.
如今的 code bases 已经大到人类无法理解了。
00:10
Millions of lines, decades of tech debt, and complexity existing tools just can't fathom.
几百万行代码、几十年的 tech debt,还有现有工具根本摸不透的 complexity。
00:16
With Blitzie, thousands of specialized agents reverse engineer the code base, mapping
有了 Blitzie,数千个 specialized agents 会 reverse engineer 这个 code base,梳理
00:21
the architecture, dependencies, and business logic.
它的 architecture、dependencies 和 business logic。
00:24
With that context, the platform then autonomously executes entire epochs, writing, validating
有了这些 context,平台就会自主执行整个 epochs,编写、验证
00:29
and testing the code for every project.
并测试每个项目的代码。
00:32
The result, Fortune 500 enterprises, are able to modernize legacy systems and ship new
结果是,Fortune 500 企业能够把 legacy systems 现代化,并交付新
00:37
features five times faster.
功能,速度快五倍。
00:40
Want to try Blitzie on your code today?
今天想在你的代码上试试 Blitzie 吗?
00:42
Unlock 1 million lines of reverse engineering and 25k lines of code generation by visiting
想解锁 100 万行 reverse engineering 和 25k 行 code generation,就访问
00:47
blitzie.com slash sandbox.
blitzie.com slash sandbox。
00:53
We've started to enter a phase of AI that's not just about making models smarter.
我们已经开始进入 AI 的一个阶段,这个阶段不只是让 models 变得更聪明。
00:58
It's also about making them economically sustainable.
这同时也是在让它们在经济上可持续。
01:02
As reasoning models consume more tokens, context windows continue to grow, and agents become
随着 reasoning models 消耗更多 tokens,context windows 持续增长,而 agents 变得
01:07
embedded in more products and workflows, the economics of these systems are becoming
随着它们嵌入更多产品和 workflow,这些系统的经济账正变得
01:11
impossible to ignore.
无法忽视。
01:13
That's given rise to a new conversation around tokenomics.
这催生了一场围绕 tokenomics 的新讨论。
01:17
How we think about the costs, incentives, and trade-offs shaping next generation of AI.
我们如何看待塑造下一代 AI 的成本、激励和 trade-offs。
01:22
One person who's been thinking deeply about this is Stanford professor and Big Spin co-founder
一直在深入思考这个问题的人,是 Stanford 教授、Big Spin 联合创始人
01:29
His recent work argues that measuring AI progress requires looking beyond model benchmarks to
他最近的研究认为,衡量 AI 进展需要超越 model benchmarks,去
01:33
ask a different question.
问一个不同的问题。
01:35
What are our tokens actually buying us?
我们的 tokens 到底给我们换来了什么?
01:39
Here's Chris explaining how he thinks about tokenomics.
接下来是 Chris 解释他如何看待 tokenomics。
01:42
Another interesting moment to be in, as we're all being made aware of the true costs of all
另一个很有意思的时刻是,我们所有人都开始意识到所有
01:47
this AI usage.
这些 AI 使用的真实成本。
01:48
The analogy here is like, I used to cost me $20 to take a ride share to the airport Uber
这里的类比就像,我以前打网约车去机场要花 20 美元,用 Uber
01:54
or Lyft, and now it costs $90, but it's more like $20 to like $500 or something, right?
或 Lyft,现在要花 90 美元,但实际更像是从 20 美元到 500 美元左右,对吧?
01:59
I think what's happening is that the big providers are testing the waters on charging us
我觉得现在的情况是,那些大型 providers 正在试水,向我们收取
02:04
the true costs, plus whatever profit they need to make as they all try to gear up for
真实成本,再加上他们需要赚取的任何利润,因为他们都在努力为……做准备
02:08
IPOs and so forth.
IPOs 等等。
02:10
And in turn, that is very quickly leading people to ask questions like, what is the return
反过来,这很快会让人们开始问这样的问题:我们买的这些 tokens 的投资回报
02:15
on investment for all these tokens that we have purchased?
到底是多少?
02:19
And it's a very tricky area to be in because what does it mean to think about value in
而且这是一个非常棘手的领域,因为在这个
02:23
this context?
语境下,思考价值到底意味着什么?
02:24
I'm Sam Charrington, and this is the Twimmel AI podcast.
我是 Sam Charrington,这里是 Twimmel AI podcast。
02:28
For over a decade, I've been exploring the ideas and innovations shaping the future
十多年来,我一直在探索那些塑造 AI 未来的想法和创新
02:32
of AI through conversations like this one that help you understand what's real, what's
通过像这样的对话,帮助你理解什么是真实的,什么是
02:37
next, and what matters.
接下来,以及真正重要的东西。
02:49
I want to say thanks for coming on.
我想说,谢谢你来做客。
02:51
I've been looking forward to this conversation.
我一直很期待这次对话。
02:54
And I think where I'd love to start us off is to really dig into your background and
我觉得,我想先带我们从这儿开始:真正深入聊聊你的背景,以及
03:02
how it kind of got you to where you are now.
这一路是怎么让你走到今天这个位置的。
03:06
Yeah, my background is in linguistics.
对,我的背景是 linguistics。
03:10
Linguistics proper, not even natural language processing.
真正的 linguistics,甚至都不是 natural language processing。
03:13
I did my PhD on, among many other things, swears, what swears are like, why we swear,
我读博的时候,研究过很多东西,其中就包括脏话——脏话是什么样的,我们为什么说脏话,
03:19
what information they encode, what kind of taboos exist around them and so forth.
它们 encode 了哪些信息,围绕它们有哪些禁忌,等等。
03:23
And that was actually the trigger that got me into NLP because I wanted a lot of data
而这其实正是让我进入 NLP 的契机,因为我想要大量
03:29
of people swearing.
关于人们说脏话的 data。
03:30
I wanted to know what the context was like, what their intentions were.
我想知道 context 是什么样的,他们的意图是什么。
03:34
So I turned to corpora.
所以我转向了 corpora。
03:36
And from there, you start using NLP toolkits to add structure to those corpora.
从那里开始,你会用 NLP toolkits 给这些 corpora 加结构。
03:41
And then after a few years, maybe you're writing your own tools for doing that work.
然后几年之后,也许你就在写自己的工具来做这项工作了。
03:45
And then when you look back after 18 years or whatever it's been, you're just an AI person
然后等你回头看,过了 18 年还是多久,你就只是个搞 AI 的人
03:51
or an NLP person.
或者搞 NLP 的人。
03:52
But that is the true story.
但这就是真实的故事。
03:53
And I feel like if I had to, I could trace the lineage of every one of my current projects
而且我感觉,如果非要的话,我可以把我现在每一个项目的脉络
03:58
back to my fascination with why we care when someone drops an f-bomb.
都追溯回我对这个问题的着迷:为什么有人爆出一句 F 词时,我们会在意。
04:05
So are you an f-bomb dropper or did you come at it from the perspective of trying to
那你是自己会说 F 词的人,还是从试图
04:10
understand these others?
理解这些其他人的角度切入的?
04:13
I think very infrequently in my life, for my linguistics class, semantics and pragmatics,
我觉得在我人生里这种情况非常少,至于我的语言学课、semantics and pragmatics,
04:20
which is about linguistic meaning, on the final day, we always do a class on swearing.
这是关于语言意义的,在最后一天,我们总会上一节讲脏话的课。
04:25
And I review the history and we kind of tie all the course themes together.
我会回顾历史,也会把课程的所有主题串起来。
04:29
My handouts for that are full of swears.
我那节课的讲义里全是脏话。
04:33
But I only swear once in the lecture.
但在讲座里我只爆一次粗口。
04:35
I present the result that people remember things better if the utterance contains a
我会展示这个结果:如果一句话里包含一个
04:40
swear because it has a kind of emotional resonance, very primitive reaction.
脏话,人们会记得更牢,因为它有一种情感共鸣,一种非常原始的反应。
04:45
And so in that moment, I pick some fact from the course, some trivial thing.
于是在那一刻,我会从课程里挑一个事实,某个很琐碎的东西。
04:49
And I restate it with a swear and then I say, all of you will remember this for eternity.
然后我用一句带脏话的话重新说一遍,接着说,你们所有人都会永远记住这个。
04:54
But other than that, I'm very shy about it in the class.
但除此之外,我在课上对这件事特别害羞。
04:59
I'm sure there is loads of research on this, but I grew up in New York City.
我确定这方面有一大堆研究,不过我在 New York City 长大。
05:07
And as a New Yorker, I think that swearing is just kind of part of my natural language
而作为一个 New Yorker,我觉得说脏话差不多就是我的 natural language 的一部分
05:12
and way of communicating.
也是我交流方式的一部分。
05:14
And I married a Midwestern girl and she doesn't tolerate it at all.
而且我娶了一个来自中西部的女孩,她完全不能容忍这个。
05:19
She doesn't do it.
她自己不说脏话。
05:20
She doesn't tolerate it.
她不能容忍这个。
05:21
She won't tolerate it for me.
她可不会容忍我这样。
05:23
And it may for, we've been married for 30 years, so I adapt quickly, apparently.
而且可能,因为——我们已经结婚30年了,所以显然我适应得很快。
05:29
But for a long time, it took a lot of restraint to change that way of communicating, particularly
但很长一段时间里,要改变那种沟通方式需要很多克制,尤其是
05:38
when I'm communicating about something that I'm excited about or emotional about or want
当我在聊让我兴奋、让我情绪激动,或者我想
05:43
to convey the importance of, it's a really interesting topic.
传达其重要性的事情时,这真的是个很有意思的话题。
05:52
Well, we're not going to turn the podcast into a podcast about swearing, but I imagine
好吧,我们不会把这个播客变成一档聊脏话的播客,但我想
05:57
there was enough research there that we could if we wanted to.
那方面的研究足够多,如果我们想的话,完全可以做。
06:00
It's a fascinating area because it gets right to the heart of the culture that we've
这是个引人入胜的领域,因为它直击我们文化的核心,而我们已经
06:04
constructed and how it relates to our usage and everything else about us.
它是怎么被建构出来的,以及它如何关联到我们的使用和我们身上的一切。
06:08
Yeah, it's fascinating that we have swears.
是啊,我们竟然有脏话这件事挺有意思的。
06:11
When the old swears lose their power, we invent new ones.
当旧的脏话失去力量时,我们就会发明新的脏话。
06:13
We pretend like nobody should use them, but as you say, people use them all the time.
我们假装好像谁都不该用它们,但就像你说的,人们一直都在用。
06:17
And it feels like an important part of being a language user that we've got them
而且感觉,作为语言使用者,我们能拥有它们
06:21
available to us.
并让它们随时为我们所用,是很重要的一部分。
06:24
And the string of questions.
还有那一连串的问题。
06:25
I'd love to hear your take on kind of a linguist in the age of modern AI, you know,
我很想听听你的看法,你知道,关于在现代 AI 时代做一个语言学家,
06:34
transformers, statistical models, you know, this is an NLP used to be kind of coming from
transformers、statistical models,你知道,NLP 过去有点像是从
06:40
a linguistic perspective.
语言学视角出发的。
06:42
And now the entire field is shifted to a statistical perspective.
而现在整个领域已经转向了 statistical perspective。
06:46
And I'd love to hear your reflections on, you know, being on the other side of that
而且我很想听听你的思考,你知道,关于身处那个
06:50
transition, as well as maybe more importantly, ways that you think that kind of the traditional
过渡的另一边,以及也许更重要的是,你认为传统的
06:56
foundational linguistics is still important to the way we think about AI today.
foundational linguistics 在哪些方面仍然对今天思考 AI 的方式很重要。
07:02
These questions are on my mind all the time, yeah, because I operate at the intersection
这些问题一直在我脑海里,是的,因为我就在交叉点上
07:07
of all these different fields.
在所有这些不同领域里。
07:08
And I will say it's useful to distinguish in this context linguistics, you know, people
而且我想说,在这个语境下,把语言学区分出来是很有用的,你知道,有些人
07:13
in my department at Stanford study language and social identity, historical linguistics,
在 Stanford 我所在的系里研究语言和社会身份、历史语言学,
07:19
the structure of language.
语言的结构。
07:21
And they're just doing scientific investigation of language as a human phenomenon.
他们只是在把语言作为一种人类现象来做科学研究。
07:25
And they are not technologists and they're not trying to inform technology.
他们不是技术专家,也不是想为技术提供参考。
07:29
So their project is interestingly impacted by technological developments for NLP people
所以,对做 NLP 的人来说,他们的项目受到技术发展的影响很有意思,
07:35
who are, of course, participating directly in the engineering project.
当然,做 NLP 的人是直接参与这个工程项目的。
07:39
They're affected in a very different way by the rise of GNI and the kind of homogeneous
他们受到 GNI 崛起以及那种
07:43
nature of the solutions that people now adopt in that space.
同质化解决方案特性的影响,方式非常不同,而人们现在在这个领域正是采用这类方案。
07:47
So for the linguists, I feel like this is the most exciting moment that anyone could
所以对语言学家来说,我觉得这是任何人都能
07:52
have dreamed of.
梦想得到的最激动人心的时刻。
07:53
I feel incredibly privileged to be alive in this moment where humans encounter for the
我觉得无比幸运,能活在这个时刻:人类第一次
07:59
very first time non-human creatures that use our language very fluently.
遇到能非常流利地使用我们语言的非人类生物。
08:05
I think it's weirding us all out, but from the point of view of understanding the human
我觉得这让我们所有人都觉得怪怪的,但从理解人类
08:09
capacity for language, what a gift because you can ask about the mechanisms which are
语言能力的角度看,这是多好的礼物啊,因为你可以去问那些机制,它们是
08:15
different from humans, but obviously sufficient for achieving a certain kind of behavioral
与人类不同,但显然足以实现某种 behavioral
08:21
performance.
performance。
08:23
We can think about them as investigative tools.
我们可以把它们看作研究工具。
08:25
I mean, we train them on the internet.
我的意思是,我们在互联网上训练它们。
08:26
They're basically incredibly powerful distributional learners, and we can learn a lot from them
它们基本上就是极其强大的 distributional learners,而且我们可以从它们身上学到很多
08:31
about the true structure of language by just looking at the kinds of things that they learn.
关于语言真正结构的东西,只要看看它们学到的那些东西就行。
08:36
And it really gets at the heart of core questions in linguistics about how much of language
而这真的触及了语言学核心问题的核心:语言学习有多少
08:41
learning is innate and the nature of our capacity and whether it's statistical or symbolic.
是 innate 的,我们这种能力的本质是什么,以及它到底是 statistical 还是 symbolic。
08:46
All those things come flooding in in a completely fresh way.
所有这些都以一种完全崭新的方式涌进来。
08:50
And so whatever your reaction to language models is, it should be a significant one, right?
所以不管你对 language models 的反应是什么,它都应该是很重大的,对吧?
08:56
This should be causing you to rethink key questions.
这应该会促使你重新思考一些关键问题。
09:00
And that's all you could hope for as a scientist that you have new angles, new perspectives,
而作为一名科学家,你能期望的也就这些了:你有了新的角度、新的视角,
09:04
new questions reopened.
新的问题被重新打开。
09:06
That's been incredible.
这真是太不可思议了。
09:09
For NLP, I think it's a more uncertain prospect because pre the arrival of like pre-trained
对 NLP 来说,我觉得它的前景更不确定,因为在像 pre-trained
09:18
models, which for me would be like the LMO model back in 2017, 2018, before that there
models 出现之前——对我来说,那就像是 2017、2018 年的 LMO model——在那之前,有
09:25
was still a lot of statistical work, of course, and we were in the deep learning era.
当然,还是有很多 statistical work,而且我们当时处在 deep learning 时代。
09:29
But you could still, for example, do a PhD that was entirely about some specific phenomenon
但你仍然可以,比如,读一个 PhD,完全围绕某个具体现象
09:34
and maybe some very specific tweak to a model.
或者也许是对模型做某个非常具体的 tweak。
09:37
So you could say, I'm going to work on summarization and I've got a new idea about how to do that
所以你可以说,我要做 summarization,我有个新想法,关于怎么
09:41
well using deep learning models and that could be your PhD.
用 deep learning 模型把它做好,那就可以是你的 PhD。
09:45
And what we started to see 2018, 2019, 2020, especially with the arrival of GPT-3, that
而我们开始在 2018、2019、2020 年看到,尤其是随着 GPT-3 的到来,那
09:51
that was a very uncertain prospect because you might wake up one morning to find that
那是一个非常不确定的前景,因为你可能某天早上一醒来就发现
09:55
you had been completely scooped, that with essentially no effort, one of these large
你已经被完全 scoop 了,就是几乎不费吹灰之力,这些大型的
09:59
pre-training runs had done better than you at the thing that you'd worked so hard on.
pre-training 实验在你花了那么多心血的那件事上,反而比你做得更好。
10:04
And that caused an interesting, probably overall productive, but interesting and challenging
而这引发了一场很有意思、整体上大概也算有成效、但既有趣又充满挑战的
10:08
crisis for people, especially students who are trying to figure out what to do next with
危机,对很多人来说尤其如此,特别是那些正在琢磨接下来该拿
10:13
their PhD research.
自己的 PhD research 怎么办的学生。
10:15
But I think all of us felt a kind of real uncertainty in that moment.
但我觉得,在那一刻,我们所有人都感受到了一种真切的不确定感。
10:18
Yeah, I remember the anxiety of that time and I always felt it was kind of expressed
是啊,我记得那时的焦虑,而且我一直觉得它某种程度上表现为
10:26
as, you know, is research and NLP fundamentally like scale-limited already need a certain degree
就是,你知道,研究和 NLP 是不是从根本上就已经是 scale-limited 了,是不是已经需要一定程度的
10:33
of scale that only a handful of organizations have to do foundational research and is everyone
scale,只有少数几个组织才拥有,才能做基础研究,那是不是每个人
10:41
else going to be relegated to like poking the beast and seeing what it does.
否则就只会沦为去戳一戳那头野兽,看看它会有什么反应。
10:48
And I'm curious, do you feel like that was an anxiety that's passed or is it still very
我很好奇,你觉得那是一种已经过去的焦虑,还是它现在依然非常强烈?
10:55
Has it panned out quite like that?
那你是怎么——你知道的——这件事对你来说最后是怎么解决的?
10:58
How do you, you know, how's it been resolved for you?
太有意思了。
11:01
Not resolved.
这是我经常和我的合作者以及学生们讨论的事情。
11:02
It's something I discuss a lot with my collaborators and with my students.
这是我经常跟我的合作者还有我的学生讨论的一件事。
11:06
We're all trying to figure this out in this moment.
这一刻,我们都还在努力把这件事弄明白。
11:09
I will say one concrete thing we did was orient a lot of our research toward interpretability.
我会说,我们做过的一件具体的事,就是把很多研究都转向了 interpretability。
11:16
Just the project of understanding how these models end up being so good at such hard tasks.
就是去理解这些 models 到底是怎么变得这么擅长这么难的任务的。
11:21
And the reason we did that is it's relatively inexpensive and it's also an area where clearly
我们这么做的原因是,它相对不那么贵,而且这也是一个领域,很明显
11:26
you would be explicitly hoping that models would get better because then there would be more
你会明确地希望 models 变得更好,因为那样就会有更多
11:33
Versus if you were doing that summarization project, you might quietly be hoping that there
而如果你在做那个 summarization 项目,你可能会暗自希望
11:37
wasn't going to be so much progress so that you could make the progress.
不会出现那么大的进展,这样你才能取得进展。
11:41
Let's hope the next model isn't good at summarization.
希望下一个 model 别太擅长 summarization。
11:43
I want to be the star of that show.
我想成为那个节目的明星。
11:45
That's, as I said, very uncertain.
就像我说的,这事儿非常不确定。
11:47
But if you're doing that kind of trip, you're like, let's get the new model released because
但如果你在做那种旅程,你会想,咱们赶紧把新的 model 发布出来吧,因为
11:49
now we're going to have even more structure to find, even more to explain.
现在我们要找的结构会更多,要解释的也会更多。
11:53
And that felt like a very productive choice.
而且感觉那是个非常有成效的选择。
11:56
I don't want to leave out the fact that it's also cheaper to do this research and that
我不想漏掉一个事实,那就是做这项研究也更便宜,而且
12:00
And then I would say that right now, a lot of us are in a moment of thinking, we should
然后我会说,现在,我们很多人都在想,我们应该
12:05
do stuff that is weird and creative and out of the mainstream.
做一些古怪、有创意、脱离主流的事情。
12:09
We should be thinking about trying to achieve the next big thing because competing with these
我们应该琢磨着去实现下一个大事件,因为跟这些
12:15
massively resourced, incredibly creative and talented teams is just not a winning game.
资源极其雄厚、极具创造力和才华的团队竞争,根本赢不了。
12:21
So let's play a different game and hope that that's, as they say, where the puck is
所以,让我们玩点不一样的,并希望那就是人们说的,冰球要去
12:25
going, not where it is.
的地方,而不是它现在所在的地方。
12:27
What are some examples of that kind of thinking?
这种思维方式有哪些例子?
12:29
We've been thinking a lot about architectures because I have a lot of complaints about current
我们一直在大量思考 architectures,因为我对当前的
12:33
architectures.
architectures。
12:34
And I would say the other main theme right now for us in my group is thinking about data.
而且我想说,现在我们组里另一个主要主题就是思考 data。
12:40
You know, data have strange and wondrous properties.
你知道,data 有一些奇怪又奇妙的特性。
12:43
I think we don't understand how data affect models.
我觉得我们并不理解 data 是如何影响 models 的。
12:46
And that has all sorts of implications for security and safety and also the nature of the learning
而这对于 security 和 safety,以及 learning 的本质,都有各种各样的影响
12:51
that these models do.
也就是这些 models 所进行的 learning。
12:52
And really, data is fundamental.
而且说真的,data 是根本。
12:53
It's all data-driven learning.
这全都是 data-driven learning。
12:55
And so telling the full causal story from data to final model state feels like it will
所以,把从数据到 final model state 的完整因果故事讲清楚,感觉会对很多问题都意义重大。
13:00
just be significant for lots of questions.
但我不想漏掉 architecture 这一块,因为我觉得大家最终都走到的那套 architecture,也就是我们做得非常深、非常大的这些 staffed transformers,效率极其低下。
13:04
But I wouldn't want to leave out the architecture one because I feel like the architecture everyone
我们会希望它们利用所有这些 depth 和 representational power,去学习处理各种东西的 modular、recursive functions,以及各种令人兴奋的东西。
13:10
has arrived at, these staffed transformers that we make very deep and very large are tremendously
但我们发现的并不是这样,而这似乎是一个真正的机会,让我们直接 level up、做得更好。
13:17
We would hope they were using all that depth and all that representational power to learn
我们本来会希望它们是在利用所有这些 depth 和所有这些 representational power 来学习
13:23
modular, recursive functions for things and all sorts of exciting stuff.
modular, recursive functions,用来处理各种东西,以及各种令人兴奋的东西。
13:28
It is not what we find and that seems like a real opportunity to just level up and do better.
但这并不是我们发现的情况,而这看起来是一个真正的机会,可以好好升级一下、做得更好。
13:33
And maybe we could get massively more capable models with half the depth and a quarter of
也许我们能得到能力强大得多的模型,只要一半的 depth 和四分之一的
13:38
the representational width.
representational width。
13:40
And that would be transformative for the economics of AI in addition to leading to all sorts
而且这对 AI 的经济性会是变革性的,同时还会带来各种
13:44
of exciting things for capabilities.
令人兴奋的 capabilities 相关成果。
13:47
It's funny and maybe a bit validating for me to hear you say that because whenever I articulate
听你这么说,我觉得挺有意思,也许还让我有点被认可,因为每当我表达
13:55
a thought in that direction, particularly with folks that are coming from the frontier
那个方向的想法时,尤其是跟来自 frontier
14:01
labs or essentially the frontier labs, I get back this feeling that you're just not
labs,或者说本质上就是 frontier labs 的人聊时,我得到的反馈是,你只是不够
14:10
bit or less impilled enough like structures, that's old school thinking, you know, you're
bitter lesson-pilled,就像,结构这东西是老派思维,你知道,你……
14:17
just trying to like train some features, just collect a lot of data, throw it at the model
就是试着 train 一些 features,收集一大堆 data,然后一股脑丢给 model
14:22
and that's all you need.
而这就是你需要做的全部。
14:23
Okay, but here's my response to them.
好,但这是我给他们的回应。
14:26
Let's say rewind to 2017, we've got the transformer.
咱们假设倒回 2017 年,我们有了 transformer。
14:30
It's got absolute positional encodings and it's got a particular structure for its MLP
它有 absolute positional encodings,而且它的 MLP
14:36
layer, which is pretty narrow and pretty dense and a certain structure to its activations
layer 有一种特定的结构,它相当窄、相当 dense,而且对它的 activations
14:42
and its layer norms.
和它的 layer norms 有某种结构。
14:46
The bitter less impilled thing to do would be to scale that up.
按照 bitter lesson 所暗示的,该做的就是把它 scale up。
14:50
But just consider, for example, how much it would cost to use the N-squared attention and
但你就想想,比如,用 N-squared attention 要花多少钱,而且
14:56
the absolute positional encodings would have a context window of 1 million.
absolute positional encodings 会有 1 million 的 context window。
15:00
This is the bitter less impilled thing, right?
这就是 bitter lesson 所暗示的做法,对吧?
15:02
Just keep scaling, but it would be absurd.
就一直 scale 下去,但这会很荒谬。
15:04
It would cost trillions of dollars to produce models that we all interact with right now.
要做出我们现在都在用的那些 models,得花数万亿美元。
15:08
What did people do instead?
那人们转而做了什么?
15:09
They thought hard about locality and they thought about how like positional encoding should
他们认真思考了 locality,也思考了 positional encoding 应该怎样
15:13
be favoring local relationships.
更偏向局部关系。
15:16
They completely rethought the MLP so that it's now wide and sparse.
他们彻底重新思考了 MLP,让它现在又宽又稀疏。
15:20
Everyone did careful work on the activation functions to make sure there weren't weird
每个人都在 activation functions 上做了细致的工作,以确保没有奇怪的
15:24
outliers so that they could quantize in a good way and so forth and so on.
outliers,这样他们就能以良好的方式进行 quantize,诸如此类,等等等等。
15:29
All of this analysis work built on intuitions about data and learning led to the model
所有这些建立在关于数据和学习的直觉之上的分析工作,最终造就了
15:33
that we have now, which is like a ship of theses compared to the 2017 transformer.
我们现在的模型;与 2017 年的 transformer 相比,它就像一艘由论文组成的船。
15:37
The only thing that survives is attention and the feed forward layer.
唯一幸存下来的是 attention 和 feed forward layer。
15:42
And I claim for you that none of that stuff is bitter less impilled.
而且我向你断言,那些东西没有一样是 Bitter Lesson 驱动的。
15:45
That was all analysis work that was meant to save based on priors and the data and priors
那全都是分析工作,本来是想基于 priors 和 data 以及 priors 来节省。
15:50
about how they knew learning would happen.
关于他们是怎么知道 learning 会发生的。
15:52
So I go back at them.
所以我就回头去怼他们。
15:53
You're not bitter less impilled enough apparently, although this is a reductio, I think.
显然,你并不 bitter,只是还不够被推动,不过我觉得这是一种 reductio。
15:58
Oh, I love this.
哦,我太喜欢这个了。
16:00
That's such a great response.
这个回应真是太棒了。
16:02
I think it also really calls out the relationship between data,
我觉得它也真的点出了 data 与
16:09
Macinturte and efficiency like core themes that you've been focused on and how they interrelate
Macinturte 和 efficiency 之间的关系,就像你一直关注的核心主题,以及它们如何相互关联
16:16
and support one another.
并且互相支持。
16:18
Yeah, absolutely.
对,绝对没错。
16:19
And this relates to one of my hot takes, you know, it's very fashionable, especially
而这和我一个暴论有关,你知道,现在特别流行,尤其是
16:22
among interpresearchers, but I think in general for people to say, we don't understand
在 interp researchers 当中,但我觉得总体来说,大家会说,我们不理解
16:26
how these models work.
这些 models 是怎么工作的。
16:27
It is also very mysterious to us.
对我们来说,它也非常神秘。
16:29
But the truth is that people in the field have very deep intuitions about how these models
但事实是,业内的人对这些 models 如何
16:34
work and that is the causal factor in us making so much progress because they could think
工作有非常深刻的直觉,而这就是我们能取得这么大进展的 causal factor,因为他们能够思考
16:39
analytically what would the structure of positional encodings and attention be so that I could
从分析上讲,positional encodings 和 attention 的结构得是什么样,我才能在 million context scale 上做到这一点?
16:44
do this at million context scale.
你只有基于深入的分析和洞察才能做到那种事,光靠猜是不行的。
16:47
You can only achieve that kind of thing based on deep analysis and insight, not by just
所以当人们说,哦,我们不懂,我就会说,我觉得你们懂的比你们表现出来的多得多。
16:51
guessing.
我觉得你们是懂的,至少不亚于我的修车师傅懂我的车是怎么运作的。
16:53
And so when people say, oh, we don't understand, I say, I think you understand much better than
确实有一些谜团,但你可以采取很多行动,并且在改进方面非常有效。
16:58
you're letting on.
你是在暗示。
16:59
I think you understand, at least as well as my car mechanic understands how my car works.
我觉得你懂,至少跟我那修车师傅懂我的车怎么运作差不多。
17:03
There are mysteries, but you can take a lot of action and be very effective in improving
有些东西仍然是谜,但你可以采取很多行动,并且在改进上非常有效。
17:08
Why do you think they say that?
你觉得他们为什么会这么说?
17:11
Why do you think they say that they don't understand the models?
你觉得他们为什么说自己不理解这些 models?
17:14
There's just got to be some payback there.
这里面肯定有什么回报。
17:15
It's probably a paradox of expertise, right?
这大概就是专业知识的悖论,对吧?
17:18
So the more you do know, the more you feel like there are also mysteries and it's hard
所以你知道得越多,就越觉得其中也还有谜团,而且很难
17:22
to step back from that and be objective and say, well, we did make a phenomenal amount
从中抽身出来、保持客观地说,嗯,我们确实取得了惊人的
17:26
of progress and that can't be just because of happenstance.
进步,而这不可能只是偶然。
17:29
That was because we know a lot, but all you see as an expert is all the things that are
那是因为我们知道很多,但作为专家,你看到的全都是那些
17:34
still to be explained.
还没被解释清楚的东西。
17:37
Partly also, it's just a narrative in the field and it does stretch back to days when I think
部分原因也在于,这只是这个领域里的一种说法,而且它确实可以追溯到那些我觉得
17:40
we had very little understanding of how these models worked and possibly because a lot
我们对这些 models 如何运作了解得还非常少的日子,也可能是因为很多
17:44
of them weren't that good that was very little to explain.
models 当时都没那么好,所以没什么可解释的。
17:47
And so that's just been slow to catch up with how much progress we have made in understanding
所以,这就一直迟迟没能跟上我们在理解方面已经取得了多大进展
17:52
the kind of intuitive human level mechanisms that these models are operating with.
也就是这些 models 运作时采用的那种直觉性的、人类水平的机制。
17:57
Also wanted to ask you about DSPY, I forgot about this as we were talking earlier, but you
另外也想问问你关于 DSPY 的事,我们之前聊的时候我忘了这个,但你
18:04
were involved in DSPY, which, well, I'll let you talk about it, but I'm curious how
参与了 DSPY,这个嘛,我会让你来讲,但我很好奇它怎么和你的研究联系起来,还有,你知道,我们聊过的一些支柱。
18:12
it connects into your research and like, you know, some of these pillars that we've talked
哦,有很多很棒的脉络。
18:18
about.
而且我能给你提供一个多么 meta 的线索啊,因为我们刚才在聊做研究要有策略。
18:19
Oh, there's lots of wonderful strands.
这确实源自我的学生 Omar Qatab。
18:21
And what a meta strand I could offer you because we were talking about being strategic
他是 DSPY 背后的远见者,现在仍然是它的负责人。
18:26
This does stem from Omar Qatab, my student.
这确实源自 Omar Qatab,我的学生。
18:29
He's the visionary behind DSPY and still its lead.
他是 DSPY 背后的远见者,现在也仍然是它的负责人。
18:33
And he just had the intuition early on that we should rethink what it means to make a scientific
他一开始就有一种直觉:我们得重新思考,做出科学贡献到底意味着什么。
18:37
contribution.
以前,我们总是从论文的角度来看,把论文当成你所能做的这类贡献的起点和终点。
18:38
Previously, we saw it in terms of papers as the beginning and the end of all of this kind
他说,我们应该转而思考项目,思考如何赋能人们。
18:43
of thing that you would contribute.
所以对他来说,论文只是一个更广泛贡献的一部分,而这个贡献实际上可能以 open source 或 open weights 的 release 为核心,让人们能够做成大事。
18:45
We should instead, he said, think about projects and about empowering people.
他说,我们应该转而思考项目,思考如何赋能人们。
18:50
And so for him, the paper is one part of a broader contribution that might actually be
所以对他来说,这篇论文只是一个更广泛贡献的一部分,而这个贡献实际上可能
18:55
centered on an open source or open weights release that would allow people to do big
以 open source 或 open weights 发布为中心,让人们能够做
19:02
And that's where you find impact and that's the nature of a contribution going forward.
而那就是你能看到影响力的地方,也是未来贡献的本质。
19:06
And DSPY is a kind of embodiment of that, although he made a similar investment with
而 DSPY 就是这一点的一种体现,尽管他在
19:10
the Colbert retrieval model.
Colbert retrieval model 上也做了类似的投入。
19:13
And then people built on what he did.
然后人们在他所做的基础上继续发展。
19:15
And then you really saw it take off where open source contributions made it easier and
然后你真的看到它起飞了,open source 的贡献让那项技术变得更容易用,而且
19:19
easier to use that technology leading to more and more impact.
越来越容易使用,从而带来了越来越多的影响力。
19:22
And of course, DSPY is another wonderful example because in investing in this community,
当然,DSPY 是另一个绝佳的例子,因为通过投资这个社区,
19:29
and in the open source resource itself, he built a huge following.
以及这个 open source 资源本身,他建立了一个庞大的追随者群体。
19:33
There are lots of startups, mine included, where the core tech stack for the LMS is built
有很多初创公司,包括我的那家,它们的 LMS 核心 tech stack 都建立在
19:37
on DSPY, and that has made life so much easier.
DSPY 上,而这让日子好过太多了。
19:41
And then of course, it was a platform for him and for us to really think in an innovative
然后当然,对他和我们来说,它也是一个平台,让我们真正以创新的
19:46
way about prompt optimization and agentic workflows and all of those things.
方式去思考 prompt optimization、agentic workflows 以及所有这些东西。
19:51
Yeah, I was thinking not too long ago the degree to which model strength as a correlate
是啊,我不久前还在想,model strength 作为与
20:00
to model size, I suppose, and capability has kind of overcome the need for an explicit
model size 相关的东西,我猜,以及 capability,在多大程度上已经有点让我们不再需要一个显式的
20:09
framework like DSPY, DSPY.
像 DSPY、DSPY 这样的 framework。
20:12
Yeah, there are kind of two levels to that.
是啊,这里面其实有两个层面。
20:14
One would be just the engineering side where DSPY is great four years ago because it's
一个就是工程侧,DSPY 在四年前很厉害,因为
20:20
kind of hard to construct the code around one of these systems in a way that's modular
围绕这些系统之一来构建 code 有点难,要做到模块化
20:25
and reproducible and so forth because pecking out something where you've got a prompt
又可复现等等,因为你要敲出某个东西,里面有 prompt
20:29
string in the middle of your code with some slots in it, it's very error prone and it leads
string 夹在你的 code 中间,里面还有一些 slots,这非常容易出错,而且会导致
20:33
to bad system designs.
糟糕的 system designs。
20:35
And DSPY solved that.
而 DSPY 解决了那个问题。
20:36
And you could think that the need for that is diminishing somewhat because now we all
你可能会觉得,对这种东西的需求正在有所减弱,因为现在我们所有人都
20:41
specify these systems in English and have the coding agents do them.
用英语来指定这些系统,然后让 coding agents 去做。
20:45
And before that, there were another hundred frameworks that solve that particular part
而在那之前,还有另外一百个框架,专门解决拼图里的那个特定
20:52
Oh, yeah, there's always competition and I think at that level of just thinking about
哦,对,竞争总是有的,而且我觉得在那个层面,单纯考虑
20:55
programming interfaces and APIs, they can all learn from each other and so DSPY learned
programming interfaces 和 APIs 的时候,它们都能互相学习,所以 DSPY 从 PyTorch 学到了
21:00
a lot from PyTorch in terms of layer wise design and the kind of modularity that introduced
很多,在 layer-wise design 以及它引入的那种 modularity 方面
21:05
and then of course, you would hope that everyone kind of slurps up all these interesting
然后当然,你会希望每个人都能把所有这些有趣的
21:09
innovations and it leads to everyone being better.
创新都吸收进去,这会让每个人都变得更好。
21:12
There's lots of evidence of that at the level of interfaces.
在 interfaces 这个层面上,有很多证据能证明这一点。
21:14
I would maintain for you that even if we have agents actually writing the code for these
我还是会坚持跟你说,即使我们有 agents 真的在为这些
21:18
systems, it's great for us and for them if they write it in something that actually expresses
系统写 code,这对我们和他们都是好事,如果他们用某种能真正表达
21:22
these systems as modular components so that we can audit them so that they can change them.
这些系统为 modular components 的东西来写,这样我们能 audit 它们,这样它们也能修改这些系统。
21:27
It just feels like good engineering practices for any agent to think in a modular way and
这感觉就像,任何 agent 用 modular 的方式思考,都是很好的 engineering practices,而
21:32
that's what DSPY encodes.
这正是 DSPY 所编码的东西。
21:35
The other side is like the prompt optimization side and I believe people have that the need
另一边就像是 prompt optimization 那一侧,而我相信人们都有那种认识:
21:40
to be careful with your prompts is diminishing over time.
要小心对待你的 prompts 的必要性正在随着时间减弱。
21:43
I understand that narrative, but people should also, for example, just run like a simple
我理解这种叙事,但人们也应该,比如说,就直接跑一个简单的
21:49
annotation study where they use a few different models or the same model a few times on slightly
annotation study 里,他们会用几个不同的 models,或者同一个 model 在略有
21:54
different data.
不同的 data 上跑几次。
21:55
They will be blown away by the amount of variation that still exists.
他们会被仍然存在的 variation 之大给震到。
22:00
To be charitable, let's say that these LMs disagree about fundamental facts about how
说得客气点,这些 LMs 在一些基本事实上都达不成一致,比如到底该怎么
22:04
to label certain texts or what kind of response to give.
label 某些文本,或者该给什么样的 response。
22:08
We all kind of slip past this because we feel like, hey, they're smart and they're good
我们都有点把这一点滑过去了,因为我们觉得,嘿,它们很聪明,表现也很好
22:11
and they're getting better, but if you quantify it, it's pretty disturbing and the next step
而且还在变得更好,但如果你把它量化出来,其实挺让人不安的,而下一步
22:15
from that is to think about having all those agents optimize a prompt so that their behavior
就是去想,让所有这些 agents 去 optimize 一个 prompt,好让它们的行为
22:20
is at least consistent and then you're right back at that DSPY vision.
至少是一致的,然后你就又回到了那个 DSPY 愿景。
22:25
It's interesting that you say that because I don't feel like that necessarily aligns
有意思的是你这么说,因为我不觉得那一定符合
22:31
with my recent experience and in particular, one thing that I've noticed that's been surprising
我最近的经历,尤其是我注意到的一件很让人意外的事,
22:38
is how well aligned, I guess, maybe that's not the right word, but how similar the responses
就是它们有多 aligned,我猜,也许这个词不太对,但就是这些回复有多相似
22:47
I get to a query across different models.
我在不同模型上针对一个 query 得到的。
22:51
For example, these are often kind of what I would call a like a casual prompt, a casual
比如说,这些通常算是我会称之为一种随意的 prompt,一种随意的
22:58
query, something that I might type in the Google and it will now generate an LLM response
query,就是我可能会在 Google 里输入的东西,然后它现在会生成一个 LLM 回复
23:06
for me and it's kind of AI mode and I'll take the same thing and put it into ChatGPT and
给我,而且这有点像 AI mode,我会把同样的东西输入到 ChatGPT 里,然后
23:15
maybe Claude.
也可能是 Claude。
23:16
It surprises me that the responses are often very, very similar, like very similar structure,
让我惊讶的是,这些回答经常非常非常相似,像是结构非常相似,
23:25
very similar facts, very similar citations and stepping back, there are lots of ways
事实非常相似,引用非常相似,而且退一步看,其实有很多种方式
23:34
that they could answer or approach these different questions but it seems like the models
它们可以回答或处理这些不同的问题,但看起来 models
23:39
or the training or the system prompts or something is all kind of converged on something that
或者 training,或者 system prompts,或者别的什么,都好像收敛到了某个东西上,
23:44
makes the models express themselves very similarly which seems to be at odds with the last
让 models 表达自己的方式非常相似,这似乎和最后
23:52
thing you said about the need to optimize prompts or the impact of the individual prompt.
你说的关于需要优化 prompts,或者单个 prompt 的影响那一点相矛盾。
23:57
I'm open-minded, but for example, we did a thing recently we were writing a grant and
我持开放态度,但举个例子,我们最近做了件事,我们当时在写一份 grant,而且
24:02
we needed a title and you want to be strategic with these titles, so we come up with a
我们需要一个标题,而你想在这些标题上讲究策略,所以我们会自己先想出一
24:05
whole bunch of them ourselves and then we all disagree on what would be the best.
大堆标题,然后我们所有人对哪个最好都意见不一。
24:09
So let's find out what the agents think, so ask a few anthropic models and a few GPT
那我们就看看 agents 怎么想,去问几个 Anthropic models 和几个 GPT
24:15
models, which of these five titles, which is the best.
models:这五个标题里,哪个最好。
24:18
So you get a different answer from all of them, along with a detailed rationale about why
结果你会从它们每个那里得到不同的答案,还附上一份详细的 rationale,解释为什么
24:23
obviously, of course, the choice that the model has made in that moment is the best one.
显然、当然,model 在那一刻做出的选择就是最好的那个。
24:28
This is great because then we can think about which one of these arguments is most persuasive,
这很棒,因为这样我们就能想想这些论证里哪一个最有说服力,
24:32
but if you are hoping for consistency at a subjective labeling task, which this is one,
但如果你在一个 subjective labeling task 上指望一致性,而这正是这样一个任务,
24:38
you can see right there that you're going to have a real problem unless you give very specific
你能看到,就在那里,你将会遇到一个真正的问题,除非你给出非常具体的
24:41
criteria and then you're kind of also constructing a prompt for them and you might want to
标准,然后你其实也是在给它们构造一个 prompt,而你可能想
24:47
manage them differently.
用不同的方式管理它们。
24:48
There is a real, I don't have evidence for this yet, but we have an intuition at Bigspin
确实有一个真实的——我现在还没有证据,但我们在 Bigspin 有一种直觉,
24:53
and the research we've done that you get a kind of paradox that the more requirements you
以及我们做过的研究显示,你会遇到一种悖论:你加的要求
24:57
add, actually the more variation you'll see because the different models will key into
越多,实际上会看到越多的变化,因为不同的 models 会抓住
25:01
different sub parts of the requirements.
要求的不同子部分。
25:04
And since they do it very concertedly, you can actually get systematically biased behavior
而且因为它们非常协同地这么做,你实际上会得到带有系统性 bias 的行为。
25:10
from something that you thought was a very good specification.
从一个你原本觉得非常好的 specification 出发。
25:14
And that again calls for this idea that what you need to do is figure out what the label
而这又回到了这个思路:你需要做的是弄清楚 label
25:17
was ought to look like and then have some automatic optimization process get the model there.
本来应该长什么样,然后让某种自动优化过程把模型带到那个状态。
25:24
And that's what things like Jepa and Meet Pro were for.
而 Jepa 和 Meet Pro 这类东西就是干这个的。
25:27
A topic that I really wanted to, a topic that I would really like to dig into with you
有一个话题我一直很想,有一个话题我特别想跟你深入聊聊
25:33
based on our previous conversation was the idea of tokenomics.
基于我们之前的对话,就是 tokenomics 这个概念。
25:38
It's something that people are talking about a lot recently.
这是最近大家聊得很多的一个东西。
25:44
I think folks that use Cloud Code, for example, have a very visceral experience with anthropic
我觉得比如用 Cloud Code 的人,对 anthropic 会有非常切身的体验
25:50
changing the terms around uses, but it's happening under the covers with all of these large providers.
改变术语的使用方式,但这正在所有这些大型 provider 的底层发生。
25:57
And so I think way more now than six months ago, we're all a little antsy with the relationship
所以我觉得现在比六个月前,我们对与这些大型 model provider 的关系以及我们获得的价值都更加焦虑不安。
26:08
we have with these big model providers and the value that we get.
我们最近写了一篇文章讨论这个问题。
26:12
We recently wrote an article about this.
稍微谈谈这与你们更广泛的研究有什么关联,同时也聊聊当你们开始深入这个领域时发现的一些事情。
26:16
Talk a little bit about how it ties into your broader research, but also some of the things
是的,又是一个身处其中的有趣时刻,因为我们所有人都开始意识到所有这些 AI 使用的真实成本。
26:25
that you found when you started to dig into this area.
当你开始深入这个领域时,你所发现的东西。
26:29
Yeah, another interesting moment to be in, as we're all being made aware of the true
是的,又是一个很有意思的时刻,因为我们都开始意识到所有这些AI使用量的真实成本。
26:36
cost of all this AI usage.
我看到Ed Zitron发了一条推文,就是一张截图,有人注意到Copilot告诉他们上个月花了500美元,如果他们继续以现在这种方式使用Copilot的新构建,下个月就会是11,000美元,这真是实打实的吓人。
26:37
I saw a tweet from Ed Zitron, just a screenshot from someone who was noticing that Copilot
这里的类比就像,我以前坐网约车去机场,Uber或Lyft,花20美元,现在要90美元,但这更像是从20美元到500美元之类的,
26:44
was telling them that their last month was $500, and if they keep up the way they are
我当时在跟他们说,他们上个月花了 $500,要是他们继续照现在这样
26:49
with Copilot's new building, it will be $11,000 in the next month, which is real sticker
用 Copilot 的新 building,下个月就会变成 $11,000,这真是实打实的
26:57
And the analogy here is like, I used to cost me $20 to take a ride share to the airport
这里的类比就像,我以前打网约车去机场要花 $20
27:03
Uber or Lyft, and now it costs $90, but it's more like $20 to like $500 or something,
Uber 或 Lyft,现在要花 $90,但这更像是从 $20 到 $500 之类的,
27:13
If only the slope will be as shallow as it were, right?
如果斜率能像以前那么平缓就好了,对吧?
27:17
We start to wish for those easier stories.
我们就开始怀念那些更轻松的故事了。
27:20
And so what will happen?
那会发生什么呢?
27:21
I mean, I think what's happening is that the big providers are testing the waters on
我的意思是,我觉得现在的情况是,那些大provider正在试探性地
27:25
charging us the true costs, plus whatever profit they need to make as they all try to
向我们收取真实成本,再加上他们为了准备IPO之类的事情所需要赚取的利润。
27:30
gear up for IPOs and so forth.
为 IPO 之类的事情做准备。
27:33
And in turn, that is very quickly leading people to ask questions like, what is the return
反过来,这也很快让人们开始问这样的问题:
27:37
on investment for all these tokens that we have purchased?
我们买的这些 tokens,投资回报率到底是多少?
27:42
And it's a very tricky area to be in because what does it mean to think about value in
而这是一个非常棘手的领域,因为在这个语境下,
27:47
this context, even if we focus in on people who are doing just coding with coding agents?
思考价值意味着什么,即使我们只聚焦在那些用 coding agents 做编程的人身上?
27:53
Can we agree on what it means to add value?
我们能对“增加价值”的含义达成一致吗?
27:55
I mean, we have a few measures in mind like making a pull request or committed lines of
我的意思是,我们心里有几个衡量标准,比如发起一个 pull request,或者 committed lines of code
28:01
code that last in the repo for a while, or documentation touched or skill files created,
能在 repo 里留存一段时间,或者 documentation 被改动,或者 skill files 被创建,
28:08
but we might also worry that that's not capturing the value for many kinds of sessions we have,
但我们也可能担心,这并没有捕捉到我们许多种 sessions 的价值,
28:14
which are more open-ended and about discovery.
这些更开放,也更关乎探索发现。
28:16
So that's the first question is just solving this value issue, right?
所以第一个问题就是解决这个价值问题,对吧?
28:20
Let's just agree on what it would mean to add value for a coding agent.
我们先就什么叫做为 coding agent 增加价值达成一致吧。
28:24
I like this line of inquiry because, to me, it's the response to this thing that drives
我喜欢这条探究思路,因为对我来说,它是对这件事的回应,这件事让我
28:33
me crazy, which is, oh, big tech company CEO.
发疯,那就是,哦,大型科技公司 CEO。
28:41
This year, 95% of our code will be generated by AI.
今年,我们 95% 的代码将由 AI 生成。
28:47
It's like, yeah, what does that really mean at that level?
就好像,是啊,在那个层面上这到底意味着什么?
28:53
Out of the details beneath there, but is that a good thing for a man?
抛开底下那些细节不谈,但那对一个人来说是好事吗?
28:59
Oh, another dimension, right?
哦,还有一个维度,对吧?
29:00
Which is that code of liability or an asset?
也就是说,那个 code 到底是负债还是资产?
29:05
And this idea of like you articulate as kind of code longevity in the code base, that's
而且你把它阐述成 code base 里的某种 code longevity,这个想法挺有意思的。
29:10
an interesting way to think about it.
可能有很多有意思的思考角度,但现在几乎没人在想。
29:11
There's probably a lot of interesting ways to think about it that very few are thinking
大概有很多有趣的思考角度,但现在很少有人会想到
29:15
about right now.
这些,就在当下。
29:16
And all of these fall victim to the standard thing that once you make it a metric, it's
而所有这些都会掉进那个老问题里:一旦你把它变成一个 metric,它
29:19
no longer useful to you.
就不再对你有用了。
29:21
If we said, well, it's completion of projects, right?
如果我们说,好吧,那就是看项目的完成情况,对吧?
29:24
Well, then everyone would just have many projects that they completed, but they could all
那大家就只会有一大堆自己完成的项目,但它们可能全都
29:28
be liabilities and add very little value.
是 liabilities,带来的价值少得可怜。
29:31
But one framework we could offer that we did in the research you alluded to is, let's
但我们可以提供的一个 framework,是我们在你提到的那项研究里做过的,那就是,让我们
29:34
think about this like economist might, so we might have like a consumer price index.
像经济学家那样来思考这件事,比如我们可能会有类似 consumer price index 的东西。
29:39
And the first step will be what's the basket of goods that we're going to consider, you
而第一步会是,我们要考虑的 basket of goods 是什么,你
29:44
In that standard land, it would be like the price of eggs and the cost of rent and other
在那种标准语境里,它就会像鸡蛋的价格、房租和其他
29:48
kinds of tangible goods.
各种有形商品。
29:50
What are engineering goods that we might track?
那我们可能追踪的 engineering goods 有哪些?
29:52
But eggs might be a summary or a pull request, a bug fix or something like that.
但鸡蛋可能是一个 summary、一个 pull request、一个 bug fix 之类的东西。
29:58
Or we could think broadly because we both use these coding agents, requirement discovery,
或者我们可以想得更宽泛,因为我们都用这些 coding agents、做 requirement discovery,
30:05
Knowledge accumulation.
Knowledge accumulation。
30:07
These are things that we don't currently track, of course, even as engineers, but might
这些是我们目前不追踪的东西,当然,即使作为工程师,但它们可能
30:11
be behind our intuition that these coding agents are making us productive, even if it's
是我们直觉背后的原因,即这些 coding agents 让我们更高效,即使它
30:15
not reflected in the PR counts or whatever, right?
没有反映在 PR 数量或其他什么上,对吧?
30:18
I mean, in a sophisticated approach, you might say, I don't want more PRs because this
我是说,在一个成熟的方法中,你可能会说,我不想要更多的 PRs,因为这
30:22
is just a certain kind of busy work that doesn't relate to the actual goals I have.
只是某种瞎忙,与我实际的目标无关。
30:27
What are the actual goals?
实际的目标是什么?
30:28
It's completing valuable projects and so forth.
是完成有价值的项目等等。
30:30
I could do it with fewer PRs, but I had, you know, really robust code.
我可以用更少的 PRs 来完成,但是,你知道,我有真正 robust code。
30:35
I'd be possibly happy with that.
我可能对此会挺满意。
30:37
So we got to figure out what the basket of goods is, but then we could start to track it relative
所以我们得搞清楚 basket of goods 是什么,但之后我们就可以开始跟踪它相对于
30:41
to token usage.
token usage 的情况。
30:42
And that would be the consumer price index.
那就会是 consumer price index。
30:44
So for any time period, we could just say, I've got my tokens spent and I've got my
所以对任何时间段,我们都可以直接说,我有我花掉的 tokens,我也有我
30:48
goods produced, tokens divided by goods produced is a pretty rough measure of the purchasing
产出的 goods,tokens 除以 goods produced 是衡量那些时间段里 token 的 purchasing
30:56
power of the tokens in those time periods.
power 的一个相当粗略的指标。
30:59
Then you would do the standard consumer price index thing of making what they call a hedonic
然后你就会做标准的 consumer price index 那套,做他们所谓的 hedonic
31:04
So you could just say, maybe quality is improving over time, so you'd pick some measure
所以你可以就说,也许质量随着时间在提升,那你会选某个指标
31:07
for that and make an adjustment to the line.
据此对那条线做个调整。
31:10
And when we did that study, we did code survival.
我们做那项研究时,用的是 code survival。
31:13
So the number of lines of code that survives more than four days in the repository, we made
就是在 repository 里存活超过四天的 lines of code 数量,我们做了
31:18
an adjustment upward because that rate is going up.
向上调整,因为这个比率在上升。
31:21
That's surprisingly short.
这短得有点意外。
31:24
Four days survival.
四天的生存期。
31:26
So again, all this is around measurement and I'm happy to just be starting this discourse
所以再说一次,这一切都是围绕 measurement 的,而我很高兴能刚开始这场讨论
31:30
because we can see it's important to the economics of AI and it seems like the work
因为我们可以看到,这对 AI 经济学很重要,而且看起来这项工作
31:33
isn't being done at a high enough rate for us to get a clear picture.
没有以足够高的速度推进,让我们无法看清全貌。
31:37
So we could make it longer and maybe the adjustment would be different.
所以我们可以把它做长一点,也许调整就会不一样。
31:40
I think currently for the data we have, which is this sweet chat benchmark, which was released
我觉得目前就我们有的数据来说,也就是这个 sweet chat benchmark,它是由
31:45
by researchers at Stanford, it's about 6,000 real coding sessions, all the metadata, everything
Stanford 的研究人员发布的,大约有 6,000 个真实的 coding sessions,所有 metadata,所有东西
31:52
What we see with Opus 4.6 usage in the time period we have, which is February to mid-April
在我们目前有的这个时间段里,也就是今年2月到4月中旬,我们从 Opus 4.6 的使用中看到的是,tokens 的购买力在下降。
31:58
of this year, a decline in the purchasing power of tokens.
这个 CPI 正在下降。
32:02
That CPI is going down.
再说一次,我只是想提出这个问题:是因为我们的商品篮子选错了,还是因为我们实际上从这些 tokens 中得到的价值更少了?
32:05
And again, I just want to open the question, is it because we have the wrong basket of
有一点我可以肯定,这背后绝对有一个因果因素:今年2月,大多数 tokens 都用于生成 code,而这跟我们刚才谈到的结果有关。
32:10
goods or is it because we're actually getting less value from these tokens?
商品,还是因为我们实际上从这些 tokens 中获得的价值更少了?
32:15
The one thing I can say that's definitely a causal factor here is that in February of
我能确定地说,这里肯定有一个因果因素,那就是
32:21
this year, most of the tokens went to producing code, which relates to the outcomes we just talked
在今年二月,大部分 tokens 都用于生成 code,这和我们刚谈到的结果有关
32:28
By mid-April, it was quite split between code generation thinking and also explanation
到四月中旬,它在 code generation 的思考和
32:34
to the user.
给用户的解释之间已经相当分裂了。
32:36
And so that split now is going to have an effect on the things we're measuring and that
所以这种分裂现在会影响我们正在衡量的东西,而那
32:40
might be cause for reflection.
可能值得反思。
32:42
There's value in those explanations that's not reflected in PRs, but might be reflected
那些解释里的价值并没有反映在 PRs 里,但可能反映
32:46
in something like knowledge discovery.
在像 knowledge discovery 这样的东西里。
32:48
And I see that coming up within the same time frame, it's become very common to now
而且我看到,在同一时间段内,这一点也出现了,现在变得非常常见的是
32:55
talk about the token efficiency of a new model that's been released with the implication
聊聊一个刚发布的新模型的 token efficiency,它暗示着
33:01
of being tokens of internal use tokens, thinking tokens versus, you know, per token
这些是 internal use tokens、thinking tokens,相对于,你知道,每个
33:09
of output, I guess is maybe a way to think about it.
output token,我猜这也许是一种理解方式。
33:11
Another fascinating dimension.
另一个很有意思的维度。
33:13
And this actually relates all the way back to the theme of efficiency for these architectures.
而这其实一路都能联系回这些架构的效率主题。
33:17
So here's a claim I'll make for you.
所以我来给你提个主张。
33:19
Based on my read of the literature on inference time scaling, what's sometimes called test
根据我对 inference time scaling 文献的阅读,它有时被称为 test
33:25
time scaling, which is just having the models generate lots of tokens at the moment that
time scaling,也就是让模型在那个时刻生成大量 token
33:30
you ask them a question.
你问他们一个问题。
33:31
So those scaling trends, everything we're seeing now is completely in line with those predictions,
所以那些 scaling trends,我们现在看到的一切都完全符合那些预测,
33:36
which is you get pretty good gains for a while with the more tokens you spend on a log
也就是说,随着你把更多 tokens 花在 log
33:42
scale.
scale 上,你会在一段时间里获得相当不错的提升。
33:43
So this is jumping up quite a lot, but you do see it reflected in performance improvements,
所以这个涨得挺多的,但你确实能看到它反映在 performance improvements 上,
33:47
but it flattens out over time.
但会随着时间逐渐平缓下来。
33:50
And it's not like this curve sky rockets.
而且并不是说这条曲线会一飞冲天。
33:54
You got to spend a lot of tokens for small gains in performance.
你得花很多 tokens,才能换来 performance 上的一点点提升。
34:02
And we're just seeing it now play out.
而我们现在只是看到它正在上演。
34:03
And when people talk about token efficiency and worry about this, I think what they're seeing
当人们谈论 token efficiency 并为此担心时,我觉得他们看到的
34:06
is just the real lesson of what we already projected from inference time scaling.
其实只是我们早就从 inference time scaling 中预测到的真正教训。
34:11
And this is independent of the approach to inference time scaling you're taking, whether
而这和你采用的 inference time scaling 方法无关,无论
34:15
it's, you know, multiple parallel, you know, multiple parallel inferences or some kind
是,你知道,多个并行,你知道,多个并行 inferences,或者某种
34:23
of Oracle or, you know, any number of other schemes, it's just fundamental to inference
Oracle 的,或者,你知道,其他任意数量的方案,这对 inference
34:29
time scaling.
time scaling 来说就是根本性的。
34:30
That's a great question, right?
这是个好问题,对吧?
34:31
I think we know that it's independent of some of those things like the parallel work
我想我们知道,它不依赖于其中一些东西,比如 parallel work
34:36
versus having it do lots of long chains, but some of the other factors you mentioned,
而不是让它做大量 long chains,但你提到的其他一些因素,
34:40
I think we just don't know.
我想我们就是不知道。
34:42
And that's why I said it relates back to the question of efficiency for these architectures.
这就是为什么我说,它又回到了这些 architectures 的 efficiency 问题。
34:46
If we made a fundamental change to how the models work, maybe these trade-offs would
如果我们对 models 的工作方式做出根本性的改变,也许这些 trade-offs 会
34:51
be very different.
会非常不一样。
34:52
I mean, after all, so all of this stuff is a kind of patch job on the fact that there's
我的意思是,毕竟,所以所有这些东西都算是在给一个事实打补丁,那就是
34:57
no recursion in the depth.
depth 里没有 recursion。
34:59
It's a fixed depth.
它就是一个固定 depth。
35:01
And so the only recursion we can get, the on open ended notion of computation is by generation.
所以,我们唯一能得到的 recursion,也就是那种开放式的 computation 概念,是通过 generation 来实现的。
35:06
But if we had models that could be recursive, maybe fewer tokens for larger gains, I think
但如果我们有可以 recursive 的模型,也许更少的 tokens 就能换来更大的收益,我觉得
35:12
Yeah, I mean, in the end, we're going to spend the cost on compute or tokens.
是啊,我的意思是,到最后,我们反正都要把成本花在 compute 或者 tokens 上。
35:17
So this might not affect our bills in the end, but it is a fascinating question.
所以这最终可能不会影响我们的账单,但这是个很吸引人的问题。
35:21
What are the true scaling laws and what's possible in this space?
真正的 scaling laws 是什么,在这个领域里什么又是可能的?
35:24
And you're right to push back.
你说得对,是该反驳一下。
35:26
We talk about these things like they were like platonic ideals of laws, scaling law and
我们谈论这些东西时,就好像它们是某种柏拉图式的理想法则、scaling law,以及
35:31
folks that, right?
诸如此类的东西,对吧?
35:33
But even for the scaling laws for pre-training, you know, there's lots to discover there.
但即便是 pre-training 的 scaling laws,你知道,那里也还有很多东西有待发现。
35:37
And many of the stories of progress are actually like transcending the scaling law.
而且很多关于进步的故事,其实都是在超越 scaling law。
35:42
We see like better improvements than those laws predicted because everyone worked so hard
我们看到比那些 scaling laws 预测的还要好的提升,因为每个人都非常努力。
35:46
behind the scenes to do very innovative things, which maybe relates to our bitter lessons
在幕后做一些非常创新的事情,这也许和我们那些惨痛教训的
35:51
Any particular example come to mind of that?
这方面有没有什么特别的例子浮现在你脑海里?
35:54
Data usage in the nature of the data really matters and overtraining the models really matters,
数据的使用方式、数据本身的特性真的很重要,而 overtraining 模型也很重要,
35:58
which is kind of pushing up against the standard scaling law presentation.
这有点是在跟标准的 scaling law 说法相抵触。
36:02
And now I'm just going to speculate, I should check on this.
现在我只是在推测,我应该去核实一下。
36:04
But things like mixture of experts might have really flipped the script on what it means
但像 mixture of experts 这样的东西,可能真的彻底改写了“计算 parameters 意味着什么”这件事
36:09
to count parameters and then turn how these laws relate.
以及这些 laws 之间如何关联。
36:11
And then I think maybe even also stuff like the context window and so forth.
然后我觉得,可能甚至还包括像 context window 之类的东西,等等。
36:15
This is another thing to check, but I just speculate that we've seen larger gains from
这是另一个需要核实的事情,但我只是猜测,我们看到的更大收益来自
36:20
pre-training than you would have predicted by those early scaling laws papers suggesting
pre-training,比你从那些早期 scaling laws 论文里会预测到的还要大,那些论文暗示
36:24
that there is some innovative thing that was happening on top of pure scaling.
在纯粹的 scaling 之上,还有一些创新性的东西在发生。
36:29
Thinking about the concept of a market basket, one kind of pushback that comes up for me is
想到 market basket 这个概念,我脑海中会冒出的一种质疑是
36:38
in the, you know, the real economy, you know, A's is different than milk is different than
在,你知道,实体经济里,你知道,A's 和牛奶不一样,和
36:46
beef, et cetera, et cetera.
牛肉不一样,等等,等等。
36:49
And they're all influenced by different factors, you know, production, for example.
而它们都受到不同因素的影响,你知道,比如生产。
36:57
Whereas what you've done with this kind of CPI basket with tokens is kind of like more
而你用这种带 tokens 的 CPI basket 做的事情,更像是
37:06
like analogies, like here's the typical bundle of work and, you know, what it requires
一种类比,比如这是典型的一组工作,而且,你知道,它需要什么
37:11
from a consumption perspective.
从消费的角度来看。
37:14
The tokens aren't fundamentally different, like they're the same tokens.
tokens 本质上没什么不同,它们就是同样的 tokens。
37:17
It's just like how much it takes to do this versus how much it takes to do that versus
就只是像做这个需要多少,对比做那个需要多少,再对比
37:21
how much it takes to do that.
做那个需要多少。
37:22
You know, tell me what I'm, what I'm missing there and, you know, what does kind of characterizing
你知道,告诉我,我,我漏了什么,而且,你知道,算是刻画
37:30
these products, you know, give you in your analysis.
这些产品,你知道,在你的分析里到底能给你带来什么。
37:34
Yeah, fascinating to think about.
是啊,想想挺有意思的。
37:36
One thing I could insert there is the tokens are different at the level of being used
我能补充的一点是,tokens 在被用于
37:40
for code generation or skill file writing or explanation or thinking, right?
code generation,或者写 skill file、解释或思考时,是不一样的,对吧?
37:46
Those are different kinds of tokens that probably do feel tangibly different to us.
那些是不同类型的 tokens,对我们来说,很可能确实能感觉到明显不同。
37:51
So is that an element in your think?
所以,这是你思考中的一个因素吗?
37:52
Because I think I was thinking from our conversation that you had 10 different, almost like tasks,
因为我觉得,我是从我们的对话里想到,你有 10 个不同的、几乎像是 tasks 的东西,
37:59
like 10 different types of tasks from the domain of code generation, which, you know,
像是来自 code generation 这个 domain 的 10 种不同类型的 tasks,你知道,
38:05
if they were all kind of largely code generation, you know, that is the part that had some
如果它们大体上全都是 code generation,你知道,那就是有某种……的那部分
38:09
dissonance for me.
对我来说有种违和感。
38:10
But if you're talking about, like if your, your basket is like creative writing versus,
但如果你说的是,比如你的,你的 basket 是像 creative writing,对比于,
38:16
you know, a few code generation things that are kind of in different versus, you know, summarization
你知道,几个 code generation 的东西,它们有点属于不同的类别,对比于,你知道,summarization
38:24
versus editorial, commenting feedback, those, you know, may be more fundamental.
对比于 editorial、commenting feedback,那些,你知道,可能更根本。
38:31
So I think this is very significant when we have done some research on this as well at
所以我觉得这非常重要,因为我们自己在这方面也做过一些研究,在
38:35
the level of what kinds of session types exist and in turn, what kinds of users are there.
这个层面上:存在哪些 session types,以及反过来,有哪些类型的用户。
38:40
So you might notice of your own behavior.
所以你可能会注意到你自己的行为。
38:42
I guess this is reflected in your comment that sometimes you want to quick check in on
我猜这也体现在你说的那句话里,就是有时候你想快速确认一下
38:47
Sometimes you want a quick PR to get fired off.
有时候你想快速发一个 PR 出去。
38:51
Sometimes you want to be in a mode of deep collaboration.
有时候你想进入一种深度协作模式。
38:54
Sometimes you're partnering with the AI.
有时候你是在和 AI 搭档。
38:55
Sometimes you're delegating the work and so forth and so on.
有时候你是在把工作委派出去,诸如此类,等等等等。
38:59
And the outcome measures that we choose should be sensitive to this.
而我们选择的 outcome measures 应该对这一点敏感。
39:02
We shouldn't penalize the agent if your chat interaction with it about some scientific
如果你的 chat interaction 是和它围绕某个 scientific 的,我们不应该惩罚 agent
39:07
question didn't lead to a PR was never on the table in the first place.
问题并没有导向一个 PR,这件事从一开始就根本不在讨论范围内。
39:12
Whereas if you're trying to get some work delegated, that's actually a coding task and
但如果你是想把一些工作委派出去,那其实是个 coding task,而且
39:16
all it does is chat with you, that would feel quite unproductive.
它做的只是跟你聊天,那就会让人觉得挺没产出的。
39:20
So we need to bring that in and that would be a higher level discovery process of what
所以我们需要把这一点引入进来,而这会是一个更高层次的 discovery process,去搞清楚人们
39:24
people are trying to do and so forth.
想做什么之类的事情。
39:27
And thinking about the notion of value, is this something that you're anticipating?
再想想 value 这个概念,这是你预期中的东西吗?
39:35
It strikes me that that's an entire research thread that one could go into.
我觉得这整个就是一条可以深入研究的 research thread。
39:41
I don't know if that's a linguist or an economist or computer scientist, probably interdisciplinary.
我不知道这算是语言学家、经济学家还是计算机科学家的事,大概是跨学科的。
39:50
Most interesting questions.
最有趣的问题。
39:52
Is that something that you're working on or was it something that you put out there
那是你正在做的东西,还是你把它抛出来,让别人接手去推进的?
39:59
for someone to take up and run with?
我不确定。
40:01
I am not sure.
我可以告诉你,这个想法的脉络是这样的:我们创办了 Big Spin 这家 startup,因为我们希望看到更多人从 AI 中受益。
40:02
I can tell you the lineage of this idea is that we founded this startup Big Spin because
不管你爱它还是恨它,它都已经来了。
40:08
we would like to see more people benefit from AI.
而且我希望这些好处能分配得更均匀。
40:12
Whether you love it or hate it, it's here.
不管你爱它还是恨它,它都已经来了。
40:15
And I would like the benefits to be more evenly distributed.
而我希望这些好处能分配得更均匀一些。
40:19
And I can tell that that will mean bringing on board many more people than currently
And I can tell that that will mean bringing on board many more people than currently
40:25
benefit from AI.
benefit from AI.
40:26
Right now, I would say that it's mostly experts deriving real value.
我能看出,这意味着要让比现在多得多的人受益于 AI。
40:31
And a lot of the world is currently even trying to figure out what this is all about as a
Right now, I would say that it's mostly experts deriving real value.
40:34
tool or an entity in their lives.
现在的话,我会说主要还是专家在获取真正的价值。
40:37
So we would like to have more access and more productivity.
And a lot of the world is currently even trying to figure out what this is all about as a
40:41
And that implies making the user experience is much better.
tool or an entity in their lives.
40:46
Trying out what interaction patterns lead to success for people, meeting them where
而世界上很多人现在甚至还在试图搞清楚,这个东西作为一个工具,或者说作为他们生活中的一个存在,到底意味着什么。
40:50
they are in a kind of adaptive way.
它们是以一种自适应式的方式。
40:52
The whole list of things that you might worry about if you were a product manager who
如果你是一个产品经理,手上有一个已经部署的 AI 产品,
40:55
had some deployed AI product.
那你可能会担心的所有事情。
40:58
And I think by that root and from that perspective, we just ended up worrying about our own
而我觉得,从那个根源、从那个视角出发,我们最后就只是在担心我们自己的
41:04
token usage increasing and wondering whether there's real value there.
token 用量在增加,并且怀疑那里是否真有价值。
41:07
And it just happened to collide actually just like three weeks ago with this emerging
而且它实际上就在大概三周前,刚好撞上了这个正在兴起的
41:12
narrative on the back.
背后的叙事。
41:13
I think of all these rumors about IPOs, about what the return on investment was.
我想到所有这些关于 IPO 的传闻,关于投资回报到底是多少。
41:18
And then all these CEOs came out and said, all our spend was enormous.
然后所有这些 CEO 都出来说,我们的支出全都大得惊人。
41:21
And we want to scale back and we're walking back our claims from a few months ago.
而且我们想 scale back,也在收回几个月前说过的话。
41:25
And that is just a fascinating thing to witness in the narrative here.
而在这整个叙事里,这真是让人叹为观止的一幕。
41:29
I'm wondering are you also, does this research also attempt to project forward?
我很好奇,你们是不是也——这项研究是不是也尝试向前推演?
41:35
In theory, you could create a model for anthropics cost and spend based on publicly available
理论上,你可以根据公开可用的
41:46
data and some presumptions and give us a sense for how close we are to paying full freight
数据和一些假设,为 Anthropic 的成本和支出建一个 model,并让我们大致了解
41:55
for our tokens versus if we're only paying 10% for our tokens.
我们离为我们的 tokens 支付全价有多近,还是说我们只为我们的 tokens 支付 10%。
42:00
You could then project what that cost might look like over time as we're paying more
然后你就能推演,随着我们付得越来越多,这个成本随时间可能会变成什么样。
42:06
and more of the full cost.
以及更多完整成本。
42:08
Yeah, I don't have again fascinating questions.
是啊,我还是没有那些引人入胜的问题。
42:10
I don't have resolving answers.
我没有能解决问题的答案。
42:11
I am glad I am not tasked in some organization with projecting spend on all of this stuff
我很庆幸自己没有在某个组织里被安排去预测所有这些事情的支出
42:16
because I think it would be basically possible.
因为我觉得这基本上是有可能做到的。
42:18
For the time period that I was describing for our little CPI experiment, anthropic change
在我为我们那个小小的 CPI 实验所描述的那段时间里,anthropic 改了
42:25
the default reasoning on the model at least two times.
model 上的 default reasoning,至少两次。
42:28
So we see like they launched it with default reasoning high.
所以我们看到,他们发布它的时候 default reasoning 设得很高。
42:32
We have a mysterious sudden rise in the token usage, which we cannot explain.
我们的 token usage 出现了一次神秘的突然上涨,我们无法解释。
42:36
And there's a new baseline.
而且现在有了一个新的 baseline。
42:37
They lowered it to medium as the default.
他们把它默认调低到了 medium。
42:39
They patched a bunch of bugs that were related to context management and then turned it back
他们修了一堆跟 context management 相关的 bug,然后又把它调回了 high。
42:43
up to high.
而且所有这些都会影响总的 token output。
42:45
And all of these things have an effect on the total token output.
你可以想象,他们还改了默认的 context window,这意味着人们能在任何时刻吞下多得多的东西。
42:47
As you can imagine, they also changed the default context window, which meant people
你可以想象,他们还改了默认的 context window,这意味着人们
42:50
could swallow up much more stuff at any given moment.
能在任何特定时刻吞下多得多的东西。
42:55
So imagine trying to predict what token spend is going to be like when you have all these
所以想象一下,要预测 token spend 会是什么样,当你有所有这些
42:59
exogenous events in addition to changes that we don't even know about and questions about
exogenous events,再加上我们甚至都不知道的变化,还有关于
43:06
where the value actually lies.
价值到底在哪里的问题。
43:09
And then the true cost of a token, the estimates very wildly for every dollar we spend
然后是一个 token 的真实成本,每花出去一美元,估算会非常剧烈地波动
43:14
that could be as low as two and as high as 20.
可能低到 2,高到 20。
43:18
And I think this is just because it's hard to factor in things like R&D and future build
我觉得这只是因为很难把像 R&D 和未来的 build
43:22
out and depreciation and all of that stuff.
out、depreciation 以及所有这些东西都考虑进去。
43:25
I think at the current moment, we just don't know, but there couldn't be a more significant
我觉得眼下我们还不知道,但对全球经济来说,基本上没有比价值在哪里、谁来付钱、付多少更重要的问题了。
43:29
question for the global economy, basically, than where the value is and who's going to
是啊。
43:35
pay and how much.
在你的文章里,你创造了 tokenflation 这个词,至少用来描述最近 token economics 的表现。
43:38
In your article, you coined the term tokenflation to describe at least the recent behavior
看起来确实在继续。
43:44
of token economics.
token economics 的。
43:47
I imagine you see that continuing.
我想你看到这种情况在持续。
43:49
Seems to be continuing.
看起来确实还在持续。
43:52
So the picture that we get from the CPI, a picture of tokenflation, yes, your token
所以从 CPI 里我们看到的图景,就是 tokenflation,是的,按我们这里能想到去衡量的一切来看,你的 token 已经买不到过去能买到的东西了。
43:56
is not buying you what it once did, according to everything we can think to measure here.
而且就算把模型变好这个因素也调整进去,对吧?
44:02
And even adjusting for models getting better, right?
这一点很关键,因为如果这只是一个模型思考更多、变得更稳健,而我们所有人从中获得的结果都在指数级变好的故事,那这些支出看起来就会完全合理,但我们看到的并不是那幅图景。
44:04
That's critical there because if it was just a story of models thinking more and being
这很关键,因为如果这只是一个关于 models 想得更多、并且变得...
44:08
more robust and we were all getting exponentially better outcomes from this, then the spend
更加健壮,而且我们都能从中获得指数级提升的结果,那么这笔支出
44:13
would look completely rational, but that's not the picture that we see.
看起来就完全合理,但我们看到的并不是这样一幅图景。
44:16
And so we have to do some hard thinking about what's going to happen and how to improve
所以我们必须认真思考一下会发生什么,以及如何改善
44:22
Let's dig into that a little bit more.
让我们再稍微深入聊一聊这个。
44:25
How would you articulate what you're seeing?
你会怎么描述你所看到的情况?
44:26
The models are getting, quote, unquote, better.
这些模型正在变得,所谓,更好。
44:31
There's a set of open questions about, are the reported ways that models are better,
有一系列开放问题:那些被报告出来的模型变好的方式,
44:39
actually reflective of some intrinsic betterness, like in that question brings, it is often
是否真的反映了某种内在的更好,就像那个问题引出的那样,它往往
44:47
about like benchmarking and learning the benchmarks overfitting that kind of thing.
是关于像 benchmarking 和 learning the benchmarks、overfitting 这类事情。
44:53
And then there's the kind of question of chattiness and the volume of thought that it requires
然后还有一个问题,就是聊天量和
45:04
a given model generation to produce an answer.
某一代模型生成一个答案所需的思考量。
45:07
What are other factors that you see?
你还看到哪些其他因素?
45:09
Yeah, we can pick that apart as well.
是的,我们也可以把这一点拆开来讲。
45:11
So and this relates to a line I've had consistently, which is that we should think
所以,这和我一直坚持的一条线有关,那就是我们应该
45:17
in terms of systems, not in terms of models.
从系统的角度来思考,而不是从模型的角度来思考。
45:19
So in the data that we've got, Sonnet, as Opus and Sonnet, 4.5 versus 4.6, those two
所以,在我们拿到的数据里,Sonnet——拿 Opus 和 Sonnet 来说——4.5 对 4.6,那两次 generation 变化,我猜,都是真正的 model 变化。
45:26
generation changes, those are real model changes, I assume.
但另一方面,你知道,这里面有个事儿,它也意味着,你不该在那些你没有足够专业能力去评估答案的领域使用这些东西,可那又恰恰是你最需要帮助和支持的地方。
45:30
I think they did something very substantive at the level of the weights.
我觉得他们在 weights 层面做了非常实质性的工作。
45:34
And everybody immediately saw that that led to like a 5x increase in token usage.
而且大家立刻就看到,这导致 token 使用量大概增加了 5 倍。
45:39
And this was related to the introduction of adaptive thinking.
而这跟 adaptive thinking 的引入有关。
45:45
That's the level shift that we already took, and maybe we're seeing improvements there
那是我们已经经历过的 level shift,也许我们在那里看到了一些改进
45:49
that are worthwhile.
这些改进是值得的。
45:50
It gets hard to say, but let's assume there was a level up in improvement.
这很难说,但我们假设改进上有一个 level up。
45:54
Then for the period that we did our CPI experiment for, that's a fixed model, Opus 4.6.
然后,在我们做 CPI 实验的那段时间里,那是一个固定模型,Opus 4.6。
46:01
So all the code improvements that we saw in the data relate to the product.
所以我们在数据里看到的所有 code 改进,都和产品有关。
46:06
This has to relate to things like them turning the knobs on the adaptive thinking, changing
这肯定和这些事情有关,比如他们去调 adaptive thinking 的那些旋钮、改变
46:10
things about the system prompt, changing things at the level of the product, and that's
system prompt 相关的东西、在产品层面做改动,而
46:14
where the improvements were.
改进就出在这些地方。
46:16
And so that shows you that even for a fixed model, we can get varied for now comes through
所以这说明,即使 model 是固定的,我们也能得到变化,现在这些变化来自
46:20
these things because they really are sophisticated engineered systems at this point.
这些东西,因为它们现在真的是很精密的工程化系统。
46:25
And so was the product in this case specifically, cloud code or, oh yeah, and so we know
所以在这个情况下,产品具体是 cloud code 吗?还是,哦对,所以我们知道
46:31
there's tons of stuff there.
那里有超多东西。
46:33
Yes, I believe we know that these are all cloud code sessions that we kept in our data.
是的,我相信我们知道,这些都是我们保存在数据里的 cloud code sessions。
46:37
Sweet chat is broader than that and involves a couple of other coding agents.
Sweet chat 比这更广,还涉及另外几个 coding agents。
46:40
But I think I can say that all our data, our cloud code sessions using Opus 4.6.
但我想我可以说,我们所有的数据,我们使用 Opus 4.6 的 cloud code sessions。
46:45
Have you seen any evidence that changes via API usage, experience, similarly dramatic
你有没有看到任何证据表明,通过 API 使用发生的变化,会经历类似戏剧性的
46:56
variation in performance?
性能波动?
47:00
To kind of control for a lot of that product level stuff, all the prompts that are hidden
为了稍微控制很多那种产品层面的东西,所有那些被隐藏的 prompts。
47:04
from us, all of those affordances.
来自我们,所有那些 affordances。
47:06
Yeah, I don't know.
是啊,我也不知道。
47:07
But that's a nice thing to think about because it gives us more things that we can control
但这是件值得思考的好事,因为它给了我们更多可以控制的东西
47:12
for and more things that are knowable.
以及更多可知的东西。
47:14
So kind of in parallel to the model evolution, there's also evolution of the user you've alluded
所以,某种程度上,和 model 的演化并行,用户也在演化,你提到过
47:22
to this a little bit about kind of your concept is that most AI users now are experts.
这一点,你的概念大概是:现在大多数 AI 用户都是专家。
47:31
Talk a little bit about the role of expertise.
稍微聊聊专业能力的作用。
47:34
I think this is also kind of echoing back to our conversation about DSPY and like prompt
我觉得这也有点像是在呼应我们之前关于 DSPY 和 prompt 的对话。
47:40
optimization.
optimization。
47:42
You've done some research into how folks are using these models and the role of AI fluency.
你做过一些研究,是关于人们如何使用这些 models,以及 AI fluency 所扮演的角色。
47:47
Tell us about that research.
跟我们讲讲那项研究吧。
47:50
First I should say, so the distribution of users across expertise levels.
首先我得说,就是用户在不同专业水平上的分布。
47:56
So I guess the nuance picture I'd offer is that the people deriving a lot of value
所以我想,我给出的更细致的图景是,那些从 AI 中获得很多价值的人
48:00
from AI in the current moment tend to be experts.
在当下往往都是专家。
48:03
It must be the case that most users of AI are beginners just because the numbers are so
情况肯定是,大多数 AI 用户都是初学者,只是因为数字太...
48:08
large and expertise can't be that widely distributed yet.
大规模和专业能力还没法那么广泛地分布。
48:12
And that's a very interesting thing because I think probably most things are getting designed
而这非常有意思,因为我觉得大多数东西可能都是被设计成
48:16
for those experts implicitly or explicitly.
给那些专家用的,不管是有意还是无意。
48:19
But for the whole economic picture to work out, many more people need to derive value from
但要让整个经济图景成立,需要多得多的人从中获得价值,
48:24
this via one avenue or another.
通过这样或那样的途径。
48:27
And so that does shine a light on this expertise thing as a real factor.
所以这确实让专业能力这件事作为一个真正的因素凸显出来。
48:32
And the headline result there actually builds on something that anthropic did.
而那里的核心结果其实建立在 anthropic 做过的某件事之上。
48:35
They have this AI fluency index and their core observation in that work is that experts
他们有一个 AI fluency index,而他们那项工作的核心观察是,专家
48:41
display an augmentative style.
展现出一种 augmentative style。
48:43
They iterate with the AI.
他们会和 AI 反复迭代。
48:48
They change their requirements.
他们会改变自己的需求。
48:50
It's a really collaborative mode, whereas novices low fluency users delegate.
这真的是一种协作模式,而新手、低熟练度用户则会采取委派模式。
48:56
So they trust in the AI.
所以他们信任 AI。
48:58
They let it do its thing.
他们让 AI 放手去做。
48:59
They accept the responses uncritically.
他们会不加批判地接受这些回应。
49:02
And our contributors just show that this is a causal factor in success with these products
而我们的贡献者恰恰表明,这是这些产品能否成功的一个 causal factor,
49:09
Experts can do harder things more reliably as a result of all that friction they introduce,
专家能更可靠地完成更难的事,正是因为所有这些他们引入的 friction,
49:16
all that pushback.
所有这些 pushback。
49:18
Whereas novice users, they accept, but they end up accepting the wrong thing and they're
而新手用户呢,他们会接受,但最终接受的是错误的东西,而且他们
49:22
not able to level up from the basic tasks that they think to start with.
没法从他们一开始以为要做的那些基础任务 level up 上去。
49:27
And so that's obviously significant.
所以这显然很重要。
49:29
And it feels so tantalizing because pushing back is a natural human behavior.
而且它之所以这么诱人,是因为反驳本来就是人的一种自然行为。
49:35
I feel like we could encourage everyone in the world to do this.
我觉得我们可以鼓励世界上每个人都这么做。
49:37
We probably need to get them out of the mode of thinking, it's a super intelligence.
我们可能得让他们摆脱“它是 super intelligence”这种思维模式。
49:41
You should just trust it.
你就该直接相信它。
49:43
That has been the narrative for a while.
这种叙事已经持续一阵子了。
49:45
What we're seeing in the current moment and possibly for the foreseeable future is that
我们现在看到的,以及可能在可预见的未来会看到的,是
49:49
you got to complain, collaborate, introduce yourself, pushback, all that stuff that I think
你得抱怨、协作、介绍自己、反驳,所有这些我觉得
49:55
We take that for granted, right?
我们都把这当成理所当然的,对吧?
49:58
And so from a methodology perspective, how did you approach exploring this?
那从 methodology 的角度来说,你是怎么着手探索这个的?
50:04
We built on the work that Anthropic did, which they set up a nice framework with some
我们基于 Anthropic 做的工作,他们搭了一个很好的框架,和一些
50:10
independent research who were doing this kind of usability stuff.
独立研究者一起,他们当时就在做这类 usability 的东西。
50:14
And we just have an annotation protocol.
而我们只是有一套 annotation protocol。
50:16
We can talk in detail if you want about this, but at BigSpin we have lots of these best
如果你想的话,我们可以详细聊这个,但在 BigSpin,我们有很多这类 best practices
50:20
practices around having language models essentially collaborate on annotation projects to kind
围绕让 language models 基本上在 annotation projects 上协作,来 kind
50:26
of triangulate on the truth and factor out their individual biases.
去 triangulate 真相,并把他们各自的 biases 排除掉。
50:30
So we do that stuff and we apply all these fluency markers.
所以我们就做这些,然后把这些 fluency markers 全都用上。
50:33
And then separately we do a thing of estimating task complexity and looking for signs of visible
然后另外,我们还会做一件事:估算 task complexity,并寻找 visible
50:38
and invisible failures.
和 invisible failures 的迹象。
50:41
And so it's the connection between the fluency markers and the task complexity success metrics
所以,正是 fluency markers 和 task complexity success metrics 之间的联系
50:47
that was our contribution there.
那才是我们在这方面的贡献。
50:50
And that's where you can see high fluency users are the ones doing harder tasks.
从这里就能看出,high fluency users 才是那些做更难任务的人。
50:54
Paradoxically, there's more signs of failure for them because they complain, they push back,
矛盾的是,他们反而有更多 failure 的迹象,因为他们会抱怨、会反驳,
50:58
they're trying harder things.
他们在尝试更困难的事情。
51:00
But as part of all that friction, they're successful with harder things as well.
但作为所有这些摩擦的一部分,他们在更困难的事情上也取得了成功。
51:05
And have you were to try to apply this insight from the perspective of someone in an organization
而如果你要从一个组织里某个人的视角,试着应用这个洞察
51:11
that's trying to help or guide their organization to be more successful with AI?
这个人正试图帮助或引导自己的组织,让它在 AI 上更成功?
51:18
Like what do you think are the key lessons of this fluency work?
比如,你觉得这项 fluency 工作的关键经验是什么?
51:21
If it's an organ that's just starting out and wants people to figure out how this could
如果是一个刚起步的组织,想让人们弄清楚这怎么才能
51:25
be part of the organization's mission, it would just be that pushback message.
成为组织使命的一部分,那就会是那个 pushback message。
51:30
And you could do an experiment where you interact with it about something where you're
而且你可以做一个实验,和它围绕某件你正在……的事情互动
51:34
We're all an expert in something.
我们每个人都在某个方面是专家。
51:36
Engage in a discourse with one of the best models about something you're an expert in
就你擅长的某个领域,跟最好的模型之一展开一场讨论,
51:39
and see how often you feel you have to push back.
看看你有多经常觉得自己必须反驳它。
51:42
And this could be a kind of a lesson.
这可能是一种教训。
51:44
For other spheres where I don't know the answer, it might be just as errorful.
对于那些我不知道答案的其他领域,它可能也一样容易出错。
51:48
That could be a good visceral thing.
那可能会是一种很直观的切身感受。
51:50
If the org is very far along, I think the main thing to do right now is to have a team
如果这个组织已经走得很远了,我觉得现在最主要的事情就是有一个团队
51:54
of these LMs interacting to improve things.
这些 LMs 互相交互,把事情做得更好。
51:57
For example, at Big Spin, I didn't set this up.
比如,在 Big Spin,这不是我搞起来的。
52:00
Our founding engineer is very future forward on agents and he's incredible at this.
我们的创始工程师对 agents 非常前瞻,而且他在这方面特别厉害。
52:06
And when we do PRs now, the first round of review is the agents all interacting,
而现在我们做 PRs 的时候,第一轮 review 是 agents 全都在互相交互,
52:12
collaborating, disagreeing.
协作、意见不合。
52:14
They do the first round of comments.
他们做第一轮 comments。
52:15
They do the first round of code updates.
他们做第一轮 code updates。
52:17
Only after they resolve things do we look at a PR.
只有等他们把问题解决之后,我们才会看 PR。
52:20
So the final human stage should be very high value and the agents did all that work.
所以最后那个人类环节应该价值非常高,而 agents 把那些活儿都干了。
52:25
But when you have one agent do it, they often just reinforce themselves and you don't get
但如果你只让一个 agent 来做,它往往只是在自我强化,你得不到好的结果。
52:29
good outcomes.
真正有变革性的,就是那种 team of rivals 的机制。
52:30
It's that team of rivals thing that is transformative.
你知道,我觉得这挺有意思,因为一方面,这当然说得通。
52:34
You know, I think it's interesting because on the one hand, of course, that makes sense.
但另一方面,你知道,这里面有个点:它也意味着,在那些你没有足够专业能力去评估答案的领域,你就不该用这些东西;可偏偏那些地方,你才最需要帮助和支持。
52:40
But on the other hand, there's something, you know, it also implies that you shouldn't
我们其实都待在自己那条非常窄的垄里,甚至都没真正意识到。
52:46
be using these things in areas where you don't have enough expertise to evaluate the
谁知道这个花园里,外面还有什么呢。
52:52
answer, yet that's where you most need the assistance, the support.
这玩意儿真的很酷。
53:01
And again, and this is a little bit worrisome about the overall narrative around AI.
而且再次,这对围绕AI的整体叙事来说有点令人担忧。
53:06
The place where we can get around this is with software development because let's say
我们能绕开这个问题的地方是软件开发,因为假设
53:11
that I'm trying to accomplish something in a language that I don't know how to code in.
我想用一种我不会编程的语言来完成某件事。
53:15
I can have the agent do work for me because probably in the end, I can run the program
我可以让agent替我干活,因为很可能到最后,我可以运行这个程序
53:20
and look at the results.
然后看结果。
53:22
And that's what mattered to me is that I run the results and I see.
而对我来说重要的是,我运行结果,然后我看到了。
53:25
And if I don't see what I want and I can complain and we can iterate, that verification
如果我没看到我想要的,我可以抱怨,我们可以迭代,那个verification
53:31
step that doesn't imply I have comprehensive knowledge, it just implies that I know what
这一步并不意味着我有全面的知识,它只意味着我知道
53:34
I want to see in the end is so critical.
我最终想看到什么,这一点非常关键。
53:37
And I think this is a causal factor in models being so good at coding because it's like
而且我觉得,这是 models 在 coding 上如此擅长的一个 causal factor,因为它就像
53:41
the ultimate verifiable domain for them.
对它们来说,这就是终极的 verifiable domain。
53:44
But as soon as we leave that and go even into something like the legal realm where the
但一旦我们离开那个领域,甚至进入像法律领域这样的地方,那里的
53:49
requirements are strict, but they're not codified in code and they have ambiguity about
要求很严格,但它们并没有被 codified in code,而且它们还有
53:55
This whole picture falls apart and you then are back at what you just said, which is this
这整个图景就崩了,然后你又回到你刚才说的那一点,也就是这个
54:00
awful kind of paradox is like, yeah, use AI, but in the end, unless you're expert enough
那种糟糕的悖论就像是,是啊,用 AI,但到最后,除非你足够专业
54:05
to evaluate every single one of its responses, you might be in real trouble.
能评估它的每一个回应,否则你可能真会陷入大麻烦。
54:10
I don't know how to get out of this because the verification step is like, we go to trial,
我不知道怎么摆脱这一点,因为 verification step 就像是,我们得上法庭,
54:15
but this is very consequential.
但这后果非常严重。
54:18
Yeah, that's funny.
是啊,这挺好笑。
54:20
I mean, it does make me think a little bit about, you know, some of the types of errors
我是说,这确实让我稍微想到,你知道,某些类型的错误
54:26
that we're trying to avoid are factuality and, you know, there is a temptation to say,
我们正试图避免的,是 factuality,而且,你知道,会有一种冲动想说,
54:33
well, let's just throw more tokens at it.
好吧,那我们就干脆多给它扔点 tokens 吧。
54:34
Like, I'll have a critic model that evaluates everything that the, you know, is generated
就像,我会有一个 critic model,来评估所有那些,你知道,被生成出来的东西
54:42
by the primary model.
由 primary model 生成的。
54:44
But then you go back to my observation that these models tend to correlate in their responses
但然后你再回到我那个观察,就是这些 models 在它们的回答上往往会 correlate
54:51
Yeah, it's super interesting.
是啊,这特别有意思。
54:55
That's a good point.
这个观点很好。
54:56
Yeah, for my picture, we want real diversity of perspectives.
是啊,就我设想的图景而言,我们想要的是真正多元的视角。
54:59
This is just like, you know, red teaming for humans.
这就像,你知道的,对人类做 red teaming 一样。
55:02
This is most successful when you have a really diverse team of people who think creatively
当你有一个真正多元化的团队,成员都能有创意地思考时,这种方式最成功
55:06
and differently.
而且他们的想法还各不相同。
55:07
And if every one of the members of that team is thinking in a homogeneous way, they miss
而如果那个团队里的每个成员都用同质化的方式思考,他们就会漏掉
55:11
all of the crucial things.
所有关键的东西。
55:13
Same exact issue of all the code review agents are biased in the same way.
完全相同的问题:所有 code review agents 都以同样的方式有 bias。
55:16
They will miss exactly the same class of bugs.
它们会漏掉完全同一类 bugs。
55:19
And then we're all sunk.
然后我们就全完蛋了。
55:20
Yeah, I don't know how you'd encourage this diversity in the ecosystem or probably, as you
是啊,我不知道你要怎么在 ecosystem 里鼓励这种多样性,或者,可能就像你
55:24
say, converging towards some kind of one model.
说的,最终会收敛到某种单一的 model。
55:27
But I think for my picture, we need diversity.
但我觉得,按我的设想,我们需要多样性。
55:29
Yeah, we got to keep those open weights models going or something because they're the weird
是啊,我们得让那些 open weights models 继续跑下去还是怎样,因为它们确实是这个领域里怪异的
55:33
players in the space for sure, for sure.
参与者,肯定的,肯定的。
55:36
So we've talked about efficiency, interpretability, tokenomics, fluency.
所以我们已经聊了 efficiency、interpretability、tokenomics、fluency。
55:48
You're involved in a lot of different research directions, excellent, excellent, what's next
你参与了很多不同的研究方向,太棒了,太棒了,你接下来
55:53
Where do you see either where do you see this all going kind of externally, but also like
你觉得——或者说,你觉得这一切从外部来看会走向哪里,但同时也像是
55:57
where is your research going?
你的研究要往哪个方向走?
55:59
Yeah, this is great.
对,这太棒了。
56:01
And as I said before, we're trying to think in weird and creative ways about what the
而且就像我之前说的,我们正试着用古怪又有创意的方式去思考
56:05
future could hold.
未来可能会是什么样。
56:06
And I encourage my students to do this.
我也鼓励我的学生这么做。
56:08
And they're smart.
而且他们很聪明。
56:09
So they say, all right, Chris, I'll think along those lines, but what's your answer to
所以他们会说,好吧,Chris,我会按这个思路去想,但你对……的答案是什么呢?
56:13
So I do have an answer.
所以我还真有个答案。
56:14
And it's really shooting for the moon here, which would be what about the architectural
而且这里真的是在冲着月亮去,那会是什么呢,就是关于 architectural
56:18
innovation that would upend the whole story around the stack transformer and the way we
innovation,它会彻底颠覆围绕 stack transformer 的整个故事,以及我们
56:24
need to do data center build out to even get incremental gains in performance.
为了哪怕在 performance 上拿到一点增量提升,需要怎么做 data center build out。
56:28
That could be upended and it would come from some very innovative thing around maybe
那可能会被颠覆,而且它会来自某个非常创新的东西,也许是关于
56:33
recursive use of the building blocks that we've got.
对我们手里已有的 building blocks 做 recursive use。
56:37
So architectures, we should think.
所以 architectures,我们应该想一想。
56:38
And when people say, oh no, we don't need more architectures.
而当人们说,哦不,我们不需要更多 architectures 时,
56:40
The transformer is good enough.
transformer 已经够好了。
56:42
That's where we should push back as academics doing something more clever and more scrappy
这就是我们作为学者应该顶回去的地方——去做一些更聪明、更敢闯的事情,
56:47
that could change the world.
那可能会改变世界。
56:49
And the other one is thinking in the interps space, much more about data.
另一个方向是在 interps 领域里思考,更多地去关注 data。
56:54
And that's just because I want to tell the true story of how we go from data to model
而这只是因为我想讲一个真实的故事:我们怎么从 data 到 model
56:58
capabilities, but it also checks a box for me on connecting interpretability to safety.
capabilities,但它也帮我在把 interpretability 和 safety 连起来这件事上,算是完成了一个目标。
57:06
It has been hard for me to connect those two things.
对我来说,把这两件事连起来一直很难。
57:08
We have found some ways to do it, but it's not a slam dunk as a narrative, even though
我们已经找到了一些实现方法,但作为一套叙事,它并不是十拿九稳,尽管
57:12
it's a dominant narrative.
它是主流叙事。
57:13
But I will say that when we get into things like data poisoning from innocuous examples,
但我要说,当我们谈到用看似无害的样本进行 data poisoning 这类事情时,
57:19
this is probably a growing societal concern.
这很可能是一个日益严重的社会担忧。
57:22
There is evidence that with very few examples planted in a pre-training data set, you can
有证据表明,只要在 pre-training data set 里植入极少数样本,你就能
57:27
have a significant influence on the outlook and preferences and quirks of the final model.
对 final model 的观点、偏好和一些怪癖产生显著影响。
57:34
So can we detect those examples?
那么,我们能检测出那些样本吗?
57:36
What's the nature of those attacks?
那些 attacks 的本质是什么?
57:38
How well hidden could they be?
它们能藏得有多隐蔽?
57:39
What's the smallest number of examples?
最少需要多少个 examples?
57:41
And why does it happen?
那为什么会发生?
57:43
These are all going to be very pressing questions.
这些都会是非常紧迫的问题。
57:46
And so again, it's just a data-oriented question that's very alive for me in the current moment.
所以还是那句话,这只是一个 data-oriented 的问题,而此刻它对我来说非常鲜活。
57:51
On the architecture front, is there research that you're seeing or doing that is as yet
在 architecture 方面,有没有你看到或正在做的研究,目前还
58:02
under the radar that you think is promising and are underappreciated?
没被太多人注意到,但你觉得很有前景、也被低估了?
58:08
I think you had my student, Julie Colledion, and she is an advocate for bite-level models,
我记得你邀请过我的学生 Julie Colledion,她是 bite-level models 的倡导者,
58:13
essentially tokenizer free models.
本质上就是 tokenizer free 的模型。
58:15
I think that's a big part of the future.
我觉得这会是未来的很大一部分。
58:17
It's a critical thing if you want to have truly multilingual models that are also equitable
如果你想拥有真正 multilingual 的模型,而且这些模型还要公平,
58:21
in terms of how many tokens they charge us for, getting back to that earlier theme.
也就是从它们按多少 tokens 向我们收费这个角度来说,回到之前那个主题。
58:25
But also, Julie's perspective is that this is speculative, but I think there's something
但另外,Julie 的观点是,这还是推测性的,不过我觉得这里面有点
58:29
to this, that it's a kind of inference-time scaling, because you do more compute at test time,
道理,它是一种 inference-time scaling,因为你在 test time 会做更多 compute,
58:36
because you have more tokens, and there are four more opportunities to build on interesting
因为你有更多 tokens,而且还有另外四个机会,可以在有趣的
58:42
So that could be a big part of the future, and the other one would be recursive architectures,
所以那可能会是未来很重要的一部分,另一个就是 recursive architectures,
58:48
But if you want to go all the way out, you could think, why do we always assume we're going
但如果你想彻底往远处想,你可以想想,为什么我们总是假设我们会做 gradient-based learning?
58:52
to do gradient-based learning?
那有很多替代方案,但没人去探索,因为所有人都把它当成不言自明的事。
58:53
There are lots of alternatives to that, and nobody is exploring them, because everyone
我们都待在自己非常窄的那一垄里,甚至都没真正意识到这一点。
58:58
takes it as a truism.
谁知道这个花园里,外面还有什么呢。
58:59
We're all in our very narrow row here without even really realizing it.
我们其实都待在自己那一窄垄里,甚至都没真正意识到。
59:03
Who knows what's outside in this garden.
谁知道这个花园外面有什么。
59:06
It's very risky as a research bet, because only one in a thousand of these ideas will
从研究下注的角度看,这风险非常大,因为这类想法里只有千分之一会真正有回报。
59:11
pay off.
但如果你不打算冒这种风险,那做学术研究者又有什么意义呢?
59:13
But what's the point of being an academic researcher if you're not going to take that kind of risk?
这正是我们的定位所在。
59:18
That's what we're positioned to do.
嗯,Chris,非常感谢你来参加,并分享了一点你正在做的事情。
59:20
Well, Chris, thanks so much for jumping on and sharing a bit about what you're working
这真的很酷。
59:25
It's very cool stuff.
这真是很酷的东西。
59:27
What a wonderful conversation.
真是一次精彩的对话。
59:28
It gave me lots of new things to think about.
它让我有很多新的东西可以思考。