Anthropic Research · Alignment Series
Constitutional AI 2026: Training Claude to Self-Critique
Constitutional AI 2026:训练 Claude 自我批判
Anthropic walks through the evolution of Constitutional AI 2026 — from principle lists to full AI self-play workflows.
Anthropic 详解 Constitutional AI 2026 版本的演进:从原则列表到 AI 自我对练的完整工作流。
00:00
00:00
All right, let's get into it.
好了,我们开始吧。
00:02
Today, we are taking a really close look
今天,我们要深入剖析
00:04
at one of the biggest names in the AI game right now.
当前人工智能领域最炙手可热的巨头之一——
00:07
Claude, the AI assistant from Anthropic.
来自Anthropic公司的AI助手Claude。
00:10
You know, with hundreds of AI tools out there,
要知道,市面上已有数百种AI工具,
00:12
it feels like a new one pops up every day.
几乎每天都有新面孔出现。
00:14
So the big question is, what makes Claude so special?
那么关键问题来了:Claude究竟有何独特之处?
00:17
What's its secret sauce?
它的制胜秘诀又是什么?
00:19
So at its heart, Claude is an AI assistant
其核心而言,克劳德是一款由名为Anthropic的研究公司打造的人工智能助手。
00:22
built by a research company called Anthropic.
该公司在推出它时,怀揣着一个极其清晰的使命:
00:25
And when they launched it, they had this super clear mission.
让它有用,让它诚实——
00:28
Make it helpful, make it honest,
而最关键的一点是,让它无害。
00:30
and this is the big one, make it harmless.
这种对安全性的全面考量,没错,
00:33
And that whole idea of safety, yeah,
正是我们今天要反复探讨的主题。
00:35
that's a theme we're gonna be hitting on a lot today.
好了,先说重点。
00:38
Okay, first things first.
好,先说第一点。
00:39
If you've never used Claude before, you're gonna love this.
如果你之前没用过Claude,你一定会爱上它的。
00:42
Getting started is actually super, super simple.
上手其实超级、超级简单。
00:45
Let's just walk through the steps right now.
现在咱们就来一步步操作。
00:47
Seriously, it couldn't be easier.
说真的,没有比这更容易的了。
00:49
You just pop over to Claude.ai, sign up using your email
你只需打开Claude.ai,用邮箱注册,
00:52
and click through a quick verification.
再快速完成验证。
00:54
And bam, you're in.
然后“啪”一下,就进去了。
00:56
The chat window is right there, ready for you
聊天窗口就在那儿,随时等你开聊。
00:58
to type in your very first prompt.
以下是中文翻译:
01:00
Now, here's a quick tip.
这里有个小技巧。
01:01
And honestly, this goes for pretty much any AI you use,
**输入你的第一个提示词。**
01:04
but it's especially true for Claude.
现在,这里有个小技巧。
01:06
Be specific.
说实话,这适用于你使用的几乎所有AI,
01:08
The clearer and more detailed your prompts are,
但对Claude来说尤其如此:
01:10
the better your results are gonna be.
**要具体。**
01:12
Don't think of it like typing a few words
你的提示词越清晰、越详细,
01:14
into a search engine.
以下是中文翻译:
01:15
Think of it more like you're giving detailed instructions
把它想象成你在给专业人士、学生、研究人员下达详细指令。
01:17
to a really smart assistant who's ready to help.
**翻译成搜索引擎的查询语句。**
01:21
Okay, now we're getting to one of Claude's real superpowers.
把它想象成你在给一个非常聪明的助手下达详细指令,
01:24
This is the feature that makes it an absolute beast
这个助手随时准备提供帮助。
01:26
for professionals, students, researchers,
你可以直接扔一份又密又无聊的报告给它。
01:29
really anyone who deals with a lot of information.
好了,现在我们开始触及克劳德真正的超能力之一了。
01:32
We're talking about its ability
正是这个功能,让它成为专业人士、学生、研究人员——
01:33
to handle massive amounts of text.
处理海量文本的关键,
01:36
So, this all boils down to something called
归根结底在于一个概念——
01:38
the context window.
上下文窗口。
01:40
The easiest way to think about it
最直观的理解方式,
01:41
is like the AI's short-term memory.
就是把它看作AI的短期记忆。
01:44
And Claude's is huge.
而克劳德的记忆容量惊人——
01:46
We're talking 200,000 tokens,
足足20万个词元,
01:48
which are just tiny pieces of words.
也就是构成词语的微小单元。
01:50
But what does that actually mean?
但这到底意味着什么呢?
01:52
Well, to put it in perspective,
这么说吧,打个比方——
01:54
it's like you could hand Claude a 500-page book
就像你可以递给克劳德一本500页的书,
01:56
and it would remember the whole thing.
它能把整本书一字不差地记住,贯穿你整个对话。
01:58
Word for word for your entire chat.
那它在现实世界有什么用呢?
02:01
And the real world uses for this?
哦,用处可大了。
02:02
Oh man, they're huge.
你可以把一份枯燥冗长的报告扔给它处理。
02:04
You can throw a dense boring report at it
而且你知道最棒的是什么吗?
02:06
and get a perfect summary in seconds.
并在一秒内获得完美摘要。
02:08
Or find that one needle and a haystack piece
或从海量PDF中精准定位那根“针”——
02:11
of information inside a giant PDF.
哪怕它藏在浩如烟海的文档里。
02:13
You can even analyze hundreds of pages of code
你甚至能分析数百页代码,
02:15
or pull out the most critical points
或从冗长复杂的合同中提炼关键要点。
02:17
from a long complicated contract.
这完全是颠覆性的变革。
02:19
It's a total game changer.
而你知道最棒的是什么吗?
02:20
And you know what the best part is?
立刻就能针对那份文档提问。
02:22
Using this feature is so easy.
使用这个功能非常简单。
02:24
You just look for that little paperclip icon
你只需在聊天界面中找到那个小小的回形针图标,
02:26
right there in the chat.
点击它,上传你手头的任何文件——无论是PDF、
02:27
Click it, upload whatever you've got, a PDF,
文本文件还是Word文档,就搞定了。
02:30
a text file, a Word doc, and that's it.
然后你就可以直接开始针对这份文档提问。
02:32
You can start asking questions
真的就这么简单。
02:34
about that document right away.
就这么简单。
02:35
It really is that simple.
真的就这么简单。
02:37
Okay, before we get into the really fascinating stuff,
好的,在我们深入探讨那些真正引人入胜的内容之前——
02:40
the philosophy behind how Claude keeps itself safe,
也就是克劳德如何确保自身安全背后的哲学原理——
02:42
just a quick pause.
先稍作停顿。
02:43
If you're getting something useful out of this,
如果你觉得这些内容对你有所帮助,
02:45
it would be awesome if you could take a second
那将非常棒,如果你能花一秒钟
02:46
to hit that like button, maybe share it
点个赞,或许再分享给
02:48
with someone who'd find it interesting,
觉得有趣的人,
02:50
and of course, follow us so you don't miss out
当然,也请关注我们,这样你就不会错过后续内容。
02:51
on what's next.
接下来要聊什么?
02:53
All right, now we're going to talk about
好了,现在我们要谈谈
02:54
what really, truly sets Claude apart from the pack.
究竟是什么让Claude真正从同类中脱颖而出。
02:58
It's all about its ethical framework.
关键在于它的伦理框架。
03:00
And this isn't some feature they tacked on at the end.
这可不是什么事后添加的功能。
03:02
No, this is baked into its DNA
不,这是刻在它基因里的东西,
03:04
and it's all thanks to this incredible concept
而这一切都归功于一个了不起的概念——
03:06
they call constitutional AI.
他们称之为“宪法式人工智能”。
03:09
So what in the world is constitutional AI?
那么,宪法人工智能到底是什么呢?
03:13
Well, its anthropics totally unique way
这是Anthropic公司独树一帜的方法,
03:15
of building a safe AI from the very beginning.
旨在从源头构建安全的人工智能。
03:19
See, instead of having humans constantly tell the AI
不同于让人类不断告诉AI
03:22
that was a good answer or that was a bad answer,
“这个回答好”或“那个回答差”,
03:25
they give the AI a rulebook, a set of core principles,
他们给AI一本规则手册,一套核心原则,
03:28
literally a constitution that it has to follow.
确切地说,是一部它必须遵守的“宪法”。
03:31
And this is where it gets really different
而真正的独特之处,正在于此。
03:33
from how most other models are trained.
这与大多数其他模型的训练方式不同。
03:35
You see, most models use this thing called RLHF,
你看,大多数模型采用一种名为RLHF的技术,
03:38
reinforcement learning from human feedback
即基于人类反馈的强化学习——
03:40
where people are constantly grading the AI's responses.
人们会不断对AI的回应进行评分。
03:43
But Claude, it learns to supervise itself.
但Claude不同,它学会了自我监督。
03:46
It literally uses its own constitution
它实际上会依据自身的“宪法”
03:48
to check its work, critique its answers, and make corrections.
来检查自己的输出、评判自己的答案,并做出修正。
03:51
It's basically its own ethics coach.
它本质上就是自己的伦理导师。
03:53
So how do they do this?
那么,它们是如何做到的呢?
03:54
It's basically a two-step process.
这基本上是一个两步过程。
03:57
First, in what's called the supervised phase,
首先,在所谓的监督阶段,
03:59
the model just practices.
模型只是进行练习。
04:00
It generates answers and then immediately checks them
它生成答案,然后立即对照其“宪法”进行核对,
04:03
against its constitution, correcting itself as it goes.
并在此过程中自我修正。
04:06
Then, in the second phase, the reinforcement stage,
接着,在第二阶段,即强化阶段,
04:08
it uses all that self-critique to get better and better,
它利用所有这些自我批评来不断改进,变得越来越好。
04:11
constantly reinforcing the good constitution-friendly habits.
不断强化有益于宪法的良好习惯。
04:15
And we're not talking about a few simple do's and don'ts here.
这里我们讨论的并非几条简单的注意事项。
04:18
This constitution is a serious document.
这部宪法是一份严肃的文件。
04:20
We're talking about 75 different articles,
它包含75条不同的条款,
04:23
pulling from some really heavy hitting sources,
借鉴了极具分量的来源,
04:25
like the UN's Universal Declaration of Human Rights.
例如联合国的《世界人权宣言》。
04:28
So this gives it a really solid, globally aware, ethical foundation.
因此,这赋予了它一个非常坚实、具有全球视野的道德基础。
04:32
Okay, so we've talked about how to get started
好了,我们已经讨论了如何开始。
04:34
and we've talked about the deep ethical principles it runs on.
我们探讨过它所遵循的深层伦理原则。
04:37
But what does this actually mean for you today
但这对今天的你——
04:39
trying to figure out which Claude model to use?
正纠结该用哪个Claude模型——究竟意味着什么?
04:41
Let's break down the options.
我们来逐一分析选项。
04:43
The Claude 3 family basically has a tool for every job.
Claude 3系列基本能应对各种任务。
04:46
You've got Haiku, think of it as the sprinter.
比如Haiku,你可以把它想象成短跑选手:
04:48
It's super fast, really cost-effective,
速度极快,性价比超高,
04:51
and perfect for quick, simple tasks.
非常适合处理快速、简单的任务。
04:53
Then there's Sonnet.
然后是「十四行诗」。
04:54
This is your all-rounder hitting that sweet spot
这是一款全能型选手,在速度与性能之间找到了完美的平衡点。
04:56
with a perfect balance of speed and performance.
接着是「杰作」。
04:59
And then, there's Opus.
「杰作」是重量级冠军。
05:01
Opus is the heavyweight champion.
它是你为最复杂、最烧脑的分析和研究而准备的利器。
05:02
It's the one you bring out for the most complex,
所以,当你退后一步,纵观全局时——
05:04
brain-melting analysis and research.
脑洞大开的分析与研究。
05:07
So when you step back and look at the whole picture,
所以当你退一步,纵观全局,
05:09
Claude really isn't just another chatbot.
克劳德绝非又一款普通的聊天机器人。
05:11
It's designed to be more like a partner.
它被设计得更像一位伙伴——
05:13
One that's incredibly powerful, totally focused on making
强大无比,全心致力于提升你的效率,
05:16
you more productive, and most importantly,
而最重要的是,它建立在极其坚实的伦理基础之上。
05:18
built on this really strong ethical foundation.
这便引出了核心要点:
05:21
And that really brings us to the big takeaway here.
Anthropic 从根本上试图
05:24
Anthropic is fundamentally trying
将道德指南针直接嵌入人工智能本身。
05:25
to build a moral compass right into the AI itself.
把道德指南针直接嵌入AI本身。
05:28
So I want to leave you with a question to think about.
因此,我想留给大家一个问题去思考:
05:30
What does it mean for the future?
这对未来意味着什么?
05:32
For how we work and live with these machines,
对我们如何与这些机器共事、共存意味着什么——
05:34
if we can give them an ethical constitution from the start?
如果我们从一开始就能赋予它们一套道德准则?
05:37
It's definitely something to chew on.
这绝对值得深思。
05:39
Thanks for tuning in.
感谢收看。