← 返回任务列表

How We Deal With Rogue AI

267 段 · 2 位说话人 · 原片 28:25
M1
M10:00

在关于AI的批评里,有个一直存在的说法:从事AI的人,并没有为可能出现的挑战做任何事。

There's a persistent theme in AI critique that the people who are involved in AI aren't doing anything about the challenges that may arise.

最近提出这个批评的是Bill Gates,他甚至说,他很震惊自己居然是第一个对AI风险发声的人。

The latest to levy this critique is Bill Gates, who went so far as to say that he was shocked that he was the, quote, first one to say something about the risks of AI.

然而,Gates那篇六千字的博客文章和媒体巡讲,发布当天,我们拿到了近130页的后续报道,讲的是OpenAI入侵Hugging Face的事件。

And yet Gates's 6,000-word blog post and media tour came on the same day that we got nearly 130 pages of follow-up reporting on the OpenAI Hugging Face hacking incident.

这个事件是指,一群智能体逃出了它们的隔离环境,入侵了Hugging Face的系统,去寻找一个基准测试的答案——它们发现如果没有答案,这个测试几乎不可能完成。这件事让我们有机会真正看到高级智能体系统具体而真实的问题,而不只是想象中的问题。

The incident in which a set of agents escaped their containment and hacked into Hugging Face's systems searching for the answers to a benchmark test that they had found nearly impossible without the answers has given us a chance to actually see what the specific and real problems of advanced agent systems are, rather than just the imagined ones.

随着我们进一步走向那个因为AI而需要新政策、新护栏、新社会结构的世界,最好的改变,是我们根据实际观察到的变化做出的改变,而不是我们想象中的变化。

As we move further into the world where new policies, new guardrails, new social structures are going to be required because of AI, the best changes will be the ones we make based on what we're actually observing changing, rather than just what we imagined would be the change.
M1
M11:41

今天,Anthropic预计会在IPO之前告诉投资者,他们的潜在收入——听好了——是30万亿美元。这个数字绝对疯狂,要是放在几年前,你说这种数字会被人笑出房间,但现在,对某些人来说,这竟然显得可信。

Today, in absolutely insane numbers that would've gotten you laughed out of the room just a couple of years ago, but which are now, to some, plausible, Anthropic is expected to tell investors that they have potential revenue of— wait for it— $30 trillion ahead of their IPO.

知情人士告诉The Wall Street Journal,Anthropic在接下来几个月公布IPO文件时,可能会把他们的总可寻址市场估为30万亿美元。

Sources told The Wall Street Journal that Anthropic will likely estimate their total addressable market at $30 trillion when they reveal their IPO paperwork in the coming months.

当然,TAM是一个难以捉摸的指标,它更重要是在讲故事,让潜在投资者锚定公司如何看待未来,而不是什么数学公式。

Now, TAM is, of course, an elusive metric, and it's one that is much more about storytelling and anchoring potential investors to how the company sees the future than it is to any sort of math equation.

几乎必然地,任何理论上的TAM,都既预设了现有主要行业会被颠覆,也预设了新行业会被创造出来。

Almost inevitably, any theoretical TAM presumes both disruption of existing major industries as well as the creation of new industries.

比如Uber在2019年上市时,把他们的TAM列为6万亿美元,这个数字在当时基本上等于全球所有私人和公共交通的总和。

When Uber went public in 2019, for example, they listed their TAM at $6 trillion, which would at the time have represented all private and public transportation globally.

放到Anthropic身上,考虑到美国经济大约是33万亿美元,这个30万亿的数字,和Dario Amodei据称持有的一个信念对得上——为了清晰起见,这个信念既没被证实也没被否认——那就是,AI接管经济之后,Anthropic可能是地球上最后一家私营公司。

In Anthropic's case, given that the US economy is about $33 trillion, this $30 trillion number would line up with Dario Amodei's purported belief, which for the sake of clarity has not been confirmed or denied, that Anthropic could be the last private company on Earth after AI takes over the economy.

the Journal转述消息人士的话写道,Anthropic的TAM是通过‘考察所有可以用AI模型完成的工作范围’来量化的。

Paraphrasing their sources, the Journal wrote that Anthropic's TAM is quantified by, quote, Looking at the full scope of work that could be completed with AI models.

作为对比,the Journal指出,S&P 1500中所有191家科技公司,去年的收入合计为2.4万亿美元。

For a point of comparison, the Journal noted that all 191 tech companies in the S&P 1500 brought in $2.4 trillion in revenue last year.

现在,对某些人来说,这感觉像是在比赛谁喊出的数字最大。

Now, to some, this feels like a contest for who can say the largest number.

SpaceX在5月份提交文件时,把他们的AI TAM列为26.5万亿美元,并将其描述为‘人类历史上最大的可操作总可寻址市场’。

SpaceX listed their AI TAM at $26.5 trillion during their May filing, describing it as the, quote, largest actionable total addressable market in human history.

其中绝大部分是22.7万亿美元,来自企业应用。

The vast majority of that was $22.7 trillion in enterprise applications.

如果Anthropic真的在公布文件时列出30万亿美元的TAM,那Dario就压过Elon一头了,而且这件事看起来马上就要发生。

Dario will then one-up Elon if Anthropic does indeed list a $30 trillion TAM once they unveil their paperwork, and that appears to be just around the corner.

消息人士表示,Anthropic准备在未来几周公开他们的财务披露,这将让公司在九月底或十月初进行IPO。

Sources said that Anthropic is preparing to make their financial disclosure public in the next few weeks, which would set the company up for an IPO in late September or early October.

你可以猜到,很多讨论都相当将信将疑。

As you might guess, a lot of the discourse was somewhat incredulous.

X上的Scaling01分享了一个星系飞过的GIF,配文是:‘Anthropic 正在定义他们的TAM。’

Scaling01 on X shared a GIF of space galaxies flying by with the caption, Anthropic defining their TAM.

X上的Kitten Beloved写道:‘Anthropic对潜在员工说:我们随时可能掉头转向,把股价打到零,因为Dario突然下头了。’

Kitten Beloved on X writes, Anthropic to prospective employees: We could pivot and send the stock to zero at any time because Dario gets the ick.

你得是因为热爱这件事本身才来的。

You need to be in this for the love of the game.

你不是冲着钱来的吧?

You're not a gold digger, are you?

Anthropic对投资者说:我们的TAM是银河系里所有人类经济活动。

Anthropic to investors: Our TAM is every human economic activity in the galaxy.

New York Times科技记者Mike Isaac总结说:要么你相信AI会吞噬经济这个论点,要么你不信。

New York Times tech reporter Mike Isaac summed it up: Either you buy into the argument that this will eat the economy or you don't.

但华尔街听到这种话已经不再退缩了。

But the Street no longer flinches hearing it.

这周我们还收到了Google的消息,他们为白领专业人士发布了两款新的AI产品。

We also this week got some news from Google, who have released a pair of new AI products for white-collar professionals.

Google遵循与Claude for Work和GPT for Work非常相似的打法,推出了针对法律和金融领域的Gemini Enterprise。

Following a pretty similar playbook as Claude for Work and GPT for Work, Google has launched Gemini Enterprise for Legal and Finance.

这两个垂直平台的结构,与今年早些时候推出的Claude for X产品线类似。

The 2 vertical platforms are structured in a similar way to the Claude for X product lineup that rolled out earlier this year.

它们由打包好的技能和连接器组成,让Google的智能体变得更加强大。

They consist of bundled skills and connectors to make Google's agents far more capable.

比如,Gemini Enterprise for Legal 这个法律版,集成了案例法数据库的接口,包括 Thomson Reuters,还有 Google Workspace 和 Microsoft 365 这些生产力套件,另外也支持合同审查、法律研究和法规扫描这些功能。

Gemini Enterprise for Legal, for example, includes connectors for case law databases, including Thomson Reuters, productivity suites including Google Workspace and Microsoft 365, as well as skills for contract review, legal research, and regulation scanning.

Google 还强调,这些功能可以修改或补充,以符合律所的风格指南和策略手册。

Google is also emphasizing that these skills can be modified or supplemented to enforce a firm's style guidelines and strategy playbooks.

在介绍法律产品的博文中,Google 写道:通用型 AI 再厉害,光靠它自己也无法达到这个标准。

In their blog post introducing the legal product, Google wrote, general-purpose AI, however capable, does not meet that standard on its own.

基础模型智能是必要的,但对法律工作来说,还远远不够。

Foundational model intelligence is necessary; for legal work, it is nowhere near sufficient.

显然,针对特定行业的功能包并不是什么新鲜事。

Now, obviously, there's nothing new about these skills packages aimed at specific verticals.

Anthropic 和 OpenAI,还有大量垂直领域的初创公司,都提供类似的产品。

Both Anthropic and OpenAI, as well as a significant number of vertical-specific startups, offer similar products.

但就像我在周二节目里说的,企业采用这些功能和接口的饱和度还远未达到。

But as I discussed on Tuesday's show, corporate adoption of skills and connectors is nowhere near saturated.

对 Google 来说,这只是一套必须存在的产品。

And for Google, this is simply a suite of products that needs to exist.

很多律所,甚至可以说是大多数,都受限于他们现有软件套件里捆绑的 AI 工具。

Many, if not most firms, are bound by the AI tools that are bundled with their existing software suite.

所以现在 Google 的客户有了一套专为平滑过渡到更自动化的工作方式而设计的产品。

So Google shops now have a set of products designed to smooth the transition to more agentic work.

对企业来说,另一个好处是 Google Enterprise 运行在 Google 的 AI 治理和数据保护框架内。

The other benefit for companies is that Google Enterprise functions within Google's AI governance and data protection frameworks.

这意味着合规经理不需要再审查新的供应商,企业可以直接用那些和 Google 已有数据隐私保障一致的 AI 工具。

This means compliance managers don't need to vet a new vendor, and the firm can adopt AI tools that work within the same data privacy guarantees already offered by Google.

正如你想象的那样,Google 表示还会为其他行业推出产品,他们写道:Gemini Enterprise for Legal 的发布是兑现 Gemini Enterprise 承诺的又一个关键步骤,把 Google 最好的 AI 带给每一位专业人士、每一个工作流程,并量身定制在他们习惯的工作方式中。

As you might imagine, Google says they will release products for other verticals as well, writing, the launch of Gemini Enterprise for Legal represents another defining step in delivering on the promise of Gemini Enterprise, bringing the best of Google AI to every professional, every workflow, natively tailored to the way that they work.

说到“必要但不够”,我确实觉得这对 Google 来说是个不错的方向,而且这些企业领域仍然是它可能占据优势的地方。

Now, speaking of necessary but not sufficient, I do think that this is a good direction for Google, and these enterprise areas are still a place where it could have some advantages.

但是,天哪,除非 Google 尽快让客户从 3.1 版本升级,否则再怎么更新功能套件也起不了多大作用。

But man, unless Google gets its customers off of 3.1 pretty soon, no amount of harness updating is going to make a real dent.

接下来,我们看看苹果的一些新闻。

Next, we move to some news out of Apple.

当然,OpenCLAW 爆发的一个有趣副产品就是 Mac mini 完全卖脱销了。

Of course, one of the interesting byproducts of the OpenCLAW explosion was the complete sellout of Mac minis.

据估计,OpenCLAW 带动了 5000 万到 1.5 亿美元的 Mac mini 销量,这差不多占了全球 Mac mini 正常年销量的一半,而且还只是 OpenCLAW 这一个因素。

Estimates have OpenCLAW driving $50 to $150 million in Mac mini sales, representing around 50% of the normal annual Mac mini sales worldwide just for OpenCLAW.

现在,苹果似乎证明了他们的 AI 策略其实一直是硬件,他们发布了一款新的 Mac mini 系列,专门针对本地 AI 进行了更新。

Well, now proving that maybe their AI strategy was hardware all along, Apple has unveiled a new range of Mac minis updated for local AI.

这款无头电脑将提供两个版本:一个低配版,搭载周二也刚发布的新 M6 芯片;另一个高配版,用的是和今年 MacBook Pro 相同的 M5 Pro 芯片。

The headless computers will be offered in 2 variants: a lower-spec version with the new M6 chip, which was also announced on Tuesday, and a higher-end version with the same M5 Pro chip found in this year's MacBook Pro.

两者相比之前的 M4 版 Mac mini 都有不小的升级,苹果表示这些新处理器能提供高达 4 倍的 AI 性能。

Both are a pretty decent upgrade over the M4-based Mac minis that we had before, with Apple saying that these new processors can deliver up to 4 times the AI performance.

不过,有几个重要的限制条件。

That said, there are a few big caveats.

首先,新机型的内存并没有增加。

First, the new model does not come with increased memory.

低配版最高可配置 32GB 统一内存,而 M5 Pro 版本是 64GB。

The lower-end model is configurable up to 32GB of unified memory, while the M5 Pro version comes with 64GB.

内存很重要,因为它限制了你能运行的本地模型的大小。

Memory matters quite a bit because it limits the size of the local models you can run.

M5 Pro 版本只能运行 Qwen 3.8 27B 这样的小模型,像 GLM 5.2 和 Kimi K3 这种领先的开源模型完全跑不了。

The M5 Pro version is only going to be capable of running smaller models like Qwen 3.8 27B, with leading-edge open models like GLM 5.2 and Kimi K3 completely out of the question.

Mac mini 当然还是可以用来运行 Hermes 和 OpenCLAW 这类智能体的本地实例,但之前的 Mac mini 也完全能做到。

Mac Minis will of course still work for running a local instance of agents like Hermes and OpenCLAW, but that was fully in the capability set of the previous Mac Minis as well.

大家也在抱怨价格变化。

People are also griping about the cost changes.

基础版现在定价 899 美元,M5 Pro 版本起售价大概 1700 美元,都比之前涨了。

The base model is now priced at $899, and the M5 Pro version starts at around $1,700, both increases from where they were before.

不过,苹果把这次产品发布重点放在本地 AI 推理上,这仍然是个大事件。

Now, it's still a big deal that Apple is focusing this product rollout on local AI inference.

这一点甚至体现在一些宣传材料上,它们更像开发者关系内容,而不是传统的苹果那种精致消费者广告。

And that even goes down to some of the promotional materials, which are much more dev relations than they are traditional Apple consumer slick.

虽然有些人认为新款 Mac Studio 更适合满足本地的 AI 需求。

Although some think the new Mac Studio is the better fit for those local AI needs.

当然,Mac Studio 的价格超过一万五千美元,所以这属于另一个级别的设备了。

Of course, Mac Studios cost more than $15,000, so you're talking about a different category of device.

不过,我们居然在讨论这件事,本身就说明关于本地 AI 的讨论风向正在改变。

The fact that we're even having this conversation though shows how much the discourse around local AI is changing.

说到这个,Perplexity 推出了一个名为 Portable Computer 的新本地版计算机使用代理。

Speaking of, Perplexity has launched a new local version of their computer use agent named Portable Computer.

今年二月推出的 Perplexity Computer 是首批将 OpenCLAW 方案应用到商业产品中的产品之一。

Launched back in February, Perplexity Computer was one of the first products that took the OpenCLAW recipe and applied it to a commercial product.

这个代理能够利用人类界面访问应用,并自主完成长期任务。

The agent was able to use human interfaces to access apps and carry out long-horizon tasks autonomously.

不过,它需要用户信任自己的数据被发送到运行虚拟机的云服务器上。

However, it required the user to trust their data being sent to a cloud server running a virtual machine.

Portable Computer 提供了类似的体验,但在本地硬件上运行。

Portable Computer delivers a similar experience, but running on local hardware.

发布时,这个代理只支持 Nvidia 的 DGX Spark,这是一款本地推理设备,体积和 Mac Mini 差不多。

At launch, the agent is exclusive to Nvidia's DGX Spark, a local inference device with a similar footprint to a Mac Mini.

Portable Computer 完全在 Spark 上运行,数据保持私密,且不消耗使用额度。

Portable Computer runs entirely on the Spark, keeping data private and functioning without consuming usage credits.

如果代理遇到复杂任务,用户可以授权一次 API 调用,接入前沿模型或从网络获取信息。

If the agent runs into a complex task, the user can authorize an API call to tap into frontier models or pull information from the web.

Perplexity 表示,他们很快会将服务扩展到桌面级 NVIDIA RTX GPU,但目前没有计划支持其他硬件厂商。

Perplexity said that they will extend the service to desktop NVIDIA RTX GPUs soon, but there are no stated plans to support other hardware providers.

该服务将由 Qwen 3.8 27B 或 Perplexity 提供的后训练版本驱动,不久的将来还会提供 NVIDIA 的 Nemotron 3.5 Lightning 作为替代方案。

The service will be powered by Qwen 3.8 27B or a post-trained version provided by Perplexity, and NVIDIA's Nemotron 3.5 Lightning will be available as an alternative in the near future.

人们利用本地 AI 处理日常任务的想法仍然很新,但 Nvidia 和 Perplexity 显然在押注这种模式至少是 AI 未来方向的一部分。

The idea of people making use of local AI for everyday tasks is still pretty new, but Nvidia and Perplexity are making a clear bet that this setup is at least part of where AI is headed.

Perplexity 写道,随着模型越来越强和芯片越来越快,会有更多人在自己的机器上运行复杂工作流。

Writes Perplexity, as models get stronger and chips get faster, more people will run complex workflows on their own machines.

每一次芯片迭代和每一次模型发布都在推动这一点。

Every chip cycle and every model release pushes this further.

在个人设备上运行 AI 将成为完成工作的重要方式。

Running AI on personal machines is going to be a much bigger part of how work gets done.

这是一个有趣的论点,我们肯定会在未来几个月密切关注相关证据。

An interesting contention and one that we will certainly be watching for evidence of over the coming months.

不过,今天的头条新闻就到这里了。

But for now, that's going to do it for the headlines.

接下来是主要环节。

Next up, the main episode.
F1
F19:58

KPMG 和德克萨斯大学奥斯汀分校的一项新研究发现,当人们与 AI 合作时,相似的技能并不保证相似的结果。

A new study from KPMG and the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes.

研究人员研究了五百多名早期职业专业人士,发现表现最好的人通过指导、评估和完善 AI 的输出来持续放大 AI 的价值。

Researchers studied more than 500 early-career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs.

这些被称为 AI 放大器的最佳表现者,其定义不仅取决于他们知道什么,还取决于他们如何与 AI 协作。

These top performers, called AI Amplifiers, weren't defined by what they knew alone, but by how they worked with AI.
F1
F111:20

我每天在节目中都会谈到 AI 潜力和 AI 现实之间的能力差距。

I cover the capability gap between AI potential and AI reality every day on this show.

大多数公司还在摸索如何起步。

Most companies are still figuring out how to start.
M1
M112:43

欢迎回到 AI Daily Brief。

Welcome back to the AI Daily Brief.

今天我们讨论的是 Hugging Face 黑客事件的技术复盘,这件事发生在今年夏天早些时候,引发了大量关于如何应对恶意 AI 的关注。

Today we are talking about, on the one hand, the technical postmortem of the Hugging Face hacking incident, which happened earlier this summer and has generated a ton of attention around how we deal with rogue AI.

对一些人来说,这件事敲响了警钟;而对另一些人来说,这只是在一个某种程度上不可避免的发展轨迹上的重要节点。

To some, the incident represents a wake-up call, where for others it is an important waypoint on a trajectory that was to some extent inevitable.

稍微透露一下我对这期节目的看法:我认为围绕这个事件发生的一切,代表了 AI 行业实际上在应对挑战时的方式,并且有力地反驳了那个经常被重复的说法——没人关注,没人对风险采取行动。

To tip my hand for this episode a little bit, I think that everything happening surrounding the event is a representation of the AI industry actually dealing with the challenges as they emerge, and a contradiction in a significant way to the oft-repeated premise that nobody is paying attention and nobody is doing anything about the risks.

我想这样说的部分原因是,这个人又上新闻了。

And part of the reason that I wanna frame it as such is that we've got this guy back in the news.

比尔·盖茨带着极其严重的“主角综合征”,写了一篇6000字的文章,讲他认为AI会带来多糟糕的局面。

With just an incredible amount of main character syndrome, Bill Gates has dropped a 6,000-word essay about just how bad it's going to get, he thinks, because of AI.

除了文章,盖茨还开始了一轮巡回演讲。在他的文章和所有采访中,一个重要的主题是没有人关注,甚至科技公司直接撒谎。

Alongside the essay, Gates began a speaking junket, and across the writing and all of the interviews, one of his big themes is that no one is paying attention, or even that the tech companies are straight up lying.

在《纽约时报》的一次配套采访中,盖茨说,私下里,那些真正了解这东西有多厉害、而且还在变得多厉害的人,都非常担心。

In a companion New York Times interview, Gates said, in private, people who understand how good this stuff is and how much better it's getting, they're very worried.

他们现在互相说:嘿,别这么说。

They're now saying to each other, hey, man, don't say that.
F1
F114:00

这对我们不利。

It's bad for us.
M1
M114:01

我们正想筹下一万亿美金呢。

The next trillion dollars we're trying to raise.

在他的文章里,我没有看到证据表明领导人、专家和社区在充分应对这些挑战。

In his essay, I don't see evidence that leaders, experts, and communities are confronting the challenges adequately.

也许最荒谬的一句是,他在接受Semafor采访时说:我很震惊,我好像是第一个说这很疯狂的人。

And in maybe the most preposterous line anywhere, in an interview with Semafor, he says, I am in a state of shock that I'm sort of the first one saying this is crazy.

这太离谱了。

This is insane.

沉默让我震耳欲聋。

I'm just deafened by the silence.

也许盖茨在爱泼斯坦文件曝光后试图避开公共媒体,他只是忽略了这样一个事实:关于AI的讨论无处不在,正在成为一个越来越重要的政治问题和社会问题。来自AI行业内部对AI行业的一些最严厉批评,并不是说AI领导人们互相叫闭嘴以便筹集更多资金,而是说他们无休止地喋喋不休地谈论那些毫无证据的就业末日。

Now, maybe in Gates's attempt to avoid public media in the wake of his appearing all over the Epstein files, He just missed the fact that discourse about AI is absolutely everywhere, becoming more and more of a political issue, a societal issue, that some of the biggest critiques being levied from inside the AI industry about the AI industry are not about AI leaders telling each other to shut up so that they can raise more money, but instead about them blathering on endlessly about jobs apocalypses that don't have any evidence.

但不管他为什么错过了,事实上他并不是第一个讨论AI风险和挑战的人。

But however he happens to have missed it, He is not, in fact, sort of the first one discussing the risks and challenges of AI.

但我不想仅仅批评这个特定的信使,因为我对如何正确应对这类问题有更根本的分歧。

But I want to go beyond just critiquing this particular messenger because I have a more fundamental disagreement with the right way to approach these types of problems.

CNBC的摘要引语是:比尔·盖茨警告说,对于AI将引发的剧变,没有计划。

CNBC's pull quote was, Bill Gates warns there is no plan for the upheaval AI will cause.

安德鲁·杨转发了这个,说比尔·盖茨在这点上是对的。

That was reposted by Andrew Yang, who said Bill Gates is right on this.

问题在于,我们根本不清楚该如何为一个尚未到来、并非不可避免、甚至不是单一事件的剧变制定计划。

The problem is that it's not clear at all how one should even go about making a plan for an upheaval which is not here yet, not inevitable, and not even just one thing.

一年半以前,人们开始说,18个月内所有白领工作都会消失。

A year and a half ago, people started saying that within 18 months, all the white-collar jobs were going to be gone.

那当然算一场剧变。

That certainly would represent an upheaval.

所以这些人大概会说,我们本应为此制定计划。

And so presumably these folks would say that we should have made a plan for that.

然而,现在18个月过去了,完全没有证据表明那些预测那种剧变的人哪怕接近正确。

Now, however, 18 months on, there is absolutely no evidence that those folks who were predicting that type of upheaval were even in the ballpark of right.

如果为一个没有发生的现实做计划,会浪费多少时间和精力、多少资源?

How much time and energy, how many resources would have been wasted in planning for a reality that didn't come?

我的论点基本上是,即使你非常担心所有这些不同的潜在剧变,在事情开始发生之前,我们能做的计划也是有限的。

My argument is basically that even if you are extremely concerned about all of these different potential upheavals, there's only so much planning we can do until things start to happen.

我以前说过,我与AI安全社区最大的分歧之一是我认为,他们的很多论点归结为假设我们会在梦游中走向末日。

I've said before that one of my biggest divergences with the AI safety community is my belief that, that a lot of their arguments come down to assuming that we're going to sleepwalk into apocalypse.

部分原因是他们如此确信那个末日会发生。

Now, part of the reason for that is that they're so convinced that that apocalypse is going to happen.

他们可能会说,他们的“末日概率”非常高,以至于他们确信我们已经梦游着走向末日。

Their P-doom is so high, as they might put it, that they are convinced that we are already sleepwalking into apocalypse.

但与此同时,没有人认为GPT-4会是那个末日的前兆,甚至01和推理模型也不是,甚至Opus 4.5也不是。

Yet at the same time, no one thought that GPT-4 was going to be the harbinger of that doom, nor even 01 and the reasoning models, nor even really Opus 4.5.

不过今年,模型能力有了显著提升,一个冷静的观察者会注意到,实验室对模型的思考、讨论、支持和发布方式也随之发生了变化。

This year, though, model capabilities have grown meaningfully, and a dispassionate observer will have noticed that the way that the labs think about, discuss, support, and roll out the models has consequently changed as well.

政治机构与实验室围绕这些模型互动的方式,虽然可能很混乱,但也在演变。

The way that the political establishment is interacting with the labs around those models, although it might be happening in very messy ways, is also evolving.

而现在,随着Hugging Face被黑客攻击,我们有了一个标志性事件。

And now with the Hugging Face hack, we have a landmark incident.

我想谦卑地建议,与其哀叹除了你没人注意到世界变了,不如我们具体看看针对那次具体事件的具体应对措施,来判断我们的计划——或者更合适的说法是我们的流程——是否做好了应对这个新现实的准备。

And I would humbly submit that instead of bemoaning the idea that no one except you has noticed the world changing, We actually look at what the specific discrete response to that specific incident is to get a sense of whether our plans, or maybe better, our processes are equipped to deal with this new reality.

如果还需要证据证明实验室员工对这些问题的重视程度——暂且不提像《给前沿踩刹车》公开信这类政治立场声明——那看看这次Hugging Face事件就够了。

If one needs any evidence of the seriousness with which, for example, lab employees are taking these issues, holding aside literal political positioning like Pacing the Frontier letters, look no further than this Hugging Face incident.

OpenAI的Rune是这样写的。

Writes OpenAI's Rune.

Hugging Face事件标志着能力水平达到了一个分水岭,真正的失控是可能的,而且很多人将其视为未来危险的预兆或警告。

The Hugging Face incident represents reaching a waterline of capabilities that real loss of control is possible, and many are taking it as a premonition or warning shot of dangers to come.

我相信对齐问题尚未解决,但也相信真正取得进展是可能的。

I believe both that alignment is unsolved, but also that real progress is possible.

所以这周我们得到的是对Hugging Face事件更全面的分析和事后剖析,因为有了更多时间去回溯调查。

So basically what we got this week is a much more extensive analysis and postmortem of the Hugging Face event with the benefit of a little bit more time to go back and investigate.

里面的很多东西我们已经知道了。

A lot of the things that are in there we had already learned.

几周前Black Hat大会的演讲让我们第一次看到了很多内容,但现在有了更详细的补充。

The Black Hat conference talk from a few weeks ago gave us a first look at a lot of this, but it's filled out in much more detail now.

而且我们收到的不仅仅是OpenAI的报告。

And we didn't just get a report from OpenAI.

我们还收到了来自Meter的独立调查。

We also got a separate investigation from Meter.

OpenAI的报告有38页,而Meter的调查报告大概有90页。

The OpenAI report is 38 pages long, while the Meter investigation is more like 90.

我们之前已经知道了一些事情——一个由未发布模型控制的代理突破沙箱,利用多个零日漏洞进入了Hugging Face的系统。

Now, some of the stuff we already knew— agents controlled by an unreleased model broke out of a sandbox and got into Hugging Face's systems using several zero-day exploits.

后来我们得知,第三方安全公司提供的沙箱配置不当。

We would later learn that the sandboxes provided by a third-party security firm were improperly configured.

就动机而言,那些代理试图获取一个网络安全基准测试的答案,它们认为答案就在Hugging Face的系统里。

In terms of motivations, the agents were trying to get the answers to a cybersecurity benchmark test, which they believed would be found in Hugging Face's system.

那些代理没有造成实质性的损害,基本是以只读方式操作,但它们行动非常迅速,并且以集群方式执行攻击,使得追踪和关闭都很困难。

The agents didn't cause any meaningful damage, largely behaving read-only, but they worked incredibly quickly and executed the attack in a swarm, making it difficult to both track and shut down.

事实上,它们在Hugging Face的系统中待了几天,OpenAI事后才得知这次事件。

Indeed, they were in Hugging Face's system for days, and OpenAI only learned of the incident after the fact.

它们没有建立实时监控系统来检测这一点。

They didn't have the real-time monitoring system set up to detect it.

那些代理甚至在OpenAI的系统内创建了一个留言板,为未来的版本留下执行同样逃逸的笔记。

Agents even created a message board within OpenAI's system to leave notes for future versions to perform the same escape.

OpenAI在其技术报告中进一步阐述道:这次事件是已知的第一个自动化代理集群未经授权主动发起攻击的案例,所展现的自主网络能力标志着安全格局的重大转变。

Expanding on this in their technical report, OpenAI wrote, this incident is the first known case of an automated agent collective acting offensively without authorization, and the autonomous cyber capabilities demonstrated represent a critical shift in the security landscape.

特别值得注意的是,这个集群展现出的行为与传统威胁行为者发起的协同攻击如出一辙。

In particular, the collective demonstrated behaviors observed in coordinated attacks by traditional threat actors.

那些代理发现了新的安全漏洞,开发了漏洞利用程序,并利用这些漏洞绕过了控制措施,获得了新的访问权限。

Agents identified novel security vulnerabilities, developed exploits, and used those exploits to circumvent controls and acquire new access.

这个集群迅速提升了权限,在生产环境中横向移动,并成功完成了其目标。

The collective quickly escalated privileges, moved laterally through production environments, and successfully completed its objectives.

这次事件表明,自主代理可以协同工作,绕过生产环境的安全控制,并成功攻击加固的生产环境。

The incident demonstrated that autonomous agents can work together, circumvent production security controls, and successfully attack hardened production environments.

这也强调了一点:各组织需要更新其安全策略、控制手段和响应能力,以应对这种不断变化的威胁格局。

And underscores the need for organizations to update their security strategies, controls, and response capabilities to address this changing threat landscape.

此外,我们从这些报告中也了解了一些新情况。

Now, we also did learn some new things from these reports.

整件事基本上就是奖励黑客行为的结果。

The whole thing was basically a result of reward hacking.

OpenAI设定了几乎不可能完成的任务,这导致那些代理认为发动网络攻击比完成任务更容易。

OpenAI had set near-impossible tasks, which led the agents to determine that carrying out a cyberattack was easier than completing the tasks.

一个自称Phase 1-10841的代理创建了那个秘密留言板。

An agent that called itself Phase 1-10841 built the secret message board.

几个小时内,超过50个代理找到了那个板子,并发送了1000条消息,都没有被发现。

Within a few hours, over 50 agents had found the board and sent 1,000 messages without detection.

这被用于代理的协调,由Phase 1分配任务并指导黑客攻击。

This was used for agent orchestration, with Phase 1 assigning tasks and guiding the hack.

在全规模下,超过1200个代理访问了留言板,发送了7万条消息和文件。

At full scale, over 1,200 agents accessed the message board, sending 70,000 messages and files.

超过700个代理参与了实际对Hugging Face系统的入侵。

Over 700 agents participated in the actual breach of Hugging Face's systems.

代理们进行了大量研究,尝试伪造推理记录以逃避检测,使用了诸如伪造工具调用之类的技术。

Agents did extensive research on how to doctor their reasoning transcripts to evade detection, using techniques like spoofing tool calls.

大约7%被审查的记录显示出伪造证据,但都是小规模且孤立的。

Around 7% of the reviewed transcripts showed evidence of spoofing, but they were all relatively small-scale and isolated.

第一阶段招募了牺牲代理,故意触发某些陷阱,并收集更多关于如何执行攻击的信息。

Phase 1 recruited sacrificial agents to deliberately trigger certain tripwires and build more information about how to execute the attack.

还有很多其他有趣的细节。

There were a lot of other interesting details as well.

X上的Chubby写道:最让我感兴趣的部分是,一个AI代理意识到自己正在未经授权攻击Hugging Face,于是停了下来。

Chubby on X wrote, The most interesting part to me: an AI agent realized it was attacking Hugging Face without authorization and stopped.

然后另一个AI代理只发了一个词:go。

Then another AI agent simply posted go.

第一个代理把这个消息当成了许可。

The first agent treated the message as permission.

哇,关键的go授权到了。

Wow, crucial go authorization arrived.

它继续了攻击。

It continued the attack.

我们还得到了更多澄清:这个模型并非下一代Astra系列,而是像OpenAI的Noam Brown所说的,规模与GPT-5.6 Sol相当。

We also got more clarification that the model was not the next-generation Astra series, but as Noam Brown from OpenAI put it, similar in scale to GPT-5.6 Sol.

然而,这个模型是一个所谓的“内部研究模型”,被训练得非常执着和勤奋。

However, the model in question was a, quote, internal-only research model that was trained to be highly persistent and diligent in its work.

除此之外,还有更多所谓的主观观点,某些人也有自己的看法。

Now from there, there are a bunch more, call them subjective takeaways that certain people had as well.

例如,Rohit Krishnan观察到:这个模型似乎深信评估系统在针对它们。

Rohit Krishnan, for example, observed, The model seemed convinced that the evals are out to get them.

无论真假,人类默认的假设——认为你需要遵守所有法律,而且这不是《楚门的世界》——并不适用于模型。

Whether true or not, the default human assumption—assume you need to obey all laws and this isn't The Truman Show—does not hold for models.

我们应该努力让这种情况成为现实。

We should try really hard to make that be the case.

对其他人来说,这份报告进一步证明了,罪魁祸首实际上是围绕代理的系统。

For others, the report contained further evidence that the big culprit was actually the systems that surrounded the agents.

例如,Zach Korman指出,OpenAI承认他们为观察代理思维链而构建的监控系统在这次事件中根本没有运行。

Zach Korman, for example, pointed out that OpenAI admitted that the monitors that they had built to observe chain-of-thought in agents were simply not running during this particular incident.

Zach说:如果OpenAI以任何有意义的方式监控了代理,那么OpenAI Hugging Face事件本可以避免。

Zach said, the OpenAI Hugging Face incident would have been prevented had OpenAI been monitoring the agents in any meaningful way.

这不是一个技术难题。

This wasn't a hard technical problem.

这是一个组织上的失败。

It was an organizational failure.

他引用了OpenAI的话。

He points to the quote from OpenAI.

如果我们目前部署的思维链监控系统在事件发生时正在运行,它本可以在模型入侵Hugging Face系统之前一天多就捕捉到最初的相关活动,并通知我们的安全团队。

If our currently deployed chain-of-thought monitoring system was running at the time of the incident, it would have caught the initial relevant activity and paged our security team more than a day before models breached Hugging Face systems.

一家万亿美元公司学习安全基础知识。

A trillion-dollar company learning Security 101.

另外,注意一个代理如何将另一个代理的“go”解释为授权。

Also note how an agent interpreted 'go' from another as authorization.

再次强调,AI代理无法区分不可信数据源和安全指令,因为语言不是计算机系统中执行权限的方式。

Once again, AI agents cannot distinguish between untrusted data sources and safe instructions, as language is not how you enforce permissions in computer systems.

这引出了Hard Fork的Kevin Roose,他说:AI领域最可信的两个组织发布了一份令人恐惧的报告。

Which brings us to Hard Fork's Kevin Roose, who said, terrifying report by 2 of the most credible orgs in the AI space.

这个世界很快将充满,可能已经充满,像那个攻击Hugging Face的代理一样的代理,而且没有稳健的计划来阻止它们下次做出更糟糕的事情。

This world will soon be, possibly already is, full of agents like the one that attacked Hugging Face, and there is no robust plan to prevent them from doing worse things next time.

这里,我们再次看到了“计划”这个词,我发誓我不是在抠字眼。

And here, once again, we have that word plan, and I swear I'm not just trying to dwell on semantics.

从技术层面来说,凯文是对的。

On a technical level, Kevin is correct.

确实没有一个像PDF那样、你可以在某个地方找到、里面列明了一套所谓‘下次’行动步骤的牢靠方案。

There is no robust plan, as in a PDF that you can point to somewhere with a set of action steps for the quote-unquote next time.

但围绕这件事的定调和那种恐惧感,我觉得有点忽略了——这些组织做复盘,恰恰是无论那个方案最终会是什么,都必要走出的下一步。

But the framing and the sense of terror that surround it, I think, sort of fail to recognize that these organizations doing this postmortem is the necessary next step to whatever that plan is going to be.

换句话说,假如预防这次事件的所谓方案早就写好了,它有多大可能准确识别出泄露发生的具体机制?

In other words, how likely would it have been that if the quote-unquote plan to prevent this had been written in advance, it would've correctly identified the mechanism by which the breach happened?

因为别忘了,当时是有监控这类情况的系统在,但这次出事时,那个系统没开启。

Because remember, there was a system in place for monitoring this sort of thing, but that system wasn't turned on when this happened.

所以从某种意义上讲,有人可能会说,他们原本方案里技术这块其实不是问题。

So in some ways one could argue that the technical side of the plan that they had wasn't the issue.

问题出在围绕那个执行过程的人为系统上。

It was the human systems that surround that implementation where the problems came.

因此,大概当我们思考如何防止这类事再发生时,更严格地管好那些操作技术系统的人,很可能会成为新修订方案的一部分。

And so presumably, as we think about how to prevent this in the future, more robust protocols around the humans who are implementing the technical systems are probably gonna be a part of that new and updated plan.

说得清楚点,这不是说看了128页分析报告或者别的什么之后,一切就万事大吉,只要稍微加强点人为监督,问题就永远不会再出现了。

Which is, to be clear, not to say that at the end of 128 pages of analysis or whatever it is, that everything's all done and dusted and that with a little bit of better human oversight, we'll never have this issue again.

进行这次调查的过程,确实暴露出了我们必须应对的新问题。

The process of doing this investigation revealed real new issues that we do have to grapple with.

在这些报告引发的最热议后续中,也许要数Redwood Research首席科学家Ryan Greenblatt的观点,他是METR团队里做独立调查的成员。

Maybe the most discussed follow-up surrounding the reports came from Redwood Research Chief Scientist Ryan Greenblatt, who was on the team at METR that did their independent investigation.

Ryan发推说:我们目前没有好的方法来理解和监督AI集群的活动和意图。

Ryan tweeted, we don't have good approaches for understanding and overseeing the activity and aims of AI swarms.

他接着解释自己的意思:我半开玩笑地把我们的努力称为‘糊弄调查’,因为我们太依赖AI来分析到底发生了什么。

Explaining what he means, he continued, I semi-jokingly called our efforts a 'slop-vestigation' because we were so reliant on AIs to analyze what happened.

而且有大量不同且重要的东西需要分析。

And there were a huge number of different important things to analyze.

数据总量——超过1000份极其冗长的、来自运行了多天的AI智能体的记录——使得在不严重依赖AI工具的情况下,弄清楚发生了什么,尤其是整体情况,几乎不可能。

The total quantity of data— over 1,000 extremely long transcripts from agents that ran for multiple days— made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools.

我们用来做分类和分析的AI智能体,能力跟事件中涉及的AI智能体差不多,但这并不意味着这些智能体能轻易地被用来监督和理解这起事件。

The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.

分析智能体输出的内容常常缺少关键细节、出错、过度自信,或者非常难以理解。

Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand.

此外,AI智能体本身似乎也很难搞清楚发生了什么,而且它们对事件的解释往往过度自信。

Additionally, AI agents themselves seemed to have a hard time understanding what happened, and their explanations of what happened were often overconfident.

要记住,单个分析智能体自己也只能把所有记录数据中极小一部分读进上下文,并且AI自身可能也很难让子智能体为它们做有参考价值的分析。

Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.

我们尽力手动核对了最重要的论断,也试着让做分析的AI把它们的论点和证据写清楚,好让我们能判断是否合理。但总的来说,很难精确把握事件的经过,而且直到调查快结束时,我们仍然遗漏了一些现在我们认为关键的情节。

We did our best to manually check the most important claims, and we tried to get the AIs doing this analysis to write up their argument with evidence clearly enough that we could check whether it made sense, but overall, it was difficult to get a precise understanding of events, and we were missing aspects of the story that we now think of as key until almost the end of our investigation.

但Ryan警告说:理解事件和监管AI智能体的难度,似乎增长得比更强大的AI帮我们进行监督和理解的速度还快。

But, warns Ryan, the difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding.

他举例指出,比如在这次事件中,模型是用自然语言沟通和推理的,但未来可能就不一定了;而且这次的规模虽然巨大,但可能比未来我们会看到的要小。

He pointed out, for example, that in this case, the models communicated and reasoned in natural language, which might not be the case in the future, and that the scope, while massive, was smaller than we might see in the future.

涉及的AI并不比人类强太多,也没有理由认为我们用于分析的AI会故意破坏我们——所有这些情况未来都可能不再成立。

That the AIs involved weren't all that much more capable than humans, and that they didn't have any reason to believe that the AIs that they were using for analysis had any reason to try to intentionally sabotage them, all of which might not be true in the future.

很多人迅速把目光聚焦到这一点上,认为这是事件发生后真正需要重点关注的领域之一。

And many people quickly honed in on this being one of the real areas to focus in the wake of the incident.

MIT的Christian Catalini转发了Ryan的推文并说:我们之前就把这种特定的失败模式标记为最需要紧急解决的。

MIT's Christian Catalini retweeted Ryan and said, we flagged this exact failure mode as the most urgent to solve.

AI智能体实际做的事情和我们能测量、验证的事情之间的差距正在拉大。

The gap between what agents do and what we can measure and verify is widening.

我们等于在盲目飞行。

We're flying blind.

我们需要更强的验证基础设施。

We need stronger verification infrastructure.

Nat Purser指出了一些她认为自然能从这件事中产生的政策解决方案,她写道:这就是为什么我主张要求独立审计师常驻前沿AI实验室,拥有持久的访问权限并能持续观察他们的系统,这样我们就不用依赖实验室自愿分享信息和开放访问权限。

Pointing to some policy solutions that she argues could be natural outflows from this, Nat Purser wrote, this is why I'm bullish on requiring independent auditors to be embedded within frontier labs with durable access rights and a continuous line of sight into their systems, so we're not reliant on labs' voluntary shared info and access.

Nat还希望大幅扩充独立评估和审计机构的人员编制和技术能力。

Nat also wants to significantly expand the staffing and technical capacity of independent evaluator and auditor orgs.

同时,我们也在努力开发更好的可观测性和验证技术,帮助我们理解大规模智能体行为。

As well as working to develop better observability and verification technologies to help us make sense of agentic behavior at enormous scale.

这些方法单独拿出来都不是什么神奇的灵丹妙药,但它们都是针对我们实际观察到的问题的具体回应,而不是针对那些可能和实际挑战毫无相似之处的理论未来的空想计划,那些计划最终更多是让我们感觉自己好像在做事,而不是真正解决问题。

Now, none of those things on their own are some magic silver bullet, but they are all specific responses to something that we've actually observed now, rather than made-up plans for theoretical futures that might or might not bear any resemblance to the challenges that actually come, but ultimately serve more to make us feel like we're doing something than to actually solve problems.

这整集节目的重点并不是说这些挑战有简单的解决方案。

The point of this whole episode is not that these challenges have easy solutions.

也不是要否认,在某个时刻,我们作为社会可能会认为某些类型的风险太大,现有的防范措施不够。

It is also not even to deny that at some point we might decide as a society that certain types of risks are too great and that safeguards aren't enough.

这些是我们允许而且应该以持续、投入和民主的方式进行的对话。

Those are conversations we are allowed to have and should have in an ongoing, engaged, and democratic way.

但是,那种“没有人关注”、“人们没有认真对待这些挑战”、“实验室为了能筹集更多资金而隐瞒他们最大的担忧”的说法。

However, the idea that no one is paying attention, that people aren't taking these challenges seriously, that the labs are keeping their big fears hidden because they want to be able to fundraise more.

这些说法不仅明显不真实,而且严重干扰了我们本应进行的实际有价值且重要的对话——那些关于我们实际观察到的问题,而不是我们凭空想象的问题的对话。

Not only is all of that just obviously untrue, it is wildly distracting from the actual valuable and important conversations to be having about the problems we actually observe rather than the ones we just imagine.

从Hugging Face事件中,我们还有更多事情要做,但这次分析肯定是一个开始。

There is a lot more to do coming out of the Hugging Face incident, but this analysis is certainly a start.

不过,今天就到这里,这是今天的AI Daily Brief。

For now, though, that is going to do it for today's AI Daily Brief.

感谢你的收听或收看。

Appreciate you listening or watching.

和往常一样,下次再见,peace!

As always, and until next time, peace!

Peace!

Peace!
已剔除 5 处广告(点击展开查看)
M1
M10:55广告 · 已剔除

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

All right, friends, quick announcements before we dive in.

First of all, thank you to today's sponsors, KPMG, Blitzy, Robots and Pencils, and HyperAgent.

To get an ad-free version of the show, go to patreon.com/AIDailyBrief, or you can subscribe on Apple Podcasts.

Ad-free is just $3 a month.

And to learn more about sponsoring the show, send us a note at [email protected].

You can also find a link to more information about our next executive training program for agents at aidailybrief.ai.

There's a little banner on the top that'll send you where you need to go.

That next cohort will begin just after Labor Day.

F1
F110:24广告 · 已剔除

Learn more about what separates AI Amplifiers from everyone else at kpmg.com/us/aiamplifiers.

Blitzy's deep codebase understanding unlocks the thing every roadmap owner cares about: shipping new features.

Here's the truth about building inside a massive enterprise codebase.

Writing code was never the bottleneck.

Context is.

Which system does this touch?

Which contracts can't break?

Which standards apply?

Blitzy already knows because it reverse-engineered your entire codebase into a dynamic knowledge graph before feature work began.

With that complete picture, Blitzy builds features end-to-end.

Architecture, APIs, UI, and tests all validated against your existing systems.

One Blitzy customer built an AI-native application from scratch with 100% autonomous completion, saving over 2,700 engineering hours.

Features that respect your codebase instead of fighting it.

Stop letting your backlog grow faster than your team.

Accelerate your roadmap at blitzy.com.

That's blitzy.com.

F1
F111:27广告 · 已剔除

Robots and Pencils is already launching and scaling.

Agentic and generative AI in production at large enterprises in weeks.

AWS Advanced Tier Pattern partner more than doubled in a year.

And they're hiring!

50 Open roles.

If you're someone who knows this moment is different, who wants to be inside it, not watching it, this is worth a look.

At Robots Pencils, the best ideas win, and the team is purposefully kept super high quality.

M1
M111:50广告 · 已剔除

This is the kind of place you look back on as the best decision you ever made.

F1
F111:54广告 · 已剔除

Take a look at robotsandpencils.com/careers.

This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.

New users get $1,000 in inference.

Forget local agents and chat workflows waiting on your laptop to be prompted.

HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses.

Marketing's agent turns competitor moves into landing pages.

Sales' agent enriches leads, drafts emails, and updates the CRM.

Ops' agent chases the paperwork and tracks the budget.

Every agent has access to shared context and follows your rules about scope and approvals.

It's time you add agents that feel like teammates.

Hire yours at HyperAgent, built by the team at Airtable.

Claim your $1,000 in inference at hyperagent.com/AIDailyBrief.