← 返回任务列表

What the Top AI Users Are Doing Differently

258 段 · 2 位说话人 · 原片 27:38
M1
M10:00

普通AI用户和最先进的AI用户之间一直存在差距,但天哪,这个差距已经变大了。

There's always been a gap between an average AI user and the most advanced AI users, but my goodness, has that gap grown.

在最近发布的研究中,OpenAI显示,最先进用户和普通AI用户之间的差距已经从一月份的2.6倍增长到六月底的8.3倍。

In recently released research, OpenAI showed that the gap between the most advanced users and the average AI user had grown from 2.6x back in January to 8.3x by the end of June.

换句话说,最先进的AI用户使用的AI量是普通用户的8倍。

In other words, the most advanced users of AI were using 8 times as much AI as were their average counterparts.

原因当然是AI代理。

The reason, of course, is agents.

今年年初,基于代理的使用场景开始变得可行,并显著提升了AI能承担的工作的难度、复杂性和重要性。

At the beginning of the year, agentic use cases became viable and significantly upgraded the difficulty, complexity, and importance of the work that AI could take on.

顶级用户已经一头扎进去,弄明白了如何大幅提高他们从AI使用中获取的价值。

The top users have jumped in headfirst, figuring out how to significantly increase the value they get from their AI usage.

而普通用户呢,根本就没有这样做。

The average users, on the other hand, just haven't.

但随着高级用户利用AI代理接管越来越有价值的工作,这个差距只会越来越大。

But as the power users use agents to take on increasingly valuable work, that gap is just poised to grow.

《AI每日简报》是一个关于AI领域最重要新闻和讨论的每日播客和视频节目。

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.

好了朋友们,在我们深入之前先快速说几句公告。

All right, friends, quick announcements before we dive in.
M1
M11:21

你也可以在aidailybrief.ai网站上找到相关信息。

You can also find information on the aidailybrief.ai site.
M1
M11:48

根据The Information看到的路线图文件,Meta正在完成他们的消费级AI代理的收尾工作,预计在未来几周内发布。

According to roadmap documents viewed by The Information, Meta is putting the finishing touches on their consumer agent ahead of release in the coming weeks.

这个消息我们已经听说了一段时间,但随着产品即将上线,我们得到了更多细节。

Now, this is something we've been hearing about for a while, but we're getting more details as the product becomes imminent.

在内部,这个产品被称为Hatch,听起来有点像Grok系列,提供一种更精简的OpenAI风格代理体验。

Internally, the product is known as Hatch and sounds like it could be sort of in the Grok family of delivering a more streamlined version of an OpenAI-style agent experience.

据报道,该公司正在考虑将Hatch作为新的AI代理订阅的一部分,这可能为高使用量账户定下每月200美元的价格。

The company is reportedly looking at using Hatch as part of a new AI agent subscription, which could justify a $200 a month price tag for high-usage accounts.

Meta还计划在WhatsApp上推出一个新平台,以便更好地集成第三方AI代理。

Meta also plans to launch a new platform on WhatsApp to allow better integration for third-party agents.

该平台据说允许多个代理通过WhatsApp消息相互协调,这再次模仿了Grok的一些功能。

The platform will reportedly allow multiple agents to coordinate with each other using WhatsApp messages, again, which is mirroring some of the functionality, like I said, of Grok.

说实在的,我完全不觉得这是Meta在抄袭Grok。

For what it's worth, this doesn't strike me at all as Meta cribbing off of Grok.

我觉得这更像是我们未来会经常看到的交互模式。

I think these are just interaction patterns that we're likely to see more of.

推出可能会最早在本周开始,先向一小部分客户提供预览。

The rollout could begin as soon as this week as a preview to a smaller group of customers.

最后,在Meta模型新闻方面,一个名为Watermelon的更大模型正在准备于10月发布。

Finally, in Meta model news, a larger model known as Watermelon is being prepared for an October launch.

今年7月,Meta的AI首席执行官Yann LeCun告诉员工,Watermelon在内部基准测试中已经赶上了GPT-4.5。

Back in July, Meta's AI CEO Yann LeCun told staff that Watermelon had already caught up with GPT-4.5 on internal benchmarks.

显然,前沿技术已经随着GPT-5.6的发布大幅推进,到10月可能又会进一步变化。

Now, obviously, the frontier has moved forward substantially with the release of GPT-5.6 and will likely move again by the time October arrives.

所以我们得看看Watermelon是否真的能跟上步伐,还是继续落入Meta接近前沿却从未真正达到的范畴。

So we'll see whether Watermelon can actually keep pace or continues to fall in the column of Meta getting closer to the frontier without actually reaching it.

说到GrokBot,你们很多人测试它的一个障碍就是它极高的价格。

Speaking of GrokBot, One of the barriers for a lot of you guys testing it has been its extreme premium pricing.

实际上,不仅仅是价格高,而且定价还有点混乱。

In fact, it wasn't just high pricing, it was kind of confusing pricing.

起初,用户不确定是否需要同时订阅Cursor Ultra和Super Grok Heavy,那样每月总共要花500美元才能使用该服务。

Initially, users weren't sure if they had to subscribe to both Cursor Ultra and Super Grok Heavy, which would be a total of $500 a month to access the service.

但后来很多人,包括我自己在内,只通过Super Grok就拿到了,而这本身也不算便宜的订阅。

But then many, myself included, were able to get it just through Super Grok, which is itself not a cheap subscription.

但很明显,这是一次有意限制速度的发布,以确保系统不会崩溃,并希望之后能降低价格。

But it was very clear that this was an intentionally rate-limiting launch to make sure that things didn't go down, with the hopeful anticipation that prices would be reduced later.

从本周开始,GrokBot已包含在每月60美元的Cursor Pro订阅和每月100美元的Super Grok订阅中。

As of this week, GrokBot is included in the $60 a month Cursor Pro subscription, as well as the $100 a month Super Grok subscription.

说到降价,OpenAI 也在下调 GPT-5.6 SOL 的价格。

Speaking of dropping prices, OpenAI is also dropping prices for GPT-5.6 SOL.

通过 API 访问 SOL 现在每百万输入 token 花费 4 美元,每百万输出 token 花费 20 美元,之前分别是 5 美元和 30 美元。

Accessing SOL over the API will now cost $4 per million input tokens and $20 per million output tokens, down from $5 and $30 respectively.

上个月末,LUNA 和 TERRA 的价格也已经下调了。

Costs for LUNA and TERRA were already cut late last month.

很多人在猜测,这是 OpenAI 试图在 Anthropic 进行 IPO 之前对其施加压力。

Many are speculating that this is OpenAI trying to put pressure on Anthropic ahead of their IPO.

我的猜测是,不管这带来什么额外好处,OpenAI 更可能只是意识到他们面临一系列新的挑战,也就是我们每周在节目中讨论的那些。特别是企业客户不再会为了一切都使用最昂贵、最前沿的模型,而是必须以更复杂的方式思考他们的整个模型栈。

My guess is that for whatever ancillary benefit that might have, OpenAI is more likely to just be realizing that they've got a new set of challenges based on what we talk about every week here on this show, that especially business customers are not just going to use the most expensive state-of-the-art model for everything anymore and have to think in more sophisticated ways about their complete model stack.

如果 OpenAI 有足够的算力来更便宜地交付前沿模型,那看来他们决定这么做是有道理的。

To the extent that OpenAI has the compute to deliver their frontier models cheaper, it seems like they've decided it makes sense to do so.

接着昨天的消息,Business Insider 报道,Hugging Face 正在以 130 亿美元的估值寻求收购。

Now, following up on news from yesterday, Business Insider reported that Hugging Face was courting an acquisition at a $13 billion valuation.

据与 The Information 交流的消息人士称,可能支持他们高要价的部分原因是,公司现在的年化收入超过 1.5 亿美元,比两个月前增长了 50%。

And according to sources speaking with the information, part of what might justify their high asking price is that the company is now generating more than $150 million in annualized revenue, which is up 50% from 2 months ago.

相对于编码代理创业公司,那个数字可能显得很低,但 Hugging Face 是一家特意不专注于创收的公司。

Now, that number may seem low relative to, for example, the coding agent startups, but Hugging Face is a company that has specifically not been focused on generating revenue.

97% 的用户完全免费访问该平台,包括下载最新的模型权重。

97% Of users access the platform entirely for free, including downloading the latest model weights.

Hugging Face 的主要收入来源是高级账户和企业级账户,这些账户通过与超大规模云服务商合作提供推理服务。

The primary revenue drivers for Hugging Face are premium and enterprise-grade accounts serving inference in partnerships with hyperscale clouds.

基本上,到目前为止,平台的盈利部分一直是为了补贴开放模型的免费托管和分发。

Essentially, up until now, the profit-seeking segments of the platform have existed to subsidize the free hosting and distribution of open models.

六月,CEO Clem Delangue 表示,高级账户数量在上半年翻了一番,基于这一点,平台正在接近盈利。

In June, CEO Clem Delangue said that the number of premium accounts had doubled in the first half of the year, and based on that, the platform was approaching profitability.

对于收购方来说,这很可能不是一种严格的基于收入倍数的谈判。

For acquirers, this is very likely not a strict revenue multiple type of conversation.

总结一下这个逻辑,Jess Fields 写道,Hugging Face 的估值应该和 Cursor 一样高,远超过 130 亿美元,可能是其3到4倍。

Summing up the logic, Jess Fields writes, Hugging Face should be worth as much as Cursor is, way more than $13 billion, maybe 3 to 4 times that.

考虑到 Hugging Face 是开放权重运动挑战前沿把关的中坚力量,它在整个经济中占据着独特的强大地位。

Considering that Hugging Face is the backbone of the open weights challenging frontier gatekeeping, it occupies a uniquely powerful position in the entire economy.

现在,人们猜测可能与 Hugging Face 产生有趣契合的公司之一,当然就是 Nvidia。

Now, one of the companies that people are speculating on might be an interesting fit for Hugging Face is, of course, Nvidia.

就算没有其他原因,他们现在似乎每一笔收购都在讨论之列。

If for no other reason than they seem to be in the conversation for every acquisition right now.

事实上,最近几天关于 Nvidia 的交易动作的报道非常多,有些人在猜测 Jensen 到底在构建什么。

In fact, with a flurry of reporting around Nvidia's dealmaking in recent days, some are wondering just what Jensen is building.

过去一周,我们听说 Nvidia 签署了一项协议,从 Poolside 获得技术许可并引进人才,收购数据标注公司 Merkur 的股份,并且可能以令人费解的 300 亿美元估值投资 Perplexity。

Over the past week, we've heard that Nvidia signed a deal to license technology and acquire talent from Poolside, buy a stake in data labeling company Merkur, and potentially invest in Perplexity at a potentially perplexing $30 billion valuation.

更不用说其他对 NeoClouds 的股权投资、土地和电力交易,以及数据中心的后备支持。

That's to say nothing of other equity investments in NeoClouds, land and power deals, and data center backstops.

一些人已经开始将 Nvidia 视为计算领域的中央银行,它既支撑着 AI 经济,也设定着关键资源的价格。

Some have started to conceptualize Nvidia as the central bank of compute, standing behind the AI economy as well as setting the price of the key resource.

The Information 的 Martin Peers 将 Jensen 的做法与 John Malone 进行了比较,John Malone 在1980年代建立了庞大的有线电视帝国。

Martin Peers of The Information compared Jensen's approach to that of John Malone, who built up a giant cable TV empire in the 1980s.

另一个粗略的比较点可能是 Google 在2010年代中期通过其控股公司 Alphabet 的做法。当时 Google 重组了公司,创建了 Other Bets 部门,用来容纳对 Waymo 的登月式投资以及 Google Ventures。

Another rough comparison point might be Google's approach with their holding company Alphabet in the mid-2010s, when Google restructured the company and created their Other Bets division to house moonshot investments in Waymo as well as Google Ventures.

基本想法是将他们在互联网广告中获得的巨额收益再投资到更广泛的科技生态系统中。

The basic idea was to reinvest their massive earnings from internet advertising into the broader tech ecosystem.

当然,随后其中很多投资现在都得到了回报。

And of course, subsequently, a lot of those bets have now paid off.

Nvidia 的做法显然不同。

Nvidia's approach is obviously different.

Nvidia 没有分散投资于下一代科技,而是专注于 AI 生态系统。

Rather than fanning out across next-generation tech, Nvidia is sticking to the AI ecosystem.

尽管如此,他们仍在建立一组强大的其他投资组合。

Still, they are building up a formidable portfolio of other bets.

在3月的最近一次财报电话会议上,他们报告称已在私营公司投资423亿美元,而本周晚些时候他们发布财报时,这个数字肯定会更高。

During their last earnings call in March, they reported $42.3 billion invested in private companies, and the number will certainly be higher when they report later this week.

在咱们急着喊什么循环交易之前,有一点很重要得认识到:这部分反映出英伟达在自身业务再投资方面已经差不多到顶了。

And before we scream up and down shouting circular dealmaking, one thing that's important to recognize is that part of this is a reflection that Nvidia is more or less tapped out when it comes to reinvesting in their own business.

英伟达自己不运营晶圆厂,所以现阶段芯片收入的增长受限于其供应商网络的产能限制。

Nvidia does not operate their own fabs, so at this stage, growing chip revenue is limited by constraints across their network of suppliers.

换句话说,就是他们自己控制不了的瓶颈。

In other words, constraints they can't control.

通过把投资转向外部,英伟达支撑着整个AI生态,这反过来又能确保他们的收入在未来很多年里保持强劲。

By turning their investments outwards, Nvidia supports the entire AI economy, and that in turn ensures that their revenues can stay strong for years to come.

至少看起来是这么个目标,同时他们在其他押注上野心也变得更大了。

At least that appears to be the goal, as they get even more ambitious with their other bets.

不过芯片仍然是重头戏。在地球另一边,台湾检方起诉了9个人,涉及芯片走私,其中包括一名英伟达经理。

Still, chips remain the big game, and over on the other side of the world, Taiwanese prosecutors have charged 9 people for chip smuggling, including one Nvidia manager.

周一,台湾方面宣布了9项起诉,涉及一个将尖端Blackwell 300系统走私到中国的计划。

On Monday, the Taiwanese announced 9 indictments in relation to a scheme to smuggle cutting-edge Blackwell 300 systems into China.

除了被指认为英伟达经销业务经理的那个人之外,另外两人是英伟达合作伙伴SuperMicro的员工。

In addition to the person who was identified as a manager in Nvidia's distribution business, 2 others worked for Nvidia partner Supermicro.

今年早些时候,SuperMicro的一位联合创始人就因另一起走私案被起诉。

Earlier this year, a Supermicro co-founder was charged in relation to another smuggling incident.

英伟达和SuperMicro都表示,问题仅限于少数违规员工,他们正在配合当局调查。

Both Nvidia and Supermicro have indicated the problems are contained to a few rogue employees and that they're working with authorities.

让你了解下这事的规模:该团伙据称订购了130台装有Blackwell 300的SuperMicro服务器。检方称有74台已交付给中国买家,另外56台被台湾官员截获。

To give you a sense of the scale of this incident, the group allegedly ordered 130 Supermicro servers containing Blackwell 300s, with prosecutors claiming that 74 were delivered to buyers in China, while another shipment of 56 were stopped by Taiwanese officials.

总共算下来也就不到一万块芯片,这虽然不算少,但绝对不足以搭建一个前沿训练集群。

In total, you're talking about less than 10,000 chips, which is not nothing, but certainly not enough to build a frontier training cluster.

最后今天,我们来看看一家超大规模云服务商的员工使用AI的情况。

Lastly today, an interesting peek under the hood around how much AI one hyperscaler's employees are using.

Business Insider拿到了一份微软内部表格,员工在其中自愿填报关键指标,包括薪资、奖金和AI使用量。

Business Insider got hold of an internal spreadsheet where Microsoft employees self-report key metrics including salary, bonuses, and AI usage.

这不是一个关于刷token量的故事,因为BI发现微软内部AI使用量与薪酬或晋升之间没有相关性。

Now, this is not a story about token maxing, as BI found no correlation between token burn and financial compensation or promotions at Microsoft.

无论是跨部门还是部门内部,AI使用情况都极不均匀。

Both across and within different departments, AI usage was extremely jagged.

在Azure部门,月度AI花费的范围从1美元到7500美元。

In the Azure department, the range of monthly AI spend was $1 to $7,500.

在云与AI部门,最高端达到了15000美元。

In the Cloud and AI division, the top end went all the way up to $15,000.

而在微软客户与合作伙伴解决方案部门,至少有一个人花费了28000美元。

And in Microsoft Customer and Partner Solutions, there was at least one person who ran up a bill of $28,000.

所有部门的中位数AI使用量则更集中一些。

The median AI usage across all the different departments was a little more clustered together.

8个部门中有7个的中位数AI使用量在150到500美元之间。

7 of the 8 had median AI usage of right around $150 to $500.

Core AI是个例外,他们中位数AI使用量是975美元。

With Core AI being the outlier where their median AI usage was $975.

重要的是,这些全都是自愿填报的数据。

Now, importantly, this is all voluntary self-reporting.

这个表格是为了让员工自愿分享薪酬数据,以促进薪酬透明度而维护的。

The spreadsheet is maintained to allow staff to voluntarily share compensation figures in an attempt to promote pay transparency.

微软总共22.3万名员工中,只有极少数人贡献了这份数据。

Only a tiny sliver of Microsoft's 223,000 employees overall contribute to this chart.

实际上只有600名美国员工,其中仅350人报告了AI使用情况。

Just 600 US employees, in fact, only 350 of whom reported their AI usage.

不过,这也能让你大概了解像微软这样的公司内部一些比较突出的AI用户在使用规模上是什么情况,我觉得这正好为我们今天的主节目做个引子,现在就开始。

Still, it gives you a sense of the type of magnitude you're seeing among perhaps the more prominent AI users inside a company like Microsoft, which I think is actually the perfect segue into our main episode, which we will begin now.
M1
M113:38

欢迎回到AI每日简报。

Welcome back to the AI Daily Brief.

现在AI叙事方面一件最好的事,至少在我看来,是我们开始看到人们不再无条件接受“AI显然会摧毁工作”这种前提了。

One of the best things that's happening right now when it comes to AI narratives, at least, is that we're starting to see a shift away from the accepted without question kind of premise that AI is obviously going to be job destroying.

常听节目的听众会知道,我的立场是很明确的。

Now, regular listeners will know my position on this is pretty clear.

我绝不是盲目乐观,谈到两种完全不同的工作范式之间的转变时,潜在挑战的规模非常大。

I am not Pollyannish at all about the potential scale of challenge when it comes to the transition between 2 totally different work paradigms.

你不可避免地会看到某些类型的岗位,这种新技术会让它们变得不再必要,这会给个人带来真正的冲击,而社会最好做好准备去支持那些受到影响的人。

You are inevitably going to see certain types of roles that this new category of technology will obviate the need for that will cause real personal disruption that society would do well to be ready to support the people affected.

但是,认为会有某种剧烈而快速的就业末日,这个想法从来都不准确。

The idea, however, that there was going to be some radical and rapid jobs apocalypse was never accurate.

关于这一点,Sam Altman 越来越不遗余力地——不只是改变了他的立场,还解释说他相信自己是错的,并试图解释他为什么认为自己错了。

On this point, Sam Altman has increasingly gone out of his way to not just have changed his position, but to explain that he believes that he was wrong and to try to explain why he thinks he was wrong.

在最近的一个播客采访中,他说:我原以为当我们到了 GPT-4 的时候,也就是 2023 年,在那之后很快就会有更大的颠覆,软件公司会立刻被抢走,但实际并没有。

In a recent podcast interview, he said, I thought when we got to GPT-4, which was back in 2023, that very quickly after that, there was going to be much more disruption, software businesses up for grabs right away than it turned out to be.

我觉得我在几件事上错了,但其中一个就是速度。

I think I was wrong about a few things, but one in terms of the speed.

经济本身就有巨大的惯性。

The economy just has so much inertia.

人们继续做同样的事,从同一家公司买东西,希望用同样的方式使用他们的工具。

People keep doing the same things, buying from the same company, wanting to use their tools the same way.

其实我认为这在很多方面是积极的,它会让这个巨大的转型进行得更平滑、更缓慢。

I think this is actually a positive in many ways, and it's going to make this big transition go smoother and slower.

我对此很感激,但这意味着我们所有人都对时间表太乐观了。

I'm grateful for it, but it means we've all been too ambitious on timelines.

即使有这项令人难以置信的技术,社会和经济也会更慢地适应。

Even with this incredible technology, society and the economy will adapt more slowly.

换句话说,就像我喜欢思考的那样,当你有企业的时候,谁还需要暂停 AI 运动呢?

In other words, as I like to think about it, who needs a pause AI movement when you've got corporations?

但有趣的是,Sam 实际上走得更远,他认识到,虽然是的,制度惯性可能是推广放缓的最大驱动因素,但这种惯性也会在个人层面上发生,即使是对那些非常高级的用户也一样。

What's interesting though is that Sam actually went farther and recognized that while yes, institutional inertia is perhaps the biggest driver of the slowdown in the rollout, that, that inertia happens on an individual level, even with really advanced users as well.

在同一次采访中,他说:我觉得自己最心理矛盾的地方是,20 年来我一直用同样的方式使用电脑。

In that same interview, he said, the thing that feels most psychologically inconsistent about myself is that I have for 20 years been using computers the same way.

现在我有了一个神奇的东西叫 Codex。

I now have a magic thing called Codex.

你也有。

So do you.

每个人都有了。

So does everybody.

这意味着我应该完全以不同的方式使用我的电脑。

That means I should completely be using my computer in a different way.

我不应该再点来点去,从一个聊天应用复制粘贴到另一个。

I should not be clicking around, pasting from one messaging app to another.

我不应该再盲目地滚动邮件,试图找出哪一封打开起来最不痛苦。

I should not be scrolling mindlessly through my emails trying to figure out which one is the least painful to open.

我不应该再像以前那样,保持待办事项列表,做那些机械的电脑任务。

I should not be keeping a to-do list and doing these rote computer tasks the same way I have for so long.

然而,我脑子里有一种编码,认为做这些事就是工作和高效的表现。

And yet there's something in my mind that is encoded that doing this kind of stuff is what it means to work and be productive.

如果你问我,我绝不会说我喜欢那样做。

If you asked me, I would never say I like doing it that way.

事实上,我会说相反的话,而且我是认真的。

In fact, I'd say the opposite and I'd mean it.

但根据显示出来的偏好,我现在有更好的方法,却仍然用老方法。

But by revealed preference, I have a better way to do it now, and I still do it the old way.

这完全说不通,除非我其实偷偷喜欢它,或者觉得那样挺好。

It makes no sense other than I must secretly like it or feel good about it.

但我认为甚至也不是什么隐秘的满足感。

I don't think it's even some secret satisfaction, though.

我只是觉得我们都会陷在固有的做事模式里。

I just think we all get stuck in the patterns of how we've always done things.

我们建立了智力的肌肉记忆,而一想到要放弃它去建立一种新的、更不舒服的肌肉记忆,往往就显得很累人——因为我们确切知道用老方法做需要多长时间。

We build intellectual mind muscle memory, and the thought of undoing that to go build a new type of mind muscle that is much less comfortable often appears exhausting when we know exactly how long it will take us to do it the old way.

而且我们可以现在就做完,然后继续去做我们真正想做的事,这就是为什么花时间去看看那些已经打破自己肌肉记忆的人是怎么做事的,如此重要。

And we could just get that done right now and move on to whatever it is we actually wanna be doing, which is why it's so important to spend time looking at how people who have broken out of their mind muscle memory are doing things.

有意思的是,几周前,OpenAI发布了一些关于这个问题的研究。

And interestingly, a couple of weeks ago, OpenAI published some research about exactly that.

我不知道他们为什么没大肆宣传这篇文章,但如果我错过了某些数据,关于前沿公司怎么跟别人做法不一样,那东西宣传得远远不够。

Now, I don't know why they weren't screaming about this article from the rafters, but if I missed something that's putting numbers around how frontier firms are doing things differently than others, You know, that thing was not promoted nearly widely enough.

总之,我是通过A16Z把它当成每周图表重新发布时才注意到这些数据的。

In any case, I only noticed the data when A16Z reposted it as part of their charts of the week.

吸引我和其他人注意的那张图,是按企业职位分类显示Codex用户增长的图表。

The chart that grabbed mine and many others' attention was this one showing Codex user growth by enterprise job title.

这张图以今年2月初为基准,在那段时间里,虽然所有职位都在增长,但非技术类职位的增速最快。

It was indexed back to the beginning of February of this year, and in that time, while every role has gone up, it is the non-technical roles that have grown the fastest.

当然,一部分原因是他们的起点较低,但举个例来说,工程和技术人员的用量增长了5倍,而财务和会计领域的用量增长了20倍。

Now, of course, part of this is because they had a lower starting point, But to give you some examples, while use among engineering and technical practitioners is up 5x in that time, use in finance and accounting is up 20x.

市场和传播领域则是26倍。

In marketing and communications, it's 26x.

人事招聘,还有销售和客户管理,增长了41倍。

In people and recruiting, and separately sales and account management, it's up 41x.

在法务领域,Codex的使用量从2月份开始增长了108倍。

And in legal, Codex usage is 108x from where it was back in February.

关于律师这块,我半开玩笑地说,Spellbook的Scott Stevenson引用了Jack Newton的话说,LLM对律师来说就像电子表格对会计一样。

Joking but not really joking about the lawyer side of this, Spellbook's Scott Stevenson quoted Jack Newton saying, LLMs are for lawyers what spreadsheets were for accountants.

但这并不是这里唯一的数据。

But that was hardly the only data that was available here.

这篇文章的主要信号,题目叫《企业信号:前沿公司在做什么不同的事》,分为两部分。

The big throughline signal in this piece, which was called Enterprise Signals: What Frontier Firms Are Doing Differently, was 2 parts.

第一,更多工作被委派给了代理。

First, more work is being delegated to agents.

第二,正因为如此,产生了一种复合效应,领先的公司正在拉大与落后公司之间的距离。

And second, because of that, there is a compounding effect where the firms that are farthest along are getting farther away from the firms that are behind.

换句话说,代理式使用正在复合它们的领先优势。

In other words, agentic use compounds their lead.

重要的是,这不是模型本身的问题。

And importantly, this is not a model question.

这关乎OpenAI所说的,如何让这些模型发挥作用,从辅助转向委派,给代理提供上下文和工具来完成复杂任务,并把代理式AI的应用范围扩展到软件开发之外。

It's about, as OpenAI puts it, how they put those models to work, moving from assistance to delegation, giving agents the context and tools to complete complex tasks, and accelerating agentic AI beyond software development.

那我们来聊几个最有趣的数据。

So let's talk about a few of the most interesting numbers.

一年前的这个时候,也就是2025年8月,GPT-5发布时,企业以任何方式与ChatGPT生态系统交互所产生的输出token中,ChatGPT使用和代理式使用的比例基本上全是ChatGPT的token,几乎没有代理式的。

A year ago at this time, in August of '25, when GPT-5 was announced, the balance between ChatGPT usage and agentic usage as measured by the output tokens produced by the enterprise interacting with the ChatGPT ecosystem in any way, was basically 100% ChatGPT tokens and not agentic tokens.

在10月到2月之间,我们开始看到实际代理时代的初步迹象。

Between October and February, the first glimpses of the actual agentic era started to be seen.

10月,Codex正式全面开放,企业输出token中只有低个位数百分比属于代理式类别。

Codex becomes generally available in October, and a low single-digit percentage of enterprise output tokens are now in that agentic category.

12月,GPT-5.2发布,我们看到了一点提升,延续到2月初。

In December, GPT-5.2 comes out, And we see a bit of a bump that extends up into the beginning of February.

2月,当Codex应用在macOS上推出时,按企业输出token衡量,ChatGPT和代理式使用的比例是87%对13%。

In February, when the Codex app launches for macOS, the percentage balance between ChatGPT and agentic usage, again, as measured by enterprise output tokens, was 87% ChatGPT to 13% agentic.

但之后,代理式使用案例开始起飞。

But then from there, the agentic use cases just take off.

到了3月,Codex在Windows上发布,GPT-5.4也出来了。

We get to March, Codex is released for Windows, and GPT-5.4 comes out.

现在我们到了73%的ChatGPT,27%的代理式。

And we're now at 73% ChatGPT, 27% agentic.

4月底,仅仅一个月后,GPT-5.5发布,我们看到了翻转,代理式输出token突然占到了53%。

Towards the end of April, just a month later, GPT-5.5 comes out and we get the flippening where all of a sudden agentic output tokens are representing 53%.

到了6月,这个数据集结束的时候,我们降到36%,而代理式token上升到64%。

And by June, when this dataset ends, we're down to 36% while agentic tokens are up to 64%.

请注意,这并不意味着突然间64%的人坐下来使用OpenAI产品工作时,都在和代理打交道。

Now, keep in mind, this does not mean that all of a sudden 64% of the times that someone sits down to use an OpenAI product at work, they're doing something with agents.

相反,这意味着如果你用所有提示产生的输出token数量来算,也就是把输出token当作工作量或工作量的代理指标,那么其中大部分,在这些统计被记录时几乎有三分之二,现在都是代理式工作。

Instead, what this means is that if you use the amount of output tokens that the aggregate set of prompts lead to, in other words, if you use output tokens as a proxy for the amount or volume of work being done, the preponderance of it, almost 2/3 by the time these statistics were captured, is now agentic work.

所以第一点是,代理型AI的使用在增加;第二点是,领先的AI用户,也就是那些使用代理最多、最擅长的企业,与普通企业用户之间的差距正在扩大。

So point one was that agentic use is up, and point 2 was that the gap between the leading AI users i.e., the ones using agents the most and the best, and the general enterprise users was getting wider.

OpenAI将前沿企业定义为每月使用量排名前10%的企业,衡量标准是每个活跃用户生成的输出token数;而普通企业则位于第45到第55百分位之间。

OpenAI defines frontier firms as those in the top 10% of usage in a month as measured by output tokens per active user, with the average firm to be between the 45th and 55th percentile.

最好用数字来说明:从2025年4月左右开始,整个2025年,普通企业与前沿企业之间每个活跃用户的输出token差距只有大约2倍——也就是说,前沿企业的普通用户使用的输出token大约是普通企业普通用户的两倍。

And providing the best numerical advice, you can see that going back to about April of 2025, throughout the year of 2025, the gap in output tokens per active user between the typical firm and the frontier firm was only about 2x, as in the average user at a frontier firm used about twice as many output tokens as an average user at an average firm.

这个差距在10月左右开始扩大,到1月时达到了2.6倍。

The gap started to widen around October, and in January stood at 2.6x.

现在这个差距彻底爆发了,普通企业与前沿企业之间的差距已经达到了8.3倍。

The gap now has absolutely exploded, with the distance between the typical firm and the frontier firm now at 8.3x.

总体来看,虽然普通企业使用的token数量比一年半前增加了一倍左右,但前沿企业使用的token数量是一年半前的17倍。

Overall, while the average firm is using about twice as many tokens as they did a year and a half ago, frontier firms are using 17 times as many tokens as they were a year and a half ago.

当然,部分原因是他们更多地使用了代理,但另一部分原因是他们更擅长使用代理。

Now, part of this is, of course, that they're just using agents more, but part of it is that they're also better at using agents.

为了衡量这一点,OpenAI统计了每周活跃用户中使用插件或技能的人数。

To measure this, OpenAI looked at the number of weekly active users who use either plugins or skills.

插件当然是一组能力集,可以连接到其他应用或数据;而技能则是可重复使用的指令,有助于处理常见的工作流程。

Plugins are, of course, capability sets that can connect to other applications or data, whereas skills are reusable instructions that can help with common workflows.

在普通企业中,大约9%的每周活跃用户使用插件,只有大约3%使用技能。

At typical firms, about 9% of weekly active users are using plugins, and only about 3% are using skills.

而在前沿企业,也就是排名前10%的企业中,19%使用技能,21%使用插件——这并不意味着这些前沿企业的使用率已经见顶。

At frontier firms, which is again the top 10% of enterprises, 19% are using skills and 21% are using plugins, which is not to say that those frontier firms have topped out in terms of their usage.

作为对比,目前OpenAI内部93%的员工使用技能,95%使用插件。

By way of comparison, at OpenAI right now, 93% of their employees are using skills and 95% are using plugins.

但回到前沿企业与普通企业员工之间的差距:使用插件的人数差距是2.3倍,使用技能的人数差距超过6倍。

But still, coming back to the gap between employees at Frontier and typical firms, you're talking about 2 and a third times as many people using plugins and more than 6 times as many people using skills.

至于他们把这些技术用在哪些工作上——正如我们节目开头的图表所显示的,代理型AI用例增长最快的是来自软件和工程之外的知识工作者。

And in terms of what work they're deploying this towards, as we saw from that chart that kicked off this show, the fastest growth in these agentic use cases is coming from knowledge workers that are outside software and engineering functions.

OpenAI是这样诊断的:软件行业之所以先行一步,是有原因的。

OpenAI diagnoses it like this: software, they said, moved first for a reason.

代码库为代理提供了清晰的上下文,测试让输出更容易验证,编程方面的进步也有助于加速AI的研究和开发。

Codebases give agents clear context, tests make outputs easier to verify, and progress in coding helps accelerate AI research and development.

相比之下,他们认为通用知识工作方面的进展较慢,因为许多任务提供的上下文有限,难以明确指定,而且缺乏验证结果的清晰标准。

In contrast, they say, progress in general knowledge work has been slower because many tasks provide limited context, can be difficult to specify, and lack clear criteria for verifying the result.

但持续扩展的规模、强化学习的进步,以及针对性地提升在GDPVal等评估中的表现,正在让更多现实世界的任务、工具和工作环境被前沿模型所触及。

But continued scaling, advances in reinforcement learning, and targeted efforts to improve performance on evaluations such as GDPVal are bringing more real-world tasks, tools, and work environments within the reach of frontier models.

因此,自今年年初以来,代理型AI越来越多地在通用知识工作者中找到了产品市场契合点。

As a result, agentic AI has increasingly found product market fit with general knowledge workers since the beginning of the year.

你看,OpenAI正在努力让他们的模型和工具更好地服务于知识工作,这很好,但代理型AI在非软件工程知识工作者中增长的原因并不是OpenAI做了什么。

Look, it is great that OpenAI is working hard to have their models and harnesses work better for knowledge work, but stuff that OpenAI has done is not the reason that agentic use has grown among these non-software engineering knowledge workers.

代理型AI在通用知识工作者中增长的原因,是我们开始摸索出那些能让代理在我们自己的环境中真正发挥作用的模式。

The reason that agentic use has grown among general knowledge workers is that we've started to figure out the patterns that actually allow agents to thrive in our own contexts.

那么这些模式是什么呢?

So what are those patterns?

为此,我结合了OpenAI关于各部门在聊天和代理型AI中使用案例的研究,以及我们在AIDB和Superintelligent的自身经验,来描绘一下那些更高级的代理型AI在实际中是什么样子的。

From this, I'm combining the research from OpenAI about the types of use cases both in chat and agentic across departments, plus our own experience at both AIDB and at Superintelligent to provide a little bit of a picture of what that more advanced agentic use actually looks like in practice.

在OpenAI的图表中你可以看到,从聊天到代理型AI,工作类型发生了巨大转变。

You can see in OpenAI's chart that there's a massive shift in the type of work when you move between chat and agentic.

需要说明的是,这并不意味着用聊天完成的工作没有价值。

And to be clear, this doesn't mean that the type of work being done with chat is not valuable.

写作与沟通支持、知识检索与搜索——这些事情确实带来了很多价值。

Support for writing and communications, knowledge retrieval and search— those things do bring a lot of value.

但代理型AI将许多个人工作流程整合起来,转化为系统级的工作,影响范围不再局限于个人。

But agentic takes a lot of those individual workflows and instead moves them into systems-level work that can impact more than just the individual.

你几乎可以想象一个用例阶梯。

You can almost think about a use case ladder.

最底层是生成。

At the base is generation.

比如起草邮件、报告、Excel公式之类的。

Think drafting an email, a report, an Excel formula, things like that.

第二层是综合能力,能够将不同的数据来源整合起来,生成一个更完整的、基于这些信息的结果。

On the second rung is synthesis, being able to take disparate data sources and produce something more complete that is informed by them.

再往上走,就到了执行层,AI实际上被派去与现有系统交互,并在其中执行操作。这当然与下一层的维护紧密相关——在维护层,智能体不仅负责执行特定任务,还要长期维护一个系统。

Moving farther up the ladder, we have execution, where the AI is actually being tasked with interacting with existing systems and doing things within them, which of course is closely tied to the next level of maintenance, where agents are tasked not just with executing something specific, but maintaining a system over time.

比如说,在法律领域,人们依赖的上下文当然是合同、政策文件、判例,以及任何相关的对手方历史。

So for example, if you're looking in legal, The context that people are drawing on is, of course, things like contracts, policy documents, precedent, as well as any relevant counterparty history.

你会看到的智能体工作模式包括:比较条款、标记差异、起草修订标记、记录决策、监控承诺履行。

The types of agentic work patterns you're going to see are around things like comparing terms, flagging deviations, drafting redlines, recording decisions, monitoring commitments.

人类将继续负责关键条款的谈判、设定风险承受度、批准例外情况和最终措辞。

Humans will continue to negotiate the material terms, to set risk tolerance, to approve exceptions and final language.

所以基本上,这是一种分工:智能体负责覆盖面和协调工作。

So basically, you have a division of labor where the agent is handling coverage and coordination.

而人类仍然掌握风险判断和问责权。

While people still own risk judgment and accountability.

如果你回顾一下法律领域通过聊天或智能体完成的工作类别差异,就会发现:在聊天模式中,写作占法律工作的57%,紧随其后的是知识检索,占20.5%。

If you go back and look at the difference in the categories of work in the legal field that are being done via chat or agentic, writing makes up a full 57% of legal work in chat, followed closely by knowledge retrieval at 20.5%.

系统操作几乎可以忽略不计,只有0.2%。

System operation barely registers at 0.2%.

然后你转向智能体模式,写作下降到16.2%,知识检索下降到8.3%,而一大批新类别大幅涌现。

Now you move over into agentic and you've got writing down to 16.2%, knowledge retrieval down to 8.3%, And a whole bunch of new categories coming online in a huge way.

分类与提取跃升至4.6%。

Classification and extraction jumps to 4.6%.

系统操作跃升至17.7%。

System operations jumps to 17.7%.

工作流自动化跃升至7.7%。

Workflow automation jumps to 7.7%.

而编码,即实际构建应用程序,尽管这些人不是软件工程师,也跃升至32.9%。

And coding, actually building applications, even though these aren't the software engineers, jumps to 32.9%.

基本上,你在其他每个部门也能看到同样的模式。

And you see this pattern in basically every other department as well.

写作和知识检索,再加上教育和指导,仍然是聊天模式token的主要用途。

Writing and knowledge retrieval with a side of education and guidance remain the preponderance of chat tokens.

而智能体模式则进入了更深入的系统集成工作。

While Agentic gets into deeper systems integration work.

我认为,有一件事将把这一进程推向新高度,那就是多玩家和团队AI的出现,而不仅仅是单个AI。

Now, one thing that I think is going to supercharge this to the next level is the emergence of multiplayer and team AI as opposed to just individual AI.

我认为,尽管你看到这些前沿用户逐渐进入执行和系统维护的更高层次,但大多数智能体仍然在各自独立的孤岛中运作。

I think that even though you are seeing these frontier users move more into these higher-order tiers of execution and maintenance of systems, most of these agents are still operating within individual silos.

而我相信,下一代收益的很大一部分将来自不同团队的交汇点。

And my belief is that where a lot of the next generation of gains are going to come from is actually at the intersection of different teams.

归根结底,这些都不那么令人惊讶。

Ultimately, none of this is all that surprising.

一段时间以来,人们已经清楚,2026年是智能体真正落地的一年。

It has been clear for a while that 2026 was the year that agents became real.

但天哪,数字不会说谎。

But boy, the numbers do not lie.

尽管连前沿公司也才刚刚开始摸索,但他们似乎在加速前进,与普通公司之间的差距越来越大。这对于那些尚未大规模部署智能体应用的公司来说,应该是一记警钟。

And while even the frontier firms are still just barely beginning to figure it out, the fact that they seem to be racing ahead and putting more and more distance between themselves and the average firms should be a wake-up call for those who aren't deploying agentic uses at scale yet.
M1
M127:29

无论如何,OpenAI这次做得非常棒。

In any case, this is great stuff from OpenAI.

请继续这样做,并且下次更大力地推广。

Please keep doing this and please promote it more heavily next time.

当然,对于各位听众,一如既往地感谢你们的收听。

And of course, for you guys, appreciate you listening as always.

下次再见,peace。

Until next time, peace.
已剔除 12 处广告(点击展开查看)
M1
M11:06广告 · 已剔除

First of all, thank you to today's sponsors, KPMG, Blitzy, Harbor, and HyperAgent.

To get an ad-free version of the show, go to patreon.com/AIDailyBrief, or you can subscribe on Apple Podcasts.

And to learn more about sponsoring the show, send us a note at [email protected].

M1
M11:25广告 · 已剔除

While you're there, you can also find a link to our free webinar and hands-on lab, Agentic Loops for Knowledge Workers.

That is happening on Wednesday.

And even if you can't make it, if you register, we will send you the recording after.

And of course, if you were looking for a little bit more hands-on support, our next executive catch-up and executive agent leadership program is starting in a couple of weeks.

And you can find a link to that program from the top of aidailybrief.ai.

2
说话人 210:28广告 · 已剔除

A new study from KPMG and the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes.

Researchers studied more than 500 early career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs.

These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI.

Learn more about what separates AI amplifiers from everyone else at kpmg.com/us/aiamplifiers.

Blitzy deeply understands your codebase before it writes code.

Here's the first place that pays off: security in the age of AI.

Vulnerabilities don't live in isolation.

They live buried inside millions of lines of interconnected code where patching one thing quietly breaks 3 others.

That's why surface-level scans fail.

Blitzy starts from its knowledge graph of your entire application, identifies and surfaces CVEs across the full estate, proactively recommends patches, and can execute the PR.

Each fix is grounded in how your systems connect and validate so nothing new breaks, and the knowledge graph dynamically updates, keeping you ahead of an ever-accelerating threat landscape.

M1
M111:36广告 · 已剔除

One Blitzy customer resolved 21 active CVEs across 6 core microservices in 4 days.

2
说话人 211:41广告 · 已剔除

Zero compile errors, every validation scan clean, months of planned work fixed in less than a week.

Security remediation grounded in real architectural context at the speed of compute.

M1
M111:50广告 · 已剔除

Harden your codebase at blitzy.com.

2
说话人 211:52广告 · 已剔除

That's B-L-I-T-Z-Y dot com.

M1
M111:56广告 · 已剔除

If you listen to this show, you likely have a thesis.

2
说话人 211:58广告 · 已剔除

Maybe it's enterprise adoption, maybe it's compute, maybe it's a specific lab.

Harbor Capital's AI Lab Ecosystem ETFs let you express it via 5 actively managed ETFs.

Each seeking exposure to the ecosystem around one major lab: Anthropic, OpenAI, DeepMind, Meta, or SpaceX AI.

Your view of the AI race in ETF form.

Harbor Capital Advisors' AI Lab Ecosystem ETF suite gives investors a way to invest in the AI ecosystem they believe is best positioned for success.

Search Harbor AI Lab Ecosystems ETFs wherever you invest or follow @HarborCapital on X to learn more.

Visit harborcapital.com for a prospectus containing investment objectives, risks, fees, expenses, and other important information.

Read and consider it carefully before investing.

Risks include principal loss and artificial intelligence-related risks.

Harbor ETFs are distributed by Foresight Fund Services, LLC.

M1
M112:42广告 · 已剔除

Harbor is not affiliated with AI Daily Brief, and the funds are not affiliated with, sponsored by, or endorsed by any AI lab.

2
说话人 212:47广告 · 已剔除

This is a paid advertisement and not personalized investment advice.

Investing involves risk, including possible loss of principal.

This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.

New users get $1,000 in inference.

Forget local agents and chat workflows waiting on your laptop to be prompted.

HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses.

Marketing's agent turns competitor moves into landing pages.

Sales' agent enriches leads, drafts emails, and updates the CRM.

Ops' agent chases the paperwork and tracks the budget.

Every agent has access to shared context and follows your rules about scope and approvals.

It's time you add agents that feel like teammates.

Hire yours at HyperAgent, built by the team at Airtable.

Claim your $1,000 in inference at hyperagent.com/AIDailyBrief.

M1
M127:21广告 · 已剔除

This is probably where I should insert a shill for our training programs over at Superintelligent, but if you are a regular listener, you will already know that those are there and available for you.