← 返回任务列表

The AI Model Tier List

285 段 · 1 位说话人 · 原片 29:05
M1
M10:00

过去一提到先进的 AI 模型,大家关心的就只有一件事:谁领先。

It used to be that when it came to advanced AI models, all that anyone cared about was who was in the lead.

到底是 Anthropic 的模型,还是 OpenAI 的模型,或者 Google 的模型,是市面上最好的?

Was the model from Anthropic or OpenAI or Google the best one out there?

而且它是不是好到足以让我必须马上换过去?

And was it better enough that it meant that I needed to switch right away?

但现在,情况变得复杂多了。

These days, things are getting a lot more sophisticated.

不光是这些模型都已经达到某个关键门槛,能做到的事情比过去那些模型多得多;我们在个人、小团队和企业层面使用 AI 的量也大到了一定程度,由此进入了一个新阶段:人们和公司不再只考虑能力,也开始考虑模型效率,以及如何搭建完整的模型架构或者模型栈,让合适的任务找到合适的模型。

Not only have all of these models reached a certain critical threshold where they can just do a lot more than any of those models used to be able to do, the sheer volume at which we are using AI on both individual, small team, and enterprise levels has created a new moment where people and companies are thinking not only about capabilities, but also model efficiency, and how they put together complete model architectures or model stacks that can allow for the right tasks to find the right models.

今天我们会看几个方面,看看这个新阶段是如何体现在数据里的,同时也会分析一位很受欢迎的 AI YouTuber 做的 AI 模型分级榜。

Today we're looking at a few ways in which that new moment is showing up in the numbers, as well as analyzing a popular AI YouTuber's AI model tier list.

The AI Daily Brief 是一档每日更新的 podcast 和视频节目,关注 AI 领域最重要的新闻和讨论。

The AI Daily Brief is a daily podcast and video about the most important news and discussions in AI.
M1
M11:45

今天贯穿新闻头条和正片内容的一个副主题,是开放模型在整个模型栈里越来越重要的位置。

The subtheme that's going to run through both the headlines and the main episode today is about the growing place of open models in the overall model stack.

而这也正是我们第一条新闻背后的潜台词:Hugging Face 显然正在接触潜在收购方。

And that is certainly the subtext of our first story, which is Hugging Face apparently courting acquisition partners.

Business Insider 报道说,Hugging Face 正在寻求以一百三十亿美元的价格退出。

Business Insider reports that Hugging Face is seeking a $13 billion exit.

消息人士说,他们已经聘请了一家投资银行来接收报价,但目前还没有达成任何交易。

Sources say they've engaged an investment bank to field offers, but no deal has been reached as of yet.

这家公司上一轮融资还要追溯到 2023 年,当时估值是四十五亿美元。

The company's last round came all the way back in 2023 at a valuation of $4.5 billion.

那一轮融资的参与方包括 Google、Amazon、Nvidia、Intel 和 Salesforce。

That round saw participation from Google, Amazon, Nvidia, Intel, and Salesforce.

从那以后,这个平台的重要性当然只是在不断上升。

Since then, the platform has of course only grown in prominence.

它一开始是一个让开发者、研究人员和爱好者探索开放模型的地方;这些模型当然在很多不同方面都很有意思、也很重要,但当时并不真正属于专业用户或者商业用户会考虑的范围。

It started off as a place for developers and researchers and enthusiasts to explore open models that, while of course they were interesting and important in a variety of different ways, weren't really in the consideration set for professional or business type of users.

当然,在过去一年里,开放模型和前沿模型之间的差距已经缩小了;开放模型跨过了一些关键门槛,使它们能够被整合进严肃的商业工作流里。

Over the past year, of course, the gap between open models and frontier has closed, with open models crossing critical thresholds that allow them to be integrated into serious business workflows.

伴随着这个变化,Hugging Face 已经变成了一块关键基础设施,承载着最新发布的模型,而这些模型可能会大幅改变 AI 工作的完成方式。

In and around that change, Hugging Face has become a critical piece of infrastructure, hosting the latest model drops that can dramatically change how AI work gets done.

AI 评论者 Rohan Paul 写道,Hugging Face 现在托管了超过两百万个模型、一百五十万个数据集,以及一百五十万个 AI 应用。

AI commentator Rohan Paul wrote, Hugging Face now hosts more than 2 million models, 1.5 million datasets, and 1.5 million AI apps.

买家买到的,会是围绕这些资产的分发层,再加上帮助开发者找到一个 artifact、判断它是否安全、并把它投入生产环境的工作流。

A buyer would be acquiring the distribution layer around those assets, plus the workflow that helps developers find an artifact, judge whether it is safe, and put it into production.

随着开放模型越来越多,这个协调层就会变得越来越难以替代。

As open models multiply, that coordination layer becomes harder to replace.

Stripe 收购 OpenRouter,以及现在市场对 Hugging Face 的新兴趣,看起来都像是在押注:未来会是一个模型长期碎片化的世界。

Both Stripe's purchase of OpenRouter and this new interest in Hugging Face look like a bet on persistent model fragmentation as the future.

说实话,对 Nvidia 来说,这会很有道理。

To be honest, for Nvidia, it would make a lot of sense.

而且进一步说,如果 Nvidia 收购 Hugging Face,并且真的利用到那些数据,他们很容易在今年年底前推出一个能击败中国的 open-weight 模型。

And expands, if Nvidia acquires Hugging Face and actually taps into that data, they could easily drop an open-weight model that beats China before the end of the year.

当然,Nvidia 有一个没有受到足够关注的产品线,就是他们的 Nemotron 系列模型。

Certainly, it is the case that one of the underfollowed Nvidia products is their Nemotron series of models.

但如果你有在留意,就一定会感觉到,Nvidia 正在越来越认真地把开放模型视为竞争栈中的重要组成部分,这也会让这种交易变得相当有意思。

But if you are paying attention, you certainly get the sense that Nvidia is getting more and more serious about open models as a major piece of the competitive stack, which could make this type of deal pretty interesting.

进一步支撑这个想法的是,周四,独立科技记者 Eric Newcomer 报道说,Poolside 已经接受了一项实质上相当于 Nvidia 部分收购的交易。

Adding some further heft to that idea, on Thursday, independent tech journalist Eric Newcomer reported that Poolside had accepted what amounted to a partial acquisition deal from Nvidia.

Nvidia 将支付六十亿美元,获得一项非独家许可协议,以访问 Poolside 的技术;与此同时,还会以一百二十亿美元估值进行十亿美元的股权投资。

Nvidia will pay $6 billion for a non-exclusive licensing deal to access Poolside's technology alongside a billion-dollar equity investment at a $12 billion valuation.

Poolside 成立于 2023 年,由一位前 GitHub CTO 创办,目标是训练面向软件开发的开源 foundation models。

Poolside was founded in 2023 by a former GitHub CTO to train open-source foundation models geared towards software development.

作为交易的一部分,Nvidia 将从 Poolside 挖走一百多名工程师,让他们参与未来版本的——没错,正是——Nemotron 模型的开发。

As part of the deal, Nvidia will hire over 100 Poolside engineers away from the company to work on future iterations of their— yep, exactly— Nemotron models.

消息人士说,这基本上是 Poolside 工程团队的大部分成员;不过根据发给 Poolside 投资者的信,原话是:“这不是一次收购,也不是一次 acquihire。”

Sources said that this is the bulk of Poolside's engineering team, but according to the letter sent to Poolside investors, quote, this is not an acquisition and it is not an acquihire.

一个关键区别是,跟近几年其他大型 acquihire 交易不一样,创始人和核心领导层会继续留在 Poolside,并且会继续运营这家创业公司,重点放在一些还没有具体说明的研究项目上。

A key distinction is that unlike other huge acquihire deals in recent years, the founders and key leaders will remain at Poolside and will continue operating the startup with a focus on unspecified research projects.

消息人士说,计划是给 Nemotron 团队扩充人手,试图打造世界上最强大的开放模型,尤其是要和 DeepSeek、Moonshot 这样的中国实验室竞争。

Sources said the plan was to staff up the Nemotron team for an attempt to build the world's most powerful open models to rival Chinese labs like DeepSeek and Moonshot specifically.

在同一封写给股东的信里,Poolside 的创始人写道,这笔交易的目的,是创造一个未来,让 AGI,引用原话,“不是由少数人控制的封闭技术,而是由很多人在开放环境中共同构建的技术。”

In that same letter to shareholders, Poolside's founders wrote that the deal was intended to create a future where AGI, quote, would not be a closed technology controlled by a few, but one built by many out in the open.

哇,这还挺让人意外的。

Wow, this is kind of a shock.

据我理解,Nvidia 买下的是 Poolside 里面模型工厂的那一部分,而且很多员工、研究人员都收到了 Nvidia 的 offer。

From what I understand, Nvidia bought the model factory part of Poolside, and a lot of employees, researchers, got offers from Nvidia.

创始人继续留在 Poolside,这一点不太常见。

Founders staying at Poolside is unusual.

我在想,他们会不会就变成一家 neocloud 或者算力提供商,因为我在这里没看到任何关于 PIC,也就是 Poolside Infrastructure Company 的提法。

Wondering if they will just become a neocloud/compute provider since I don't see any mention of PIC, Poolside Infrastructure Company, here.

The Wall Street Journal 报道说,这笔交易是在最近几周匆忙谈成的,原因是一轮融资失败了。

The Wall Street Journal reports that the deal came together in a hurry over recent weeks as a result of a busted fundraising round.

Poolside 的创始人写信给股东说,去年年底,我们只有一个六周的窗口期,要筹集二十亿美元,用来支付一个四万块 GB300 组成的集群,这个集群原本一月份就要上线。

Poolside founders wrote to shareholders, at the end of last year, we had a 6-week window in which to raise $2 billion to pay for a 40,000 GB300 cluster coming online in January.

我们没能及时完成融资,所以失去了这个集群。

We didn't close it in time and we lost the cluster.

他们说,他们振作起来,重新投入工作,但很快意识到,最早到明年,他们的算力和资本就会耗尽。

They said they dusted themselves off and got back to work, but quickly realized that they would run out of compute and capital as soon as next year.

在他们看来,Nvidia 是继续推进打造前沿开放 coding model 这项工作的完美合作伙伴。

In their view, Nvidia was the perfect partner to carry on the work of building a frontier open coding model.

现在,为了进一步说明 Nvidia 正在更深入地进入模型训练这件事,上周 The Information 报道说,Nvidia 正在参与数据标注创业公司 Mercor 的最新一轮融资,而且值得注意的是,Nvidia 在他们最近两个 Nemotron 模型上做 reinforcement learning 时,用的就是 Mercor。

Now, just to add further heft to the idea that Nvidia is going deeper on model training, last week The Information reported that the company is taking part in the latest fundraising round for data labeling startup Mercor, and notably Nvidia used Mercor for reinforcement learning on their last 2 Nemotron models.

周日晚上,The Information 又补充报道称,Nvidia 也在参与 Perplexity 的新一轮融资。

On Sunday night, The Information added reporting that Nvidia is also participating in a new fundraising round for Perplexity.

这一轮融资会让 Perplexity 的估值达到三百亿美元,比他们差不多一年前上一轮融资时高出百分之五十。

The round would value Perplexity at $30 billion, a 50% markup from their last fundraising round almost a year ago.

消息人士说,Nvidia 一开始感兴趣的是一项 licensing deal,这样他们可以雇用一些员工,但现在他们接受的是一笔普通的股权投资。

Sources said that Nvidia was initially interested in a licensing deal that would allow them to hire some staff, but are settling for a normal equity investment.

我觉得 AI 圈里那些爱讨论的人,肯定会对这件事有很多话要说,所以我预计我们还会听到更多相关消息。但重点是,从所有这些交易来看,很明显,Nvidia 在认真考虑研究、人才、训练数据,以及应用层,因为他们看到了 open-source frontier 的重要性正在上升。

Now I think the chattering classes in the AI world are going to have a lot to say about this one, so I would expect we'll hear more about it, but the point is that it's very clear that across all of these deals Nvidia is putting serious consideration into research, talent, training data, and the app layer, as they look at the growing importance of the open-source frontier.

现在回到 Nvidia 的核心业务。

Now back to Nvidia's core business.

The Information 又报道称,Nvidia 已经开始通知客户,顶级 Grace Blackwell 和 Vera Rubin 芯片的价格最高将上涨百分之十七。

The Information again reports that Nvidia has begun notifying customers that the price for top-end Grace Blackwell and Vera Rubin chips will increase by as much as 17%.

这次调整适用于已经下单、计划明年交付的芯片。

The change applies to chips already ordered and set to be delivered next year.

The Information 写道,一整柜七十二块芯片的 Vera Rubin 预计价格会达到八百万美元,这会让建设一 gigawatt 算力的成本增加五十亿美元。

The price for a full 72-chip rack of Vera Rubins is expected to reach $8 million, adding $5 billion to the cost of building a gigawatt of compute, writes The Information.

目前还不清楚,购买 Nvidia 芯片的云服务商会自己消化一部分涨价,还是把成本转嫁给租用这些芯片的客户。

It isn't clear whether cloud providers that buy Nvidia chips will eat some of the price hikes or pass the cost to customers that rent the chips.

一位了解这次涨价情况的人说,云服务商几乎肯定需要把涨价转嫁给他们的客户。

One person with knowledge of the price hike said cloud providers will almost certainly need to pass on the increases to their customers.

Bloomberg 认为,这次涨价源于内存成本的不断飙升。

Bloomberg suggests the price increase stems from the spiraling costs of memory.

Nvidia 已经削减了一些 Vera Rubin 系统中将包含的内存容量,但这并没有让它们免受成本压力的影响。

Nvidia already trimmed the amount of memory to be included on some Vera Rubin systems, but that hasn't made them immune to cost pressures.

总体来看,这似乎进一步证实,各家公司都在为内存短缺做准备,而这种短缺会一直延续到明年很晚的时候,甚至更久。

Overall, it seems like further confirmation that companies are positioning for a memory shortage that will stretch deep into next year or even longer.

说到为了应对 AI-flation 而提前布局,Alibaba 通过一次创纪录的股票发行筹集了一百亿美元。

Now speaking of positioning to deal with AI-flation, Alibaba has raised $10 billion in a record-breaking share sale.

这次二级市场股票发行是在周五收盘时完成的,成为香港市场同类发行中规模最大的一次。

The secondary share sale was executed on Friday at the market close, completing the largest offering of its kind in the Hong Kong market.

周一早盘,股价一度下跌百分之十,这是自去年四月以来最大的盘中跌幅。

Shares were down as much as 10% on Monday morning, their largest intraday drop since April of last year.

这次股票发行表明,中国正在加快 AI 建设,并开始从所有可用来源调动资金。

The sale suggests that China is ramping up their AI buildout and starting to pull capital from every available source.

Union Bancaire Privée 的董事总经理 Vey-Sern Ling 指出,这不同于 Alibaba 过去对股票供应的严格管理,他问道,为什么不发债券呢?

Vey-Sern Ling, the managing director at Union Bancaire Privée, noted this is a departure from Alibaba's tight management of share supply, asking, why not bonds?

这让我觉得,他们在 AI 投资上可能需要比我们预期更多的资金,而且他们可能也在急着跑到其他公司前面。

It tells me that they may need more funds than we expect for AI investments, and also that they may be rushing to be ahead of other companies.

Big Short 投资者 Michael Burry 对 Alibaba 跟随美国科技巨头加入 AI CapEx 大战这件事,态度非常直言不讳。

Big Short investor Michael Burry was outspoken on Alibaba following the U.S. tech giants into the AI CapEx wars.

他在一篇 Substack 文章里写道,Alibaba 正在美国这场商品化、低成本 LLM 的血战中取得很大进展。

In a Substack post, he wrote, Alibaba is making serious inroads in the commodity low-cost LLM bloodbath in the U.S.

作为一股颠覆力量,这令人印象深刻,而且我相信这种势头会持续下去,但我不能支持增发股票。

It is impressive as a disruptive force, and I believe this will continue, but I cannot bless share issuances.

这对 Alibaba 来说又是一个新的范式,而它的投入资本回报率会继续下降。

This is a new paradigm again for Alibaba, and its return on invested capital will continue to fall.

在中国市场的其他地方,一场大规模 IPO 标志着 humanoid robot 炒作周期的开始。

Elsewhere in the Chinese markets, a massive IPO marked the beginning of the humanoid robot hype cycle.

Unitree Robotics 周三在 Shanghai Stock Exchange 上市,募资九亿美元,上市首日市值达到九十亿美元。

Unitree Robotics went public on Wednesday on the Shanghai Stock Exchange, raising $900 million and debuting with a market cap of $9 billion.

看起来这次发行定价明显偏低,因为股票在首个交易日暴涨了超过百分之四百六十。

It appears that the offering was severely underpriced, with the stock surging more than 460% on the first day of trading.

Bloomberg Intelligence 分析师 Ian Ma 表示,Unitree 的首日暴涨表明,市场对中国 embodied AI 板块有强烈兴趣。

Bloomberg Intelligence analyst Ian Ma said Unitree's debut surge signals strong appetite for China's embodied AI sector.

IPO 募集到的资金应该会加速 AI 的开发和商业化。

IPO proceeds should accelerate AI development and commercialization.

不过,The Information 也指出,中国 IPO 首日大涨其实并不算特别罕见。

Now, The Information does note that a huge day one pop isn't all that unusual for Chinese IPOs.

事实上,这已经是今年第四只首日涨幅超过百分之四百的 IPO 了。

In fact, this is now the fourth IPO this year that rose by more than 400% on day one.

一系列监管护栏会通过限制卖出,从机制上推高首日表现;不过中国市场还有一个特点,就是规模较小的公司上市后能获得大得多的回报,这和美国这边的情况有点不一样。今天最后,还有一个有点违背主流叙事的事情。

A range of regulatory guardrails help boost day one performance mechanically by limiting selling, but the Chinese market also features smaller companies going public with much larger returns, which is a little bit different than the scenario here in the U.S. Lastly today, a bit of a narrative violation.

至少 Dr. Dre 并不担心 AI 会接管音乐行业。

Dr. Dre at least isn't worried about AI taking over the music industry.

在 New York Times 的一篇人物报道中,这位说唱传奇和他的长期制作人 Jimmy Iovine 表示,他们认为 AI 对音乐是有好处的。

In a profile in the New York Times, the rap legend and his longtime producer Jimmy Iovine said that they believe that AI is good for music.

Iovine 说,我非常支持在音乐创作中使用 AI。

Said Iovine, I'm very pro-AI in music creation.

我完全看不到有什么坏处。

I don't see the downside at all.

当然会有一些烂音乐。

There will be some crappy music.

现在也有烂音乐啊。

There's crappy music now.

在录音室里,当有才华的人拥有 AI,他们会做出更好的唱片。

In the studio, when gifted people have AI, they're going to make better records.

与此同时,Iovine 也承认,AI 公司拥有世界历史上最糟糕的公共关系。

At the same time, Iovine acknowledged, the AI companies have the worst public relations in the history of the world.

我不了解整个世界历史,但我们就这么说吧。

I don't know the history of the world, but let's just put it this way.

它们的沟通能力太糟糕了。

They have terrible communication skills.

这就是为什么大家都这么激动、这么反对。

That's why everybody is all up in arms.

Dre 完全同意,还补充说,我不觉得它是一种威胁。

Dre agreed completely, adding, I don't see it as a threat.

我觉得只有那些在创作上有困难的人,才会把它看成威胁。

I think the only people that see it as a threat are the people who have trouble creating.

几天前,我和几个人聊过。

I had a discussion with a few people a few days ago.

他们反对 AI,我就说,好吧,你们听起来就像那种当年 drum machine 刚出来时会反对它的人,或者会反对 synthesizers 的人,对吧?

They were against AI, and I'm like, okay, you sound like the person who would have been against the drum machine when it came out, or synthesizers, right?

它是一个新的创作工具。

It's a new tool for creativity.

有些人害怕学习新东西。

Some people are afraid of learning new things.

但我是主动拥抱它。

I'm embracing it.

我都等不及想看看这件事接下来会怎么发展。

I can't wait to see what's going to happen with this.

Dre 说,他在工作里大量使用 AI,尤其是想看看模型会用什么不一样的方式来做,有点像音乐版的头脑风暴。

Dre said that he is extensively using AI in his work, particularly to see how the model might do it differently, kind of the musical equivalent of brainstorming.

Iovine 提到,Timbaland 也在使用这些工具,他还评论说,外面其实有很多偷偷用 AI 的制作人。

Iovine noted that Timbaland is also making use of the tools, commenting, there's a lot of closet AI producers out there.

Dre 补充说,这个说法挺准确的。

Dre added, that's a good way to put it.

他们确实在用。

They're using it.

他们只是不想承认而已。

They just don't want to admit it.

我觉得这是一次特别有意思的采访,尤其是因为对我来说,音乐一直都提供了一些最有力的理由,让我不用担心 AI 会侵犯创造力。

I think it's a super interesting interview, particularly because music to me has always provided some of the best reason to not be concerned about AI infringing on creativity.

如果你对这个讨论感兴趣,可以去找找我去年和 Rick Rubin 做的那期采访,我们在里面聊到,为什么他也把 AI 仅仅看作是下一代工具之一,是优秀音乐人会用来创作优秀音乐的工具。

If you're interested in that discussion, go dig up the interview I did with Rick Rubin from last year, where we get into why he as well views AI simply as a tool in the next generation of things that great musicians are going to use to create great music.

不过今天的头条新闻就先到这里。

For now though, that's going to do it for today's headlines.

接下来是本期的主要内容。

Next up, the main episode.
M1
M114:05

欢迎回到 AI Daily Brief。

Welcome back to the AI Daily Brief.

现在你在社交媒体上很常见的一类内容,就是 tier list,分级榜单。

One very common kind of content that you see on social media these days is the tier list.

其实早在社交媒体出现之前,人们就一直很喜欢各种榜单。

Even back since before social media became a thing, people have always loved lists.

这也是为什么会有 Billboard 榜单、Forbes 榜单,还有那么多类似的例子。

It's why there's a Billboard and a Forbes list and so many other examples.

但在互联网上,尤其是在短视频时代,我们真的特别特别喜欢把各种东西放进 tier list 里分级。

But on the internet, especially in the short-form video era, we really, really love putting things into tier lists.

换句话说,就是按照某种 A、B、C、D 这样的评分等级来给它们排名。

In other words, ranking them on a sort of grading A, B, C, D type of scale.

最顶上的是 S-tier,至于 S 代表什么,不同的人说法不一样,有人说是 supreme,有人说是 superior,也有人说它什么都不代表,就是 S-tier,你自然知道 S-tier 是什么意思。

With the very top being S-tier, which depending on who you ask stands for either supreme or superior or just nothing and just S-tier, and you just know what S-tier means.

上周末,AI 创业者兼内容创作者 Theo 做了一份 AI 模型 tier list,而这种东西嘛,一发出来就引发了大量讨论。

Over the weekend, AI entrepreneur and content creator Theo put together an AI model tier list, and as they do, it generated a ton of discussion.

在榜单最上面,他把 Fable 5 放在了 S-tier;GPT-5.6 Sol 在 A;Kimi K3、DeepSeek V4 Flash 和 GPT-5.6 Luna 在 B;Grok 4.6 和 MuseSpark 1.2 在 C。然后再往下,对,MuseSpark 1.2 下面,在 D-tier 的是 Opus 5、Sonnet 5、GLM-5 3、GPT-5.6 Terra,以及 Cursor/Anysphere 的 Composer 2.5。

At the top of the list, he had Fable 5 in S-tier, GPT-5.6 Sol was in A, Kimi K3, DeepSeek V4 Flash, and GPT-5.6 Luna were in B, Grok 4.6 and MuseSpark 1.2 were in C. Then below, yes, MuseSpark 1.2, down in D tier were Opus 5, Sonnet 5, GLM-5 3, GPT-5.6 Terra, and Cursor/Anysphere's Composer 2.5.

DeepSeek V4 Pro 在 F-tier,而再往下,在一个单独的、很惨的低于 F 的等级里,叫 Google tier,里面是 Gemini 3.7 Flash 和 Gemini 3.1 Pro。

DeepSeek V4 Pro was in F tier, and down in their own sad tier below F, called the Google tier, was Gemini 3.7 Flash and Gemini 3.1 Pro.

那么今天我们就来探讨一下模型 tier list 这个想法。

Now we're going to explore this idea of a model tier list today.

不只是因为争论这个很有意思,虽然确实有意思,更是因为现在正在发生的一件大事,就是我们的模型栈正在变得越来越多样化。

Not just because it's fun to debate, although it is, but because one of the main things that's happening right now is a diversification of our model stacks.

这当然已经发生在个人层面,而且也越来越多地发生在企业层面。很多组织不再只是选这个模型或者那个模型,而是在搭建一种基础设施,可以根据不同需求和不同任务,在不同模型之间切换。

This is certainly happening on an individual level, and increasingly it is happening on a business level where organizations aren't simply picking one model or another, but building an infrastructure that can move between models based on different needs and different tasks.

甚至已经有一类公司就是专门做这件事的,也就是 router 公司。

There is even a category of businesses that are made to do exactly this, the router companies.

其中最有名的 OpenRouter,刚刚被 Stripe 以七十亿美元收购。

The best known of which, OpenRouter, was just acquired by Stripe for $7 billion.

甚至主流媒体也开始注意到,AI 模型之战已经不再只是关于纯粹的最先进水平了,当然,他们的理解方式还是很不完整。

And even mainstream media is picking up on the idea that the AI model war is no longer just about the pure state of the art, although of course they're doing it in a very incomplete kind of way.

你这个周末可能已经在社交媒体上看到 Financial Times 的这张图到处在传。

You might have seen this chart from the Financial Times flying around social media this weekend.

这张图的标题是,Anthropic 最好的模型 Fable 5,销售表现有限。

The header of the chart is Anthropic's best model, Fable 5, has drawn limited sales.

而且它显示,在企业花在 Anthropic 上的支出里,Opus 4.8 仍然是遥遥领先、最占主导地位的模型。

And it shows that across business spend on Anthropic, Opus 4.8 remains by far the most dominant model.

过去几周,随着 Opus 5 上线,它的使用量也已经超过了 Fable 5。

In the last few weeks, as Opus 5 has come online, it has also outpaced Fable 5.

事实上,就目前来看,Sonnet 4.6 和 Fable 5 处在相当接近的水平。

In fact, at the moment, Sonnet 4.6 and Fable 5 are at pretty common levels.

那对有些人来说,这非常令人意外。

Now, for some folks, this is very surprising.

投资人 Dan Robinson 写道:“这让我挺惊讶的,也让我重新思考一些假设。

Investor Dan Robinson wrote, "This is pretty surprising to me and makes me rethink some assumptions.

难道这么多企业用例真的已经被 Opus 满足到饱和了吗?

Are so many enterprise use cases really saturated by Opus?

哪怕是更简单的任务,我也很难想象会有人不想要 frontier intelligence。”

I can't really imagine not wanting frontier intelligence even for simpler tasks."

他的评论区也反映了围绕这张图在 X 和其他地方流传的很多讨论。

Now, his comment section reflects a lot of the discourse about this chart that's flown around X and other places.

也就是说,很多人非常笃定地认为,企业总体上是在非常有意识地决定不买 Fable,因为它太贵了;但他们要么没有理解这些数据来自什么背景,要么其实并没有任何真正的企业 AI 实践经验。

Which is to say that it's confidently sure that businesses in general are making a very conscious decision not to buy Fable because it's too expensive without either A, understanding the context of where this data comes from, or B, having any real experience with what AI in the enterprise actually involves.

这些数据来自 Ramp AI Index,大约两周前由 Ramp 的首席经济学家 Ara Kharazian 分享出来。

This data comes from the Ramp AI Index and was shared by Ramp's lead economist, Ara Kharazian, about two weeks ago.

Ramp 团队很棒,他们做的 AI 经济分析也确实很好,对整个行业非常有价值。

Now, the Ramp team is great, and the work they do putting out economic analysis of AI is really good and incredibly valuable to the industry.

但这一次,在我看来很明显,他们的分析跑偏了。

But with this one, it was pretty clear to me that they had missed the analysis.

Ara 在介绍这张图的时候,加了一句总结:“一个强大到曾经被短暂禁用的模型,然而企业并不觉得它值这个价。”

When Ara introduced the chart, he added the summary statement, "A model so powerful it was briefly banned, and yet businesses don't think it's worth the price."

但问题是,我并不认为企业在有意识地判断 Fable 5 不值这个价,和这件事有多大关系。

Except I don't think that businesses making a conscious decision that Fable 5 isn't worth the price has very much to do with this at all.

这当然可能是其中一部分原因,但这次判断里完全漏掉的一点是,Fable 5 有一项三十天的数据保留政策。

It certainly might be a part of it, but one thing that was completely missed in the diagnosis was the fact that Fable 5 has a 30-day data retention policy.

这是这个模型在被政府关停后重新上线时,附带的临时安全措施的一部分。

It was part of the provisional safeguards that came with it when the model came back online after being shut down by the government.

单凭这一点,就足以让大量企业直接说,绝对不行。

That all on its own is enough for a huge number of enterprises to say absolutely not.

当那些没有这类数据保留政策的模型依然很强大时,为了升级到下一个模型而引发各种严重的信息安全和数据方面的担忧,根本说不过去。

There's just no way that causing all sorts of serious infosec and data concerns justifies upgrading to the next model when the models that don't have that data retention policy are still quite powerful.

如果你需要证据证明这确实是很大一部分原因,那就看看过去一周 OpenAI 多么积极地在宣传他们针对 frontier models 的零数据保留政策。

And if you need evidence that this is in fact a big part of this, just look at how aggressively in the past week OpenAI has been pushing their zero data retention policies for frontier models.

不过也要替 Ramp 的 Aaron 说句公道话,他后来确实又回来补充说,有很多员工回复说,他们不被允许使用 Claude,因为 Anthropic 被要求为了美国政府的安全检查,把 prompts 保留三十天。

Now, to Aaron from Ramp's credit, he actually came back later and said, a lot of replies from employees who say they aren't allowed to use Claude because Anthropic is required to retain prompts for 30 days for US government safety checks.

而且这里还不止这一点。

And that's not the only thing here.

正如 Simon Willison 指出的,这些数据不仅来自 Ramp,而 Ramp 本身是一家极度偏技术前沿的公司,和它互动的也基本都是同样非常偏技术前沿的公司;更具体地说,这些数据还来自一个 token 和支出管理产品,而使用这个产品的用户,本来就是想尽量降低成本。

As Simon Willison points out, this data not only comes from Ramp, which is an extremely tech-forward company that only other pretty extremely tech-forward companies are interacting with, but comes specifically from a token and spend management product that users are using to try to minimize costs.

Simon 写道,Ramp 的整体数据本来就存在选择偏差,而这组数据的选择偏差更严重。

Simon writes, Ramp data overall suffers from selection bias, and this data suffers from it even more so.

这是来自他们的 token 和支出管理产品,所以用户天然就更倾向于关注成本控制。

This is from their token and spend management product, so users are predisposed to focus on cost control.

Claude 对大多数任务来说,根本不是最具成本效益的选择。

Claude simply isn't cost-effective for most tasks.

这里也暴露出了创业圈的一个盲点:觉得企业总体上还没有采用一个才发布几个月的模型很令人震惊,这种想法其实忽略了大多数企业行动起来有多么缓慢,简直像冰川一样。

There's also the startup world blind spot showing through here, where the idea that it's shocking that enterprises in general haven't adopted a model that's just a few months old kind of misses the glacial pace at which most enterprises move.

虽然听起来可能很惊人,但我每天都会听到有人说,他们还在用 GPT-4.1,以及九个月前的其他模型,因为他们公司只给他们开放了这些模型。当然,这并不是说领先指标没有显示出企业确实正在变得更懂模型,也在搭建更完整的模型组合。

Shocking though it may be, I hear from people every single day who are still using GPT-4.1 and other models from 9 months ago because that's what their companies give them access to, which is not to say that the leading indicators don't suggest that enterprises are in fact getting more model fluent and building more complete model stacks.

比如这周,The Information 就报道了 AT&T,并介绍了他们试图用 open source models 来降低 AI 账单的做法。

This week, for example, The Information profiled AT&T and reported on their attempt to use open source models to cut down on their AI bills.

根据数据科学副总裁 Mark Austin 的说法,AT&T 的计划是在未来几年让对 OpenAI 和 Anthropic 的支出保持不变,然后逐步用 open models 来替代这部分使用。

AT&T's plan, according to Vice President of Data Science Mark Austin, is to hold spending with OpenAI and Anthropic flat over coming years and slowly supplant that use with open models.

这家公司大约有十万名员工,并且已经把 AI 嵌入到各个部门的工作流程里,范围从编程、财务分析,一直到 HR 和客户支持。

The company has around 100,000 staff and has embedded AI into workflows across every department, ranging from coding and financial analysis to HR and customer support.

AT&T 使用 AI 的绝大部分场景都是内部用途,而且公司声称,他们已经在用 open models 来处理员工百分之四十的 AI 查询。

The vast majority of AT&T's AI use is internal, and the company claims that they are already using open models to service 40% of employees' AI queries.

他们计划在未来几年,把这个比例逐步提高到百分之六十到百分之七十之间。

They plan to ratchet that percentage up to between 60% and 70% over the coming years.

Austin 说,他发现 open models 已经和 Anthropic 或 OpenAI 的上一代模型一样好,甚至更好,而那些上一代模型本来就已经足够胜任这些任务了。

Austin said that he's found that open models are just as good or better than previous generation models from Anthropic or OpenAI, which were already up to the task.

AT&T 仍然会把 frontier models 用在生成代码这类高级任务上,但对于更简单的用例,比如总结一个 PR,AT&T 现在用的是 open model。

AT&T still uses frontier models for advanced tasks like generating code, but for simpler use cases like summarizing a PR, AT&T is now using an open model.

Austin 说,我们预计这只会在未来继续变得更好。

Austin said, we expect that to just keep getting better going forward.

顺便说一句,值得注意的是,AT&T 使用的模型包括 Nvidia 的 Nemotron,也包括来自 Meta 和 Google 的 open models。

By the way, it's worth noting that the models that AT&T is using include Nvidia's Nemotron, as well as open models from Meta and Google.

AT&T 也在大量使用 model routers,来进一步节省成本。

AT&T is also making extensive use of model routers to drive further savings.

在 AI 编程方面,Austin 说,使用 router 已经让成本最多下降了百分之五十六,而质量只下降了百分之二。

For AI coding, Austin said that the use of a router has decreased costs by as much as 56%, while quality only fell 2%.

那说到和中国的竞争,虽然他们目前还没有使用任何中国模型,但他们正在分析把这些模型纳入组合会带来的风险。

Now, when it comes to competition with China, although at the moment they're not using any Chinese models, they are analyzing the risks of including them in the mix.

而有一件事可能会改变他们怎么看这个权衡,那就是 Austin 提到,转向 open models 之后,他们就可以把一部分服务部署在自己的数据中心里,而这些数据中心配备了 Nvidia 和 AMD 的芯片。

And one thing which could change how they view that equation is that Austin noted that switching to open models allowed them to host part of the service in their own data centers stocked with Nvidia and AMD chips.

这通常比从 cloud providers 那里租算力更便宜。

Which was often cheaper than renting compute from cloud providers.

重点是,思考不同模型彼此相比各自擅长什么,其实不只是一个虚荣的练习,而且会成为企业越来越常做的事情,哪怕 Financial Times 只是拿了一张图来强化他们早已有的叙事。

The point being that thinking about what different models are good for compared to one another is in fact more than a vanity exercise and will be something that enterprises do more of, even if the Financial Times is just grabbing a chart that they can use to reinforce their preexisting narratives.

现在回到 Theo 的榜单。

Now back to Theo's list.

Theo 不只是发布了这个榜单,他还配套发了一个视频。

Theo didn't only publish the list, he put out a companion video with it.

在我们进入具体观点之前,先挑这个视频里的几个重点来说说。我们先聊 GPT-5.1-Codex 被放在 A 级,以及 Claude 4.5 被放在 S 级。

And to give a few of the highlights from that before we get into the takes, let's talk first about GPT-5.1-Codex at A tier and Claude 4.5 at S tier.

关于 5.1-Codex,他说,它能做到一些我以前从来没想过 AI 居然能做到的事情。

On 5.1-Codex, he says, it's capable of things I never thought AI would ever be able to do.

你能用它做的事情简直不可思议。

It's unbelievable what you could do.

它是我大多数时候处理大多数事情时默认使用的模型,但它不是我用过的最聪明的模型。

It's my default model I use for most things most of the time, but it's not the most intelligent model I use.

当我要写那些我真的希望合并进去的重要代码时,它仍然不是我最喜欢的选择。

It's still not my favorite for writing important code I actually hope to merge.

那说到 Opus,他写道,Opus 知道的东西比我接触过的任何模型都多。

Now on Opus, he writes, Opus knows more than any model I've interacted with.

它的思考细致程度令人难以置信。

Unbelievably thoughtful.

他指出,它有时候还是会在一些事情上栽跟头,也会碰一些不该碰的东西;它会走一些没必要的捷径,偶尔也会忘记自己正在做什么。

He notes that it still trips over things and touches things that it shouldn't sometimes, that it takes unnecessary shortcuts and occasionally loses track of what it's doing.

他把 Opus 4.5 称为一个需要被驯服的天才,而 GPT-5.1 Codex 则是一个稍微笨一点、但会严格照你说的去做的机器人。

He calls Opus 4.5 a genius that needs to be tamed, whereas GPT-5.1 Codex is a slightly dumber robot that does exactly what you tell it.

有意思的是,虽然他把 Opus 4.5 评为唯一的 S 级,高于 GPT-5.1 Codex 的 A 级,但他说,如果我必须选一个,我会选 Codex。

Interestingly, even though he rated Opus 4.5 as the only S-tier above GPT-5.1 Codex's A-tier, he said, If I had to pick, I would pick Codex.

它是我默认使用的模型。

It's the model I default to.

如果没有 Codex,我会比没有 Opus 更想念它。

I would miss Codex more than Opus.

但 Opus 是最好的模型。

But Opus is the best model.

它是那个会写出我想合并的代码的模型。

It's the model that writes code I want to merge.

它是那个我会信任它去复查其他东西产出的工作的模型。

It's the model I trust to double-check work from other things.

它也是那个我会拿来讨论困难而深入的问题的模型,关于我想构建的东西,或者我想探索的领域。

It's the model I talk to about hard, deep things with things I want to build or areas I want to explore.

Opus 4.5 是下一代。

Opus 4.5 is the next generation.

GPT-5.1 Codex 是一个难以置信的模型,感觉像下一代,但它其实还是建立在上一代技术之上的。

GPT-5.1 Codex is an unbelievable model that feels next generation while still being built on the last generation of tech.

Opus 就像公司里那个天才,没人想跟他一起工作,但也没人想开掉他,因为他是那里最聪明的人。

Opus is that genius at the company that no one wants to work with, but no one wants to fire because they're the smartest person there.

如果你学会怎么跟他配合,那效果就太惊人了。

If you learn how to work with them, it's incredible.

那这件事特别有意思的地方在于,这跟我现在的体验非常相似。

Now, what's super interesting about this is that this is pretty similar to my experience right now.

任何一天、任何时刻,我都在这两个模型之间来回切换。

On any given day, at any given moment, I am jockeying between these two.

而且很多任务,我会同时在这两个模型里启动,来回试几轮之后,再决定我要重点用哪一个,通常会是 Opus,但也不总是。

And for many tasks, I initiate the task in both of them, and after a little bit of back and forth, decide which one I want to hone in on, which tends to be but is not always Opus.

不过在某种程度上,比 A 档和 S 档更有意思的,是他怎么给其他模型排名。

To some extent though, what's way more interesting than the A and S tier is how he ranks the other models.

因为其他模型并不是想跟 GPT-5.1 Codex 和 Opus 4.5 竞争。

Because the other models aren't trying to compete with GPT-5.1 Codex and Opus 4.5.

它们本来就是用来做不同事情的。

They are meant to do different things.

一个很好的例子是 nano。它在三个 GPT-5.1 模型里被描述为能力最弱的那个,但他把它排在理论上更均衡、处在中间位置的 mini 模型前面好几档。

A really great example of this is that nano, which is presented as the least capable of the three GPT-5.1 models, he has ranked a couple tiers ahead of the theoretically balanced middle mini model.

关于 B 档的 nano,他说它在那里不是因为写代码,而是因为用他的话说,它聪明、快速,而且擅长一堆杂七杂八的事情。

Of nano at the B tier, he says it's not there because of coding, but because it is, in his words, smart, fast, and good at a bunch of random stuff.

他说,单看调用次数,nano 可能是我用得最多的模型,不是因为我用它写代码,而是因为我围绕代码在做一堆其他事情。

Nano, he says, is probably my most used model by sheer calls to it, not because I'm doing code with it, but I'm doing a bunch of other stuff with my code.

他指的这些事情包括给代码分类、从 GitHub 拉取内容、读取内容之类的。

The things that he's referring to are things like categorizing code, pulling from GitHub, reading content.

基本上,对于那些不可逆的事情,他不信任它。

Basically, he doesn't trust it with things that aren't reversible.

与此同时,关于 mini,他写道,它处在一个特别奇怪的位置。

Now meanwhile, of mini, he writes, Fits in such a weird place.

很多这些指标,用 nano 可以便宜得多地拿到。

A lot of these numbers can be gotten for much cheaper with nano.

我宁愿在 high 设置下用 GPT-5.1,因为它会快得多,因为它生成的 token 更少。

I'd rather use GPT-5.1 on high because it's going to be much faster because it generates fewer tokens.

我从来没有为了任何事情选择过 Terra,而且如果很多人会选它,我会觉得很意外。

I have never chosen Terra for anything, and I would be surprised if many people do.

它在价格表上看起来合理,但对我来说,在现实里并不合理。

It makes sense on a pricing chart, but doesn't make sense in reality for me.

我觉得有意思的是,这反映出一件事:我们现在才刚刚进入这种模型栈和复杂模型架构的阶段,公司开始真正认真地思考,不同任务该用不同模型。所以我们开始看到模型设计里更有意识的取舍,公司竞争的也不只是最前沿水平,而是各种类型的性能效率,取决于它们希望人们拿自己的模型来做什么。

And I think what's interesting and what this reflects is that because we are just now coming into this model stack and complex model architecture type of moment where companies are realistically thinking about different models for different tasks, we're starting to get more conscientious tradeoffs in model design with companies actually competing not just at the state of the art, but for various types of performance efficiencies based on what they hope people will do with their model.

而随着这种转变发生,我觉得很可能会看到很多模型落在一种有点诡异的中间地带:它们既不是值得为高价买单的前沿 state-of-the-art 模型,同时也不是其他使用场景里最高效、最快的模型。

And as that transition happens, it's likely to me that you see a lot of models fall in kind of an uncanny middle, where they are neither frontier state-of-the-art models worth the premium that they cost, but are also not the most efficient or fast models for other types of use cases.

比如说,虽然 Theo 喜欢 Kimi K2,但他在视频里提醒大家,它并不像人们看起来以为的那么便宜。仅仅因为它是 open weight,并不代表它更便宜。事实上,考虑到 Sonnet 用更少的 token 能做更多事,在 extra high 设置下,它的成本还比 Sonnet 略高一点。

For example, although Theo likes Kimi K2, he reminded people in his video that it's not as cheap as people seem to think, that just because it's open weight doesn't mean it's cheaper, and that in fact, it costs slightly more than Sonnet on extra high, given that Sonnet does more with fewer tokens.

那至于其他人的回应,你会感觉很多人只是在给自己个人最喜欢的模型站台。

Now, in terms of other people's responses, you get the impression that a lot of folks are just shilling for their personal favorite.

而且考虑到这类分析很多都来自 X,你也可以想象,最常见的评论之一就是 Grok 4 Fast 应该排得更高。

And given that a lot of this analysis comes from X, as you might imagine, one of the most common commentaries was that Grok 4 Fast needed to be higher.

不过还有另一条分析思路来自 Noah。他说,说实话,自从 Sonnet 4.5 之后,我完全不理解现在人们到底是怎么形成对模型的看法的。

And yet one other strand of analysis came from Noah, who said, I have zero understanding of how people develop opinions about models now ever since Sonnet 4.5, to be honest.

它们都很棒啊,兄弟。

They're all fantastic, bro.

我希望我们能达到 AGI 的原因之一,就是我已经受够了这些模型鉴赏家不停地对模型之间那些微妙的、似是而非的差异发表高见。

One of the reasons I hope we reach AGI is that I'm tired of these model connoisseurs opining nonstop about subtle pseudo-differences between the models.

这开始变得像是在争论你最喜欢什么颜色,或者你最喜欢哪只 Pokémon。

This is starting to look like arguing about your favorite color or your favorite Pokémon.

几年之后,我们会回头笑这一切的。

In a few years, we will laugh about all this.

就连 Theo 在他的视频里也说,不管你用的是像 Fable 这种昂贵的顶级产品,还是像 DeepSeek V3 Flash 这种便宜得让人意外、效果也不错的东西,基本上都不太会出错。

Even Theo in his video says, Whether you're using expensive best-in-class stuff like Fable or surprisingly cheap and effective stuff like DeepSeek V3 Flash, it's kind of hard to go wrong.

我不认为现在用一个 tier list 是比较模型的最佳方式,因为可以比较的维度实在太多了。

I don't think a tier list is the best way to compare models right now because there's so many axes to compare on.

比如任务能力、成本、token 效率、速度等等。

Task capability, cost, token efficiency, speed, etc.

而且在很多方面,比 tier list 更有意思的,是这些模型组合起来之后的整体构成。

And in many ways, what becomes more interesting than the tier list is the combined composition.

而关于这种构成可能是什么样、以及它正在怎么变化,一个有意思的数据来源是 Vercel。

And an interesting source of data for what that composition might look like and how it's changing comes from Vercel.

Vercel 的 CEO Guillermo Rauch 最近展示了,在他们的 AI Gateway 产品上,open weight 模型和 closed weight 模型之间的使用比例是如何变化的。

Vercel CEO Guillermo Rauch recently showed how the balance of open weight models versus closed weight models had shifted on their AI Gateway product.

从六月二十四日的 token 使用占比来看,也就是几个月前,closed model token 大约占百分之七十二,而 open model token 大约占百分之二十八。

In terms of the share of tokens used on June 24th, a couple of months ago, closed model tokens represented around 72%, while open model tokens represented around 28%.

两个月之后,这个比例基本反过来了,closed 降到百分之三十八,open 上升到百分之六十二。

Two months later, that ratio has largely flipped with closed at 38% and open up to 62%.

当然,如果要说的话,Vercel Gateway 的数据会更明显地偏向开发者,因为它本来就被定位成一个面向开发者的 AI routing 工具。

Now, if anything, the Vercel Gateway data is going to be even more heavily biased towards developers as it is specifically positioned as an AI routing tool for developers.

但即便如此,在一些人看来,它仍然显示了风向可能正在往哪里吹。

And yet to some, it still shows where the winds might be blowing.

投资人 Gavin Baker 分享了这张图,并说,又有更多数据显示,open source AI 正在从 OpenAI 和 Anthropic 手里抢份额。

Investor Gavin Baker shared the chart and said, more data that open source AI is taking share from OpenAI and Anthropic.

考虑到 OpenAI 和 Anthropic 加起来在七月份还加速增长了,这一点非常令人印象深刻。

Super impressive given that the sum of OpenAI and Anthropic accelerated in July.

所以净 token 需求和 AI infra 需求的加速,比我们在 frontier 层看到的加速还要更明显。

So net token and AI infra demand accelerated even more than the acceleration we saw at the frontier.

按照 Gavin 的判断,open source AI 抢占份额,对 AI infrastructure 需求是利好,因为它会压低模型层的利润率,而且一个 open source token 生产出来所需要的算力,和一个 frontier token 一样多。

In Gavin's estimation, open source AI taking share is positive for AI infrastructure demand as it lowers margins at the model layer, and an open source token costs just as much compute to produce as a frontier token.

open source AI inference 没有任何东西是免费的。

Nothing about open source AI inference is free.

Gavin 预测说,在我看来,最有可能的最终状态是,closed frontier token 占经济价值的百分之六十到百分之九十,但只占 token 数量的百分之十五到百分之二十五。

Gavin predicts most likely end state, in my opinion, is that closed frontier tokens are 60 to 90% of economic value, but only 15 to 25% of tokens.

投资人 Daniel Newman 进一步强调这个结论,他补充说,Gavin 这里有两个非常重要的点。

Doubling down on the conclusion, investor Daniel Newman adds, two very important points here from Gavin.

第一,open-source models 的使用率和消耗量会高于 closed source。

One, open-source models will be the highest utilization and consumption over closed source.

第二,Frontier 仍然会拿走大部分经济收益,因为高端 intelligence 本来就能获得经济溢价。

Two, Frontier will still realize most of the economics because premium intelligence commands an economic premium.

MIT 的 Christian Catalini 认为,实际上它会分成三种不同的类型。

MIT's Christian Catalini thinks that it will actually split in three different ways.

第一类支出是 cheap generalist,也就是商品化的 open-weight models。

The first spend category is cheap generalist, which is the commodity open-weight models.

光谱的另一端是 state-of-the-art generalist,也就是来自 closed labs 的那些 token。

On the other end of the spectrum is the state-of-the-art generalist, i.e., the tokens from the closed labs.

然后在中间,也就是他新增的这个类别里,是他所说的 state-of-the-art specialists,也就是把 open weights 和企业专有上下文结合起来的模型。

And then in the middle, in the category he's adding, is what he calls the state-of-the-art specialists, those that combine open weights with enterprise proprietary context.

当然,Microsoft 的模型显然不是 open weights,但这似乎正是 Microsoft 通过他们的 Microsoft AI Foundry 产品在推进的一类思路。这个产品允许公司使用自己的数据,在他们的 MAI models 基础上进行 post-train 和构建。

Now, while obviously Microsoft's models are not open weights, this is the type of thesis that Microsoft seems to be pursuing with their Microsoft AI Foundry product, which allows companies to use their own data to post-train and build on the base of their MAI models.

不过我觉得,关于这件事在所有企业里到底会有多普遍,还有很多合理的讨论空间。

Although I think there's a lot of reasonable debate to be had around just how common that will be across all enterprises.

可以明确的是,我们已经不再生活在一个只关心“最好的模型是什么”的世界里了。

What's clear is that we don't live in a world anymore where the only thing that matters is what's the best model.

接下来,理解不同模型因为不同原因适合放在哪里,会变得越来越重要。

Increasingly, it will be important to understand where different models fit for different reasons.

即便是企业,没错,它们行动更慢,也更倾向于和单一生态系统保持更紧密的连接,它们大概也会想要建立一些环境,让小规模用户群可以测试各种不同方法,去寻找这些新型效率提升。

And even enterprises that, yes, move more slowly and stay a little bit more connected to a single ecosystem are probably going to want to set up environments where small groups of users can test various approaches to look for these new types of efficiencies.

但这是否意味着,我们再也不会因为最新的 state-of-the-art 模型发布而感到兴奋了呢?

But does this mean that the days of getting excited about the latest state-of-the-art model release are gone?

我们还得再看看。

We'll have to see.

现在有很多人在聊,说 GPT-5.1 很快就要来了,不过看起来 OpenAI 的 open-weight model 已经推迟到九月了,所以我们很快就有机会看到结果。

A lot of chatter that GPT-5.1 is coming shortly, although it appears that OpenAI's open-weight model has been delayed until September, so we'll have a chance to find out soon.

好了,今天的 AI Daily Brief 就到这里。

For now, that's going to do it for today's AI Daily Brief.

一如既往,感谢你的收听或者观看,我们下次再见,peace!

Appreciate you listening or watching as always, and until next time, peace!
已剔除 2 处广告(点击展开查看)
M1
M11:00广告 · 已剔除

All right, friends, quick announcements before we dive in.

First of all, thank you to today's sponsors, KPMG, Rackspace, Blitzy, and HyperAgent.

To get an ad-free version of the show, go to patreon.com/AIDailyBrief, or you can subscribe on Apple Podcasts.

If you want to learn more about sponsoring the show, send us a note at [email protected].

While you're at aidailybrief.ai, you can find out what else is going on in the community.

Superintelligence's next round of agent training programs for executives is kicking off at the beginning of September, and there's a link to register for those.

And this week on Wednesday, we have a free webinar and hands-on lab, Agentic Loops for Knowledge Workers, which will try to take a thing that has been very buzzy and hypey in developer circles and make it relevant for all of you non-developers.

Again, you can find all of that at aidailybrief.ai.

M1
M111:09广告 · 已剔除

A new study from KPMG and the University of Texas at Austin found that when people work with AI, similar skills don't guarantee similar outcomes.

Researchers studied more than 500 early-career professionals and found that the best performers consistently amplified the value of AI by guiding, evaluating, and refining its outputs.

These top performers, called AI amplifiers, weren't defined by what they knew alone, but by how they worked with AI.

Learn more about what separates AI amplifiers from everyone else at kpmg.com/us/ai.

One of the more interesting shifts in enterprise AI right now is how quickly the conversation is moving towards infrastructure and operations.

As AI moves into core workflows, regulated data environments, and agentic systems, enterprises need governed infrastructure and inference that can operate reliably day to day with clear operational accountability built in from the start.

As those systems scale, the operating model increasingly becomes part of the AI strategy itself.

Rackspace Technology is the operator of the full enterprise AI stack, from agents to infrastructure across private cloud, hybrid cloud, and edge environments.

Rackspace builds and operates governed AI infrastructure, inference, and production AI systems for organizations where sovereignty, compliance, and uptime are non-negotiable.

Their forward-deployed engineers stay embedded beyond deployment to help operationalize and run AI in live environments.

To learn more about where enterprise AI runs and outcomes scale, go to rackspace.com.

Every AI coding tool on the market does the same thing first: it starts writing code.

Blitzy does the opposite.

Before writing a single line, Blitzy spends days reverse engineering your entire codebase.

Thousands of agents ingest millions of lines, mapping every dependency, every undocumented constraint, every architectural decision made over the last decade.

The result is a dynamic knowledge graph that understands your software the way a principal engineer would after 30 years in the building.

Other tools guess at context with grep searches and Markdown files.

Blitzy never guesses.

It builds true understanding first, then delivers over 80% of entire software epics autonomously.

Validated, end-to-end tested, production-grade pull requests.

That's why Fortune 500 engineering teams trust Blitzy with the codebases that matter most.

See for yourself at blitzy.com.

That's B-L-I-T-Z-Y dot com.

This episode of the AI Daily Brief is brought to you by HyperAgent, where you run fleets of agents your team can manage together.

New users get $1,000 in inference.

Forget local agents and chat workflows waiting on your laptop to be prompted.

HyperAgent deploys always-on agents in the cloud, doing real work across the tools your team already uses.

Marketing's agent turns competitor moves into landing pages.

Sales' agent enriches leads, drafts emails, and updates the CRM.

Ops' agent chases the paperwork and tracks the budget.

Every agent has access to shared context and follows your rules about scope and approvals.

It's time you add agents that feel like teammates.

Hire yours at HyperAgent, built by the team at Airtable.

Claim your $1,000 in inference at hyperagent.com/AIDailyBrief.