← 返回任务列表

Neil Movva: Making AI 10x Cheaper

1130 段 · 3 位说话人 · 原片 83:34
M1
M10:00

我的工作就是把 token 的成本压到人类能做到的最低。

My job is to make the tokens as cheap as humanly possible.

我会做到这一点,而且我会利用我能用到的每一层技术栈来实现。

I will achieve that and I will do it through every layer in the stack available to me.

我特别喜欢供给侧的那些杠杆。

I love the supply-side levers.

我会用上每一块芯片,用上每一种电力来源,还会用上美国每一块适合干这事的地。

I will use every chip, I'll use every source of power, and I will use every piece of land in the United States that's suitable for this.

我们现在还是把 AI 当作一个咨询起来很贵的人,只有遇到难题才去问它。

We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question.

这不是看待智能的正确方式。

That's not the way to think about intelligence.

机器能思考这件事太不可思议了,我们应该想办法让尽可能多的人用上它。

It's incredible that the machine can think and we should try to get that into as many hands as, as many people as possible.
M2
M20:38

我觉得在这种对话刚开始的时候,最重要的是把话说明白,就是你到底在造什么,现在它能做什么。

I think it's important early in these conversations to just say the thing, like literally what you're building and what it does today.

所以先给我们一个大概的定位吧,简单描述一下你正在构建的系统是什么,以及它为什么应该存在。

So maybe just orient us there with a, with a brief description, like literally what the system is that you're building and why it should exist.
M1
M10:51

Sail Research 就是一个 token 工厂。

Sail Research is a token factory.

我们有一个 API,任何人都可以发请求过来,使用大型语言模型,开源的大语言模型,做他们想做的任何任务。

We have an API where anyone can send us requests where they can use large language models, open source large language models for any task they want.

我们会以市场上无人能敌的价格把这些 token 提供给对方。

We will serve those tokens to them at a price that is unbeatable in the market.

同时我们也支持他们在上面构建 AI 代理。

We also support their ability to build agents on top of this.

我们托管了叫作 Sailboxes 的东西,这是云端长期运行的代理虚拟机,专为那些运行几小时、几天甚至几周的代理设计的。

We host what we call Sailboxes, which are long-running agent virtual machines hosted in the cloud that are designed for agents that run for hours, days, or weeks.
M2
M21:17

所以你应该被看作是一家和提供不同类型推理服务的其他公司同级别的公司。

And so I should think about you as a peer company to others that serve different kinds of inference.

你提供的是某一种特定的推理服务,而你的目标就是成为绝对最便宜的供应商,让某种智能应用成为可能。

You're serving one specific kind of inference, and your goal is to be the absolute cheapest provider and enabler of a certain kind of use of intelligence.
M1
M11:30

没错。

Exactly.

我们公司的主题是富足。

The theme of our company is abundance.

我们希望把这种新的智能商品以几乎每个行业都能承受的成本,送到尽可能多的人手里。

We want to deliver this new commodity of intelligence to as many people as possible at a, at a cost that is sustainable for almost every industry.

我们认为,当一样东西便宜了 10 倍,它就会成为一个全新的产品类别。

We think that whenever you make something 10 times cheaper, it's a new product category.

而我们的目标就是让 token 也实现这一点。

And we aspire to do that for tokens.

我们认为机器能思考这件事意义太深远了。

We think it's so profound that that the machine can think.

而现在我们的任务就是让世界上尽可能多的机器去思考。

And now our job is to make as many machines as possible in the world work towards thinking.
M2
M21:53

那么,如果今天讨论的主题是 token 成本,你觉得 token 成本是思考这件事的正确维度吗?

So if you think about the theme of the day being token costs, is token cost the right way to think about this?

还是说,你会有别的表述方式?

Like, is there some other way you'd put it?
M1
M12:01

一开始,绝对是 token 成本。

To start with, absolutely token cost.

现在我的北极星目标就是做到行业内每 token 成本最低,而且要遥遥领先。

Today my North Star is I want to have the lowest cost per token in the industry and do that by a mile.

我不觉得 token 是最终的工作单元或者智能单位,但它是我们目前用的东西。

I don't think tokens are the final unit of work or intelligence, but they are what we use today.

所以这很直接。

And so it's very straightforward.

我觉得过了 token 之后,就会慢慢转向更多基于结果的方向,这个方向还比较模糊。

I think after tokens you start to move more towards More outcomes, which is like a vague direction.

比如,你可以想象,现在当你通过一个智能体消耗 token 时,你实际上控制不了智能体为了推理用了多少 token。

You can imagine, for example, today when you consume tokens through an agent, you don't actually control how many tokens the agent reasons for.

它可以推理一段时间,或者调用一定数量的工具。

It can reason for a certain amount of time or it can call a certain number of tools.

而且我认为,越来越地,我们会让智能体完成一些工作任务,尽可能多地尝试。

And increasingly, I think we will have agents do some unit of work, take as many shots on goal as they can.

而它们用了多少 token 来达到目标,其实是一个依赖于任务的变量。

And however many tokens they use to get there is going to be kind of a dependent variable depending on the task.
M2
M22:41

嗯哼。

Mm-hmm.
M1
M12:42

所以你想想看,比如智能体自己管理 token 预算,而不是公司为工程师每月能用多少 token 设定预算。

So you think about like agents that self-administer a token budget as opposed to a company setting a budget for how many tokens engineers can spend per month.
M2
M22:50

为什么你认为这里有机会可以抓住?

Why is there an opportunity that you can tackle?

现在整个世界似乎都围绕着更多、更好、更快、更便宜的 token 来发展。

It seems like the entire world is oriented around more, better, faster, cheaper tokens right now.

全世界都在非常积极地解决这个问题。

It seems like the world is trying to solve this problem very aggressively.
M1
M13:02

对。

Yeah.
M2
M23:03

你看到的独特机遇是什么?市场可能不够高效。我认为有两件事是我们公司的顺风。

What was the unique opening that you saw that's maybe the market's not being efficient and it's I think there's 2 things that are tailwinds for our company.
M1
M13:12

第一个必须是开源的发展。

One has got to be the rise of open source.

我得先谈谈这个。

I had to talk about that first.

我认为我们开始看到越来越多的客户和更广泛的市场在意拥有智能。

I think we are starting to see an increasing number of our customers and the broader market care about owning intelligence.

他们希望对所依赖的东西有控制权和自主权。

They want to have control, sovereignty over the thing that they depend on.

因此,这就为定制化模型,甚至那些任何人都无法从你手中夺走的普通开源模型,创造了一个更加繁荣的市场。

And so that created a much more robust market for customized models or even just like these vanilla open source models that no one can ever take away from you.

你总是拥有模型权重,你总是有权以任何方式部署它们。

You always have the weights, you always have the right to deploy them however you like.

在这个世界里,过去几年已经出现了一个相当繁荣的市场,用于大规模服务这些模型。

In that world, there's been a reasonably robust market for the past couple of years serving these models at large scale.

挑战在于所有这些公司——你可以随便挑,比如Base10、Fireworks、Together——它们都专注于低延迟推理。

The challenge is all those companies, you could take your pick, Base10, Fireworks, Together, they all focus on low latency inference.

它们是被一个非常重要的客户Cursor拉向那个方向的。

And they were pulled in that direction by one very important customer, Cursor.

而且我认为大约一年前那是正确的选择。

And I think that that was the right choice about a year ago.

而到了六个月前,情况开始变得像是低延迟可能并不是你对智能体的唯一要求了。

And as of 6 months ago, it started to look like maybe low latency wasn't the only thing you wanted from an agent.

你希望有更持久的执行,处理更长周期的任务。

You wanted more persistence, more long horizon tasks.

而现在对我来说非常明显,智能体推理的未来就是长周期任务。

And now it's to me very obvious that the future of agentic inference is long horizon tasks.

你会让机器一次运行几个小时甚至几天。

You're going to run the machine for hours or days at a time.

它以每秒100个token的速度输出也没什么关系。

It doesn't matter if it spits out tokens at 100 tokens per second.

也许每秒10个token也完全可以,只要随之而来的是相应的优势和效率。

Maybe 10 is just fine if that comes with corresponding advantages and efficiency.
M2
M24:22

你为什么对这个这么有信心?

Why are you so confident in that?

对我来说,我好像什么都希望越快越好。

To me, it seems like I want everything as fast as possible.
M1
M14:27

当你在等待的时候,你绝对应该得到最快的速度。

When you're waiting on it, you absolutely deserve the fastest answer possible.

我的诀窍就是不想让你等着它。

My trick is I don't want you to be waiting on it.

我希望它是主动出击的。

I want it to be proactive.

我希望它是在后台默默运行的。

I want it to be in the background.

有一种说法是,最好的延迟就是没有延迟。

One way to say it is the best latency is no latency at all.

你早上醒来,活儿昨晚已经干完了。

When you wake up in the morning, the work's already been done overnight.

你甚至都不需要开口提要求。

You didn't even have to ask for it.

这就是理想状态。

That's the dream.

我们还没完全做到那一步。

We're not quite there yet.

但更重要的是,我觉得你在提示 agent 并等待回复的过程中参与得越多,实际上你就越成了帮 agent 多干活或少干活的瓶颈。

But more importantly, I think the more you're in the loop as you prompt agents and wait for a response, in fact, you're the bottleneck in helping the agent do more or less work.

我们希望 agent 能按照人类的节奏来运作。

What we'd like is the agent to operate on more human timescales.

你不会每五分钟就去管一下你的同事吧。

You don't manage your colleagues every 5 minutes.

你给他们布置一项高层次的任务,然后可能每天或更可能每周回来检查一次进度。

You ask them to do a high-level task and you come back and check in maybe every day, but more likely once a week.

而这在我看来就是人机协作的未来,更像人类的节奏。

And that to me is the future of human-agent collaboration, more like human timescales.
M2
M25:08

多讲讲那些早期迹象,让你觉得这事正在发生,所以你该去创办这家公司。

Say more about the early indications that this is happening and therefore you should be building this company.
M1
M15:13

嗯,首先最重要的一点是 test-time compute scaling 这个概念。

Well, so the first and most important thing is the idea of test-time compute scaling.

就是给 agent 更多时间,它能给出更好的答案。

The idea that you can give an agent more time and it will give you a better answer.

这个概念大约两年前就被提出来了。

So that was theorized about 2 years ago now.

但直到去年 Opus 4.5 出现,我们才真正能把它当成一个可靠的方向。

And it wasn't really something that we could actually bet on until I would say late last year with Opus 4.5.

Opus 4.5 是第一个适合处理较长时间跨度任务的 agent。

Opus 4.5 was the first agent that was at all suitable for longer horizon tasks.

它刚出来的时候表现一般,但你看看最近的新模型以及我们在开源方面的工作,就会发现 agent 已经能连续运行一个小时了。

And it was pretty mediocre when it first came out, but you look at the more recent models and what we've done on open source as well, and you see that agents are capable of running for an hour at a time.

我不敢说能运行好几天,但一个小时肯定没问题。

I wouldn't say it's days, but definitely an hour is quite suitable today.

所以光是看到平均轮次或任务时长越来越长,不需要太多数据点你就能画出一条指数曲线。

And so just seeing that average turn or task length get longer and longer, it doesn't take many points to have you draw out the exponential.

然后你就会发现,让 agent 运行更长时间是值得的。

And see that agents are worth running for longer periods.
M2
M25:58

你觉得三年后,长时间运行的 agent 会占到多少市场份额?

What do you think will be the market share of long-running agents in 3 years or something like this?
M1
M16:03

我喜欢这个市场,因为它没有上限。

I love this market because it's unbounded.

没有人在流程中参与,所以你可以在后台消耗任意多的 token,而不受人类注意力的限制。

There's no human in the loop, so you can consume as many tokens as you like in the background versus human attention span.

如果你让我在 Codex 或 Claude Code 上消耗 10 倍的 token,我其实不确定自己还能不能做到。

If you tell me to consume 10x as many tokens at Codex or at Claude Code, I'm actually not sure if I can anymore.

因为我已经在流程里了,我对着电脑的时候,大部分时间都在编程。

I'm already in the loop and locked in coding for most of the day that I'm at the laptop.

真正的上限是没有的,是在后台或主动模式下能消耗多少 token。

What is unbounded is how many tokens can be consumed in the background or proactively.

所以长远来看,我觉得今年结束时,后台和实时工作负载大概会各占一半,但我预计这个比例会变成九比一,后台占大头。

So long-term, I think we're going to end this year at maybe 50/50 background and real-time workloads, but I see this going to 90/10 in favor of background.
M2
M26:33

有哪些事情,你最喜欢的例子是什么,作为后台任务比作为有人参与的循环任务完成得更好?

What are the sorts of things, what are your favorite examples of something that gets accomplished much better as a background task than as a human-in-the-loop task?
M1
M16:40

大多数深度研究,大多数你想得到明确答案的问题,不是基于100个来源,也不是1000个来源,而是基于一万个或更多来源。

Most deep research, most questions where you want to have a definitive answer over not 100 sources, not 1,000 sources, but 10,000 sources or more.

如果你想建立一个权威的信息索引,比如说,我们的一个客户Parallel Web Systems就想做这件事,他们想要建立覆盖整个互联网的索引,并且实时监控互联网的变化。

If you want to build an authoritative index of information, like for example, one of our customers, Parallel Web Systems, seeks to do, they want to build an index over the whole internet and they want to monitor the internet in real time for changes.

这是那种疯狂的艾字节规模的任务,你需要一种非常不同的智能或者说是规模的智能才能实现。

That is the kind of crazy exabyte-scale task that you need a very different kind of intelligence or scale of intelligence to achieve.

深度研究是我们的首要类别。

Deep research is a top category for us.

而且我们越来越看到网络安全也在朝着这个方向发展。

And then increasingly we see cybersecurity following this direction.

如果你想一想,是的,你可以生成很多代码,但是破坏同一段代码的方式比生成它的方式要多出指数级。

If you think about, yes, there's so much code you can generate, but there's exponentially more ways to break that same code than to generate that code.

而且有一些很棒的客户正在非常努力地寻找能够破解任何软件并主动修补它们的智能体。

And there are some great customers out there who are working very hard to find agents that can break any piece of software and proactively patch them.

所以,比如当Fable刚发布或者Mythos刚发布的时候,基本上网络安全社区里有一种推动力,要对每一行我们写过的代码运行Fable。

So when Fable first came out, for example, or Mythos first came out, basically there was this push in the cybersecurity community to run Fable against every line of code we've ever written.

用二十种不同的方式寻找漏洞,意思是你要寻找内存错误、业务逻辑错误、网络漏洞等等所有这些。

And look for bugs in 20 different ways, meaning you're looking for both memory errors, you're looking for business logic errors, and looking for network vulnerabilities, all these things.

而这些实际上都是你会编写专门的智能体来处理的事情。

And these are all actually things that you would write specialized agents for.

你不会只是让Fable看一遍源代码,而是让它实际搭建环境,让你可以对这些应用程序进行渗透测试。

You wouldn't just have Fable look at the source code once, you'd have it actually set up environments where you can pen test these applications.

到某个时候,人们开始开这个玩笑,说安全变成了工作量证明。

And at some point, people started to make this joke that security has become proof of work.

当你想要安全的软件时,这实际上变成了一个问题:你在Anthropic的API上花了多少钱?

When you want secure software, it's really a question of how many dollars did you spend on Anthropic's APIs?

尝试攻破你的软件。

Trying to break into your software.

这是衡量它有多安全的最佳指标,因为这是世界上最好的工具。

That is the best indication for how secure it is because that's the best tool in the world.

而且我们越来越发现,这里的智能前沿是相当不平坦的。

And increasingly we found that the frontier of intelligence here is quite jagged.

并不是说Fable能找到软件中所有漏洞的超集。

It's not the case that Fable finds a superset of all bugs in software.

你会用很小的模型找到一些大型模型找不到的漏洞。

You would find some bugs with a very small model that you don't find with a large model.

你会用Haiku找到一些Fable找不到的漏洞,反之亦然。

You'd find some bugs with Haiku that you wouldn't find with Fable and vice versa.

所以这鼓励了这种非常多样化的采样方法,并尝试构建自主破解软件的网络安全智能体。

So it encouraged this very diverse approach to sampling and trying to build cybersecurity agents that break software autonomously.

这样你就可以修补它们。

Such that you can patch them.
M2
M28:35

如果你稍微推测和想象一下,那些长期运行、非常廉价、非常持久的智能体能够实现哪些事情,我们讨论了一些非常实际的例子,比如深度研究、网络安全等等。

If you were to get sort of like speculative and imaginative about the sorts of things that long-running, very cheap, very long-running agents can enable, we talked about some very practical examples, deep research, cybersecurity, et cetera.

但如果你对这些用例想得更梦幻一点,这种推理将会解锁一种新的产品类别。

But if you get a little bit dreamier about the use cases, a new product category that this sort of inference will unlock.

我想问题就是,那又怎样?

And I guess the question is just like, so what?

比如,如果你取得了最大程度的成功呢?

Like what if you're maximally successful?

稍微想象一下这可能会带来什么。

Dream a little bit about what that might enable.
M1
M19:03

是的,绝对。

Yeah, absolutely.

所以对于个人用户来说,我最兴奋的是这种主动智能体的概念。

So I think for individual users, what I'm excited about most is this idea of proactive intelligent agents.

你可以想象一个一直在后台运行的Siri,它了解你一天中收到的所有邮件、所有短信,对你的生活有一个更全面的百科全书式的视角,并且知道如何在那生活中提供帮助。

You can imagine a Siri that is running in the background all the time to understand all the emails you received in a day, all the text messages you receive in a day, and it has a much more encyclopedic view of your life and how to be helpful in that life.

目前还有一些点解决方案,所以最终你还是得做大量的提示工作。

Right now, there's still point solutions, and so you have to, you end up doing a lot of prompting.

Siri 不是很主动。

Siri is not very proactive.

这个问题,我们可以通过充足的推理来解决。

It's something that we can fix with abundant inference.

如果你足够信任机器,觉得它可靠又可信任——比如隐私方面——那你甚至可以想象,机器能理解你是如何与它互动的,并主动为你展示下一步操作。

If you trust the machine enough that it's reliable and also trustworthy, as in private, you might even imagine the machine can understand how you interact with it and proactively surface your next action.

每当你打开手机,我们能不能建立一个好的模型,预测你接下来要做什么?

Whenever you open your phone, can we build a good model of what you're going to do next?

而我的估计是,是的,我们完全可以。

And my estimation is yes, we totally can.
M2
M29:46

而关键就在于极其廉价的智能。

And the key to that is incredibly cheap intelligence.
M1
M19:48

你必须愿意在没有回报承诺的情况下花费 token。

You have to be willing to spend tokens without any promise of return.

从长远来看,我们拥有一种可以解决任何可验证问题的智能形式。

That's the long lens view to take on this is that we have a form of intelligence that can tackle any verifiable problem.

任何可验证问题,意味着大多数软件。

Any verifiable problem means most software.

也意味着很多形式化的数学证明之类的东西。

It means a lot of formal math proofs and similar.

还可能意味着科学发现。

And it could also mean scientific discovery.

这些都是相对可验证的问题。

These are all relatively verifiable problems.

而所有这些事情目前都有个美元成本,本质上是个隐藏成本。

And all those things currently have a dollar cost attached to them, essentially, that's a hidden one.

就像你需要投入多少 token 才能让它运转起来?

It's like how many tokens could you possibly harness to make this work?

我们其实已经开始让这些长周期任务的美元成本变得可见且合理。

And we have actually started to bring it within view, a dollar cost for these long horizon tasks that is reasonable.

不是百万级,而是千元级,甚至在不久的将来可能降到几百甚至几十美元,就能获得任何科学问题或研究问题的明确答案。

It's not millions, it's thousands, and maybe it could be hundreds or even tens of dollars in the near future to have a definitive answer to any scientific question, to any research problem.
M2
M212:04

所以如果我们畅想那个未来,基本上我们最后只受限于人们能提出的问题。

And so if we dream about that future where we then become limited just by the questions that people can ask, basically.
M1
M112:12

差不多。

Pretty much.

我们能提出的问题,模型已经处于可以处理一个高层问题并追踪到底的边缘了。

The questions we can ask, the models are on the cusp of basically taking even a high-level question and chasing it down.

每一个可能的后续问题,你都可以让模型自己去处理。

Every possible follow-up, you can have the model essentially take that on its own.

问题就是,你愿意花多少 token?

And the question is, what is your token budget?

我们会解决 token 预算的问题。

And we will solve the token budget problem.
M2
M212:27

那不可验证的任务呢?

What about non-verifiable tasks?
M1
M112:29

我基本上把人类品味这一类全都归到那个类别里了。

I put basically the entire category of human taste into that category.

我们还没有解决人类品味的问题。

We have not solved human taste yet.

我也不确定从根本上能否解决。

And I don't know that it fundamentally can be.

我期待能有惊喜,但我们目前专注于非常定量的问题。

I'm excited to be surprised here, but we are focused on very quantitative problems.

我们把写作的质量、艺术的美感留给人类。

We leave the quality of writing, we leave the beauty of art to people.
M2
M212:47

好了,现在我们来聊聊你希望最终构建的那个非常巧妙的解决方案栈,用来打造一个巨大的 token 工厂,成为一个提供极其廉价智能的供应商。

All right, now let's talk about the very clever stack of solutions that you hope to build ultimately to have this giant token factory, extremely low-cost intelligence supplier of extremely low-cost intelligence.

我觉得你可以从三个层面来思考这个问题:软件、硬件和电力。

I think you think about this in terms of 3 levels: software, hardware, and power.

那就说说你的整体计划吧,怎么去应对这个和大多数人想法都不同的挑战?

Talk through what your master plan is to approach this challenge that's so different from what others are thinking about doing?
M1
M113:09

我们一直都得从软件入手。

We always had to start with software.

在现在的芯片和数据中心上,有哪些机会可以提升效率?

Where is the opportunity on today's chips with today's data centers to improve efficiency?

我们做的第一件事,就是尝试围绕 GPU 的峰值效率搭建整个 LLM 软件栈,也就是说,我们用的是 NVIDIA 的 GPU。

And the first thing we did was we tried to build the entire LLM software stack around peak GPU efficiency, meaning we're using NVIDIA GPUs.

我们想从同一块芯片上榨出的 token 数量比世界上任何人都多。

We wanted to squeeze out more tokens from the same chip than anyone else in the world.

这要从最底层的编程开始,也就是 kernel。

And that starts with the lowest level of programming, kernels.

这其实是我的老本行。

It's actually my background.

我这辈子,不对,是我整个职业生涯,都在和 GPU 以及 kernel 打交道。

I spent my whole life, actually my whole professional life working on GPUs and kernels.

我的第一份工作就在 NVIDIA。

NVIDIA was my first job.

我上大学的时候,亲眼见证了 Tensor Core 是怎么一步步赢得它在芯片上的位置的。

While I was in college and I got to see how the Tensor Cores got to earn their right to be on the chip.

那是在 2016 年。

This is back in 2016.
M2
M213:44

给外行人解释一下这到底意味着什么。

Just describe what that means for the layperson.
M1
M113:47

好,Tensor Core 是 GPU 上一个专门加速矩阵乘法的单元。

So, okay, Tensor Core is a specialized unit on the GPU that accelerates matrix multiplication.

就这么简单。

Simple as that.

Tensor Core 随时间演进的历程很长,我们后面再细说。

There's been a long history of how we evolved that Tensor Core over time that we'll get into.
M2
M213:56

那为什么矩阵乘法这么重要?

And why is matrix multiplication so important?
M1
M113:58

问得好。

That's a great question.

其实我也不能说宇宙真理就是矩阵乘法成了计算的基本单位。

I actually, I cannot say that there is a divine truth of the universe that explains why matrix multiplies seem to be the atomic unit of computation.

但有人跟我这么解释过:矩阵乘法是一种非常简洁的方式,能把两组数字混合在一起,让它们以有趣的方式相互作用。

But one way I've heard it described to me is, well, it's a really succinct way to mix 2 blocks of numbers together and have them interact in some interesting way.

关于这点我也只能说到这儿了。

That's as much as I can say about it.

线性代数恰好能很紧凑地表示任意关系和数据,这确实很方便。

It is really convenient that linear algebra turns out to be a very compact representation of arbitrary relationships and data.

NVIDIA 是一家很棒的图形公司,在 GPU 和游戏图形领域占据市场主导地位已经很久了。

So NVIDIA, great graphics company, obviously has had market share dominance in GPUs and gaming graphics for quite some time.

然后从 2010 年代中期开始,他们启动了一些秘密项目,让图形处理器更适合他们正在关注的机器学习任务。

And then starting in the mid-2010s, they started to actually start these skunkworks projects to make the graphics processor more suitable for machine learning tasks that they were tracking.

我记得我在 NVIDIA 的时候,还读过一些经理的实验笔记。

I remember actually reading some of the lab notebooks of some of my managers when I was at NVIDIA.

他们那时候会去参加像 ICML 或 NeurIPS 这样的小型机器学习会议,然后记下这些论文,心想:哦,这个深度学习的东西好像开始流行起来了。

They would visit these small ML conferences like ICML or NeurIPS at the time, and they would just take note of these papers like, oh, this deep learning thing seems to be catching on.

特别有意思的是,这些研究生用游戏级的 NVIDIA GPU 来训练他们的大模型。

And what's really interesting is that these grad students are using gaming NVIDIA GPUs in order to train their large models.

我们应该深入了解一下,看看是怎么回事。

We should double-click on this and figure out what's going on here.

到了 2015、2016 年,至少 Jensen 有决心去加码,觉得我们芯片的这种用途只会越来越大。

And by 2015, 2016, at least Jensen had the conviction to kind of double down on, hey, this usage of our models is only— or of our chips is only going to grow.

让我们开始把越来越多宝贵的硅面积分配给这种正在兴起的能力。

Let's start allocating more and more precious silicon die area to this capability that seems to be emerging.

我们把第一版Tensor Cores放到芯片上。

Let's put the first version of Tensor Cores on the chip.

所以我们说的是拿这个游戏芯片——它本来是用来在屏幕上画像素的——然后改造它来做矩阵乘法。

So we're talking about taking this gaming chip, which is designed for painting pixels on a screen and adapting it to do matrix multiplies.

那时候还很早,你基本上要跟图形团队竞争。在任何芯片公司,你要更多的硅片面积,总是有竞争的。

And it was early and you would be competing against the graphics teams essentially when you ask for more silicon area at any chip company, there's always competition for that.

设计师们对那个保护得非常小心。

It is something that the designers guard so carefully.

你绝对不想投资到错误的技术上,因为那是机会成本,你本来可以把资源分配到其他功能上的。

You don't ever want to invest in the wrong technology because that's opportunity cost that you could have allocated to some other functionality.

所以我们拼命争取,最后只拿到了很小的一点芯片面积,大概5%到10%吧,给第一代这样的芯片,用来对基本的卷积做一定加速,而卷积是当时计算机视觉模型的基本运算。

And so we kind of fought and tooth and nail and got just a tiny bit of die area, maybe like 5, 10%, something like that for the first generation of these chips to get some amount of acceleration for basic convolutions, which were the fundamental operation for computer vision models in the day.

然后我们有一个软件团队,试图把芯片的所有性能榨干。

And then we had a software team that was trying to squeeze all the performance we could out of the chip.

我想在那个软件团队——也就是我工作的地方——那实际上教会了我最多关于NVIDIA的那种精神。他们有一个术语叫'speed of light'。

And I think on that software team, which is where I worked, that's what actually taught me the most about just the ethos that NVIDIA has around. They have this term called speed of light.

他们做任何硬件,都会去追求光速。

They always chase the speed of light for any piece of hardware that they make.

这在每个工程师的脑子里根深蒂固:如果机器能做到,我们就会把机器推到极限,直到它做到我们认为它能做到的。

It is so ingrained in every engineer's mind that if the machine can do it, we're going to push the machine to the frontier until it does what we think it can do.
M2
M216:29

光速就是可能性的极限边缘。

And the speed of light is the edge of what's possible.
M1
M116:31

光速就是可能性的边缘。

The speed of light is the edge of what's possible.

没错。

Exactly.

如果我们认为芯片能以这个频率运行,每个周期能产生这么多乘法运算,我们就一定要达到那个目标。

If we think the chip can run at this frequency and produce this many multiplies per cycle, we're going to get there.

我们会打破每一个瓶颈,达到那个性能峰值。

We're going to break every bottleneck and get to that peak level of performance.

所以直到今天,我跟我所有的工程师说,我们在追求100%的speed of light。

And so to this day, I tell all my engineers, like, we're chasing 100% Speed of light.

我不在乎跟竞争对手相比的相对数字。

I don't care about relative numbers versus the competition.

我只在乎绝对数字。

I only care about absolute numbers.

我们能在芯片上做到什么,以及我们怎么实现它?

What are we able to do on the chip and how do we achieve that?
M2
M216:51

在我们结束你在Nvidia的那段经历之前,除了那个文化触点之外,还有没有别的真正改变了你思考方式的东西,或者当时那个公司运营方式、它的文化里最让你印象深刻的?

Before we leave that chapter of your time at Nvidia, anything else beyond that cultural touchpoint that really like changed the way you think about things or that stood out the most about how the business ran back then or its culture?
M1
M117:02

我有好多关于Nvidia的故事。

I have a ton of stories about Nvidia.

我们可以——我可以给你讲几个。

We can— I could tell you a few of them.

我最喜欢的一个是,从任职时间来看,2015、2016年和我一起在Nvidia工作的很多人今天还在那里。

One of my favorites is that on the tenure side, a lot of people I worked with in Nvidia in 2015, 2016 are still there today.

那家公司的员工留存率高得惊人,而且坦率地说,这些是硅片领域最好的工程师,至少是我整个职业生涯中合作过的最好的。

That company has incredible retention and these are the best engineers, frankly, on the silicon side, at least I've worked with in my whole career.

他们极其、极其有动力,充满热情。

They're extremely, extremely motivated and passionate.

他们相信并行计算这个概念,经历了它的各种形态,也喜欢看着芯片不断进化。

They believed in parallel computing as a concept through its various incarnations and have loved seeing the chip evolve.

这是他们毕生的事业,而且他们在那个方向上非常非常有能力。

This is their life's work and they're extremely, extremely competent in that direction.

他们同时也是一家非常节俭的公司。

They're also a very frugal company.

NVIDIA,还有所有——我想是2008年之后所有的硅谷公司,都在员工福利上做了一些削减。

NVIDIA and all, I guess, all the Silicon Valley companies after 2008, they had some cutbacks in like perks.

比如说,没有免费午餐了。

So no free lunch, for example.

NVIDIA 更进一步了。

NVIDIA took it one step further.

冰箱里没有免费牛奶。

There was no free milk in the fridge.

所以你在 NVIDIA 想喝咖啡加牛奶的话,每个月得交一块钱给牛奶俱乐部,他们会用 Costco 的牛奶把冰箱填满。

So if you wanted to drink coffee at NVIDIA and you wanted some milk, you actually had to chip in $1 every month to the milk club, and the milk club would stock Costco milk in the fridge.

我记得清清楚楚。

And I remember that distinctly.

我们在 Sail 不搞这个,但这种节俭文化渗透了整个公司。

We don't do that at Sail, but it's a frugality that permeates the company.
M2
M217:58

从那段经历出来,你就学会了怎么通过软件更高效地利用底层硬件。

And so coming out of this time there, you get this experience of what it's like to develop more efficient usage of the underlying hardware through software.
M3
M318:06

对。

Yes.
M2
M218:06

那把它跟今天的环境联系起来呢?

And so link that to today's environment.
M1
M118:10

嗯,没错。

Yeah, absolutely.

我觉得 GPU 本质上就是个吞吐量机器。

So I think the GPU is fundamentally a throughput machine.

GPU 最开心的时候,就是你给它一堆活干,让它满负荷地跑计算单元。

The GPU is happiest when you give it a lot of work to do and let it chew through that work at peak utilization of its compute units.

但这并不是过去十年我们使用 AI 的方式,对吧?

But that's actually not the way that we've taken AI in the last 10 years?

我们实际上把 AI 推向了交互式聊天工具,这是今天最常见的用法。

We've really pushed AI to be an interactive chatbot tool is the most common form of AI usage today.

在这种场景下,你特别关心的是尽快把答案吐给键盘前面的人。

And in that world, you care a lot about actually spitting answers out to the person at the keyboard as quickly as possible.

你说得对,别让用户等,我要的就是越快越好。

To your point about don't make the user wait, I want things as fast as possible.

这对 GPU 来说就很有意思了。

And so that's actually quite interesting for the GPU.

你想快速吐出 token 的时候,很难让 GPU 进入那种计算单元满载的 happy path。

It's very difficult to put the GPU in its happy path of being fully compute utilized when you're trying to spit out tokens quickly.

GPU 上有一个根本的权衡:要么侧重吞吐量,要么优化延迟。

There's a fundamental trade-off on the GPU between being throughput-oriented or latency-optimized.

因为用法是聊天机器人导向的,大家都选了延迟优化。

And everyone has chosen latency optimization because the shape of usage was chatbot-oriented.

我相信这是未来一年我们能看到的最深刻的变化。

I believe that's the most profound change we're going to see in the next year.

我们会从聊天机器人转向更主动或后台运行的 agent。

We're going to move away from chatbots to more proactive or background agents.

在那个世界里,围绕吞吐量来构建整个体系就合理多了。

And in that world, it makes a lot more sense to build a stack around throughput.
M2
M219:10

你能从技术上解释一下为什么吞吐量和延迟之间的权衡无法打破吗?

Can you explain technically why the trade-off between throughput and latency is unbreakable?

为什么同一块硬件上我们不能两者兼得?

Why can't we have both from the same hardware?
M1
M119:17

这在几乎所有你能想到的系统里都是很基础的问题。

It's quite foundational in almost every system that you could ever possibly look at.

总是有个权衡:要么让少量数据尽快通过系统,留出足够的缓冲空间;要么就宽通道慢处理。

There's always a trade-off between getting a small amount of data through the system as quickly as possible and leaving a lot of buffer room for that, or trying to run wide and slow.

就像窄通道快跑对宽通道慢跑,这是计算机科学里的经典权衡。

Like narrow and fast or wide and slow is like a classic trade-off in all computer science.

但具体到 GPU,我觉得有个东西值得关注,就是 GPU 上的 batching 概念。

But for GPUs specifically, I think there's one thing to focus on, which is there's this concept of like batching on the GPU.

我们想把很多用户的工作打包成一个 batch,然后在 GPU 上一次性跑完。

We want to group many users' work together into a batch that we can run all at once on the GPU.

这就是 GPU 的并行处理。

That's the parallel processing of the GPU.

我们希望能有大量并行的工作要做。

We'd like to have a lot of parallel work to do.

但问题是,当你一起运行一大批次计算时,实际上你做的总工作量是更多的。

The thing is though, you're doing net more work when you run a large batch of compute together.

所以你可能填满了所有单元,但在每一步,当你在GPU上携带一批工作前进时,要做的工作反而更多了。

And so you might be filling all the units, but every step along the way, as you carry a batch of work through the GPU, there's more work to be done.

因此,批次中的任何一个token或任何一个用户的请求,都会在GPU上花费更长的时间,因为要和其他人的流量一起被处理。

And so any individual token or any individual user's request in that batch, it's going to spend a longer time on the GPU being carried with other people's traffic.

也许可以这么说,如果你想去旧金山市中心,你可以坐公交,也可以坐私人交通工具。

Maybe the way to say it is, if you want to get downtown in SF, you can take the bus or you can take a private transit.

私人交通会有一条直达路线,直线距离最短,或者完全按照你想走的路线从A到B。

And the private transit is going to have its own direct path as the crow flies or using exactly the roads that you want from point A to point B.

而公交车要服务更多乘客,它必须从根本上做一些对所有人都有效的事情。

A bus, it's going to have to serve many more people and it has to fundamentally do something that works for everyone.

所以它会走更慢的路线,而且会停下来等其他人上下车。

And so it takes a slower path and it stops and waits for other people to get on and off.

我觉得公交车和汽车的类比还挺准确的。

I think the bus versus car analogy is pretty accurate.
M2
M220:39

这是个很好的类比。

And it's a great analogy.

所以你们要做的事情的第一步,就是在NVIDIA GPU上打造出最好的公交车。

And so step one for what you're trying to do is create the best possible bus on top of NVIDIA GPUs.

那是你们的第一步——完全正确。

That's step one of your—that's exactly right.
M1
M120:49

这意味着我们会探索不同的并行方案之类的东西。

It means we explore things like different parallelism schemes.

也许我可以再给你举个例子,在NVIDIA GPU上,他们真正创新且做得非常出色的一点是GPU之间的NVLink互连。

Maybe that's another example I can give you is with NVIDIA GPUs, one of the things that they've really innovated on and done a great job with is the NVLink interconnect between GPUs.

事实上,NVLink系统非常出色,如果你有一个大型矩阵乘法想要更快地执行,你可以把这个矩阵乘法切成两半,然后分散到2个甚至最多8个GPU上。

And in fact, that NVLink system is so good that you can, if you have a large matrix multiply that you want to perform faster, you can actually cut that matrix multiply in half and shard it across 2 or more, up to 8 GPUs.

比如说NVIDIA GPU,让它们各自处理那个大型矩阵乘法的一部分,最后再把结果连接起来,汇总到一起。

Let's say NVIDIA GPUs, and have them all work on pieces of that larger matrix multiply and have them connect their results together at the end, reduce their results back together in the end.

这是一个非常非常好的方法来降低操作的最低延迟。

And this is a great, great way to cut the minimum latency of an operation.

每个GPU现在只做原来的八分之一的工作,所以它能更快完成,但并不是快8倍。

Each GPU is now doing 1/8 as much work, let's say, and therefore it can finish faster, but not 8 times faster.

这是次线性扩展。

It's sublinear scaling.

你会用8倍的硬件,但不会得到8倍的速度。

You'll use 8 times more hardware, but you won't get 8 times the speed.

你可能只获得4到5倍的速度。

You might get like 4 to 5x the speed.

你得不到强扩展性。

You're not going to get strong scaling.

这是因为通信开销。

And this is because of communication overhead.

这是因为每个GPU在处理更小的数据块时,效率会比处理更大的数据块稍微低一些。

It's because every GPU is going to be a little bit less efficient working on a smaller tile of work than a larger tile of work.

所以这是唯一能加速的方法。

And so it's the only way to speed up.

如果你想要尽可能低的延迟,你可以这么做,但这不是我会做的选择。

If you want the minimum latency possible, you can do that, but it is not the choice I would make.

比如说,嗯。

For example, Mm-hmm.

我更倾向于使用不同的并行方案,比如专家并行或流水线并行。

I would prefer to use a different parallelism scheme like expert parallelism or pipeline parallelism.

我们可能会做一些有趣的事情来重叠和隐藏通信延迟。

And we may do interesting things to overlap and hide the communication latency.

在低延迟服务中,你做到这一点的能力就会较弱。

In a way that you would have less ability to do that for a low latency service.
M2
M222:12

所以,NVLink 这项技术是不是应该理解为它主要提升延迟性能?

So is the right way to think about NVLink as a technology which improves latency performance?
M1
M122:16

对。

Yes.
M2
M222:17

只是延迟性能?

And only latency performance?
M1
M122:18

这其实会引出下一段,也就是我们公司具体做了哪些不一样的事情。

Which will segue into the next segment of what we, you know, do differently as a company.

但没错,NVLink 对于低延迟推理来说,我觉得是必须的。

But yes, NVLink is mandatory, I would say, for low latency inference.

NVIDIA 在低延迟推理方面确实很厉害。

So NVIDIA is excellent at low latency inference.

但我告诉你,我们并不太关心低延迟推理。

And I'm telling you that we don't really care that much about low latency inference.

那这又意味着什么呢?

So where does that leave us?

我其实不指望其他公司能很快搞明白 NVLink。

Well, I think I'm not holding my breath for other companies broadly to figure out NVLink quickly.

这是个很有挑战性的技术。

It's challenging technology to figure out.

很难规模化。

It's hard to scale.

也很难产品化。

It's hard to productionize.

所以如果我拿到的是其他厂商的芯片,它在基础计算组件上表现不错,矩阵乘法依然能做得很好。

And so if I do have some other vendor's chip and it is good at the foundational compute components, it can still do matrix multiplies really well.

但它没法快速把计算结果在芯片之间传递。

It just can't communicate those results across its peers quickly.

那么,也许这种其他芯片在我的架构里可以充当一个非常划算的计算单元,每美元的计算性能很高。

Well, maybe there's room for that other chip in my stack as a really, really good compute per dollar option.

而我大多数情况下真正优化的就是:这块芯片有多少 FLOPS,以及我每小时运营、拥有和运营它的成本是多少。

And that's what I actually optimize for in most cases is how many FLOPS does this chip have and how much is it going to cost me per hour to operate, to own and operate?

所以确实有其他芯片在每美元 FLOPS 上比 NVIDIA 更高,但它们的互联能力可能没那么强。

And so there are other chips that definitely rank higher than NVIDIA on FLOPS per dollar, but they may not have as much interconnect.

所以我的任务就是搞清楚要用什么并行方案,让这块芯片适合推理。

And so it's my job to figure out what parallelism scheme am I going to use that's going to make this chip suitable for inference.

不会用张量并行。

It's not going to be tensor parallelism.

NVIDIA 基本是张量并行的必备,但其他技术也可能行得通。

NVIDIA is basically mandatory for that, but other techniques may work well.
M2
M223:27

那在离开延迟这个话题之前,你能谈谈像 Cerebras 或者其他能实现极快运算的公司吗?

So before we leave the latency part of the story, can you comment on companies like Cerebras or others that can perform incredibly fast operations?

我很好奇,你怎么看这些方法、这些公司,以及未来可能发生什么。

I'm curious, like, what you think about those approaches, those companies, what might happen in the future.

你对未来专注于超低延迟的硬件有什么预测?

What is your prediction for the future of very low-latency-focused hardware?
M1
M123:43

Cerebras、Groq,还有另外几家刚走出隐身模式的公司,我觉得它们做了一个很有意思的赌注:不只是造另一个 GPU,而是造一种不同类型的加速器,专注于不同的存储层次。

Cerebras, Groq, and a couple others that are coming out of stealth now, I think, have made a very interesting bet on not just building another GPU, but actually building a different kind of accelerator that focuses on a different memory hierarchy.

它们想最大化芯片上的 SRAM,把它用作非常快的存储,来存放权重和 KV cache。

They want to maximize the amount of SRAM on the chip and use that as very, very fast memory for weights and KV cache.

SRAM 和 DRAM,给芯片做存储有两种方式。

So SRAM versus DRAM, there's 2 ways to make memory for a chip.

一种是把存储直接集成在逻辑芯片上,也就是你告诉台积电,我要在芯片上放这么多兆字节的存储。

One is to integrate the memory on the logic die itself, like meaning you tell TSMC, I want this many megabytes of storage on my chip.

有办法做到这一点。

And there's a way to build that.

台积电有一套标准单元库你可以用,你直接打印出一堆 SRAM 单元就行。

TSMC has a standard cell library you can use and you can just print out a bunch of cells of SRAM.

SRAM 的问题在于它占用了硅片很大的面积。

The problem with SRAM is it takes a lot of area on the silicon die.

所以如果你想做一个大芯片,比如说英伟达的Blackwell,800平方毫米,如果整个芯片都用SRAM,那容量大概也就个位数GB。

So if you want to build a large die, like let's say the NVIDIA Blackwell at 800 millimeters square, if you made that whole die SRAM, it would be in the maybe like single-digit gigabytes.

并不是特别大的数据存储量。

It's not a crazy amount of data storage.

相比之下,如果你愿意换一种完全不同的工艺技术。

Compare that to if you're willing to take a different process technology entirely.

那就不是台积电了,而是美光、SK海力士、三星。

So not TSMC anymore, but now Micron, SK Hynix, Samsung.

他们生产DRAM,这是一种完全不同的内存制造方式,更侧重于电容器而不是晶体管单元。

They build DRAM, which is a whole different way to build memory that's more focused on capacitors than transistor cells.

SRAM的标准构建方式叫做6T晶体管单元。

So SRAM, the standard way to build SRAM is what's called the 6T transistor cell.

这是一种稳定的晶体管结构,你可以写入一个比特,然后它会保持这个状态,不需要持续管理,当然还是需要供电,但它是静态保持的。

It's a stable transistor arrangement that allows you to write a bit to it and then it holds that state in that bit regardless of whether you keep applying, well, you had to apply some power, but it's holding that bit without any sort of like active management.

它是静态的。

It's static.

现在是动态RAM,也就是DRAM。

Now dynamic RAM, DRAM.

它之所以叫动态,是因为写入数据时,你是把电荷写到一个电容上。

It's dynamic because what you do to write some data is you write a charge onto a capacitor.

而一旦你把电荷写进那个电容,电荷就开始消散了。

And as soon as you write that charge into that capacitor, the charge is dissipating.

它在泄漏。

It's leaking.

所以DRAM的动态部分在于,你必须每50毫秒左右刷新你写过的每一个比特。

And so the dynamic part of DRAM is that you must, every 50 milliseconds or so, refresh every bit you've written.

所以你相当于同时在抛几十亿个球,也就是几十亿个比特,必须由一个内存控制器来管理,它不停地读取和刷新DRAM上的每个比特。

So you're constantly juggling billions of balls in the air, essentially billions of bits have to be managed by a memory controller, which is reading and refreshing every bit on the DRAM.

但这样做的好处是,你可以实现高得多的密度。

Now the benefit of that is you can get much, much higher density.

而且这是一种完全不同的工艺技术。

And it's a whole different process technology.

这里面有非常多的权衡,所以我们把DRAM制造分给了完全不同的公司,比如美光、SK海力士和三星。

There's a ton of different trade-offs, hence why we split the DRAM manufacturing into an entirely different company like Micron, SK Hynix, and Samsung.

这些是全球做这件事最顶尖的公司。

These are the best companies in the world to do this.

他们生产DRAM。

They build DRAM.

如果你从这些公司拿到DRAM,把它们堆叠成很多层,然后印刷或焊接在英伟达的主逻辑芯片周围,你就能得到几百GB的容量。

And if you take DRAM from those companies and you stack it into many layers and you kind of print them or solder them around the main logic die that you get from NVIDIA, you can now get hundreds of gigabytes.

比如Blackwell有288GB的HBM容量。

Like Blackwell has 288 gigabytes of HBM capacity.

都分布在逻辑芯片的周围。

Around the logic die.

而逻辑芯片本身可能只有大概500MB的SRAM。

And the logic die itself maybe only has like 500 megabytes of SRAM.

所以DRAM和SRAM的密度差距可能有好几个数量级,大概是3个数量级。

So it's possibly multiple orders of magnitude, 3 orders of magnitude difference in density for DRAM versus SRAM.

好,那咱们回到Cerebras。

Okay, so let's go back to Cerebras.

他们在做什么?

What are they doing?

他们看到了这个问题。

Well, they see this problem.

在芯片上提高SRAM密度并没有太明显的办法。

There's not really an obvious way to increase SRAM density on the chip.

但SRAM的好处是,它离真正做计算的逻辑门非常近,算术逻辑单元就在它要读取的SRAM旁边。

But the thing with SRAM is because it's so physically close to the logic gates that actually do the computation, the arithmetic logic units are right next to the SRAM that they're going to pull from.

做矩阵乘法的计算单元从SRAM读取数据,速度快得惊人。

The compute units that are doing the matrix multiplies can pull data from SRAM at just mind-boggling speeds.

Cerebras 宣称每秒能处理 21 PB,也就是他们的晶圆级引擎 3 能达到每秒 21 PB。

Cerebras quotes petabytes per second, 21 petabytes per second for their Wafer Scale Engine 3.

相比之下,NVIDIA Blackwell 上的 HBM 大概只有每秒 10 TB 左右。

And so compare that to HBM on an NVIDIA Blackwell is 10 terabytes per second or so in that range.

所以还是差了好几个数量级,容量更大但带宽比例上反而更低,差不多就是这样。

So once again, many orders of magnitude difference, more capacity, but proportionally less bandwidth, essentially.
M3
M327:04

对。

Yeah.
M1
M127:04

所以 Cerebras 的做法是,他们打算尽可能多地使用这些芯片。

And so what Cerebras does is they say that we're gonna take as many of these dies as we can.

他们不打算局限在台积电给我们设定的那个 800 平方毫米的掩模限制里。

We're not gonna limit ourselves to the 800 millimeter reticle limit that TSMC, 800 square millimeter limit that TSMC imposes on us.

他们要直接拿整个晶圆,并且让每个芯片通过划片线跟其他所有芯片相连。

We're gonna take the entire wafer and have actually every die connect to every other die over scribe lines.

他们就是想尽量在整个晶圆上堆更多的 SRAM。

And we're just going to try to get as much SRAM as we can on the whole wafer.

那么每个晶圆大概能拿到 50 GB 的 SRAM。

And we can get to, let's say, 50 gigabytes of SRAM per wafer.

然后他们会把多个晶圆像流水线那样叠在一起,或者用类似的方式。

And then we're going to stack many wafers together in a pipeline or similar.

这样一来,最多就能有 TB 级的内存,而且速度非常非常快。

And now we can have up to a terabyte of memory, very, very fast memory.

所有这些工作,就是为了让每个晶圆能以每秒 21 PB 的速度从 SRAM 读取数据。

And you do all that work just to get to the ability to read data from SRAM at 21 petabytes per second per wafer.

所以现在你能以极高的 tokens per second 来运行这些语言模型,因为你可以在大约一毫秒之类的时间里,把一个像 Kimi 这样的大模型的所有参数数据都搬进搬出芯片,或者说是搬进搬出逻辑核心。

Therefore, you can now serve these language models at extremely high tokens per second because you can move the entire parameter count of a large model like Kimi, you can move all that data in and off the chip, or sorry, in and off the logic cores in about a millisecond or something like that.

就是这样。

So there you go.

你有望达到每秒 1000 个 token。

You have a path to 1,000 tokens per second.
M2
M228:07

那么你对这个市场细分领域有什么预测?

And so what is your prediction for that segment of the market?
M1
M128:09

我觉得他们的结果会是某种混合形式。

Okay, so I think what happens to them is some hybrid sort of outcome.

我们必须把 Cerebras 芯片——这个在快速内存访问方面非常非常擅长的芯片——跟一个内存容量更大的东西搭配起来。

We had to pair the Cerebras chip where it's very strong.

因为虽然你可以把像 KIMI 这样 1 万亿参数的模型放到很多 Cerebras 晶圆上,但要处理 KV 缓存就没那么容易了。

It's very, very good at fast access to memory, with something that has more capacity for memory.

KV 缓存这个东西会随着用户使用模型越来越多而增长。

Because it's true that you can take a 1 trillion parameter model like KIMI and fit it on a large number of Cerebras wafers, but you can't do something about the KV cache very easily.

而且它始终是动态变化的。

The KV cache is something that grows as people use the model more.

你甚至不知道你需要多少 KV 缓存。

And that is always dynamic.

这完全取决于你的用户数量,以及你想服务多少用户。

You don't even know how much KV cache you're going to need.

你能用简单的话解释一下 KV 缓存吗?

It depends on what your users, how many users you have and how many users you want to serve.
M2
M228:45

可以。

Can you explain KV cache just like in basic terms?
M1
M128:47

KV 缓存是这样的,每次你用语言模型时,每个你输入的 token 实际上在你对话期间会一直留在模型的上下文窗口里。

Yeah.

所以如果我们对话了 10 万个 token,那么第 10 万零 1 个 token 仍然在之前的对话里,模型会参考所有过去的对话历史来更好地预测接下来要说什么。

So KV Cache, whenever you use a language model, every token you send through the language model actually stays in the context window of the language model for as long as you're having a conversation.

所以那个 KV 缓存就是一块内存。

So if we talk for 100,000 tokens, the 101,000th token is still in the conversation behind us and the model is referencing all the past conversation history in order to make better predictions about what the next thing we're going to say is.

你必须为你传入语言模型的每个 token 存储一个表示,而且它常常比模型本身的权重还要大。

And so that KV cache is a bunch of memory.

所以这个KV缓存就是一堆内存。你得为每个通过语言模型的token存储一个表示,而且它经常比模型本身的权重还要大。

You have to store a representation for every token that you send through the language model, and it frequently gets to be larger than the weights of the model themselves.

你可以这么理解:模型权重里存的是固化的知识,而 KV 缓存里存的是我们当前对话的动态知识。

You have this crystallized knowledge in the model weights, and you have the dynamic knowledge of the exact conversation we're having in the KV cache is the way I like to think about it.
M2
M229:35

所以有时候人们会观察到,聊到深入的时候内容开始变差,就是因为某种技术问题。

And this is why sometimes people would observe deep in a conversation, things start to degrade because there's some sort of technical problem.
M1
M129:41

对,KV 缓存在这方面挺有意思的。

Yeah, so the KV cache is quite interesting in that regard.

KV 缓存精确地记录了之前所有内容。

The KV cache is an exact representation of everything that came before.

我们把对话中见过的所有信息都存起来了。

We store all the information that we've seen in the conversation.

不过,在训练的时候,模型主要不是用超长上下文对话来训练的。

However, during training, the model did not get trained primarily on very long context conversations.

它主要是在八千 token 或者一万六千 token 的对话上训练的。

It got trained primarily on, let's say, 8,000 token conversations or 16,000 token conversations.

所以如果你把模型拉到二十万 token,虽然在那样的上下文长度上也有一些训练,但那不是模型的核心强项。

So if you take the model to 200,000 tokens, there was some training that happened at that context length, but it's not the model's core strength.

所以前沿实验室一直面临一个挑战:怎么让模型在一万 token 时跟二十万 token 时一样聪明?

And so there's always been a challenge for the Frontier Labs to figure out How do we make the model exactly as intelligent at 10,000 tokens as we expect it to be at 200,000 tokens?

这将会是一场持久战。

And it's going to be a perennial battle for us.

我们提出百万级上下文窗口这个概念已经好几年了。

We've had 1 million context windows as a concept for years now.

Anthropic 好像是最先达到一百万上下文窗口长度的。

Anthropic was, I think, the first to hit that 1 million context window length.

我在自己的 Claude code 里,远没到一百万字上下文长度的时候,就还在用 /compact 命令。

I still use /compact in my Claude code well before 1 million context length.

我觉得真正用到那么长其实并不好。

I don't think it's actually great to hit the full length.
M2
M230:34

所以这些极快、极低延迟的方法最终都会受这个因素的限制。

And so these extremely fast, extremely low latency approaches ultimately are limited by this factor.
M1
M130:40

没错,权重的存储你想怎么做都行。

Yes, you can do whatever you want for the weights.

权重存储做到无敌的性能是完全可以的。

It's very possible to have unbeatable performance on weight storage.

但是 KV 缓存会给你造成大麻烦。

However, the KB cache is going to be a big thorn in your side.
M2
M230:49

那么三五年后,你觉得这类芯片会扮演什么角色?

And so 3 years from now, 5 years from now, what role do you think these kinds of chips play?

比如在异构芯片市场上,它们能占多少份额?

Like what sort of market share do they have in the heterogeneous chip market?
M1
M130:57

Cerebras 和 Groq,可能还有其他几家,你应该把它们看作加速器。

Cerebras and Groq and maybe a couple others, you should think of them as accelerators.

它们真正擅长的是和更传统的、带有片外内存的 GPU 类设备一起使用。

What they are really good at is being used in conjunction with a more traditional GPU-like device that critically has this off-chip memory built in.

你需要片外内存来保证容量,片内内存来保证速度。

You want off-chip memory for capacity and on-chip memory for speed.

我们想把这两者结合起来。

We want to hybridize these 2 things.

拿 transformer 来说,如果你把上下文长度拉到一百万,就会出现两种情况:一个是计算密集的阶段,也就是我们说的 MLP 里的矩阵乘法,模型大部分的知识和世界知识都编码在这里。

So if you take transformers in the limit, you take a transformer to a million context length, what ends up happening is you have this compute-bound stage, which is the actual matrix multiplies for what we call the MLP, which is where most of the model's knowledge, world knowledge is encoded.

另一个是注意力层,它负责动态适应当前的对话。

And then you have the attention layer, which is where we are dynamically adapting to the current conversation.

在极限情况下,注意力层通常是内存瓶颈,而 MLP 在足够大的 batch size 下是计算瓶颈。

Attention in the limit is usually memory bound and the MLP in the limit is compute bound at large enough batch size.

我觉得 transformer 的原罪就是把一个本质上极度内存密集的层,紧挨着一个计算密集的层。

And I would say the original sin of transformers is that you've taken this extremely fundamentally memory bound layer and juxtaposed it right next to a compute bound layer.

很难让单一芯片同时擅长计算操作和内存操作。

It is very difficult to have a single chip that is good at both compute operations and memory operations.

GPU 在这方面比较平衡,但你本来必须在两者中选一个。

The GPU is quite balanced in this regard, but you had to choose one or the other.

Cerebras 在做矩阵乘法之类的运算时,内存访问速度非常快,特别适合在 Cerebras 芯片上托管 MLP,也就是那些权重。

Cerebras has a very fast memory access for something like a matrix multiply, and it's really good to host the MLP, the weights essentially on the Cerebras chip.

但 GPU 有能力扩展到非常长的上下文长度。

But the GPU has the capacity to scale to really long context lengths.

所以你可能想把注意力机制放在 GPU 上,把 MLP 放在 Cerebras 芯片上。

And so you would like to put the attention possibly on the GPU and the MLP on the Cerebras chip.

而且我相信这就是 NVIDIA 和 Groq 正在做的事情。

And I believe this is what's happening with NVIDIA and Groq.
M2
M232:25

你能花一分钟聊聊 Transformer 吗?

Can you riff for a minute just on transformers?

你对一些核心概念的讲解一直都很棒,所以想请你给那些对 2017 年这项创新不太熟悉的人讲讲,它的优势和劣势是什么,以及你认为它是否会继续主导 AI 的未来,或者说还会是未来的主流架构吗?

And yeah, you've been so good at explaining some of the core concepts just for people that, again, aren't deeply familiar with what this innovation was in 2017, like what its strengths and weaknesses are and whether or not you think it will remain the dominant architecture or a dominant architecture for the future of AI.
M1
M132:45

它的作用是让我们能够非常有效地在无监督数据上进行学习。

What it did was it allowed us to learn on unsupervised data really effectively.

因为 Transformer 归根结底就是,它能够处理任意序列——任意数据序列——并试图在其中找到模式。

Because transformers, what they're all about at the end of the day is taking any sequence, any arbitrary sequence of data and trying to find patterns in that data.

最关键的是注意力机制,这是 Transformer 的招牌组件,它让模型能够动态调整,聚焦于它认为序列中最相关的部分。

And they critically, the attention operation, which is the headline component of transformers, it allows the model to dynamically adapt to what it thinks is the most relevant component of the sequence.

每当你通过 Transformer 运行一步,你实际上都是在重新调整之前看过的输入的权重,然后决定哪些信息对你的下一个预测最有用。

Every step you take through a transformer, you are essentially reweighting the input that you looked at before and figuring out which is most relevant for your next prediction.

所以它非常适合学习任意的序列数据。

And so it's extremely amenable to learning arbitrary sequence data.

而我们日常产生的最有趣的序列数据就是语言。

And the most interesting sequences of data that we produce on a regular basis is language.

这就是为什么我们能在语言领域取得统治地位。

And that's how we got to dominance in the language regime.

但如果站得更远一点看,我认为 Transformer 真正做得好的一点是它们能够规模化。

But to zoom out even further, I think what transformers really did well is that they scaled.

Transformer 没有带任何人为的先验假设。

Transformers make no such human prior.

Transformer 只是说:嗯,数据序列里一定有某种模式。

Transformers just say, well, there's going to be a pattern in the sequence of data.

如果有模式,我就能找到它。

And if there is a pattern, I'm going to find it.

我会往这个问题上堆越来越多的参数,直到它工作为止。

I'm going to throw more and more parameters at this problem until it works.

Transformer 也从很多计算机视觉的工作中受益。

And transformers benefit from a lot of the computer vision work too.

比如计算机视觉的一个挑战是,我们很难从几十万个参数(就像支持向量机这种线性模型,或者其他传统机器学习模型所用的几千个参数)跳出来。

For example, one of the challenges in computer vision was we had a hard time going from hundreds of thousands of parameters, which you get for linear models like support vector machines or other legacy machine learning models.

那些模型只有几千个参数。

Those had on the thousands of parameters.

然后我们进入了深度学习时代,计算机视觉模型达到了几千万个参数。

Then we got to deep learning and got to tens of millions of parameters with computer vision.

当时最大的模型大概有 1.5 亿个参数,这在计算机视觉里已经算很大的模型了。

The biggest models were around Like 150 million parameters was a huge model for computer vision.

而现在我们经常讨论的是万亿级别的参数。

And now we routinely talk about trillions of parameters.

而 Transformer 正是从百万参数跨越到万亿参数的关键。

And transformers are the link to go from millions to trillions of parameters.
M2
M234:22

那么如果我认为规模化的关键单元是数据和算力,是不是有理由觉得 Transformer 会一直存在下去?毕竟我们现在擅长获取的就是这两样东西。

And so if I think about the important units of scaling being data and compute, does it stand to reason then that you think transformers will just stick around because that's the thing that we're good at getting more of, those 2 things?
M1
M134:32

嗯,数据还是个开放性问题,但算力肯定是。

Well, data is an open question, but compute, sure.

对,Transformer 就像——它们就像超级海绵。

Yeah, transformers are so— they're just such great sponges.

你知道,如果你给一个 Transformer 增加 10 倍的算力,它在某些方面就会对数级地提升。

You know, like you increase the compute, available to a transformer by 10x and you'll get some log improvement somewhere.

而且到目前为止,规模定律确实有效。

And so far the scaling laws really work.

它们真的非常漂亮。

They're really quite beautiful.

那么回到刚才那个点,我觉得,transformer 到底擅长做什么?

And to the point about, I guess, what do transformers do really well?

它们几乎能适用于你能给到的任何数据集。

They extend to almost any dataset you can throw at them.

它们是极其强大的通用学习器。

They're extremely powerful general learners.

我觉得 transformer 比其他我们试过想替代 attention 的技术更有用的是,transformer 能够表示你想要的任意一对关系。

And I think what's especially useful about transformers over other techniques that we've tried to replace attention is transformers represent Any pairwise relationship that you want.

序列中的任意 token 都可以关注到序列中的任意其他 token。

Any token in the sequence can attend to any other token in the sequence.

所以如果序列中哪怕存在任何关系,你用 transformer 都能找到它。

So if there's any relationship that's in the sequence at all, you're going to find it with transformer.

不过,有时候你可能并不需要全对全的建模。

Now, it may be the case that you don't need all-to-all modeling.

你不需要每个 token 都去看其他所有 token。

You don't need every token to look at every other token.

但如果你需要的话,transformer 给你提供了这个选项。

But if you need to, transformers give you that option.

而直到我们找到更好的方法来缩减这个空间,让信息建模更有选择性之前。

And until we know a better way to prune that space down, a better way to have information modeling be more selective.

Attention 是一个非常非常好的操作。

Attention is a very, very good operation.

这又是我们在计算机视觉时代学到的一个技巧。

This is another kind of trick that we learned in the computer vision days.

Karpathy 以前有句话是这么说的:如果你有一个新数据集要训练模型,你的首要目标应该是让模型参数过充裕,然后试着过拟合你手里的数据——这样能证明数据中存在你可以建模或记忆的关系,证明你的学习算法有效,证明你可以把知识注入到模型里。

One of the old Karpathy sayings is that if you have a new dataset that you want to train a model for, your first goal should be to overparameterize the model and try to overfit the data that you have to prove that there is a relationship that you can model or memorize, that your learning algorithm works, that you can instill knowledge into the model.

一旦你能过拟合,接下来就可以做压缩。

Once you can overfit, then you can compress.

而压缩就是你获得泛化能力的方法。

And the compression is how you get generalization.

你并不想真的死记硬背你面前的数据。

You don't want to actually memorize the data that you have in front of you.

你想做的是泛化。

You want to generalize.

所以,一旦你过拟合了数据集,就可以反过来工作,试着找到那些能用最少参数数量去适配的通用模式。

And therefore, once you overfit the dataset, then you can kind of work backwards and try to find the general patterns that fit into the smallest parameter count possible.
M2
M236:14

你对数据的未来有什么预测?顺便谈谈数据在整个故事里的重要性?

What's your prediction for the future of data and riff on the importance of data in this whole story?
M1
M136:18

我喜欢这个说法:互联网是一次性数据补贴。

I like the phrase that the internet was a one-time subsidy on data.

我们是免费拿到的。

We got it for free.

数据质量非常高,大约有30万亿 token 的高质量文本。

It's extremely high quality, about 30 trillion tokens of high quality text.

如果你对好文本的定义更宽泛一点,那就是300万亿 token。

300 trillion tokens if you take a wider view on what qualifies as good text.

基本上我们已经全部看过了。

And we've basically looked at it all already.

到目前为止,模型已经把整个互联网看了无数遍,在互联网人类数据上能做的事情已经不多了。

Models have seen the entire internet many times over at this point, and there is not a whole lot more to be done on human data from the internet.

在我看来,数据的下一个阶段基本上是模型通过强化学习环境(gym)来自我提升。

The next phase of data in my mind is model self-improvement through RL environment gyms, basically.

事实上,我们甚至不再从获取更多的随机用户与 AI 互动中受益。

In fact, we don't even benefit from getting more random user interactions with AI.

以前我们非常关心的一种新数据是用户使用 ChatGPT 时的互动数据,他们给 ChatGPT 提供喜欢或不喜欢的反馈信号。

It used to be that the new type of data that we cared about a lot was the interaction data from people using ChatGPT and giving ChatGPT signals on what they liked and didn't like.

我现在喜欢这样一个观点:我们提供的中位数模型已经远比一个随机的人类反馈者要先进,以至于从随机人类偏好——或者说人类无条件的偏好——得到的信号实际上已经没有任何价值了。

I like the argument now that the median model that we serve is so much more advanced than a random human giving feedback that the signal you get from random human preference, or I guess unconditioned human preference, is not actually worth anything anymore.

到这个时候,你需要专家的偏好反馈。

You want expert human preference at this point.

模型已经超出——

The model has out—
M2
M237:21

对。

Yeah.

普通人。

Everyday Joe.
M1
M137:23

对,普通人。

Yeah, Everyday Joe.

没错。

Exactly.

所以我对数据未来的看法是,给模型一个硬性的、可验证的任务,让它在一个像健身房一样相对隔离的环境里运行,面对一个可以不断推进的问题,并且能衡量它在这个问题上有没有进展。

So the future of data to me is giving the model a hard verifiable task and letting it run in this gym where it's kind of isolated and it just has a problem that it can make progress on and get measurement of whether it made progress on that problem or not.

你可以想象,编程问题就属于这一类。

You can imagine coding problems are in this category.

数学问题也属于这一类。

Math problems are also in this category.

而且越来越多的情况下,我们基本上可以给智能体一台电脑,让它像人类员工一样工作。

And increasingly more and more we can just give the agent a computer essentially and have it act like it's a human worker.

然后只需要告诉它,它离目标结果有没有推进。

And just give it feedback on whether it's making progress towards the target outcome.

那个环境本身就变成了数据。

That environment becomes the data.

我觉得这不算什么超级独特的见解,但从我目前看到的情况来说,效果真的非常非常好。

I think this is not a super differentiator to take, but it's been really, really productive from what I've seen so far.
M2
M238:01

那你觉得这个路子会持续很长一段时间,还是说,打个比方,如果我把互联网看作一个大区块,那这个就是另一个大区块,它会风光一阵子,等我们把它全都吃透,就得转向别的东西了?

And you think that just goes on for a really long period of time, or is that another, like, if I think about the internet as this one big block, like this is another big block that will have its day in the sun and we'll kind of get it all and then we'll have to move on to something else?
M1
M138:16

我觉得其实比那更深刻。

I think it's actually more profound than that.

基本上,这个想法是,如果你想要通用人工智能,最好的办法就是不断叠加专用智能,直到没有空白需要填补。

Basically, the idea is that if you want artificial general intelligence, the best way to get there is to just keep stacking specialized intelligences until you have no more gaps to fill.

这里的关键测试是,你要确保这个做法能奏效,唯一要做的就是让任务变得可验证。

And the test here, the only thing you need to make sure you do to make this work is you must make sure that your task is verifiable.

你得给模型一个自我评分系统。

You need to give the model a self-grading system.

如果你有这个,那你就有了在任何任务上自我改进的方法。

If you have that, you have the recipe for self-improvement on any task you like.

我觉得你们已经看到这一点在顶级实验室的投入上体现出来了。

I think you've seen this held up by the way Frontier Labs spend.

他们以前花那么多钱在数据上。

They used to spend that much on data.

现在他们花更多钱在强化学习环境上了。

Now they spend a lot more on RL environments.

而这些环境正好抓住了在可验证任务上进行递归自我改进的这种关系。

And these environments absolutely capture that relationship of recursive self-improvement on a verifiable task.
M2
M238:55

好。

Okay.

嗯,我很高兴我们跑偏了,聊了些小话题,但回到你最初的目标——通过软件更好地控制硬件层面,让现有硬件更高效。

Now, so I like that we've veered off in different little sidecars here, but coming back to your initial task of making existing hardware more efficient by being more in control of what's going on at the hardware level through software.

那你就继续讲讲你已经做了什么、想做什么,然后我们再跳到硬件,最后再聊能源。

So yeah, just keep going on what you've done so far and what you want to do, and then we're going to jump to hardware and then jump to energy finally.
M1
M139:19

没问题。

Sounds good.

嗯,我提到了内核。

So yeah, I mentioned kernels.

这挺让人惊讶的。

It's surprising.

大家都觉得内核已经搞定了。

People think kernels are done.

有一些很厉害的人,比如Tri Dao,他们写了非常优秀的kernel,这些kernel构成了现代深度学习的基石,我们所有的现代深度学习都是建立在flash attention上的。

There are great people like Tri Dao who write excellent kernels and they form the bedrock of all of our modern deep learning is built on flash attention.

现代的transformer都是建立在flash attention上的。

Modern transformers are built on flash attention.

但是如果你稍微偏离了这条顺风顺水的路径,就会出现一个新模型,它嵌入位置信息的方式略有不同,比如RoPE系统的变化。

But if you deviate from the happy path at all, there's a new model that comes out that has a slightly different way to embed positional information, like the change of the RoPE system.

突然之间,我们已有的那个kernel就不适用于这个新模型了。

Suddenly the kernel that we had is not suitable for this new model.

然后我们可能就得给这个kernel打一个补丁。

And we may have to make a patch to this kernel.

我不敢说我们到了必须从头发明新kernel的阶段,但有能力快速修改现有的GPU kernel——抱歉,是kernel——顺便说一句,kernel是一个通用术语,指你在GPU上运行的任何程序。

I wouldn't say we're in the phase where we have to invent new kernels from scratch, but having the ability to quickly modify existing GPU kernel, sorry, a kernel, and by the way, it's a general term for any program you run on the GPU.

所以从历史上看,kernel往往会被放进一个库里面,每个kernel都有非常非常明确的用途。

And so historically kernels tend to be put into a library where every kernel has a very, very scoped purpose.

通常你会有一个做矩阵乘法的kernel。

Typically you have a kernel for a matrix multiply.

你还有另一个kernel,甚至是为了像加法这么简单的事情。

You have another kernel for even something as simple as addition.

你想把两个tensor加在一起,那又是另一个kernel。

You want to add 2 tensors together, that's another kernel.

然后渐渐地,我们开始把这些kernel融合在一起。

And then increasingly we've started to fuse those kernels together.

所以如果我做一个矩阵乘法,然后想把它加到另一个也同样乘过的矩阵上,也许这两个就可以合并成一个kernel,然后我把这些操作融合起来,这样就不用为了做加法而把数据写到DRAM里再读回来,也许我就能轻松实现。

So if I do a matrix multiply and then I want to add it to another matrix that I've also multiplied, maybe those 2 become one kernel and I just fuse the operations where instead of writing the data out to DRAM and then reading it back in just to do the addition, maybe I can just do this easily.
M2
M240:33

为什么现在还是人类在做这件事?

Why are humans still doing this?

这看起来像是AI会特别擅长的事情——设计更高效的kernel。

It seems like the sort of thing that AIs would be exceptionally good at engineering, more efficient kernels.

也许这正是我们要走的方向,只是我们还没真正走到那一步。

Maybe that's where we're going and we're just not quite there yet.

但如果我们还没走到那一步,那这真的是我们要走的方向吗?

But if we aren't there yet, is that where we're going?

如果我们还没到那一步,那为什么还是人类在做这件事?

If we're not there yet, why are humans still doing this?

为什么Treehouse这么有名?

Why is Treehouse so well known?

这个名字我听说过。

It's a name I know.
M1
M140:51

我不想替Tree说话,但他教给我的是,你不一定非要再手写kernel了。

I don't want to speak for Tree, but what he taught me was you shouldn't write kernels by hand anymore necessarily.

我喜欢说我们是在白板上写kernel。

I like to say we write kernels on the whiteboard.

我们走到白板前,描述我们认为机器应该做什么。

We go to the whiteboard, we describe what we think the machine should be doing.

然后我们把这个用自然语言简洁地描述给一个模型。

Then we succinctly describe that in natural language to a model.

然后模型就能够执行,它会说,好的,这是我的输入和输出。

And then the model is able to do the execution of, okay, here's my input and output.

这就是我们想把工作分配到GPU上的策略。

Here is the strategy of how we want to dispatch this work onto the GPU.

我要去把这个工作实现出来。

I'm going to go implement this work.
M2
M241:17

所以我们是做概念设计。

So we're doing the conceptual design.
M1
M141:18

没错。

Exactly.

而且我觉得,我不太确定为什么模型做这个并不出色。

And that I think I'm not sure exactly why models are not superb at doing this.

我不认为这是我们的护城河或者类似的东西。

I don't think this is like our moat or anything like that.

我敢说六个月后,我们在内核工程上会有更好的模型。

I'm sure in 6 months' time we'll have much better models on kernel engineering.

而且我敢说,那些实验室会告诉你,他们已经用全自动的方式做了很多内核工程。

And I'm sure the labs would tell you that they already do a lot of their kernel engineering in a fully automated way.
M2
M241:34

所以软件作为一种优势——如果我把软件理解为在接近光速效率下使用底层硬件——从长远来看,它不会成为你们公司这样的优势。

And so software as an edge, if I think about software as maximally near speed of light efficient usage of an underlying piece of hardware, is going to trend towards not being an advantage for a company like yours over time.
M1
M141:47

我觉得没错。

I think that's right.

像 Mythos 或者 GPT-5.6 SOL 这样的浪潮,能带动所有船只。

The rising tide of something like Mythos or GPT-5.6 SOL, that lifts all boats.

确实如此。

It really does.

我其实不认为有必要专门说我们只把模型做得更好用于内核工程。

I actually don't think there's a point in specializing to say we work on making the model better for just kernel engineering.

我觉得这实际上不是编程中最有意义的一个子集,特别是内核工程。

I think that's actually not the most meaningful subset of coding in general, kernel engineering in particular.

也许你可以往提示里注入一些特权信息,这样就能引导模型更好地编写内核。

Maybe there's some privileged information you inject into the prompt that's a useful way to steer the model to be better at writing kernels.

但总的来说,是的,就这种能力而言,我们都依赖前沿进展。

But broadly speaking, yes, we're all downstream of the frontier in terms of this capability.
M2
M242:22

我一直很喜欢能源史上的一个现象。

I always love this from the history of energy.

总有一个钟摆在摆动:一边是原始资源,比如煤;另一边是,一块煤里有多少能量能被我们利用,利用率是多少。

There's always this pendulum between the raw source, let's say coal, and then if there's a certain amount of energy available inside of a hunk of coal, what percent of it we can harness and use.

能源史的一个重要部分,就是把那个数字从10%提高到95%左右。

And a big part of the history of energy was getting that number from 10% to 95% or whatever.

我们处在什么位置?

Where are we in that?

就像,假如把 Blackwell 或类似的东西看作一块煤,你觉得我们现在利用率大概是多少?

Like, if I just think about a Blackwell or something and Blackwell is the piece of coal, like, what percent do you think we're at?

我的意思是,今天对现有硬件的利用效率有多高?

Like, how efficiently can we use an existing piece today?
M1
M142:55

有很多不同的方式来分析这个问题。

There's a lot of different ways to analyze that.

我觉得在某些层面,当GPU执行它最擅长的任务——也就是大维度的矩阵乘法时,我们在优化性能上做得非常高效。

I think in some level we are really efficient at optimizing the performance when the GPU is doing the thing that it's most happy doing, which is a large dimension matrix multiply.

那个运算能达到70%到80%的峰值利用率,限制因素不是软件,而是功耗。

That operation runs at 70-80% of peak utilization and it's limited not by software, but by power.

NVIDIA 标称的峰值FLOPS有点乐观。

The way NVIDIA quotes peak FLOPS is a little optimistic.

因为功耗限制,你永远达不到那个数。

You never hit that because of power throttling.

因为发热。

Because of heat.

对,没错。

Yeah, exactly.

热管理。

Thermals.

就算70%到80%吧。

Let's say 70-80%.

已经饱和了。

It's saturated.

那已经很不错了。

That's pretty good.

但实际上,在Transformer中,大部分时间你并不处于那个愉快的路径——也就是做大批量的矩阵乘法。

But in practice, you don't spend the majority of your time in a transformer in that happy path where you're doing a large batch matrix multiply.

所以我们的工作本质上就是构建芯片周围的环境,确保GPU始终有大批量任务可以处理。

And so our job is to basically build the engine around the chip such that we are feeding the GPU these large batches of work at all times.
M2
M243:39

嗯。

Yeah.
M1
M143:40

过去一年里,GPU领域最深刻的一个转变就是,你不再是一次只编程一个GPU了。

And one of the most profound transitions we've had in the GPU world in the last year has been this moving of, you know, you don't program one GPU at a time anymore.

你得考虑整个机架,甚至要考虑整个集群,整个数据中心。

You should think about the whole rack and maybe you should think about the whole cluster, the entire data center at a time.

而NVIDIA,再次,他们已经开始发货的不再只是单个GPU或单个主板,而是整个机架系统,这是他们推荐的方式。

And with NVIDIA, again, they've started shipping not just a single GPU or a single motherboard, but actually the whole rack system is something that they prescribe.

他们叫它NVL72。

They call it NVL72.

他们最新的芯片Grace Blackwell 300,是以一个72台的机架形式出货的。

Their latest chip, the Grace Blackwell 300, that ships as a rack of 72 units.

现在大家都在竞争,看谁能以最高效的方式编程整个机架规模的计算机。

And it's an open race to figure out who can program the whole rack-scale computer as efficiently as possible.

而我相信,这种计算形态是未来效率和速度的方向。

And my belief is that that shape of compute is the future of both efficiency and speed.

实际上,NVIDIA做得很好,如果你想要最低的延迟,你就应该用这个芯片。

In fact, NVIDIA does a great job of, if you want the lowest possible latency, you should be using that chip.

如果你想要最高的吞吐量,你大概也应该用这个芯片。

And if you want the highest possible throughput, you should probably also be using that chip.

至少目前是这样。

As of right now.

这一切都归结为,这是一种非常新的编程范式。

And it all comes down to like, this is a very new paradigm of programming.
M2
M244:32

我听说的一种说法是,现在最好的芯片,比如Blackwell,它的市场就像毒品市场一样。

One of the things you hear is that the market for the best chips, Blackwells, let's say, is like a drug market or something right now.

人们为了尽可能多地搞到它们,各种疯狂的事情都在发生,因为实在太短缺了。

Like there's all sorts of fascinating things happening to get as many of them as possible because everyone's so short.

我很想听听你对这个比喻的看法。

I'd love you to react to that analogy.

实际情况真的是这样吗?

Like, is that what it feels like?
M1
M144:47

百分之百是这样。

100%.
M2
M244:48

那你能也说说那些不是最顶尖芯片的市场情况吗?

But then also to talk about what the market is like for like not the bleeding edge chips.

比如我如果愿意接受一个稍微差一点或者中等差一点的芯片。

Like if I, willing to accept a slightly or moderately inferior chip.

那个市场是什么样的?

What's that market like?

给我们讲讲那个世界吧。

Like, let us into that world.
M1
M145:00

好。

Yeah.

行。

Okay.

说几点。

A couple of things.

第一,是的,基本上NVIDIA对他们所有芯片都有长期考量。

Number one, yes, basically Nvidia has a long-term view on all their chips.

他们看到对Blackwell芯片的巨大需求,本可以像其他供应商过去那样,直接涨价来满足市场,供需曲线自然会调整。

They see this immense demand for the Blackwell chips and they could do what other suppliers have done in the past, which is like just crank prices and meet the market, you know, supply and demand curves will correct.

它们会在某个点相交,最终大家从技术上来说都会更满意。

They'll intersect at some point and everyone will be technically happier.

但NVIDIA认为,如果让资金最雄厚的买家把芯片都买走了,长期来看可能对自己不利,因为那个客户最终可能会积累更大的权力。

But Nvidia sees the, if they just let the most deep pockets buy all the chips, that maybe hurts them in the long term if that customer ends up accruing a lot more power.

他们明白,现在的计算能力就是权力。

They understand that compute is power today.

所以他们在分配算力时非常具有战略眼光。

And so they're quite strategic about how they allocate compute.

这是我的第一个想法。

That's the first thought.

第二个想法是,人际关系非常重要。

The second thought is that relationships matter a lot.

没人愿意看到一个新创公司下了一个巨大的芯片租赁订单,说什么‘哦,是的,我要租一万块Blackwells,租三到五年。’

Nobody wants to have a huge order of a chip rental come in from this new startup that says, oh yeah, I'm going to rent 10,000 Blackwells for 3 years or 5 years.

这个新创公司通常才运营了几个月。

The startup has only been operating for months typically.

谁知道他们值不值这个钱?

Who knows that they're good for the money?

如今,要说服别人给你提供算力是相当有挑战的,要么需要非常好的人际关系,要么就得有惊人的财务支持,才能在Nvidia那边实现。

The way you convince someone to give you access to compute is quite challenging these days and requires some pretty either great relationships or just incredible financial backing to make this happen on the Nvidia side.

这都是因为稀缺性太高,而需求又高得离谱。

And it's all because the scarcity is so high and demand is just off the charts.

至于其他芯片,我甚至不会说它们更差。

Now for other chips, I wouldn't even call them inferior.

我喜欢说,没有不好的芯片,只有不好的定价。

I like to say there's no bad chips, there's only bad pricing.

只要价格合适,我就能让任何芯片发挥作用。

And I will make any chip work at the right price.

这有点像公司的精神之一。

That's like kind of one of the ethos of the company.

我们来聊聊AMD。

Let's talk about AMD.

对,AMD,我认为总体上是很好的芯片。

Yeah, AMD, I think great chips overall.

挑战在于人们不太明白怎么很好地编写它们的程序。

The challenge is that people don't understand how to program them very well.

所以,你知道,我之前跟你聊过,我们有一个非常棒的kernel团队。

So, you know, I've been talking to you about how we have such a great kernel team.

我们非常认真地压榨硬件的性能。

We're so serious about squeezing the performance out of the hardware.

坦率地说,Nvidia在自己的芯片上做得相当好。

Nvidia is pretty good at doing that for their own chips, frankly.

我们能压榨出一些额外性能,但实际上在其他芯片上还有更多可做的,因为供应商在开箱即用地提供最佳kernels方面做得比Nvidia少一些。

There's some alpha that we can squeeze out, but actually there's a lot more to be done on other chips because the vendor does a little bit less work than Nvidia does to make the best kernels out of the box.

或者对我来说更好的是,有额外性能可榨取,而且别人都认为AMD不如Nvidia。

Or even better for me, there is alpha and just like other people have this perception that AMD is not as good as Nvidia.

这对我来说真是好消息。

That's music to my ears.

我很高兴他们不重视这种芯片,让我能尽可能多地买。

I'm very happy for them to sleep on this chip and for me to buy as much as I can.

现在,我觉得这个说法其实不再那么准确了。

Now, I think that that's not actually super true anymore.

我认为AMD在一些大买家中间其实还挺受欢迎的。

I think AMD is actually somewhat popular amongst some large buyers.

我知道公开的,Meta和OpenAI买了大量AMD芯片。

I think publicly Meta and OpenAI have bought a ton of AMD chips.
M3
M347:10

有意思。

Interesting.
M1
M147:10

所以我们越来越多地看到,所有的AMD供应也在被分配,但还有一大批其他公司正在涌现。

And so we're increasingly seeing that all the AMD supply is also being allocated, but there's a long tail of other companies that are popping up.

不过真正的新公司很不错,比如Etched、SambaNova或者D-Matrix。

Yet net new companies are great, such as Etched or SambaNova or D-Matrix.

所有这些公司都在涌现,我认为它们的主要挑战是规模。

All these companies are popping up, and I think the main challenge for them is scale.

他们真的能从台积电拿到足够的晶圆配额来生产芯片并推向市场吗?

Can they actually get enough wafer allocation from TSMC to pump out chips to make it into the market?

但当然,如果市场上有新的芯片,我想尽快知道,并评估我们是否能够购买。

But certainly if there's a new chip on the market, I'd like to know about it as quickly as possible and evaluate whether we can buy.

相当一部分的供应。

A good fraction of that supply.
M2
M247:41

所以对你来说,这本质上是一种套利。

And so it's fundamentally an arbitrage for you.

如果你能更擅长从那些不太受关注的芯片中榨出性能,你就可以转手赚取差价,这可以成为一个很好的生意。

Like if you can be much better at eking out performance from chips that have received less attention, you can then resell that at a margin and it can be a great business.
M1
M147:52

没错。

Exactly.

没错。

Exactly.

而且我觉得,并不是说其他人有技术难题,没法把这些芯片用得好。

And I think that it's not the case that everyone else is just, you know, has a skill issue that they can't, you know, make these chips work as well.

我们在这方面挺有能力的。

I think we're quite competent at this.

我们可能是世界上最好的团队之一,能够使用多种架构的芯片,并在那些意想不到的地方积极追求性能提升。

I think we're probably one of the best teams in the world to use multiple silicon architectures and be quite aggressive in chasing down performance in unlikely places.

但我觉得,关键在于我们愿意围绕新芯片快速构建软件栈。

But yeah, I think it's the speed at which we're willing to kind of build our stack around a new chip.

我们并没有太多历史包袱,比如数据中心供应商只能固定使用某类芯片,部署新芯片会很麻烦。

We don't have a huge amount of incumbency around, well, our data center providers are only stuck with this class of chip and it's gonna be a huge pain for us to deploy these net new chips.

我们有一些非常有创意的数据中心合作伙伴,他们愿意快速行动,而且还有一类新的合作伙伴可以介绍。

We have some very creative data center partners who are willing to move very quickly and there's a new class of those that we can talk about.

最重要的是,我们不怕挑战。

And most importantly, we don't shy away from the challenge.

坦率地说,这很大程度上就是愿意说,是的,我们喜欢TPU。

That's frankly a big part of this is just saying, Yes, we love TPUs.

我们会让TPU发挥作用。

We're going to make TPUs work.

是的,我们喜欢Trainium。

Yes, we love Trainium.

我们会让Trainium发挥作用。

We're going to make Trainium work.

如果它不能很容易地运作,我们也会想办法,通过异构服务系统来适配。

And if it doesn't work easily, we're going to find a way to fit it in with a heterogeneous serving system.

它总会有用武之地的。

It will have a place.

每种芯片都有它的比较优势。

Every chip has a comparative advantage.

我们得找到那个优势,然后从那个方向去榨取潜力。

We had to find that advantage and then squeeze it in that direction.
M2
M248:52

在我们进入硬件、数据中心、能源这些更精彩的话题之前,我想先聊聊你对投资者群体忧虑的看法。

Just as an interlude before we get to hardware, data centers, energy, etc., which will be a really fun part of the conversation, I'd love you to talk about your perception of the investor class's worry.

比如,你看看存储器股票。

Like, you look at memory stocks.
M1
M149:06

是的。

Yes.
M2
M249:07

或者我最近喜欢看一张图,显示半导体在标普500指数中的占比。

Or my current favorite is you look at the chart that plots the percent of the S&P 500 that's semiconductors.

历史上,大概是2%、3%、4%。

Historically, it was like 2%, 3%, 4%.

现在到了19%、20%、21%。

Now it's 19%, 20%, 21%.

看起来,如果你研究市场历史,总会看到一些东西在某段时间冲到疯狂的高点,然后又跌回到长期平均水平。

And it just sort of looks like if you're a student of market history, you get all these things through time that have sort of reached some crazy near-term peak and then collapsed back to long-term norms.

我很好奇,这怎么让所有投资者都感到担忧的。

And I'm curious how that has all investors worried.
M1
M149:30

是啊。

Yeah.
M2
M249:31

所以,很多人从Micron和SK Hynix这类公司赚了很多钱。

So, you know, a lot of people made a lot of money in Micron and SK Hynix and companies like this.

但每个人都觉得,长期来看,计算是一种商品,它不可能占到全世界总市值的四分之一或五分之一。

But everyone feels like, oh, these, you know, in the long term, like compute's a commodity and it will not represent a quarter or fifth of the entire market capitalization of the world.

所以他们很害怕。

And so they're scared.

这就是那个背景。

And that's the setup.

每个人都承认,嗯,存在巨大的短缺,但大家又都觉得,哎,我们会搞定的,这些东西会回到资本市场里它们正常的位置。

Everyone acknowledges that, like, there's a huge shortage, but everyone sort of feels like, ah, we'll figure it out and these things will revert back down to their normal place in capital markets.

我很好奇你怎么看这种说法。

I'm curious what you think about that narrative.
M1
M150:02

有一点,我不太算是历史的学生,更像是历史的一部分。

One thing, I'm less of a student of history and more of a member of history.

我1997年出生,我妈妈在Intel工作,正好赶上2000年前夕,还有互联网泡沫的兴起和崩盘。

I was born in 1997 and my mom worked at Intel in the run-up to the year 2000 and the dot-com boom and crash.

我记得当时Cisco是全球最有价值的公司,Intel也紧随其后。

And I remember the time where Cisco was the most valuable company in the world and Intel was close behind.

我主要是把25年前那段历史和今天做对比。

I mean, I mostly draw parallels to that period of history from 25 years ago to today.

我觉得主要区别在于,历史上很多网络设备的投资都是投机性的。

And I think the main difference is that a lot of the investment in networking equipment historically was speculative.

我们预期了未来用户的需求,但那些用户从未出现。

We anticipated this future demand for users that never came.

我觉得token消耗,或者说广义上的AI消耗,有趣的地方在于它不再是投机性的了。

And I think what's interesting about token consumption or AI consumption broadly is that it's no longer speculative.

人们购买token是因为它们对当下就有价值。

People buy tokens because they're immediately valuable to them.

你不会囤积token,你会立刻使用它们。

You don't hoard tokens, you use them immediately.

这也和我们两年前的情况很不一样,2023、2024年的时候,Hopper系列的芯片出现了供应紧张。

This is also even different from what we had 2 years ago where there was a supply crunch for Hopper generation chips in 2023, 2024.

那个时期,支出都是面向训练的,而训练本质上就是投机性的。

In that period, it was all training-oriented spend, and training is inherently speculative.

现在则是——每个人都在对自己在云计算上的支出设置上限。

Now it's— everyone is instituting caps on how much you can spend on cloud compute.

这是一个非常、非常不同的世界,我们现在谈论的是推理支出,并且预测推理支出会增长。

It's a very, very different world to be talking about inference spend and predicting inference spend to go up.

我确实认为推理支出会单调递增。

I do think inference spend monotonically increases.

推理支出没有任何投机成分。

There's no speculation on inference spend.
M2
M252:44

现在回到你对硬件的看法。

Coming back now to your take on hardware.

我对单元级别很感兴趣。

And so the unit level is interesting to me.

比如聊聊芯片、机架、集群。

Like talk about chips, talk about racks, talk about clusters.

我很想聊一聊数据中心。

I'd love to talk about data centers.

你说过你有一些很有意思的合作伙伴在做很酷的事情。

And you said you've had some interesting partners doing some cool things.

跟我们讲讲你眼中数据中心的现在和未来吧。

Talk to us about the present and future of data centers as you see it.
M1
M153:04

嗯。

Yep.
M2
M253:05

因为这显然对于能够服务所有这些推理需求来说至关重要,这个领域需要大量的创新。

Because this seems like obviously a critical thing for being able to serve all this inference is like lots of innovation in this part of the world.

而且显然你也专注于这方面。

And obviously you're focused on it.
M1
M153:12

我觉得我们对话中反复出现的一个主题就是:什么是训练?什么是推理?

I think one of the themes in our conversation has come back to what is training versus inference?

这两类之间的区别到底是什么?

Like what is the difference between these 2 categories?

那么,两年前以训练为主,而现在以推理为主,这有什么不同呢?

And what was different about 2 years ago being training-oriented and today being inference-oriented?

我觉得整个AI产业链中最保守的玩家,肯定是基础设施提供商,不管是数据中心还是更保守的,像TSMC这样的芯片基础设施。

And I think the most conservative players in the entire AI stack have got to be the infra players, whether that's data centers or even more conservative is TSMC, the chip infra people.

以前的数据中心,比如AI数据中心,都是为训练而建的。

And so data centers historically were built like AI data centers, they were built for training.

而训练是一种包含推理的超集负载。

And training is a superset workload over inference.

你可以把任何训练集群拿来做推理,但反过来就不一定行。

You can make any training cluster work for inference, but maybe not vice versa.

这里的区别在于网络。

And the difference there is networking.

你在芯片之间的带宽上投入多少,你需要多大的集群?

How much do you invest in bandwidth between chips and how large of a cluster do you need?

实际上,数据中心在某种程度上存在规模不经济。

There's actually a diseconomy of scale to data centers in some way.

比如,在一个数据中心里建10万张GPU,比建1万张要贵得多、难得多,比建1000张更是如此。

Like it's way more expensive and difficult to build 100,000 GPUs in one data center than it is to build 10,000, than it is to build 1,000.

现在我们只讨论你有多少兆瓦或吉瓦。

And now we just talk about how many megawatts or gigawatts do you have?

基本上,在美国要建一个吉瓦级别的数据中心,现在已经不容易了。

And basically there's no way to build a gigawatt data center in the United States easily anymore.

就连100兆瓦都越来越难。

Even 100 megawatts is increasingly hard.

除非你是非常特殊的客户,否则基本不可能。

It's basically impossible unless you're a very special set of customers.

10兆瓦大概在今天还勉强可行。

10 Megawatts is probably on the edge of what's possible today.

而1兆瓦,我觉得是充足的。

And 1 megawatt, I would argue, is plentiful.

所以市场上有大量可用的总电力,但问题是它们不集中。

So there's this incredible floor on the market where you can find lots of aggregate power, but it will not be concentrated.

这对那些建训练数据中心的人来说没什么吸引力,因为他们默认——

And that was not interesting to anyone who's building training data centers because you just assume—.
M2
M254:36

必须全部集中在一个地方。

It needs to all be in one spot.
M1
M154:36

必须全部集中在一个地方。

It needs to all be in one spot.

没人想跨数据中心做训练。

No one wants to deal with cross-data center training.

所以市场有些滞后。

So the market has some lag in it.

我觉得市场仍然认为,我们需要到处去找那些100兆瓦和10兆瓦的数据中心。

I think that the market still assumes that we have to go shake down those 100-megawatt and 10-megawatt data centers wherever we can find them.

我仍然从很多数据中心开发商那里听到这种态度。

It's still the attitude I hear from a lot of data center developers.

但越来越多地,我们看到一些新的思考者意识到,推理更适合这些分布式的一兆瓦数据中心。

But increasingly we're seeing a few new thinkers realize that inference is gonna be suitable for these distributed 1-megawatt data centers.

我们非常同意这一点。

And we're quite in agreement with that.

我们很乐意在美国各地购买小规模的计算资源,作为我们的推理集群。

And we are very happy to buy small pools of compute across the United States and use that as our inference fleet.
M2
M255:08

给我们讲讲,1兆瓦和10兆瓦在物理尺寸上大概是什么概念?

Give us a sense of literal physical size of 1 megawatt versus 10.
M1
M155:12

嗯,随着液冷技术的出现,这个问题变得有点专业了。

Yeah, well, so this got really wonky with the advent of liquid cooling.

现在你可以把极高的功率密度塞进一个物理机架里。

Now you can pack insane levels of power density into a single physical rack.

比如一兆瓦的计算能力,你可能会想象一个巨大的数据机房,就像一个巨大的仓库。

Like a megawatt of compute, you'd imagine this like massive data hall, like a huge warehouse basically.

现在你实际上可以把那压缩到大摡8个机架的算力。

And now you can actually pack that into around like 8 racks worth of compute.

每个机架大概就一个冰箱那么大。

Each rack is about the size of a refrigerator.

你可以想象一下,8个并排摆在那里的场景。

You can just imagine 8 of them lined up.
M2
M255:38

嗯。

Yeah.

那大概是一兆瓦。

That's a megawatt.

所以你的想法是,你希望帮助构建的未来,就是一大堆不同的芯片可以被一起使用。

And so your view would be that the future that you want to help build is a whole bunch of different chips that can be used together.
M3
M355:46

是的。

Yes.
M2
M255:46

你可以去买,作为买家,你会想方设法从每块芯片上榨出最多的性能。

That you can buy, you know, you're a buyer to eke out the most per chip.

然后这些芯片可以被组合在非常小的数据中心里,专门做推理。

And that those chips can then be coupled in very small data centers to just do inference.

而这两步——一大堆五花八门的计算资源,其中有些比它本应便宜的价格还要便宜,再加上你从中能榨出更多性能的能力,然后在一个数据中心里用小型单元来部署——就等于更便宜的智能。

And that those 2 steps of a whole bunch of random compute, some of which is cheaper than it should be, your ability to eke more out of it, and then small units of expression in a data center equals way cheaper intelligence.
M1
M156:12

我当然这么认为。

I certainly think so.

是的,有很多途径可以获得更便宜的FLOPS。如果你能对你拿到的资源灵活处理,那么我描述我们做的事情时,会这么说:我们会买世界上任何地方、任何时长的任何芯片。

Yes, there's a lot of ways to access cheaper FLOPS. If you're able to be creative with what you take, and so one of the ways that I describe what we do is we will buy any chip anywhere in the world for any duration of time.

这种灵活性和流动性,我认为目前没有其他人能做到。

That is a level of flexibility and liquidity that I think no one else has right now.

我们非常积极地把钱花在我们自己相信的事情上,我们真的会接受任何算力,然后想办法在我们的集群里把它用起来。

We're very aggressive about putting our money where our mouth is, and we will really take any capacity and find a way to make it work in our fleet.

这也是我们今天优势的一大部分。

And that is a big part of our advantage today.

长期来看,我们必须通过投资这些别人会持怀疑态度的数据中心来创造更多优势,因为,你知道,当你建立一支由1000个小型数据中心组成的队伍,而不是一个大型的千兆瓦数据中心,会发生什么呢?

And long term, we have to create more of that advantage by investing in these data centers that other people are gonna be skeptical of because, you know, what's gonna happen when you set up this army of 1,000 small data centers versus the one big gigawatt data center?

嗯,有几件事。

Well, few things.

你不会有电力冗余,Frank,很多时候是这样。

You're not gonna have power redundancy, Frank, quite often.

你现场不会有后备柴油发电机。

You're not gonna have backup diesel generators on site.

那些都非常贵。

Those are all very expensive.

我们砍掉了所有这些开销。

We cut all that overhead.

很多情况下,我们甚至不会有冗余的网络。

We're not even gonna have redundant networking in a lot of cases.

我们会把这些设施放在有良好电力接入的地方,一个单一的电力来源,然后我们会挖一条光纤线路。

We're gonna put these in facilities where we have good access to power, a single source of power, and we're gonna trench one line of fiber.

连接到这些数据中心,但我们不会有3条光纤线路,然后冗余、故障切换和SLA这些。

To these data centers, but we're not gonna have like 3 lines of fiber with redundancy and failover and SLAs.

它有时候就是会断掉。

It's just gonna go down sometimes.

实际上,如果其中一些的可用性只有95%,我也不会惊讶。

In fact, I won't be surprised if some of them get down to like 95% uptime.
M2
M257:19

那挺糟糕的。

Which is bad.
M1
M157:20

非常糟糕。

Very bad.

对行业里其他任何人来说,那是致命的,灾难性的。

That's fatal, atrocious for anyone else in the space.
M2
M257:22

在大型的共享经济里根本活不下去。

Couldn't survive in a big gig economy.
M1
M157:25

一个只有95%正常运作时间的数据中心基本是没买家要的。

You'd have basically zero buyers for a data center that has 95% uptime.

我就是那个第一个买家。

I'm that first buyer.

我愿意买95%正常运作时间的服务。

I will buy 95% uptime.
M2
M257:32

这是因为我们这个后台引擎的特性。

And the reason for that is because of this background engine thing.

如果后台有任务在跑,你们就不在意了,对吗?

That if there's things running in the background, you don't care?
M1
M157:37

不完全对。

Partially.

其实有两个原因。

It's actually 2 things.

第一,我们有一个非常稳健的控制层,只要故障不是多个数据中心同时出问题,单个数据中心里的任何一次故障它都能处理好,我可以直接把工作负载挪到别的地方去,那我就完全没问题。

One is that we have a really robust control plane that is gonna be fine handling any single failure in any single data center, as long as it's not correlated with other data centers and I can just move the workload somewhere else, I'm cool with that.

故障的发生是有一定频率的,而我对一个95%正常运作时间的数据中心、98%或99%的数据中心的满意度基本上是线性的。

The failures happen at some rate and I am basically linearly happy with a data center that's 95% uptime versus 98% versus 99%.

对我而言,就是线性地好或差。

It's just linearly good or bad for me.

不过,确实需要我之前提到的那个异步机制,因为我们服务的是那些长期运行的代理。因为一旦请求失败,我就得去找一个新的GPU来重新放这个请求。

Now, you do need that async piece that I mentioned of we serve these long horizon agents because what happens when a request fails is that I'm going to have to go find a new GPU to put that request on.

这就意味着在那个代理工作的单一轮次里,它本来要运行一个小时,结果因为GPU被拿走了而遇到了障碍。

And that means that for that single turn of the agent's work, it's working for an hour, but then it hits a roadblock because its GPU got pulled away.

在那个时刻,那个代理可能要多等一两分钟、三分钟,甚至十分钟的延迟。

In that moment in time, that agent is going to experience maybe an extra minute or 2 or 3, maybe even 10 of latency.

但我的理由是,我的客户并不在乎,因为他们的代理已经运行了好几个小时。

But my argument is that my customers don't care because their agent was running for hours.
M2
M258:32

他们都在睡觉呢。

They're sleeping.
M1
M158:33

那没关系。

It doesn't matter.

偶尔单个轮次变长一点,根本无所谓。

It doesn't matter if like a single turn occasionally becomes a little bit longer.

所以我们告诉客户:听着,我们的平均吞吐量会非常有竞争力,但我们的P99、也就是99百分位的延迟,我们没法控制。

So we tell our customers, look, our average throughput is gonna be very competitive, but our P99, our 99th percentile latency, it's not gonna be controlled.

这没办法做到。

It cannot be.

但作为交换,我会给你无可匹敌的经济效益。

And in return, I'll give you unbeatable economics.

我认为这非常适合后台代理的场景。

And I think that's the right fit for background agents.
M2
M258:52

聊聊电力这个类别吧。

Talk about power as a category.

你看到了什么有趣或创新的东西?

What are you seeing that's interesting, innovative?

你觉得这个领域会怎么发展?

Where do you think this goes?
M1
M158:57

好,我刚才说我想要数据中心有95%的正常运作时间。

Okay, so I said I want 95% uptime on my data centers.

如果价格合适,我甚至能接受80%的正常运作时间吗?

Could I even take 80% uptime at the right price?

可能可以。

Probably.

那这意味着什么呢?

And what does that mean?

嗯,我是加州人。

Well, I'm a son of California.

我爱太阳能和风能。

I love solar and wind.

我觉得太阳能和风能在美国还没有被充分利用起来。

I think solar and wind power is way undertapped in the United States.

而挑战一直在于这种间歇性。

And the challenge has always been this intermittency.

你甚至会觉得太阳能和风能不适合用于数据中心,因为你需要持续的基载电力,而它们却是间歇性的能源。

You would even consider solar and wind unsuitable for data centers because you have a persistent base load and an intermittent power source.

那你打算怎么办呢?

What are you going to do?

嗯,我觉得我们其实离解决这个问题并不远了。

Well, I think that we're actually not that far from solving that problem.

我完全能够容忍数据中心停电几天甚至几周,这听起来像是数据中心噩梦般的场景——因为没风、云层遮蔽、山谷里雾气弥漫,导致长时间停电。

I am totally capable of tolerating an outage from a data center that's measured in even days or weeks, which is like the worst case nightmare scenario for a data center is that we're going to have a long-term outage because the wind doesn't blow and the clouds are in the sky, fog is hanging over the valley for some time.

那就是最坏的情况。

That's worst case scenario.

但实际上这高度可预测,我只需要在发生这种情况时,从世界其他地方调用算力就行。

It's in fact highly predictable and I can just call in capacity in some other place in the world whenever that happens.

我只需要模拟天气,就能知道我的数据中心什么时候会离线。

I'll just model the weather and figure out when my data centers are going to be offline.

把工作负载转移到别处,就没事了。

Move my workload somewhere else and it's fine.

关键在于,这能让我获得别人不愿意碰的电力资源,因为这种停电处理起来太麻烦了。

The trick is that it's gonna give me better access to power that no one else is gonna touch because it is so annoying to deal with that kind of outage.

而且如果我的芯片够便宜,它们很可能不是 Nvidia 的机架。

And if my chips are cheap enough, they're probably not gonna be Nvidia racks.

而且如果我的芯片够便宜,我也不介意闲置芯片的资本成本。

And if my chips are cheap enough, I don't mind the capital cost of having idle chips.
M2
M260:09

对。

Yeah.

我听说你把这整个系统描述成一种拾荒策略。

So I've heard you describe this entire system as like a scavenger strategy.
M1
M160:13

是的,没错。

Yeah, that's right.
M2
M260:14

那你能稍微展开讲讲这个比喻吗?

Is that, yeah, unpack that analogy a little bit.
M1
M160:16

嗯,首先我们拾荒芯片,然后为这些芯片拾荒电力。

Well, first we scavenge chips, and then we scavenge power for those chips.

思路是,在这两种情况下,我都不想跟 Anthropic 或 OpenAI 竞价争夺算力。

The idea is in both cases, I do not want to be bidding against Anthropic or OpenAI for compute capacity.

我赢不了他们,我也不想赢。

I'm not going to win against them and I don't want to.

我想更有创意一点,用他们目前还看不上眼的供给。

I want to be more creative and use the supply that they don't find legible today.

随着时间的推移,我积累了足够的总供给量。

And over time, I amass enough aggregate supply.

我永远得不到集中的供给。

I'm never going to get concentrated supply.

我只能得到分散的总供给。

I will only get aggregate supply.

随着时间的推移,我建造了在经济学上无法匹敌的分散式工厂。

And over time, I build my aggregate factory that is unbeatable in economics.

我们正在建造一座工厂。

We are building a factory.

我们试图建造世界上最好的钢铁厂,但它是通过小型钢厂来完成的,而不是大型一体化钢厂。

We're trying to build the best steel factory in the world, but it will come through mini mills, not through large monolithic steel plants.
M2
M260:52

如果我设想不同的版本,比如你可以垂直整合到什么程度,一个极端版本就是你拥有所有东西,这样就成了一个资本非常密集的业务。

And if I imagine the different versions of this, like how vertically integrated you can be, one version would be, the extreme would be you own everything so that it's a very capital-intensive business.

你拥有电力来源,建造数据中心,设计自己的芯片,掌控能最大化芯片效能的软件,然后把最终的 token 卖给用户。

You own the power source, you build the data centers, you design your own chips, control the software that ekes the most out of those chips, and you sell the end finished token to your user.

比如,你的用户就是我,而你拥有整个堆栈。

Like, your user is me, and you just own the whole stack.
M1
M161:19

嗯。

Yeah.
M2
M261:20

但你可以想象这个业务还有很多其他变体,你知道的,你可以随便划条线,想怎么划都行。

But you can imagine many other permutations of the business where, you know, you could, whatever, you draw the line anywhere.

你也可以做得非常轻资产,什么都不拥有,就充当所有这些事情的协调层,像个虚拟的拾荒者。

You could be incredibly capital light, own nothing, and just be like the coordination plane across all this stuff, the virtual scavenger.
M1
M161:34

对。

Right.
M2
M261:35

你怎么看待这个问题,就是该选择哪一类业务?

How do you think about that question of like which type of these businesses to be?
M1
M161:39

你知道吗,我内心其实有两部分在回应这个问题。

You know, there's actually 2 parts of me to receive that question.

一部分是公司的CEO,每天都需要工作,要尽可能可持续、尽可能快地增长。

One is the CEO of a company that needs to work every day and grow as sustainable and as quickly as it possibly can.

另一部分是创始人,创始人更有想象力,就是热爱这些东西。

The other is a founder, and the founder is much more imaginative and just loves this stuff.

我内心的创始人想什么都做。

The founder in me wants to do everything.

这就是我的全部人生。

This is my entire life.

我这辈子都在思考芯片、电力、能源。

I spent my entire life thinking about chips, power, energy.

我关心的就只是这些东西。

Like all I care about is this stuff.

所以我当然想做到最大程度的雄心勃勃。

So of course I want to be maximally ambitious.

我永远不想停下来。

I don't want to stop ever.

我永远不会停下来,直到我构建出一个从头到尾最高效的系统。

I will never stop until I have built the most efficient system from soup to nuts.
M2
M262:08

你基本上就是在做现实版的Factorio。

You're doing real-life Factorio, basically.
M1
M162:10

确实如此。

Very much so.
M2
M262:11

为了智能。

For intelligence.
M1
M162:11

确实如此。

Very much so.

所以这就是从感性角度给出的硬核答案。

So that's like the emotional from the hard answer.

从CEO的角度,我觉得我们必须更务实一些。

On the CEO side, I think we have to be more pragmatic.

我觉得要拥有所有东西所需投入的资金,就像你说的,简直是疯狂。

I think that the capital we're looking at for owning everything is, like you said, it's insane.

对,软件的杠杆很高,所以我们得从软件入手。

Yeah, software has high leverage, so we have to start with software.

但最终,我们是自己发电,还是能和公用事业公司签下好的购电协议?

But ultimately, do we own power generation or can we get great power purchase agreements with utilities.

我更倾向于让别人专注于他们一直擅长的事情,然后看看我们能不能达到这个规模。

I'm more inclined to pursue letting other people specialize in the things that they're historically good at and then see if we can get to this scale.

我把它看作这样:我想达到一个规模,让我有资格把这件事纳入我们的羽翼之下。

I think of it as like, I want to get to the scale where I earn the right to take this under our wing.

我绝对认为,在整个技术栈中到处都有效率提升的空间,只要你能打破一个假设——我要从他们那里购买的人,他们对客户是谁有自己的假设,而我可能打破了这些假设。

I absolutely think that there's efficiencies to be gained everywhere in the stack if you can break the assumption that people I would be buying from, they made assumptions about who their customers would be and I maybe break those assumptions.

这是一个相当乐观的看法。

It's a pretty optimistic view.

我觉得这之所以可能,是因为我们实际上是在尝试为计算史上最大的计算市场提供资金支持。

I think it's only possible because we're actually trying to underwrite the largest market for compute in the history of computing.

我们实际上将要投入数十亿、数万亿美元的投资到推理上。

We are actually going to build so many billions, trillions of dollars of investment into inference.

因为有了这个专注点,我们就有理由去构建很多专门为推理定制的组件。

And because of that focus, it makes sense to build a lot of things that are custom for inference.

而我的工作就是去找所有可能做这种定制的地方。

And it's my job to seek all the places where that's possible.

然后,当这些地方对我来说以及我的合作伙伴变得很明显时,我就会让他们为我打造定制化的东西。

And then as they become obvious to me and my partners, I will get my partners to build custom things for me.

如果他们做不了,我就自己来。

And if they can't do it for me, I will do it myself.
M2
M263:27

如果你要宏观地审视整个系统——软件、硬件、能源等等——然后排序一下,你认为目前我们在生成有用的智能token方面,哪些环节效率最低?

If you had to just zoom out on this entire system, software, hardware, energy, et cetera, and stack rank the places that you think that we are the most inefficient today at producing useful intelligent tokens.

这个排序长什么样?

What does that list look like?
M1
M163:42

我认为算力扩展其实非常高效,意思是给我更多的FLOPS,我就会用掉更多的FLOPS。

I think compute scaling is actually very efficient, as in you give me more FLOPS and I will use more FLOPS.

而且我觉得我们目前在FLOPS的使用上已经相当精打细算了。

And I would say we're actually fairly judicious already with our use of FLOPS.

如果你看看现代的MoE模型,很少有模型密度超过10%,也就是说,只有10%的专家被激活了。

If you look at a modern MoE model, there are very few models that are more than 10% dense, meaning 10% of the possible number of experts you can activate are activated.

而且我觉得前沿模型更接近1%。

And I think the frontier models are closer to 1%.

所以已经相当稀疏了。

So fairly sparse already.

我不认为我们在MoE这方面浪费太多。

I don't think that we're wasting too much on the MoE side.

人们研究MoE已经有一段时间了。

People have been working with MoEs for quite some time.

他们在压榨MoE性能方面已经很擅长了。

They're pretty good at squeezing MoEs.

我们做得不好的地方是注意力机制及其对内存的使用。

Where we are not good is attention and its use of memory specifically.

目前的KV缓存压缩程度很低。

The KV cache is quite uncompressed right now.

我觉得如果你看看KV缓存中的熵,还差得远。

I think if you look at the entropy in the KV cache, it's nowhere near.

它物不所值。

It's not earning its keep.

每个token我们会在KV缓存中存储好几KB的数据。

We're storing many kilobytes of data in the KV cache per token.

这个量级可能差了一两个数量级。

And that's probably off by an order of magnitude or 2.

我不知道前沿实验室在做什么,但DeepSeek确实发表了一些非常有趣的工作来进一步压缩它,而且他们进展不错。

And I don't know what the Frontier Labs do, but DeepSeek certainly publishes really interesting work to compress that further and further and they're making good progress.

我认为他们每年能在这方面实现一个数量级的进步,这本身就说明还有很多提升空间。

And I think the fact that they're able to make an order of magnitude in progress here every year or so signals that there's a lot more room to go.

我猜这些都是微观层面上的。

I guess this is all on the micro scale.

如果把视野拉得更宽,我觉得我们实际上根本没有有效地调度我们的算力。

If you zoom out further, I think that we actually don't marshal our compute effectively at all.

比如说,全球有这么多算力。

Like we have all this compute in the world.

英伟达今年要出货500万块Blackwell芯片。

Nvidia is pumping out 5 million Blackwell chips this year.

它们都去哪了?

Where are they all going?

它们是否始终都在被使用?

Are they all being used at all, all the time?

我当然表示怀疑。

I certainly doubt it.

我认为在某种程度上,我们只是需要更好的全球算力编排。

I think that at some level we just need better orchestration of compute across the world.

这非常难做到,因为大量算力流入了私有算力池,永远不会被公开使用。

This is very difficult to do because a lot of compute disappears into private pools of compute that will never see the light of day.

而那些GPU只能闲置着,非常可惜。

And those GPUs sit very sadly idle.

看到那些GPU,明明投入了硅片和电力,却只是闲置着,我心里真的很难受。

It pains me physically to see that those GPUs are just, silicon and power went into that and it's just sitting idle.

我想改变这个状况。

And I want to fix that.

我们如何把全球算力组织协调成共享资源,并更高效地利用它。

How we organize and orchestrate the world's compute as a shared resource and pack it more efficiently.

我估计,我们都会嘲笑xAI在集群的FLOP利用率上有些问题,但现实是,世界其他地方的状况要糟糕得多。

I would estimate that we all make fun of xAI for having some challenges with total FLOP utilization on its clusters, but the reality for the rest of the world is it's far worse.

大量的GPU要么堆在仓库里,要么被分配给特定客户的私有池,根本用不上。

A ton of GPUs just sit in warehouses or sit in private pools allocated to a specific customer, just don't get utilized.
M2
M265:47

你直接针对这个效率问题下手。

You're attacking the efficiency of that very directly.
M1
M165:49

这样有效得多。

That's way more effective.
M2
M265:50

对。

Yeah.

那晶圆厂呢?

What about fabs?

就是说,你觉得晶圆厂本身未来的发展会怎样?

Like, what do you think is the future of fabs themselves?

我想大家都在想,那些存储公司、台积电、英特尔和其他厂商,它们到底能不能——它们要如何扩大产能呢?

Like, I think everyone is wondering, will the memory companies, will TSMC, will Intel and others be able to— how will they expand capacity basically?

我们会在美国这里建厂吗?

Will we do it here in the US?

对,就是聊一聊芯片制造本身。

Yeah, riff on fabrication of chips themselves.

比如,如果我们打个响指就能让今天的芯片库存变成一百倍,那token价格大概会便宜很多。

Like if we could just snap our fingers and have 100 times the chips in the stock today, we'd probably have way cheaper tokens.

所以,这似乎是挺重要的一个方面,想听听你的看法。

So yeah, that seems like an important part of the universe to hear your view on.
M1
M166:25

嗯,这挺有意思的。

Yeah, well, it's interesting.

所有东西都是相互平衡增长的,对吧?

Everything grows in balance with each other, right?

如果我们打个响指,把所有这些东西都翻倍,你可能会解决台积电的瓶颈,但马上又会遇到另一个瓶颈。

If we snap our fingers and double all those things, you might fix a TSMC bottleneck, but you're just going to run into another bottleneck.

你多生产了20%的芯片,立刻就会有另一个瓶颈出现。

You make 20% more chips, then you have another bottleneck immediately.

不过,有意思的是,他们觉得什么是必须交付的,比如他们把什么当作一个不变的要求,认为他们的客户,也就是我,一定会想要,而我却觉得这更像是一种灵活的关系。

I will say though, it is interesting what they consider to be a must deliver, like what they consider to be like an invariant that their customers, me, are always going to want versus what I think of as like a more fluid relationship.

我觉得,如果晶圆厂能向我展示更多它们的权衡取舍,我就能更明智地决定自己能做什么。

I think that if the fab exposes more of their trade-offs to me, I'm able to make more intelligent decisions about what I think I can do.

最有趣的例子之一就是,任何晶圆厂生产出来的芯片,最差的和最好的之间差距很大。

One of the most interesting examples here is that any fab has a lot of spread in their worst chip that comes out of the production line and the best chip that comes out of the production line.

芯片制造过程中存在很大的差异性。

There's a lot of variance in how chips are made.

那么问题来了,像台积电这样的公司,它们会非常努力地去收紧我们所说的工艺角。

And then the question is, if you have a company like TSMC, they work very, very hard to tighten what we call these process corners.

我们希望最差的芯片在特性上尽可能接近最好的芯片。

We want to keep the worst chip as close in characterization to the best chip.

然后它们会尽最大努力来实现这一点。

And then they go to great lengths to make that possible.

但这意味着他们在工艺中增加了许多我可能根本不需要的控制步骤。

But that means that they are adding a lot of controls in the process that maybe I don't need.

也许我其实愿意为那颗最差的芯片找到一个用武之地。

Maybe I'm actually willing to find a place for that worst chip.

你不需要把过程控制收得太紧,因为那样更花时间和成本。

You don't need to tighten the process control as much, which takes more time and cost.

也许我愿意接受更多的不良品。

Maybe I'm willing to take a lot more rejects.

对我们来说,这是一种更全面的优化,涉及芯片的成本、芯片的供应,还有电力成本以及我们能放置它们的地方。

And I think for us, it's like a more holistic optimization around there's, you know, cost of the dies, supply of the dies, and then the cost of power and places where we can put them.

我的整个目标其实是大幅扩大全美国的电力供应,这样很多原本没资格在数据中心里占一席之地的芯片,也能有地方安放。

And my whole goal is to actually so dramatically expand the supply of power across the United States that I have a home for a lot of chips that otherwise would not have earned their place in a data center.
M2
M267:59

我们能聊聊你如何设计自己公司的体系吗?

Can we talk about how you design the system of your own business?

你学到了哪些教训?

What lessons have you learned?

你谈到了一些关于英伟达的有趣教训,但带我了解一下你的文化,以及你如何构建一个以这个为北极星的团队和业务。

You talked about some of the interesting Nvidia lessons, but bring me into the culture and how you structure a team and a business where this is the North Star.
M1
M168:11

我觉得这里面有很多“极限思维”。

I think there's a lot of in-the-limit thinking.

我们不会担心刚开始做一个模型时效率不高这种眼前的问题。

We don't worry about the immediate nature of when we start working on a model, the efficiency is not going to be very good.

但我们会想一个月、六个月或一年后我们能达到什么程度。

But we think about where we could end up in a month or 6 months or a year's time.

我们不认为我们所用的机器状态是固定不变的。

We don't accept the state of the machines we work on as fixed.

即使是像 Blackwell 这样的芯片,如果我们认为存在某个瓶颈阻碍我们达到这种性能,对我来说,非常重要的是我们要很好地理解和描述它。

Even something like the Blackwell chip, if we think that there's some bottleneck that is holding us back from achieving this performance, I mean, it's very important to me that we understand and characterize that very well.

然后把它写下来,这样我们一方面可以告诉我们的朋友 NVIDIA,另一方面也能在将来购买芯片时记住这些要点。

And write it down so we can both, A, tell NVIDIA about it, who are friends, and also to basically keep this in mind for future chips that we buy.

我们希望学到那些我们认为对个人或公司长期来说基本不变的东西,并把它们融入到未来的决策中。

We want to learn things that are what we think are essentially invariant for us or the company long-term and fold that into future decisions that we make.

我们非常注重协作。

We're very collaborative.

我认为我们寻找的最重要的特质之一,是那种既是好学生又是好老师的人。

I think one of the most important traits that we look for are people who either, who are both good students and great teachers.

我们团队里很多人大学时当过助教,很喜欢这种分享知识的方式。

A lot of our people on the team were TAs in college and loved the experience of sharing knowledge in this way.

我们经常做白板讨论,我认为这种每个人都有可教的东西、也有可学的东西的合作环境对我们极其重要。

We do whiteboard sessions all the time, and I think the collegial environment where everyone has something to teach and something to learn is extremely important for us.
M2
M269:25

你希望招聘的人具备哪些特质,让你觉得他们能在三年后、当更多事情由机器处理时依然适应这种工作环境?

What are the attributes of people that you would want to hire that you think will be resilient to the work environment 3 years from now when more stuff is handled by machines?
M1
M169:35

好奇心。

Curiosity.

百分之百是好奇心。

It's 100% curiosity.

我唯一没法教的是对性能的热爱,那种喜欢钻研机器运行的每一微秒、理解当时机器上发生着什么的热爱。

The one thing I cannot teach is love for performance, love for digging into every microsecond that the machine is working and understanding what's happening on the machine at that time.

对我来说,这是性能工程师最重要的特质,也是我寻找的东西。

That to me is the most important trait for a performance engineer and it's what I look for.

我不看重大量的 AI 经验。

I don't look for lots of AI experience.

我完全不看重 CUDA 经验。

I don't look for CUDA experience at all.

那其实是个很大的误导。

That's actually a huge red herring.

我是说,CUDA 或 GPU 作为一个概念,在过去五年里已经演变了很多。

I mean, CUDA as a concept or GPUs as a concept have evolved so much in the last 5 years.

要求十年经验毫无意义。

There's no point asking for 10 years of experience.

我想教这些,但我没法教对性能工程的热爱。

I want to teach that, but I cannot teach the love for performance engineering.

那才是我追寻的。

That is what I seek.
M2
M270:10

你能不能逐个评估一下主要的大模型实验室,同时再讲讲闭源作为一个类别和开源之间的关系,以及你觉得现在正在发生什么、未来会发生什么?

Can you give your assessment of the major labs one by one, but also then the relationship of closed source as a category to open source and what you think is happening and will happen?
M1
M170:21

简单来说,我觉得各大实验室为了比其他人领先三到六个月,付出了极高的代价。

In a line, I would say the labs pay an immense premium to be 3 to 6 months ahead of everything else.

而且我觉得这很可能还是值得的。

And I think that's probably still worth it.

我认为OpenAI和Anthropic做他们正在做的事情,是完全合理的。

I think it makes perfect sense for OpenAI and Anthropic to do what they do.

有一个关于蒸馏的敏感话题,我觉得这是闭源和开源前沿模型之间关系的一个非常核心的部分。

There's a sensitive topic around distillation, which I think is a very core piece of the relationship between closed and open frontier.

对此我想提出一个不同的观点——有一种看法认为蒸馏就是偷窃,如果你在前沿模型的输出上进行蒸馏,就等于从它们那里拿走了东西。

And I'd like to offer an alternative view on that, which is there is the sense that distillation is theft, that you are taking something from the frontier models when you distill on their outputs.

但实际上,即使你没有这个意图,即使你从来不去爬取Anthropic的数据,我想指出一点:我们在互联网上发布的成果中,AI生成的内容占比越来越大。

And in fact, even if that's not your intent, even if you don't ever try to scrape data from Anthropic, one thing I'll offer is that an increasingly large percentage of the artifacts we put out on the internet, are AI-generated.

单看GitHub,你觉得过去一年创建的代码仓库里,有多少是Claude Code写的?

Even if you just look at GitHub alone, what percentage of repos created in the last year do we think were created by Claude code?

你觉得这算是蒸馏吗?

Do we consider that to be distillation?

因为可能就用这么多就够了。

Because that's probably all we need.

如果只拿GitHub上开源且你认为质量好的代码的输出作为训练数据,我完全不会惊讶你能训练出一个Fable级别的模型。

I would not be surprised if you could train a Fable-class model only on the outputs of code you consider good on GitHub that's open source.

当然,如果我们认为用户拥有他们与AI交互后产出的内容,并且他们选择把这些内容放到GitHub上——很多人确实这么做了——那长期来看隐性蒸馏就不可避免。

And certainly if we take the position that users own the outputs of their interaction with AI and they choose to put that up on GitHub, which a lot of them do, we're going to have latent distillation for a long time.

我觉得这从根本上来说是无法阻止的。

It seems fundamentally impossible for me.

我认为从根上说,没有办法阻止信息或模型能力的传播。

I don't think it's fundamentally possible to prevent the diffusion of information or model capabilities.

它一定会发生。

It will happen.

问题只是速度有多快。

The question is just how fast.
M2
M271:40

那么问题就变成了:规模定律和进步定律会不会永远成立,或者至少持续很长一段时间?

And so then the question becomes, do scaling and improvement laws hold forever or for a really long period of time?

如果它们成立,那领先三到六个月就有价值,并且这种优势能持续多久就会持续多久。

And if they do, then there's value to being 3 to 6 months ahead and that will just last as long as it lasts.

相比非常便宜的的开源token,他们可以为自己的token收取很高的溢价。

And they can charge a huge premium for those tokens relative to a very cheap open source token.

这么想对吗?

Is that the right way to think about it?
M1
M171:58

我觉得有可能。

I think it's possible.

但我不确定领先三到六个月的这个溢价能持续那么久。

I don't know that the premium for being 3 to 6 months ahead is going to last that long.

你看那些企业级部署,它们可不会按三到六个月的节奏来更新。

I mean, if you look at like enterprise deployments, they don't move at 3 to 6 month speed.

很多企业很可能还在用4.6、Opus 4.6或者Opus 4.7。

A lot of enterprises are probably still on like 4.6, Opus 4.6 or Opus 4.7.

它们不会很快采用最前沿的东西。

They don't adopt the bleeding edge rapidly.

在做任何变更的时候,人们都会有各种顾虑。

There's a lot of questions that people have around rolling out any change at all.

而且我觉得我们现在只是刚开始触及皮毛,远远不到能说谁已经赢了这场竞赛的时候。

And I think we're just so early in scratching the surface that I don't think there's any way to call a winner in this race.

我当然也不认为这是一场有终点的竞赛。

And certainly I don't even think this is a race that can be decided ever.

它永远是一个持续的过程。

There's always, it's a continual process.

从根本上说,我觉得开源永远不会消失。

And fundamentally, I don't think open source ever goes away.

如果有一个领导者退出,新的领导者就会进来。

If there's a vacuum because one leader steps out, a new leader will step in.

驱动力太强了,而且还有很多顺风因素。

There's too much incentive and too much, there's a lot of tailwinds too.

训练一个 frontier-class 模型变得越来越容易了。

It's just, it gets easier every day to train a frontier-class model.
M2
M272:44

那你对未来有什么期望?

And so your hope of what the future looks like is what?

比如封闭和开放之间怎么平衡?

Like what balance between closed and open?

模型公司因为拥有整个技术栈的优势而包揽一切,这之间怎么平衡?

What balance between model companies doing everything because they have the advantage of owning the stack or whatever?

Anthropic 能做到。

Anthropic can do that.

它就像新的谷歌。

It's like the new Google.

谷歌也会这么干或者类似的事情。

Google will just do that or something.

你希望未来是什么样?

What do you hope the future looks like?
M1
M173:01

我想要大量的 token 和多样化的使用方式。

I want abundant tokens and diverse harnesses.

我希望每个人都能构建自己的使用方式。

I want everyone to build their own harness.
M2
M273:07

每个公司。

Every company.
M1
M173:07

每个公司,甚至每个用户。

Every company, every user even.

让 agent 成为你自己的。

Make the agent your own.

我觉得我们离那种程度的定制化和能力已经不远了。

I think we're not that far away from that level of customization and capability.

我希望人们拥有自己的智能,并且我希望那种智能是定制的,可能不是通过 weight fine-tuning,而是通过更多的 in-context learning。

I want people to own their intelligence and I want that intelligence to be customized, probably not through weight fine-tuning, but probably through more in-context learning.

这是一个更技术性的细节,但实现这种富足未来的基础基本上就是廉价的 token。

That's a more technical detail, but the underlying input to this abundant future is basically cheap tokens.

我的工作就是尽可能让 token 变得便宜。

My job is to make the tokens as cheap as humanly possible.

我会做到这一点,而且我会通过我所能触及的每一层技术栈来实现。

I will achieve that and I will do it through every layer in the stack available to me.

我喜欢供给侧杠杆。

I love the supply-side levers.

我会用上每一种芯片。

I will use every chip.

我会用上每一种电力来源,以及美国每一块适合做这件事的土地。

I'll use every source of power and I will use every piece of land in the United States that's suitable for this.

作为回报,人们会有动力去探索拥有富足智能是什么感觉。

And in return, people will have the incentive to explore what it's like to have abundant intelligence.

我们现在仍然把 agent 当作一个咨询很贵的人,只有遇到难题的时候才去请教。

We still treat the agent as a person that is expensive to consult and you should ask them when you have a hard question.

这不是看待智能的正确方式。

That's not the way to think about intelligence.

机器能够思考,这太不可思议了,我们应该努力让尽可能多的人都能用上它。

It's incredible that the machine can think and we should try to get that into as many hands of as many people as possible.
M2
M274:04

你坐在一个如此独特的位置,对于如何实现这个未来有着独特的视角。

You sit in such a unique seat and you have such a unique perspective on what you're trying to do to make this future a reality.

你觉得你跟那些消息很灵通、对这个领域很感兴趣的朋友们相比,最不一样的观点是什么?

What do you think are your most divergent views of the world versus your friends who are really well-informed and interested in this stuff?

你的哪些想法会让朋友们觉得你像个外星人?

What ideas of yours make your friends look at you like you have 3 heads?
M1
M174:21

大部分想法是关于芯片的,我觉得。

Most of the ideas on chips, I would say.

当我谈到搭建定制芯片时,他们问我,那有什么不同?

When I talk about building custom chips and they ask me, so what's different?

基本上,就是避开HBM短缺的问题,专注于更极端地把数据转移到其他形式的存储器上,比如flash。

Basically, it's about sidestepping the HBM shortage and focusing on more extreme offload to other forms of memory, such as flash.

我对这个想法非常热衷。

I'm quite passionate about that idea.

我们团队所有人都知道,我一直在强调我们必须改变模型架构,以便更大规模地把KV缓存转移到flash上。

Everyone on my team knows that I keep banging the drum around what we have to change about the model architecture to make offloading KV cache to flash work at a much greater level.

我一直在白板上画这个。

And I'm whiteboarding that all the time.

这在推理社区里是众所周知的。

That's in the community of inference people.

对于如果你围绕每秒1到10个token来设计一个系统能做什么,我们有一些不同的看法,那是我们的北极星目标。

We have some divergent views on what you can do if you design a system around serving at 1 to 10 tokens per second, which is our whole North Star.

更广泛地说,我觉得,有一个更大的问题:你该怎么做?

More broadly, I think, there is this larger sense around what do you do?

人们如何每天消耗一万亿个token?

How do people consume a trillion tokens per day?

那就是我们想创造的世界,让他们有能力做到这一点。

That's the world we want to create, the capability for them to do that.
M2
M275:11

一万亿个token是什么概念?

What's a trillion tokens?

给我们一个具体的量级概念。

Ground us in how much that is.
M1
M175:13

一万亿个token。

A trillion tokens.

好吧。

Well, okay.

按OpenAI的定价,那至少是500万美元,至少,对于5.5或5.6来说。

At OpenAI pricing, that's at least $5 million at the very least for 5.5 or 5.6.

对,我觉得用美元来衡量可能是最直观的方式。

Yeah, I think the dollars is probably the most relevant way to look at it.
M2
M275:22

对,这是一种看待指标的方式。

Yeah, it's a way to look at metric.
M1
M175:23

对。

Yeah.

那是几百万美元。

It's millions of dollars.
M2
M275:24

对。

Yeah.

所以,我们每天每人消耗目前价值500万美元的东西,那会是一个什么样的世界?

So what's the world in which we consume what currently costs $5 million per person per day?
M1
M175:29

对,我的意思是,我们要求每个token的成本至少提高3到6个数量级。

Yeah, I mean, we were asking for at least 3 to 6 orders of magnitude improvement in cost per token.

如果降到5000美元,你大概就会有客户了。

Get that into $5,000, you probably have some customers.

事实上,我认为对于某些规模的模型来说,一万亿个token的成本已经接近几万美元了。

And in fact, I would argue that for some size of model, we are approaching a trillion tokens being measured in tens of thousands of dollars.

那是你可以想象为单个任务运行的东西。

And that's something that you could imagine running for a single job.
M2
M275:51

你有没有担心过,普通人就是不能也不会这样做?

Are you at all worried that just like the average person just can't and won't do that?

就像他们现在不会用自己的大脑那样做。

Like doesn't do that now with their own brain.

就像世界上实际上没有那么多的智力需求。

Like there actually isn't that much demand for intelligence in the world.
M1
M176:01

我永远不会相信这一点。

I never will believe in that.

世界上对智能的需求永远存在。

There's always demand for intelligence in the world.

我认为,作为产品社区,我们面临的挑战是如何提供这种智能的入口和途径。

I think that the way and the on-ramps to that intelligence are our challenge as a product, you know, community.

我不是做产品的,所以我不敢说我有最好的愿景。

I'm not a product person, so I cannot say I have the best vision.
M2
M276:13

对,你想赋能那些人。

Yeah, you want to enable those people.
M1
M176:14

但我想赋能那些人。

But I want to enable those people.

我希望他们永远不会因为觉得“我的免费用户用不了”或者“我付不起这么多token”而被束缚。

I want them to never be held back by the sense that, oh, my free tier users cannot use, or I can't afford to give them this many tokens.

我经常从客户那里听到这种话。

I hear that from my customers all the time.

我们想解决这个问题。

We want to fix that.
M2
M276:25

那反过来呢?

What about the inverse question?

不是问你什么最疯狂,而是问你什么大家公认的东西其实是错的?

Not what you think is craziest, but like what consensus thing you think is wrong?
M1
M176:31

我一直反复思考的一个问题就是NVIDIA。

One of the things I keep coming back to is this question of NVIDIA.

短期内我看好NVIDIA。

I am bullish on NVIDIA in the short term.

而且NVIDIA,你永远不该跟他们对着干。

And NVIDIA, you should never bet against them.

他们总能自我革新。

They're always going to reinvent themselves.

但本质上,有件事会让人们惊讶,那就是如果你看Hopper到Blackwell再到Rubin,然后做同类对比,比如bfloat16乘法的每瓦性能,其实提升并没有那么大。

But fundamentally, I think one thing that surprises people is when I tell them that, hey, if you look at Hopper to Blackwell to Rubin, and you compare like for like, like what is the performance per watt of a bfloat16 multiply?

提升并没有那么大。

It hasn't improved all that much.

再往前推一步,看台积电。

Or even you take that one step further, go to TSMC.

如果你看台积电的5纳米、4纳米、3纳米和2纳米,这些芯片的每瓦性能变化并不大。

If you look at TSMC 5 nanometer versus 4 versus 3 versus 2, the performance per watt on these chips doesn't change like a dramatic amount.

所以后果就是人们在地缘政治上很焦虑,比如如果我们因某种原因失去台积电的供应会怎样。

So the consequence of this is people lose their minds over geopolitics, like what would happen if we lost access to TSMC for any reason?

而我的反直觉看法是,其实没那么糟。

And my contrarian take is that it wouldn't be that bad.

供应肯定会受到冲击,但西方最好的制程,比如Intel,其实差距没那么大,最差也就是每瓦性能差两倍左右。

Supply would take a shock for sure, but the best processes that we have in the West, like Intel, not that far behind, at worst, like maybe 2x worse performance per watt.

如果你跟着芯片战争的那些说法走,你会觉得差距很大,但实际上差距要小得多。

And the gap is just far smaller than you would make it out to be if you follow like the chip war dialogue.
M2
M277:32

在AI领域,还有什么事情不是你正在做的路径上的,也就是说不是你这个系统里需要去做的部分,但最让你感兴趣?

What else is happening in the AI world that is not in your path, meaning it's not like a component of this whole system that you would end up doing something in, that interests you most?
M1
M177:42

嗯,我们是完全在下游做模型服务。

Well, we're fully downstream of models.

所以模型设计者们可以决定他们的架构怎么设计。

So the model people get to decide how to design their architectures.

我对OpenAI或Anthropic只有非常微弱的影响,甚至可以说没有,我只能祈祷他们往对我有利的方向走。

And I have only very light, I mean, I don't have any input to OpenAI or Anthropic, but I can only pray that they go in the direction that is amenable to me.

或者我必须尽力预测他们可能的方向,然后据此构建我的推理架构,包括软件和硬件选择。

Or I have to do my best to predict where I think they're going to go and build my serving architecture accordingly, both software and hardware choices.

我觉得他们在某种意义上玩的是最有趣的游戏。

They have, I think, the most interesting game in some ways to play.

这又回到了机器思考的深刻性,以及决定使用稀疏注意力还是密集注意力有多重要,或者使用不同的数据类型有多重要。

Once again, this is getting back to the profundity of the machine thinking and how consequential it is to decide to use something like sparse attention versus dense attention or how consequential it is to use a different data type.

我们之前用bfloat16训练,但现在可以用FP8或FP4,或者更低精度的数据类型。

We were training in bfloat16, but now we can train in FP8 or FP4 or lower precision data types.

这感觉就是个随意的选择,但它对我能用什么芯片、该怎么造硬件、以及怎么思考计算的未来,都有深远的影响。

That is just an arbitrary choice, it feels like, but it has profound implications for what chips I can use and how I should build my hardware and think about the future of compute.
M2
M278:35

如果一百个创业者聚在一个房间里,每个人都想搞个新的计算初创公司。

If you had 100 entrepreneurs in a room, all of whom wanted to create some new compute startup.
M1
M178:40

嗯。

Yeah.
M2
M278:41

假设他们特别想做硬件芯片、系统、机架之类的。

And let's say they specifically wanted to make hardware chips or systems or racks or whatever.

你会给他们什么建议,关于怎么定位他们的公司,或者公司类型,而不是具体的技术赌注?

What advice would you give them on how to orient their companies or the type of company, not the specific choice they're making on a tech bet or something like this?

因为看起来我们什么都会试一遍,这对世界是好事。

'Cause it seems like we're gonna try everything and that will be great for the world.

有些东西会成功。

Some stuff will work.

但如果你必须给他们建议,关于怎么定位他们的业务才能在这个即将到来的世界里成功,你会说什么?

But if you had to give them advice on how to orient their business to be successful in this coming world, What advice would you give them?
M1
M179:04

关键全在供应链的瓶颈上。

It's all about the bottlenecks on the supply chain.

所以你得先说服我,或者说服投资人,你明白现代芯片供应中那三到五个核心瓶颈。

So you need to first convince me or convince an investor that you understand the like 3 to 5 bottlenecks that dictate modern chip supply.

有台积电的晶圆产能、HBM 的产能,还有先进封装。

There's TSMC wafer capacity, there's HBM capacity, and there's like advanced packaging.

也许第四个是电力。

And maybe a 4th one would be power.

比如说,你从哪搞到电?

Like, where will you get the power?

你怎么搭这些机架?

How will you build these racks?

我想听到的是,你对这四个瓶颈都得有很好的答案。

And I want to hear like, you should have a great answer to each of those 4 bottlenecks.

以及你怎么绕过它们,因为说到底这都是套利。

And how you're going to work around them because it's all arbitrage at the end of the day.

你造芯片是因为你觉得英伟达做了一些他们很难改的选择,这确实没错。

You're building a chip because you think that Nvidia has made some choices that are difficult for them to change, which is true.

英伟达做了很多很难改的选择。

Nvidia makes a lot of choices that are difficult for them to change.

他们不是完美的,只是非常平衡。

They're not perfect, they're just really well balanced.

所以你要做得尖锐。

And so you want to be spiky.

你得选一个点,然后说,我觉得他们低估了HBM短缺的影响。

You want to pick something and say, I think they've underpriced the impact of how short we're going to be on HBM.

我们要朝另一个方向猛推,顺便说一句,我觉得这很可能就是最该攻击的点。

We're going to push really hard in this other direction instead, which as an aside, I do think is probably the thing to attack most.

为什么?

Why?

没有简单的方法能快速增加大量内存晶圆厂。

There's no easy way to bring on a lot more fabs of memory.

而且那些家伙一直——

And those guys have been—.
M2
M280:02

所以还得等一段时间才能有——

So it's just going to be a while until we have—.
M1
M180:03

还得等一段时间。

It's going to be a while.

嗯。

Yeah.

嗯。

Yeah.

博伊西那帮人可不喜欢为了周期性市场投大笔资本支出。

The boys in Boise don't love huge CapEx for cyclical.

他们在这方面吃过很多次亏了。

They've been burned on that many times.
M2
M280:11

但可以想象,由于这种短缺,世界是不是就会通过让系统里其他所有东西都变得更高效率来绕开它呢?

But conceivably, like, because of that shortage, the world is just going to route around it by making everything else in the system more efficient?
M1
M180:18

我觉得他们会把其他所有东西都弄得更贵。

I think they're going to make everything else more expensive.

我觉得iPhone会削减内存。

I think that iPhones will cut their memory.

iPhone的价格会上涨。

iPhones are going to go up in price.

然后想办法应对。

And going to deal with it.
M2
M280:25

你觉得为什么Nvidia不直接做到最后一步,卖token呢?

Why doesn't Nvidia go all the way to the end and sell tokens, do you think?
M1
M180:30

Nvidia在这方面真的很聪明。

Nvidia is really smart about this.

他们不跟客户竞争。

They don't compete with their customers.

Nvidia对任何事情都看得很长远。

Nvidia takes the long view on everything.

他们为什么不干脆搞个NeoCloud呢?

Why don't they even start with the NeoCloud?

他们为什么不直接私下卖算力呢?

Why don't they just sell compute out the back door?

嗯,Nvidia真的很厉害。

Well, Nvidia is really good.

Jensen很擅长让他的朋友成为亿万富翁。

Jensen is really good at making his friends billionaires.

他把CoreWeave做成了十亿美元级别的公司,甚至是几百亿美元的公司,他没必要去破坏这种友好的关系。

He's made CoreWeave a billion-dollar company, a many-billion-dollar company, and there's no need for him to kind of destroy that goodwill.

他想要建立一个多样化的社区,包括各种neo cloud和推理服务提供商,他们都在争着为Nvidia创造需求,这样如果其中任何一家决定——比如说——垂直整合或者转用AMD或者其他选择,他还有另外三家随时准备着、渴望填补那个位置的人。

Like he wants to create a diverse community of neo clouds and inference providers who are all jockeying to create demand for Nvidia, such that if any one of them decides to, I don't know, vertically integrate or go with AMD or any other option, he's got three more people ready and hungry to fill that position.

让他的买家之间有竞争,这太棒了。

It's great to have competition amongst his buyers.
M2
M281:09

我最后最喜欢问大家的问题是:别人为你做过的最善意的事情是什么?

My favorite closing question for everyone is: What is the kindest thing that anyone's ever done for you?
M1
M181:13

最善意的事情啊。

The kindest thing.

我的意思是,我脑子里第一个想到的就是这些年遇到的各位导师。

I mean, my immediate first thought is like all the mentors that I've had over the years.

这是一种很少见的人,

It's a rare person who.

他们从自己的日程中抽出大量时间,几乎把它当成自己的个人兴趣,来确保你理解某件事,或者教你一些东西,或者给你灌输某种他们觉得你快要理解但就差临门一脚的价值观念。

Takes a lot of time out of their schedule and makes it like their personal interest essentially to make sure that you understand something or teach you something or like ingrain some value in you that they think that you're on the cusp of understanding, but just push you over the line for understanding.

我之前提到的Nvidia的那些人,他们给我灌输了对性能工程的热爱,还有我大学的教授们,我记得大二时的导师,我当时是个非常没耐心的学生。

A lot of the people at Nvidia that I mentioned earlier who instilled that love of performance engineering in me, but also my professors in college who I remember like my advisor in sophomore year, I was very impatient student.

所以我去他的答疑时间,跟他说,我想造AI芯片。

So I show up at his office hours and say, I want to build AI chips.

我知道我想做什么。

I know what I want to do.

我为什么要浪费时间上这些网络和操作系统的基础课?

Why am I wasting time taking all these other basic classes in networking and operating systems?

他看了看我,然后说,他把整个技术栈都摆出来,让我看到了理解这个拼图中每一块的重要性,或者说美妙之处。

And he just looked at me and said, he laid out basically the whole stack and showed me the depth of, or the beauty of understanding every piece in the puzzle.

他把我原本只想专注于系统某一部分的思路完全扭转过来,说能够真正理解从门级硅片到构建一个大型互联网服务整个技术栈的人非常罕见。

He took my entire path of trying to focus on one piece of the system and said that it's so rare that someone can actually understand the entire stack from the gate-level silicon all the way to building a great internet-scale service.

而且,你知道吗,你应该立志成为那种在一生中能达到那种理解水平的人。

And, you know, you should aspire to be someone who over the course of your lifetime achieves that level of understanding.

这是非常非常罕见的特质。

It is such a rare, rare trait.

而且,你知道,追求那种专业水平是非常崇高的。

And, you know, that level of expertise is so noble to chase.

而且我觉得这个想法一直伴随着我。

And I think that stays with me quite a bit.

这不是评论,而是建议。

Not a comment, but an advisory.
M2
M282:35

Neil,这次对话太棒了。

Neil, amazing conversation.

非常感谢你的时间。

Thanks so much for your time.
M1
M182:36

非常感谢你邀请我。

Thank you so much for having me.
M3
M382:42

你知道微小的优势是怎么随着时间复利增长的吗?

You know how small advantages compound over time?

这在投资中是这样,在你如何经营公司时也是如此。

That's true in investing and just as true in how you run your company.

你的支出系统就是你的资本配置策略。

Your spending system is your capital allocation strategy.
已剔除 5 处广告(点击展开查看)
M3
M310:41广告 · 已剔除

Ramp is the only platform built to make your finance team leaner, faster, and better, saving businesses 5% annually on average so you can stay focused on growth.

Ramp customers grow revenue 3.2 times faster than the average American business.

Visa, Vercel, Cursor, Stripe, Notion, ElevenLabs, Shopify, and 70,000 other businesses all now run on Ramp.

Mine does too, and so should yours.

Learn more at ramp.com/invest.

OpenAI, Cursor, Anthropic, Perplexity, and Vercel all have something in common.

They all use WorkOS.

To achieve enterprise adoption at scale, you have to deliver on core capabilities like SSO, SCIM, RBAC, and audit logs.

Instead of spending months building these mission-critical capabilities yourself, you can just use WorkOS APIs to gain all of them on day zero.

That's why so many of the top AI teams you hear about already run on WorkOS.

WorkOS is the fastest way to become enterprise-ready and stay focused on what matters most—.

M1
M111:34广告 · 已剔除

Your product.

M3
M311:35广告 · 已剔除

Visit workos.com to get started.

Felix by Rogo is a personal finance agent that turns a single prompt into finished, client-ready work using your firm's own templates, context, and standards.

Send Felix an email like, take these comments and turn them for me, or update my tracker with the context of these emails.

And Felix sends back finished PowerPoint decks, Excel models, and sourced research.

Felix works the way your team already does, delivering work quickly and accurately around the clock.

Learn more at rogo.ai/felix.

M3
M351:11广告 · 已剔除

Vanta automates security and compliance for over 16,000 fast-moving companies like Ramp, Cursor, and Harvey, keeping them audit-ready around the clock.

It's the number one agentic trust platform, and it now helps companies like yours watch for the risks that show up between audits across your vendors, your AI tools, and your whole environment.

Every new tool your team signs up for, every vendor that turns on AI features, is an opportunity for something to go wrong.

And most security programs weren't built for AI's pace of growth.

The Vanta Agent works like a 24/7 GRC engineer in the background, finding issues, drafting fixes for you, and cutting vendor assessment time by up to 50%.

Whether you're a fast-growing startup or a global enterprise, Vanta helps you earn and prove trust.

Invest Like the Best listeners get a special offer of $1,000 off Vanta at vanta.com/invest.

Ridgeline is the first end-to-end system of record with embedded AI for investment management firms running portfolio accounting, reconciliation, reporting, trading, and compliance all on one unified platform.

Firms are moving off legacy technology and onto Ridgeline because of how far ahead Ridgeline's AI features are compared to anything else in investment management software.

I've been hearing from a lot of investment managers about AI, and they fall roughly into 2 camps, with some unsure where to even start and others convinced they can build their own order management system over just a weekend.

The reality is that running an investment firm will always require governance, controls, and a single source of truth for your data, and no amount of AI enthusiasm changes that requirement.

If you're serious about your firm's AI strategy, Ridgeline should be part of that conversation.

And you can request a demo at ridgeline.ai.

M3
M382:51广告 · 已剔除

Ramp makes it smarter by default.

Better data, better decisions, better economics over time.

See how at ramp.com/invest.

As your business grows, Vanta scales with you, automating compliance and giving you a single source of truth for security and risk.

Learn more at vanta.com/invest.

The best AI and software companies from OpenAI to Cursor to Perplexity use WorkOS to become enterprise ready overnight, not in months.

Visit workos.com to skip the unglamorous infrastructure work and focus on your product.

Ridgeline is redefining asset management technology as a true partner, not just a software vendor.

They've helped firms 5x and scale, enabling faster growth, smarter operations, and a competitive edge.

Visit ridgeline.ai to see what they can unlock for you.