前沿实验室真正的解药有三:推理市场将以量补价、推理数据反哺模型、以及 Claude Code 这类工具栈的用户黏性。
真正该害怕的是网络安全:Hugging Face 被 AI 智能体攻破后,因美国模型护栏锁死,被迫用中国开源模型做防御分析——「这太荒唐了」。
There's a story I tell about my first day in STRT-431 at Kellogg School of Management, the introductory class that every first-year MBA was required to take; I leafed through the readings and case studies and was dismayed that there weren't any tech companies on the docket. Me being me, I spoke to the professor after class wondering why, and was told that the goal of the course was not to necessarily learn about specific industries, but rather to uncover broadly applicable universal principles that could be applied to any company in any industry.
我常讲一个故事:在凯洛格管理学院(Kellogg School of Management)上 STRT-431 的第一天——那是每个一年级 MBA 的必修入门课——我翻了翻阅读材料和案例清单,很失望地发现课表上一家科技公司都没有。我这人就这样,下课去找教授问为什么,得到的回答是:这门课的目标不是学习某个具体行业,而是提炼放之四海而皆准的普遍原理,适用于任何行业的任何公司。
I did not, as I usually tell the story, find this very satisfactory: to me the nature of tech, particularly the fact that software and distribution had zero marginal costs (and zero transaction costs), was something fundamentally different; putting in zeroes in formulas tends to wreak havoc! I soon realized, however, that that was my opportunity. The fundamental insight undergirding Aggregation Theory is that zero marginal costs leads to fundamentally different value chains than people once expected from the Internet: centralization and scale in a world where controlling demand mattered more than distributing supply.
What is fascinating about AI, however, is the extent to which those old universal principles are coming back to the forefront. That was never more apparent than this past weekend, when arguments raged on X about the implications of Kimi K3, another open weights model out of China, approaching the state-of-the-art in terms of capabilities. The long and short of it is this: marginal costs are back in a big way, both in terms of short-term implications of state-of-the-art free models, and in terms of the long-term structure of the industry.
然而,AI 的迷人之处恰恰在于:那些古老的普遍原理正在多大程度上重回舞台中央。上个周末这一点再明显不过——X 上吵翻了天,争论的是又一款来自中国的开源权重模型 Kimi K3 能力逼近最前沿意味着什么。长话短说:边际成本大举回归了,无论是最先进的免费模型带来的短期冲击,还是行业长期结构,都是如此。
销售成本 vs 研发COGS Versus R&D
One of the most common misconceptions undergirding discussion of open weights models is that they are cheaper — free, even. After all, you can just download the weights, and skip the time and expense and capabilities necessary to create your own model. That is, of course, true, but the "free" in this case is a reference to the amount you need to spend on research and development; R&D is a fixed expense that is independent of the revenue you generate. If you spend $1 million in R&D, it doesn't matter if you do $100 thousand in revenue or $100 million; you still spent $1 million on R&D (it does, of course, impact your profitability).
What is related to revenue is COGS — cost of goods sold — and COGS is real for AI in a way it hasn't been for software for a very long time. Specifically, running inference on a model — whether that model be Kimi or Fable — costs money, and the amount of money an AI provider spends on inference is, at least in most business models, directly correlated to revenue. To reuse the above example, generating $100 million versus $100 thousand in revenue will likely require 1,000x COGS. In concrete terms, if it costs 50 cents to generate the tokens that drive $1 in revenue, then $100 million in revenue will have $50 million in COGS; $100 thousand in revenue will only have $50 thousand in COGS.
真正与收入挂钩的是 COGS——销售成本(cost of goods sold)——而 COGS 对 AI 来说是实打实的存在,软件行业已经很久没面对过这种局面了。具体来说,跑模型推理——不管这个模型是 Kimi 还是 Fable——是要花钱的,而 AI 服务商在推理上的花费,至少在大多数商业模式里,与收入直接相关。沿用上面的例子:做 1 亿美元收入和做 10 万美元收入,前者需要的 COGS 大约是后者的 1000 倍。说得更具体些:如果产生 1 美元收入需要烧掉 5 毛钱的 token,那么 1 亿美元收入就背着 5000 万美元的 COGS;10 万美元收入只背 5 万美元。
The point in terms of open weight models is that they are not free to serve. Kimi K3 costs $3 per million input tokens, and $15 per million output tokens; that is cheaper than Sol's $5 per million input tokens and $30 per million output tokens, but that might not even be the right measurement.
Nvidia CEO Jensen Huang has described what Nvidia is building as "token factories", and from Nvidia's perspective that framing makes sense. Nvidia GPUs are model agnostic: they generate tokens, and do so in the fastest and most efficient way possible. That leads to measurements like tokens-per-second, time-to-first-token, tokens-per-watt, token cost, etc., and Huang argues that these metrics will be the basis for decision-making.
This is a framing that definitely made sense during the first paradigm of AI, the ChatGPT era, when tokens were delivered straight to the end user. The second paradigm of AI, however, the reasoning era, confounds this measurement. Reasoning entails an explosion in chain-of-thought tokens, and different models need different amounts of reasoning tokens to arrive at the right answer. Kimi, for example, reportedly uses significantly more tokens than Sol, rendering its price advantage moot. Agents introduce a similar dynamic: some models are more efficient than others in terms of the number of tokens they need to execute agentic workflows.
这套框架在 AI 的第一个范式——ChatGPT 时代——确实成立,那时 token 直接交付给终端用户。但 AI 的第二个范式,也就是推理(reasoning)时代,把这套度量搅乱了。推理意味着思维链 token 的爆炸式增长,而不同模型得出正确答案所需的推理 token 数量并不相同。比如据报道,Kimi 消耗的 token 明显多于 Sol,这就让它的价格优势失去了意义。智能体(Agent)带来了类似的问题:在执行智能体工作流所需的 token 数量上,有些模型比别的模型更高效。
What this means is that tokens are not a commodity. The defining characteristic of a commodity is that it is fungible: a gallon of oil is a gallon of oil; a ton of copper is a ton of copper; a bushel of wheat is a bushel of wheat. A token from one model, however, is not the same as a token from another model. What is fungible is what is constructed from tokens, which is to say intelligence. In other words, if both Kimi and Sol generated the right answer, then that answer is fungible; the difference in tokens generated to get to that right answer is a contributor to a difference in COGS.
这意味着 token 不是大宗商品。大宗商品的定义性特征是可互换性:一加仑石油就是一加仑石油,一吨铜就是一吨铜,一蒲式耳小麦就是一蒲式耳小麦。但一个模型产出的 token,和另一个模型产出的 token 并不相同。真正可互换的,是由 token 构建出来的东西——也就是智能。换句话说,如果 Kimi 和 Sol 都给出了正确答案,这个答案就是可互换的;而为得到这个正确答案各自消耗了多少 token,则构成了 COGS 差异的一部分。
The COGS for intelligence is a function of a few different factors: model footprint (the weights and runtime state determine how much expensive memory and how many accelerators are required to host each serving replica); inference efficiency (architectural choices like Mixture-of-Experts reduce computation per generated token); memory efficiency (architectural choices can reduce KV cache requirements, allowing more concurrent requests and better GPU utilization); serving efficiency (batching, scheduling, prefix caching, and other inference optimizations maximize utilization and share work across requests); and token efficiency (the fewer tokens required to reach a correct answer, the lower the inference cost).
The reason this matters is that we are rapidly approaching a state in which intelligence for many economically beneficial tasks is in fact a commodity. Anyone building a basic CRUD app, for example, can likely do so using models from multiple providers. And, in a commodity market, the route to profitability is not through charging higher prices — again, you can (or will soon be able to) make the exact same app using multiple models — but rather through having a superior cost structure.
It's worth stepping through the mechanics here, because, as I noted a few months ago in Amazon's Durability, the dynamics of commodity markets are not something people in tech are generally familiar with. In commodity markets, everyone charges the same price, because everyone is selling the same thing; that price is determined by supply and demand. The demand for a commodity is a function of price elasticity: the cheaper the commodity, the more demand there is for it, and vice-versa. The supply for a commodity is a function of the marginal cost of producing the commodity.
The key thing to understand is that the marginal cost of producing the commodity differs by supplier. What this means in practice is that the supplier with the worst cost structure ends up selling the commodity at their marginal cost (if they can produce at all); the profits of everyone else depend on the extent to which their cost structure is better than the marginal supplier. As an example: Supplier A can produce 10 units of the commodity for $10 each; Supplier B can produce 10 units for $15 each; Supplier C can produce 10 units for $20 each. Let's assume the price elasticity is such that there is demand for 25 units of the commodity at $20.
要理解的关键是:不同供应商生产同一商品的边际成本不同。这在实践中意味着,成本结构最差的供应商最终只能按自己的边际成本卖货(如果它还生产得出来的话);其他所有人的利润,都取决于他们的成本结构比那个边际供应商好多少。举个例子:供应商 A 能以每件 10 美元生产 10 件,供应商 B 每件 15 美元生产 10 件,供应商 C 每件 20 美元生产 10 件。假设价格弹性使得在 20 美元的价位上,市场需要 25 件商品。
That means: Supplier A will sell 10 units for $20, earning $10/unit; Supplier B will sell 10 units for $20, earning $5/unit; Supplier C will sell 5 units for $20, earning $0/unit. This isn't precisely right: the reason why Supplier C will bear the shortfall is because Suppliers A and B will be able to slightly undercut them in price, which will of course affect demand (which is elastic), but it makes the point. Supplier A has a great business, Supplier B has a good business, and Supplier C is going to go bankrupt.
Bankruptcy risk is where fixed costs come back to the forefront: Supplier C has both fixed costs (like potentially R&D spend) and also may have taken on debt to finance the equipment necessary to produce the commodity. It can't price its commodity with these costs in mind — remember, the market-clearing price approximates the marginal cost of the highest-cost unit needed to satisfy demand — but those costs can absolutely drive the supplier out of business. And, if that supplier goes out of business, then prices go up, until another supplier decides to enter (or the other suppliers expand).
固定成本正是在破产风险这里重回前台的:供应商 C 既有固定成本(比如研发开支),又可能为购买生产设备背了债。它定价时没法把这些成本算进去——记住,市场出清价格约等于满足需求所需的最高成本那一单位的边际成本——但这些成本绝对可以把它逼出市场。而一旦它出局,价格就会上涨,直到有新的供应商决定进场(或者存量供应商扩产)。
智能市场The Intelligence Market
Let's bring this back to models. Right now, none of the above analysis applies because demand exceeds supply for frontier models, and supply is limited by a lack of compute. This compute shortage doesn't just mean that a compute supplier like Nvidia makes very large margins, but also that Nvidia's customers, like SpaceXAI, can turn around and resell compute at high margins as well to a company like Anthropic. Anthropic, meanwhile, can pay the markup because they can sell tokens with a higher markup still.
It's not just excess demand that gives Anthropic great margins, however: Anthropic and OpenAI likely have among the lowest costs per unit of frontier-quality intelligence, thanks to model capability, serving scale, and token efficiency. They are serving models at a particular capability level for months before their competitors, and are simultaneously applying the best models to optimizing those costs.
It's also worth noting that the market is not yet treating intelligence like a commodity: demand is for Anthropic and OpenAI specifically, and much less for models that aren't as good (thus SpaceXAI and Meta selling capacity to Anthropic); one way to think about the push for optimizing cost is that that is a function of defining jobs-to-be-done by intelligence level, such that intelligence buyers can create a market where intelligence is commoditized. In the long run, however, whoever is on the frontier is the best placed to dominate non-frontier markets as well, which are just the frontier minus n-months, i.e. months in which the frontier model makers have been optimizing their cost of serving.
还值得指出的是,市场目前还没有把智能当作大宗商品:需求点名要 Anthropic 和 OpenAI,对不够好的模型需求寥寥(所以 SpaceXAI 和 Meta 才把算力卖给 Anthropic)。理解"优化成本"这条路线的一种方式是:它相当于按智能水平来定义"待办任务"(jobs-to-be-done),好让智能的买方创造出一个智能被商品化的市场。但长期看,谁站在前沿,谁就最有条件统治非前沿市场——非前沿市场无非"前沿减去 n 个月",而那 n 个月正是前沿模型厂商优化服务成本的时间。
All of this is to say that I think the reaction to Kimi and Chinese models generally is pretty over-blown, at least from an economic perspective. Right now there is a price umbrella that is downstream of the lack of compute; I highly doubt that Chinese models are cheaper to serve on a marginal cost basis, they just seem cheaper because Anthropic and OpenAI are so supply constrained that they are charging far more than they would if there were sufficient supply to meet the demand for intelligence.
Why, then, do the model makers in particular seem so panicked about Chinese models? First, I think the frontier labs are anchored in a world where training costs dominated their financial modeling. As long as training consumed more GPUs than inference, it was critical to maximize inference revenue to help fund the next training run, which meant charging very high prices for inference.
Going forward, however, I expect the inference market to grow much faster than training costs (and that includes the assumption that training costs will continue to skyrocket), which means they really can make it up in volume. It wasn't clear this would be the case as recently as eight months ago, but the agent paradigm unlock is so massive that frontier labs should have more confidence that they can not just survive but thrive with lower prices (once they have sufficient compute).
Second, intelligence isn't in fact a perfect commodity, in part because applied intelligence makes itself smarter. Specifically, whoever is running inference is also collecting data, and that data goes into making the next iteration of the model better. This is, on one hand, all the more reason for the frontier labs to lower prices and increase usage as more compute comes online; on the other hand, this is why companies like Microsoft are increasingly obsessed with helping companies run their own models. That is much more viable if Chinese models are a viable alternative.
Third, the other way that frontier labs can not only differentiate from Chinese models but also from each other is by continuing to integrate up into the customer experience. It's striking the extent to which Claude Code and Codex are proving to be quite sticky; whichever harness you start working with is likely to be the one you stick with, and that figures to be even more the case with non-technical users. And, in the long run, this imperative to move up the stack does mean that frontier models are absolutely a threat to software providers, including Microsoft. On the flipside, the extent to which software companies who currently own the customer experience have access to competitive models is the extent to which they may be able to resist the encroachment of the frontier labs.
Finally, the ideological angle of Anthropic in particular is impossible to ignore. This is a company that believes only it can be entrusted with AI, and the existence of open weights alternatives strikes a fatal blow to that presumption.
By the same token, don't expect China to do anything about distillation attacks on the frontier labs. I think it is mistaken to attribute all of the success of Chinese labs to distillation, but it's just as much of a mistake to pretend like distillation doesn't give Chinese labs a big advantage. That advantage has really come to bear in the last year as post-training reinforcement learning has become increasingly crucial to model performance. Instead of having to fashion reinforcement learning environments from scratch, Chinese labs can simply use frontier labs models as teachers, allowing for rapid improvement at much lower costs (this is not the only reason why Chinese models are cheaper to develop, but it's a big one).
What is interesting is that one of the most important use cases for Chinese models in the West is itself distillation. Thinking Machines, for example, which just released an open-weight model, relies on Chinese models to solve the cold start problem for reinforcement learning. Dean Meyer and Konstantine Buhler wrote an excellent article on X explaining that distillation means that Western open weight models are fundamentally disadvantaged relative to China:
有意思的是,中国模型在西方最重要的用途之一,本身就是蒸馏。比如刚刚发布了一款开源权重模型的 Thinking Machines,就依靠中国模型来解决强化学习的冷启动问题。Dean Meyer 和 Konstantine Buhler 在 X 上写了一篇出色的文章,解释蒸馏为何意味着西方开源权重模型相对中国处于根本性的劣势:
Distillation does not explain China's entire open-model lead. Chinese labs have world-class researchers, substantial compute, strong pre-trained models, software-hardware codesign, and rapidly improving post-training capabilities. But distillation compresses the costly final gap between a strong base and a near-frontier system. Even if distillation represents a smaller share of a Chinese model's total capability, it represents a meaningful share of its advantage over American open models.
New enforcement mechanisms will make large-scale distillation harder, slower, and more expensive for Chinese companies. However, enforcement will not eliminate distillation backed by state actors. Every Western frontier advance therefore creates another teacher for Chinese labs. Western builders must either reproduce those capabilities independently or wait to learn from Chinese models. This gap gives Chinese labs a recurring structural advantage over Western companies.
This is a point that bears repeating: because U.S. open weight model makers must follow the frontier labs' terms of service, they (1) are worse than Chinese alternatives and (2) end up distilling the distillation, just with a detour through Chinese labs. Wouldn't it be better if western open weight model makers could go to the source?
To that end, here's an even more interesting question around distillation: why exactly is it bad? After all, what are large language models but the distillation of all of the knowledge on the open Internet, scraped by the frontier labs and distilled into the models that are themselves being distilled? Who is exactly being wronged here?
In fact, this paradox is the solution. I believe that open weight models are good for innovation (and, per the above, I think that labs on the frontier will be fine), but it's a problem to be dependent on China. The U.S. should pass a law that (1) makes explicit that collecting data for training models is fair use, and (2) bars terms of service that forbid distillation, for U.S. companies at a minimum. Stopping distillation — which is literally just querying the API — is nearly impossible; the U.S. should go the other way and lean into a new copyright policy that both indemnifies the labs and also guarantees that what they learned fuels further innovation for everyone else.
This entire Article has been an exercise in defusing overreaction to Kimi K3 specifically and Chinese open weight models generally; however, there is one reason to be concerned, and that is cybersecurity. Consider this story from The Stack:
整篇文章都是在消解对 Kimi K3、乃至对中国开源权重模型的过度反应;但确实有一个值得担忧的理由,那就是网络安全。看看 The Stack 的这个报道:
Hugging Face said its production infrastructure was breached by an "autonomous" AI agent system early last week. The platform's security team were initially stymied in their incident response (IR) by unnamed US LLM frontier model guardrails "which cannot distinguish an incident responder from an attacker," they said. So Hugging Face's defenders turned instead to the open-source GLM 5.2 model from China's Z.ai lab – running it on their own infrastructure to analyse the 17,000+ logs, or footprints, that the attackers left behind.
拥抱脸(Hugging Face)称其生产基础设施上周早些时候被一个"自主"AI 智能体系统攻破。该平台的安全团队称,他们在事件响应(IR)之初被某个不具名的美国前沿大模型的护栏卡住了——"它无法区分事件响应者和攻击者"。于是,Hugging Face 的防守方转而使用中国智谱(Z.ai)实验室的开源模型 GLM 5.2,在自己的基础设施上运行它,分析攻击者留下的 17000 多条日志与痕迹。
That's a striking public admission for the New York-headquartered Hugging Face, which lets users collaborate on models, datasets and applications, and which this summer hit the $100 million ARR mark. In an incident report, the company recommended that defenders "have a capable model you can run on your own infrastructure [our italics] vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment."
对于总部位于纽约、让用户在模型、数据集和应用上协作、今夏刚达到 1 亿美元年度经常性收入(ARR)的 Hugging Face 来说,这是一个令人吃惊的公开承认。该公司在事件报告中建议防守方:"在事件发生之前,就准备好一个经过审查的、能在你自己的基础设施上[斜体为原文所加]运行的强模型——既避免被护栏锁死,也避免攻击者的数据和凭证离开你的环境。"
It's difficult to overstate how wrong-headed the Trump administration's panicked response to Anthropic's release of Fable was, particularly since it exacerbated Anthropic's worst tendencies in terms of assuming only they can be trusted with powerful AI. In a world with only one AI, it might make sense to reserve the most powerful cybersecurity capabilities for the U.S. government and trusted allies; however, that's not the world we live in.
怎么强调都不为过:特朗普政府对 Anthropic 发布 Fable 的恐慌式应对错得有多离谱——尤其它还加剧了 Anthropic 最糟糕的倾向,即认定只有自己才配被托付强大的 AI。在一个只有一家 AI 的世界里,把最强的网络安全能力保留给美国政府和可信盟友或许说得通;但我们并不住在那个世界里。
There are and will be models eminently capable of mounting cybersecurity attacks on existing infrastructure, and those models will be — already are — widely available. The best defense — the only viable defense, in fact — will be to make sure defenders have access to the best models as well. Right now defenders are effectively banned from using Fable or Sol for cybersecurity because of Trump administration directives; that means the best alternative is using models from a country which has been trying to weaken our cyber defenses for years. This is insane!
现在存在、将来还会有完全有能力对现有基础设施发动网络攻击的模型,而这些模型将会——其实已经——随处可得。最好的防御——事实上也是唯一可行的防御——是确保防守方同样能用到最好的模型。眼下,由于特朗普政府的指令,防守方实际上被禁止将 Fable 或 Sol 用于网络安全;这意味着最好的替代方案,是使用一个多年来一直试图削弱我们网络防御的国家所造的模型。这太荒唐了!
The better course is clear: first, loosen Fable and Sol restrictions on cybersecurity, and second, ensure that U.S. open weight model makers are on an equal playing field with China. Yes, the frontier labs will kick and scream about this, but the Administration should realize that listening to their histrionics has led the U.S. to a position where U.S. companies are dependent on China for their defenses. Let the frontier labs win by being better; don't let them define safety or security, or pull up the ladder of humanity's collective knowledge.
更好的路线很清楚:第一,放宽 Fable 和 Sol 在网络安全上的限制;第二,确保美国开源权重模型厂商与中国站在同一起跑线上。是的,前沿实验室会为此大吵大闹,但政府应该意识到:听信他们的表演,已经把美国带到了一个荒唐的境地——美国公司的防御要依赖中国。让前沿实验室靠"更好"去赢;别让他们定义什么叫安全,也别让他们抽走人类集体知识的梯子。