AI 悖论:一边自动化一切,一边员工翻倍。Shipper 的解释——AI 在「人类专业知识的残渣」上训练,只会解决已被解决过的问题;遍地「差不多对」的泔水被重新定价后,真正值钱的是能针对具体情境想透问题的专家。AI 抬高了地板,也抬高了天花板。
写作的铅笔线:对外署名文章与信息型指南分开对待,后者他乐见 AI 合写——「反正很大一部分会被 AI 读」;他引用马克·吐温与打字机的故事:每代人都觉得新工具「没有人情味」。批判性阅读:Shipper 是被访者立场,「写作者都在用 AI」属个人观察,其公司产品亦建立在所评模型之上,存在利益相关性(作者已在文首披露其伴侣任职 Anthropic)。
My fiancé works at Anthropic, whose models Every reviews and builds on, and which comes up below.
Earlier in our miniseries on productivity in the AI era, Replit's Amjad Masad described a "self-driving company" where most engineers don't look at the code anymore. For our penultimate episode, I wanted to visit someone running a version of that experiment in a somewhat unexpected place: the media business. Dan Shipper runs a website that reviews new technology alongside a product lab that's building it — and for some time now I've wondered what it's like to work in a place like that.
Shipper is the co-founder and CEO of Every, which he launched in 2020 with Nathan Baschez as a bundle of business newsletters. Today Every is a publication about AI that draws attention across the industry — for Shipper's column Chain of Thought, his podcast AI & I, and especially for the "vibe checks" in which he and his coworkers get early access to new frontier models and put them through their paces before their general release.
西珀是 Every 的联合创始人兼 CEO。2020 年,他和内森·巴谢兹(Nathan Baschez)以一组商业通讯的形式创办了这家公司。如今,Every 是一家备受行业关注的 AI 媒体——有西珀的专栏《Chain of Thought》、他的播客《AI & I》,而最有名的当属他们的「vibe check」:他和同事们提前拿到新的前沿模型,在其正式发布前做足全方位的实测。
At the same time, Every is also a product studio. The roughly 30-person company offers Cora, an email assistant; Sparkle, a file organizer; Spiral, a writing tool; and Monologue, a dictation app — all of which are bundled with the journalism into a $20-a-month subscription. Shipper says AI now writes essentially all of the company's code, while humans still (mostly) write the essays.
But AI is changing the way that the company writes. Shipper told me Every has tried to clone the taste of its editor in chief, Kate Lee, by collecting a dataset of 30,000 of her historical edits, using it to build a copy-editing agent, and back-testing it against her past work. It's an effort to capture the expertise of a single employee and distribute it more broadly throughout the enterprise — a preview, I think, of how more businesses will think about the relationship between AI and employees in the years to come.
但 AI 正在改变这家公司的写作方式。西珀告诉我,Every 尝试「克隆」其总编辑凯特·李(Kate Lee)的品味:他们收集了她过去 3 万条修改记录作为数据集,用它构建了一个文字编辑智能体,并用她过往的工作做了回测。这是把一名员工的专业能力提取出来、分发到整个组织的尝试——在我看来,它预示了未来几年更多企业将如何思考 AI 与员工的关系。
Shipper was also candid about what it's like to publish a critical review of a frontier model from a lab the company depends on, arguing that Every's role as an arbiter may be one of its most durable assets. "No one trusts a model company to tell you where they objectively sit," he said.
More controversially, Shipper told me that far more writers are integrating AI into their workflows than will admit it publicly. "I think there is a real dirty secret right now, which is that almost every writer is using it," he said. "Just, most of them are not saying so."
更有争议的是,西珀告诉我,把 AI 纳入工作流程的写作者远比公开承认的多得多。「我认为现在有一个真正的、肮脏的秘密:几乎每个写作者都在用 AI,」他说,「只是大多数人不说。」
I also had to ask Shipper a question at the heart of our miniseries: if AI automates the work, why does Every keep hiring? The company doubled from about 15 people to around 30 over the past year while loudly automating everything it can. His explanation — that AI is "trained on the residue of human expertise," but can't see beyond it — may be good news for jobs in general, at least for as long as it holds true.
我还必须问西珀一个位于本系列核心的问题:如果 AI 把工作都自动化了,Every 为什么还在招人?这家公司一边高调地把一切能自动化的都自动化,一边在一年里从约 15 人扩张到约 30 人。他的解释是——AI「在人类专业知识的残渣上训练」,无法看到残渣之外的东西——这对整体就业或许是个好消息,至少在这个判断还成立的时期内。
An excerpt of our conversation is below, edited for clarity and length.
以下是我们对话的节选,为清晰和篇幅做了编辑。
Casey Newton: Every does "vibe checks" of new models. I found when I worked at a site that reviewed gadgets, there was sometimes a tension between what the companies want from early reviews and what you give them as a critic. You published what I would say was a fairly critical review of Sonnet 5 — "a model pitched for everyone impresses no one." How did that affect your relationship with Anthropic, and how much did you think about that before you hit publish?
Dan Shipper: Obviously, we know a lot of people at OpenAI and Anthropic, and you never want to be totally mean to people you're friends with. But actually, even before we publish anything, they're asking, "What do you think?" Because they want to make the model better, and they know that if we don't like it, it means something — and they'd rather know beforehand, honestly, than find out from a ton of other people who use it. They would probably also prefer that we didn't publish a big thing saying this model sucks. But they know we're not trying to be mean. We just have to say what we think, and if we think that, it's pretty likely a lot of other people are going to feel that way. My goal is never to shit on them; it is to help make better AI happen, and I think we do that in partnership with them. Sometimes it can get a little bit heated every once in a while — they're like, "I don't see how you could feel that way." But that's the exception to the rule.
丹·西珀:显然,我们在 OpenAI 和 Anthropic 都认识很多人,你永远不会想对朋友太刻薄。但实际上,甚至在我们发表任何内容之前,他们就会来问:「你们觉得怎么样?」因为他们想把模型做得更好,而且他们知道,如果我们不喜欢,这说明了一些问题——说实话,他们宁愿提前知道,也不想等一大堆用户用过之后才发现。他们大概也希望我们不要发一篇大文章说这个模型很烂。但他们知道我们不是出于恶意。我们只是必须说出真实想法——如果我们是这么想的,很可能很多人也会有同感。我的目标从来不是羞辱他们,而是帮助更好的 AI 诞生,我认为我们和他们是合作关系。偶尔确实会有点火药味——他们会说「我不明白你怎么会有那种感觉」。但那是例外,不是常态。
Newton: The labs are enabling you to do these vibe checks, and you're using their models to build products. But it also seems like they are encroaching on your terrain, and everyone else's. Every ran a piece in June called "Built on Moving Ground," about the vertigo of building on models you don't control, where there's always a risk the labs will release as a feature something you spent the last year on. How do you think about that risk, and where do you see the durable value in a bundle like yours?
牛顿:实验室让你们能做这些 vibe check,你们也用他们的模型做产品。但他们似乎也在侵蚀你们的领地——以及所有人的领地。Every 今年 6 月发过一篇《Built on Moving Ground》(建在移动的地面上),讲的就是建立在自己无法掌控的模型之上的那种眩晕感:你花了一年做的东西,实验室随时可能作为一个新功能直接发布。你怎么看待这种风险?你认为你们这种订阅包的持久价值在哪里?
Shipper: It's a really good question, and I don't have an answer to it. There is just this dynamic where they're going to make their models better, and their models getting better does actually take a lot of the stuff that you build and make it less relevant. And they also have application layers, so there are different parts of the same company that are supporting you and also sort of competing with you. It's a messy dynamic.
I have a couple of feelings about this. One is that the thing we can do that no model company can do is tell you which models are good. No one trusts a model company to tell you where they objectively sit — how could they? So we have a good position as an arbiter between them, and that's not something that model progress will get rid of. And just because you make the model doesn't mean you know exactly how to use it well. I liken them a little bit to oven makers. You can make the oven, but it doesn't mean you know how to make a soufflé. Our job is to take an oven and say: what is the coolest thing that we could make that would be good? And they're like, "Cool, great, we'll make the oven better for that." But then sometimes they're like, "Well, maybe we'll make a soufflé, too."
The other part is that we live in this zone where things are moving really fast, and we can't rest in any one particular place. You have to both really want to make something awesome and high quality, and be willing to throw it out every three to six months as the capabilities change. That's a hard thing to do, but I think it's possible.
Newton: Where does AI make Every measurably more productive?
牛顿:AI 在哪些地方让 Every 的生产力有了可衡量的提升?
Shipper: We would never have been able to do almost all the things that we do without it. For a while we were maybe 12 or 15 people, and we were running six software products and a daily newsletter. That's insane. Even running a daily newsletter that grows, and that people like and read all the time, is hard. Then to add software products on top of that, without very much funding — we haven't raised very much money — it only became possible because we started to be able to get enough from a single engineer that you can have one person run an entire software product end to end. That was certainly not possible before at any real level of scale.
Now that it is, once you have one person, you start to hire more people, so we have products with more than one person on them. But you can get signal, and really serve an actual customer base with a real product, with one person. In a lot of ways you can think of, there are a lot of structural overlaps between The New York Times and Every. But the Times was only able to do the games bundle and Cooking and The Athletic after 150 years and a lot of scale, and we can start to do that much earlier and more quickly, with less money.
既然这成为可能,有了一个人跑通之后,你就会开始加人,所以我们有些产品已经有不止一个人在做。但重点是:一个人,就能拿到市场信号,就能用一个真实的产品去服务真实的客户群。你可以想见,《纽约时报》和 Every 之间有很多结构上的相似性。但《纽约时报》是在 150 年积累和巨大规模之后,才做出游戏包、Cooking 和 The Athletic 的;而我们能早得多、快得多、用少得多的钱开始做同样的事。
Newton: On the flip side, I'm curious if there's something you keep throwing models at that they're just terrible at, or that feels like a stubbornly human job.
Dan Shipper: Yes, all the time. Let me start simple. One thing we have been throwing models at for a long time only just started to work. We have an editor in chief, Kate Lee, who's fantastic, who I've been trying to automate out of a job for years in an extremely benevolent way. She does a ton of copy editing for us — she has the best copy-editing taste of anyone at the company. As the company has grown, she's no longer just copy editing articles; she's doing launch emails and landing pages, and making sure they all adhere to a standard. But her time is limited. She's an editor in chief; she has many other responsibilities.
Since GPT-3, I've been saying, I think we can make this better. And the answer has been "no, you can't" for a really long time. And it just started to work. Part of that is the models are good enough at instruction following that you can make a good enough prompt that it actually knows what to do in any given situation. Another thing is they're good enough at browser use, or computer use, that they can actually go into a Google Doc and make suggested changes, which is wild when you see it.
从 GPT-3 时代起,我就一直在说:我觉得我们能用模型把这件事做得更好。但在很长一段时间里,答案都是「不行」。直到最近它才开始管用。部分原因是模型的指令遵循能力足够好了,你能写出一个足够好的提示词,让它在任何给定情境下都知道该做什么。另一个原因是它们的浏览器操作、电脑操作能力足够好了——它们真的能进入一份 Google 文档、以「建议修改」的方式提出修改意见,亲眼看到时会觉得不可思议。
Another thing is they're good enough now that I collected a dataset of 30,000 of her historical edits, used that to make a prompt, and then back-tested the prompt on all of the previous documents to hill-climb and make it better and better. Now we have an agent internally that we use, and anytime someone has a piece they're working on, or a landing page or whatever, they just @ the Every Agent — "do a Kate copy edit on it" — and it does it. It's not perfect, but it's much better than having her do everything. And it gets better automatically over time: as it makes edits, and then she goes in and makes more edits, it automatically learns "here are the things I missed." So that's one thing that just became — we call it "compounding" — compoundable. But even copy editing, which is rules-based, is super, super complicated and not fully automatable even now.
还有一点:模型现在已经好到让我可以收集她 3 万条历史修改记录做成数据集,用它生成提示词,然后在所有过往文档上回测、迭代爬坡,让效果越来越好。现在我们内部有一个智能体,任何人手头有稿件、落地页或别的什么,只要 @ 一下 Every Agent,说一句「给它做一次凯特式文字编辑」,它就做了。它不完美,但总比所有事都让凯特亲自做强。而且它还会自动越变越好:它先做修改,凯特再进去补充修改,它就自动学到「这些是我漏掉的」。所以这件事刚刚变得——我们称之为「可复利的」(compounding)。但即便是文字编辑这种有章可循的工作,也超级、超级复杂,至今也无法完全自动化。
I see what we do less as "we're going to automate all copy editors" and actually more as: Kate has a specific set of skills as an expert inside of Every that she can only apply right now by spending her time. What we do with compounding is allow her to get some of that taste and viewpoint and set of skills into a little tool that lets her spread it throughout more of the org, where she doesn't have to spend her time to do more work. When you start seeing it that way, you're like, of course I want that — a tool I can teach my taste, so I can spend my time on higher-level, more interesting things.
我越来越少把我们做的事理解为「我们要自动化掉所有文字编辑」,而更多是这样:凯特作为 Every 内部的专家,拥有一套特定的技能,而这套技能眼下只能靠她亲自花时间才能施展。「复利化」做的事,是让她把一部分品味、视角和技能装进一个小工具里,传播到组织的更多角落——她不必再花时间,却能做更多的「工作」。一旦你开始这样看问题,你就会觉得:我当然想要这个——一个我能教会它我的品味的工具,好让我把时间花在更高阶、更有趣的事情上。
Newton: My impression is that you aren't automating anyone out of a job. In fact, I think you went from about 15 people in the middle of last year to around 30 this spring. You've doubled while automating. I think you've called this the AI paradox, where the more things you automate, the more humans you need to do more things. Did you expect to double in size?
Shipper: As a company, we try to automate everything we possibly can. So why did we double in size in terms of human employees? My ideal world is not one where I only hire agents and we have no humans — I'm not weird like that — but we don't have a ton of funding, and you would expect a company like ours to try to be efficient and not hire people unless we have to. And we've had to hire people. Part of that is that we're growing, so we can and we should. But I think there are also deeper structural reasons why automation weirdly creates more work for humans, especially for human experts.
The way that AI works is that it is trained on the residue of human expertise. It's trained on problems that have already been solved. One of the beauties of AI is that now you have this thing that knows how to solve every problem that's ever been solved, and you're trying to apply it to your problem. The interesting thing is that your problem is slightly different from any other problem that's ever been solved. What that creates is a situation where tons of people are just mashing on their keyboard — "solve my problem" — and it solves it, but only sort of. It's close, but not quite there. And that creates a ton of slop, and that's not really valuable. You have this glut of things that look impressive at first blush, but eventually you realize they're kind of worthless, and the market reprices. So what do you do now? You need an expert to come in and solve the problem for this particular situation — really think it through, and use AI to do that.
AI 的运作方式是:它在人类专业知识的残渣上训练,在已经被解决过的问题上训练。AI 的美妙之处在于,你手里有了一个知道如何解决「所有已被解决过的问题」的东西,然后你试着把它用到你的问题上。有趣的是,你的问题总跟任何已被解决过的问题略有不同。于是就形成了这样的局面:无数人在键盘上猛敲——「解决我的问题」——它确实解决了,但只是「差不多」。很接近,但没到位。这制造出大量的泔水(slop),而那并没有真正的价值。你会看到一堆乍看惊艳的东西过剩,但最终你会发现它们基本没价值,市场会重新给它们定价。那接下来怎么办?你需要一位专家进场,针对这个具体情境真正想透并解决问题——用 AI 去做,但由专家来做。
And who does that work? Because everybody can now do something that's sort of like what they do — everyone is a programmer, to some extent. But experts are involved in building systems to take the people who want to program and contribute, and make that actually productive. The same thing is true inside OpenAI: they have teams of people building infrastructure so that everyone can ask data-science questions and know that the answer is the approved thing that OpenAI would do — because if they just raw-Codexed or Clauded it, it would be something, but it wouldn't be right.
That's one thing experts do. The other thing they do is moonshot things that wouldn't have been possible previously. So it both raises the floor and it raises the ceiling, and there's much more to do than ever before. If you told me in 2020 that you could just send Fable off and vibe code an entire to-do app in a day, and asked what would happen to engineering, I'd have said, I don't know, that sounds nuts. And the reality is, we still have engineering. It's just moved up a level.
这是专家做的一件事。他们做的另一件事,是去做以前根本不可能的「登月」项目。所以 AI 既抬高了地板,也抬高了天花板——要做的事比以往任何时候都多。如果你在 2020 年告诉我,有一天你可以直接把 Fable 派出去,一天之内凭感觉 vibe code 出整个待办应用,然后问我会对工程行业有什么影响,我大概会说:不知道,听起来很疯狂。而现实是,工程还在,它只是整体上移了一层。
Newton: Let me ask about writing and AI. You're leaning very hard into having AI do as much as possible, but you're retaining some level of authorship. How do you think about how much of the writing — let's say of your vibe checks — should be a person typing words on a keyboard, and how much of it they can outsource?
牛顿:聊聊写作和 AI 吧。你非常激进地让 AI 做尽可能多的事,但你们仍保留了一定程度的「作者性」。你怎么把握这个分寸——比如你们的 vibe check,多少应该由人亲手敲键盘,多少可以外包出去?
Shipper: I will tell you, but first I want to ask: have you ever used an editor? Anyone that's going into your Google Doc and making suggested changes? Have you clicked accept on any of those changes? That is a similar dynamic — especially for public writing that has your name on it — to the appropriate use of AI. You want someone that understands what you think — and you have to know what you think, which sometimes you can get to with an AI — and is helping you create the best version of that. But it's yours, and whether or not you type the words does not really matter to me. But they have to be yours.
西珀:我会回答,但先让我反问一句:你用过编辑吗?就是进入你的 Google 文档、提出修改建议的那种人。你点过「接受」吗?这和「恰当地使用 AI」是同构的问题——对那些署着你名字的公开写作尤其如此。你需要的是一个理解你想法的人——而你自己得先知道自己在想什么,这一点有时也可以借助 AI 来抵达——他来帮你把想法打磨成最好的版本。但它是你的。字是不是你亲手敲的,对我来说并不重要;但它们必须是你的。
Newton: Well, how are they yours if you didn't type them?
牛顿:可如果字不是你敲的,它们怎么算是你的?
Shipper: How are they yours if you just pressed accept on a change that your editor made?
西珀:那如果你只是对编辑提出的修改点了「接受」,它们又怎么算是你的?
Newton: It's a fair question. But people do have really strong feelings about AI. I don't think most people would get mad about AI suggesting a different phrase — but if it wrote the first version of an entire chunk of your vibe check and you published it as is, people would have feelings about that. How much theorizing have you had to do about where the lines are — and are they drawn in pen or in pencil?
牛顿:这个问题问得公道。但人们对 AI 确实有非常强烈的情绪。我想大多数人不介意 AI 建议换个措辞——但如果它写了你 vibe check 里一整段的初稿、而你原样照发,人们是会有意见的。关于这条界线该划在哪,你做过多少推演?这条线是用钢笔画的,还是用铅笔画的?
Shipper: A lot of theorizing, a lot of trying different things. They're definitely in pencil, because things are changing. I also really want to separate out people's reactions from what I think the long-term norm is. There's a big difference between our audience and a mainstream audience, which would be much more sensitive to this. I'm thinking about what the right long-term norm is for people who are just used to this technology, for whom it feels like a part of everyday life, as opposed to this new big threatening thing.
There's a long history of this. When I started writing about writing and AI a couple of years ago, I researched the history of the typewriter. Mark Twain was the first American author who really loved the typewriter — he also, I think, spent a bunch of money and bankrupted himself trying to make typewriters a thing, so maybe not as good a businessman as he was a writer. At that time, it was a big deal for him to be into it, because people got offended if you sent them a typewritten letter. It wasn't your handwriting, so it looked like an advertisement. It looked impersonal. This is a very common thing in the history of technology. If I send my mom a text message, it doesn't feel as personal to her as a call — but a call didn't feel as personal to her mother as talking in person. So I'm trying to think about where the norm is going to go.
There's a whole range of different circumstances you have to consider, and we do different things in different circumstances. External writing that has your name on it and is from your perspective — that's different from, say, a long-form guide that is intended to be mostly informational rather than narrative-driven. There, we often include the AI as a co-writer, and I'm much more fine with having chunks of it be AI-written, because I think a lot of it's going to be AI-read. It's informational — the point is to put it in your agent and have it help you when you need it. For stuff that is from a person and feels like it's yours, that's different.
需要考虑的情境有一整个光谱,我们在不同情境下做法不同。署着你的名字、代表你个人观点的对外文章,和一篇主要提供信息、不以叙事驱动的长篇指南,是两回事。后者我们经常让 AI 作为共同作者参与,我也更不介意其中大段内容由 AI 撰写——因为我觉得它很大一部分本来就会被 AI 阅读。它是信息性的:意义在于把它丢进你的智能体,让它在你需要时帮到你。而那些来自一个具体的人、感觉上属于「你」的东西,是另一回事。
What is good writing, at its core? George Saunders says this, which I love: it is just applying your taste on every word, over and over and over again, until it is the most pure expression of what you've thought. I think you can use AI to help you with that — I do that all the time — but it requires a lot of your time. Writing — this is cliché to say — is about thinking. I often don't know what I think until I write something.
好的写作,本质是什么?乔治·桑德斯(George Saunders)有句话我特别喜欢:写作就是把你的品味施加在每一个词上,一遍又一遍,直到它成为你所想之物最纯粹的表达。我认为你可以用 AI 来帮你做这件事——我一直都在用——但它仍然需要你投入大量时间。写作——这话说滥了——关乎思考。我常常是写下一些东西之后,才知道自己在想什么。
Newton: On your show AI & I, you've interviewed lots of creatives about their creative process. This is a fraught topic in the creative community, but for those who are curious: what has separated the writers who get better with AI from the ones who lose themselves in it?
牛顿:在你的播客《AI & I》上,你采访过很多创作者,聊他们的创作过程。这个话题在创作圈很敏感,但对好奇的人来说:那些因为 AI 而变得更好的写作者,和在 AI 里迷失自我的写作者,区别到底在哪里?
Shipper: Here's the thing: I think there is a real dirty secret right now, which is that almost every writer is using it. Just, most of them are not saying so. I almost want to have a little writers-anonymous support group, to have people come and confess that they use AI. Some more than others — and some are truly still with pen and paper. George R.R. Martin still writes in DOS. Writers have very particular preferences for how they do their thing.
西珀:事情是这样的:我认为现在有一个真正的、肮脏的秘密——几乎每个写作者都在用 AI,只是大多数人不说。我简直想办一个「写作者匿名互助会」,让大家来坦白自己在用 AI。有人用得多些,有人用得少些,也有些人是真的还在用纸笔——乔治·R·R·马丁(George R.R. Martin)至今还在 DOS 里写作。写作者对创作方式都有非常独特的偏好。
Newton: But he hasn't finished a novel in like 15 years, so I'm not sure we want to be holding him up as a productivity model.
牛顿:可他差不多 15 年没写完一本小说了,所以我不太确定我们该拿他当生产力楷模。
Shipper: He's definitely not a productivity model. But the writers that do it well — it's the same thing as using AI well in general. It's like putty. You can do anything with it, and your goal is to find something that you're excited about, and then play around with it, and take risks: what if I did this? How would it work? Could it help me? As much as you can, allow yourself to get into it and take the risk, and allow it to change what it means to write for you a little bit, and know that you can go back. There's a whole new world of things that are possible that might be scary, but once you get into it, it's really awesome. It changes you, it changes the work that you do, and I think it's for the better.
但VLA路线在2026年已经显出疲态。2025年它还被视为具身智能大模型的核心范式,到了2026年,世界模型快速升温,NVIDIA在GTC 2026上力推Cosmos世界模型和Physical AI Data Factory,行业里“VLA是否过时”的争论此起彼伏。根本问题在于,端到端VLA虽然理论上具备一定泛化能力,但主要适用于短程任务,在复杂长程任务上存在明显局限——缺乏长期记忆和规划机制,容易出现遗漏步骤或逻辑混乱,最终陷入行为停滞。