批判性阅读:这是被访者主场——「最有对齐的模型」「防守者窗口」均无第三方验证;Hugging Face 事件的回答前半段明显公关化(「本来就有沙箱」),被 Thompson 当场戳破后才转入实质性回答;但正因如此,这是观察 OpenAI 如何自我叙事的一手材料。
This week's Stratechery interview is with OpenAI President and co-founder Greg Brockman. Brockman dropped out of college in 2010 to join Stripe, and rose to become the company's CTO; he left in 2015 and co-founded OpenAI, where he also served as CTO. Today, after an interesting few years, Brockman is President of OpenAI, and is the face of yesterday's announcement of Astra, OpenAI's newest model.
In this interview, recorded before the Astra announcement, we discuss Brockman's background, his time at Stripe, and the early years of OpenAI. We touch on the ChatGPT launch and the drama of 2023, and whether or not having a billion users is actually a disadvantage. We also touch on OpenAI's place on the value chain, and their competition with companies closer to consumers, like Microsoft, and their suppliers, like Nvidia. We also talk about Astra and OpenAI's stated commitment to alignment, and debate whether or not OpenAI took security seriously in the run-up to the Hugging Face incident.
This interview is lightly edited for clarity. Topics: Background, Stripe, OpenAI, OpenAI and the Turing Test, ChatGPT and OpenAI Drama, Productivity and AGI, Astra, The AI Value Chain, Cybersecurity.
Greg Brockman, welcome to Stratechery.
GB: Thank you for having me. Excited to be here.
汤普森:格雷格·布罗克曼,欢迎来到 Stratechery。
布罗克曼:谢谢邀请,很高兴来到这里。
So we obviously have a massive amount of news to get to, but given this is the first time we have talked, I don't want to pass up the usual Stratechery biography question I ask anyone. I do want to ask — do you define North Dakota as being a part of the Midwest?
GB: I do.
All right, well, a fellow Midwesterner, of course I have to spend some time there. You went to school in Boston, as they say, but before we get to there — you have an amazing resume even before you get to school, like the International Olympiad, but in chemistry. Where did the computers come in, or were computers a part of your life from the beginning?
GB: Well, computers were always in the background for me growing up. I loved to play computer games, but I was really into math, I was into science. I actually thought that I was going to potentially be an actor all the way through ninth grade — I was very into acting and performing as well, and dabbled a little bit in philosophy.
There's a curveball! I did not know that was coming, but I'm going to figure out what the connection is between what you do now and acting, but let's continue.
GB: Well, I felt like in ninth grade, I'd been the star of, or the male lead in, a play or two in my middle school, high school. And I was thinking about, I wanted to double down on something, I felt like I could be best in the world and really try to move the needle in a field, and I felt like I either had to pick a more cerebral, hard sciences approach or more of the arts and acting direction and creative route. I ended up picking the hard sciences one, because I felt like maybe that was an area where I could most make a difference in the world.
And why did you think you could most make a difference in the world by going in that direction?
GB: I guess for me it felt like — one of the things I loved about acting was actually the group aspect of it. I loved improv, you're just creatively thinking about things, you're bouncing ideas back and forth, but it also really requires being part of a team that works together super well, and that's something that's not always guaranteed. That's a hard thing to accomplish and find. The thing that I really liked about the more cerebral route is it feels like you sharpen your own skills. One thing I did learn, actually, was that even if you're great at writing code, that's not enough. It is actually about also bringing in that great team, and that's part of, I think, what my career has been about — really helping to build and shape environments and culture that are actually able to deliver great results.
One thing that's interesting, there's an aspect here about environment also shaping some of these things. You mentioned you were always the male lead in plays and dramas. I just recognize from my kids going through this era — my daughter was very into musicals and stuff for a while — there's intense competition for the female lead, and usually if there's just a competent male who's willing to volunteer, he gets the role every time. Was there a lot of competition for the male leads, or were you one of one?
GB: (laughing) I think that might explain it. I'll tell you a story, though.
What about the flip side too? Being in North Dakota, were there not a lot of people super intense at the hard sciences, and so did that almost seem like a rarer thing by the same token?
GB: Well, I'll tell you two stories. So one on the acting front. My first ever paid job was an acting gig. Mannheim Steamroller was in town in Grand Forks, North Dakota. Are you familiar with Mannheim Steamroller?
I am. Yes, absolutely.
GB: So they were putting on a concert, and they needed extras. They needed people to be these tin soldiers to walk around, because it was a Christmas holiday concert. I went to the audition, and I was this scrawny ninth grader, and there were all these big college students there. What they told everyone to do is, "Okay, everyone march in that direction", and then the casting people would compare notes and you'd have some downtime, then they'd say, "Okay, march in the other direction". What I noticed is that all these college students during that downtime were talking to each other, hanging out, and I was like, this job has one requirement, which is you're going to be six hours at attention the whole time walking around this concert. So during that downtime, I was there standing at attention, just being in character. At the end, they said, "Okay, we're selecting this person, this person, this person. Everyone else can leave". I was not picked.
GB: But on the way out, they said, "Actually, we thought you were amazing. We loved seeing how much you were dedicated to this, so we're going to actually make a new role for you". And so I got to be a gingerbread man, and they gave me a whole costume. And that was my first job. So it was a little bit of trying to be out of the box in order to try to get the job done — not always defaulting into the role, but trying to find creative ways to get it.
GB: But I think that when it came to growing up in North Dakota, one of the things that was really great was that it was possible for me to really excel and be the best in the state at different areas that I put my mind to. So I was advanced in math. I ended up going to University of North Dakota starting in 10th grade and taking a bunch of courses there. I got very into math competitions, and then I'd go to the national competition, I'd go to the national math camp, and there I would meet the best in the country. These people were so amazing, and actually, one thing that's been a real privilege and honor is that many of the people that I really looked up to at math camp now work at OpenAI. So I've gotten to see them in this new field, this new light.
布罗克曼:至于在北达科他长大,我觉得特别好的一点是:只要我真的用心,就有机会在很多领域做到全州最好。我数学很超前,十年级就开始在北达科他大学(University of North Dakota)修课。我迷上了数学竞赛,后来去全国竞赛、去全国数学营,在那里见到全国最厉害的一群人。他们太出色了——而一件让我深感荣幸的事是,当年数学营里我特别敬佩的不少人,如今就在 OpenAI 工作。我得以在这个新领域里,以新的眼光重新认识他们。
GB: But it was both that I could really chart my own course. I started doing math research, and I know that if I'd gone to some of these high-powered high schools or these much more competitive states, I think I wouldn't have stood out. I would have had to be in the standard track. So it was by being in this area that I was able to explore my interest and really march to my own tune.
So when people say they went to school in Boston, they usually mean Harvard, which is where you started. Then you switched to MIT. And then at some point you're working for Stripe. What's the sequence there? When did you meet the Collisons? What happened in Boston?
GB: Well, after high school, I took a year off, and I actually started working on a chemistry textbook, because I'd gotten very into chemistry in high school, competitive chemistry, and come up with a unique way of thinking about it. Very first principles, very mathematical, rather than memorization. I wanted to teach that, I wanted to propagate that. So I actually wrote 100 pages. It's on my website right now. I haven't finished it. I've been intending to get back to it.
That's like a retirement project, I love it.
GB: Exactly. Well, at this point, I think you can just ask Astra, and it'll do a great job.
GB: But I was trying to figure out how do I get this thing published, so I asked one of my friends who had done something similar in math, and he said, "Well, you don't have a PhD, so no one's going to publish it. So you can either self-publish" — and I was like, oh, that's a lot of work, a lot of capital — "or you can make a website and try to promote the ideas that way". And I said, "I guess I'm going to learn how to code". So I went on W3Schools. Do you remember W3Schools? Have you ever seen that website?
No, I don't think so.
GB: Okay, so this is the classic — they have an HTML tutorial, JavaScript, CSS, PHP. I read through them and I was like, "I should test this out". I remember I built a little first widget: you could click a table column, and it would sort the rows accordingly. It was the coolest feeling ever. I had this thing in my head that I was picturing, now it's in the world, and anyone can benefit from it. They don't need to understand the details behind it, it just works.
GB: I remember the very first thing I built that had users, it was actually a competitive chatbot game, and I got 1,500 hits from StumbleUpon one day. It was the most glorious feeling, there were 1,500 people who had played with my game, had hopefully enjoyed it, but they stuck around enough to play it. I was just like, "This is what I want, I want to help people, I want to benefit people and build for them". So that's what I showed up thinking I was going to do at Harvard, that really had changed for me. I thought I was going to do these more eclectic interests, and instead I was like, "I just want to build". Freshman year at Harvard, I was in this computer club, there were these two seniors who would have obscure technical debates every single time. We would all listen and say, "One day we'll understand, one day that will be us". But then they graduated, and sophomore year came.
So did you just switch over to MIT because you realized you had this focus on coding and building?
GB: That's basically right. Because sophomore year, I was running the club. I was supposed to be having the obscure technical debates, and I was like, "I'm not ready. I have so much to learn. I need to be around people who are so much better than me".
Got it.
GB: And so I spent so much time down at MIT, and I was like, it just makes sense to transfer.
汤普森:所以你转去 MIT,就是因为意识到自己只想专注于写代码、造东西?
布罗克曼:基本如此。因为大二时轮到我来带社团了——该我来进行那些高深的技术辩论了,可我心里想的是:「我还没准备好,我要学的东西太多了,我需要待在一群比我强得多的人身边。」
汤普森:明白。
布罗克曼:于是我往 MIT 跑的时间越来越多,后来觉得,转学过去才是顺理成章的事。
二、StripeStripe
Got it. So when did you meet the Collisons then?
GB: So I met them in 2010, late 2010. We had a lot of mutual friends, because John had gone to Harvard, Patrick had gone to MIT, and they were poking around. That team was poking around trying to find who's into computers at either of these schools, and my name kept coming up.
Right.
GB: So I got a reach out from the team, and I flew out. And I remember meeting Patrick, really, for me, was the moment I was like, "All right, this is someone I want to work with, I feel like we could build something great together".
So my next question was what attracted you to Stripe, but it sounds like you just answered it. What did you learn there? You were pretty early on the team, progressed very rapidly. By the time you left, you were CTO. What was the takeaway for you from Stripe itself and Stripe scaling, but also yourself growing so rapidly inside this company that is itself growing rapidly?
GB: First of all, for me, it's always been about the people. I knew that these were people that I wanted to work with, that it felt like we could learn together, we could accomplish something great, and so that was a real key. And by the way, dropping out of school twice is something that I do not wish on anyone's parents. It definitely was something that was difficult for mine, but they actually were very supportive in the end.
No, we're getting the idea, you take things to the extreme, right? Most founders drop out once, you had to do it twice. We're getting the drift.
GB: Exactly, exactly right. I remember a lot of the early days of Stripe was really about first principles thinking. We were in a domain, this credit card industry, that is very opaque, very Byzantine, it's been built up over many decades and has so much complexity that the card networks themselves run on ISO 8583, the spec from the '80s. It's a byte-oriented format, the whole thing. And we were trying to figure out, "How do we make this simple, dead simple, for the Internet era?".
汤普森:懂了,你做事就是走极端——大多数创始人只辍学一次,你非要辍两次。我们领会了。
布罗克曼:没错。我记得 Stripe 早年最重要的就是第一性原理思考。我们身处的信用卡行业极其不透明、极其错综复杂,几十年层层累积,复杂到卡组织自己还在跑 ISO 8583——那是上世纪 80 年代的规范,整个是面向字节的格式。而我们要解决的问题是:「怎样让它在互联网时代变得简单、极致简单?」
GB: So a lot of this is about deep understanding of a domain that you have no familiarity with. None of us really grew up as payments experts, but you just want to really deeply understand how it works so that you can expose the right primitives and the right APIs and abstractions. That to me was actually the core skill, and something that has been very transferable between Stripe and OpenAI. They're very similar in some ways, where you go and you have to scientifically learn about a domain. AI versus payments, obviously different in terms of the specifics of those domains — one is much more about natural science, the other is almost this system that has been built up of complexity — but they are fundamentally about understanding the underlying why of how something works and exposing it in a way that's simple and easy.
GB: I spent a lot of time on recruiting, a lot of time on culture, I think one thing I found is that I love coding, I love that feeling of flow state and just building and creating.
Right, you're legendary for these hours-long flow states and just coding endlessly. What's the longest coding session slash flow state you had while building Stripe?
GB: Oh, it all blurs together. I couldn't possibly say, but I would just say that for me, there was this 24-hour sprint that was how we actually got onto the credit card networks. That was just this really great time that all of us were there together in the office, none of us slept, and it was supposed to be an integration that was going to take nine months, we got it done overnight, and if we had missed it by a day, it would have been another month. As a startup, every day matters, so that kind of accomplishing what seems impossible otherwise, I love it. That is something that was incredibly exciting.
When you became CTO, you wrote a post saying how you were talking to other CTOs, you thought it was more of an architectural job, none of them did that, and you're like, "I feel like I lose my feedback loops, I'm not connected to the product, I need to code again, I'm going to become a coder". I'm curious, you wrote that towards the beginning of being a CTO, how long did that last? How long did you stay in touch with coding?
GB: Well, I would say probably for almost another decade. One area that I think I've grown on and that I've really learned is how to stay in touch and really help move forward a team and bring together a team, even if you yourself are not hands on keyboard.
GB: And by the way, I will say that this is something that I think is actually an important lesson almost for every software engineer now, because the act of what coding is has changed so significantly over the course of the year. I think over the next year it's going to change even more. We are all moving away from having to be the one who knows exactly which library to use and is able to craft the syntax and where the semicolons go. We are all moving to being these higher-level managers, these directors, the source of the inspiration, the vision, the judgment, the feedback. I think that shift, it's been something for me that was difficult, because you had to let go of something that I was used to and valued and really loved. But I've actually replaced it with something I love even more.
I know I'm talking to a smart guy because you stole my foreshadowing. I was going to circle back to that in a little bit, but yes, that's exactly where we're going.
Let's get to OpenAI. You're a part of the OpenAI founding team. What's your version of the story? I'm sure this could be a whole hour-long podcast, but what drew you to this space and got you guys started?
GB: Well, I have been excited about the idea of AI for a long time. I remember when I was first getting into programming, I read Alan Turing's 1950 paper on the Turing test. It's this really interesting paper, it's like 70 pages or something. It starts out by saying, "Well, what does it mean for a machine to be intelligent? I don't know what that means. Intelligent is not well-defined. So let's have a well-defined version of it".
汤普森:聊聊 OpenAI 吧。你是创始团队的一员,你的版本是什么故事?这话题肯定够做一整期播客,但——是什么把你吸引到这个领域、让你们起步的?
布罗克曼:我对 AI 这个想法着迷已经很久了。记得刚开始学编程时,我读了艾伦·图灵(Alan Turing)1950 年那篇关于图灵测试的论文。那篇文章非常有意思,大概 70 页。它开头就说:「机器有智能意味着什么?我不知道。『智能』没有明确的定义,那我们给它一个可操作的定义。」
I'm going to ask you what is AGI in a little bit. So it sounds like it's still unsettled, right?
GB: Well, there you go, yes. So Turing very smartly sidestepped the question and said, "Let's just have an operational definition of this, where if you can have a test where a human can't tell the difference between an AI speaking to them and another human, we'll define that machine as intelligent". The thing that was very interesting, though, that gets much less airtime, is he said, "Well, how are you ever going to solve this? It's just too hard to program an answer to it. You cannot write down all the rules for how to respond to questions. Instead, what if you could build a machine that learns? What if you could build what he called a child machine?". And then you teach it — you have a teacher who gives it rewards and punishments, and then you're able to actually give it intelligence and help it be able to pass this test.
汤普森:等会儿我会问你 AGI 是什么。这么说来,这个问题至今没有定论?
布罗克曼:没错。图灵非常聪明地绕开了这个问题,他说:「我们就给一个可操作的定义——如果一项测试里,人分辨不出对面是 AI 还是另一个人,我们就定义这台机器是智能的。」但论文里有个很少被人提起的有趣之处:他说,「那你打算怎么实现它?靠编程写出答案太难了,你不可能把回答问题的所有规则都写下来。那——如果你能造一台会学习的机器呢?如果你能造出他所说的『儿童机器』(child machine)呢?」然后你去教它——由一个老师给它奖励和惩罚,这样你就能真正赋予它智能,帮它通过这项测试。
GB: I remember being so struck by this idea, because as a programmer, you only make progress by deeply understanding the solution to something. There are so many problems I don't know the solution to, you don't know the solution to, no person has ever come up with a solution to, we're never going to be able to program the answer. But what if you could have a machine that could understand problems that we cannot, that could understand solutions that we could not? This isn't just about image recognition, though of course it has applied to that. This is also about questions of how do we get along better as a society? How do we structure the world? How do we ensure that the benefits of what we're creating end up lifting up everyone? These are super hard questions that humanity is not necessarily the best positioned to solve. But if a machine could understand, could look through more data, could have a deeper, richer understanding of many different fields all coming together, maybe it could solve them in ways that we could not. So I was so inspired by that idea.
GB: But it was an idea. I remember showing up at Harvard, asking my professors, "Hey, could I do some AI research?", and they showed me what the natural language processing state of the art was at the time. It was so clear to me, I was like, this is not what Turing was talking about. It's much more hard-coded, parse trees, all of those things, this is not going to scale to AGI.
布罗克曼:但那还只是个点子。我记得刚到哈佛时,我去问教授们:「我能做点 AI 研究吗?」他们给我看了当时自然语言处理的最前沿。我一眼就看明白了:这不是图灵说的东西——大量硬编码、句法分析树之类的玩意儿,这条路扩展不到 AGI。
GB: But something in the early 2010s changed, and I was watching from the outside. 2012 was AlexNet, then there was a series of other papers, and the thing that I would just keep seeing on Hacker News, it felt like every day there was a new "deep learning for X", I was just like, "What is deep learning?" — I remember going to deeplearning.org, and it just said, "Deep learning is a new approach to artificial intelligence". I'm like, I have no idea what this means, I actually knew one person in the field, I went to them and asked them to introduce me to more people in the field, and I just kept getting reintroduced to a bunch of my smartest friends from college. I was like, "Wait, that's interesting, these people are working on this, that's actually a very strong signal".
GB: So by 2015, it felt to me like there was something real happening. I was spending a lot of time as well really thinking about AI safety, thinking about the long-term future of this kind of technology, what it means to get it right, and it was more philosophy at the time. There were various writings you could find online that were very cool thought experiments, I ran a reading group at Stripe where we would talk about these things every week. So it's something I deeply cared about, thinking about if there's any way that I could help AI go slightly better than it would without me, that would be the best thing I could do with my career.
布罗克曼:所以到 2015 年,我觉得真有事情在发生了。同时我也花了很多时间思考 AI 安全,思考这种技术的长期未来、把它做对了意味着什么——那时这些更多是哲学问题。网上能找到各种很酷的思想实验文章,我在 Stripe 办了个读书会,每周讨论这些。这是我非常在乎的事——我想,如果有任何办法能让 AI 的发展比「没有我」时好上一点点,那就是我职业生涯能做的最好的事。
GB: That was all leading up to 2015, and I felt like I'd reached a milestone at Stripe where the company was going to work with or without me, it was kind of a question of, "Do I want to go the manager route?", which is what you need to get to the next phase, or, "Do I want to go start a new company?", That was something that had always motivated me, and so I decided that's what I wanted to do. As I was about to leave, Patrick said, "Why don't you go talk to Sam [Altman]", who he had introduced me to a couple years before. He said, "He's seen a lot of young people in similar situations, maybe he can give you some advice" — kind of hoping that Sam would convince me to stay. It didn't quite play that way.
(laughing) Yeah.
GB: I met up with Sam, and three minutes in, he's like, "Okay, you're clear you're out, what are you thinking about doing next?", I said, "Well, I'm thinking about doing something in AI", he said, "I'm also thinking about doing something in AI". And that was the start.
布罗克曼:时间来到 2015 年,我觉得自己在 Stripe 已经到达了一个节点:公司有我没我都能运转。问题变成了:「我要不要走管理路线?」——那是进入下一阶段的必经之路;还是「我要不要去创办一家新公司?」——那才是一直驱动我的东西。于是我决定创业。临走时,帕特里克说:「你何不去找萨姆·奥尔特曼(Sam Altman)聊聊」——他几年前介绍我们认识过。他说:「他见过很多处境相似的年轻人,也许能给你点建议。」他心里多少是希望萨姆能劝我留下。但事情没往那个方向发展。
汤普森:(笑)确实。
布罗克曼:我和萨姆见了面,开场三分钟他就说:「好,既然你铁了心要走,下一步想干什么?」我说:「我想做点 AI 方面的事。」他说:「我也正想做点 AI 方面的事。」一切就是这样开始的。
Were you on board with the whole non-profit thing? What were your thoughts on that when you set that up?
GB: Well, the idea of being a non-profit is something that Sam had proposed, and I think that there are actually a lot of very important properties, and you see ones that have really rung true to today. The technology we are building, it's just bigger than anything that's been created, it's bigger than the traditional structures and systems. There is no one corporate structure that exists that actually fully encapsulates the mission and the work we need to do.
GB: So I think starting that way made perfect sense, and there was always a question of what is it going to take to actually operationalize the mission? That's something we spent a long time really thinking about. I think we've been the company that's been the most innovative in really thinking about can you build a structure around all of the different aspects of both the commercial development that needs to happen, the distribution of benefits that needs to happen, the practical way of actually bringing forth this compute-powered economy, and doing all of that at once. So I think it's been an important element, it remains a critical element to what we do. But again, I think we've innovated so much on corporate structure around the core of this mission, and that mission is invariant.
Yeah, innovate is one way to put it. You mentioned the credit card networks, right? It's super opaque, lots of stuff from the '80s, a massive amount of path dependency that gave Stripe an opportunity, because you could abstract that all away and just be, for everyone else, "Here's an API, it'll work, don't ask questions". Now when you look back at OpenAI — it's hard to believe it's been over a decade now — could OpenAI have come about in any other way? Is there that sort of path dependency in there, or is there a, "If I went back to first principles, me, Greg Brockman, which I like to do, I would have structured this a lot differently"?
GB: I don't see any other way that we could have gotten to where we are, and I think where this mission needs us to be.
Tell me about the ChatGPT launch, because you have the turning point where you realize you need to scale, you partner with Microsoft, add the for-profit bit, and then ChatGPT comes out and it's huge. Did you have any expectations it would be as big as it was?
GB: So the thing that surprised me, my prediction error, was GPT-3.5 being something that people would really love and want. We had GPT-4 at the time, it had finished training in August or so, and we launched ChatGPT at the very end of November of '22. The thing that always happens when we have a new model is that we just latch onto it.
The old one looks terrible.
GB: All we see is just flaws in the previous one. We're just like, "Ah, this previous one is so bad, I can't imagine anyone would ever want to use it". We had like 200 testers who we'd been paying to use the pre-release ChatGPT, and again, we had to pay them to use it rather than the other way around. So there were some signs of product-market fit if you really dug in and were close to the details, but if you zoomed out, it really didn't look like we had it.
GB: But the way that we thought about it was GPT-4 clearly was going to change the world, we knew that, it was very obvious from the first moment we talked to it. I remember for that first week after GPT-4 came out of training, just feeling the reality of it. We'd been dreaming of AGI, thinking about AGI, thinking about what it might be like. But the first time you have a technology that really you can ask any question and it can give you pretty sensible answers, that got a 5 on AP Bio — all of those things, to me, felt like, okay, something is going to be different. It may not transform the world tomorrow, but over upcoming years, this technology absolutely will, and it's real now. That was very clear.
布罗克曼:但我们当时的想法是:GPT-4 显然将改变世界——从第一次和它对话的那一刻起,这就显而易见。我记得 GPT-4 训练完成后的第一周,那种「它真实存在」的体感。我们一直梦想 AGI、谈论 AGI、想象它会是什么样子。但当你第一次拥有一项技术——你可以问它任何问题,它都能给出相当靠谱的回答,能在 AP 生物考试里拿满分 5 分——这一切都让我觉得:好,世界要不一样了。它也许明天还不会改变一切,但未来几年一定会,而且它已经是现实。这一点非常清楚。
GB: So you look at the ChatGPT launch, the way I thought about it was we just need to get the infrastructure out first, so that we can have LLM-serving infrastructure that's battle-tested, that we've put our reps in. Then in March, when we launched GPT-4 — which, if you remember, we did the six-month delay between completing it and actually launching it — then we'll already have the infrastructure ready to go. But I didn't expect it to quite take off in that form, even though I expected it to do so in the future.
What was it like at that time? Was it just all hands on deck to keep the servers from melting?
GB: Oh, absolutely. So we launched into what we called a low-key research preview, and of course, it was just the full exponential, every single system you can imagine breaking, broke. Our login system became a big bottleneck, we had to do so much work to improve the login system, and you're scratching your head saying, we're building this magic AI technology, and the thing that is your bottleneck is, "Does your login actually scale?".
GB: I remember that we had a fairly inefficient set of inference kernels that were rolled out to production, and I'd actually written some more efficient things, or we had some more efficient things on the research side, and one of the big pieces of work was, "Let's actually take those optimizations, let's move them over", so a bunch of people swarmed on that problem, got it done. I think this was the general flavor of it for that first day, for that first week, for that first month, it was just scaling every system and trying to really keep up with this wave after wave of demand.
What happened in November 2023?
GB: Very complicated answer. Where do you want to start?
I don't know, I feel like I have to ask you about it. They're tied into — you took a sabbatical not too long after, was there a link between those two things?
GB: Look, I would say the way to think about it is that at the highest level, I think that 2023 really showed that there were tensions that had built up, really interpersonal tensions that had built up, that we had not sufficiently gotten ahead of. To me, that's one of the most important lessons of OpenAI, the fact that we're building technology, but it's always about the people, in good and bad ways. It means that really managing people dynamics, that is one of the most important things that we do, and if we don't get ahead of it, if we don't have the hard conversation, then that is actually where things can become much rougher. So I'm happy to drill into more details, but I think that a lot of it, if you really get there, it's not the more interesting technological things.
How much of that is tied to ChatGPT being this massive, huge hit you weren't necessarily expecting? Was there a link between those things, or do you think these tensions would have come to a head regardless?
GB: I don't think that there's a direct causal link, at least not in my view. I think that to some extent, there maybe is an underlying theme of, as our technology has progressed, everyone feels the weight of the world on them, feels the stakes on them. Actually, one of the things that's hardest is how do you just move forward? To me, the thing that I always remark upon is that the day-to-day activities that we do almost look the same as at every other company. You're still debugging some low-level issue, someone's upset at someone else because they said something, or they didn't include them in the meeting, whatever it is. It's just the human factors, the human work. But of course, the stakes are so massive.
GB: So I think that there is something that has been very important at OpenAI, and actually one of the big things that I have focused on, is really trying to not put people in positions where they feel that weight of the world and feel like they're alone in it. Really doing it together as a team, that's the critical thing, and that I think is maybe the way in which I would say that there is something — and it's not really specific to those events, but it is a consistent theme over the course of OpenAI — which is really keeping that feeling of we're doing this together, and trying to both rise to that occasion, but also make sure that we're doing all the basics and doing all those basics right. That's one way that I think we move forward.
Yeah, I mean, you've been a very vocal proponent of what I think is one of the overall philosophies of OpenAI: get things out in the world, experiment, see what happens, and react from there. Make your decisions based on empirical evidence, not theorizing about the future. That philosophy, I think you guys articulate that a lot in terms of AI, but this is my question, which I think you're kind of getting to as well, it feels like OpenAI as an organization is also this massive experiment that's being tweaked. The negative read on that is it seems like it's just veering back and forth, reorganization here, new leader there, is this an unwieldy monstrosity, or is it maybe more organic and more resilient than it's given credit for? As you look back, you say it could not be any other way, would it be better if it was a different way?
GB: First of all, it is absolutely true that we have changed and grown so much from where we started, a very different operating business, but we've been consistently the pioneer in terms of moving forward this field. That's true on safety, that's true on security, that's true on the core technology and just really thinking about the distribution of benefits. All of those areas we have focused on from the very beginning, and I think really the results speak for themselves.
汤普森:你一直大力倡导我认为是 OpenAI 整体哲学之一的东西:把东西放到真实世界里,做实验,看结果,再据此反应。基于实证而非对未来的推演来做决策。这个哲学你们在 AI 上讲得很多,但我的问题是——我猜你也在往这儿说——OpenAI 这个组织本身,也像一场不断被调参的巨大实验。负面解读是:它似乎一直在来回摇摆,这边重组、那边换将。这是一头难以驾驭的巨兽,还是说,它其实比外界认为的更有机、更有韧性?你回看时说「别无他路」,那如果换一条路,会不会更好?
布罗克曼:首先,没错,我们从起点到今天已经改变和成长了太多,早已是一家运营方式完全不同的公司;但在推动这个领域前进上,我们始终是先锋。安全上是,安保上是,核心技术上和认真思考利益分配上也是。这些领域我们从第一天起就在投入,我认为结果自己会说话。
GB: Now, that change, it's real. And it is the case that sometimes the team that you have that's right for one phase is not the right team for the next phase. One thing that I have been really focused on this year has been building up a leadership team that I'm just so excited about, thinking about this next phase and what we're going to be able to do together. So part of the theme of 2026, and one shift maybe from where we were before, is that because there are so many people in this field, because there's so much to do, and because the technology is taking off so fast and we're so compute bottlenecked, you've got to focus. You've got to really prune. You've got to pick the areas that all synergize together.
GB: So actually making the decisions on things like, "Hey Sora, amazing technology, but being in that specific, more entertainment aspect of consumer, that's not something we can prioritize relative to other things", then we'll cancel it. And then that causes downstream effects, and it's painful, it's tough to actually make these decisions, but it's all in service of really having that tight focus so that we're able to accomplish the core mission.
You're a big believer in scalability, is OpenAI itself scalable?
GB: I believe it is possibly the most scalable business ever. Yes.
I mean just internally, as far as an organization. What is not scalable? We talk about compute, we talk about data, you mentioned the human factor before. Is the ultimate alignment challenge — we think about alignment in terms of getting the AI to do what we want to do, but do you have the reverse challenge? Can you keep up from a management perspective with this space, this problem?
GB: I'd say two answers, first of all, absolutely yes. I think you can see it in how much we've matured as an organization over the past couple of years, where we were a couple of years ago is we had a lot of management debt. Again, there were a lot of areas where I think we did need to grow up, we did need to mature, but I think we've done that work. It's been hard, it's been painful, but I think we're in a so much better spot, and I feel just immensely excited about the company and our future.
汤普森:我说的是组织内部。什么是不可扩展的?我们谈算力、谈数据,你之前提到人的因素。终极的对齐难题会不会反过来——我们平常说的对齐是让 AI 做我们想让它做的事,但你们是不是有反向的挑战:从管理的角度,你们跟得上这个领域、这个问题吗?
布罗克曼:我有两个答案。首先,绝对跟得上——你可以看到我们作为组织在过去几年成熟了多少。几年前的我们背着大量「管理债」,确实有很多地方需要长大、需要成熟,但我认为那些功课我们都做了。很艰难、很痛苦,但我们现在的位置好太多了,我对公司和未来感到无比兴奋。
GB: But there's a second thing, too, which is that I think it's also worth stepping back and just recognizing that how companies run is changing. You can look at this, for example, just looking at revenue per headcount. The revenue per headcount for us and similar businesses is just off the charts relative to any previous business. There's a reason for that, you're starting to see this increased leverage you can get through this technology. And by the way, because we're making that technology and fighting to make that technology broadly available and to help so many companies, you're going to see many other companies be able to run in different ways, to be able to have that outsized revenue per head. That to me is something that is very exciting, that we are shifting what it even means to run a company and how to operate.
You mentioned cutting off Sora, and you framed it as being the entertainment aspect of consumer. ChatGPT, huge consumer hit, you made an unbelievable amount of money from consumers. But at the end of the day, how many people are willing to pay for this? How many customers actually want to be productive? Is there a bit where having such a hit in the consumer market was almost a negative, in that it was distracting, used up a lot of GPUs, and maybe you missed the boat — not missed the boat, but were late on the boat — as far as enterprise being the top focus?
GB: So we have conversations like this all the time internally, and actually, I think that's one of the strengths of OpenAI, that we really examine everything we're doing from first principles, rethink it all the time, have lots of diverse opinions and perspectives. There are some people who can take almost any angle on this argument, and they all have a point. So there's some truth to, "Hey, there's this agentic moment, we were late to it". There's also some truth to having a billion people — that's over 10% of the world population. Within the U.S., I think the number is something like maybe a third of the U.S. population uses ChatGPT every single week. Every week, that many people using your system, that is unique, there's nothing like it for this kind of advanced technology.
GB: So on the one hand, if you just think of it as, "Hey, we have advancing technology", one of the challenges with chat as a product is that it's not necessarily aligned with more intelligent models. It's not clear that people get the benefits of that directly through classic chat, if you're just using it as a search engine replacement. But I think that all of these things are going to come together and come to a head, and I think we're going to see that this billion users is an investment, that it is something that actually accrues to how models get unlocked in the future, and you're seeing the first steps towards it with ChatGPT Work and things like that, there's a bunch of nuance and complexity there, but a lot of the strategy has been to say, we've got consumer, we've got enterprise, these are two things — we don't want to do two things, we want to do one thing. We want to build one AGI, one system, one unified stack. We want it to be something you use in your personal life, work life.
布罗克曼:一方面,单看「我们的技术在不断进步」,聊天这个产品形态有个挑战:它未必和「更聪明的模型」对齐——如果用户只是把经典聊天当搜索引擎的替代品,他们未必能直接得到模型变强的好处。但我认为这些东西终将汇合、到达临界点。我们会看到,这十亿用户是一笔投资,它最终会兑现为模型未来被解锁的方式——你们在 ChatGPT Work 之类的产品上能看到最初几步。这里面有很多细微复杂之处,但我们的战略很大程度上是:消费级、企业级,这是两件事——而我们不想做两件事,只想做一件事。我们要建一个 AGI、一个系统、一套统一的技术栈,让它同时服务于你的个人生活和工作。
Right, but is there a bit about shipping the internal org chart? You come out with a new ChatGPT, a dramatic departure from the old one, it's built off of Codex. I can see the benefit for OpenAI internally, but is there a frustration that customers don't realize what they can do, so, "We're going to drop them in on the deep end, and hopefully that will help them figure it out"?
GB: I think that there's a fundamental shift happening in the industry, and you can see it with new emerging agentic products that are happening right now. I think that the core shift is you're going from chat to agentic use cases. And again, it's not just about productivity. I think that in your personal life, you want to be able to ask the thing to go book tickets for you, to be able to book your haircut, to be able to do those kinds of personal things, but you also want it to be able to give you good life advice, to be able to help you with health information.
GB: So to me, productivity is too narrow of a box. To me, consumer is too broad of a term. Enterprise is also something I think is going to shift. All these classic words, they are all going to smush together and grade together in ways that I think no one has ever built a product like that before. So my view has been that there's a change management required of how do you bring along a billion users to a new set of use cases, help them understand. And by the way, there is an unfair advantage that is possible, which is you have an AI that understands what you're trying to accomplish.
That's right.
GB: It can say, "Oh, I can actually help you more if you enable this connector, if you do it in this way". That's something where I feel like it's just an amazing thing, an amazing opportunity, and there's a lot of potential there. When I say unfair, I mean just relative to what you would be able to accomplish with classic technology. If you just compare one technology versus another, there's something unique about this one.
You mentioned the Turing angle before, and I'm glad you brought up both parts, because can AI talk like a human? Obviously, we surpassed that point a long time ago. But to me, the AGI definition — which is a fraught thing for you guys, it's finally, I think, out of your Microsoft agreement, so we don't need to worry about that angle anymore — to me, it's some connection to learning. You mentioned learning, and to what extent an LLM learned, past tense, but the challenge is does it learn on an ongoing basis?
To me, what is revolutionary about the agentic moment, the way I think about it, is really the ability to write things down. That's why the Codex/ChatGPT shift was necessary, because it gained the ability to write things down. If you write things down, you can remember things. If you can remember things, you can be tremendously more useful in all sorts of ways. The question is, is that an end state, or are we going to get an LLM that can learn continuously, and that's AGI? Am I thinking about this all wrong, or does that fit the part two of Turing's questions that he was raising?
GB: Yeah, I think that this is also a very interesting area for debate, because people do have their own definition of AGI, it's almost this blurry thing. At the beginning, we thought it'd be like, here's this point in time that everyone agrees that is the AGI, it hasn't played out like that at all.
GB: Now, I tend to take an abstracted view from the technology. So the question of, does memory have to get baked into the weights? Is this a transformer or something else? Those questions, I think, are details. The real question is, do you have a system that operates the way you would expect for a real AI, for something that can learn, that can learn from you, that can adapt to what your needs are? And the question of, is that implemented through a scratchpad that it writes down memories in? Is that implemented through soft tokens? Is that implemented some other way? All of that, to me, feels like possible answers to the question.
GB: I think it is very clear we've gone so much further with "write things down in a scratchpad" than is almost reasonable. It's actually quite amazing to see how successful it is, because there has been a lot of push — two years ago, we would have said, "Yeah, you need these super long contexts, that's the thing you need", actually, it turns out that with just "write down a scratchpad" and shorter contexts, it just goes unreasonably far.
Just write stuff down.
GB: So we'll see what the future holds in terms of improving these things. I have this belief that if you zoom out, everything's an exponential. You zoom in, you see these paradigm shifts. This, by the way, was the Ray Kurzweil view of how technology and computing works. I think it's been absolutely true for even these questions of how is memory going to work.
So you just launched Astra. We're finally here. Is this a new pre-train? Are you releasing any details about the size, the architecture? We're recording this before it's officially announced, so I haven't seen everything that you've published.
GB: So we're not talking about the internal details and architectures, things like that. But this is a huge step forward. We're talking about the fact that this is the first run that we've trained on more than 100,000 GPUs, which is an easy number to throw around, but just think about the scale of that. These data centers in some ways are these big machines that we've built in order to help deliver and create AI technology, and it's a real engineering challenge and marvel that people are able to harness that amount of compute to deliver the kinds of results that we have.
汤普森:你们刚刚发布了 Astra,终于等到这一刻。这是一次新的预训练吗?会公布规模、架构之类的细节吗?我们录制这期节目时它还没官宣,你们发布的东西我还没看到。
布罗克曼:内部细节和架构这些我们不谈。但这是一次巨大的跨越。可以说的是:这是我们第一次在超过 10 万张 GPU 上完成的训练。这个数字说起来轻巧,但想想它的体量——这些数据中心在某种意义上就是我们为创造 AI 技术而建造的巨大机器。人们能驾驭这种规模的算力、交付我们拿到的这些结果,这本身就是真正的工程挑战和工程奇迹。
GB: So part of that is about making the models more capable, but so much of the compute goes into safety and alignment, and we have so much security work that's gone around it. I think that we've done a huge amount of work to deliver this model safely. It's our most aligned model yet, which to me is something that is absolutely critical and always has been. But because the capability is so strong, alignment and safety become even more front and center in terms of everyone's work.
Your announcement post is interesting. It's very matter of fact. There's a huge number of practical use cases. The contrast to, say, your competitors' announcements is very, very large. Is your framing of AI as a tool — which I think is a fair way to put it — is that about marketing, or is that how you think about AI, as opposed to, like, creating God?
GB: I think there's a deep fundamental value that we have, and some of it actually relates to how we think about people. People are valuable not just because we can do tasks. We are valuable because we are humans, because we have feelings, because we matter. That human judgment, human oversight, human control, all of those things are absolutely critical to maintain, and to maintain forever. That is something that we believe is a core invariant.
汤普森:你们的发布文章很有意思,非常就事论事,列了大量实用场景。和你们的竞争对手——比如某些家的发布——对比非常、非常强烈。你们把 AI 框定为「工具」——我觉得这个概括是公允的——这是营销话术,还是你们真的这样看待 AI?而不是,比如说,在「造神」?
布罗克曼:我认为这来自我们内心深处的根本价值观,其中一些关乎我们如何看待人。人的价值,不只是因为人能完成任务;人的价值在于我们是人,我们有感受,我们重要。人类的判断、人类的监督、人类的控制——所有这些都绝对关键,必须被守护,而且要永远守护。我们相信这是一条核心的不变量。
GB: So when we think about what we can do to help steer the future of this technology — which in some ways is what it's all about, that is why we started this place, that is what we care about, how can we help this technology go in even a slightly more positive direction than it would without us — we think about these questions of how does this technology roll out in the world? We want it to be something that uplifts everyone, but also the question of how humans relate to technology, to computers. It's clearly changing. It's even changing in terms of just typing less, talking more to your computer, having this much more natural interface.
GB: But really, that human oversight and creativity and vision, all of those things I think are very important to preserve. So that does then bleed down to these questions of, do you talk about it like it's a person, or do you talk about it like it's a tool? Do you think about the use case? Do you think about it something differently? You can see this as almost a small thing, and I'm actually glad you pointed it out, but it's something we're very thoughtful about. The team spends a lot of time really thinking about everything we want to talk about and how we want to present this kind of work to the world.
So is this a model release, or is it a product release, or is there any difference?
GB: These things do blur together. I would say that this is first and foremost a model release, but the model is qualitatively more capable. Maybe the headline one is computer use. It's really crossed the threshold for me, computer use has always been — even from the beginning of OpenAI, I remember in November of 2015, before it even really started—
Well, that was like your first product, right? It was like playing video games or something like that.
GB: Yeah, exactly. Ah, you remember, yes. We had this vision of if you could do screen pixels, keyboard, mouse, an AI that you train end-to-end on that, it would be able to actually go and address any sort of task, anything that you want people to have help with, this AI will be able to do.
汤普森:所以这算一次模型发布,还是产品发布?两者还有区别吗?
布罗克曼:这些东西确实在相互融合。我会说这首先是一次模型发布,但模型的能力有了质变。最头条的一项大概是「计算机操作」(computer use)——它真的跨过了我心里的那道门槛。计算机操作一直都是——哪怕从 OpenAI 创立之初,我记得 2015 年 11 月,一切还没真正开始之前——
汤普森:那差不多是你们的第一个产品吧?就是打游戏之类的那个。
布罗克曼:对,没错,你还记得。我们当时有个愿景:如果能让 AI 端到端地学习屏幕像素、键盘、鼠标,它就能真正去处理任何任务——任何你希望有人帮忙的事,这个 AI 都能做。
GB: If you look at the era we've been in for the past two years, it's been a connector era. You have some pieces of software, humans can use it just fine, the AI has no access to it. So what do you do? You have to write a very specific connector that hooks up to the APIs, and not everything's exposed, so you can't do everything that you could. Then you think about that there are so many pieces of software that don't have APIs, and those are totally out of bounds.
GB: So we have this limited world where the AI is so restricted from helping you. I think that we now have the technology that's almost this universal connector. Now, that doesn't mean that all the problems are solved. You have to think about how do you have enterprise guardrails around what these AIs are doing? How do you have the appropriate oversight, management, tracking, and observability? All of those we're working on. So I would view this as a continuous process of how the product rolls out in order to help harness this capability. But it's already transforming how people do work within OpenAI, and it's really, I think, going to uplift so many companies, so many individuals.
布罗克曼:于是我们被困在一个受限的世界里,AI 想帮你却处处受限。我认为我们现在拥有的技术,几乎就是一个「万能连接器」。当然,这不意味着所有问题都解决了——你得想清楚:企业级的护栏怎么架?这些 AI 的行为怎样做恰当的监督、管理、追踪和可观测?这些我们都在做。所以我把这看成一个持续的过程:产品逐步铺开,来驾驭这项能力。但它已经在改变 OpenAI 内部的工作方式,而且我认为它将真正提升非常多的公司和个人。
七、AI 价值链The AI Value Chain
If you think about the overall value chain, there's a place where you're fighting battles on two fronts, I could see. One is you have companies like Microsoft, or other partners — if you don't want to use their name since they're still an important partner — but they want to commoditize models. They want to build the thing on top, and you can plug-and-play, shift models in and out, they're holding all the context and what's important. But at the same time, you're building these incredible capabilities that are really tied ultimately to the end user, it just goes and does the things that you want it to do. Is that just inevitably where you have to get to, to accomplish what's yours? Is there also this economic imperative — if we don't want to be commoditized, we need to get up into products and actually doing things directly connected to users?
GB: I would say that our underlying imperative is really that we want there to be more AI capability in the world. We want people to be doing more with AI, for it to help them, and we really view that we're shifting this compute-powered economy. What that means takes different forms, especially across different verticals. Sometimes we feel like we are in a position to really focus on an area and do a good job with it, or it's very core to our mission. Health is a good example. We're building something incredibly unique in health. It's actually very surprising to me how little airtime what we're doing in health gets relative to how many people it's actually helping. We have like 300 million people each week coming to ChatGPT for health queries. 300 million people, that's a huge number. Then we're also building a bottoms-up clinicians product, and we're building a top-down enterprise product for hospitals. So we have this three-sided marketplace in health, and what we're going to be able to do there is things like, you want to find people for clinical trial enrollment — that's a hard problem, but we actually may have the ability to help find people that would otherwise not be found. That both helps the patient and helps these drugs be able to move faster.
布罗克曼:要我说,我们底层的驱动力是:希望世界上的 AI 能力变得更多,希望人们用 AI 做更多事、得到它的帮助——我们真的认为自己正在推动这个算力驱动的经济转型。它在不同垂直行业会呈现不同的形态。有时我们觉得自己有位置、也有能力把某个领域真正做深做好,或者它与使命高度相关。健康医疗就是好例子。我们正在医疗领域构建某种极为独特的东西——其实让我很惊讶的是,相对它实际帮助到的人数,这件事得到的关注度少得不成比例:每周约有 3 亿人带着健康问题来找 ChatGPT。3 亿人,这是个巨大的数字。同时我们还在做一款自下而上的临床医生产品,以及一款自上而下的医院企业级产品。于是医疗成了我们的三边市场——我们能做成的事包括:为临床试验招募找到合适的受试者。这是个难题,而我们或许真的有能力找到那些原本不会被找到的人——这既帮助患者,也让新药推进得更快。
GB: One thing that we do when we go into specific verticals is we think about how do we play well with the ecosystem. It's not to say we won't compete there — we often do compete very hard — but we also really think of it as we're going to lift up all the boats, too, and how do we actually just focus on this core mission of, we have this technology, we want it to be broadly diffused, we want it to be out there. So sometimes it can be a little bit nuanced. There are always a lot of questions when we go into a specific area of exactly what we want to do, what we're set up to do and what we're not. But I think the way that we view it is that our overall goal at OpenAI benefits the more people are using AI to positive benefit.
布罗克曼:进入具体垂直领域时,我们一定会思考怎样与生态共赢。不是说我们不竞争——我们经常竞争得非常凶——但我们也真心把它看作「抬升所有的船」。核心使命始终是:我们有这项技术,我们希望它广泛扩散、无处不在。所以有时确实有些微妙——每进入一个具体领域,我们到底想做什么、有能力做什么、不做什么,总有很多问题。但在我们看来,越多人用 AI 产生正面价值,OpenAI 的整体目标就越受益。
Well, if you have the layer on top of you trying to commoditize you, there's probably an angle of you trying to commoditize the level under you. You guys just talked a lot more about your Jalapeño chip at Hot Chips. Why is Jalapeño important? Is it important beyond just saving money as far as paying for chips?
GB: I would think of it as, since 2017, we have been plugged into basically every hardware startup out there, every vendor. We talk to them, we give them feedback, we say, "Hey, here's where we see the models going, here's what we think you should do". Sometimes they listen to us, sometimes they don't listen to us, sometimes we're close partners, sometimes they don't really want to talk to us. One thing that has been very freeing about having a chip program in-house is that we're able to just go directly to the thing that we think is the best, that we think is exactly tuned for, not just necessarily what we're doing, but the aperture of where we think this technology is going.
汤普森:如果你上面那一层想把你商品化,那你们大概也在想办法把你们下面那一层商品化。你们刚在 Hot Chips 大会上大谈了自家的 Jalapeño 芯片。Jalapeño 为什么重要?它的意义超出「省芯片钱」吗?
布罗克曼:这么说吧:从 2017 年起,我们基本上和市面上每一家硬件初创公司、每一个供应商都保持着联系。我们和他们聊、给反馈,说「我们认为模型在往这个方向走,我们建议你们这样做」。有时他们听,有时不听;有时我们是亲密伙伴,有时他们不太想搭理我们。而拥有自研芯片项目带来的一种解放是:我们可以直接去做我们认为最好的东西——不仅为眼下的业务精确调优,也为我们眼中这项技术将要展开的整个光谱而调优。
GB: It was a very big investment — we have a team, an absolutely incredible team, with great leadership that has been working on this for quite some time. But again, it is also something where we work very closely with the ecosystem. We partner very closely with Nvidia as our preferred compute partner, and if you look at the size of the computers we're building and the unique computers we're building, we need Nvidia, there's no question about it, we're building these amazing training computers, we're building lots of inference with them, we're able to push their hardware actually sometimes in ways that even they didn't realize was possible.
Yeah, I heard there was a little bit of a hard pickup, maybe that made it a little harder to get very large models out in time, but it's working now, I suppose.
GB: Yes, yes.
GB: And I would say that there's something that is enabled by us having that in-house expertise, because we deeply understand things. It's one thing to be sitting on the sidelines and throwing advice over the fence, it is another if you actually have gone through the pain. A good example of this actually is AI for chip design. We've talked about this, that we've used our own model in the design of Jalapeño, it really sped things up, it got us some real wins, all the cool things. As an aside, there's a cool story there where we were coming up on a deadline, we had like a month to go, we got some optimization done with our model. We're like, "Do we spend the time to really read what it did? We know it's correct. Do we need to understand exactly what optimizations it did, or do we just spend the rest of the time getting more optimizations?" — and so we said, "You know what? We'll just get more optimizations in". So we spent that month on just running it without deeply understanding exactly all the tweaks it made. Then we went back and read it, and it actually turned out that it found a bunch of optimizations that had been on our list, but we just never would have gotten to, so that was actually a pretty cool story.
布罗克曼:而且我认为,正是自研的专业能力让一些事情成为可能——因为我们有真正深入的理解。坐在场边往墙内扔建议是一回事,自己真正趟过那些痛苦是另一回事。一个很好的例子是「用 AI 做芯片设计」:我们说过,Jalapeño 的设计用上了我们自己的模型,真的加快了进度、拿到了实实在在的成果。顺带讲个有趣的故事:有一次临近截止,还剩大约一个月,模型给出了一批优化。我们讨论:「要不要花时间仔细读它到底改了什么?我们知道结果是对的。是需要精确理解它做的每一项优化,还是干脆把剩下的时间全用来跑更多优化?」最后我们说:「就这样吧,继续跑。」于是那个月我们直接用它跑,没有深究它做的每一处调整。之后我们回头去读,发现它找到的一批优化,本来就在我们的清单上,只是我们永远排不到去做。这个故事相当酷。
GB: But now we have that expertise, we know this thing works, and we can bring that to the ecosystem. We can work closely with everyone in order to actually bring these benefits broadly, to transform hardware and do that at mass scale. So there's something about that flywheel that's absolutely critical, the chip is incredible, the team did a great job.
Is it a problem talking about it now, though, when you can't ship in volume and you still need to partner with other folks in the ecosystem to get the supply you need?
GB: Well, but this is the core, this is actually the core of everything. We think of it as — I think everything is multiplicative, everything is complementary, everything adds up. And again, it is absolutely the case that Nvidia is our preferred partner, that's not changing. In fact, we're leaning in even more with them. We're deeply, deeply grateful for that partnership, we spend a lot of time with their team, there's a lot of things that we learn from them, there are things that we hope that they can learn from us, I think that's something that doesn't change. The fact that we are able to have in-house expertise and really think about things in our own way as well, to me, that's something that's just multiplicative, I think it really benefits everyone.
You mentioned you just trusted the AI design, and that got you further down the road. Is that the answer to cybersecurity? You had some engineers give a talk at the Black Hat conference and talk about this structural problem — attackers don't need to worry about breaking things, they're trying to break things. If you're on the other side, you're worried about everything continuing to run in addition to fighting off these attacks. Do defenders need to get to the place where they just fully trust the AI?
GB: I think the hardware side is a very important case study, because there we have guardrails. We have verification, and actually, the way that we write our underlying hardware design is specifically to allow verification, so we almost co-designed the whole system.
That's like how you code, it writes the unit test first and then backs into it.
GB: That kind of thing, how you pick your language and the toolchain, the whole thing, it's all together, and it actually all adds up to a system that you can have that kind of observability and trust.
汤普森:你刚才说你们直接信任了 AI 的设计,结果走得更远。这会是网络安全的答案吗?你们有工程师在 Black Hat 大会上讲过这个结构性难题:攻击者不用担心「把事情弄坏」——他们本来就是来弄坏东西的;而防守方在抵御攻击之余,还得操心一切照常运转。防守方是不是必须走到「完全信任 AI」那一步?
布罗克曼:我认为硬件那边是个非常重要的案例,因为那里有护栏。我们有验证手段——实际上,我们底层硬件设计的写法,就是专门为了「可验证」而写的,整个系统几乎是协同设计出来的。
汤普森:就像写代码那样,先写单元测试,再反推实现。
布罗克曼:就是那种感觉——怎么选语言、怎么配工具链,所有东西都在一起,最终加总成一个你能够观测、能够信任的系统。
GB: I think it's okay for there to be some areas where you say, "I have sufficient guardrails here that it is actually okay if it's this code or optimizations that I haven't fully inspected", as long as you have the appropriate compensating controls. But I think that it is very important that you as a human do understand and feel accountability for the system you're creating. That to me is actually a core thing, back to what is it that humans are, what is unique to us, what is something that we are going to carry forward, I think accountability is a core of it. At the end of the day, you're responsible for what happens at your company.
Right, but if those on offense are not accountable, is that a structural disadvantage?
GB: So I think that this is something we think about a lot, that there is what we call this The Defender's Window. I think that we can see a little shape of the future, we have frontier capabilities that have shown the kinds of capabilities that will diffuse to threat actors. And by the way, I think the fact that this capability is not being locked up forever in a small number of labs is actually very important, it is very important that there is broad distribution of power, that is part of our mission as well. But we have the ability to have a separation in time. There's this window where defenders can get access to these capabilities, and differentially so. And my view is that it is true — there's a common wisdom in cybersecurity that offense is a technology problem, defense is a political problem. The attackers can just take something off the shelf and run with it, whereas as a defender, you have to think about your stakeholders, you have to think about your business, you have to think about how you actually get people on board, your CEO, all the executives, all those things. So I think that there is something here where defenders need that willpower.
布罗克曼:这是我们思考很多的问题——我们称之为「防守者窗口」(The Defender's Window)。我们能瞥见一点未来的形状:前沿能力已经展示出终将扩散到威胁行为者手中的那种能力。顺便说,这种能力不会永远锁在少数几个实验室里,我认为这其实非常重要——力量的广泛分布很重要,这也是我们使命的一部分。但我们有可能获得一个时间差:存在这样一个窗口期,防守方可以率先、且差异化地拿到这些能力。网络安全有句老话:进攻是技术问题,防守是政治问题。我同意——攻击者拿来现成的东西就能跑,而防守方要考虑利益相关者、考虑业务、考虑怎样真正让 CEO 和所有高管上车。所以防守方需要的是那份意志力。
GB: One thing we are recommending, and we've actually done ourselves and are talking about it publicly now, is that every company should treat this as a proactive incident. Critical business operations, proactive incident, that's your next priority, so we actually took 25% of our production engineers and put them to securing ourselves. We took our models — in fact, we took Astra, pointed it at our own systems to find vulnerabilities, and not just read the code, but really look at the end-to-end of how these things are running, so we would find real validated vulnerabilities, and then it also helped us with the remediation, patching, and fixing. So I think that you do need a shift in the energy in the ecosystem in order to stay ahead and to take advantage of this window.
Well, that's all great and fine that you're doing this now, but to me the most remarkable thing about the Hugging Face incident and the things that have come up about it is it doesn't feel like OpenAI was particularly concerned about cybersecurity. Why didn't you do this before? Hasn't the Defender's Window been open for a while, and you were also failing to take advantage of it?
GB: Well, two answers. So one is that if you look at the way that we were doing sandboxing, it was not that this workload was not sandboxed. There was actually a sandbox around it, and I think that one thing we realized is that we had—
汤普森:你们现在做这些当然很好。但对我来说,Hugging Face 事件及其后续披露里最惊人的一点是:OpenAI 看起来并没有多在乎网络安全。你们之前为什么不做?防守者窗口不是已经开了很久了吗?而你们自己也没能利用它?
布罗克曼:我有两个回答。第一,如果你看我们当时做沙箱隔离的方式——并不是这个工作负载没有沙箱,它外面确实有一层沙箱。我认为我们意识到的一点是,我们当时——
Right, which wasn't clearly sufficiently tested. Is it really a sandbox, or is there a connection to the Internet via a third party who were just thrown in there? That's the most remarkable thing about this incident. It's like, if you wanted to test for vulnerabilities, I guess you did that.
GB: It's definitely the case that the AI was able to do very creative things in order to get out and get into Hugging Face. But to me, there is a bigger thing, and I think you're pointing at the right thing, which is that since this summer, when Mythos came out, when we started to have cyber-capable models — and we even talked about our Trusted Access for Cyber program back in February, because we saw this wave coming, we wanted to really prepare for it — there is a tendency—
I know, but you talked about it in February, but you didn't point it at your sandbox, "Is my sandbox actually secure?".
GB: There is an instinct, there's a reaction to that, to say, "Let's restrict access massively, let's really put a bear hug around this, only if you can get access". And I think that to your point, because the field continues to move, it means there's time that defenders lost, there's time that people were not defending. Part of that is about access, but part of that is about how much do you put your full weight behind saying we're going to shift around this in a significant way.
GB: Now, I think that to some extent, the time is not all made equally, because we've gone from a world of cyber models being not that useful, not that differentiated, to actually being incredibly capable, incredibly powerful. We're seeing that with Astra, we've talked about how it's really saturating a bunch of these evals. I think now is the time, it's possible that a couple of months ago could have also been the time, but you would just have had much less capable models, you would have made much less progress.
GB: So I think that really estimating where we are, we are clearly there now, and I feel like that is something that we have learned, we've taken it to heart, I think that you've seen a real shift. It's actually been a cultural change in a lot of ways, an operational change, and it's not easy, because it really means that you have to have teams working together very tightly in a loop with much higher standards around how policies are set and all these things. All of that, for us, it wasn't a shift in terms of we've always cared about these aspects, but really bringing them together operationally and being able to make decisions the way that we have, I think it's been an up-level across every aspect of what we do.
Now suddenly they're able to do it, as if doing it previously would have been a waste of time, which I think is kind of a valid point. How do you avoid the trap of, "Well, the AI will be able to do this in the future, so we don't need to do it now"? Just in general, though, not even just with this.
GB: I was going to say one thing that's a specific data point, so early on, sometime in Q1, we really started thinking about, we are going to have — it's hard to know when, but we're going to have these very cyber-capable models. What is a sandbox that we could build from first principles that's as secure as you could get while being built on cloud infrastructure? And we built that. We actually took some of our best engineers and pointed them at that problem, and they sprinted on it and they produced something.
汤普森:现在你们突然能做了——仿佛之前做就是浪费时间。我觉得这个论点也有点道理。更一般地说,不只是这件事:你们怎么避免「反正 AI 未来能做到,所以现在不用做」的陷阱?
布罗克曼:我想说一个具体的数据点。早在今年一季度的某个时候,我们就开始认真思考:我们终将拥有——具体什么时候不好说,但终将拥有——网络攻防能力极强的模型。那我们能不能从第一性原理出发,在云基础设施之上造一个安全性拉到极限的沙箱?我们真把它造出来了——抽调了最好的一些工程师扑在这个问题上,冲刺攻关,做出了东西。
GB: So I think building infrastructure from first principles around what you see coming, that is something that I think is important. And to your point on when is it, "Oh, we can just let the AI do it" — again, we've seen this show before in different fields, in writing kernels, and thinking about the fact that, "Okay, we're going to be in a world where in the future the AI is going to be able to write GPU kernels very well", do the classic kinds of investments where it takes many months, sometimes a year, to get new infrastructure in place for thinking about new hardware, that kind of thing, or do we just say, "Ah, the AI will figure it out"? I think the answer is always that it takes a little bit longer than you expect for the AI to get there. But when it does, it is surprising and powerful in ways you didn't imagine.
布罗克曼:所以我认为,围绕你预见到将要到来的东西,从第一性原理出发建设基础设施,这很重要。至于你说的「什么时候可以说,哦,交给 AI 就行」——这出戏我们在别的领域已经看过,比如写 kernel:「未来 AI 会把 GPU kernel 写得很好」,那我们是做那种经典的投资——花好几个月、有时一年把新硬件的基础设施铺好——还是直接说「啊,AI 会搞定的」?我认为答案永远是:AI 到达那个水平所需的时间,总比你预期的要长一点;但当它到达时,它的强大和出人意料会超乎你的想象。
One example of this is with Astra. One thing we have found is that a number of our skills that we've built up over the course of this year, very painstakingly, to show our models the right way of doing things in OpenAI and things like that are actually now net negative for its performance.
Too many rules.
GB: Exactly. It is able to generalize better, or be able to find better ways of approaching patterns and things like that, than what we had written. So I think there's something about this where you do want to build those controls, you do want to build the deterministic infrastructure, you want to write those skills. But you also need to be prepared for, as the AI gets more capable, that some of those things, the scaffolding, will become a limiter. It's kind of like training wheels. At first, it helps you, but once you start going faster, once you have something more capable, something better, something more aligned, then it actually starts to be a hindrance.
布罗克曼:Astra 就是个例子。我们发现一件事:今年我们非常用心积累的一批「技能」(skills)——用来示范在 OpenAI 里做事的正确方式之类——如今对模型的表现反而是净负面的。
汤普森:规则太多了。
布罗克曼:没错。它比我们写下的东西更会泛化,能找到更好的处理模式的方法。所以我认为这里有个辩证:你当然要建那些控制、要建确定性的基础设施、要写那些技能;但你也必须做好准备——随着 AI 能力增强,其中一些东西、那些脚手架,会变成限制器。有点像辅助轮:一开始它帮你,但一旦你开始加速,一旦你有了更强、更好、更对齐的东西,它反而成了阻碍。
Will you ever be in a 12-hour or 24-hour coding flow state ever again?
GB: I hope so. I think there may be a day where that happens, but I will say that I have found so much joy and value in helping the team in the way that I do now. I think that for me, it's really about that mission.
Well, it's not just you, but will anyone? Because isn't AI's benefit almost that it is permanent flow state available at your command?
GB: I think that we're going to find new ways, whether it's managing agents — actually, one thing that's been so wild is seeing that software engineers are working harder than ever, because you realize if your agents aren't working, it's just time that's lost, you're never getting it back. So I think that people will achieve that flow state in ways that are kind of unimaginable right now.
At the end of the day, you're talking about this is going to be controllable, these AIs, "Don't put too many rules, they'll figure it out", if you play that out in the fullness of time, isn't that ultimately about them being uncontrollable?
GB: Well, I think this is the core of the moment, of the new phase that we're in, and in some ways I would say we're into the AGI era now. I think that is the core of this moment, where — maybe it was the previous model, maybe it's Astra, maybe it's the next model, but somewhere in there, I think we're going to cross most people's AGI threshold. Ensuring that we are pacing, ensuring that we're thinking about safety, security, alignment, and capability, all as requirements — we have standards around each of these, we want to progress them together — that is something we've always believed.
汤普森:说到底,你讲的是这些 AI 将是可控的——「别加太多规则,它们自己会弄明白」。如果把这条路推演到底,那最终不就意味着它们不可控吗?
布罗克曼:我认为这正是此刻的核心、我们所在的新阶段的核心——某种意义上,我会说我们已经进入 AGI 时代了。我认为这就是此刻的核心:也许是上一个模型,也许是 Astra,也许是下一个模型,但就在这中间的某个地方,我们会跨过大多数人的 AGI 门槛。确保我们有节奏地推进,确保我们把安全、安保、对齐和能力都当作硬性要求来思考——每一项我们都有标准,我们希望它们齐头并进——这是我们一贯的信念。
GB: But I think it's becoming very front and center that these other aspects are becoming almost the bottleneck to development. And I think that, again, is something where we've been prepared for that, we've been thinking about this, and I think we're operationalizing it in a real way. So my view is that there's a lot of progress to be made, but I think that the way that we should approach it is through increasing our standards in all of these. If you look at that, I think we see line of sight for things like monitorability. That's very key. We're bringing that in a real way. I think that we have a very good program. We have a good set of people, we have a good track record and a good mission that I think all point towards we are building systems in a way that they are controllable, and we're taking these step by step in terms of pacing.
Greg Brockman, congratulations on Astra, and yeah, can't wait to use it.
GB: Thanks so much, thank you for having me.
这也是为什么头部AI应用的更迭如此之快。点点数据AI应用先锋下载榜Top10数据显示,相较2026年4月,5月新晋上榜应用多达6款,包括ChatGPT、Piclux、AI Chat、Photo Video maker with music、Kling AI和Pivo AI;仅Refoto、Hailuo AI、VibeShort和Vidix这4款产品连续两个月在榜。