From Nonprofit Lab to Silicon Forge: The OpenAI Story and Why Its In-House Chip Bites So Hard
从非营利实验室到硅基熔炉:OpenAI 的前世今生,与自研芯片为何如此凶猛
"We are not a chip company. We are a full-stack intelligence company — and silicon is now part of the stack."
“我们不是一家芯片公司,而是一家全栈智能公司——硅片,已经是栈的一部分。” — Greg Brockman, Hot Chips 2026
Part I — Past & Present: The Arc of OpenAI
第一篇|前世今生:OpenAI 的十年弧光
1.1 2015–2018 — The Utopian Lab(理想主义实验室)
In December 2015, Sam Altman, Elon Musk, Greg Brockman, Ilya Sutskever and a handful of researchers incorporated OpenAI as a 501(c)(3) nonprofit in San Francisco, pledging $1B and a single mission: safe AGI for all humanity. No equity, no products, no CUDA. The early years produced OpenAI Gym (2016), the PPO algorithm (2017), and GPT-1 (2018) — a 117M-parameter Transformer that nobody outside NLP noticed.
2015 年 12 月,Altman、Musk、Brockman、Sutskever 等在旧金山注册为 501(c)(3) 非营利组织,许诺 10 亿美元,使命只有一句:让 AGI 安全造福全人类。没股权、没产品、没碰 CUDA。那几年只交出 Gym(2016)、PPO(2017)、GPT-1(2018,1.17 亿参数)——圈外无人知晓。
1.2 2019–2022 — Capped-Profit & The Microsoft Anchor(利润上限与微软锚定)
The 2019 pivot changed everything: a capped-profit subsidiary (OpenAI LP) sat under the nonprofit, Microsoft invested $1B, and Azure became the exclusive cloud. GPT-2 (2019) teased scale; GPT-3 (2020, 175B params) proved the scaling law; the API turned research into revenue. Then ChatGPT launched on Nov 30, 2022 — 100M users in ~2 months, the fastest consumer app in history, and the moment AI left the lab.
2019 年是分水岭:非营利母体下塞进 capped-profit 子公司 OpenAI LP,微软砸 10 亿美元,Azure 成独家云。GPT-2 试水缩放,GPT-3(2020,1750 亿参数)坐实 scaling law,API 把研究变现金。接着 2022 年 11 月 30 日 ChatGPT 上线——两月破亿,史上最快消费级应用,AI 就此出圈。
1.3 2023–2025 — Crisis, GPT-4, and the PBC Rebirth(危机、GPT-4 与公益公司重生)
November 2023 brought the five-day coup: the nonprofit board fired Altman, 770 employees threatened to resign, Microsoft blinked, Altman returned. The drama exposed a structural lie — a charity governing a billion-dollar business. By October 2025, OpenAI collapsed the contradiction into OpenAI Group PBC, a Public Benefit Corporation governed by the OpenAI Foundation (~26%), with Microsoft at ~27%, employees/investors ~47%. The same window shipped GPT-4, GPT-4o, o1, Sora, GPT-5/5.5.
2023 年 11 月上演 五天政变:非营利董事会开除 Altman,770 名员工以集体辞职相逼,微软让步,Altman 回归。闹剧撕开一个结构谎言——用慈善董事会管百亿生意。到 2025 年 10 月,OpenAI 把矛盾压进 OpenAI Group PBC(公益公司):OpenAI 基金会控股约 26%,微软约 27%,员工与投资人约 47%。同期出货 GPT-4、4o、o1、Sora、GPT-5/5.5。
1.4 2026 — The $852B Behemoth Goes Vertical(8520 亿巨头向下垂直整合)
March 2026: a 122Braiseat852B valuation (Amazon 50B,Nvidia30B, SoftBank $30B). June 2026: confidential S-1 filed for IPO. August 2026: Hot Chips stage, where OpenAI stopped being “a model company” and became “a model + silicon company.” The arc is complete — nonprofit → capped-profit → PBC → pre-IPO full-stack vertically integrated AI utility.
2026 年 3 月:1220 亿美元融资、估值 8520 亿(亚马逊 500 亿、英伟达 300 亿、软银 300 亿)。6 月:秘密递交 S-1,筹备 IPO。8 月站上 Hot Chips,OpenAI 不再是“模型公司”,而是“模型+硅片公司”。弧线收口——非营利→利润上限→PBC→IPO 前全栈垂直整合的 AI 公用事业。
Part II — Jalapeño: The Intelligence Processor
第二篇|Jalapeño:一颗“智能处理器”
On June 24, 2026, OpenAI and Broadcom unveiled Jalapeño (“Intelligence Processor”, not “AI accelerator”) — an inference-only ASIC, TSMC N3P compute die + N3E I/O chiplet, Broadcom Tomahawk networking, Celestica rack integration. 9 months RTL-to-tape-out. That is not a fast GPU cycle; that is half of Google’s TPU v1 schedule and a quarter of a typical HPC ASIC.
2026 年 6 月 24 日,OpenAI 与博通揭幕 Jalapeño(自称 Intelligence Processor 而非 AI 加速器)——纯推理 ASIC,台积电 N3P 计算 die + N3E I/O chiplet,博通 Tomahawk 互联,Celestica 做板卡机架。RTL 到 tape-out 仅 9 个月。这不是“快了一点的 GPU 周期”,而是谷歌 TPU v1 周期的一半、常规 HPC ASIC 的四分之一。
2.1 Spec Sheet(规格表)
|
Item |
Jalapeño |
Context |
|---|---|---|
|
TDP (rated / sustained) |
700W / ≤550W |
GB200=1200W, GB300=1400W |
|
Memory |
6× HBM4, 216 GiB, 15.4 TB/s |
~22 GB/s/W bandwidth efficiency |
|
Rack |
128 chips, 1.7 EFLOPS MXFP4, 27.5TB HBM4 |
Pod = 2048 ASICs |
|
Node |
TSMC 3nm (N3P/N3E) |
Same family as Blackwell, M4 |
|
Workload |
Inference only (prefill+decode, KV-cache local) |
No training |
2.2 Benchmark Reality Check(基准实测)
On SemiAnalysis InferenceX, across GPT-OSS-120B, DeepSeek R1 670B, Kimi K2.5 1T:
- Peak throughput per watt: 1.5–1.9× vs GB200/GB300
- End-to-end latency: 1.7–3.6× lower (1.65s vs 5.99s on R1)
- Interactive-agent latency regime: 2.1–4.1× faster
- Single-user decode: 1459 tok/s (Jalapeño) vs 535 tok/s (GB200), 0.69ms vs 1.87ms inter-token
- Cost: ~50% lower inference TCO vs GPU baseline; ~1.56/chip−hourvsH1001.55, Vera Rubin est. $3.61
⚠️ Caveat from SemiAnalysis itself: the fair fight is Jalapeño vs Vera Rubin, not Blackwell — a custom ASIC should beat a 2-generation-old general GPU. The number that matters is tokens per megawatt, and there Jalapeño rewrites the denominator.

Part III — Why It Is So Strong: Three Concentric Rings
第三篇|为何凶猛:三个同心圆
Ring 1 — Problem-First Silicon(场景专用:为推理从零画硅)
Jalapeño is a blank-slate design for LLM inference — systolic array instead of CUDA cores, KV-cache pinned local, prefill/decode balanced on the same die, HBM4 bandwidth matched to transformer memory traffic instead of oversold FLOPs. Nvidia Blackwell is a generalist: it carries compute inference never touches. OpenAI threw that dead weight out.
Jalapeño 是给 LLM 推理从白纸画的:脉动阵列替掉 CUDA core,KV-cache 钉在本地,prefill/decode 同 die 内配比,HBM4 带宽按 transformer 访存画像配而非堆废 FLOPs。Blackwell 是通用件,推理用不到的算力它也得带着。OpenAI 把死重全扔了。
Ring 2 — AI Designing AI’s Hardware(AI 造 AI 的硬件)
This is the recursive trick. GPT-class models + Codex sat inside the EDA loop: RTL draft, layout routing, SIMD/matrix-engine area optimization, verification triage. Result — matrix engine area −10%, SIMD unit −8%, attention/MoE blocks where model-generated code ran 1.5–1.8× faster than human-expert handwritten RTL. Nine months becomes possible because the designer is the thing being designed for.
这是递归彩蛋:GPT 级模型 + Codex 进 EDA 回路——写 RTL 草稿、布局布线、SIMD/矩阵引擎面积优化、验证分流。结果:矩阵引擎面积 −10%,SIMD −8%,注意力/MoE 模块里模型生成的代码比人类专家手写 RTL 快 1.5–1.8 倍。9 个月流片之所以可能,是因为“设计师”就是“被设计对象”本身。
Ring 3 — Full-Stack Closure(全栈闭环:模型↔服务↔硅)
OpenAI owns the model roadmap (GPT-5.x, Codex, agentic traces), the serving stack (batching, speculative decode, KV policies), and now the die. Each layer sees the other’s shape:
- Kernels tuned for Jalapeño’s memory hierarchy
- Serving latency targets baked into microarchitecture
- Power budget set by data-center MW, not by GPU SKU
That closure is what Google (TPU), Amazon (Inferentia), Meta (MTIA) chase but OpenAI reached faster because it started from the token stream, not from the datasheet.
OpenAI 同时握有模型路线图(GPT-5.x、Codex、agent 轨迹)、serving 栈(批处理、投机解码、KV 策略)、以及如今这颗 die。每层都看见另一层的形状:kernel 按 Jalapeño 存阶调、时延目标烤进微架构、功耗预算按机房兆瓦定而非 GPU SKU。这种闭环正是谷歌 TPU、亚马逊 Inferentia、Meta MTIA 在追的,但 OpenAI 从“token 流”而非“datasheet”起步,所以更快。
Part IV — Strategy, Not Sabotage
第四篇|是战略,不是背刺
Richard Ho (ex-Google TPU, heads OpenAI hardware) is explicit: Jalapeño does not replace Nvidia. Training clusters stay on Blackwell/Rubin; Jalapeño eats the inference TCO, which is now >60% of OpenAI’s compute spend. 10GW of Broadcom-built Jalapeño racks from 2026–2029, limited deployment late 2026, full ramp 2027+. Multi-vendor forever — but the marginal token now costs half.
Richard Ho(前谷歌 TPU 核心、OpenAI 硬件负责人)说得很直:Jalapeño 不替换英伟达。训练集群继续 Blackwell/Rubin;Jalapeño 吃的是推理 TCO——现已占 OpenAI 算力支出 60%+。2026–2029 年博通交付 10GW Jalapeño 机架,2026 年底小批量,2027 起放量。永远多供应商——但边际 token 成本砍半。
The deeper move: tokens-per-megawatt replaces FLOPs as the unit of AI economics. When power is the binding constraint in Stargate-scale datacenters, a 550W part that out-serves a 1400W part is not a faster chip — it is more datacenter.
更深的棋:AI 经济的单位从 FLOPs 换成 tokens-per-megawatt。当 Stargate 级机房里电力是硬约束,550W 的件干翻 1400W 的件,不只是“芯片更快”——它是凭空多出来的机房。
Closing — The Flywheel Closes
尾声|飞轮合拢
Nonprofit lab → ChatGPT → PBC → $852B pre-IPO → in-house 3nm inference silicon in 9 months. The Jalapeño story is not “OpenAI beat Nvidia.” It is that a model company became a silicon company because the model company knew the tensor shapes, the batch sizes, the latency tolerances, and the power ceiling better than any GPU vendor ever could — and then let its own models draw the transistors.
非营利实验室→ChatGPT→PBC→8520 亿准 IPO→9 个月出自研 3nm 推理硅。Jalapeño 的故事不是“OpenAI 打败英伟达”,而是一家模型公司之所以变成硅片公司,是因为它比任何 GPU 厂商都更懂自己的张量形状、batch 大小、时延容忍和电力天花板——然后让自家模型亲手画晶体管。
That is why it is strong. Not magic, not Moore’s Law, just the shortest feedback loop in semiconductor history: model → metal → model.
这才是它凶猛的根。不是玄学,不是摩尔定律施恩,只是半导体史上最短的反馈环:模型→金属→模型。
*Sources: OpenAI/Broadcom Jalapeño disclosure (Jun 24 2026), Hot Chips 2026 presentation (Aug 25 2026), SemiAnalysis InferenceX, DCD, CNBC/Brockman interview, OpenAI PBC restructuring filings 2025–2026.
此文由 怡心湖 编辑,若您觉得有益,欢迎分享转发!:首页 > 观·世界 » 从非营利实验室到硅基熔炉:OpenAI 的前世今生,与自研芯片为何如此凶猛
LRASM 弹载 AI 决策树(伪代码级)& B-
从"工业母机"到"技术主权":日本五轴
算力自主与地缘博弈:RISC-V如何成为
PyTorch是如何变成“AI界的POSIX”
CUDA的护城河与AI生态的开放之战 C
后GPU时代的基础设施之争:英特尔AI
隐形霸权与产业断层:日本科技的“咽
算力迁徙:AI时代下PC与移动互联网的