hraness

saved

OpenAI Jalapeño: Better Than Nvidia Blackwell

by Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee and Dylan PatelSemiAnalysispublished

gist

SemiAnalysis visited OpenAI's lab and verified InferenceX runs of Jalapeño, a Broadcom-partnered inference ASIC designed in about 16 months. On tokens per megawatt, its single-token-prediction results beat Blackwell and published Vera Rubin multi-token-prediction numbers; TCO per token is roughly even with Rubin until Jalapeño adds speculative decoding. The authors call it a generalized inference chip, not an OpenAI-only specializer, and note the figures are OpenAI-supplied 8k1k runs without AgentX. Production is scheduled to ramp through 2027.

ideas

  • It is a general inference chip, not an OpenAI-model specializer. SemiAnalysis ran InferenceX on DeepSeek R1, Kimi K2.5, and GPT-OSS in the lab; media overread comments about model-specific tuning.
  • Compare it to Rubin, not Blackwell. Both use HBM4, Vera Rubin is already shipping, and Jalapeño is still engineering samples, so beating Blackwell is the expected bar.
  • Power is the scoreboard. OpenAI is datacenter-power limited, so tokens per megawatt matter more than floorspace or budget; Jalapeño STP already beats Rubin's published MTP tok/MW.
  • Keep the fleet homogeneous. OpenAI skips prefill/decode disaggregation so the mix of input, cache, and output tokens can change across model eras without stranding a specialized pool.
  • Software codesign is the lead. A clean-sheet Gluon/Teacup stack plus Codex kernel search produced fast bring-up; the authors treat that pace as pressure on the CUDA software moat.

quotes

Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong

Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel, rejecting the specialized-chip story.

Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve.

Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel, stating the headline InferenceX result.

we believe that comparison to Blackwell is somewhat incomplete and unfair.

Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel, limiting the Blackwell comparison.

The CUDA moat is potentially dead given how fast OpenAI can bring up new models on their silicon.

Bryan Shan, Myron Xie, Jordan Nanos, Wega Chu, Clara Ee, and Dylan Patel, judging the software bring-up.