OpenAI's 'Jalapeno' chip clocks 104× on open-weight models — and Nvidia should be nervous
OpenAI just showed the first real results from its in-house inference chip, and the numbers change the calculus on who owns AI compute.
The open-weight vs closed-weight story just took a hard turn. This week OpenAI showed off the first performance numbers from its in-house AI inference chip — nicknamed Jalapeno — and Matt Wolfe walks through why it's more than a spec-sheet flex.
The headline: on public open-weight models, Jalapeno is hitting up to 104× the performance vendors are getting today. OpenAI deliberately benchmarked on public models so competitors can verify the numbers on their own hardware — a level of receipts that's rare in this space.
The strategic layer is more interesting than the silicon. OpenAI has spent years being one of Nvidia's biggest customers; a working in-house chip is the first credible step toward not being a customer at all. That reshapes pricing power for every AI lab, and it puts a floor under how far Nvidia's margin story can stretch.
For anyone building on top of OpenAI's API, the near-term read is: inference gets cheaper, faster, and stickier — a mild tailwind. The longer read is that "who owns the compute" is going to matter more than "who owns the model," and the labs that control both are the ones setting terms in 2027.