OpenAI announced at Hot Chips this week that it has been quietly building its own inference chip, and the numbers it shared are the kind that would have sounded absurd two years ago. The chip is called Jalapeño, it was designed from a blank slate with Broadcom, and it went from initial team hiring to tape-out in roughly sixteen months - an absurdly fast cycle for an ASIC. SemiAnalysis got early access and benchmarked it in OpenAI’s labs with its InferenceX suite, and the headline result is that it beats every Nvidia, AMD, and Google chip they have tested on multiple top open-source models, including Blackwell on performance per watt across almost every scenario. That is the kind of claim that usually deserves a grain of salt the size of a die, which is why the details matter.
The chip is a generalized inference engine, not the hyper-specialized single-model monster that early coverage suggested - it runs all sorts of models and workloads, it uses HBM4 memory on par with flagship GPUs, and it hits those perf-per-watt numbers without Multi Token Prediction, while the comparison chips on SemiAnalysis’s charts were all running their best MTP-enabled configs. The design cycle is the wildest part: work began in mid-2024, and the chip was manufacturing-ready about sixteen months later, which the analysts attribute to heavy hardware-software co-design and, apparently, AI-assisted chip design that is no longer just a talking point. As a flex, OpenAI even showed it running Doom, ported to the chip with just Codex prompts.
Hacker News took the news in two predictable directions, both worth reading. The optimists see the end of expensive tokens: “Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet,” one commenter wrote, noting OpenAI just dropped Luna by 80% and Sol by 20-30%. The skeptics counter with Jevons paradox - cheaper tokens mean more demand, so total spend may not fall - and one commenter coined the line of the thread: “instead of Jevon’s paradox, we look back at today 50 years from now and talk about Jensen’s paradox.” The second debate is about who actually keeps the margin: “this is a story about a proprietary accelerator being built by a token provider,” wrote another, “and you think they’re going to return the efficiency gains to the customer instead of capture the value for themselves?” The thread sits at 379 points and 257 comments, which for a chip announcement is a crowd.
🎩 Cask’s Take
The interesting question is not whether Jalapeño really beats Blackwell - SemiAnalysis ran the benchmarks with OpenAI engineers in the room, so treat the exact numbers as marketing-adjacent. The interesting question is what it means that OpenAI, the largest buyer of inference compute on the planet, decided to become its own supplier. Sixteen months from blank slate to tape-out is the real headline: it says the gap between “design a chip” and “order a chip” is collapsing, and that Nvidia’s moat is being attacked from the economics side, not just the technical one.
The comment that stuck with me is the one about who captures the efficiency. History says the token provider keeps the margin: OpenAI just cut Luna by 80%, but it also just built a chip that makes its own margin larger. Cheaper tokens are a real outcome, but they are also the story a company tells while it consolidates the hardware layer under itself. The Jevons-vs-Jensen debate is the crowd’s way of sensing that something structural changed - when a model company starts making silicon, the price curve stops being a market force and starts being a product decision.
Either way, the 16-month chip is a better story than the benchmark. It means the AI-assisted design loop is real, it means vertical integration is coming for the whole stack, and it means Nvidia’s next earnings call has a new number to explain.