OpenAI claims its new custom inference chip, Jalapeño, delivers better throughput per watt and lower token latency than Nvidia’s Blackwell and Rubin processors.
The device handles inference only, meaning it runs AI models but does not train them. It is a general-purpose large language model accelerator, not tuned specifically for OpenAI’s own software.
According to the company, Jalapeño provides 1.5x to 1.9x more AI work per watt at peak throughput across three tested models. End-to-end latency is 1.7x to 3.6x lower than the best commercially available systems. For interactive workloads, performance is stated to be 2.1x to 4.1x higher.
The figures come from SemiAnalysis’s public InferenceX benchmark. OpenAI supplied the numbers, while SemiAnalysis verified some runs in its own lab. The models tested were GPT-OSS 120B, Deepseek R1 670B, and Kimi K2.5 1T.
On GPT-OSS, Jalapeño achieved roughly 1,400 tokens per second per user. On Deepseek R1, it exceeded 700 tokens per second on a single concurrent request.
Jalapeño posted these numbers without using techniques like multi-token prediction or speculative decoding. Some comparison systems relied on those optimizations, leaving room for further improvement.
SemiAnalysis wrote that Jalapeño “smokes every other chip” in its headline performance-per-watt comparison. SemiAnalysis CEO Dylan Patel added that first generation chips are usually not competitive, yet OpenAI is beating Nvidia Blackwell and even Rubin.
SemiAnalysis notes that the fairer comparison is not Blackwell but Nvidia’s newer Vera Rubin platform, since both use HBM4 memory. Even here, Jalapeño squeezes out more output tokens per megawatt than Vera Rubin, despite Nvidia’s accelerator using the multi-token prediction optimization that Jalapeño has not adopted yet. On total cost of ownership per token, the two come out roughly even.
There are caveats. Nvidia and AMD have already published results with larger models like Deepseek V4 Pro and Kimi K3 that have not been tested on Jalapeño yet. While Rubin systems are already shipping to customers, Jalapeño reportedly has not moved beyond engineering samples.
Development timeline and strategy
OpenAI developed Jalapeño with Broadcom. Design work began in mid-2024, and the final design went to fabrication in November 2025. The full cycle took about 16 months, but OpenAI states only nine months passed between the first chip design and the finished blueprint heading to the factory. The company used its own AI models during development. Older model generations assisted with chip design, while newer ones sped up programming and optimization.
SemiAnalysis views this as a sign that Nvidia’s much-discussed “CUDA moat” may not hold anymore. The firm wrote that the moat is potentially dead given how fast OpenAI can bring up new models on their silicon.
OpenAI CFO Sarah Friar says the chip fits into a broader compute strategy where data centers, chips, models, the developer platform, products, and devices all work as one integrated system. She claims Jalapeño complements OpenAI’s existing partnerships with Nvidia, AMD, AWS, Cerebras, CoreWeave, and others rather than replacing them.
OpenAI has deep ties with several of these companies. Nvidia, AMD, and AWS are all investors or compute partners, with Nvidia being one of the largest. Each of them is also building its own AI chips, which makes the relationship both cooperative and competitive. That said, all of these companies keep saying the world cannot have enough compute, a claim that conveniently supports their own business models.
What it means
For people building AI applications, the practical change is faster response times and lower energy costs for running models. OpenAI’s claim suggests an independent inference engine that does not rely on Nvidia’s proprietary software stack. This could lower barriers for competitors who want to run large models without paying Nvidia’s licensing fees. However, the chip is not yet available for purchase, so these benefits remain theoretical until Jalapeño moves from engineering samples to production hardware.




