OpenAI Reveals Jalapeño Chip at Hot Chips: Benchmark Results Show Major Performance Leap
By admin | Aug 25, 2026 | 2 min read
At the Hot Chips conference on Tuesday, OpenAI offered a deeper dive into Jalapeño, unveiling the first benchmark results for the system. When tested using Semianalysis’s InferenceX benchmark, Jalapeño delivered more tokens per user and greater throughput per kilowatt compared to the leading inference processors currently on the market. “The bottom line is that the results show a very, very significant performance advance over state of the art,” said Richard Ho, OpenAI’s head of hardware, during a press call. “Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It’s very efficient to serve a lot of customers, but it can also be very low latency.”
Worth noting, that comparison is against an Nvidia Blackwell system—but by the time Jalapeño reaches full deployment, the competitive landscape may have shifted considerably. Ho estimated that Jalapeño would start rolling out “in very small volumes” at the end of 2026, with more substantial deployment following in 2027. Initially announced last October, Jalapeño was developed in close partnership with Broadcom, with OpenAI’s own models playing a role in the design process. The company intends for Jalapeño to be a multigenerational platform, enabling AI products, models, chips, and memory to be developed in tandem. This full-stack approach allowed OpenAI to tackle specific phases of the inference process that often create friction during processing. In particular, Jalapeño is engineered to reduce delays during the prefill and communication stages, which OpenAI identifies as frequent bottlenecks. “We designed Jalapeño to minimize data movement and communication delays,” the company stated in a blog post presenting the findings. “This means that model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.”
Comments
Please log in to leave a comment.
No comments yet. Be the first to comment!