OpenAI says Jalapeño tops current inference chips

At Hot Chips, OpenAI shared early Jalapeño benchmarks against Nvidia Blackwell and said the chip is built to cut latency and power use.
OpenAI has shared the first benchmark results for Jalapeño, its new chip for AI inference, at the Hot Chips conference. The company says the system delivered more tokens per user and higher throughput per kilowatt than the current state of the art in this category. That matters because inference is the part of an AI system that generates answers for users, so speed and power use directly shape how efficiently a service can run. OpenAI is positioning Jalapeño not just as a high-performance chip, but as one built to serve many customers quickly while keeping latency low.
OpenAI pits Jalapeño against Nvidia Blackwell
Richard Ho, OpenAI’s head of hardware, said in a press call that the results show a “very, very significant performance advance over state of the art.” He said Jalapeño can handle more AI work per unit of power while also returning responses more quickly.
OpenAI compared the chip with an Nvidia Blackwell system. That makes the claim notable, since Nvidia hardware remains a key reference point in AI infrastructure. Still, as IT-PUB News notes, OpenAI also said Jalapeño is not expected to reach full deployment until later, which leaves time for the competitive picture to change.
Ho said Jalapeño would be deployed at the end of 2026 “in very small volumes,” with more substantial deployment expected in 2027.
Broadcom helped build the chip around OpenAI's models
Jalapeño was first announced last October. OpenAI developed it in close collaboration with Broadcom, and the company said its own models helped during the development process.
OpenAI also described Jalapeño as part of a multigenerational platform. In practice, that means AI products, models, chips, and memory would be developed together rather than as separate pieces. The goal is to optimize the whole system instead of improving only one component at a time.
That matters because AI performance depends on more than raw compute power. If the surrounding hardware and software are not aligned, systems can lose time moving data around or waiting for different parts to communicate.
OpenAI says Jalapeño targets two common bottlenecks
OpenAI said Jalapeño is designed to reduce friction in two phases that often slow inference: prefill and communication. Prefill is the stage where the system processes the input before generating a response, while communication delays can become a bottleneck when different parts of the system need to exchange information.
According to the company, the chip is meant to minimize data movement and communication delays. OpenAI said model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase.
The point is practical efficiency, not just strong lab results. For businesses running AI products at scale, lower power use and lower latency can matter as much as headline speed numbers because they affect operating costs and user experience at the same time.
OpenAI's timeline leaves room for rivals to catch up
Even with the benchmark gains, Jalapeño is still well short of broad deployment. OpenAI’s own timeline puts a small-volume rollout at the end of 2026, with larger deployment in 2027. By then, rival chipmakers and AI infrastructure providers may have moved ahead too.
So the announcement is significant, but it does not settle the competitive race. OpenAI is signaling that it wants more control over the hardware behind its AI systems and that it is willing to design chips around specific inference bottlenecks rather than rely entirely on off-the-shelf processors.
For users, the effect would be indirect but potentially meaningful: faster responses, more efficient AI services, and systems that can handle more demand without using as much power. For the broader AI market, Jalapeño is another sign that major AI companies are pushing deeper into the hardware layer that supports their software.