OpenAI disclosed that its self-study of the AI chip Jalalapeño was primarily oriented towards a large-scale reasoning scene, emphasizing the processing of more requests and the reduction of response time at unit cost. Richard Ho, head of the company ' s hardware, stated that the test results showed a significant improvement in performance over the existing programme.
Small deployment by 2026
According to OpenAI, the current performance comparison of Jalalapeño includes Nvidia Blackwell. However, there is still a time gap between the chip and its full spread. Richard Ho anticipates that Galapeño will be deployed at a very small scale by the end of 2026, with a wider application to 2027.
This means that by the time the chips actually enter a large-scale go-live phase, the competition in the market may also have continued to overlap, at which point the actual pattern of competition remains variable.
Development in cooperation with Chase
Galapeño was first published last October and developed by OpenAI in cooperation with Chase. OpenAI also indicated that the company ' s own model was involved in the chip development process.
OpenAI plans to make Jalalapeño a multigenerational platform, not just a single chip. Along these lines, future AI products, models, chips and memory will be designed as synergistically as possible to enhance overall system efficiency.
Focus on optimization of reasoning bottlenecks
According to OpenAI, this whole-pool approach allows companies to optimise specific bottlenecks directly in the reasoning process, in particular the prefilling and communication phases. According to the company, these two links tend to slow down the pace of reasoning.
In a blog, OpenAI states that one of the design objectives of Jalapeño is to reduce data handling and communication delays. To this end, the model state, including the KV cache used to generate the response, can be placed more clearly locally and appropriate computing, memory and network resources can be mobilized at different stages of reasoning.
From the information disclosed, Galapeño ' s core selling point is not just a single-point calculus, but an overall optimization around delays in the reasoning process, energy efficiency and system synergy. This also reflects the fact that, as the generation of AI services expands, the chip competition is further extending from training capacity to reasoning costs and speed of response.
