On 25 August, OpenAI first disclosed the results of the self-study of the reasoning chip Jalalapeño. According to the company, in the open InferenceX benchmark using GPT-OSS 120B, the chip achieved a higher peak per kilowatt and lower token delay than the commercial system involved in the comparison; it also performed strongly on DeepSeek R1 and Kimi K2. OpenAI indicated that it now has a working first chip and that subsequent generations are being developed.

This release of an accurate boundary is the “first measurement results”, not that Galapeño has already laid all data centres on a large scale, let alone that training and reasoning have migrated to self-study hardware. The public articles do not give the size of the production, process, single-card power, memory capacity, system sale or full deployment schedule, nor do they indicate the full configuration of the comparative system. It can be confirmed that OpenAI has moved from the planning stage of the design of models, service software, chips, memory and network co-designs to the detectable silicone phase; what is not confirmed is that it has generally taken a lead in the cost of real total business ownership.

The peak cannot be a substitute for a constant load while the token delay is visible for each kilowatt of vomiting.

The reasoning chip competition is often simplified to produce more tokens per second, but large-scale services are also constrained by first token time, token-to-token delays, co-processing users, batch size, model accuracy, context length, memory bandwidth and power consumption. OpenAI places particular emphasis on the expansion of throughput with the established token delay and the measurement of energy efficiency per kilowatt over time, suggesting that Jalalapeño is aimed at the economic reasoning of continuous operation rather than the pursuit of a single peak.

The use of GPT-OSS 120B as an open benchmark facilitates external comparisons, as the subject of the test is not a private model not available within the company; the results of DeepSeek R1 and Kimi K2 are used to demonstrate that optimization is not just a model family. However, press releases do not synchronize with the release of sufficiently detailed raw data, software versions, precision models, batch parameters and duplicate experimental scripts. The ranking of different systems may change if there are inconsistencies in the definition of quantitative accuracy, delayed target or hang power. It would therefore be more appropriate at this stage to consider the outcome as the first party statement of performance, pending a re-examination of the InferenceX entry and third parties.

Co-optimization of chips with models reduces redundancy in general-purpose hardware to compatible loads, but also presents life-cycle risks. Model structures usually change faster than the chip design and manufacturing cycle, and hardware that is optimized today for attention, mixed experts or specific memory access can face different calculations in the future. OpenAI has to prove not only that the first chip can run fast, but also that the compiler, running time and the software warehouse can keep up with model changes.

Energy efficiency is not only equal to chip effort. The hangar network, memory, cooling, power conversion and low-utilization standby will affect the actual token consumption in the data centre. The peak is an important indicator, but does not directly extrapolate the average annual cost. The real determination of commercial value is whether the system can maintain the same advantages under genuine request fluctuations, long context, utility calls and multiple model paths.

Self-study chips add supply options and bargaining power, not end cooperation.

OpenAI has positioned Jalalapeño as a credible first-party path beyond the partner accelerator. The company clearly listed the underlying role of Microsoft and NVIDIA for its growth and indicated that the calculation combination also included AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank. Different types of training, bulk reasoning, low-delayed reasoning and permanent intelligent missions require different hardware, and enterprises place loads in the system that best matches performance and cost, rather than just a chip.

For OpenAI, the strategic value of self-study chips is three layers. First, to optimize its own high-frequency reasoning pattern and reduce the cost of marginal services; second, to increase controllable capacity routes in times of supply constraints; and third, to maintain procurement bargaining power through alternative options. Even if the percentage of Galapeño is low, the price of external accelerators and the pace of deployment may be affected as long as it is able to carry enough stable loads.

This control remains dependent on manufacturing, sealing, HBM, network and data centre construction partners. The so-called “self-study” usually means that the architecture and system design are company-driven and do not mean that all physical links are completed by OpenAI. Supply chain probabilities, software migration costs and new generational overlaps will determine whether it can move from measuring chips to stable production. The article does not disclose the date of production, so that the words “work silicones already exist” should not be written as “full commercial”.

For cloud services and API users, Galapeño does not necessarily change the product interface in the short term. Higher energy efficiency may create space for lower delays, higher quotas or lower unit costs, but transmission to prices is determined by a combination of capacity, demand and product strategies. Users should focus more on service levels, model consistency and privacy boundaries than on the assumption that self-study chips will automatically lead to price reductions.

The real change in this disclosure is the position of OpenAI in the calculator: It is no longer just a large accelerator buyer, but is beginning to possess measurable first-party reasoning hardware. What needs to be tested in the next phase is not the “leading” in advocacy, but the re-emergence of open benchmarks, the long-term stabilization of real clusters, the scale of deployment, and the ability of companies to turn self-study hardware into a lasting rather than a one-off economic lever, while continuing to use multiple partners.

Source: OpenAI, The Full Stack behind anticipatory, 25 August 2026, https://openai.com/index/the-Full-stack-behind-abundant-intelligence/