OpenAI has published the first test results for Jalapeño — its first custom inference chip (ASIC), developed jointly with Broadcom. At the Hot Chips conference, the company said that in SemiAnalysis's InferenceX benchmark, the chip outperforms NVIDIA GB200/GB300 superchips in energy efficiency and end-to-end latency. The figures are not yet independently verified, so this is a vendor claim rather than a proven result.

image
image

What happened

On August 25, 2026, at the Hot Chips conference, OpenAI presented measurable results for Jalapeño — a custom ASIC for inference, i.e., running already-trained models. The chip was developed jointly with Broadcom and was first shown in June 2026. According to SemiAnalysis's InferenceX benchmark on the GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T models, Jalapeño delivered 1.5–1.9× more AI work per watt and 1.7–3.6× lower end-to-end latency than the best previously recorded results on NVIDIA GB200/GB300 superchips. The chip's TDP is about 700 W — roughly half that of NVIDIA's flagship system. Small-batch deployment is planned for late 2026, with volume ramp-up in 2027; second- and third-generation chips are already in development.

Context

Inference — serving already-trained models — is the most expensive part of the AI economy: the cost of each request is determined by price per token and latency. Jalapeño is the first inference ASIC from a frontier-model developer to pass Hot Chips and be measured using a third-party methodology: InferenceX was developed by SemiAnalysis, which increases confidence in the figures compared to typical vendor claims. Important context — the relationship with NVIDIA: the company recently agreed to provide OpenAI up to $105 billion in financing for data centers, and OpenAI emphasizes that it does not intend to fully move away from partner GPUs. Jalapeño complements existing GPU infrastructure rather than replacing it, and it is a step toward OpenAI's vertical integration into inference hardware. Additionally: the tests were run not only on OpenAI models but also on open models from other developers, showing that the ASIC is not tied to a proprietary stack.

Why this matters for the industry

If the InferenceX figures are confirmed, OpenAI's inference cost will fall, directly strengthening the company's margin and its pricing position against API competitors, as well as its negotiating position against NVIDIA. For the industry as a whole, this is a strong signal: the model of “inference ASIC + training GPU” could become an industry standard, after which other major labs will follow OpenAI with their own custom chips. The effect is already visible at the expectations level: the narrative of falling inference costs is taking hold, and pressure on API pricing is starting before the chip reaches scale. By 2027, with volume ramp-up and the Gen 2 launch, NVIDIA's share of inference workloads could shrink noticeably.

Why this matters for users

There is nothing to integrate right now: there are no new public features, prices, or APIs for the chip, and Jalapeño is not yet in production, so for end users this is a vendor claim rather than a confirmed fact. In the future, this hardware will handle requests in ChatGPT and the OpenAI API: lower latency, more responsive agents, and voice modes are expected. When deployment begins, OpenAI's internal scenarios will see the effect first, followed by possible API price cuts, higher limits, and the first data on real latency. For developers, the value is already in planning: low-latency interactive scenarios and voice modes that were previously too expensive could become viable within the next one to two years.

What is still unknown / limitations

All figures were announced by OpenAI itself: the company published the results on its own, no fully independent re-measurements on GB200/GB300 have been provided at the time of publication, and the full methodology has not been disclosed. The comparison was made against the best previously recorded results on GB200/GB300 at the time of the test, but the baseline configurations — batch size, quantization and precision, software stack version, system topology — are not disclosed in the sources, so interpreting the 1.7–3.6× latency range is risky. The TDP of about 700 W is based on Tom's Hardware data, not OpenAI's technical datasheet. There are no public prices or APIs for the chip, so the metrics should be treated as vendor-claimed rather than proven.

Sources

Author

Look at AI, editorial team