product-updateproduct-update 5 min read

OpenAI Posts Jalapeño InferenceX Benchmarks at Hot Chips

OpenAI published Jalapeño's first custom-chip results on 25 August 2026 and presented them at Hot Chips the same day. On SemiAnalysis' InferenceX benchmark it claimed 1.5 to 1.9 times more AI work per watt than the best recorded Nvidia GB200 or GB300 systems.

PC

PromptCrates Editorial

Staff Writer

0 0
OpenAI Posts Jalapeño InferenceX Benchmarks at Hot Chips

OpenAI published Jalapeño's first custom-chip results on 25 August 2026 and presented them at Hot Chips the same day. On SemiAnalysis' InferenceX benchmark it claimed 1.5 to 1.9 times more AI work per watt than the best recorded Nvidia GB200 or GB300 systems.

What Jalapeño is, and where the InferenceX numbers came from

Jalapeño is an inference ASIC built with Broadcom. It is not a training-chip announcement. The 25 August post is the first public scorecard for that ASIC.

OpenAI hardware VP Richard Ho said it delivers both higher throughput and lower latency, a pair existing systems usually trade off. That sentence is the frame. One number that only wins watts, or one number that only wins lag, would not match the claim. The post tries to show both.

The yardstick is SemiAnalysis' public InferenceX benchmark. OpenAI said Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T than the best recorded Nvidia GB200 or GB300 systems.

Treat the table as vendor scores on a public bench, not a third-party audit of OpenAI's serving stack. InferenceX is SemiAnalysis' benchmark. The comparison systems are the best recorded Nvidia GB200 or GB300 results OpenAI cited. They are not a promise about every rack in the field.

If you already track Nvidia's 15% AI server price notices for 2027, keep that as a rack-BOM story. Jalapeño is OpenAI's own inference chip. One is a 2027 price letter to builders. The other is a 2026 scorecard at Hot Chips. They can both be true in the same week.

How Ho split throughput, latency, and interactive loads

The same InferenceX tests showed 1.7 to 3.6 times lower end-to-end latency across those three models, and 2.1 to 4.1 times higher performance on highly interactive workloads.

Those are three bands. Peak work per watt is 1.5 to 1.9 times. End-to-end latency is 1.7 to 3.6 times lower. Interactive performance is 2.1 to 4.1 times higher. Do not flatten them into one multiplier. A slide that says "up to 4.1x" without naming the interactive band is a different claim.

The 4.1x figure lives in the interactive band. The 1.9x figure lives in peak work per watt. The 3.6x figure lives in end-to-end latency. Quote the band with the number.

If your team already routes coding work through GPT-5.6 Sol, Terra, and Luna in AWS Kiro, do not put Jalapeño in that picker. Kiro is a coding-agent surface with its own credits. This post is about the serving chip under OpenAI's own infrastructure.

Small volumes in 2026, Nvidia stays, and why the chip moves less data

The Verge and TechCrunch both quote Ho saying Jalapeño will deploy in very small volumes by the end of 2026, with volume ramping in 2027. OpenAI did not say how many chips it will ship.

Very small volumes is the 2026 line. Ramp is the 2027 line. Missing unit counts are a fact, not a gap you should fill. Do not invent a wafer start, a rack count, or a megawatt figure.

Ho said OpenAI will not replace its entire chip lineup with Jalapeño and will keep using partners including Nvidia while it develops second and third generations. Generation one is the chip with numbers. Generations two and three are already in the roadmap. Nvidia remains in the mix, as the partners line is written.

That is an add, not a swap. Teams that budget as if OpenAI is leaving Nvidia this year are writing a sentence Ho did not say.

OpenAI said its own models helped design and program the chip, and that it minimizes data movement so KV cache and other model state can stay local across prefill and decode. Prefill and decode are the two phases of serving. Keeping KV cache local is the named method for not paying a movement tax between them.

Do not turn that into a floorplan. The 25 August note does not give a process node, a package wattage, or a memory capacity. It gives a design intent: less data movement; state stays local through both phases; models helped design and program the ASIC.

How to pin the InferenceX scorecard in a skill

Put the bands in the wrapper, not in a chat line that says "OpenAI is faster now." A skill prompt that collapses 1.5x, 3.6x, and 4.1x into one slogan is the pattern this post is meant to retire.

The skill should say: Jalapeño first results published 25 August 2026 and presented at Hot Chips; inference ASIC built with Broadcom; Richard Ho claims higher throughput and lower latency together; InferenceX on GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T versus best recorded Nvidia GB200 or GB300; 1.5 to 1.9 times more AI work per watt at peak; 1.7 to 3.6 times lower end-to-end latency; 2.1 to 4.1 times higher performance on highly interactive workloads; very small volumes by end of 2026, ramp in 2027; unit count not disclosed; Nvidia and other partners stay; second and third generations in development; models helped design and program the chip; KV cache and model state kept local across prefill and decode.

Do not write a 2026 fleet replacement. Do not write a public chip price. Do not write a wattage OpenAI did not publish in this announcement.

Sources

OpenAIJalapeñoInferenceXHot ChipsBroadcomNvidiaRichard HoSemiAnalysisGB200GB300

Related articles