The next AI speed race may be won at inference—not training.

OpenAI says Jalapeño, its first custom inference chip co-developed with Broadcom, delivered 1.5–1.9× the peak throughput per kilowatt and 43–72% lower end-to-end latency than selected Nvidia GB200 and GB300 systems across three open models.

Jalapeño delivered more output per watt

Across three open-model tests, OpenAI reports 1.5–1.9× the peak throughput per kilowatt of the named Nvidia comparison systems.

These are OpenAI-published results from pre-deployment hardware on three fixed workloads, not an independent production evaluation or a general result for every Nvidia system.

View benchmark values and method
Jalapeño delivered more output per watt
ModelPeak throughput per watt
GPT-OSS 120B1.9
DeepSeek R1 670B1.7
Kimi K2.5 1T1.5
OpenAI: Jalapeño’s first results OpenAI compared peak mixed tokens per second per kilowatt on nominal 8K-input, 1K-output InferenceX workloads, using published package power ratings for normalization.

More work. Less waiting.

Inference is the part of AI that users feel. Latency shapes how quickly an answer starts and finishes; throughput determines how many answers the same power budget can serve.

Those goals often pull in opposite directions. In these tests, OpenAI reports that one architecture improved both. At ChatGPT scale, that can turn chip design into a product advantage: faster interactions, more capacity, and less energy per unit of AI work.

A first result, not a final verdict

The numbers come from fixed InferenceX workloads with a nominal 8,000-token input, 1,000-token output, and single-token prediction. OpenAI supplied and published the results from pre-deployment hardware; they do not establish that Jalapeño leads every Nvidia system, model, batch size, or production workload.

OpenAI says deployment begins by the end of 2026 while qualification and software work continue. It also expects to keep using commercial accelerators. The useful signal is narrower: a first-generation chip tuned around one operator’s inference stack can already compete on both speed and energy.

Blackwell is not suddenly obsolete. But inference is becoming a full-stack contest—and custom silicon can now shape the experience users notice, not just the bill behind it.

Sources and boundaries

The benchmark figures are OpenAI-reported and workload-specific. The source pages below provide the data, test method, and deployment context.