OpenAI's Jalapeño chip outperforms Nvidia Blackwell in AI inference benchmarks, deployment begins 2026

Reviewed byNidhi Govil

15 Sources

Share

OpenAI revealed detailed benchmarks for its Jalapeño custom AI chip at Hot Chips conference, showing 1.5-1.9x better performance per watt and 1.7-3.6x lower latency than Nvidia's Blackwell systems. Developed with Broadcom, the chip deploys in small volumes by end of 2026, with full production ramping in 2027.

News article

OpenAI Jalapeño Chip Delivers Superior AI Inference Performance

OpenAI shared comprehensive benchmark results for its Jalapeño custom AI chip at the Hot Chips conference on Tuesday, demonstrating significant performance advantages over current state-of-the-art AI accelerator systems

1

. Testing on Semianalysis's InferenceX benchmark revealed that the Jalapeño chip delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T models compared to Nvidia Blackwell systems

2

. Richard Ho, OpenAI's head of hardware, emphasized that the results show a very significant performance advance, with Jalapeño capable of serving more AI inference workloads per unit of power while returning responses more quickly

1

.

The custom AI chip achieved 1.7 to 3.6 times lower end-to-end latency across the three tested models, meaning users will experience faster responses, more responsive agents, and more reliable access as demand grows

2

. For ultra-low-latency inference workloads, OpenAI claims its chips deliver 2.1 to 4.1 times faster performance

4

. Ho stated that Jalapeño offers the best of both worlds with lower latency and higher throughput, as AI systems typically have to make a trade-off between the two

2

.

Broadcom Collaboration and Deployment Timeline

First announced in June, the Jalapeño chip was developed through close Broadcom collaboration, with OpenAI's own models assisting in the development process

1

. The Application-Specific Integrated Circuit (ASIC) is designed specifically for AI inference—the process of running a trained AI model to complete a task or deploy an agent

2

. OpenAI plans to deploy Jalapeño in small volumes by the end of 2026, with more significant deployment ramping up throughout 2027

1

2

. The company is already working on second and third generations of the chip, planning to make Jalapeño a multigenerational platform that allows AI products, models, chips and memory to be developed in concert

1

.

Technical Architecture Optimizes Data Movement

The Jalapeño AI accelerator features a massive design with 216 GB of HBM4 memory and up to 15.4 TB/s of bandwidth, delivering up to 3.4 MXFP8 PFLOPS and up to 13.4 MXFP4 PFLOPS at 700W

3

. The chip operates at 1.70 GHz on silicon already running in OpenAI's labs, with plans to increase clocks to 1.80 GHz

3

. At rack-scale architecture level, each system with 128 accelerators packs 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 memory, and just shy of 2 petabytes per second of memory bandwidth

4

.

OpenAI designed Jalapeño to minimize data movement and communication delays using a memory-sliced, NUMA-style architecture

3

. The chip has 64 core slices, each paired with its own HBM slice to guarantee predictable latency and bandwidth while avoiding conflicts associated with unified memory subsystems

3

. This arrangement lets frequently used operands remain close to compute resources that need them. Model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase

1

4

.

Competitive Landscape and Market Implications

Analysts view the OpenAI Jalapeño chip as a threat to Nvidia margins in the fast-growing inference market. Adrien Sanchez, technology analyst at Yole Group, noted that a hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency, representing a threat to Nvidia's inference margins, which is the field growing the most at the moment

5

. The custom silicon could reduce OpenAI's reliance on Nvidia over time for inference workloads, though Nvidia GPUs will remain important for compute-intensive workloads like large-scale model training given their broad programmability, performance, and software ecosystem

5

.

SemiAnalysis noted that while Jalapeño beat Blackwell on performance per watt in nearly all tested scenarios, the comparison is somewhat incomplete because Jalapeño uses newer HBM4 memory technology

5

. Nvidia's Rubin platform, which also uses HBM4 memory, represents a more like-for-like comparison, with Vera Rubin systems already shipping to customers while OpenAI still has time before anything beyond engineering samples of Jalapeño becomes available

5

. Despite this, Ho confirmed that OpenAI doesn't expect to replace its entire chip lineup with Jalapeño, saying its overall compute strategy includes very good partners like Nvidia

2

. OpenAI still needs compute for training, and the highly programmable nature of GPUs means deployment on AMD and Nvidia first, then transition to in-house custom silicon later

4

. Omdia expects custom ASIC chips like Jalapeño to exceed GPUs in volume by 2028, representing the biggest competitive threat to Nvidia as about half the capital expenditure on AI infrastructure comes from hyperscale cloud providers who either have a custom chip program or could reasonably have one

5

.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved