OpenAI Jalapeño Chip Delivers 1.9x More Work Per Watt Than Nvidia GB300, Benchmarks Show

Reviewed byNidhi Govil

7 Sources

Share

OpenAI unveiled benchmark results for its custom AI inference chip Jalapeño at Hot Chips conference, showing it outperformed Nvidia GB300 with 1.5-1.9x more work per watt and significantly lower latency. Developed with Broadcom, the chip deploys late 2026 in small volumes, ramping up in 2027.

OpenAI Unveils Jalapeño Chip Performance at Hot Chips Conference

OpenAI shared detailed benchmark results for its custom AI inference chip Jalapeño at the Hot Chips conference on Tuesday, revealing performance metrics that position it ahead of current market leaders. The OpenAI Jalapeño chip, developed with Broadcom, demonstrated 1.5 to 1.9 times more AI work per watt at peak throughput compared to Nvidia's GB200 and GB300 Blackwell system.

1

2

Richard Ho, OpenAI's head of hardware, described the results as showing "a very, very significant performance advance over state of the art" during a press briefing.

1

Tested using SemiAnalysis's InferenceX benchmark platform, the custom inference chip delivered faster AI responses across three public models: GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T.

2

5

The AI inference chip achieved 1.7 to 3.6 times lower end-to-end latency across these models, with ultra-low-latency inference showing 2.1 to 4.1 times higher performance for highly interactive workloads.

3

5

Power Efficiency and Technical Architecture Drive Performance

The custom AI chip operates at 700 watts, though measured sustained power remained at or below 550 watts during tested workloads.

5

This power efficiency represents a critical advantage, as electricity constitutes the dominant running cost in data centers.

4

Ho emphasized that Jalapeño offers "the best of both worlds" by combining lower latency and higher throughput, capabilities that AI systems typically must trade off between.

2

Source: TechCrunch

Source: TechCrunch

Each Jalapeño system features 128 accelerators delivering 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 memory, and nearly 2 petabytes per second of memory bandwidth.

3

The rack-scale architecture mirrors designs from Nvidia's NVL72 and AMD Helios systems. Individual accelerators produce 13.4 petaFLOPS at MXFP4, fed by 216 GB of HBM4 memory delivering 15.4 TB/s of memory bandwidth.

3

Designed to Minimize Data Movement and Communication Delays

OpenAI engineered Jalapeño specifically to address bottlenecks in AI inference processing, particularly during prefill and decode phases. The chip minimizes data movement by keeping model state, including the key-value cache used during response generation, local to processing resources.

1

5

"We designed Jalapeño to minimize data movement and communication delays," OpenAI explained, noting the system activates the right combination of compute, memory, and networking for each inference phase.

1

The architecture handles both compute-intensive prefill operations and memory-bandwidth-dependent decode operations within the same design, avoiding the typical tradeoff between throughput and latency.

5

This becomes particularly relevant for AI agents performing sequential inference steps, where small delays compound into longer overall task times.

5

AI-Assisted Development and Deployment Timeline

OpenAI employed its own AI models during chip development, using them to explore implementations, shorten design cycles, and optimize arithmetic circuits.

5

The team moved from initial design to tapeout in nine months, then used Codex with GPT-Astra to bring three open-weight models to high performance on Jalapeño within two months.

5

For selected GPT-OSS attention and mixture-of-experts blocks, AI-generated implementations ran 1.5 to 1.8 times faster than human-written code.

5

Source: The Next Web

Source: The Next Web

Ho estimated Jalapeño would deploy at the end of 2026 "in very small volumes," with more significant deployment coming in 2027.

1

2

OpenAI plans to make Jalapeño a multigenerational platform, with second- and third-generation designs already in development.

1

5

Competition Context and Strategic Positioning

While Jalapeño outperformed Nvidia GB300 in testing, it was not evaluated against Vera Rubin, Nvidia's newer generation that recently began shipping.

4

The comparison also excluded speculative decoding techniques that can boost inference performance.

3

AMD's MI455X and Nvidia's Rubin GPUs, expected to ramp production in early 2027, are optimized for both training and AI inference, whereas Jalapeño focuses exclusively on inference.

3

Ho emphasized OpenAI's continued partnership with Nvidia, stating "Nvidia is a really good partner, and we continue to need a lot of Nvidia."

4

OpenAI still requires compute for training, and the programmable nature of GPUs means the company will likely deploy on AMD and Nvidia first before transitioning to in-house silicon.

3

Every major AI company is pursuing similar strategies, with Anthropic designing custom silicon and Google discussing custom inference chips with Marvell.

4

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved