15 Sources
[1]
OpenAI's Jalapeño chip is built for fast inference at scale, benchmarks show
At the Hot Chips conference on Tuesday, OpenAI shared a more detailed look at Jalapeño, including the first batch of benchmark results for the new system. Tested on Semianalysis's InferenceX benchmark, Jalapeño registered both more tokens per user and more throughput per kilowatt than the currently
[2]
OpenAI says its Jalapeño chip can power faster AI responses than the competition
OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to a blog post published on Tuesday. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the "best of both worlds" with
[3]
Hot Chips 2026: OpenAI's Jalapeño AI ASIC unpacked -- accelerator developed using AI achieves efficiency and throughput gains against power-hungry Blackwell
OpenAI made quite a splash back in June, when it unveiled its 'Jalapeño' AI accelerator and revealed that the chip reached tape-out in just nine months. At the Hot Chips conference, OpenAI disclosed more details about the architecture of its Jalapeño inference processor as well as shared its target
[4]
OpenAI's upcoming Jalapeño chip looks like it'll be an inference beast
OpenAI offered its closest look yet at its spicy new Jalapeño AI accelerator at the annual Hot Chips semiconductor development conference at Stanford on Tuesday. The chips, first teased earlier this year, were developed in collaboration with Broadcom, and are the first in a series of custom
[5]
OpenAI's Jalapeño AI chip brings new 'threat' to Nvidia margins as custom silicon gains ground
* OpenAI's first custom AI chip, Jalapeño, is expected to begin deployment in the company's computing infrastructure by the end of the year. * Analysts told CNBC that it could put pressure on Nvidia in the fast-growing inference market and reduce OpenAI's reliance on Nvidia for some workloads. *
[6]
OpenAI's Jalapeno chip outperformed the GB300 on power and speed, according to OpenAI
Jalapeno leads the GB300 on work per watt and response speed at 700 watts, though it was not tested against Vera Rubin and cannot train models at all OpenAI says its Jalapeno inference chip, developed with Broadcom, outperformed Nvidia's GB300 on AI work per unit of power and on response speed. It
[7]
OpenAI chip beats Nvidia systems with 1.9x more work per watt
OpenAI says its first custom inference chip, Jalapeño, can deliver more AI work per watt while also reducing response times, pointing to a hardware design aimed at handling increasingly demanding model workloads more efficiently. The company tested Jalapeño across three public models: GPT-OSS
[8]
OpenAI says its Jalapeño chip bests Nvidia, others
Why it matters: It's part of a trend of AI model creators developing their own chips to cut costs, meet demand and reduce dependence on industry giant Nvidia. State of play: OpenAI said that the benchmarks show Jalapeño outperforming chips from Nvidia and others when running DeepSeek R1, the
[9]
OpenAI Jalapeño chip beats Nvidia GB300 in benchmark tests
The results were presented at the Hot Chips conference at Stanford University. Richard Ho, OpenAI's head of hardware, said on a press call that Jalapeño delivers greater AI output for every watt consumed while also cutting the time users wait for responses. The chip was tested using SemiAnalysis'
[10]
OpenAI says its new Jalapeno chip can make AI faster while using less power
OpenAI unveiled its new Jalapeno AI chip at a recent conference. This processor significantly boosts AI response speeds and reduces power consumption. The chip delivers improved performance per watt and lower end-to-end latency. OpenAI plans to deploy Jalapeno in its infrastructure by the end of
[11]
OpenAI's First-Gen Jalapeno ASIC Blows Competition Out Of The Park, Performs 1.5x to 1.9x More Work Per Kilowatt Than NVIDIA's Blackwell Chips, While Threatening The CUDA Moat
OpenAI seems to have used its current spate of models that run on NVIDIA GPUs to design the Jalapeno chip, and paired it with its own Gluon kernel programming language, thereby daring to threaten NVIDIA's legendary CUDA moat. How the tables have turned! The architecture of OpenAI's Jalapeno
[12]
Hot Chips: OpenAI's Jalapeño First Results Place it Ahead of Nvidia Blackwell
Open-AI brought out its first custom-designed chip in collaboration with Broadcom with a clear focus in inference processing Hot Chips Symposium, a conference held annually at Stanford activity that brings together companies in the field of electronics, turned out to be the best location for
[13]
OpenAI Jalapeño Chip: Up to 4.1x Faster AI Performance
OpenAI unveiled results for Jalapeño, its custom AI processor developed with Broadcom, on August 25, 2026. The targets faster AI inference through specialized hardware, memory and networking. OpenAI plans deployment inside its computing infrastructure by the end of 2026. The company says the
[14]
OpenAI's Jalapeño chip puts the pressure on Nvidia in inference
According to SemiAnalysis, Jalapeño beats Nvidia's Blackwell systems on performance per watt in nearly every scenario tested. The comparison remains imperfect, however, because OpenAI's chip uses HBM4 memory, bringing its profile closer to Nvidia's next Rubin platform. Jalapeño is expected to begin
[15]
OpenAI claims its Jalapeno chip outperforms Nvidia GB300, promises faster and cheaper AI
The company is already working on a second-generation Jalapeno chip. OpenAI has claimed that its Jalapeno AI chip has outperformed Nvidia's current systems in key tests. The chip was tested for how much AI work it can handle for the amount of power it uses, as well as how quickly it can respond to
Share
Copy Link
OpenAI revealed detailed benchmarks for its Jalapeño custom AI chip at Hot Chips conference, showing 1.5-1.9x better performance per watt and 1.7-3.6x lower latency than Nvidia's Blackwell systems. Developed with Broadcom, the chip deploys in small volumes by end of 2026, with full production ramping in 2027.

OpenAI shared comprehensive benchmark results for its Jalapeño custom AI chip at the Hot Chips conference on Tuesday, demonstrating significant performance advantages over current state-of-the-art AI accelerator systems
1
. Testing on Semianalysis's InferenceX benchmark revealed that the Jalapeño chip delivered 1.5 to 1.9 times more AI work per watt across GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T models compared to Nvidia Blackwell systems2
. Richard Ho, OpenAI's head of hardware, emphasized that the results show a very significant performance advance, with Jalapeño capable of serving more AI inference workloads per unit of power while returning responses more quickly1
.The custom AI chip achieved 1.7 to 3.6 times lower end-to-end latency across the three tested models, meaning users will experience faster responses, more responsive agents, and more reliable access as demand grows
2
. For ultra-low-latency inference workloads, OpenAI claims its chips deliver 2.1 to 4.1 times faster performance4
. Ho stated that Jalapeño offers the best of both worlds with lower latency and higher throughput, as AI systems typically have to make a trade-off between the two2
.First announced in June, the Jalapeño chip was developed through close Broadcom collaboration, with OpenAI's own models assisting in the development process
1
. The Application-Specific Integrated Circuit (ASIC) is designed specifically for AI inference—the process of running a trained AI model to complete a task or deploy an agent2
. OpenAI plans to deploy Jalapeño in small volumes by the end of 2026, with more significant deployment ramping up throughout 20271
2
. The company is already working on second and third generations of the chip, planning to make Jalapeño a multigenerational platform that allows AI products, models, chips and memory to be developed in concert1
.The Jalapeño AI accelerator features a massive design with 216 GB of HBM4 memory and up to 15.4 TB/s of bandwidth, delivering up to 3.4 MXFP8 PFLOPS and up to 13.4 MXFP4 PFLOPS at 700W
3
. The chip operates at 1.70 GHz on silicon already running in OpenAI's labs, with plans to increase clocks to 1.80 GHz3
. At rack-scale architecture level, each system with 128 accelerators packs 1.7 exaFLOPS of 4-bit compute, 27.5 TB of HBM4 memory, and just shy of 2 petabytes per second of memory bandwidth4
.OpenAI designed Jalapeño to minimize data movement and communication delays using a memory-sliced, NUMA-style architecture
3
. The chip has 64 core slices, each paired with its own HBM slice to guarantee predictable latency and bandwidth while avoiding conflicts associated with unified memory subsystems3
. This arrangement lets frequently used operands remain close to compute resources that need them. Model state, including the KV cache used while generating a response, can be explicitly placed and kept local while the system activates the right combination of compute, memory, and networking for each inference phase1
4
.Related Stories
Analysts view the OpenAI Jalapeño chip as a threat to Nvidia margins in the fast-growing inference market. Adrien Sanchez, technology analyst at Yole Group, noted that a hyperscaler-designed chip can now match or beat Nvidia's Blackwell-class GPUs on inference efficiency, representing a threat to Nvidia's inference margins, which is the field growing the most at the moment
5
. The custom silicon could reduce OpenAI's reliance on Nvidia over time for inference workloads, though Nvidia GPUs will remain important for compute-intensive workloads like large-scale model training given their broad programmability, performance, and software ecosystem5
.SemiAnalysis noted that while Jalapeño beat Blackwell on performance per watt in nearly all tested scenarios, the comparison is somewhat incomplete because Jalapeño uses newer HBM4 memory technology
5
. Nvidia's Rubin platform, which also uses HBM4 memory, represents a more like-for-like comparison, with Vera Rubin systems already shipping to customers while OpenAI still has time before anything beyond engineering samples of Jalapeño becomes available5
. Despite this, Ho confirmed that OpenAI doesn't expect to replace its entire chip lineup with Jalapeño, saying its overall compute strategy includes very good partners like Nvidia2
. OpenAI still needs compute for training, and the highly programmable nature of GPUs means deployment on AMD and Nvidia first, then transition to in-house custom silicon later4
. Omdia expects custom ASIC chips like Jalapeño to exceed GPUs in volume by 2028, representing the biggest competitive threat to Nvidia as about half the capital expenditure on AI infrastructure comes from hyperscale cloud providers who either have a custom chip program or could reasonably have one5
.Summarized by
Navi
[3]
24 Jun 2026•Technology

03 Jan 2026•Technology

24 Aug 2026•Technology

1
Science and Research

2
Technology

3
Policy and Regulation
