14 Sources
[1]
OpenAI sidesteps Nvidia with unusually fast coding model on plate-sized chips
On Thursday, OpenAI released its first production AI model to run on non-Nvidia hardware, deploying the new GPT-5.3-Codex-Spark coding model on chips from Cerebras. The model delivers code at more than 1,000 tokens (chunks of data) per second, which is reported to be roughly 15 times faster than
[2]
A new version of OpenAI's Codex is powered by a new dedicated chip | TechCrunch
On Thursday, OpenAI announced the release of a light-weight version of its agentic coding tool Codex, the latest model of which OpenAI launched earlier this month. GPT-5.3-Codex-Spark is described by the company as a "smaller version" of that model, one that is designed for faster inference. To
[3]
OpenAI's new Spark model codes 15x faster than GPT-5.3-Codex - but there's a catch
Runs on Cerebras WSE-3 chips for a latency-first Codex serving tier. The Codex team at OpenAI is on fire. Less than two weeks after releasing a dedicated agent-based Codex app for Macs, and only a week after releasing the faster and more steerable GPT-5.3-Codex language model, OpenAI is counting
[4]
OpenAI Debuts First Model Using Chips From Nvidia Rival Cerebras
The new Codex model marks OpenAI's latest attempt to best AI rivals such as Alphabet Inc.'s Google and Anthropic PBC in the market for AI coding assistants. OpenAI is releasing its first artificial intelligence model that runs on chips from semiconductor startup Cerebras Systems Inc., part of a
[5]
OpenAI unveils first model running on Cerebras silicon
GPT-5.3-Codex-Spark may be a mouthfull, but it's certainly fast at 1,000 Tok/s running on Nvidia rival's CS3 accelerators Nvidia and AMD can take a seat. On Thursday, OpenAI unveiled GPT-5.3-Codex-Spark, its first model that will run on Cerebras Systems' dinner-place-sized AI accelerators, which
[6]
OpenAI deploys Cerebras chips for 15x faster code generation in first major move beyond Nvidia
OpenAI on Thursday launched GPT-5.3-Codex-Spark, a stripped-down coding model engineered for near-instantaneous response times, marking the company's first significant inference partnership outside its traditional Nvidia-dominated infrastructure. The model runs on hardware from Cerebras Systems, a
[7]
OpenAI's rapid GPT-5.3-Codex model moves beyond simple coding tasks - SiliconANGLE
OpenAI's rapid GPT-5.3-Codex model moves beyond simple coding tasks OpenAI Group PBC today released a lightweight version of its popular agentic coding tool GPT-5.3-Codex, which is designed for more rapid inference, the process by which artificial intelligence models take actions in response to
[8]
OpenAI launches GPT-5.3-Codex-Spark for ultra-fast real-time coding
On Thursday, OpenAI announced GPT-5.3-Codex-Spark, a lightweight version of its agentic coding tool Codex, which the company launched earlier this month. Powered by Cerebras' Wafer Scale Engine 3 chip, Spark enables faster inference as the first milestone in OpenAI's multi-year partnership with
[9]
OpenAI's Fast New Model Aims to Push Vibe Coding Toward Warp Speed
On February 12, in a blog post on its website, OpenAI said that the new model is called GPT‑5.3‑Codex‑Spark, and it's been optimized for one thing in particular: speed. On X, OpenAI cofounder and CEO Sam Altman said that the model is incredibly fast, and "sparks joy for me." Unlike OpenAI's other
[10]
OpenAI Says This Is Its First AI Model That Can Code in Real-Time
* The AI model is available in Codex, CLI, and IDE extension * OpenAI says Codex-Spark can process 1,000 tokens per second * Codex-Spark has a context window of 128K tokens OpenAI's focus on Codex has intensified in 2026. The San Francisco-based artificial intelligence (AI) giant released the
[11]
OpenAI GPT-3.5 Codex Spark Reaches 1,000 Tokens per Second for Coding
GPT-3.5 Codex Spark, as overviewed by Prompt Engineering, is a specialized AI model designed for speed and efficiency in real-time coding and agentic tasks. Capable of processing up to 1,000 tokens per second, it achieves this remarkable performance through custom hardware developed in
[12]
OpenAI's Launches New Version of Codex with a Dedicated Processor Underneath
OpenAI has released a light-weight version of Codex - its agentic coding tool - which it has described as a smaller version of what it had launched earlier this month. The latest version, the company claims, provides faster inference and it does this through a dedicated processor that it acquired
[13]
GPT-5.3 Codex Spark: Is 15x speed worth the reasoning trade-off?
In a move that caught the developer community off-guard on February 12, 2026, OpenAI launched GPT-5.3 Codex Spark. This isn't just another incremental update; it's a radical departure from the "bigger is better" philosophy. Powered by a high-stakes partnership with Cerebras, Spark is a lightweight,
[14]
OpenAI unveils ultra-fast GPT 5.3 Codex Spark model for real-time coding
GPT-5.3-Codex-Spark is currently rolling out as a research preview. OpenAI has introduced GPT-5.3-Codex-Spark, a new AI model built specifically for real-time coding. The new model is currently rolling out as a research preview and is a smaller version of GPT-5.3-Codex. It can generate more than
Share
Copy Link
OpenAI released GPT-5.3-Codex-Spark, its first production AI model running on non-Nvidia hardware from Cerebras Systems. The lightweight coding model delivers over 1,000 tokens per second—roughly 15 times faster than its predecessor—using Cerebras' wafer-scale chips packed with 4 trillion transistors. The release marks the first milestone in OpenAI's $10 billion partnership with the AI chipmaker.
OpenAI released GPT-5.3-Codex-Spark on Thursday, marking a significant shift in the company's infrastructure strategy. The new AI coding model represents OpenAI's first production deployment on non-Nvidia hardware, running instead on chips from Cerebras Systems
1
. The lightweight model delivers code at more than 1,000 tokens per second, which OpenAI reports is roughly 15 times faster than its predecessor3
. To put this speed in context, OpenAI's fastest models on Nvidia hardware deliver significantly lower throughput: GPT-4o reaches roughly 147 tokens per second, o3-mini hits about 167, and GPT-4o mini clocks around 521
.
Source: ZDNet
Codex-Spark is available as a research preview to ChatGPT Pro subscribers at $200 per month through the Codex app, command-line interface, and VS Code extension
1
. OpenAI is also rolling out API access to select design partners. The model ships with a 128,000-token context window and handles text only at launch.The release represents the first milestone in OpenAI's hardware partnership with Cerebras, announced last month as a multi-year agreement worth over $10 billion
2
. Codex-Spark runs on Cerebras' Wafer Scale Engine 3, the company's third-generation waferscale megachip decked out with 4 trillion transistors2
. The dinner-plate-sized AI accelerators feature some of the world's fastest on-chip memory, using SRAM that is roughly 1,000 times faster than the HBM4 found on Nvidia's upcoming Rubin GPUs5
.
Source: The Register
"Cerebras has been a great engineering partner, and we're excited about adding fast inference as a new platform capability," said Sachin Katti, head of compute at OpenAI
1
. Sean Lie, CTO and Co-Founder of Cerebras, emphasized the potential for discovering "new interaction patterns, new use cases, and a fundamentally different model experience" through fast inference2
.OpenAI built Codex-Spark specifically for real-time coding rather than the heavyweight agentic tasks handled by the full GPT-5.3-Codex model launched earlier this month
1
. The company tuned the model for speed over depth of knowledge, creating what it describes as a "daily productivity driver" for rapid prototyping2
. OpenAI reduced overhead per client-server roundtrip by 80 percent, per-token overhead by 30 percent, and time-to-first-token by 50 percent through session initialization and streaming optimizations3
.
Source: SiliconANGLE
The model supports interruption and redirection mid-task, enabling tight iteration loops for developers who need to adjust instructions quickly
3
. It defaults to lightweight, targeted edits and doesn't automatically run tests unless requested. OpenAI says Cerebras' chips excel at assisting "workflows that demand extremely low latency"2
.Related Stories
The Cerebras deployment is part of OpenAI's broader push to diversify hardware suppliers and meet growing computing needs . OpenAI struck a blockbuster agreement in October with AMD to deploy 6 gigawatts' worth of GPU over multiple years, and agreed to buy custom chips and networking components from Broadcom . An OpenAI spokesperson emphasized that the company's partnership with Nvidia remains "foundational" and that OpenAI is "anchoring on Nvidia as the core of our training and inference stack, while deliberately expanding the ecosystem around it" .
OpenAI noted that "GPUs remain foundational across our training and inference pipelines and deliver the most cost effective tokens for broad usage" while Cerebras complements that foundation for low-latency workflows
5
. The company suggests that as Cerebras brings more compute online, it will bring larger models to the platform for users willing to pay a premium for high-speed inference5
.On SWE-Bench Pro and Terminal-Bench 2.0, two benchmarks for evaluating software engineering ability, Codex-Spark reportedly outperforms the older GPT-5.1-Codex-mini while completing tasks in a fraction of the time
1
. The model delivers greater accuracy than GPT-5.1-Codex-Mini in Terminal-Bench 2.0 while being much faster than the smarter GPT-5.3-Codex model5
. The new model marks OpenAI's latest attempt to compete with AI rivals such as Alphabet's Google and Anthropic, which are vying for dominance in the rapidly growing market for AI coding assistants . Codex has more than 1 million weekly active users, OpenAI said .Cerebras raised $1 billion in fresh capital at a valuation of $23 billion last week and has announced intentions to pursue an IPO
2
. Sam Altman hinted at the launch in a tweet, saying "We have a special thing launching to Codex users on the Pro plan later today. It sparks joy for me"2
.Summarized by
Navi
[5]
02 Feb 2026•Technology

24 Apr 2026•Technology
19 Nov 2025•Technology

1
Technology

2
Technology

3
Policy and Regulation
