5 Sources
[1]
Hot Chips 2026: Intel dives deep on Crescent Island AI accelerator -- larger caches and deeper XMX engines target maximum AI FLOPS per watt
Xe3P focuses on the most compute-intensive phases of AI inference. Intel shared more details of its Crescent Island AI accelerator, powered by the Xe3P architecture, at the Hot Chips symposium this week. Unlike Nvidia's Rubin and AMD's MI455X GPUs, which are high-power, exclusively liquid-cooled chips with massive pools of HBM4 memory that provide maximum performance across both AI training and inference workloads, Crescent Island is designed to fit into a lower-power, inference-first niche in the AI accelerator market. As a refresher, Crescent Island is a 350W air-cooled PCIe card that uses up to 480 GB of LPDDR5X memory, meaning it can be deployed in traditional servers without exotic power and cooling requirements. We've already learned about some of Crescent Island's DNA from past disclosures, but Intel went deeper into the chip's architectural details at Hot Chips. Crescent Island is built up from four Xe3P slices, each containing eight Xe Cores, for a total of 32. Each Xe Core has eight Xe Vector Engines and eight XMX matrix accelerators, for a total of 256 of each resource. The Xe3 graphics architecture, as seen on Intel's Panther Lake processors, already modified the capacity and flexibility of the GPU cache hierarchy to improve utilization and decrease performance-sapping register spills, and Xe3P further refines that hierarchy. The Xe3P Xe Core has twice the amount of general register file space for working data versus Battlemage. Each Xe Core now has 1MB of general-purpose register file space, up from 512KB on Battlemage and Xe2. In addition, Xe3P offers 512KB of L1 cache or shared local memory per Xe Core, a structure that started out at 256KB on Battlemage and grew by approximately 1.33x on Panther Lake's Xe3 GPU. The chip also has 32MB of shared L2 cache. These expanded caches are meant to serve the chip's larger matrix accelerators on its AI compute-focused mission. Xe3P boasts a larger systolic depth in its XMX engines than past Xe GPU designs. Xe3P's XMX systolic engines are a 16-deep design, meaning they can process matrices in much larger chunks than the four-deep systolic design of Xe2 and Xe3. Nvidia doesn't discuss the architecture of its Tensor Cores in anywhere near this level of detail, but as an AI inference-focused part, the fact that Xe3P can theoretically work on more elements at once during general matrix-multiply operations is an important capability boost for Crescent Island's inference ambitions. Intel is also prioritizing a broad range of data types with this chip, from FP4 formats with microscaling support (aka MXFP4) all the way to what it describes as full-rate double-precision (via 64 FP64 FMA units per Xe Core). FP64 isn't widely used in AI workloads, but Intel says that the inclusion of full-rate processing for that data type makes Crescent Island useful as a converged high-performance computing and AI chip. Each Xe Core also supports sigmoid and tanh transcendental functions, which are important to a variety of operations during AI inference, especially the softmax function. AMD and Nvidia have prioritized the performance of these functions in their recent architectures as well, so the fact that Xe3P offers support for them is key for its AI-first initiatives. While Crescent Island does have a media codec block featuring four encoders and decoders to help serve up video to multimodal AI models, gamers hoping for a glimpse of future Arc cards won't find it with this product, as graphics-specific functionality like RT cores has been omitted from this chip to preserve die area for compute functionality. As a data-center-focused part, Crescent Island offers a full suite of reliability, availability, and serviceability features, including ECC and parity protection across the die and a range of memory reliability features. As for the specific applications that Crescent Island will target, Intel highlights the rise of mixture-of-experts models paired with speculative decoding as a new class of workload that Crescent Island can serve well. Speculative decoding strategies vary, but in general, they use a fast, lightweight mechanism to create drafts of future tokens that the main model can then be used to accept or reject, potentially improving decode performance. Not every draft token generated this way will be approved, but much like speculative execution in CPUs, it helps produce useful work from compute resources that would otherwise be left idle. As model serving recipes pursue more aggressive drafting mechanisms, more compute is required to generate those draft tokens. At a high level, that understanding changes the common perception of decode as being a mostly memory-bandwidth-bound operation. As an LPDDR5X-powered chip, Crescent Island won't have the eye-popping bandwidth of HBM-backed accelerators at its disposal for maximum performance with traditional autoregressive decode, so any help it can get from these speculative methods will be helpful. Overall, Intel claims that Crescent Island is built to offer high FLOPS per watt and that it's optimized for compute-bound workloads like prefill (aka prompt processing and KV cache construction). Intel's emphasis on those areas of AI performance, as well as heterogeneous deployments, suggests that this chip could have a niche alongside HBM-backed accelerators whose resources are best used for decode operations. Intel and its partner SambaNova could both stand to benefit from such an arrangement, as that company's SN50 inference accelerators are explicitly built to benefit from disaggregated prefill processing powered by GPUs. SN50 racks and Crescent Island are both meant to serve as lower-power, air-cooled systems that customers can deploy in existing data centers without dramatic upgrades to power or cooling infrastructure, so there is broad synergy in the shape of those products. Intel still isn't discussing just how many theoretical compute FLOPS to expect from Crescent Island, nor is it disclosing memory bandwidth figures. But the architectural decisions it's shared so far -- getting lots of data close to the compute engines of the chip and processing more of it at once in a relatively narrow power envelope -- seem sound in a world where the company is still trying to reset its AI ambitions after a string of high-profile product failures and cancellations. Intel has promised Crescent Island for a second-half 2026 time frame, and the clock is ticking on that launch window, so we're eager to learn more about the chip's final specifications, as well as customer and partner wins, when that launch does occur.
[2]
Intel confirms Crescent Island GPU will pack 32 Xe3P cores and up to 480GB of memory
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. The big picture: Intel announced its Crescent Island data center GPU last year before sharing more details at Computex 2026 in June. The company has now revealed all the key specifications, including core count, cache, memory, and more. Team Blue also emphasized that Crescent Island is optimized for performance per watt and improved AI model support rather than raw compute performance. At the Hot Chips 2026 symposium in Palo Alto this week, Intel revealed that Crescent Island will feature 32 Xe3P cores, 256 XMX engines, 32MB of unified L2 cache, a PCIe Gen5 x16 interface, and a 350W TDP for the air-cooled model. Intel will ship Crescent Island accelerators with 160GB of LPDDR5X memory, but the chip will support ODM designs with up to 480GB of VRAM. Xe3P offers a marked improvement over both Xe2 and Xe3 in terms of AI capabilities. The XMX units in the new architecture have a 16-deep systolic array, compared with the four-deep design used in their predecessors. The new Xe Vector Engine also gains support for FP8 and FP4 precision, as well as microscaling formats, while doubling the GRF to 1MB per Xe core. Other notable changes include a more capable memory subsystem, packing 512KB of L1 cache/SLM memory per Xe core, as well as full-rate FP64 support with 64 FMAs per Xe core. Intel also noted that the new architecture comes with an industry-standard AI software stack that is open, upstreamed, and Day 0-ready for Crescent Island. Crescent Island is specifically designed for agentic AI rather than gaming, so it does not include any 3D or ray-tracing hardware found in Intel's consumer graphics chips. Instead, Team Blue has used the freed-up space to pack eight Vector Engines and eight XMX Engines into each Xe3P core, for a total of 256 of each in the full GPU. The Xe3P graphics architecture is an AI-focused variant of the standard Xe3 architecture used in Core Ultra 300-series "Panther Lake" chips for laptops and compact desktops. Xe3P is also expected to power Intel's Arc-C series GPUs for client devices. Customer sampling for Crescent Island is expected to begin in the current quarter, with the official launch slated for 2027.
[3]
Intel Crescent Island AI Accelerator Supports Up to 480 GB Memory
Intel has shared further technical details about Crescent Island, a dedicated AI inference accelerator based on its Xe3P compute architecture. Presented at Hot Chips 2026, the card combines 32 Xe3P cores with unusually large LPDDR5X memory configurations intended for local large-model workloads. Although Intel describes Crescent Island using its Xe architecture, the product is not a conventional graphics card. The silicon does not include the graphics and display components needed to render images or drive a monitor. Its resources are instead reserved for vector and matrix calculations used by AI inference, including long-context and agentic workloads. Crescent Island comes in an air-cooled PCIe add-in-card format with a rated power of 350 W. Intel is targeting compatible PCs, workstations, and edge data-center systems. The design should be easier to integrate than accelerators that require dedicated liquid cooling, although host-system requirements and the physical dimensions of the board have not been detailed. The processor contains 32 Xe3P cores, each equipped with eight vector engines and eight XMX engines. This creates chip-wide totals of 256 vector engines and 256 XMX matrix engines. Intel has paired these units with 32 MB of unified L2 cache. Each core also contains a 1 MB register file and 512 KB of L1 cache, providing substantial local storage for active computational data. A redesigned XMX pipeline is one of the main architectural changes in Xe3P. Intel expanded the matrix engine from the four-stage systolic design used with regular Xe3 to a 16-stage structure. The accelerator supports FP8 and FP4 calculations, both of which can improve inference efficiency by representing model weights and operations at lower precision. Actual gains will vary according to the model and the way it has been quantized. Intel-branded Crescent Island cards will include 160 GB of LPDDR5X memory. Manufacturing partners can develop higher-capacity versions containing as much as 480 GB. Intel has not yet stated the memory speed, interface width, or available bandwidth, making it difficult to compare the memory subsystem directly with HBM-equipped data-center accelerators. The maximum capacity could allow a multi-card workstation to hold exceptionally large models locally. Four 480 GB cards would offer a combined 1.92 TB of memory. Intel presented this as enough capacity to approach the requirements of a trillion-parameter model. Efficiently using such a configuration would still depend on model partitioning, PCIe or card-to-card communication performance, and support within the selected AI framework. Intel says support for major AI frameworks will be available when Crescent Island launches. The company has not disclosed clock speeds, theoretical compute performance, PCIe specifications, multi-card interconnect details, pricing, or a firm retail date. These missing figures will be needed to judge how the product compares with existing workstation and data-center inference hardware. SpecificationIntel Crescent Island ArchitectureIntel Xe3P Intended workloadAI inference Xe3P cores32 Vector engines256 XMX engines256 Vector engines per core8 XMX engines per core8 L1 cache512 KB per core Unified L2 cache32 MB Register file1 MB per core Supported formatsFP4 and FP8 Intel card memory160 GB LPDDR5X Maximum partner-card memory480 GB LPDDR5X Board power350 W CoolingAir cooling Form factorPCIe add-in card Display and rendering hardwareNot included Source: intel
[4]
Intel 'Crescent Island' GPUs set to have up to 480 GB LPDDR5X memory with 32 Xe3P cores
Intel used Hot Chips 2026 to lay out full specs for Crescent Island, its inference-focused GPU built on the new Xe3P architecture. The chip has 32 Xe3P cores split across four compute slices, working out to 256 Vector Engines and 256 XMX matrix engines. It's paired with 32MB of unified L2 cache and up to 480GB of LPDDR5X memory, though the reference Intel card ships with 160GB. Partners building their own ODM versions can push that number higher. Crescent Island skips HBM entirely, which is the interesting part. NVIDIA's Rubin and AMD's MI450X both lean on HBM4 for their next AI accelerators, but that memory has gotten harder to source and pricier as demand climbs. Intel is betting LPDDR5X gets it close enough on capacity (480GB beats MI450X's 432GB and Rubin's 288GB on paper) while running at a fraction of the power. The card is rated for 350W and is air-cooled, so it fits into standard PCIe Gen5 x16 servers without needing liquid cooling. The Xe3P core looks different from Xe2 under the hood. XMX units moved from a 4-deep systolic array to 16-deep, the general register file doubled to 1MB per core, and L1/SLM cache is now 512KB per core. Intel also stripped out the 3D and ray-tracing hardware you'd find on a gaming GPU, freeing up die space for AI compute instead. Support spans FP4 up through FP64. On the software side, Intel is leaning on an open stack: vLLM, SGLang, llm-d, and even NVIDIA's own Dynamo framework are supported, alongside KV-cache-aware routing for long-context agentic workloads. It should be clear that this isn't a preview of a future Arc gaming card. Xe3P was originally slated for the canceled Celestial GPUs before Intel folded the architecture into Crescent Island, so gamers shouldn't expect much from today's reveal, even if Xe3P eventually reaches the Arc C-series. What it does show is that Intel now has a real answer to NVIDIA and AMD in the inference accelerator market, at a lower price and power point than either. Customer sampling is expected in the second half of 2026, with a full launch in 2027.
[5]
Intel Crescent Island GPUs Pack Up To 32 Xe3P Cores, Optimized For Agentic AI With Low-Cost LPDDR5X That Reaches Up To 480 GB Capacity
Intel has given the rundown on its next-gen AI inference accelerator, Crescent Island, which packs 32 Xe3 cores and 480 GB of LPDDR5X memory. Intel To Enter Agentic AI Accelerator Space With Xe3P-Powered Crescent Island GPUs, Offering Low-Cost & Low Power Through LP5X Memory & 350W TDPs The Intel Crescent Island GPU is based on the brand-new Xe3P architecture, which is the same graphics architecture that was teased by the company during its Panther Lake and Xe3 deep dives. The new silicon features of Crescent Island include: * Agentic AI Ready: High memory capacity and sustained compute * LPDDR5x: Higher density, lower power. Enhanced reliability with no BW/capacity loss * AI Optimized: Featuring larger context lengths, higher memory, and performance/TCO * Prefill Optimized: High FLOP/W, compute-bound workloads Software innovations for Crescent Island include: * Broad model support: Supports language, reasoning, multi-modal, and diffusion models * KV Cache Efficiency: Enables KV-cache-aware routing and offloading for scalable inference * Open framework ecosystem: Built to work with vLLM, SGLANg, llm-d, and NVIDIA Dynamo * Heterogeneous agent orchestration: Runs across heterogeneous infrastructure without changing agent code The new GPU architecture will be a further upgrade over the Xe3 architecture, and for clients, the architecture will be featured on the next-gen Arc family, the Arc C-Series. But Xe3P is going to be even more scalable, from client iGPUs to data center AI GPUs. * Xe3P microarchitecture with optimized performance-per-watt * 480GB of LPDDR5X memory * Support for a broad range of data types, ideal for "tokens-as-a-service" providers and * inference use cases Intel Crescent Island will be both power- and cost-optimized. It will be targeted at air-cooled data center solutions and workstations and will be aimed at AI inference workloads. According to Intel, the Xe3P graphics architecture used for Crescent Island will be optimized for performance per watt. Competitors such as NVIDIA and AMD are offering their data center AI solutions with top-grade HBM memory, such as HBM3E, and are already shipping the first HBM4 solutions for next-gen parts such as Rubin and MI450X. But at the same time, sourcing HBM has become difficult due to increased demand, and that has also led to higher prices. Just for comparison: * Intel Crescent Island - Up To 480 GB (LPDDR5X) * AMD Instinct MI450X - Up To 432 GB (HBM4) * NVIDIA Vera Rubin - Up To 288 GB (HBM4) Xe3P IP Is Optimized For AI The Intel Crescent Island GPUs will pack the 3rd Generation Xe architecture codenamed Xe3P. Intel says that this architecture leverages a multigenerational Xe install base with improvements in efficiency and software compatibility. The Crescent Island GPUs leveraging Xe3P are optimized for AI. On the architectural levels, the 3rd Gen Xe Core "Xe3P" packs 8 Vector Engines and 8 XMX engines per Xe core. The Xe Vector Engine now includes FP8+FP4 Precision, 3-way co-issue, an extended math & FP64 (full rate 64 FMA/XeCore) capability, while the Xe Matrix Extensions (XMX) units offer 16-deep systolic array. There's also 1 MB of GRD per Xe Core. The Crescent Island GPU is made up of four Compute Slices. Each slice packs 8 cores, so that gives us a total of 32 Xe3P cores, 256 Vector Engines, and 256 XMX units. Each Xe core retains load/store units, a dedicated I$, and L1$/SLM cache. All Compute slices are connected to a large L2 cache. For the memory subsystem, Intel's Xe3P "Crescent Island" GPU packs 512 KB of L1$/SLM per Xe Core or 16 MB in total, 32 MB of Unified L2 cache, and up to 480 GB of LPDDR5X memory. Intel-branded solutions will feature 160 GB of LPDDR5X memory. LPDDR5X Saves Power & Costs For Crescent Island GPUs Leveraging LPDDR5X memory can give Intel a big edge in the cost/performance segment. Furthermore, the architecture will support a broad range of data types that are ideal for "Tokens-as-a-service" providers and inference use cases. This is why Intel is going big with Crescent Island. Starting with the specifications, the Intel Crescent Island Data Center GPU will feature up to 480 GB of LPDDR5X capacity. The reference Intel PCIe card will feature 160 GB of LP5X memory, but partners will be given the freedom to build ODM-branded cards with flexible options up to 480 GB. The use of LP5X memory also cuts down power significantly, with just 350W of TDP on the air-cooled PCIe version. Intel states that LP5X memory enables a densely packed channel design, which enables a significant bandwidth increase. We have already seen initial PCIe boards with support for up to 160 GB of memory, so the partner board should be even more interesting. Another thing about Crescent Island is that traditional GPU blocks, such as graphics or 3D support. The chip solely focuses on GPGPU, and that freed more die area for additional AI compute. Plus, the chip is optimized for Perf/Watt and Perf/TCO. The GPU will be backed by a wide-range data format support from FP4 up to FP64, and will also support a fully Open and Robust software stack. Intel is already evaluating its open and unified software stack for heterogeneous AI systems with its existing Arc Pro B-series lineup, so future iterations will be able to access these optimizations early on. Intel is currently targeting customer sampling for its Crescent Island GPU for the 2H of 2026, so we'll definitely learn more about the GPU in the coming months. Follow Wccftech on Google to get more of our news coverage in your feeds.
Share
Copy Link
Intel detailed its Crescent Island AI accelerator at Hot Chips 2026, featuring 32 Xe3P cores and up to 480GB LPDDR5X memory. The 350W air-cooled PCIe card targets inference-first workloads and agentic AI, bypassing expensive HBM memory to offer a cost-efficient alternative to NVIDIA and AMD's liquid-cooled solutions.
Intel unveiled comprehensive specifications for its Crescent Island AI accelerator at Hot Chips 2026, positioning the chip as a cost-efficient alternative in the data center GPU market
1
2
. Unlike NVIDIA's Rubin with 288GB HBM4 or AMD's MI450X with 432GB HBM4, Intel Crescent Island uses LPDDR5X memory to achieve up to 480GB capacity while maintaining a 350W TDP in an air-cooled PCIe form factor4
. This design choice allows deployment in traditional servers without exotic power and cooling requirements, addressing the growing challenge of HBM sourcing and pricing5
.
Source: Wccftech
The Xe3P architecture powering Crescent Island represents a focused evolution for inference-first workloads. The chip contains 32 Xe3P cores organized across four compute slices, delivering 256 Vector Engines and 256 XMX matrix engines
2
. Each Xe Core houses eight Vector Engines and eight XMX matrix accelerators, with Intel stripping out graphics and ray-tracing hardware to maximize die area for AI compute functionality1
. The XMX engines feature a 16-deep systolic array, a substantial upgrade from the four-deep design in Xe2 and Xe3, enabling processing of matrices in larger chunks during general matrix-multiply operations2
.
Source: Tom's Hardware
Intel significantly expanded the cache hierarchy to improve utilization and reduce performance-sapping register spills. Each Xe Core now contains 1MB of general-purpose register file space, double the 512KB found in Battlemage and Xe2
1
. The L1 cache or shared local memory per Xe Core reaches 512KB, up from 256KB on Battlemage, totaling 16MB across all cores3
. The chip also features 32MB of unified L2 cache2
. These expanded caches serve the larger matrix accelerators on the chip's AI compute-focused mission, targeting maximum AI FLOPS per watt rather than raw performance1
.Crescent Island supports a comprehensive range of data types from FP4 formats with microscaling support (MXFP4) to full-rate double-precision via 64 FP64 FMA units per Xe Core
1
. The inclusion of FP8 and FP4 precision can improve inference efficiency by representing model weights at lower precision3
. Each Xe Core supports sigmoid and tanh transcendental functions, critical for softmax operations during AI inference, matching capabilities prioritized by AMD and NVIDIA in recent architectures1
. While FP64 isn't widely used in AI workloads, Intel positions this support as making Crescent Island useful as a converged high-performance computing and AI chip1
.Intel-branded Crescent Island cards will ship with 160GB of LPDDR5X memory, but the architecture supports ODM designs with up to 480GB capacity
2
3
. This maximum capacity exceeds AMD's MI450X at 432GB HBM4 and NVIDIA's Rubin at 288GB HBM44
. A four-card workstation configuration could deliver a combined 1.92TB of memory, approaching requirements for trillion-parameter models3
. The LPDDR5X choice enables densely packed channel design for significant bandwidth increases while consuming a fraction of the power required by HBM solutions5
. The chip connects via PCIe Gen5 x16 interface2
.Related Stories
Intel designed Crescent Island specifically for agentic AI workloads and mixture-of-experts models paired with speculative decoding
1
5
. Speculative decoding uses fast, lightweight mechanisms to create draft tokens that the main model accepts or rejects, improving decode performance by producing useful work from otherwise idle compute resources1
. As model serving recipes pursue more aggressive drafting mechanisms, additional compute becomes necessary to generate draft tokens, shifting decode from a memory-bandwidth-bound operation toward compute intensity1
. The chip supports KV cache-aware routing for long-context agentic workloads5
.Intel emphasizes an open AI software stack that is upstreamed and Day 0-ready for Crescent Island
4
. The chip supports major AI frameworks including vLLM, SGLang, llm-d, and NVIDIA's Dynamo framework4
5
. The architecture enables heterogeneous agent orchestration across infrastructure without changing agent code5
. Intel positions the solution as ideal for tokens-as-a-service providers and inference use cases5
. The chip includes four media codec encoders and decoders to serve video to multimodal AI models1
.
Source: Guru3D
Crescent Island targets a lower-power, cost-optimized niche between high-end liquid-cooled accelerators and entry-level solutions
1
. The air-cooled design fits standard PCIe servers without requiring exotic cooling infrastructure, addressing deployment challenges faced by HBM4-based competitors3
. Customer sampling is expected to begin in the second half of 2026, with full launch planned for 20272
4
. The Xe3P architecture was originally planned for the canceled Celestial gaming GPUs before being redirected to Crescent Island, though it may eventually appear in Arc C-series client GPUs4
. Intel has not disclosed clock speeds, theoretical compute performance, pricing, or multi-card interconnect details3
.Source: TechSpot
Summarized by
Navi
[1]
[4]
20 May 2026•Technology

14 Oct 2025•Technology
01 Jun 2026•Technology

1
Technology

2
Policy and Regulation

3
Technology
