2 Sources
[1]
Your next CPU could be the biggest AI upgrade in years -- here's why
AI is fast approaching a fork in the road. The spiraling cost of memory for both personal and cloud platforms, rising compute costs for training next-gen models, and high user demand for cloud platforms have made AI quite expensive to both provide and use. Then there's the growing privacy concerns
[2]
Arm, Intel and AMD update upcoming CPUs: a big boost for on-device assistants
Windows 11 PCs and phones should gain faster AI and battery life Arm, Intel, and AMD are reworking their next chip designs so more generative AI can run directly on phones and laptops. You'll start seeing that in the next wave of Windows 11 PCs and next-generation smartphones. The main shift is
Share
Copy Link
Arm, Intel, and AMD are redesigning their next-generation CPUs to handle AI workloads directly, moving computation away from cloud servers and power-hungry GPUs. The shift promises faster on-device assistants, better privacy, and significantly improved battery life for phones and laptops running Windows 11 and beyond.
Arm, Intel, and AMD are fundamentally rethinking how CPUs handle artificial intelligence, introducing native matrix computation capabilities that could reshape how we interact with AI on phones and laptops
1
2
. Instead of routing every request to cloud platforms or relying exclusively on dedicated accelerators, the next wave of processors will execute generative AI tasks on phones and laptops directly on the CPU alongside NPUs. This CPU AI approach addresses spiraling memory costs, privacy concerns about sensitive data sent to chatbots, and the growing demand for running AI models locally1
.
Source: Android Authority
The technical foundation for this AI upgrade began with Arm's introduction of SVE2 in 2021 as part of Armv9, which brought flexible vector processing capabilities. But the real transformation arrived with SME2, which extends the CPU with dedicated matrix execution mode and hardware support for GEMM-style operations specifically designed for transformer and LLM inference workloads
1
. SME2 delivers 3x to 5x better performance over older CPUs without acceleration support while consuming minimal power and chip space, according to Arm. The technology already ships in MediaTek's Dimensity 9500 processor and will appear in more chipsets built from Arm's C1 Ultra cores1
.In a significant development, AMD and Intel announced a joint venture to bring AI Compute Extensions (ACE) to future CPUs, introducing native matrix instructions to the x86 instruction set architecture
1
2
. ACE supports INT4 sub-byte types along with BF16 and FP16 data types, building on existing AVX instructions to deliver consistent CPU-based AI acceleration across vendors. This standardization means developers can target AI workloads without wrestling with proprietary GPU or NPU APIs, dramatically simplifying the path to bringing AI features to consumer products1
.The performance gains are substantial. AMD's Zen 5-based Ryzen AI 300 chips can reach up to 50 TOPS, approximately three times the previous generation
2
. Intel's Lunar Lake and Arrow Lake architectures aim for roughly triple the neural performance of earlier chips, with Lunar Lake expected to exceed 40 TOPS for Copilot+ PCs2
. On mobile platforms, Arm's Cortex-X925 delivers a 41% AI performance improvement, supported by Kleidi libraries for PyTorch and TensorFlow that make CPU acceleration more accessible to developers2
.The shift toward CPU for AI inference carries practical benefits beyond raw speed. Industry estimates suggest many NPU workloads consume 5 to 10W compared with roughly 30 to 40W on a GPU
2
. If these figures hold, users could see approximately 1.5 to 3 extra hours of laptop battery life when running on-device assistants, writing tools, or image processing features2
. Running AI models locally also means messages, documents, and photos stay on your device rather than traveling to cloud servers, addressing growing privacy concerns1
2
.While GPUs will remain dominant for large-scale training and running the largest models due to their massively parallel architecture, CPUs are rapidly becoming far more capable for low-latency inference workloads
1
. The evolution from vector SIMD into architectures that natively support tensor and matrix computation represents a fundamental shift in processor design. Developers gain a consistent target across phones, laptops, and PCs without maintaining multiple code paths for different accelerators1
.Related Stories
These capabilities will arrive in upcoming Windows 11 PCs and next-generation smartphones rather than as downloadable software updates
2
. Existing phones and laptops won't benefit from these architectural improvements, meaning the real impact comes with your next hardware upgrade1
. For users who regularly interact with assistants, writing tools, or image features, the combination of reduced lag, better privacy, and extended battery life could make this one of the most meaningful upgrades in years. However, waiting for independent battery and speed tests before weighing manufacturer claims remains advisable2
.Summarized by
Navi
[1]
1
Science and Research

2
Policy and Regulation

3
Technology