Apple explores PrismML's compression tech to run AI models on iPhone without the cloud

Reviewed byNidhi Govil

10 Sources

Share

Apple is evaluating technology from PrismML that could shrink powerful AI models from 54GB to under 4GB, enabling them to run entirely on an iPhone. The Caltech spinout's compression approach could reduce Apple's cloud computing costs, enhance privacy, and allow Siri to process more requests locally without internet connectivity.

Apple Evaluates PrismML's Breakthrough in AI Model Compression

Apple is in early-stage discussions with PrismML, a Caltech spinout backed by Khosla Ventures, about technology that could fundamentally change how AI models for iPhone operate. PrismML CEO Babak Hassibi confirmed to CNBC that Apple and other companies are evaluating the startup's models, measuring their speed, energy efficiency, and performance on devices

1

. "They're really evaluating our technology right now," Hassibi said of Apple, characterizing the discussions as very early but noting that "things are progressing nicely"

1

. The timing aligns with Apple's ongoing efforts to make Siri more competitive while addressing one of the central constraints facing its AI strategy: most capable models require too much memory and processing power to run AI models on iPhone hardware alone.

Source: AppleInsider

Source: AppleInsider

PrismML Releases Bonsai 27B with Dramatic Size Reduction

On Tuesday, PrismML publicly released Bonsai 27B, compressed versions of the Alibaba Qwen model that demonstrate the startup's core innovation

1

. The company reduced the model from roughly 54GB to less than 4GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer

1

. This represents a significant advance in on-device AI capabilities. Unlike Apple's AFM 3 Core Advanced model with 20 billion parameters that uses a sparse architecture where only 1 billion to 4 billion parameters are active at a time, all 27 billion parameters in PrismML's compressed AI models remain active simultaneously[4](https://www.macrumor ar-147521 s.com/2026/07/09/apple-prismml-larger-on-device-ai-models/). The release came one day after Apple opened the public beta of iOS 27, giving iPhone owners their first broad access to the company's long-delayed Siri overhaul

1

.

How AI Model Compression Works and Its Trade-offs

PrismML achieves dramatic size reduction by drastically simplifying how internal information is stored, reducing each value from 16 bits to just one or three possible values

1

. Hassibi compared it to the chip industry's move from eight-bit to four-bit computing, but said PrismML takes it a step further

1

. The startup reports that compressed AI models use between 10 and 15 times less memory usage, generate responses six to eight times faster, and consume three to six times less energy efficiency than conventional versions running on existing hardware

1

. However, there is a trade-off. PrismML's models typically lose a few percentage points of overall performance, with factual recall weakening before skills such as reasoning, math, and coding

1

. PrismML says its builds keep about 95% of full performance in the ternary version and 90% in the 1-bit one

2

.

Strategic Implications for Apple Intelligence Features and Privacy

The technology could address critical challenges in Apple's AI strategy by enabling more on-device processing rather than relying on Private Cloud Compute servers. Carolina Milanesi, president and principal analyst at Creative Strategies, said smaller models could let Apple move more demanding features onto the iPhone, including computational photography, video generation, and health or fitness tools that rely on sensitive personal data

1

. "The more you can do on device, the better it is," she said, pointing to health and medication data that users would want to keep private

1

. Running more AI directly on the iPhone would reduce latency associated with sending data to remote servers, lower cloud-computing costs, and support Apple's privacy pitch

1

. It would also allow certain Apple Intelligence features to work without an internet connection, addressing privacy concerns that have become central to Apple's positioning against competitors like OpenAI and Anthropic.

Source: AppleInsider

Source: AppleInsider

Technical Implementation and Device Compatibility

Bonsai 27B runs natively on Apple devices including Mac, iPhone, and iPad via Apple's MLX framework, and on NVIDIA GPUs via CUDA through custom low-bit kernels built for its hybrid-attention architecture

3

. PrismML ships two versions under a free Apache 2.0 License: a ternary build that runs on a laptop, and a smaller 1-bit build at about 3.9GB designed to fit within the memory budget of an iPhone 17 Pro

2

. Fitting a phone is a stricter gate than storage numbers suggest, since a phone never exposes its full memory to an app—a 12GB iPhone offers about 6GB for the model to use for on-device processing, and the model shares that budget with its KV cache and activations

3

. At about 4GB, 1-bit Bonsai 27B is the first to pass through with room to work

3

.

Analyst Perspectives on Edge AI and Future Roadmap

Analysts urged caution about real-world performance. Tarun Pathak of Counterpoint Research said the real test would be millions of queries across thousands of devices, while Phil Solis of IDC said power use was the biggest open question, since a model that runs often could still drain a battery

2

. The release also feeds a debate over whether efficiency gains will cut demand for memory and data-center chips. Gil Luria, an analyst at D.A. Davidson, said shrinking models would not remove the need for processors but would simply move some of them from data centers onto phones, part of a broader shift toward edge AI

2

. Hassibi said Google's open-source Gemma model is next in the pipeline, followed by much larger models, including those from frontier labs that today generally require datacenter hardware

1

. The technology could ultimately extend well beyond phones and laptops to robotics, autonomous systems, and other products that need to make decisions quickly without relying on a cloud connection

1

. The California Institute of Technology owns the underlying patents and licenses them exclusively to PrismML, which raised a $16.25 million seed round in March backed by Khosla Ventures and other investors

1

.

[5]

AppleInsider

|

AppleInsider.com

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved