10 Sources
[1]
Apple in talks with startup that shrinks AI models to run on an iPhone
If PrismML's claims hold up in real-world testing, the technology could reshape demand for memory and datacenter compute -- though analysts say AI will still require plenty of chips. Apple is in talks with a small Silicon Valley company that says it can shrink powerful artificial intelligence models enough to run directly on an iPhone, the startup's CEO told CNBC. PrismML, a Khosla Ventures-backed spinout from the California Institute of Technology, publicly released compressed versions of Alibaba's open-source Qwen model on Tuesday. The company said it reduced the model from roughly 54 GB to less than 4 GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer. PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been evaluating the startup's models and measuring their speed, energy efficiency and performance on devices. "They're really evaluating our technology right now," Hassibi said of Apple. He characterized the discussions as very early and said it remains unclear where they will lead, but that "things are progressing nicely." The release comes one day after Apple opened the public beta of iOS 27, giving iPhone owners their first broad access to the company's long-delayed overhaul of Siri. Apple is trying to make Siri more competitive with assistants from OpenAI and Anthropic while keeping more personal information and AI processing on the device. The company's approach could address one of the central constraints facing Apple's AI strategy. The most capable models typically require too much memory and processing power to run on a smartphone. Apple can send complex requests to cloud-based models, but running more AI directly on the iPhone would reduce the delay associated with sending data to a remote server, lower cloud-computing costs and support the company's privacy pitch. It would also allow certain features to work without an internet connection. Carolina Milanesi, president and principal analyst at Creative Strategies, said smaller models could let Apple move more demanding features onto the iPhone, including computational photography, video generation and health or fitness tools that rely on sensitive personal data. "The more you can do on device, the better it is," she said, pointing to health and medication data that users would want to keep private. PrismML said it shrinks AI models by drastically simplifying how their internal information is stored -- reducing each value from 16 bits to just one or three possible values. That significantly cuts the memory required to store and operate the model. Hassibi compared it to the chip industry's move from eight-bit to four-bit computing, but takes it a step further. The startup said the compressed models use between 10 and 15 times less memory, generate responses six to eight times faster and consume three to six times less energy than conventional versions running on existing hardware. Hassibi did acknowledged there is a trade-off, however. PrismML's models typically lose a few percentage points of overall performance, with factual recall weakening before skills such as reasoning, math and coding, he said. PrismML is releasing two compressed versions of the model for free. They are designed to run on everyday devices, including iPhones, MacBooks and Nvidia-powered PCs. The technology emerged from Hassibi's research group at Caltech. The university owns the underlying patents and licenses them exclusively to PrismML. In March, the company raised a $16.25 million seed round backed by Khosla Ventures and other investors. Hassibi said Google's open-source Gemma model is next in the pipeline, followed by much larger models, including those from frontier labs that today generally require datacenter hardware. The technology, according to PrismML, could ultimately extend well beyond phones and laptops to robotics, autonomous systems and other products that need to make decisions quickly without relying on a cloud connection. "It's very important that the intelligence be local and that it can run fast," he said. Apple already runs parts of its AI system locally, including translation, some summarization and features tied closely to personal information. More complex requests are routed to Apple's private cloud infrastructure or outside models. Horace Dediu, founder of Asymco, said Apple is likely trying to keep the large majority of common Siri interactions on-device while reserving the most demanding tasks for the cloud. The advantage is not simply using less memory, he said, but fitting a more capable model within the same physical limits. "They're trying to figure out how big a model and how clever a model they can fit on the device," Dediu said. Keeping common requests local gives Apple lower latency, greater privacy and potentially lower licensing and cloud costs. Apple may have an advantage in putting these models to work because it designs the iPhone's chips and software together, giving it tighter control over how AI runs on the device. But analysts cautioned that PrismML's claims still need to be proven outside controlled demonstrations. Tarun Pathak, research director at Counterpoint Research, said the model's performance on lengthy prompts, battery consumption during multitasking and reliability across millions of requests will be critical. "The ultimate test will be millions of queries, thousands of device combinations and robust testing at scale," Pathak said. Phil Solis, who leads IDC's research on client processors, said power consumption may be the biggest open question. A model that is capable enough to be used frequently -- or continuously in the background for agent-like tasks -- could drain a phone's battery even if it requires less memory. PrismML's release also comes during an intense debate over whether improvements in AI efficiency could eventually reduce demand for memory chips and expensive datacenter infrastructure. Memory has become one of the biggest constraints and costs across consumer electronics and AI servers. Morgan Stanley estimates Apple's average dynamic random access memory cost per bit could rise roughly 190% year over year in fiscal 2027, with NAND costs up about 180%. NAND is typically used in flash drives and solid state drives. The firm expects Apple to raise the starting price of comparable iPhone 18 models by about $200 to protect margins. PrismML said its approach could allow a cloud model that normally requires eight GPUs to run on one, while also allowing models that once required a server to move onto phones and laptops. That could reduce the amount of memory or computing capacity needed for a given AI task. But it does not necessarily mean overall chip demand will fall. Gil Luria, an analyst at D.A. Davidson, said shrinking models would not eliminate the need for processors or memory. It could simply move more of those chips from datacenters into phones and other devices. "It's not that you're not going to need the chip," Luria said. "You're still going to need the GPU, and you're still going to need the memory." He added that running AI on individual devices can actually be less efficient than using shared datacenter infrastructure because chips in phones may sit idle much of the time. Efficiency breakthroughs can also lead to more use rather than lower spending, as cheaper and faster AI enables new products and prompts consumers to run models more often. Still, the market has been quick to punish anything that suggests AI may need less memory than expected. Micron shares plunged in March after Google published its TurboQuant paper on cutting memory use without hurting model performance, though the stock later recovered. PrismML's public release gives everyday users and investors a chance to test whether its claimed gains hold up outside the lab. And for Apple, running more capable AI directly on the iPhone could help the company improve Siri without abandoning the privacy and hardware integration that distinguish its products. "The combination of cloud and on-device AI can serve a more complete, efficient and privacy-centric AI experience," Counterpoint's Pathak said. "Complex tasks will be offloaded to the cloud, whereas sensitive, latency-critical and privacy-relevant tasks will be executed on-device." Choose CNBC as your preferred source on Google and never miss a moment from the most trusted name in business news.
[2]
Apple eyes PrismML's on-device AI for the iPhone
Apple is in early talks with PrismML, a Caltech spinout whose on-device AI compression shrinks a 27-billion-parameter model from 54GB to under 4GB, small enough to run on an iPhone. CEO Babak Hassibi told CNBC the talks are early, and Apple has not commented. Apple is in early talks with PrismML, a startup that shrinks large AI models to run directly on a phone. The company's on-device AI pitch could help Apple keep more of Siri's work off the cloud. PrismML chief executive Babak Hassibi told CNBC that Apple and other companies were evaluating its technology. He called the discussions very early and said it was unclear where they would lead, but that "things are progressing nicely." Apple did not comment. The Information first reported Apple's interest last week. The startup is a Khosla Ventures-backed spinout from the California Institute of Technology. Caltech owns the underlying patents and licenses them exclusively to PrismML. The company raised a $16.25 million seed round in March. What PrismML built On Tuesday, PrismML released Bonsai 27B. It is a compressed build of Alibaba's open-source Qwen model, not a new one trained from scratch. The company shrank it from roughly 54GB to as little as 3.9GB. PrismML ships two versions under a free licence. A ternary build runs on a laptop. A smaller 1-bit build, about 3.9GB, is designed to fit within the memory budget of an iPhone 17 Pro. PrismML says it is the first model of that size to run on a phone. The trick is how the model stores its internal values. PrismML reduces each one from 16 bits to just one or three possible values. It says this cuts memory use by 10 to 15 times, speeds up responses by six to eight times, and lowers energy use by three to six times. There is a cost. Hassibi said the compressed models lose a few percentage points of performance. Factual recall weakens first, he said, before skills such as reasoning, maths, and coding. PrismML says its builds keep about 95% of full performance in the ternary version and 90% in the 1-bit one. Why Apple cares The timing is not an accident. PrismML released the model a day after Apple opened the public beta of iOS 27, which carries its long-delayed Siri overhaul. Apple is trying to make Siri competitive with assistants from OpenAI and Anthropic. Running more AI on the device would help. Apple already sends complex requests to cloud models. Keeping more work local would cut delay, lower cloud costs, and support the company's privacy pitch. Some features would also work offline. There is a cost angle too. Morgan Stanley estimates Apple's memory costs could climb sharply in its 2027 financial year. The bank expects the company to raise iPhone prices to protect margins. Smaller models help Apple fit capable AI into tight hardware without paying for more memory. Claims still to be proven Analysts urged caution. Tarun Pathak of Counterpoint Research said the real test would be millions of queries across thousands of devices. Phil Solis of IDC said power use was the biggest open question, since a model that runs often could still drain a battery. The release also feeds a debate over whether efficiency gains will cut demand for memory and data-centre chips. Gil Luria, an analyst at D.A. Davidson, said shrinking models would not remove the need for processors. It would simply move some of them from data centres onto phones, part of a broader shift toward edge AI. Hassibi said Google's open-source Gemma model is next in the pipeline, followed by larger frontier models. "It's very important that the intelligence be local and that it can run fast," he said.
[3]
PrismML releases Bonsai 27B, claiming first major AI model of its size fit for iPhone
PrismML is back in the news today after the AI startup's CEO told CNBC that Apple is looking at the company's technology. The company has also released its Bonsai 27B model that it says runs on iPhone, iPad, and Mac. PrismML first made headlines last week when The Information reported on the AI startup. That report included mention of PrismML holding meetings with Apple about "ways it could use its technology." The report also noted PrismML's Bonsai 27B model, scheduled for release today. As expected, PrismML has now released that model: Bonsai 27B reaches up to 163 tok/s in 1-bit and 134 tok/s in Ternary on an NVIDIA GeForce RTX 5090. On an M5 Max, it reaches up to 87 tok/s in 1-bit and 58 tok/s in Ternary. Fitting a phone is a stricter gate than storage numbers suggest. A phone never exposes its full memory to an app - a 12 GB iPhone offers about 6 GB for the model to use on-device, and the model shares that budget with its KV cache and activations. No conventional build of a 27B model comes close to clearing it. At about 4 GB, 1-bit Bonsai 27B is the first to pass through with room to work. The press release specifically mentions running the new model natively on Apple hardware: Bonsai 27B runs natively on Apple devices (Mac, iPhone, iPad) via MLX and on NVIDIA GPUs via CUDA, through custom low-bit kernels built for its hybrid-attention architecture. Model weights are available today under the Apache 2.0 License. With this release, we're offering a free, limited-time developer preview API so developers can easily try our model. Meanwhile, PrismML CEO Babak Hassibi tells CNBC that Apple is interested in PrismML's tech: PrismML CEO Babak Hassibi told CNBC that Apple and other companies have been evaluating the startup's models and measuring their speed, energy efficiency and performance on devices. "They're really evaluating our technology right now," Hassibi said of Apple. He characterized the discussions as very early and said it remains unclear where they will lead, but that "things are progressing nicely." Apple did not immediately respond to a request for comment. You can learn more about PrismML's Bonsai 27B model here, and read the CNBC piece in full here. 9to5Mac's Take It's worth separating two things here. We'll leave it to the AI experts to evaluate PrismML's claims and technology. As for the Apple connection, it's clear that the AI startup is touting communication with Apple as a way to generate buzz around its new release. It's generally wise to keep any serious discussions between startups and Apple on the quieter side. Otherwise, an environment for strong skepticism is created. At any rate, PrismML has managed to capture attention from The Information and CNBC around their latest release. Short-term mission accomplished.
[4]
Apple Exploring Ways to Run Much Larger AI Models Directly on iPhones
Apple has held meetings with PrismML about ways it could use the startup's technology to run much larger AI models directly on iPhones, according to The Information. The report said PrismML has managed to shrink down Alibaba's open-source large language model Qwen 3.6 to run entirely on an iPhone 17 Pro. The model has 27 billion parameters, which is larger than Apple's on-device AFM 3 Core Advanced model with 20 billion parameters. Apple's model powers iOS 27 enhancements such as Siri AI's more expressive voices and improved systemwide dictation on iPhone 17 Pro and iPhone Air models. Unlike with AFM 3 Core Advanced, all of Qwen 3.6's parameters can be active at the same time. "One new on-device Apple model has 20 billion parameters but uses a so-called sparse architecture, in which only 1 billion to 4 billion parameters are active at a time," the report said, in reference to AFM 3 Core Advanced. "In the case of PrismML's on-device model, all 27 billion parameters are active at the same time." Larger models running directly on iPhones would allow for more Apple Intelligence features to run on device instead of on Apple's Private Cloud Compute servers, which could reduce Apple's costs and further enhance user privacy.
[5]
AppleInsider.com
Five days after initial reports that there was an AI-model shrinking technology by PrismML suitable for iPhones, the company has confirmed it is talking to Apple over the use of its technology. One of the major problems with AI processing is the need to manage massive models that most normal computers and smartphones cannot easily handle alone. On July 9, the start-up PrismML surfaced with a solution to shrink down the models to a more manageable size. A few days later, on July 14, the startup confirmed the talks were underway. Speaking to CNBC, startup CEO Babak Hassibi said that Apple and other companies are evaluating its models. "They're really evaluating our technology right now," he said of Apple. Hassibi didn't go into detail about what the talks are about, but that they are very early and are "progressing nicely." These talks could range from simply licensing the technology from PrismML to outright offering to buy the startup completely. Small models, big deal For Apple, the technology that PrismML has developed could be extremely advantageous. As Apple is keen to keep as much AI processing on-device as possible, it is limited in what it can do at the moment due to size constraints. PrismML's work allows it to shrink the size of models down considerably. The work appears to make it feasable to get it to a size where it could run in an iPhone's memory, without needing cloud access.
[6]
Report: Apple interested in startup that runs giant AI models on iPhone without servers
The Information reports that Apple may be interested in PrismML's technology. The firm is focused on shrinking AI models that generally require servers to function, offering on-device functionality with comparable intelligence. PrismML may be key to unlocking more powerful models on-device "The startup, PrismML, said it has shrunk down Qwen 3.6, an open-source large language model developed by Chinese internet giant Alibaba, to run on an iPhone 17 Pro," per The Information. "The model has 27 billion parameters, which are roughly similar to the synapses in a brain and can help determine the complexity of the data a model can process. In contrast, most models that run on mobile phones have only a few billion parameters active at a time." The report goes on to say that PrismML plans to release its open-source model on Tuesday, July 14, and that it's capable of tasks like software development. Later in the piece, The Information reports that Apple has already held meetings with PrismML about "ways it could use its technology," citing people familiar with the talks. In terms of acquisitions, Apple's biggest secretive AI play has been acquiring Q.ai, a startup that reportedly went for $2 billion. Meanwhile, Apple worked with Google to develop Siri AI, the modern version of Siri that will debut with iOS 27. You can read the piece in full here. You can learn more about PrismML here.
[7]
Meet Bonsai: The First 27B AI Model That Fits on Your Phone
Apple is in early talks with PrismML about the underlying compression technology, per CNBC, with the company targeting a compressed Gemma model next in the pipeline. I models eat up a lot of memory. A 27-billion-parameter AI model, considered medium-sized by industry standards, needs roughly 54 GB of memory to run on half precision. Most laptops can't hold that. Some desktop rigs can't either. Earlier this week, PrismML released one at 3.9 GB -- small enough to fit on an iPhone. Parameters are the number of dials and tweaks a model can handle. The more parameters, the denser and more capable a model is. Bonsai 27B is the first 27B-class model to clear the memory ceiling of a consumer smartphone, running at 11 tokens per second on an iPhone 17 Pro Max. (Tokens are the basic unit of information that AI models can handle and produce.) The ternary variant, at 5.9 GB, hits around 26 tokens per second on an M5 Pro laptop. Both are free under Apache 2.0. The compression method, built on Caltech intellectual property, reduces each model weight from 16 bits of floating-point precision to a single sign -- +1 or -1 in the binary build, one of three values in the ternary. Each group of 128 weights shares a 16-bit scaling factor, landing the binary variant at 1.125 bits per weight: 14 times smaller than the full-precision original. The ternary model adds a zero state for slightly more expressive power and settles at 1.71 bits. In easier terms, this means a ternary AI model uses only three settings for each internal value -- negative, zero, or positive -- while a standard AI can choose from about 65,000 settings. PrismML did that without losing much of the output quality. What makes this different from conventional "low-bit" models is that nothing gets a higher-precision escape hatch: embeddings, attention, and the full language model head are all compressed end-to-end. Most quantized builds keep certain sensitive layers at full precision, which ends up increasing their size as a tradeoff for better quality. Bonsai doesn't play that game. This is the second major release in the family. In March, PrismML shipped Bonsai 8B, a 1.15 GB model that proved the 1-bit architecture could survive at 8 billion parameters without its reasoning collapsing. The jump to 27 billion is where the stakes change -- that scale is where sustained chain-of-thought reasoning, reliable tool use, and multi-step agentic behavior actually emerge consistently -- the things smaller models still fumble. Benchmarks Across 15 benchmarks evaluated in thinking mode on NVIDIA H100 GPUs -- spanning knowledge, math, coding, and tool use -- Ternary Bonsai 27B averages 80.49, or 94.6% of the full-precision model. The 1-bit variant hits 76.11. Overall, on benchmarks, the models perform much better than Gemma 4 or Qwen 3.6 in terms of how much potential they offer for their size. The models are pretty good for what they offer, and considering how little resources they require, they take small hardware (smartphones and lower end PCs) to another level in terms of capabilities. AIME25 and AIME26, modeled on the American Invitational Mathematics Examination, come in 93.7% for Ternary Bonsai 27B versus 95.3% for the much bigger Qwen 3.6B. Bonsai scores 86 points in codig vs 88 for Qwen 3.6 and 77% on general knowledge vs 83 for Qwen 3.6. The model also uses a hybrid attention backbone where roughly 75% of the layers are linear rather than full quadratic attention. That architecture is what makes a 262K-token context window practical on-device -- something a standard attention stack would make prohibitively expensive on phone hardware. We tested it We ran Bonsai 27B ourselves. Coding takes iteration: single-shot prompts won't compete with cloud frontier models. Being local and free makes that irrelevant. For our Zombie Type game -- a first-person typing-horror browser game -- two vibe coding rounds produced clean collision detection, proper scoring logic, and graphics that held together. The model grasps structure early; the second pass refines rather than rebuilds. Interestingly enough, some models (like the skeletons) looked more elaborate than the ones from GPT 5.6 Sol. It doesn't mean it's better by any means, just that on this task it produced a cute skeleton whereas the AI king made a poorer stylistic choice. The game is available for testing here. Creative writing is a more qualified story, and the criteria is more subjective. Roughly speaking, the results aren't particularly imaginative if you have a zero-shot prompt in mind. That said, Bonsai produces stories with consistent internal logic, pacing, and arc -- better, or on par with Claude Haiku or even Sonnet on lower effort on comparable prompts. For a model that runs entirely on your own hardware with no API costs, that's a lot to say. The story it created can be found in our Github repository. PrismML also ships a DSpark speculative decoding layer alongside the model -- a lightweight drafter that proposes blocks of candidate tokens, which the main model verifies in a single forward pass rather than generating token-by-token. On an H100 that adds a 1.37x throughput boost with no change in output quality, since verification preserves the exact output distribution. On Apple Silicon it's not yet enabled by default, but for GPU serving it's a real gain. Apple's interest adds a commercial dimension. PrismML CEO Babak Hassibi confirmed to CNBC that the company is in early talks with Apple, which is evaluating the compression technology for potential on-device use. Hassibi said a compressed Gemma model is next in the pipeline, followed by larger frontier models; 1-bit Bonsai 27B is available for free download now under Apache 2.0. If you need a primer on running models like this locally, check out our guide.
[8]
AppleInsider.com
PrismML, a startup that uses mathematical wizardry to shrink the size of large language models, could be used to bring more advanced AI capabilities to iPhones in the future. The company has already been able to compress the 54GB Qwen 3.6 model to just 4GB. And, importantly, the technology that it uses doesn't impact the model's performance. Apple is already keen to find ways to shrink LLMs that would normally run on servers. The company prefers such models to run on-device, but has so far been limited by the huge amount of data that more advanced models require. Now, according to a report by The Information, Apple has identified PrismML as a potential partner in its AI-shrinking efforts. Billions of parameters The Qwen 3.6 model is so large because of its 27 billion parameters, a figure that allows it to perform more advanced tasks than smaller models. The models that run on phones normally have just a few billion parameters, limiting their capabilities. It's reported that the Qwen 3.6 model that PrismML has already made to work on an iPhone is capable of complex chat and reasoning. It can also power fully autonomous agents. While Apple is already working to shrink models for iPhone use, performance has been an issue. But PrismML's technology has no such issue, making it particularly appealing. With that in mind, it's no surprise that Apple would take notice. The report claims the company has already held meetings with PrismML to discuss the ways its technology could be used to bring more advanced AI technologies to the iPhone. Ultimately, there's no guarantee the two companies will work together. And even if they do, it's unknown when iPhone and Siri AI owners could expect to benefit from the partnership.
[9]
A Startup Says It Shrunk an AI Model by 93%. Apple Wants to Talk.
Honey, I shrunk the AI. That's startup PrismML's message to Apple, and it could solve one of the company's biggest headaches. PrismML, a Caltech spinout backed by Khosla Ventures, says it shrank a 54GB AI model down to under 4GB, small enough to run all 27 billion of its parameters directly on an iPhone 15 or newer, CNBC reports. "They're really evaluating our technology right now," PrismML CEO Babak Hassibi said of Apple, calling the talks early but promising. The company raised a $16.25 million seed round in March and just released two free compressed AI models built to run on iPhones, MacBooks and Nvidia-powered PCs. The breakthrough could let Apple run more powerful AI locally instead of relying on the cloud, boosting speed, privacy and battery life for a revamped Siri. Analysts caution the claims still need real-world testing, and shrinking AI models won't necessarily shrink the broader industry's appetite for chips.
[10]
Apple eyeing startup PrismML to bring massive AI models direct to iPhone By Investing.com
Apple is in talks with PrismML, a stealthy California Institute of Technology (Caltech) spinout, as it looks to drastically reduce its reliance on the cloud and run heavyweight AI entirely on-device. The discussions, first reported by The Information, come on the heels of a massive technical breakthrough from the startup: PrismML successfully compressed Alibaba's 27-billion-parameter open-source Qwen 3.6 model to run locally on an iPhone 17 Pro. Typically, cramming a model of that caliber onto a smartphone requires a "sparse architecture" -- meaning the phone only wakes up a fraction of the AI's brain at any given moment to keep the device from melting. PrismML completely flips the script. * The Shrink: Squeezed the model from a massive 54 GB down to less than 4 GB. * The Brainpower: Keeps all 27 billion parameters active simultaneously without sacrificing benchmark performance. * The Core Tech: Built on ultra-dense 1-bit and ternary weight architectures (reducing memory footprint up to 14x while running up to 8x faster). Why this matters for Apple: Local AI means total user privacy, lightning-fast response times, and zero reliance on a cellular connection. It also saves Apple from paying astronomical cloud server bills. At WWDC, Apple made waves by introducing a revamped Siri architecture. However, many of those advanced features still rely on sending data off-device to Google's Gemini models in the cloud. By actively hunting for acquisition targets like PrismML, Apple is aiming to shift those heavy AI workloads back into your pocket. While rivals like Meta, Microsoft, and Amazon are committing hundreds of billions to build out massive AI data centers, Apple's ultimate power move might just be making the data center obsolete for the everyday user.
Share
Copy Link
Apple is evaluating technology from PrismML that could shrink powerful AI models from 54GB to under 4GB, enabling them to run entirely on an iPhone. The Caltech spinout's compression approach could reduce Apple's cloud computing costs, enhance privacy, and allow Siri to process more requests locally without internet connectivity.
Apple is in early-stage discussions with PrismML, a Caltech spinout backed by Khosla Ventures, about technology that could fundamentally change how AI models for iPhone operate. PrismML CEO Babak Hassibi confirmed to CNBC that Apple and other companies are evaluating the startup's models, measuring their speed, energy efficiency, and performance on devices
1
. "They're really evaluating our technology right now," Hassibi said of Apple, characterizing the discussions as very early but noting that "things are progressing nicely"1
. The timing aligns with Apple's ongoing efforts to make Siri more competitive while addressing one of the central constraints facing its AI strategy: most capable models require too much memory and processing power to run AI models on iPhone hardware alone.
Source: AppleInsider
On Tuesday, PrismML publicly released Bonsai 27B, compressed versions of the Alibaba Qwen model that demonstrate the startup's core innovation
1
. The company reduced the model from roughly 54GB to less than 4GB, allowing all 27 billion of its parameters to run on an iPhone 15 or newer1
. This represents a significant advance in on-device AI capabilities. Unlike Apple's AFM 3 Core Advanced model with 20 billion parameters that uses a sparse architecture where only 1 billion to 4 billion parameters are active at a time, all 27 billion parameters in PrismML's compressed AI models remain active simultaneously[4](https://www.macrumor ar-147521 s.com/2026/07/09/apple-prismml-larger-on-device-ai-models/). The release came one day after Apple opened the public beta of iOS 27, giving iPhone owners their first broad access to the company's long-delayed Siri overhaul1
.PrismML achieves dramatic size reduction by drastically simplifying how internal information is stored, reducing each value from 16 bits to just one or three possible values
1
. Hassibi compared it to the chip industry's move from eight-bit to four-bit computing, but said PrismML takes it a step further1
. The startup reports that compressed AI models use between 10 and 15 times less memory usage, generate responses six to eight times faster, and consume three to six times less energy efficiency than conventional versions running on existing hardware1
. However, there is a trade-off. PrismML's models typically lose a few percentage points of overall performance, with factual recall weakening before skills such as reasoning, math, and coding1
. PrismML says its builds keep about 95% of full performance in the ternary version and 90% in the 1-bit one2
.The technology could address critical challenges in Apple's AI strategy by enabling more on-device processing rather than relying on Private Cloud Compute servers. Carolina Milanesi, president and principal analyst at Creative Strategies, said smaller models could let Apple move more demanding features onto the iPhone, including computational photography, video generation, and health or fitness tools that rely on sensitive personal data
1
. "The more you can do on device, the better it is," she said, pointing to health and medication data that users would want to keep private1
. Running more AI directly on the iPhone would reduce latency associated with sending data to remote servers, lower cloud-computing costs, and support Apple's privacy pitch1
. It would also allow certain Apple Intelligence features to work without an internet connection, addressing privacy concerns that have become central to Apple's positioning against competitors like OpenAI and Anthropic.
Source: AppleInsider
Related Stories
Bonsai 27B runs natively on Apple devices including Mac, iPhone, and iPad via Apple's MLX framework, and on NVIDIA GPUs via CUDA through custom low-bit kernels built for its hybrid-attention architecture
3
. PrismML ships two versions under a free Apache 2.0 License: a ternary build that runs on a laptop, and a smaller 1-bit build at about 3.9GB designed to fit within the memory budget of an iPhone 17 Pro2
. Fitting a phone is a stricter gate than storage numbers suggest, since a phone never exposes its full memory to an app—a 12GB iPhone offers about 6GB for the model to use for on-device processing, and the model shares that budget with its KV cache and activations3
. At about 4GB, 1-bit Bonsai 27B is the first to pass through with room to work3
.Analysts urged caution about real-world performance. Tarun Pathak of Counterpoint Research said the real test would be millions of queries across thousands of devices, while Phil Solis of IDC said power use was the biggest open question, since a model that runs often could still drain a battery
2
. The release also feeds a debate over whether efficiency gains will cut demand for memory and data-center chips. Gil Luria, an analyst at D.A. Davidson, said shrinking models would not remove the need for processors but would simply move some of them from data centers onto phones, part of a broader shift toward edge AI2
. Hassibi said Google's open-source Gemma model is next in the pipeline, followed by much larger models, including those from frontier labs that today generally require datacenter hardware1
. The technology could ultimately extend well beyond phones and laptops to robotics, autonomous systems, and other products that need to make decisions quickly without relying on a cloud connection1
. The California Institute of Technology owns the underlying patents and licenses them exclusively to PrismML, which raised a $16.25 million seed round in March backed by Khosla Ventures and other investors1
.Summarized by
Navi
[2]
[5]
08 Feb 2025•Business and Economy

02 Jan 2026•Technology

09 Oct 2025•Technology
