PrismML released Bonsai 2 27B, compressing Alibaba's Qwen3.8 27B model from 56 GB to just 5.9 GB while retaining 98% of its capabilities. Using ternary weights compression, the startup is making powerful AI accessible on PCs and high-end smartphones without cloud dependency.

PrismML Releases Bonsai 2 27B With Breakthrough AI Compression

PrismML, an AI lab founded by Caltech researchers, released Bonsai 2 27B on Thursday, marking a significant advance in compressed AI model technology

1

. The startup compressed Alibaba's widely-used open-source Qwen3.8 27B model from approximately 56 GB down to just 5.9 GB, achieving a 9x to 10x reduction in memory while retaining 98% of the original model's aggregate benchmark scores

2

. This compression makes the AI model small enough to run on PCs and high-end smartphones, fundamentally changing how users can access powerful AI capabilities without relying on cloud computing.

Led by CEO Babak Hassibi, a Caltech professor and compression technology expert, PrismML has raised a $22.25 million seed round backed by Khosla Ventures, Cerberus Capital, and Caltech

1

. The company also counts Ion Stoica, co-founder of Databricks and director of Berkeley's Sky Computing Lab, as an advisor. The first Bonsai model, released just months ago in March with 95% performance retention, has already been downloaded over 11 million times, with PrismML's smaller models accumulating another 2.6 million downloads

1

.

How Ternary Compression Technique Achieves Dramatic Size Reduction

PrismML achieves this breakthrough using a ternary compression technique that fundamentally reimagines how AI model weights are stored

1

. Weights are the numerical parameters that control how models process information and generate outputs. In traditional full-size models, each weight requires 16 bits of storage. PrismML's approach simplifies these ternary weights down to just three values: +1, -1, or 0, represented by three bits

2

. This dramatic reduction in storage requirements for each weight allows the model to occupy significantly less space while maintaining intelligence.

Unlike standard quantization methods that typically strip away accuracy, knowledge, and systematic capabilities during compression, PrismML's ternary approach preserves model performance more effectively

2

. Qwen3.8's minimal memory footprint using conventional compression sits around 9.4 GB, but Bonsai 2 27B pushes well below that threshold at 5.9 GB. The improvement from the first Bonsai's 95% performance retention to Bonsai 2's 98% demonstrates the rapid advancement of this AI compression technology.

Benchmark Performance Shows Bonsai 2 Punches Above Its Weight Class

Bonsai 2 27B demonstrated competitive performance across multiple benchmark categories despite its compressed size

2

. On agentic tasks and tool calling, it scored 77.6 compared to Qwen3.8's 79.8, staying within 3 points of the original. For coding benchmarks across HumanEval+, LiveCodeBench v6, MBPP+, and BigCodeBench, Bonsai 2 achieved aggregate scores of 81.6 versus Qwen3.8's 82.2. On knowledge and reasoning tests including MMLU-Redux, GPQA Diamond, and AA-LCR, it scored 82.7 compared to the original's 81.3, actually outperforming in this category.

Hashibi acknowledged that achieving perfect 100% benchmark performance parity remains uncertain, as compression will likely always have some impact

1

. However, he noted that perfect parity is "fairly academic" since uncompressed models aren't perfectly accurate either, and benchmarks don't perfectly reflect actual tasks. A 2% degradation is unlikely to meaningfully affect real-world performance, especially considering that surrounding software infrastructure significantly impacts accuracy.

Consumer Hardware Can Now Run Powerful AI Models Locally

Source: TechCrunch

Source: TechCrunch

Bonsai 2 27B can run on consumer hardware including Nvidia GeForce GTX 5090 cards without quantization, reaching 143 tokens per second, and on Apple M5 Max chips at 46.8 tokens per second

2

. The model runs on Nvidia GPUs via CUDA and on Apple devices, including Mac, iPhone, and iPad, via MLX through low-bit kernels. Model weights are available under Apache 2.0 license, making them freely accessible for developers and researchers.

The ability to run on consumer hardware enables local inference, eliminating the need to send data to cloud servers. "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud," said Ion Stoica

1

. This privacy-preserving approach addresses growing concerns about data sharing with third parties and helps organizations meet strict privacy regulations while improving security.

Energy Efficiency and Practical Applications Transform AI Deployment

Bonsai 2 27B consumes just 0.714 megawatt-hours per token, making it 40% more energy-efficient than other 8B models running at full precision

2

. This energy efficiency, combined with low-latency applications enabled by local processing, opens new deployment scenarios. Simple tasks like translation, summarization, and search organization can run on-device, while complex long-horizon task comprehension and research requiring more computational power can still be routed to cloud-based models when necessary.

This hybrid approach lets users and enterprises balance cost, privacy, and capability requirements. Sensitive information can be processed locally using the compressed AI model on PCs and high-end smartphones, while computationally intensive work scales up to cloud computing resources. PrismML is rumored to be in talks with Apple, though Hassibi declined to comment on this speculation

1

.

PrismML Plans Even Larger Compressed Models in Coming Months

Source: SiliconANGLE

Source: SiliconANGLE

Hashibi revealed that PrismML's next goal involves applying the ternary compression technique to even larger models in the several-hundred-billion-parameter range, expected within the next couple of months

1

. He anticipates that retaining intelligence will actually become easier with larger models: "There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it's easier to get to 100%."

While PrismML isn't the only company working on compression technology—Multiverse Computing, founded by a professor from Spain's Donostia International Physics Center, is another notable competitor—Hassibi maintains that PrismML's approach is unique in how little performance degradation occurs

1

. The rapid improvement from Bonsai 1 to Bonsai 2, combined with millions of downloads and backing from prominent investors and advisors, positions PrismML as a significant player in making advanced multimodal AI model capabilities accessible on consumer devices.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved