PrismML, founded by Caltech researchers, has compressed AI models to fit on smartphones and smart glasses. The startup's Bonsai 2 27B shrinks Qwen3.8 from 56GB to 5.9GB while retaining 98% performance. Qualcomm showcased the 1-bit version for AR glasses, enabling real-time vision queries without cloud dependency.

PrismML Demonstrates AI Compression Breakthrough with Qualcomm Integration

PrismML, an AI lab founded by Caltech researchers and advised by UC Berkeley's Ion Stoica, showcased its 1-bit Bonsai LLM running on Qualcomm's Snapdragon AR1 Gen 1 Platform at the Snapdragon Summit

1

. The demonstration marks a significant step toward bringing powerful AI models to AI smart glasses and wearable devices. The tiny LLM enables users to ask real-time questions about what they're seeing, with all processing happening locally on consumer hardware rather than in the cloud.

Source: TechCrunch

Source: TechCrunch

Led by CEO Babak Hassibi, a Caltech professor specializing in AI compression technologies, PrismML has raised $22.25 million in seed funding from Khosla Ventures, Cerberus Capital, and Caltech

2

. The startup's core mission centers on creating open-weight AI that can run locally on devices, offering an alternative to proprietary AI labs that require massive computing resources and raise privacy concerns.

Bonsai 2 27B Achieves 9x Compression with Minimal Performance Loss

PrismML released Bonsai 2 27B, a compact multimodal AI model that compresses Qwen3.8 27B from approximately 56GB down to just 5.9GB—a 9x to 10x reduction in memory

2

3

. This makes the model small enough to run on PCs and potentially high-end smartphones. The breakthrough lies in retaining 98% of Qwen's aggregate benchmark scores, a significant improvement from the first Bonsai model released in March that matched 95% performance.

Source: TechCrunch

Source: TechCrunch

The original Bonsai model has been downloaded over 11 million times, with PrismML's smaller models accumulating another 2.6 million downloads

2

. On specific benchmarks, Bonsai 2 demonstrated near-parity with Qwen3.8, scoring 77.6 versus 79.8 on agentic and tool calling tasks, 81.6 versus 82.2 for coding across HumanEval+, LiveCodeBench v6, MBPP+ and BigCodeBench, and 82.7 versus 81.3 for knowledge and reasoning

3

.

Ternary Compression Technique Enables Radical Size Reduction

PrismML achieves this compression through a ternary compression technique that shrinks the "weights" making up a model

2

. While standard models require 16 bits per weight, PrismML's approach uses ternary weights simplified to three values: +1, -1, or 0. This dramatically reduces the storage space needed for each weight without catastrophic performance degradation that typically accompanies other compression methods like quantization.

Hassibi explained that larger models actually offer more room for compression: "There is more room to be able to compress them without losing the intelligence. So I would just say, as a general trend, for larger models, it's easier to get to 100%"

2

. The startup plans to apply this technique to models in the several-hundred-billion-parameter range within the next couple of months.

Performance Metrics Show Energy Efficiency and Speed Gains

Bonsai 2 can run locally on devices with impressive performance metrics. On an Nvidia GeForce GTX 5090 card without quantization, the model reaches 143 tokens per second, while achieving 46.8 tokens per second on Apple's M5 Max chip

3

. The model consumes just 0.714 megawatt-hours per token, making it 40% more energy-efficient than other 8B models running at full precision.

The model runs on Nvidia GPUs via CUDA and on Apple devices including Mac, iPhone and iPad via MLX through low-bit kernels. Model weights are available under Apache 2.0 license, supporting PrismML's vision of open-weight AI accessible to developers and researchers

3

.

Local Inference Addresses Privacy and Latency Concerns

Ion Stoica emphasized the transformative potential of running advanced models on users' devices: "You are going to have intelligence at your fingertips, and it's going to be free because it's going to run on the device you already bought. It's also going to be private, because you're not going to send it to the cloud"

2

.

Source: SiliconANGLE

Source: SiliconANGLE

Local inference eliminates the need to send sensitive data across the internet to third-party servers, addressing privacy concerns while reducing latency for low-latency applications. Users and enterprises can run simple tasks like translation, summarization, and search organization on-device with privacy-preserving benefits, while scaling up to cloud-based models only for complex, high-touch work. This hybrid approach helps meet strict privacy regulations and improves security by keeping prompts and responses local.

Watch for PrismML's rumored discussions with Apple and the release of even larger compressed models in coming months. The Qualcomm Snapdragon platform integration signals growing industry interest, though no smart glasses running PrismML have been announced yet.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved