4 Sources
[1]
Meta releases compact versions of Llama 3.2 AI models
The new quantised models are 56pc smaller and use 41pc less memory when compared to the full-size model released last month. Meta has released compact versions of its lightweight Llama 3.2 1B and 3B models that are small enough to run effectively on mobile devices. The Facebook owner, in an
[2]
Meta Releases Quantized Llama 3.2 with 4x Inference Speed on Android Phones
These models reduce memory usage by an average of 41% and decrease model size by 56% compared to the initial BF16 format. Meta has introduced quantized versions of its Llama 3.2 models, enhancing on-device AI performance with up to four times faster inference speeds, a 56% model size reduction,
[3]
Meta debuts slimmed-down Llama models for low-powered devices - SiliconANGLE
Meta debuts slimmed-down Llama models for low-powered devices Meta Platforms Inc. is striving to make its popular open-source large language models more accessible with the release of "quantized" versions of the Llama 3.2 1B and Llama 3B models, designed to run on low-powered devices. The Llama
[4]
Introducing quantized Llama models with increased speed and a reduced memory footprint
Similar improvements were observed for the 3B model. See results section for more details. At Connect 2024 last month, we open sourced Llama 3.2 1B and 3B -- our smallest models yet -- to address the demand for on-device and edge deployments. Since their release, we've seen not just how the
Share
Copy Link
Meta has released compact versions of its Llama 3.2 1B and 3B AI models, optimized for mobile devices with reduced size and memory usage while maintaining performance.

Meta has unveiled quantized versions of its Llama 3.2 1B and 3B AI models, marking a significant advancement in on-device artificial intelligence capabilities. These compact models, designed to run efficiently on mobile devices, offer improved performance while maintaining the quality and safety standards of their original counterparts
1
2
.The new quantized models boast impressive enhancements:
These improvements enable the models to operate effectively on resource-constrained devices, such as smartphones
3
4
.Meta employed two primary quantization techniques to achieve these results:
Quantization-Aware Training (QAT) with LoRA adaptors: This method optimizes performance in low-precision environments while prioritizing accuracy
2
4
.SpinQuant: A technique that focuses on model portability, allowing for substantial compression without compromising inference quality
2
4
.The development of these quantized models involved close collaboration with industry leaders:
3
.1
4
.3
.This collaborative effort ensures that the models are well-suited for a wide range of mobile devices and can leverage specific hardware capabilities for optimal performance
3
4
.The quantized Llama 3.2 models open up new possibilities for on-device AI applications, including:
1
3
Related Stories
Meta is exploring additional performance gains through Neural Processing Unit (NPU) support, working with partners to integrate NPU functionalities within the ExecuTorch open-source ecosystem. This effort aims to further optimize the quantized models for a broader range of devices
2
4
.The quantized Llama 3.2 1B and 3B models are now available for download from Llama.com and Hugging Face. This release allows developers to create unique AI experiences with enhanced privacy, as all interactions can take place directly on the user's device
3
4
.The release of these optimized models represents a significant step towards making advanced AI capabilities more accessible on everyday devices. By reducing the computational and memory requirements, Meta is enabling a wider range of applications and use cases for on-device AI, potentially accelerating innovation in mobile AI technologies
1
2
3
4
.Summarized by
Navi
[1]
1
Technology

2
Policy and Regulation

3
Technology
