2 Sources
[1]
New Apple AI model generates 3D scenes from just three images - 9to5Mac
Apple's Machine Learning team, in collaboration with researchers from Nanjing University and The Hong Kong University of Science and Technology, has announced an interesting 3D AI model called Matrix3D. This so-called Large Photogrammetry Model is able to reconstruct 3D objects and scenes from
[2]
Apple's New Matrix3D Model Can Turn Flat Images Into Dynamic 3D Scenes
The model was developed in partnership with Nanjing University and HKUST Apple researchers released a new artificial intelligence (AI) model that can generate 3D views from multiple 2D images. The large language model (LLM), dubbed Matrix3D, was developed by the company's Machine Learning team, in
Share
Copy Link
Apple's Machine Learning team, in collaboration with researchers from Nanjing University and HKUST, has developed Matrix3D, an innovative AI model that can generate detailed 3D scenes from just three 2D images.

Apple's Machine Learning team, in collaboration with researchers from Nanjing University and The Hong Kong University of Science and Technology, has introduced Matrix3D, a groundbreaking AI model that revolutionizes the process of generating 3D scenes from 2D images
1
2
. This Large Photogrammetry Model represents a significant leap forward in the field of artificial intelligence and computer vision.Matrix3D stands out from existing 3D rendering models by unifying the entire pipeline into a single process. Unlike current methods that rely on multiple models for various subtasks, Matrix3D performs pose estimation, depth prediction, and novel view synthesis all within a single large language model (LLM)
2
. This unified approach not only streamlines the workflow but also enhances accuracy by eliminating potential errors that can occur when transitioning between different models1
.The researchers employed a novel masked learning strategy to train Matrix3D, drawing inspiration from early Transformer-based AI systems that paved the way for models like ChatGPT. This technique involves randomly hiding parts of the input data during training, compelling the model to learn how to fill in the gaps
1
. This approach enables Matrix3D to train effectively even with smaller or incomplete datasets, enhancing its versatility and robustness.Matrix3D's ability to generate detailed 3D reconstructions of objects and entire environments from just three input images is particularly noteworthy
1
. This capability could have far-reaching implications for various applications, including potential integration with immersive technologies like the Apple Vision Pro1
.The model is based on a multimodal diffusion transformer (DiT) architecture, allowing it to integrate data across multiple modalities such as image data, camera parameters, and depth maps
2
. This sophisticated architecture enables Matrix3D to process complex inputs and generate accurate 3D representations.Related Stories
In a move that could accelerate further research and development in this field, Apple has made Matrix3D available to the open-source community. Researchers and developers can now download, modify, and redistribute the model via Apple's GitHub repository under a permissive license
2
. This decision reflects Apple's commitment to fostering innovation and collaboration in the AI community.While the full extent of Matrix3D's applications remains to be explored, its ability to generate 3D scenes from minimal input could have significant implications for various industries. From augmented reality and virtual reality to urban planning and digital twin technology, the potential use cases for this technology are vast and exciting.
Summarized by
Navi
17 Mar 2026•Technology

18 Dec 2025•Technology

07 Oct 2024•Technology

1
Technology

2
Technology

3
Policy and Regulation
