3 Sources
[1]
New Apple model recreates 3D objects with realistic lighting effects - 9to5Mac
Apple researchers have created an IA model that reconstructs a 3D object from a single image, while keeping reflections, highlights, and other effects consistent across different viewing angles. Here are the details. While the concept of latent space in machine learning is not exactly new, it has
[2]
Apple can create 3D objects with realistic lighting effects from a single image with their new AI model
9to5mac.com can now tell us something interesting. Apple's researchers have created an AI model that reconstructs a 3D object from a single image, while "keeping reflections, highlights, and other effects consistent across different viewing angles". In Apple's new study, titled LiTo: Surface Light
[3]
Apple's New LiTo AI Turns Photos into Hyperreal 3D Objects: Here's How it Works
Apple has recently launched LiTo, a new AI model that can reconstruct 3D objects from an image while accurately preserving lighting effects like reflections and highlights. The end product is realistic and better than previous techniques. The model transforms visual information into numerical data
Share
Copy Link
Apple researchers have developed LiTo, an AI model that reconstructs 3D objects from a single image while preserving realistic lighting effects like reflections and highlights across different viewing angles. The Surface Light Field Tokenization approach uses latent space to jointly model object geometry and view-dependent appearance, outperforming existing methods that typically require multiple images.
Apple researchers have developed a groundbreaking AI model called LiTo that reconstructs hyperreal 3D objects from a single image while maintaining realistic lighting effects across different viewing angles
1
. The study, titled Surface Light Field Tokenization, introduces a novel approach that jointly models object geometry and view-dependent appearance within a unified framework. Unlike most prior works that focus on either reconstructing 3D geometry or predicting view-independent diffuse appearance, LiTo captures complex visual phenomena including specular highlights and Fresnel reflections under varying lighting conditions1
.
Source: Analytics Insight
The machine learning model achieves this feat by leveraging latent space, a mathematical representation that stores information about both an object's physical structure and how light interacts with its surface
3
. The process involves an encoder-decoder architecture where an encoder first compresses the image into a compact representation, then a decoder reconstructs it as a 3D object complete with shadows, reflections, and lighting changes3
. What distinguishes this approach is its ability to generate 3D objects from a single image, eliminating the need for more common methods that require images from different angles to enable 3D reconstruction2
.To train the AI model, Apple researchers selected thousands of objects rendered from 150 different viewing angles and 3 lighting conditions
1
. Rather than feeding all this information directly into the system, they randomly selected small subsets of these samples and compressed them into a latent representation. The decoder was then trained to reconstruct the full object and its appearance under different angles and light conditions from just that subset of data1
. Through this training process, the system learned to capture both the object's geometry and how its appearance changes depending on viewing direction. Subsequently, another model was trained to take a single image of an object and predict the corresponding latent representation, enabling the decoder to reconstruct the full 3D object with view-dependent effects1
.Related Stories
Apple published reconstruction comparisons between LiTo and an existing model called TRELLIS on the project page, demonstrating superior performance in capturing realistic lighting effects
1
. The ability to reconstruct 3D objects from a single image with accurate reflections, highlights, and other effects consistent across different viewing angles represents a significant advancement in computer vision and 3D modeling2
. This technology could have wide-ranging applications in augmented reality, product visualization, e-commerce, and digital content creation, particularly as Apple continues to develop its Vision Pro spatial computing platform. The research demonstrates how leveraging surface light field samples through RGB-depth images enables more accurate representation of complex lighting interactions, potentially setting a new standard for single-image 3D reconstruction methods.
Source: 9to5Mac
Summarized by
Navi
[2]
[3]
1
Technology

2
Technology

3
Science and Research
