2 Sources
[1]
AI learns how vision and sound are connected, without human intervention
Caption: The researchers split the audio into smaller windows before the model computes its representations of the data, so it generates separate representations that correspond to each smaller window of audio. Pictured is a figure showing the separate representations of "speech" and "toot"
[2]
AI learns how vision and sound are connected, without human intervention
Humans naturally learn by making connections between sight and sound. For instance, we can watch someone playing the cello and recognize that the cellist's movements are generating the music we hear. A new approach developed by researchers from MIT and elsewhere improves an AI model's ability to
Share
Copy Link
MIT and IBM researchers have created an improved AI model called CAV-MAE Sync that can learn to associate audio and visual data from video clips without human labels, potentially revolutionizing multimodal content curation and robotics.
Researchers from MIT and other institutions have developed a groundbreaking AI model that can learn to connect visual and auditory information without human intervention. This advancement mimics the natural human ability to associate sights and sounds, such as linking a cellist's movements to the music being produced
1
.
Source: MIT
The new model, called CAV-MAE Sync, builds upon previous work and introduces several key improvements:
1
.2
.1
.
Source: Tech Xplore
The model processes unlabeled video clips, encoding visual and audio data separately into representations called tokens. It then learns to map corresponding pairs of audio and visual tokens close together within its internal representation space
2
.This technology has several promising applications:
1
.2
.1
.Related Stories
The study was conducted by a collaborative team including:
1
2
The research will be presented at the upcoming Conference on Computer Vision and Pattern Recognition (CVPR 2025) in Nashville
2
.This advancement in AI's ability to process multimodal information could have far-reaching implications. As Andrew Rouditchenko states, "We are building AI systems that can process the world like humans do, in terms of having both audio and visual information coming in at once and being able to seamlessly process both modalities"
1
. This development brings us one step closer to AI systems that can interpret the world in a more human-like manner, potentially revolutionizing various fields from entertainment to robotics.Summarized by
Navi
24 Jan 2025•Science and Research

10 Dec 2024•Science and Research

23 Oct 2024•Science and Research

1
Technology

2
Policy and Regulation

3
Technology
