2 Sources
[1]
World's largest open-source multimodal dataset delivers 17x training efficiency, unlocking enterprise AI that connects documents, audio and video
AI models are only as good as the data they're trained on. That data generally needs to be labeled, curated and organized before models can learn from it in an effective way. One of the big missing links in the AI ecosystem has been the availability of a large high-quality open-source multimodal
[2]
Encord creates a new method for training powerful multimodal AI models on a single GPU - SiliconANGLE
Encord creates a new method for training powerful multimodal AI models on a single GPU Artificial intelligence data annotation startup Encord, officially known as Cord Technologies Inc., wants to break down barriers to training multimodal AI models. To do that, it has just released what it says
Share
Copy Link
Encord introduces EMM-1, the largest open-source multimodal dataset, and EBind, a novel training methodology. This breakthrough enables efficient training of powerful multimodal AI models on a single GPU, potentially democratizing access to advanced AI technologies.
In a significant leap forward for the AI industry, data labeling platform vendor Encord has introduced EMM-1, the world's largest open-source multimodal dataset, alongside a novel training methodology called EBind. This development promises to democratize access to multimodal AI and revolutionize the way AI models are trained and deployed
1
2
.The EMM-1 dataset comprises an impressive 1 billion data pairs and 100M data groups across five modalities: text, image, video, audio, and 3D point clouds. This dataset is a staggering 100 times larger than the next comparable multimodal dataset, operating at petabyte scale with terabytes of raw data and over 1 million human annotations
1
.
Source: VentureBeat
Encord's EBind methodology, which prioritizes data quality over raw computational power, has achieved remarkable results. A compact 1.8 billion parameter model trained using EBind matched the performance of models up to 17 times larger, while dramatically reducing training time from days to hours on a single GPU
1
2
.Encord's success is not just about scale, but also about addressing critical issues in AI training. The company focused on solving the problem of data leakage between training and evaluation sets, which can artificially inflate model performance metrics. By employing hierarchical clustering techniques, Encord ensured clean separation while maintaining representative distribution across data types
1
.EBind builds upon OpenAI's CLIP (Contrastive Language-Image Pre-training) approach, extending it from two modalities to five. This architectural choice prioritizes parameter efficiency by using a single base model with one encoder per modality, instead of deploying separate specialized models for each modality pair
1
.The introduction of EMM-1 and EBind has significant implications for enterprise AI applications. Multimodal models enable use cases that span different data types, allowing organizations to search and retrieve across various systems simultaneously, including content management platforms, communication tools, learning management systems, and databases
1
.Related Stories
Encord's innovations aim to break down barriers to training multimodal AI models, making them accessible to developers and companies of all sizes. By reducing the time and computational resources required for training, Encord is leveling the playing field, allowing smaller startups to compete with tech giants in the AI space
2
.Early access to the dataset and methodology has garnered positive reactions from industry professionals. Charlotte Bax, CEO of British vision AI startup Captur Ltd., praised the dataset's potential for improving image quality measures and handling edge cases in on-device models
2
.Encord's President, Ulrik Stig Hansen, predicts that future AI innovation will be driven more by data quality than by raw computing power. This shift in focus could reshape the competitive landscape in the AI industry, favoring organizations that excel in data curation and dataset construction
2
.Summarized by
Navi
1
Technology

2
Policy and Regulation

3
Technology
