5 Sources
[1]
Apple releases huge AI dataset for image editing research - 9to5Mac
Apple has released Pico-Banana-400K, a 400,000-image research dataset which, interestingly, was built using Google's Gemini-2.5 models. Here are the details. Apple's research team has published an interesting study called "Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing". In
[2]
Apple's New AI Dataset Aims to Improve Photo Editing Models
Apple researchers have released Pico-Banana-400K, a comprehensive dataset of 400,000 curated images that's been specifically designed to improve how AI systems edit photos based on text prompts. The massive dataset aims to address what Apple describes as a gap in current AI image editing training.
[3]
How Apple Plans to Improve AI Image Editors
The company used Google's Nano Banana model to develop this dataset, as well as Gemini-2.5. Apple might be dead last in the AI race -- at least when you consider competition from companies like OpenAI, Google, and Meta -- but that doesn't mean the company isn't working on the tech. In fact, it
[4]
Apple's Pico-Banana-400K dataset could redefine how AI learns to edit images
The dataset was built using an automated pipeline powered by Google's Nano-Banana and Gemini-2.5-Pro, eliminating the need for human annotators. Apple has released Pico-Banana-400K, a massive, high-quality dataset of nearly 400,000 image editing examples. The new dataset, detailed in an academic
[5]
Apple Wants to Help the World Build Nano Banana-Like AI Models
Apple's dataset comes with a non-commercial research license Apple researchers have released a large-scale dataset to help others develop image editing artificial intelligence (AI) models. Dubbed Pico-Banana-400K, the dataset contains 4,00,000 real images and their AI-edited counterparts that can
Share
Copy Link
Apple has released Pico-Banana-400K, a comprehensive dataset of 400,000 curated images designed to improve AI-powered text-guided image editing models. Built using Google's Nano-Banana and Gemini models, the open-source dataset addresses critical gaps in AI training data.
Apple has released Pico-Banana-400K, a comprehensive dataset containing 400,000 curated images specifically designed to advance AI-powered image editing research
1
. The release marks a significant departure from Apple's typically closed approach to AI development, as the company makes this resource freely available to researchers worldwide under a non-commercial research license2
.
Source: 9to5Mac
The dataset addresses what Apple researchers describe as a critical gap in current AI training resources. According to their published study titled "Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing," existing datasets often rely on synthetic generations from proprietary models or limited human-curated subsets, frequently exhibiting domain shifts and inconsistent quality control
1
.In an interesting twist, Apple built this dataset using Google's AI technologies, specifically the Gemini-2.5-Flash-Image model (also known as Nano-Banana) and Gemini-2.5-Pro
3
. The researchers sourced real photographs from the OpenImages dataset, selecting images to ensure coverage of humans, objects, and textual scenes1
.The construction process involved creating a sophisticated automated pipeline that eliminated the need for human annotators, representing a cost-effective approach estimated at approximately $100,000
4
. Apple developed a comprehensive taxonomy of 35 different edit types grouped into eight categories, ranging from basic color changes to complex transformations such as converting people into Pixar-style characters or LEGO figures2
.
Source: MacRumors
The dataset's quality control system represents a significant innovation in AI training data curation. Gemini-2.5-Pro served as an automated judge, evaluating each edit based on four criteria: Instruction Compliance (40%), Seamlessness (25%), Preservation Balance (20%), and Technical Quality (15%)
4
. Edits scoring above a 0.7 threshold were labeled as successful, while failed attempts were retained as negative examples to help models learn from mistakes.The research revealed clear performance patterns across different edit types. Global edits and stylization achieved the highest success rates, with strong artistic style transfers reaching 93% success
3
. However, precise tasks requiring spatial control or symbolic understanding proved more challenging, with font style changes achieving only 58% success and object relocation managing just 59%4
.Related Stories
Pico-Banana-400K is organized into three specialized subsets designed to address different research needs. The dataset includes 258,000 single-edit examples for basic training, 56,000 preference pairs comparing successful and failed edits, and 72,000 multi-turn sequences showing how images evolve through multiple consecutive edits
2
. This structure supports various research approaches, from basic model training to advanced preference learning and multi-step editing scenarios5
.The dataset is currently available on GitHub and can be accessed by any researcher for non-commercial purposes
5
. This open approach contrasts sharply with Apple's typical product development strategy and comes at a time when the company faces challenges with its own AI initiatives, including delays to the promised Siri overhaul announced in 20245
.
Source: Gadgets 360
Summarized by
Navi
[3]
23 Sept 2025•Technology
08 Jun 2026•Technology

22 Aug 2025•Technology

1
Technology

2
Technology

3
Science and Research
