Black Forest Labs unveils FLUX 3 multimodal AI to generate video, images, and robot actions

2 Sources

Share

Black Forest Labs launched FLUX 3, its first multimodal AI model capable of generating 20-second video with audio and images from single text prompts. The system uses a unified architecture trained across multiple modalities and extends to robotic vision through FLUX-mimic. While early benchmarks show strong performance against competitors, the model enters limited early access with no pricing details or open-weight release yet.

Black Forest Labs Launches FLUX 3 as First Multimodal Video Generator

Black Forest Labs has launched FLUX 3, marking a significant expansion beyond the image generation capabilities that built the company's reputation . The Freiburg, Germany-based AI lab developed this multimodal AI model to generate images and video up to 20 seconds long with synchronized audio from a single text prompt

2

. Unlike previous FLUX releases focused solely on still images, this represents the company's first public video generation model.

Source: Decrypt

Source: Decrypt

CEO Robin Rombach explained the strategic shift: "A model that only learns images can only generate images"

2

. The company trained FLUX 3 jointly across image, video, and audio modalities rather than assembling separate models behind a common interface, creating what BFL calls visual intelligence -- models "that can perceive, predict, and act across physical and digital environments"

1

.

FLUX 3 Outperforms Runway Gen-4.5 and Luma Ray 3.2 in Early Testing

In preliminary head-to-head preference testing on 10-second, 720p text-to-video clips with audio, FLUX 3 demonstrated competitive performance against established players. Human reviewers preferred FLUX 3 over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77% of evaluations

1

2

. The model also showed slight advantages over Gemini Omni and Seedance 2.0, winning 52% of those matchups.

However, BFL labels these results as "preliminary evaluation of an early FLUX 3 candidate," meaning the numbers describe a pre-release checkpoint rather than the shipping model

1

. The company has not announced pricing, production service-level commitments, evaluation methodology, sample sizes, or rater counts, preventing enterprise buyers from calculating total cost of ownership or independently reproducing the video comparisons.

Limited Early Access Program Delays Broader Availability

FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the upcoming open-source FLUX 3 Dev

1

. FLUX 3 Video, which supports 20-second video with audio generation, and FLUX 3 Action are entering a gated early access program that requires BFL approval. There is currently no public access through BFL's API or partner APIs, though FLUX 3 Image will roll out in the coming weeks, followed by general availability

2

.

Notably absent from this launch are downloadable weights or an open-source license. The open-weight Dev version isn't due until later in 2026

2

. This delay disappoints developers accustomed to receiving locally deployable FLUX variants alongside major announcements, particularly given the role open-weight models played in FLUX's initial adoption.

FLUX-mimic Brings AI-Generated Motion Predictions to Manufacturing

Beyond content creation, BFL is positioning FLUX 3 as a foundation for robotic vision and robot actions through FLUX-mimic, developed with Zurich-based mimic robotics

2

. The system takes FLUX 3's video-prediction engine and adds a lightweight decoder that translates the model's understanding of movement into actual robot motions. Audi is already testing the technology on tasks like fitting flexible door seals, work that conventional automation has struggled to handle.

"Audi represents the kind of manufacturing partner we built FLUX-mimic for," said mimic co-founder Stephan-Daniel Gravert

2

. Audi's Christoph Schneider noted the robots now "solve complex soft-body manipulation work" that older machines couldn't manage. The full system reacts in approximately 101 milliseconds, approaching human visual reflexes. This unified architecture approach suggests BFL's bet that learning to predict video also means learning the physics underneath it -- weight, contact, timing -- which machines need to navigate the physical world.

Open-Source and Multimodal AI Strategy Faces New Competitive Pressure

Founded in August 2024 by veteran researchers who helped build the original Stable Diffusion models at Stability AI, Black Forest Labs quickly established dominance in open-source image generation

2

. The open-source Flux Dev and Schnell models claimed the "best open source image generator" title that AI artists expected Stable Diffusion 3.5 to reclaim. FLUX 1.1 Pro topped the Artificial Analysis image arena in October 2024.

Yet the competitive landscape has shifted. Alibaba's Z-Image Turbo dethroned the original Flux in late 2025, matching its quality on lower-end consumer graphics cards. One CivitAI user wrote at the time: "This is what SD3 was supposed to be"

2

. The 52% preference rate against Gemini Omni presents another challenge, as Google's multimodal offering is generally available via API at $0.10 per second

1

. Without published pricing or full benchmark methodology, enterprises cannot yet compare FLUX 3's value proposition directly.

Today's Top Stories

© 2026 TheOutpost.AI All rights reserved