2 Sources
[1]
Black Forest Labs launches FLUX 3 capable of generating images and 20-second video with audio -- but in limited release to start
Black Forest Labs (BFL) is expanding its FLUX family beyond image generation with today's launch of FLUX 3, a multimodal frontier model trained to understand and generate images, or combined audio/video clips up to 20 seconds from a single prompt -- and to extend the same underlying architecture to robotic vision and actions. The Freiburg, Germany-based AI lab says FLUX 3 is jointly trained across those modalities rather than assembling separate image, video and audio models behind a common interface. That distinction is central to the company's pitch: BFL wants enterprises to think about creative generation, simulation, computer use and robotics as connected applications of a single capability it calls visual intelligence -- models, in the company's words, "that can perceive, predict, and act across physical and digital environments." This release marks BFL's first public video generation model. FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action and the upcoming, open source FLUX 3 Dev. FLUX 3 Video, with optional native audio generation, and FLUX 3 Action are entering a gated "Early Access" program now, to which anyone can apply, but which BFL must approve. There is presently no public access through BFL's application programming interface (API) or those of partners yet, but the company says FLUX 3 Image will roll out in the coming weeks, followed by general availability. The limited initial availability rollout echoes the release strategies of new models from other frontier labs in the U.S. lately, including Anthropic and OpenAI, though those were ostensibly for security concerns and due to government request. What the company has not announced is pricing, production service-level commitments, evaluation methodology, sample sizes, rater counts or any image-model benchmarks at all. Enterprise buyers therefore cannot yet calculate total cost of ownership or independently reproduce the video comparisons. Another big notable omission: FLUX 3 is not launching with downloadable weights at this time, nor an open source license. BFL says faster and open-weight versions will arrive later this year, and its technical blog names FLUX 3 Dev as "open-weight access to a multimodal backbone, for content creation (video, audio and image) and action prediction" -- a considerably broader commitment than any previous FLUX Dev release, all of which covered images only. But it arrives last in the sequence. Developers accustomed to receiving a locally deployable FLUX variant alongside -- or soon after -- a major model announcement will have to wait. That delay does not negate the company's commitment, but it is disappointing given the role open weights have played in FLUX's adoption thus far. Flux 3 is rated higher than the competition, but missing pricing and benchmarking details may prevent rapid enterprise adoption BFL has published several benchmark comparisons, but they're qualified as preliminary -- with full benchmark results and methodology to be published later during broader general availability. In early head-to-head preference testing on 10-second, 720p text-to-video clips with audio, the company says FLUX 3 was preferred over Luma Ray 3.2 in 93% of comparisons, Runway Gen-4.5 in 77%, Grok Imagine Video in 69%, Kling v3 Pro in 60%, Happy Horse v1 in 59%, Happy Horse 1.1 in 57%, and both Seedance 2.0 and Google's Gemini Omni Flash in 52%. One caveat travels with every one of those figures, and it comes from BFL itself. The chart carrying the results is labeled a "preliminary evaluation of an early FLUX 3 candidate" -- meaning the numbers describe a pre-release checkpoint rather than the model now entering early access. That cuts both ways: the shipping model may perform better, but nothing published today measures what customers will actually call. Luma Ray 3.2 and Runway Gen-4.5, where FLUX 3 posted 93% and 77%, are the softest comparisons on the list -- established products, but not the models currently setting the pace in independent video rankings. Those are real wins, and they are the ones least likely to change an enterprise shortlist. Seedance 2.0, at 52%, is a statistical coin flip against a model most Western enterprises cannot currently procure. ByteDance indefinitely postponed Seedance 2.0's international rollout after Netflix, Warner Bros., Disney, Paramount and Sony sent legal threats over alleged systematic copyright infringement, and that suspension remains in place. Tying a frozen product is neither a strong claim nor a damaging one. Gemini Omni Flash, also at 52%, matters much more. Omni is the closest large-platform analogue to what FLUX 3 is attempting -- multimodal input, video and audio-aware creation, conversational editing -- and by BFL's own measurement, the two are indistinguishable on 10-second text-to-video quality. Google's advantage in that matchup is that Omni is generally available via Google's Gemini API for $0.10 per second of generated 720p video, or a 10-second clip for around. One regional wrinkle matters for a German company's home market. Editing uploaded video is unavailable to Omni Flash users in the European Economic Area, Switzerland and the United Kingdom, though editing video the model itself generated is permitted. A European enterprise that wants to run its existing footage through a generative editing pass cannot currently do so on Omni Flash. Here's a rough guide for enterprises considering which video models to rely upon: One architecture for media generation and physical action FLUX 3 builds on Self-Flow, BFL's method for aligning multimodal understanding and generation within one architecture, publicized back in March 2026. The company says it significantly scaled up compute and data to train across video, images and audio simultaneously, and that testing showed video generation and action prediction do not require separate foundations -- the same architecture could be extended to action prediction without sacrificing what it learned from video. "We place vision at the center of our approach because it is the most signal-rich medium of the physical world. Images convey structure, images and video teach spatial relationships, video teaches dynamics, and actions reveal causal relationships. But vision alone is not the complete picture," said Robin Rombach, co-founder and CEO of BFL, in a pre-release statement provided to VentureBeat. "True intelligence means perceiving the world: predicting how it will change, taking action, and learning from the results. Joint training within one unified architecture is what will get us there, because each training modality strengthens the others. Audio conveys timing, prosody, and physical events that elude vision. Language conveys goals, abstractions, and instructions that pixels cannot easily express." He put the case more bluntly elsewhere in the announcement: "You can't cheat reality. A model that only learns images can only generate images. But the world is not made of still frames. It moves, sounds, changes, and responds." BFL says FLUX 3 targets creative tooling, media, design, e-commerce and physical AI, supporting video generation with synchronized audio, precise image editing, product and material consistency across motion, multilingual generation and robotic action prediction. It is already being tested by Canva, Burda, Magnific (formerly Freepik), Krea and Picsart. For creative software companies, the appeal is consolidation. A single foundation could potentially support storyboarding, image editing, product rendering, video variation and localization without repeatedly translating assets and instructions between disconnected models. For robotics teams, the potential value is data efficiency. Models that already encode motion, object behavior and physical change may need less task-specific robot training than systems starting from raw demonstrations. What FLUX 3 Video can actually do The video tier is the most concretely specified part of the launch, and it settles a question that had been circulating as rumor: FLUX 3 generates clips of up to 20 seconds with audio in a single generation. Every video output comes with native audio. For comparison, HappyHorse 1.0 tops out at 15 seconds of 1080p with synchronized audio -- though BFL has not stated what resolution its 20-second clips run at, and its published evaluations were conducted at 720p. Still, a 20-second long clip from a single prompt is among the longest yet achieved, matching OpenAI's discontinued Sora model. The capability list BFL published covers: * Text-to-video generation. * Image-to-video generation, either animating from a starting frame or using images as visual references. * Video-to-video generation from a reference clip, carrying elements such as a specific character into a new scene or context. * Generative video-audio continuation from existing video and audio input. * Keyframe-to-video generation for controlled transitions between defined moments. * Multilingual dialogue. * A broad range of visual styles and aspect ratios, from candid camcorder footage to animation and cinematics. * Typography generation and animated design. * Agentic chaining of individual clips into longer, multi-shot sequences. That last item is the one enterprise video teams should look at hardest. BFL claims the capabilities combine to produce sequences lasting several minutes, with visual references keeping characters consistent across scenes. If that holds up under production conditions, it addresses the constraint that has kept generative video out of most commercial pipelines: not clip quality, but continuity across shots. It is also the capability where competition is most direct. HappyHorse 1.1's headline upgrade is R2V, or Reference-to-Video, which accepts multiple character reference images to hold identity stable across generated footage -- the same problem, approached at the input layer rather than through agentic clip chaining. Alibaba also claims zero-drift lip sync and has specifically targeted the artifacts that mark commercial AI video as synthetic, including facial oiliness and over-sharpening. Character consistency is where this category is being contested, and both companies know it. BFL says FLUX 3 Video is already particularly strong at human facial expressions, associating sounds with physical events, and multilingual output. On the image side, the company says preliminary evaluations conducted during midtraining show significant improvement over earlier FLUX versions in complex prompt handling and text generation, including high-accuracy text in multiple languages. It published no image benchmarks or win rates. FLUX-mimic tests whether video models can become robot models BFL is applying its unified-architecture thesis through FLUX-mimic, a video-action model built on FLUX 3 and developed with Swiss firm Mimic Robotics, one of the first partners to receive early access. The technical blog describes two distinct routes to action prediction: integrating native action prediction directly into FLUX 3, scaling up the initial Self-Flow work; and using the pretrained video backbone as a dynamics-aware foundation from which specialized action models can be finetuned with limited task-specific data. FLUX-mimic is the second route -- the FLUX 3 backbone combined with mimic's robot-learning and production-deployment expertise in dexterous manipulation. FLUX-mimic is designed for general-purpose robotic manipulation: helping robots understand a visual scene, predict the consequences of an action, and adapt to new tasks with far less task-specific data. BFL and Mimic Robotics say that depending on task difficulty, the model can be finetuned for a specific manipulation task with as little as 30 minutes of robot data, where prior approaches have required 30 or more hours. "The hardest part of robotics is data," said Elvis Nava, CTO of Mimic Robotics, in a statement provided to VentureBeat. "Every new task normally means hours of a robot repeating itself. Because FLUX-mimic is built on top of frontier video models that already understand how the physical world behaves, it picks up a new task in minutes, not days. This way, we can leapfrog the current state of the art in robot learning." BFL argues that a model trained only on images cannot understand a world that "moves, sounds, changes, and responds," and that physical understanding is what produces convincing generated footage. Google makes a nearly identical claim for Gemini Omni. Its developer documentation cites "world knowledge" that combines "an understanding of physics" with Gemini's grasp of history, science and cultural context. Its marketing is blunter still: "Most AI models just predict the next pixel to build a narrative or an image. Gemini Omni is different," the company posted in June, crediting the model with "an intuitive understanding of forces like gravity, kinetic energy, and fluid dynamics for more realistic movements that follow real-world logic." The practical consequence for enterprise buyers is that world-model language is not a differentiator. Two of the three leading video systems now market physical understanding as their central advantage, and neither has published a benchmark that measures it. There is no standard test for whether generated water behaves like water, whether a dropped object falls at a plausible rate, or whether a sound arrives when the impact does. Human preference ratings capture some of it indirectly. Nothing else on offer captures it at all. Open weights helped make FLUX an industry standard BFL officially launched in summer 2024 and gained a name for itself in the AI industry in the intervening two years for its commitment to open sourcing high-quality AI image models beloved by developers, creatives, and enterprises. The company's founders, including Rombach, Andreas Blattmann and Patrick Esser, previously helped create VQGAN, latent diffusion and Stable Diffusion, the latter the open source technology that kicked off broad AI generation capabilities for the masses and currently used by many AI image generators and companies. That reach translated into commercial distribution. FLUX models now power generative features inside Adobe Photoshop, Picsart and Nous Research's Hermes Agent, among other platforms, and the company cites film director Martin Scorsese among professional users. Wired magazine described Black Forest Labs as a relatively small company that nevertheless became a leading competitor to Silicon Valley's largest AI labs, with FLUX models ranking near the top of image benchmarks and becoming some of the most downloaded text-to-image models on AI code sharing community Hugging Face. The company says it now runs a 100-person team across Freiburg and San Francisco. FLUX.1 Dev, FLUX.1 Kontext Dev, FLUX.1 Fill Dev and related control models, released shortly after the firm's launch, gave researchers and creative-tool developers access to downloadable checkpoints, local inference and integrations with frameworks including Hugging Face Diffusers and ComfyUI. FLUX.1 Kontext Dev, for example, was released as an open-weight model for research and noncommercial use, with generated outputs permitted for commercial purposes under the applicable license. The company continued that pattern with FLUX.2 Dev in late 2025, a 32-billion-parameter open-weight model combining generation and multi-reference editing. Black Forest Labs called it the strongest open-weight image generation and editing model available at launch and released weights, reference inference code and optimized implementations for consumer Nvidia GPUs. FLUX 3 Dev raises the stakes on that evaluation. Previous Dev releases were image models. This one is described as a multimodal backbone spanning video, audio, image and action prediction -- meaning a single license will govern whether a company can locally deploy a model that touches both content production and physical machinery. BFL hasn't yet shared information about its license, the parameter count, quantizations or hardware requirements. The company frames open weights as an enterprise feature rather than a community gesture, arguing they enable secure, low-latency local deployment for applications like robotic control systems and let teams adapt FLUX 3 to their own data, products and workflows. The financial backing behind FLUX 3 is worth noting alongside the technical claims. Black Forest Labs is valued at $3.25 billion and has raised more than $450 million from investors including a16z, AMP, Salesforce Ventures, Nvidia, General Catalyst, Adobe Ventures, Figma Ventures, Canva and Deutsche Telekom's T.Capital.
[2]
Black Forest Labs Unveils FLUX 3 AI: Ditches Stills for Video -- And Robot Hands
Only the open-weight "Dev" version is planned for later in 2026; Video and Action stay behind APIs and partner access for now, with Image following in the coming weeks. Black Forest Labs released FLUX 3 on Thursday, and for the first time, the company's flagship model generates video instead of just still images. The German AI lab, known for the FLUX line of image generators, trained the new system on images, video, and audio at once, inside one shared system. That's what is known as multimodality: one model learning several types of information together instead of separate tools bolted side by side. The video side is the headline feature. FLUX 3 produces clips up to 20 seconds long, with audio generated alongside the picture and synced to what's happening on screen -- dialogue, sound effects, ambient noise. In early evaluations, human reviewers preferred FLUX 3's output over Runway Gen-4.5 in 77% of head-to-head comparisons and over Luma Ray 3.2 in 93%. It seems to be slightly better than Gemini Omni and Seedance, beating those models in 52% of the evaluations. Of course, that's a preference test, not a fixed scoring rubric: evaluators simply watch two clips and pick the one that looks and sounds more convincing, and BFL counts how often FLUX 3 wins. Other than that, the model seems to be very competent on still images too, following its legacy. BFL shared a few images, and FLUX 3 seems to be very versatile and capable of generating a broad variety of styles beyond photorealism. BFL frames this as more than a content tool. "A model that only learns images can only generate images," said co-founder and CEO Robin Rombach. The company's bet is that learning to predict video also means learning the physics underneath it -- weight, contact, timing -- which is exactly what a machine needs to move through the physical world. That bet has a name: FLUX-mimic. Built with Zurich-based mimic robotics, it takes FLUX 3's video-prediction engine and adds a lightweight "decoder" -- a small add-on component that translates the model's internal sense of how things move into actual robot motions. Car maker Audi is already testing it on tasks like fitting flexible door seals, work that conventional automation has struggled to handle. "Audi represents the kind of manufacturing partner we built FLUX-mimic for," said mimic co-founder Stephan-Daniel Gravert. Audi's Christoph Schneider said the robots now "solve complex soft-body manipulation work" that older machines couldn't touch. BFL says the full system reacts in about 101 milliseconds, in the neighborhood of human visual reflexes. FLUX's rise didn't happen in a vacuum. Founded in August 2024 by veteran researchers who'd helped build the original Stable Diffusion models at Stability AI, Black Forest Labs launched Flux models that beat MidJourney and outclassed Stability's own underwhelming Stable Diffusion 3. The open-source Flux Dev and Schnell models grabbed the "best open source image generator" title that AI artists had expected Stable Diffusion 3.5, Stability's do-over, to eventually reclaim. It never did. FLUX 1.1 Pro went on to top the Artificial Analysis image arena that October. That one wasn't open source, though. BFL released FLUX.2 in November 2025 but it wasn't as popular. The open-source crown held by the original Flux lasted until Alibaba's Z-Image Turbo dethroned it in late 2025, matching its quality on lower end consumer graphics cards. "This is what SD3 was supposed to be," one CivitAI user wrote at the time. FLUX 3 is BFL's comeback, and it isn't fully open yet. Video and Action are in early access now through APIs and select partners, mimic robotics among them, with image generation following "in the coming weeks," per BFL. The open-weight Dev version, the only tier BFL plans to release for local use, isn't due until later in 2026.
Share
Copy Link
Black Forest Labs launched FLUX 3, its first multimodal AI model capable of generating 20-second video with audio and images from single text prompts. The system uses a unified architecture trained across multiple modalities and extends to robotic vision through FLUX-mimic. While early benchmarks show strong performance against competitors, the model enters limited early access with no pricing details or open-weight release yet.
Black Forest Labs has launched FLUX 3, marking a significant expansion beyond the image generation capabilities that built the company's reputation . The Freiburg, Germany-based AI lab developed this multimodal AI model to generate images and video up to 20 seconds long with synchronized audio from a single text prompt
2
. Unlike previous FLUX releases focused solely on still images, this represents the company's first public video generation model.
Source: Decrypt
CEO Robin Rombach explained the strategic shift: "A model that only learns images can only generate images"
2
. The company trained FLUX 3 jointly across image, video, and audio modalities rather than assembling separate models behind a common interface, creating what BFL calls visual intelligence -- models "that can perceive, predict, and act across physical and digital environments"1
.In preliminary head-to-head preference testing on 10-second, 720p text-to-video clips with audio, FLUX 3 demonstrated competitive performance against established players. Human reviewers preferred FLUX 3 over Luma Ray 3.2 in 93% of comparisons and over Runway Gen-4.5 in 77% of evaluations
1
2
. The model also showed slight advantages over Gemini Omni and Seedance 2.0, winning 52% of those matchups.However, BFL labels these results as "preliminary evaluation of an early FLUX 3 candidate," meaning the numbers describe a pre-release checkpoint rather than the shipping model
1
. The company has not announced pricing, production service-level commitments, evaluation methodology, sample sizes, or rater counts, preventing enterprise buyers from calculating total cost of ownership or independently reproducing the video comparisons.FLUX 3 will be offered through four product lines: FLUX 3 Video, FLUX 3 Image, FLUX 3 Action, and the upcoming open-source FLUX 3 Dev
1
. FLUX 3 Video, which supports 20-second video with audio generation, and FLUX 3 Action are entering a gated early access program that requires BFL approval. There is currently no public access through BFL's API or partner APIs, though FLUX 3 Image will roll out in the coming weeks, followed by general availability2
.Notably absent from this launch are downloadable weights or an open-source license. The open-weight Dev version isn't due until later in 2026
2
. This delay disappoints developers accustomed to receiving locally deployable FLUX variants alongside major announcements, particularly given the role open-weight models played in FLUX's initial adoption.Related Stories
Beyond content creation, BFL is positioning FLUX 3 as a foundation for robotic vision and robot actions through FLUX-mimic, developed with Zurich-based mimic robotics
2
. The system takes FLUX 3's video-prediction engine and adds a lightweight decoder that translates the model's understanding of movement into actual robot motions. Audi is already testing the technology on tasks like fitting flexible door seals, work that conventional automation has struggled to handle."Audi represents the kind of manufacturing partner we built FLUX-mimic for," said mimic co-founder Stephan-Daniel Gravert
2
. Audi's Christoph Schneider noted the robots now "solve complex soft-body manipulation work" that older machines couldn't manage. The full system reacts in approximately 101 milliseconds, approaching human visual reflexes. This unified architecture approach suggests BFL's bet that learning to predict video also means learning the physics underneath it -- weight, contact, timing -- which machines need to navigate the physical world.Founded in August 2024 by veteran researchers who helped build the original Stable Diffusion models at Stability AI, Black Forest Labs quickly established dominance in open-source image generation
2
. The open-source Flux Dev and Schnell models claimed the "best open source image generator" title that AI artists expected Stable Diffusion 3.5 to reclaim. FLUX 1.1 Pro topped the Artificial Analysis image arena in October 2024.Yet the competitive landscape has shifted. Alibaba's Z-Image Turbo dethroned the original Flux in late 2025, matching its quality on lower-end consumer graphics cards. One CivitAI user wrote at the time: "This is what SD3 was supposed to be"
2
. The 52% preference rate against Gemini Omni presents another challenge, as Google's multimodal offering is generally available via API at $0.10 per second1
. Without published pricing or full benchmark methodology, enterprises cannot yet compare FLUX 3's value proposition directly.Summarized by
Navi
[1]
1
Technology

2
Policy and Regulation

3
Science and Research
