3 Sources
[1]
Netflix's Void AI can remove objects from video and show how scenes evolve without them
Serving tech enthusiasts for over 25 years. TechSpot means tech analysis and advice you can trust. What just happened? Top-tier video editing suites can seamlessly remove objects from scenes, even generating realistic shadows and reflections for the freshly removed elements. However, these tools
[2]
Netflix's new AI doesn't create videos -- it rewrites reality (and it's open source)
Netflix challenges Sora with its new open source AI that transforms real footage I've spent a lot of time testing every AI video tool that hits the market, from OpenAI's Sora to the latest Runway updates. Usually, the pitch is the same: "Type a prompt, get a movie." But Netflix just quietly
[3]
Netflix's VOID AI removes objects while preserving real-world motion
The system analyzes interactions, then regenerates footage so actions still make sense. Netflix is detailing an AI video tool that goes beyond simple cleanup. Its system, called VOID, cuts elements from footage while keeping everything else behaving in a way that still feels grounded. That marks
Share
Copy Link
Netflix released VOID, an open-source AI model that removes objects from video while understanding physics and causality. Unlike traditional editing tools, VOID rewrites scenes to account for missing elements—erasing a car crash removes the debris and fire, deleting a person from a pool dive leaves the water undisturbed. The model could eliminate costly reshoots for studios.
Netflix has released an open-source AI model called VOID that fundamentally changes how objects can be removed from video footage
1
. Short for Video Object and Interaction Deletion, this advanced video editing tool doesn't just erase unwanted elements—it understands physics and causality to rewrite entire scenes as if the deleted object never existed2
. The Void AI represents a significant departure from traditional generative AI video tools like Sora and Runway, which focus on creating new content from text prompts rather than intelligently modifying existing footage.
Source: Tom's Guide
While conventional editing tools can remove objects from video, they struggle when deleted elements involve significant interactions like collisions or support
3
. VOID solves this by treating edits as chain reactions that preserve real-world motion. In demonstrations available on GitHub, the model showcases impressive capabilities: removing a person holding a guitar causes the instrument to fall naturally to the ground, while erasing one car from a head-on collision eliminates the resulting fire, debris, and damage as if the accident never occurred1
. The system analyzes cause and effect relationships, then performs physics-aware sequence reconstruction to maintain believable behavior throughout the edited footage.Source: TechSpot
To achieve this level of sophistication, Netflix trained the model using Kubric and Humoto to generate thousands of paired datasets showing counterfactual object removals
1
. During inference, a vision-language model identifies parts of the scene impacted by the removed object, which then guides a diffusion model to fill gaps with counterfactual data. This approach allows VOID to apply learned rules about physical interactions rather than simply copying patterns from existing footage3
. The model uses a 5-billion parameter version of CogVideoX and employs a proprietary "quadmask" system to determine which aspects of the physics need recalculation2
.Related Stories
Netflix made the open-source AI model available on Hugging Face under an Apache 2.0 license, allowing anyone to access this technology
2
. However, running VOID requires substantial computing power—at least 40GB of VRAM using GPUs like NVIDIA A100 or H100. In a survey of 25 individuals, VOID was preferred over competing tools like ProPainter, Rose, DiffuEraser, and Generative Omnimatte nearly 65 percent of the time1
. For studios, this represents massive cost-saving potential by eliminating expensive reshoots. The infamous "Game of Thrones" Starbucks cup incident, which required frame-by-frame digital surgery, could now be fixed seamlessly in post-production2
.While VOID remains a research system detailed in a 19-page arXiv paper rather than a commercial product, its capabilities raise important questions about video authenticity
3
. The ability to remove objects from video while maintaining perfect physical consistency means visual evidence may no longer serve as reliable proof. As this technology scales to handle more complex scenarios with denser setups and longer sequences, the line between captured reality and edited footage becomes increasingly blurred. Studios should watch for integration into professional workflows, while audiences may need to reconsider how they evaluate video authenticity in what experts are calling the era of editable reality2
.Summarized by
Navi
[3]
1
Technology

2
Technology

3
Policy and Regulation
