3 Sources
[1]
Nvidia RTX 5090 reset bug prompts $1,000 reward for a fix -- cards become completely unresponsive and require a reboot after virtualization reset bug, also impacts RTX PRO 6000
CloudRift and community reports suggest a reset failure on Nvidia's new Blackwell GPUs that bricks the card until the machine is power-cycled. Nvidia's new RTX 5090 and RTX PRO 6000 GPUs are reportedly being plagued by a reproducible virtualization reset bug that can leave the cards completely
[2]
RTX 5090 and RTX PRO 6000 GPU have a new bug: need a full system reboot after virtualization
TL;DR: NVIDIA's GeForce RTX 5090 and RTX PRO 6000 GPUs face a critical virtualization bug causing system crashes and unresponsiveness after days of VM use, requiring full reboots. NVIDIA acknowledges the issue, affecting AI workloads, and is actively developing a fix while offering a $1000 bug
[3]
NVIDIA's High-End GeForce RTX 5090 & RTX PRO 6000 GPUs Reportedly Bricked by Virtualization Bug, Requiring Full System Reboot to Recover
It seems like NVIDIA's flagship GPUs, the GeForce RTX 5090 and the RTX PRO 6000, have encountered a new bug that involves unresponsiveness under virtualization. NVIDIA's Flagship Blackwell GPUs Are Becoming 'Unresponsive' After Extensive VM Usage CloudRift, a GPU cloud for developers, was the
Share
Copy Link
Nvidia's RTX 5090 and RTX PRO 6000 GPUs are experiencing a severe virtualization reset bug that renders the cards unresponsive, requiring system reboots. The issue affects AI workloads and virtualization setups, with Nvidia working on a fix.
Nvidia's flagship GPUs, the RTX 5090 and RTX PRO 6000, are reportedly plagued by a critical virtualization reset bug that renders the cards unresponsive and requires a full system reboot to recover
1
2
. This issue is particularly concerning for AI workloads and virtualization setups, prompting a $1,000 bug bounty for anyone who can identify a fix or root cause.
Source: Tom's Hardware
The problem occurs after a GPU has been passed through to a virtual machine (VM) using KVM and VFIO. When the guest is shut down or the GPU is reassigned, the host issues a PCIe function-level reset (FLR), which is a standard cleanup procedure for passthrough devices
1
. However, instead of returning to a known-good state, the GPU fails to respond, becoming unreadable to system tools and requiring a complete power cycle of the machine to restore normal operation.CloudRift, a GPU cloud provider, has reported encountering this issue on multiple Blackwell-equipped systems in production
1
2
. The bug is particularly problematic for multi-tenant AI workloads and home lab setups using virtualization, as a single card failure can take down the entire host system. Users across various forums, including Proxmox and Level1Techs, have reported similar behaviors, with some experiencing complete host hangs after shutting down Windows guests1
3
.
Source: Wccftech
Related Stories
Nvidia has reportedly acknowledged the issue and is actively working on developing a fix
2
3
. In the meantime, CloudRift has issued a $1,000 public bug bounty for anyone able to identify a fix or root cause of the problem1
2
. The community is actively engaged in finding solutions, with users experimenting with various BIOS settings and configurations to mitigate the issue.This bug highlights the challenges faced by cutting-edge hardware in complex virtualization environments. As GPUs become increasingly crucial for AI and machine learning workloads, such issues can have significant implications for businesses relying on these technologies. The incident also underscores the importance of thorough testing and rapid response to hardware-level bugs in the fast-paced world of GPU development and deployment.
Summarized by
Navi
[2]
19 Mar 2025•Technology

05 Mar 2025•Technology

08 Sept 2025•Technology

1
Technology

2
Policy and Regulation

3
Technology
