2 Sources
[1]
High-severity Nvidia bug could crash GPU monitoring on exposed servers
Researchers found thousands of GPU servers exposing Nvidia's DCGM Exporter to the internet, with hundreds potentially vulnerable to a high-severity flaw that could let unauthenticated attackers crash the GPU monitoring service and disrupt AI workloads. DCGM Exporters read telemetry from the GPUs
[2]
This NVIDIA security flaw could let attackers crash GPU monitoring services
Sanket Mungase is a freelance tech writer focused on the latest tech, Android, and Windows updates. When he is not writing, you can find him digging into interesting details about cars. We're used to hearing about GPUs overheating or power connectors melting. But this time, the threat comes from
Share
Copy Link
Researchers discovered a high-severity Nvidia bug in DCGM Exporter that could allow unauthenticated attackers to crash GPU monitoring services on exposed servers. Internet scans revealed 2,100 GPU servers exposing telemetry data from over 12,000 GPUs worth $100 million, with 44% located in the US. Nvidia has released a fix for CVE-2026-47483.
A critical Nvidia security flaw in the company's GPU monitoring service has left thousands of servers vulnerable to attacks that could crash GPU monitoring and disrupt AI workloads. Security researchers at datacenter security startup Lava discovered the high-severity Nvidia bug, tracked as CVE-2026-47483, which earned a CVSS 8.2 rating and affects the widely-used DCGM Exporter tool
1
. The vulnerability allows unauthenticated attackers to send multiple requests to debugging endpoints, forcing servers to consume excessive memory and CPU power, ultimately leading to a denial of service attack that crashes the monitoring service2
.Michael Katchinskiy, a researcher at Lava who discovered and reported the flaw, conducted internet-wide scans between March and May that uncovered approximately 2,100 GPU servers exposing DCGM Exporter metrics without authentication. These exposed GPU servers revealed telemetry data from over 12,000 GPU UUIDs, representing an estimated $100 million in hardware across roughly 300 organizations
1
. The geographic distribution showed 5,274 GPUs—44 percent of the total—located in the United States, making it the most affected region2
.DCGM Exporter functions as a health monitor for Nvidia GPUs, collecting critical telemetry data including temperature, memory usage, power consumption, performance metrics, and error events. Each GPU carries a unique identifier, and this information is exposed in plaintext over HTTP without requiring authentication
1
. The vulnerability stems from DCGM Exporter accidentally exposing Go's /debug/pprof/ built-in profiling tool alongside standard metrics. According to Katchinskiy, "About a quarter of the exposed DCGM hosts in our scan were serving Go's /debug/pprof/ profiling endpoints alongside /metrics"2
. These debugging endpoints collect runtime performance data including CPU usage, memory allocations, goroutine states, and blocking events. With enough concurrent unauthenticated requests to these pprof endpoints, attackers could force the exporter to run out of memory and crash, cutting off visibility into GPU health and activity while potentially affecting AI training or inference workloads running on the same infrastructure1
.
Source: The Register
The exposed systems included some of Nvidia's most advanced and expensive AI infrastructure. Researchers identified Nvidia Blackwell Ultra B300 GPUs, H200s, and H100 processors used for large-scale AI workloads, alongside consumer-grade RTX 5090 and 4090 systems
1
. Consumer GPUs and mining farms accounted for approximately 35 percent of the exposed systems, demonstrating the vulnerability extended beyond enterprise AI infrastructure2
. The publicly exposed monitoring services affected customer infrastructure across neocloud and GPU cloud providers including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean1
.Related Stories
Beyond enabling denial of service attacks, the exposed telemetry data provides attackers with detailed intelligence for reconnaissance operations. The plaintext HTTP exposure reveals information useful for mapping GPU infrastructure, identifying potentially vulnerable systems, and monitoring workload activity patterns
1
. Lava's investigation also examined Prometheus Node Exporter exposures, finding 12,096 public hosts revealing data on server models, operating systems, firmware versions, hostnames, storage paths, and networking hardware commonly deployed in GPU clusters. This information exposes how AI infrastructure environments are built and configured, allowing attackers to match systems to known vulnerabilities1
.Nvidia addressed the Nvidia security flaw in September by releasing DCGM Exporter version 4.8.2, and administrators should upgrade immediately to this version or later
1
. Security experts recommend that DCGM Exporter, Node Exporter, and Prometheus services should never be directly reachable from the public internet. Organizations must restrict these monitoring tools to authorized monitoring infrastructure only and keep them on private networks protected by firewalls2
. Lava reported all exposures to affected providers, and according to Katchinskiy, these providers worked with customers to address the security gaps. "The findings highlight a growing security gap in AI infrastructure: companies are spending millions on GPUs while leaving critical systems exposed," Katchinskiy emphasized, noting that "those exposures can reveal how AI environments are built and, in some cases, allow attackers to disrupt them"1
. With GPU prices soaring amid the ongoing AI boom, securing these valuable AI security assets has become more critical than ever2
.Summarized by
Navi
27 Sept 2024

05 Aug 2025•Technology

13 Jul 2025•Technology

1
Technology

2
Policy and Regulation

3
Science and Research
