Nvidia GPU Monitoring Flaw Exposes AI Infrastructure to Crash Attacks
A high-severity vulnerability in Nvidia's DCGM Exporter could allow unauthenticated attackers to crash GPU monitoring services, disrupting AI workloads.

Researchers have uncovered a critical vulnerability in Nvidia's Data Center GPU Manager (DCGM) Exporter, a tool used for monitoring GPU performance and health. The flaw, tracked as CVE-2026-47483 and rated with a CVSS score of 8.2, could permit unauthenticated attackers to crash the monitoring service, thereby disrupting vital AI training and inference workloads.
The DCGM Exporter collects detailed telemetry data from GPUs, including hardware specifications, utilization metrics, memory usage, power consumption, and error events. This information is exposed over HTTP without requiring any authentication, providing attackers with valuable reconnaissance data. This includes mapping out GPU infrastructure, identifying potentially vulnerable systems, and monitoring the activity of AI workloads.
Michael Katchinskiy, a researcher at Lava, discovered and reported the vulnerability. His investigation revealed that thousands of GPU servers worldwide were exposing the DCGM Exporter to the public internet. Between March and May, scans identified approximately 2,100 GPU servers with exposed DCGM Exporter metrics, encompassing over 12,000 unique GPU identifiers (UUIDs). These exposed systems belonged to around 300 organizations, with nearly half located in the United States, representing an estimated $100 million in hardware.
The affected hardware included high-end GPUs crucial for large-scale AI operations, such as Nvidia Blackwell Ultra B300, H200, and H100 models, as well as consumer-grade RTX 5090 and 4090 cards. The Lava team also observed that about 25 percent of these exposed DCGM hosts also revealed data from Go's built-in /debug/pprof/ profiling tool, which can expose runtime performance data like CPU and memory usage.
Exploitation of CVE-2026-47483 can occur through a sufficient volume of concurrent unauthenticated requests to the exporter. This can overwhelm the exporter, causing it to run out of memory and crash. Such a crash would not only cut off visibility into GPU health and activity but could also directly impact ongoing AI training or inference tasks running on the affected servers.
Nvidia has addressed this vulnerability by releasing a fix in version 4.8.2 of the DCGM Exporter. Operators are strongly advised to upgrade to this version or a later release to mitigate the risk. In addition to the DCGM Exporter, the Lava researchers also found widespread exposure of Prometheus Node Exporter services, which monitor server hardware and operating systems, further highlighting potential reconnaissance opportunities for attackers.
The security implications extend to cloud providers and their customers. Lava reported these exposures to affected providers, including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean, who have reportedly worked with their customers to rectify the issues. The findings underscore a significant security gap in AI infrastructure, where substantial investments in hardware are sometimes coupled with inadequate security practices for monitoring and management services.
Lava recommends that DCGM Exporter, Node Exporter, and Prometheus services should not be directly accessible from the public internet. Instead, access should be restricted to authorized internal monitoring infrastructure. This proactive measure is crucial for protecting sensitive AI environments from potential disruption and reconnaissance by malicious actors.