Documentation

This commit is contained in:
2026-09-30 01:21:46 -07:00
parent fdf90ef790
commit 812e0f89f3
+1 -1
View File
@@ -36,7 +36,7 @@ Access to the `/dev/i2c` device files, which means either:
Add the make flag `USE_NVML=1` and the it will also display the main GPU temperature ("GPU1") as reported by the NVIDIA driver. It will also display the performance cap/clock reason and memory controller utilization. This requires the NVIDIA management library (NVML) to be installed.
### VRAM and Hotspot temperature
Run as root and you can also read the VRAM and "hotspot" temperatures. These require access the BAR0 through the `/sys/bus/pci/devices/*/resource0` device files. These sensors are **extremely** undocumented so I can't say anything about their accuracy. If you have a Blackwell (50 series) card you can read all 4 die temp sensors, a 'system' temperature, and every individual VRAM chip temp (add your card in blackwell-temps.c).
Run as root and you can also read the VRAM and "hotspot" temperatures. These require access the BAR0 through the `/sys/bus/pci/devices/*/resource0` device files. These sensors are **extremely** undocumented so I can't say anything about their accuracy. If you have a Blackwell (50 series) card you can read all 4 die temp sensors, a 'system' temperature, and every individual VRAM chip temp (add your card in blackwell-temps.c). You will probably require the kernel parameter `iomem=relaxed`.
### Datacenter GPU Manager metrics
Add the make flag `USE_DCGM=1` and it can display real-time metrics from supported Nvidia GPUs like SM occupancy, integer/fp pipe usage, tensor usage, and DRAM bandwidth utilization. This requires a Quadro/Workstation or a Tesla/Datacenter card, it does not work on GeForce cards.