Skip to content

Report real GPU memory on NVIDIA unified memory platforms - #511

Merged
Syllo merged 2 commits into
Syllo:masterfrom
sudoingX:fix/uma-gpu-memory-reporting
Sep 27, 2026
Merged

Syllo merged 2 commits into
Syllo:masterfrom
sudoingX:fix/uma-gpu-memory-reporting

Conversation

@sudoingX

Copy link
Copy Markdown

The problem

On UMA platforms such as DGX Spark, NVML returns NVML_ERROR_NOT_SUPPORTED for the framebuffer query, so nvtop falls back to the Linux system memory counters. The current fallback in set_unified_system_memory_info() reports:

total = MemTotal
used  = MemTotal - MemAvailable

Those describe the host, not the GPU. Measured on a DGX Spark (GB10, 121.690 GiB unified):

situation NVML says the GPU holds nvtop reports
idle GPU, 30 GiB of host allocations nothing resident 34.375 GiB used
vLLM worker serving 89.276 GiB 98.695 GiB used

In the first case the GPU has nothing on it at all. In the second, nvtop overstates by more than 9 GiB of unrelated host consumption. The total also never moves off MemTotal, so the headroom figure ignores everything else on the machine.

This is the regression described in #449.

Reproducing without a DGX Spark

This needs no GPU workload. On any machine nvtop treats as unified memory, allocate and touch host memory:

b = bytearray(30 * 1024**3)
for i in range(0, len(b), 4096):
    b[i] = 1

The reported GPU used memory climbs with host consumption while the GPU stays idle.

The fix

NVML does report per-process allocations on these platforms even though it does not report a framebuffer size. Sum the compute and graphics allocations for the device and use that as the used memory, then report total as that sum plus MemAvailable, which is the memory the GPU can actually reach. MemAvailable alone becomes the free memory.

The change is confined to the existing if (has_unified_memory) branch. The dedicated framebuffer path is untouched.

Verification

Same Spark, model resident, reading NVML immediately before and after each nvtop sample:

NVML nvtop
89.789 GiB 89.790 GiB
89.789 GiB 89.790 GiB
89.789 GiB 89.790 GiB

Also checked:

  • builds clean with all backends enabled, no new warnings
  • dedicated VRAM reporting verified separately on an RTX 5090
  • process list, temperature, power and utilization panels unaffected
  • RSS flat across 180 refresh cycles, the helper frees on both the success and failure paths

One note for review

The helper enumerates processes a second time per refresh. nvtop's order is refresh_dynamic_info, then get_running_processes, then fix_dynamic_info_from_process_info, and this code runs in the first phase, before the process list exists.

A tighter version would set only free_memory here and let fix_dynamic_info_from_process_info compute used from the already collected processes. That needs the shared needGPUMemory condition in extract_gpuinfo.c to stop requiring a valid total_memory first, which affects every vendor, so I kept this patch self contained. Happy to do it that way instead if you prefer.

Fixes #449

The unified-memory path summed the running processes' GPU allocations
inline and reported total = used + MemAvailable. That duplicated the
generic summation already performed in
gpuinfo_fix_dynamic_info_from_process_info() (and did it worse: it missed
MPS processes and had no guard against NVML_VALUE_NOT_AVAILABLE), while
also redefining total_memory as a moving quantity rather than the system
capacity.

Report total = MemTotal (same as every other backend) and free =
MemAvailable, and let the generic pass sum the process allocations into
used_memory. The remaining used + free != total residual is host
(non-GPU) residency, which the old code wrongly reported as GPU memory.

Record the existing unified-memory detection in
static_info.memory_shared_with_host so the generic pass can tell such
devices apart. For those it initializes used to 0 so an idle device
reports 0 rather than unknown, and derives mem_util_rate for the graph
from the reconstructed sum. Devices that do not set the flag are
unaffected, including those that leave used unset for other reasons
(e.g. Apple without per-process accounting).

This removes sum_process_gpu_memory() entirely.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

DGX Spark support worked in 3.3.0, has been broken since 3.3.1

2 participants