Skip to content

fix: fall back to hwmon junction temperature on SMU13.0.6 - #513

Merged
Syllo merged 3 commits into
Syllo:masterfrom
jason34105533:SMU13.0.6-family
Sep 27, 2026
Merged

Syllo merged 3 commits into
Syllo:masterfrom
jason34105533:SMU13.0.6-family

Conversation

@jason34105533

@jason34105533 jason34105533 commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Summary

MI300A(SMU13.0.6) only implements AMDGPU_PP_SENSOR_HOTSPOT_TEMP and AMDGPU_PP_SENSOR_MEM_TEMP temperature queries. It doesn't provide AMDGPU_PP_SENSOR_GPU_TEMP orAMDGPU_PP_SENSOR_EDGE_TEMP ioctl. So, originally nvtop would return N/A on GPU temp.

Change

When the existing GPU_TEMP ioctl is unavailable, this change reads hwmon sensors in this order: edge, then junction, then the legacy unlabeled temp1_input fallback. On MI300A this uses temp2_input (junction), consistent with rocm-smi.

Not really confirmed that whether this issue is only on APU or also occurs on other SMU13.0.6 device, I think the later is more likely.

Validation

  • Built successfully with CMake.
  • Confirmed rocm-smi reports MI300A junction temperature.

SMU13.0.6 only implements AMDGPU_PP_SENSOR_HOTSPOT_TEMP and AMDGPU_PP_SENSOR_MEM_TEMP temperature queries. It doesn't provide AMDGPU_PP_SENSOR_GPU_TEMP or AMDGPU_PP_SENSOR_EDGE_TEMP ioctl. So, originally nvtop would return N/A on GPU temp.

rocm-smi reports the hwmon junction temperature, so use it as the fallback when the existing GPU_TEMP ioctl is unavailable.

When the GPU_TEMP ioctl fails, look for hwmon edge first, then junction. This selects temp2_input on MI300A while preserving the usual edge reading elsewhere. Retain temp1_input as a fallback for older hwmon devices without sensor labels.
# Conflicts:
#	src/extract_gpuinfo_amdgpu.c
Read the hwmon GPU temperature from an open FILE* that is rewound and
re-read on every refresh instead of calling nvtop_device_get_sysattr_value().

sd_device_get_sysattr_value() caches sysattr values inside the sd_device
object for its lifetime (this is why the Intel driver builds throwaway
"noncached" devices for its dynamic readings), so going through the
long-lived hwmonDevice returned the first sample on every subsequent
refresh and the reported GPU temperature froze after the first poll. An
open fd is not affected by that cache.

The sensor is still resolved once during init (prefer "edge", then
"junction", then the unlabeled temp1_input), so the per-refresh label
scan is removed as well. This matches how the fan speed and power cap
files are already handled.
@Syllo
Syllo merged commit ed4a572 into Syllo:master Sep 27, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants