Hi, An important follow-up arrived shortly after my previous message.
The same separate gfxhub failure has now reproduced on the Fedora 7.1.3-200 kernel with amdgpu.dcdebugmask=0x810, after about 30.5 hours of uptime. The sequence was: [gfxhub] page fault, process Xorg, ring 24, vmid 5 GCVM_L2_PROTECTION_FAULT_STATUS: 0x00501031 Faulty UTCL2 client ID: TCP PERMISSION_FAULTS: 0x3 ring gfx_0.0.0 timeout Ring gfx_0.0.0 reset succeeded Mesa then aborted Xorg in amdgpu_ctx_set_sw_reset_status(), ending the X11 session and returning to SDDM. As in the v7.2-rc3 incident, there were no DMUB, DMCUB, flip_done, or commit-wait timeout messages. I captured a second complete 5.4 MB AMDGPU devcoredump. The fault status, client, Xorg context, Mesa abort path, and ring-reset sequence match the v7.2-rc3 incident. This corrects any impression that the machine is stable against all GPU faults with dcdebugmask=0x810. The workaround still appears to avoid the original DMCUB/page-flip timeout signature, but it does not prevent this separate gfxhub VM fault. The reproduction on both kernels also makes it unlikely that the gfxhub fault was introduced by this patch series. Both devcoredumps and the matching kernel logs are available if useful. Regards, LetsDank
