[llvm] [AMDGPU] Fix performFMACombine FDOT2 fold ignoring denormal mode (PR #205101)
Wooseok Lee via llvm-commits
llvm-commits at lists.llvm.org
Mon Aug 3 06:42:16 PDT 2026
wooseoklee wrote:
There were so many confusions about denormal behavior and I found that the confusion is from different hardware behavior and doc mismatch. Based on the current test,
GPU intruction ieee=1 ieee=0 denorm_flush
gfx90a v_dot2c_f32_f16 zero zero zero
gfx90a v_fma_mix_f32 normal zero zero
gfx942 v_dot2c/v_dot2_f32_f16 normal normal normal
gfx942 v_fma_mix_f32 normal normal zero
gfx950 v_dot2c/v_dot2_f32_f16 normal normal normal
gfx950 v_fma_mix_f32 normal normal zero
gfx1100 v_dot2c/v_dot2_f32_f16 normal normal normal
gfx1100 v_fma_mix_f32 normal normal zero
gfx1201 v_dot2_f32_f16 normal normal normal
gfx1201 v_fma_mix_f32 normal normal zero
So, gfx90a shows different behavior while the rest shows same behavior. Yet, the isa doc says
GPU Arch ISA Doc statement for dot2 Hardware result Match?
gfx90a CDNA2 "V_DOT2 instructions do not support denormal and rounding modes. They always flush input and output denorms." zero in all modes ✓
gfx942 CDNA3 Same statement as CDNA2 (copy-pasted) normal in all modes ✗
gfx950 CDNA4 Same statement as CDNA2 (copy-pasted) normal in all modes ✗
gfx1100 RDNA3 "DOT2_F16_F16 supports denorms" normal in all modes ✓
gfx1201 RDNA4 No explicit statement normal in all modes N/A
So far, we can draw following logic,
gfx90a only: allow fold when f32 denorm = PreserveSign
Everything else: allow fold only when f32 denorm = IEEE
We need to confirm this behavior on CDNA1, RDNA1, and RDNA2 as well.
https://github.com/llvm/llvm-project/pull/205101
More information about the llvm-commits
mailing list