[llvm] [AMDGPU] Fix performFMACombine FDOT2 fold ignoring denormal mode (PR #205101)

Wooseok Lee via llvm-commits llvm-commits at lists.llvm.org
Mon Aug 3 06:42:16 PDT 2026


wooseoklee wrote:

There were so many confusions about denormal behavior and I found that the confusion is from different hardware behavior and doc mismatch. Based on the current test, 

GPU	      intruction                         ieee=1	ieee=0	denorm_flush
gfx90a	v_dot2c_f32_f16	       zero	zero	zero
gfx90a	v_fma_mix_f32	normal	zero	zero
gfx942	v_dot2c/v_dot2_f32_f16	normal	normal	normal
gfx942	v_fma_mix_f32	normal	normal	zero
gfx950	v_dot2c/v_dot2_f32_f16	normal	normal	normal
gfx950	v_fma_mix_f32	normal	normal	zero
gfx1100	v_dot2c/v_dot2_f32_f16	normal	normal	normal
gfx1100	v_fma_mix_f32	normal	normal	zero
gfx1201	v_dot2_f32_f16	normal	normal	normal
gfx1201	v_fma_mix_f32	normal	normal	zero

So, gfx90a shows different behavior while the rest shows same behavior. Yet, the isa doc says 

GPU	Arch	ISA Doc statement for dot2	Hardware result	Match?
gfx90a	CDNA2	"V_DOT2 instructions do not support denormal and rounding modes. They always flush input and output denorms."	zero in all modes	✓
gfx942	CDNA3	Same statement as CDNA2 (copy-pasted)	normal in all modes	✗
gfx950	CDNA4	Same statement as CDNA2 (copy-pasted)	normal in all modes	✗
gfx1100	RDNA3	"DOT2_F16_F16 supports denorms"	normal in all modes	✓
gfx1201	RDNA4	No explicit statement	normal in all modes	N/A

So far, we can draw following logic,

gfx90a only: allow fold when f32 denorm = PreserveSign
Everything else: allow fold only when f32 denorm = IEEE

We need to confirm this behavior on CDNA1, RDNA1, and RDNA2 as well.

https://github.com/llvm/llvm-project/pull/205101


More information about the llvm-commits mailing list