[llvm] [AMDGPU] Fix isCanonicalized to require canonical input for NaN-propa… (PR #208733)

Wooseok Lee via llvm-commits llvm-commits at lists.llvm.org
Fri Jul 24 10:46:06 PDT 2026


wooseoklee wrote:

After vibe coding and experimenting on different machines, it turns out following. After discussing with Brendon, I will probably drop this PR.

Test: v_rcp_f32, v_rsq_f32, v_sqrt_f32, v_exp_f32, v_log_f32, v_frexp_mant_f32
Input: sNaN with payloads 0x00001, 0xfffff, 0x00001 (negative), 0x00042
Modes tested: ieee=1 (default compute), ieee=0 (DX10/shader)

GPU               | Architecture   | TRANS ops result   | v_frexp_mant result  | ieee=0 vs ieee=1
gfx90a (MI210)    | CDNA2        | qNaN + payload   | qNaN + payload      | No difference
gfx950  (MI350)   | CDNA4        | qNaN + payload   | qNaN + payload      | No difference
gfx1100 (RX7900) | RDNA3        | qNaN + payload   | qNaN + payload      | No difference
gfx1201 (RX9070) | RDNA4        | qNaN + payload   | qNaN + payload      | No difference

All results: quiet bit set, payload bits preserved, sign bit preserved.

Conclusions:
1. All TRANS ops (v_rcp, v_rsq, v_sqrt, v_exp, v_log) unconditionally quiet
   sNaN across all tested generations (CDNA2, CDNA4, RDNA3, RDNA4).
2. v_frexp_mant_f32 also quiets sNaN on all hardware, contradicting the
   CDNA2 ISA pseudocode (D.f = S0.f for NaN).
3. IEEE mode bit has no observable effect on sNaN quieting for any of
   these ops on any tested hardware.
4. isCanonicalized returning true for all these ops was correct in
   practice across all tested hardware.

https://github.com/llvm/llvm-project/pull/208733


More information about the llvm-commits mailing list