[llvm] [NVPTX] Add atom.add.noftz support, avoid CAS loops when noftz is available (PR #223569)
Akshay Deodhar via llvm-commits
llvm-commits at lists.llvm.org
Mon Sep 21 18:04:05 PDT 2026
================
@@ -3299,7 +3301,13 @@ defm INT_PTX_ATOM_ADD_32 : F_ATOMIC_2<I32RT, atomic_load_add_i32, "add", "u32",
defm INT_PTX_ATOM_ADD_64 : F_ATOMIC_2<I64RT, atomic_load_add_i64, "add", "u64", atomic_load_add>;
defm INT_PTX_ATOM_ADD_F16 : F_ATOMIC_2<F16RT, atomic_load_fadd, "add.noftz", "f16", atomic_load_fadd, [SM70, PTX63]>;
-defm INT_PTX_ATOM_ADD_BF16 : F_ATOMIC_2<BF16RT, atomic_load_fadd, "add.noftz", "bf16", atomic_load_fadd, [SM90]>;
+defm INT_PTX_ATOM_ADD_BF16 : F_ATOMIC_2<BF16RT, atomic_load_fadd, "add.noftz", "bf16", atomic_load_fadd, [SM90, PTX78]>;
+
+// If the function denormal mode is IEEE, and noftz is supported, lower to NoFTZ
+defm INT_PTX_ATOM_ADD_NOFTZ_F32 : F_ATOMIC_2<F32RT, atomic_load_fadd, "add.noftz", "f32", atomic_load_fadd, [hasAtomAddNoFTZ, doNoF32FTZ, noLegacyF32AtomAdd]>;
+// Otherwise, lower to the default native instruction. Passes prior to ISel
----------------
akshayrdeodhar wrote:
Nope, I mean ISel, because this pattern kicks in at ISel. AtomicExpand *is* that pass prior to ISel which will produce a CAS loop if there is a mismatch.
https://github.com/llvm/llvm-project/pull/223569
More information about the llvm-commits
mailing list