[llvm] [NVPTX] Respect FTZ flag when lowering atomicrmw fadd. (PR #200732)

via llvm-commits llvm-commits at lists.llvm.org
Wed Jun 3 04:14:16 PDT 2026


================
@@ -108,6 +108,16 @@ static cl::opt<bool> UsePrecSqrtF32(
     cl::desc("NVPTX Specific: 0 use sqrt.approx, 1 use sqrt.rn."),
     cl::init(true));
 
+// PTX atom.add.f32 has fixed FTZ behavior that may not match the function's
+// (see shouldExpandAtomicRMWInIR), so by default we fall back to a CAS loop
+// when they disagree. This flag is an escape hatch to use atom.add anyway,
+// trading correct denormal handling for the speed of the native instruction.
+static cl::opt<bool> AllowFTZAtomics(
+    "nvptx-allow-ftz-atomics", cl::Hidden,
+    cl::desc("NVPTX Specific: Lower atomicrmw fadd to atom.add even when its "
+             "FTZ behavior does not match the function's denormal mode."),
+    cl::init(false));
----------------
gonzalobg wrote:

What happens if an application:
- compiles LLVM Module A without this flag (noftz atomics),
- compiles LLVM Module B with this flag (ftz atomics),
- B calls A::fn, and
- LTO is enabled so that A::fn is inlined into B

?

An alternative is for NVPTX to correctly compile the IR that it is given, and if a frontend wants some of the atomics it generates to use FTZ, that frontend just sets FTZ for those atomics. What are the downsides of this approach?

https://github.com/llvm/llvm-project/pull/200732


More information about the llvm-commits mailing list