[llvm] [NVPTX] Respect FTZ flag when lowering atomicrmw fadd. (PR #200732)
via llvm-commits
llvm-commits at lists.llvm.org
Wed Jun 3 04:14:16 PDT 2026
================
@@ -108,6 +108,16 @@ static cl::opt<bool> UsePrecSqrtF32(
cl::desc("NVPTX Specific: 0 use sqrt.approx, 1 use sqrt.rn."),
cl::init(true));
+// PTX atom.add.f32 has fixed FTZ behavior that may not match the function's
+// (see shouldExpandAtomicRMWInIR), so by default we fall back to a CAS loop
+// when they disagree. This flag is an escape hatch to use atom.add anyway,
+// trading correct denormal handling for the speed of the native instruction.
+static cl::opt<bool> AllowFTZAtomics(
+ "nvptx-allow-ftz-atomics", cl::Hidden,
+ cl::desc("NVPTX Specific: Lower atomicrmw fadd to atom.add even when its "
+ "FTZ behavior does not match the function's denormal mode."),
+ cl::init(false));
----------------
gonzalobg wrote:
What happens if an application:
- compiles LLVM Module A without this flag (noftz atomics),
- compiles LLVM Module B with this flag (ftz atomics),
- B calls A::fn, and
- LTO is enabled so that A::fn is inlined into B
?
An alternative is for NVPTX to correctly compile the IR that it is given, and if a frontend wants some of the atomics it generates to use FTZ, that frontend just sets FTZ for those atomics. What are the downsides of this approach?
https://github.com/llvm/llvm-project/pull/200732
More information about the llvm-commits
mailing list