[llvm] [NVPTX] Respect FTZ flag when lowering atomicrmw fadd. (PR #200732)
Justin Lebar via llvm-commits
llvm-commits at lists.llvm.org
Wed Jun 3 07:52:24 PDT 2026
================
@@ -7466,22 +7476,54 @@ NVPTXTargetLowering::AtomicExpansionKind
NVPTXTargetLowering::shouldExpandAtomicRMWInIR(const AtomicRMWInst *AI) const {
Type *Ty = AI->getValOperand()->getType();
- if (AI->isFloatingPointOperation()) {
- if (AI->getOperation() == AtomicRMWInst::BinOp::FAdd) {
- if (Ty->isHalfTy() && STI.getSmVersion() >= 70 &&
- STI.getPTXVersion() >= 63)
- return AtomicExpansionKind::None;
- if (Ty->isBFloatTy() && STI.getSmVersion() >= 90 &&
- STI.getPTXVersion() >= 78)
- return AtomicExpansionKind::None;
- if (Ty->isFloatTy())
- return AtomicExpansionKind::None;
- if (Ty->isDoubleTy() && STI.hasAtomAddF64())
+ // Try to lower LLVM atomicrmw fadd to PTX atomic.add. This is complicated
+ // by the weird FTZ behavior PTX atom.add has:
+ // - atom.add.f32 on global memory flushes denormals
+ // - atom.add.f32 on shared memory does not flush denormals
+ // - atom.add.f16 and atomic.add.bf16 never flush denormals
+ //
+ // We lower to atom.add only if the function's FTZ behavior matches that of
+ // atom.add; otherwise, we lower to a CAS loop. But we always allow
+ // atomic.add.bf16; even though it never flushes denormals, we never flush
+ // bf16 denormals when doing regular arithmetic, even when FTZ is enabled.
+ if (AI->isFloatingPointOperation() &&
+ AI->getOperation() == AtomicRMWInst::BinOp::FAdd) {
+ const bool FTZ =
+ AI->getFunction()->getDenormalMode(APFloat::IEEEsingle()).Output ==
----------------
jlebar wrote:
I'm preserving the old behavior here, but I agree in principle we could check both and emit a lowering error (probably not here, but somewhere) if the denormal mode is anything other than `ieee|ieee` or `preservesign|preservesign`.
https://github.com/llvm/llvm-project/pull/200732
More information about the llvm-commits
mailing list