[llvm] [NVPTX] Honor !atomic.ignore.denormal.mode on atomicrmw fadd (PR #217586)

Christian Sigg via llvm-commits llvm-commits at lists.llvm.org
Wed Sep 16 00:10:59 PDT 2026


chsigg wrote:

Thanks @akshayrdeodhar! I checked out #223569 and having `!atomic.ignore.denormal.mode` behave like your `fast` option when compiling with `-nvptx-ftz-atomics=strict` makes total sense.

Since `shouldExpandAtomicRMWInIR` only decides whether to emit a CAS loop, and ISel in your patch already selects `.noftz` when available under `strict`, combining them in `shouldExpandAtomicRMWInIR` is just:
```cpp
const bool IgnoreFTZMismatch =
    FTZAtomics != AtomicAddBehavior::Strict ||
    AI->hasMetadata(LLVMContext::MD_atomic_ignore_denormal_mode);
```
When the metadata is present, we skip the CAS loop and ISel naturally emits `atom.add.noftz.f32` if supported in IEEE mode, or falls back to `atom.add`.

The per-instruction metadata is still needed even with PTX 9.4 because PTX doesn't have `.ftz` atomic adds for shared/generic `f32` or for `f16`. So under `strict` + FTZ mode (`-fcuda-flush-denormals-to-zero`), CUDA's `atomicAdd()` builtins (which take generic pointers) would still get expanded to CAS loops without this metadata.

I updated the test here for `.param::func` and noted the PTX 9.4 `.noftz` behavior in the description. Either PR can land first and the other should be a trivial rebase.

https://github.com/llvm/llvm-project/pull/217586


More information about the llvm-commits mailing list