[llvm] [NVPTX] Set default value of nvptx-allow-ftz-atomics to true (PR #206154)

Justin Lebar via llvm-commits llvm-commits at lists.llvm.org
Fri Jun 26 12:49:15 PDT 2026


jlebar wrote:

To add some clarity that I don't think is obvious from the PR description:

The issue here is with code that does something like this:

```
addr = ...;  // all threads have same address
atomic_add(addr);
return *addr;
```

If atomic_add is issued in a convergent manner (i.e. all threads run it at the same time), then all threads will read the same value for *addr.  If we change atomic_add into a CAS loop, then different threads may read different values for *addr.

Any code which relies on getting the same value for *addr has a race condition.  Moreover because atomics are not convergent operations in LLVM, the compiler is free to expose this race even in the case where the atomic is *not* lowered to a CAS loop.  For example the compiler is free to do the following transformation

```
fn orig() {
  if (foo) { ... }
  else { ... }
  atomic_add(addr);
  return *addr;
}

fn transformed() {
  if (foo) { ...; atomic_add(addr); }
  else { ...; atomic_add(addr); }
  return *addr;
}
```

This would expose the same problem in the user code.

If we want not to expose this problem in user code, we'd need to mark atomic_add as convergent.  If we don't mark atomic_add as convergent, and someone in LLVM checks in a change that makes us more likely to make a transformation like above, are we going to revert that change too?  If not, why not?  What is the principle here?

Personally I don't think that we should disable a correctness fix (or any other fix) in order to preserve the behavior of programs that have race conditions.  You have a clearly buggy program, the compiler change exposes the bug, that is not our problem.

Disabling this correctness fix for *performance* reasons is a different story and I think that's kind of what https://github.com/llvm/llvm-project/pull/206154#issuecomment-4812649206 is getting at.  (The theory being, the correctness is not *that* big and in many cases not even noticeable.  I'm not 100% sure that's true, but I think it *could* be correct.)

Art says that he wants to flip it back once there's a PTX version that supports native ftz and non-ftz PTX atomic ops. Apparently this is coming soon.  But I am skeptical that will actually solve his problem, for two reasons.

1. You can only flip the default once you drop support for the old PTX versions.  (Unless you're willing to have different values for this flag for old vs new ptx, but that seems *extremely* weird.)  So the flip will actually take a while.

2. The idea that we can flip this once we have the new PTX atomics seems to be assuming that the PTX atomics will not introduce thread divergence.  I have not heard someone guarantee that.  (Indeed, after Volta, *most* PTX instructions are allowed to introduce divergence.)  Many PTX atomics already get lowered into SASS CAS loops; my expectation is that these new-to-PTX ops will also be SASS CAS loops and not native hardware instructions.  So either (a) the SASS CAS loops will have the same problem as the LLVM CAS loops, or (b) the SASS CAS loops will explicitly re-sync the thread after the atomic, in which case we can get the same behavior in LLVM by emitting code to do this too, if that's the behavior we want.  (But we'd eat a performance penalty.)

Having said all this, I don't object to this flag flip.  It just doesn't make a lot of sense to me.

https://github.com/llvm/llvm-project/pull/206154


More information about the llvm-commits mailing list