[llvm] [IR] Add an "atomicity" operand bundle (PR #223298)
Krzysztof Drewniak via llvm-commits
llvm-commits at lists.llvm.org
Tue Sep 15 10:10:16 PDT 2026
krzysz00 wrote:
To write down my version of the context, since I was one of the reviewers on the changes that led to this
Because of their fiddly addressing modes, operations on buffer resources (`V#`) on AMDGPU (and, even more aggressively, things like images or samplers - which aren't even pointers, but that's out of scope here) are represented with intrinsic calls. There is a programmer convenience (being ported over from graphics compilers) where you can use a normal-looking pointer (`ptr addrspace(7)`) to encode simple usage of these buffer resources.
Whether or not we're going via ptr(7), these buffers support atomic operations - `llvm.amdgcn.raw.ptr.buffer.atomic.fadd.T` for example. (The `.struct.ptr.` variants are basically the same intrinsic but you can do two-level indexing / swizzling, and can't be represented with ptr(7)).
Currently:
1. The atomic-ness of these intrinsics isn't captured by LLVM generally (which hasn't caused too many issues, since LLVM doesn't do much optimization-wise to random intrinsics)
2. These intrinsics have no way to record the atomicity parameters (scope and strength) - which means we also don't have a good way to distinguish atomic and non-atomic load, and also
3. This means that we're having to expand atomic operations in ptr addrspace(7) to *fences* instead of being more precise about it
4. The PR that spawned this was showing how that fence expansion wasn't working correctly (there were some cache writeback instructions missing on atomic operations that should have had them)
After much iteration, the solution I workshopped with @aobolensk (and @arsenm was there) was that we should attach scope/workgroup information to the intrinsics so we can persist it into the backend.
My initial solution was to just stick metadata arguments on, but @arsenm was opposed, so we settled on an operand bundle carrying the atomicity parameters.
Further review led to the realization that there wasn't actually anything AMD-specific about "hey, this is an atomic intrinsic, here's its atomicity parameters" and me having the thought that we might not be the only people wanting to record this ... *and* @aobolensk noticing that, if we had this general `atomicity` bundle, that pessimistic fence insertion on things like ptr addrspace(7) lowering would no longer be needed because LLVM would see `call float @llvm.amdgcn.raw.ptr.buffer.load.f32(ptr addrspace(8) %rsrc, i32 %offset, i32 0, i32 0) "atomicity" {!"seq_cst", !"workgroup"}` as atomic, giving us a clean lowering from `load seq_cst float, ptr addrspace(7) %p, syncscope("workgroup"`, where `%p = {%rsrc, %offset}`, along with letting users call the atomic intrinsic directly.
Does this bit of procedural history help @nikic?
(And if this isn't restricted to intrinsics, maybe it should be, now that I think about it)
https://github.com/llvm/llvm-project/pull/223298
More information about the llvm-commits
mailing list