[llvm] [CodeGen][AMDGPU] Allow elementwise atomic load/store at element alignment (PR #219906)

Matt Arsenault via llvm-commits llvm-commits at lists.llvm.org
Sat Sep 5 10:41:03 PDT 2026


================
@@ -2361,6 +2361,22 @@ bool SITargetLowering::allowsMisalignedMemoryAccesses(
                                             Alignment, Flags, IsFast);
 }
 
+bool SITargetLowering::isAtomicAlignmentSupported(Align Alignment,
+                                                  uint64_t SizeInBytes,
+                                                  uint64_t ElementSizeInBytes,
+                                                  unsigned AddrSpace) const {
+  // Each naturally aligned dword is separately atomic. Only global and flat
+  // opt in: the LDS 4 byte rule in allowsMisalignedMemoryAccessesImpl assumes
+  // lowering to ds_read2_b32, which never happens for an atomic.
+  if (ElementSizeInBytes < SizeInBytes && ElementSizeInBytes >= 4 &&
+      (AMDGPU::isExtendedGlobalAddrSpace(AddrSpace) ||
+       AddrSpace == AMDGPUAS::FLAT_ADDRESS))
----------------
arsenm wrote:

Flat can hit an LDS address, but I assume if the flat instruction guarantees atomicity it has to make sure it works for all addresses 

https://github.com/llvm/llvm-project/pull/219906


More information about the llvm-commits mailing list