[llvm] [CodeGen][AMDGPU] Allow elementwise atomic load/store at element alignment (PR #219906)
Matt Arsenault via llvm-commits
llvm-commits at lists.llvm.org
Sat Sep 5 10:41:03 PDT 2026
================
@@ -2361,6 +2361,22 @@ bool SITargetLowering::allowsMisalignedMemoryAccesses(
Alignment, Flags, IsFast);
}
+bool SITargetLowering::isAtomicAlignmentSupported(Align Alignment,
+ uint64_t SizeInBytes,
+ uint64_t ElementSizeInBytes,
+ unsigned AddrSpace) const {
+ // Each naturally aligned dword is separately atomic. Only global and flat
+ // opt in: the LDS 4 byte rule in allowsMisalignedMemoryAccessesImpl assumes
+ // lowering to ds_read2_b32, which never happens for an atomic.
+ if (ElementSizeInBytes < SizeInBytes && ElementSizeInBytes >= 4 &&
+ (AMDGPU::isExtendedGlobalAddrSpace(AddrSpace) ||
+ AddrSpace == AMDGPUAS::FLAT_ADDRESS))
----------------
arsenm wrote:
Flat can hit an LDS address, but I assume if the flat instruction guarantees atomicity it has to make sure it works for all addresses
https://github.com/llvm/llvm-project/pull/219906
More information about the llvm-commits
mailing list