[llvm-branch-commits] [llvm] [AMDGPU][SIMemoryLegalizer] Lower single-wave workgroup scope to wavefront in absence of LDSDMA (PR #224065)

Sameer Sahasrabuddhe via llvm-branch-commits llvm-branch-commits at lists.llvm.org
Tue Sep 22 01:32:21 PDT 2026


================
@@ -7550,6 +7553,21 @@ treated as non-atomic.
 A memory synchronization scope wider than work-group is not meaningful for the
 group (LDS) address space and is treated as work-group.
 
+When a work-group's maximum flat work-group size does not exceed the wavefront
+size, the work-group fits within a single wavefront. So long as no asynchronous
+operations occur, the LLVM ``workgroup`` synchronization scope is equivalent to
+its ``wavefront`` scope.
+
+If the compiler can determine these conditions (e.g., through the function attributes
+``amdgpu-flat-work-group-size`` and ``amdgpu-no-async``), the AMDGPU backend
+optimizes ``workgroup`` scope operations by lowering them to
+``wavefront``-scoped machine instructions.
+
+This optimization applies to atomic ``load``, ``store``, ``atomicrmw``, and
----------------
ssahasra wrote:

Are there operations that take a scope argument, but are not subject to this optimization?

https://github.com/llvm/llvm-project/pull/224065


More information about the llvm-branch-commits mailing list