[llvm-branch-commits] [llvm] [AMDGPU][SIMemoryLegalizer] Lower single-wave workgroup scope to wavefront in absence of LDSDMA (PR #224065)
Sameer Sahasrabuddhe via llvm-branch-commits
llvm-branch-commits at lists.llvm.org
Tue Sep 22 01:32:21 PDT 2026
================
@@ -7550,6 +7553,21 @@ treated as non-atomic.
A memory synchronization scope wider than work-group is not meaningful for the
group (LDS) address space and is treated as work-group.
+When a work-group's maximum flat work-group size does not exceed the wavefront
+size, the work-group fits within a single wavefront. So long as no asynchronous
+operations occur, the LLVM ``workgroup`` synchronization scope is equivalent to
+its ``wavefront`` scope.
+
+If the compiler can determine these conditions (e.g., through the function attributes
+``amdgpu-flat-work-group-size`` and ``amdgpu-no-async``), the AMDGPU backend
+optimizes ``workgroup`` scope operations by lowering them to
+``wavefront``-scoped machine instructions.
+
+This optimization applies to atomic ``load``, ``store``, ``atomicrmw``, and
----------------
ssahasra wrote:
Are there operations that take a scope argument, but are not subject to this optimization?
https://github.com/llvm/llvm-project/pull/224065
More information about the llvm-branch-commits
mailing list