[llvm] [AMDGPU] getMaxLocalMemSizeWithWaveCount: Align to LDS granularity (PR #219172)
via llvm-commits
llvm-commits at lists.llvm.org
Thu Aug 27 05:23:57 PDT 2026
llvmorg-github-actions[bot] wrote:
<!--LLVM PR SUMMARY COMMENT-->
@llvm/pr-subscribers-backend-amdgpu
Author: Frederik Harwath (frederik-h)
<details>
<summary>Changes</summary>
The AMDGPUSutarget class was recently extended by a LDSAllocationGranularity
field in PR #<!-- -->205637. This was added to make the occupancy calculation in
getOccupancyWithWorkgroupSize more precise by rounding the LDS
allocation size to the allocation granularity. The
getMaxLocalMemSizeWithWaveCount does an "inverse" computation,
computing available LDS from occupancy. This should also round to a
multiple of the LDS allocation granularity since otherwise the
available LDS may be overestimated. This leads to incorrect alloca
promotions in the AMDGPUPromoteAlloca pass.
Add rounding to getMaxLocalMemSizeWithWaveCount and add a test
demonstrating the impact on AMDGPUPromoteAlloca.
---
Full diff: https://github.com/llvm/llvm-project/pull/219172.diff
3 Files Affected:
- (modified) llvm/lib/Target/AMDGPU/AMDGPUSubtarget.cpp (+2-1)
- (modified) llvm/test/CodeGen/AMDGPU/large-work-group-promote-alloca.ll (+2-2)
- (added) llvm/test/CodeGen/AMDGPU/promote-alloca-lds-granularity-limit.ll (+37)
``````````diff
diff --git a/llvm/lib/Target/AMDGPU/AMDGPUSubtarget.cpp b/llvm/lib/Target/AMDGPU/AMDGPUSubtarget.cpp
index 87515d22ed422..eb90a658e06aa 100644
--- a/llvm/lib/Target/AMDGPU/AMDGPUSubtarget.cpp
+++ b/llvm/lib/Target/AMDGPU/AMDGPUSubtarget.cpp
@@ -46,7 +46,8 @@ AMDGPUSubtarget::getMaxLocalMemSizeWithWaveCount(unsigned NWaves,
const unsigned WorkGroupsPerCU =
std::max(1u, (NWaves * getEUsPerCU()) / WavesPerWorkgroup);
- return getLocalMemorySize() / WorkGroupsPerCU;
+ const unsigned Granularity = std::max(LDSAllocationGranularity, 1u);
+ return alignDown(getLocalMemorySize() / WorkGroupsPerCU, Granularity);
}
std::pair<unsigned, unsigned> AMDGPUSubtarget::getOccupancyWithWorkGroupSizes(
diff --git a/llvm/test/CodeGen/AMDGPU/large-work-group-promote-alloca.ll b/llvm/test/CodeGen/AMDGPU/large-work-group-promote-alloca.ll
index 32b356a514cbd..dc688121b4130 100644
--- a/llvm/test/CodeGen/AMDGPU/large-work-group-promote-alloca.ll
+++ b/llvm/test/CodeGen/AMDGPU/large-work-group-promote-alloca.ll
@@ -116,7 +116,7 @@ entry:
; SI-LABEL: @occupancy_6(
; CI-LABEL: @occupancy_6(
; SI: alloca
-; CI-NOT: alloca
+; CI: alloca [42 x i8]
define amdgpu_kernel void @occupancy_6(ptr addrspace(1) nocapture %out, ptr addrspace(1) nocapture %in) #5 {
entry:
%stack = alloca [42 x i8], align 4, addrspace(5)
@@ -216,7 +216,7 @@ entry:
; SI-LABEL: @occupancy_9(
; CI-LABEL: @occupancy_9(
; SI: alloca
-; CI-NOT: alloca
+; CI: alloca [28 x i8]
define amdgpu_kernel void @occupancy_9(ptr addrspace(1) nocapture %out, ptr addrspace(1) nocapture %in) #7 {
entry:
%stack = alloca [28 x i8], align 4, addrspace(5)
diff --git a/llvm/test/CodeGen/AMDGPU/promote-alloca-lds-granularity-limit.ll b/llvm/test/CodeGen/AMDGPU/promote-alloca-lds-granularity-limit.ll
new file mode 100644
index 0000000000000..98741f40e0366
--- /dev/null
+++ b/llvm/test/CodeGen/AMDGPU/promote-alloca-lds-granularity-limit.ll
@@ -0,0 +1,37 @@
+; RUN: opt -S -mtriple=amdgpu9.50-amd-amdhsa -passes=amdgpu-promote-alloca -disable-promote-alloca-to-vector < %s | FileCheck %s
+
+ at lds_12800 = internal addrspace(3) global [12800 x i8] poison, align 16
+attributes #0 = { "amdgpu-flat-work-group-size"="64,64" }
+
+; This is a regression test for a bug in getMaxLocalMemSizeWithWaveCount
+; which did not round to the LDS allocation block size, leading
+; AMDGPUPromoteAlloca pass to overestimate the available LDS.
+
+; We have LocalMemorySize = 163840, allowing for floor(163840/ 12800)
+; = 12 workgroups and occupancy 12 / 4 = 3 with the 12800 bytes of LDS
+; usage and one wave per workgroup.
+
+; Without aligning down to the LDS granularity of 1280 in
+; getMaxLocalMemSizeWithWaveCount, the limit in alloca promotion is
+; 163840 / 12 = 13653 bytes which led to the promotion of the alloca
+; in the function.
+
+; With rounding down, the promotion limit is 12800 bytes. The alloca
+; would add 64 * 10 bytes which exceeds the promotion limit.
+; This prevents the promotion of the alloca.
+
+; CHECK-LABEL: @test(
+; CHECK: %stack = alloca [10 x i8], align 1, addrspace(5)
+
+define amdgpu_kernel void @test(ptr addrspace(1) %out, i32 %idx) #0 {
+
+ %stack = alloca [10 x i8], align 1, addrspace(5)
+ %lds.ptr = getelementptr inbounds [12800 x i8], ptr addrspace(3) @lds_12800, i32 0, i32 0
+ store volatile i8 1, ptr addrspace(3) %lds.ptr, align 1
+
+ %arrayidx = getelementptr inbounds [10 x i8], ptr addrspace(5) %stack, i32 0, i32 %idx
+ store i8 7, ptr addrspace(5) %arrayidx, align 1
+ %load = load i8, ptr addrspace(5) %arrayidx, align 1
+ store i8 %load, ptr addrspace(1) %out, align 1
+ ret void
+}
\ No newline at end of file
``````````
</details>
https://github.com/llvm/llvm-project/pull/219172
More information about the llvm-commits
mailing list