[llvm] 6cc0cb6 - AMDGPU: Don't limit VGPR usage based on occupancy in dVGPR mode (#185981)
via llvm-commits
llvm-commits at lists.llvm.org
Mon Mar 16 06:39:32 PDT 2026
Author: Jannik Silvanus
Date: 2026-03-16T14:39:26+01:00
New Revision: 6cc0cb69e486c1b8a5156dd525ea4628c7952238
URL: https://github.com/llvm/llvm-project/commit/6cc0cb69e486c1b8a5156dd525ea4628c7952238
DIFF: https://github.com/llvm/llvm-project/commit/6cc0cb69e486c1b8a5156dd525ea4628c7952238.diff
LOG: AMDGPU: Don't limit VGPR usage based on occupancy in dVGPR mode (#185981)
The maximum VGPR usage of a shader is limited based on the target
occupancy,
ensuring that the targeted number of waves actually fit onto a CU/WGP.
However, in dynamic VGPR mode, we should not do that, because VGPRs are
allocated
dynamically at runtime, and there are no static constraints based on
occupancy.
Fix that in this patch.
Also fixup the getMinNumVGPRs helper to behave consistently by always
returning
zero in dVGPR mode.
This also fixes a problem where AMDGPUAsmPrinter bumps the VGPR usage to
at least
the result of getMinNumVGPRs, per my understanding in order to avoid an
occupancy
that is higher than the occupancy target. That was causing incorrect
(too high)
VGPR usages in dVGPR mode with medium-sized workgroups (say 768).
Added:
Modified:
llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
llvm/unittests/Target/AMDGPU/AMDGPUUnitTests.cpp
Removed:
################################################################################
diff --git a/llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp b/llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
index 9160a42b7b37c..dc56d746e1a8e 100644
--- a/llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
+++ b/llvm/lib/Target/AMDGPU/Utils/AMDGPUBaseInfo.cpp
@@ -1495,6 +1495,14 @@ unsigned getMinNumVGPRs(const MCSubtargetInfo *STI, unsigned WavesPerEU,
unsigned DynamicVGPRBlockSize) {
assert(WavesPerEU != 0);
+ // In dynamic VGPR mode, (static) occupancy does not depend on VGPR usage,
+ // so getMaxNumVGPRs does not depend on WavesPerEU, and thus we need to return
+ // zero because there is no nonzero VGPR usage N where going below N
+ // achieves higher (static) occupancy.
+ bool DynamicVGPREnabled = (DynamicVGPRBlockSize != 0);
+ if (DynamicVGPREnabled)
+ return 0;
+
unsigned MaxWavesPerEU = getMaxWavesPerEU(STI);
if (WavesPerEU >= MaxWavesPerEU)
return 0;
@@ -1522,9 +1530,13 @@ unsigned getMaxNumVGPRs(const MCSubtargetInfo *STI, unsigned WavesPerEU,
unsigned DynamicVGPRBlockSize) {
assert(WavesPerEU != 0);
+ // In dynamic VGPR mode, WavesPerEU does not imply a VGPR limit.
+ bool DynamicVGPREnabled = (DynamicVGPRBlockSize != 0);
unsigned MaxNumVGPRs =
- alignDown(getTotalNumVGPRs(STI) / WavesPerEU,
- getVGPRAllocGranule(STI, DynamicVGPRBlockSize));
+ DynamicVGPREnabled
+ ? getTotalNumVGPRs(STI)
+ : alignDown(getTotalNumVGPRs(STI) / WavesPerEU,
+ getVGPRAllocGranule(STI, DynamicVGPRBlockSize));
unsigned AddressableNumVGPRs =
getAddressableNumVGPRs(STI, DynamicVGPRBlockSize);
return std::min(MaxNumVGPRs, AddressableNumVGPRs);
diff --git a/llvm/unittests/Target/AMDGPU/AMDGPUUnitTests.cpp b/llvm/unittests/Target/AMDGPU/AMDGPUUnitTests.cpp
index 593991c71d706..81982d0217f71 100644
--- a/llvm/unittests/Target/AMDGPU/AMDGPUUnitTests.cpp
+++ b/llvm/unittests/Target/AMDGPU/AMDGPUUnitTests.cpp
@@ -168,6 +168,12 @@ static void testDynamicVGPRLimits(StringRef CPUName, StringRef FS,
<< CPUName << " dynamic VGPR block size " << DynamicVGPRBlockSize
<< ":\nOcc MinVGPR MaxVGPR\n"
<< Table.str() << '\n';
+ // In dVGPR mode, max VGPR limits do not depend on occupancy:
+ EXPECT_EQ(ST.getMaxNumVGPRs(1, DynamicVGPRBlockSize),
+ ST.getMaxNumVGPRs(ST.getMaxWavesPerEU(), DynamicVGPRBlockSize));
+ EXPECT_EQ(ST.getMinNumVGPRs(1, DynamicVGPRBlockSize), 0u);
+ EXPECT_EQ(ST.getMinNumVGPRs(ST.getMaxWavesPerEU(), DynamicVGPRBlockSize),
+ 0u);
};
testWithBlockSize(16);
More information about the llvm-commits
mailing list